All case studies
Multi-Agent AICredit Risk

OpenCAM — Autonomous Maker-Checker Framework

Engineered a dual-agent LLM pipeline with deterministic policy gating, targeting a cut in CAM drafting time from a business day to 15–30 minutes.

The takeaway: A real correctness bug: a debt-free company's DSCR silently computed as 0 instead of undefined, caught before it could reach a credit committee. Concrete proof the governance model works in practice.

View Source60+ PRs merged36+ issues closed

Role

Product Manager — strategy, prompt architecture, policy rules; directed Claude Code for all implementation

Timeline

Sep 2026 · ongoing

Stack

Python · Claude Code · Anthropic API · Deterministic policy engine

Status

Early-stage · validated with one analyst to date

Problem

What was broken

I watched my wife, a corporate credit analyst, prepare Credit Assessment Memorandums every day against hard credit-committee deadlines — manual financial spreading, then drafting a narrative under time pressure with real reputational risk if a number was wrong. She's running several deals at once, and borrower financials arrive as PDFs or spreadsheets from relationship directors, often before the accounts are even filed with Companies House.

Early attempts at using an LLM to draft the memo directly were faster but untrustworthy: grounding was only ever a "cite your source" prompt instruction, not an enforced rule, so an unsupported claim could reach committee undetected — and a codebase audit later caught a debt-free company's DSCR being silently computed as 0 instead of undefined, which would have wrongly flagged a healthy borrower as a covenant breach.

  • One rigid spreading schema couldn't handle real deal variety — a borrower embedding Depreciation inside Cost of Goods Sold, a real house convention, broke the single schema the MVP shipped with.
  • The pipeline assumed every deal wanted the full automation stack — there was no way to ask for just the qualitative research while spreading a non-standard deal by hand.

Scope

What I scoped into MVP v1, and what I explicitly didn't

Must have

MVP — shipped in PRs #1–20

  • CAM template system
  • Spreading engine with formula validation
  • .docx export
  • The Maker-Checker governance loop itself
  • Deterministic policy engine layered on LLM narrative
  • Critical correctness/security fixes before wider use

Should have

Fast-follow — shipped in PR #24

  • Wiring policy checks into the primary slash-command interface, not just the headless script

Could have

Post-MVP depth — ✓/✗ shows what's actually shipped since

  • Forward-year projections & stress testing
  • Conditions Subsequent tracking, Net Debt/EBITDA & FCF ratios
  • Source-citation hyperlinking
  • AML/sanctions/PEP screening & ESG scoring

Won't have

Explicitly deferred

  • Multi-currency/FX support — today's desk is GBP-only
  • Full covenant step-down/cure-period modeling — scoped down to just Conditions Subsequent tracking

Product

What it does

A

Maker-Checker Governance Loop

Two independent agents — an Underwriter drafts, a Risk Reviewer audits — with no shared reasoning context, and the power to only downgrade a verdict, never upgrade one. Maker and Checker can run on different underlying models, so drafting and audit don't share the same blind spots.

B

Deterministic Policy & Compliance Engine

Every covenant is evaluated PASS / FAIL / UNRESOLVABLE against a ratio computed from raw financials — never silently defaulted. Every reported figure is checked against ground-truth financials within a 0.5% tolerance before the Checker even sees it.

C

Institutional Credit Policy Referencing

A one-time /calibrate-policy command derives a fork-local credit policy from an institution's own policy documents. The Underwriter treats it as advisory drafting guidance; the Risk Reviewer treats it as a mandatory, independently-verified audit obligation — any violation is REJECTED-worthy regardless of what the Underwriter declared, the same asymmetric-weight pattern used everywhere else in the governance model.

D

Financial Spreading & Auditable Excel Export

Every ratio (TNW, EBITDA, DSCR, Gross Leverage, Net Debt/EBITDA, FCF Conversion %) is computed straight from the same raw line items shown in the workbook, with formulas generated from a label-based row layout so a reorder can't silently break a reference. An analyst can also supply figures already spread against their own institution's template instead: the CAM then carries an explicit caveat disclosing the spreading wasn't independently recomputed, a real reduction in audit guarantee.

E

Confidentiality-by-Design

Every artifact derived from a user's real business — calibration samples, templates, deal state, output — writes only to git-ignored paths, enforced as a build rule, not audited in after the fact.

F

Standalone Research Workflow (/research)

A separate command for an analyst who just needs the qualitative picture — company and sector research, Go/No-Go screening — without running the full CAM pipeline. It writes the same state.json keys the full pipeline would, so a deal that later needs a full CAM can continue straight in without redoing anything, and exports its own standalone Research Brief through a script kept deliberately separate from the full CAM exporter.

Workflow

How the pipeline actually connects

Two zones. One-time setup, not run per deal: /calibrate (writing style + CAM template, per deal type) and /calibrate-policy (institution credit policy, org-wide) - both just pre-existing config, not wired into any specific deal. Per-deal execution: every deal starts at /triage or /research. The /triage lane (blue, solid - committed once chosen) runs straight through /spread, /commercial, /collateral, /project, /assemble. The /research lane (teal) is solid only for what's guaranteed - /research itself and its standalone Research Brief export; everything past that (rejoining at /spread, skipping /commercial, continuing to /collateral, /project, /assemble) is dashed because an analyst may never come back to finish a full CAM. /spread also exports a Spreading Template, and /assemble exports the Full CAM (.docx + .xlsx).ONE-TIME SETUP — NOT RUN PER DEALPER-DEAL EXECUTION/calibrate/calibrate-policyresearch lane —bypasses /commercial/triage/research/spread/commercial(/triage path only)/collateral/project/assembleWResearchBriefXSpreadingTemplateWXFull CAM.docx + .xlsx/triage lane/research lanecommittedoptional, may never run

Two entry points, one shared spine. Every deal starts at /triage for a full CAM or /research for a standalone qualitative brief. Both write the same state, so a deal can continue into the full pipeline later without redoing work. A deal continuing from /research rejoins directly at /spread and skips /commercial, since /research already produced that output itself.

/calibrate and /calibrate-policy (above the divider) are one-time, org-level setup: /calibrate derives house writing style and a CAM template; /calibrate-policy derives an institution's own credit policy. Every deal afterward just reads whichever is already in place.

Strategy

Validation & strategic context

This framework is early-stage, built and validated with one target user — my wife, a corporate credit analyst. I'd rather show the reasoning openly than present these as more settled than they are.

Built to augment, not replace

My wife never had access to a platform like nCino or Moody's CreditLens at her workplace, so this isn't a competitive swap-out — it fills a gap that simply existed for her. It's also deliberately scoped to speed up a first draft with an inspectable audit trail, not replace her judgment the way larger, better-resourced platforms are attempting.

Unit economics

Slash commands shell out to the same local policy-check modules as the headless script — no separate metered API bill per step beyond the analyst's existing Claude Code access, against a target of multiple hours of analyst time per deal if the 15–30 minute goal holds at scale.

Enterprise adoption

Because it runs locally and writes nothing but git-ignored files, an analyst or a bank could use it without sending a single confidential figure to a third-party service, and without IT needing to approve a new vendor integration into core banking systems.

Process

How I worked it

1

Started from watching a real workflow break down

Modeled the primary persona directly on my wife's own day as a credit analyst — the actual friction was manual spreading and narrative drafting under deadline pressure, not a problem I picked because it sounded interesting.

2

Shipped the Maker-Checker loop in one day

Scoped the MVP tightly — 20 merged PRs in a single day — around one correct, end-to-end loop before going deep on any single feature: template system, spreading engine, .docx export, and the governance loop itself.

3

Ran a dedicated gap-analysis audit, sequenced by risk

Once the MVP worked, I audited the live codebase for 19 further issues and merged fixes in risk-weighted severity order — correctness and security bugs first, cosmetic issues after, feature work last.

4

Reprioritized around real feedback from a live deal

Two items came directly from my wife hitting the pipeline's limits on a real deal — a rigid spreading schema and an all-or-nothing automation model — and I pulled both to the top of the queue and shipped them the same week.

Decisions

Calls I made, and why

No shared context between Maker and Checker

A single agent auditing its own draft agrees with itself. The Underwriter and Risk Reviewer share no reasoning context, and the Reviewer can only downgrade a verdict — its value comes from auditing cold.

The model narrates; code computes

Every covenant is evaluated PASS / FAIL / UNRESOLVABLE against a ratio computed deterministically, never silently defaulted. That's what fixes the debt-free-DSCR bug for good, plus a second, near-identical bug where Provisions and Other Long-Term Liabilities silently never reached total_liabilities. The model can narrate a number, but it can never produce one.

Shipped the bounded piece, flagged the rest

One issue bundled four separable asks together. Rather than push all four through unreviewed, I shipped just the well-bounded piece and explicitly disclosed the other three as deferred, not quietly dropped.

Reopened an issue I'd already closed

I'd marked Issue #31 resolved once Maker/Checker could run on different models in code. A later review caught that the setting was never actually switched on in production — the audit wasn't independent. I reopened it and only re-closed once I'd verified it working end to end.

Fixed what tests couldn't catch, same day

My first live production run leaked markdown formatting artifacts and raw internal JSON into the exported Word document. I fixed both within the day (PR #40) and codified durable prompt guidance alongside the fix, so the same class of gap couldn't recur.

Severity order beats arrival order

Working the 19-issue gap-analysis backlog, I shipped the highest-criticality correctness and security fixes first (PR #41: date-resume logic, a race condition, a hardcoded model), then lower-severity display bugs (PR #42), then feature work — any fix that could change a credit decision landed before anything cosmetic did.

Caught Claude Code skipping my own instruction, mid-deal

Running a real deal, Claude Code declared 20+ source citations across triage and commercial research but never once called the script that actually saves the underlying material, despite the instruction being right there in the command files I'd written myself. I caught it by asking directly where the material was; the sources folder didn't exist until I had it backfilled by hand. Fixed by code-enforcing that a declared citation has something saved behind it, so the same gap can't happen again unnoticed, this time in my own oversight of the AI doing the drafting.

A "lightweight" path had zero independent audit

/research, the standalone qualitative-brief command, never ran the Risk Reviewer: a Go/No-Go legal screen carried real decision weight with no independent check. Giving it a Checker pass surfaced a second, sharper bug: the existing compliance checker assumes a full CAM's shape, so pointed at a research-only brief it would have silently returned "compliant: true" with zero reasons, regardless of whether the brief was actually sound. Nothing matched what the checks look for. Caught by a manual smoke test before the fix shipped.

See how this compares to the other project's delivery model

Outcomes

What changed

15–30 min

time-to-first-draft target, from ~1 business day

439+

passing tests (up from 188 at MVP)

0

financial figures the model is allowed to compute itself

  • Trust turned out to be a product feature: the debt-free-DSCR catch is concrete evidence that code-level verification, not model judgment, is what makes the loop actually trustworthy.
  • Real usage surfaced the two highest-value roadmap items faster than the original audit backlog did — ship it, use it, let real friction drive the backlog.

Next horizon

What I'm scoping for round 2 ("Hardening Cycle 2")

MVP v1 shipped and is already in real use. This next MoSCoW pass follows it, in the same priority order as round 1: fix foundational gaps before adding polish.

Must have

Foundational correctness/architecture gaps

  • Fix the dangling Parent/UBO Guideline cross-reference — the template already points to guidance that was never written (#96)
  • Wire the primary interface through to the tested financial-formula code, not around it (#98)
  • Extend the ground-truth figures schema to cover metrics the framework already requires citing — Working Capital Days, collateral exposure (#99)

Should have

Fast-follow depth

  • Systematic Parent/UBO research guidance — full ownership chain, ownership percentages, recent ownership changes, and the materiality judgment call for when it warrants a full Ultimate Parent section (#96)

Could have

Post-round-2 depth, if capacity allows

  • HoldCo/OpCo group/subsidiary financial consolidation (#48)
  • Render Group/Parent/UBO structure as a tree diagram instead of prose (#113)
  • Charts/graphs in CAMs — sector trends, SWOT, positioning, stock price (#114)

Won't have

Deferred for this round

  • Full covenant step-down/cure-period modeling (#33) — same bundled-scope call as MVP v1; may be worth revisiting if I ever open this up for wider adoption beyond my own use
  • FX/multi-currency support (#49) — today's desk is still GBP-only; the kind of gap that would gate wider adoption if this were ever forked for a multi-currency desk
  • AML/sanctions/PEP screening & ESG scoring (#35) — my wife's desk has a separately-owned AML team whose system already supplies this as an input; building it into OpenCAM would duplicate, not fill, a gap

MoSCoW above is how I scoped round 2 itself; this board is how I sequence what's inside it — live from the repo's own issue labels, so it never goes stale the way a hand-maintained backlog table would. Click a point to open the real issue.

Supporting visuals

Inside the build

Merged pull requests on GitHub, from earlier in the build — every change reviewed before merge, none pushed direct to main. The live count above has grown since this snapshot; the discipline it shows hasn't.

The original 19-issue gap-analysis audit, tracked and closed on GitHub as real issues, not a private todo list — the first of what's now a recurring practice.

A snapshot of feature requests tracked under GitHub's enhancement label, from earlier in the build — the spreading-schema and automation-model fixes that came from real usage, not a backlog guess. Some have since been reclassified as tech debt as the labeling taxonomy matured; the live backlog board above reflects the current split.

The financial spreading output — every ratio traceable to a raw line item, DSCR shown per year rather than as a single number. Synthetic demo data, generated for this writeup.

The CAM's Facility Conditions section, with a grounding note disclosing exactly which inputs came from the relationship team, unverified against a live feed. Synthetic demo data.

The Collateral section's LGD/coverage workout — average vs. max LGD, cover %, all computed from the same disclosed inputs. Synthetic demo data.