The Pilot Already Worked. Nothing Was Linked to It.
- Northbridge Components processes roughly 1,400 supplier invoices a month. About 38% fail the ERP's three-way match and drop into a manual exception queue.
- Each exception takes 4.2 days to clear — of which 2.8 days is a buyer not replying to an email, not anyone doing work.
- Northbridge's IT function had already run an agent pilot on invoice extraction eight months earlier. It read PDFs accurately, everyone agreed it was impressive, and it was still running in a sandbox with nothing attached to it.
- Meridian Capital Partners flagged the queue during a quarterly value-creation review — not because the pilot had failed, but because nobody could say what it had changed.
That is the ordinary shape of the problem. The capability was never in doubt. What was missing was the link: from a working agent to a business number someone signs for, with permissions, escalation, and an audit trail attached to it.
What "success" had to mean here
The KPI Contract locked one distinction before Day 0: a working agent isn't the deliverable. A cycle-time number that holds after we leave is.
- Floor: 1.0-day average cycle time on exception invoices — a credible win, with 60% of manual handoffs removed.
- Stretch: 0.5 days, with exception volume itself falling and a complete audit trail on 100% of runs.
- Never contracted: agent count, model choice, or "percentage of invoices touched by AI."
Everything below was built against that line.
Three Candidates. One Bounded Enough to Own.
Phase 0 ran as an Executive Lab in its agentic-workflow variant — unmetered, off the clock, with Northbridge's CFO, controller, and AP lead in the room for a single day. Three candidate workflows surfaced. Each was scored on GUSTO-Winnability, and the two that lost were documented as rejected, with the reason, rather than quietly dropped.
| Candidate workflow | Bounded? | Binary outcome? | Auditable? | Verdict |
|---|---|---|---|---|
| Invoice exception handling | Yes — one queue, one trigger | Yes — matched or routed | Yes — every run leaves evidence | Selected |
| Vendor onboarding | No — spans three functions | Partly | Yes | Rejected |
| Board-pack drafting | Yes | No — quality is a judgement | No defensible trail | Rejected |
Board-pack drafting failed on a different test. An agent can draft a board pack; nobody can say afterwards whether the draft was right, and there is no artifact an auditor could inspect. A workflow with no binary outcome and no defensible trail is a productivity tool, not a Governed Workflow Unit — a useful thing to buy, and the wrong thing to run a 90-day engagement against.
One Number, Signed Before the Clock Started
Phase 0 is the only unmetered phase, and it ends at a signature. Four metric families were locked — the AgentThrust library, inherited from Chain Reactor with the mechanism untouched — each with a baseline, a floor, and a stretch. Northbridge's controller was named as the single owner of the number.
| Metric family | Baseline (Day 0) | Floor | Stretch | Measured from |
|---|---|---|---|---|
| Cycle-time reduction | 4.2 days | 1.0 day | 0.5 days | Per-step cycle time on the Workflow Map |
| Manual handoffs eliminated | 4 per exception | 60% | 75% | Lane crossings involving a human |
| Exception volume | 38% of invoices | 15% | 10% | Cases routed to the escalation path |
| Audit completeness | Partial — free-text ERP notes | 100% | 100% | Runs with a complete, retrievable trail |
The anticipated Risk Tier was set provisionally at Tier 2 — Augmented, on the expectation that a human would still sign anything above a monetary threshold. That expectation was confirmed in Phase 1, once the map existed to confirm it against.
The Spreadsheet Nobody Had Mentioned
Day 4 of Phase 1, four hours on a shared canvas, with the AP lead and one Northbridge engineer in the room. The rule is that you map the current state exactly as it runs today — exceptions and workarounds included — before designing anything.
The documented process had four steps. The real one had seven, and the third of them was a shared spreadsheet that lived outside the ERP, maintained by one clerk, which no process document mentioned and which the earlier pilot had never seen.
| Field | What the session captured |
|---|---|
| 1 · Trigger | Event-driven: a supplier invoice arrives in the shared AP mailbox with a PO that does not reconcile. |
| 2 · Lanes | Three: AP Clerk (Human), ERP & Intake (System), Matching Agent (Agent). The Agent lane was empty in the current-state map — the pilot ran beside the process, not inside it. |
| 3 · Steps | Seven, not four. Three of them workarounds, including the spreadsheet. |
| 4 · Gateways | One deterministic branch today (matched / not matched). In the governed design it becomes an Agentic Gateway at a $5,000 threshold, collaboration mode leader-driven. |
| 5 · Exceptions | Low confidence or missing data routes to the AP Clerk lane regardless of invoice value. This field became the escalation path verbatim. |
| 6 · Data Flows | Extracted lines and PO lines, both internal-confidential, never leaving the ERP boundary. One workaround was emailing invoice PDFs outside the ERP — flagged here, closed in the design. |
| 7 · Cycle Time & Volume | 4.2 days end to end, ≈1,400 invoices / month, 38% exception rate. This is the KPI Contract baseline — captured here and nowhere else. |
| 8 · Systems Touched | ERP (three-way match, PO table, payment run), invoice-intake mailbox, and the spreadsheet. This list is what the stack decision waits on. |
Governance Was Read Off the Map, Not Designed
No new oversight scale was invented. Each Agent Task was classified into one of CATO's three existing enterprise Risk Tiers, based on what the step touches — and the workflow as a whole is governed at its highest step.
The permissions specification came straight out of fields 6 and 8: read the intake mailbox and the ERP PO table; write only to the match queue. Nothing else. The escalation path came straight out of field 5, unaltered. Neither was designed in a separate workshop, because neither needed to be.
Nothing Was Purchased Until the Map Was Finished
Two orchestration platforms had already been demoed to Northbridge before Creativa arrived, and one had a quote attached. Both were set aside until Systems Touched was complete — because choosing the stack before the map is finished is a named anti-pattern, not a head start.
| Decision layer | Options considered | Chosen | Because |
|---|---|---|---|
| Orchestration | Enterprise agent platform (licensed) · cross-runtime framework · lighter build on frontier-model APIs | Lighter build on frontier-model APIs | Northbridge has one engineer who can maintain this. Tool sophistication is gated by maturity, not by what is newest. |
| Interoperability | Cross-system agent communication standards now, or single-agent now and add later | Single-agent now | One agent, one workflow. Multi-agent protocols were deferred to the second wave, where they were actually needed. |
| Build / license / SaaS | Decided per Agent Task, not per engagement | Build on the existing ERP API | The ERP was already licensed and already in Systems Touched. No new platform was purchased. |
| Cost guardrails | — | $1,900 / month ceiling, alert at 70% | Agentic workloads have a materially less predictable cost profile than SaaS seats. Every recommendation ships with a ceiling. |
| Observability | New tooling vs. what the client already runs | Existing log stack, extended | New tooling is recommended only where a genuine gap exists against the Risk Tier requirement. There wasn't one. |
Phase 1 closed on Day 20 with every step on the map carrying a named owner. Two lanes had come back from the first session marked "TBD"; the session did not close until both were named. A placeholder in a lane is a red flag, not something to fill in later.
Real Invoices, Real Permissions, From Day 21
No sandbox. The Matching Agent went live on Day 21 against every mismatched-PO invoice arriving that week — not a cherry-picked subset, and not a shadow mode running alongside the humans. A Governed Workflow Unit that has only ever seen synthetic cases has not been tested.
The pod followed Chain Reactor's Misfit Team pattern — a small cross-functional group with explicit air cover to bypass legacy sign-off during the sprint — plus the one AgentThrust addition: a named technical owner accountable for the orchestration stack and the audit logging, present from Phase 1 onward rather than handed the work at go-live.
Exception Volume Was Reported as Loudly as Cycle Time
Four metric families, every week, without softening. The dashboard was built as Northbridge's — it had to keep working after Day 90 — and the metric that got the most attention was deliberately not the flattering one.
Exception volume, week by week
Share of invoices routed to the AP Clerk lane. The contracted floor was 15%; the stretch was 10%.
Contracted floor — 15%, cleared in week 4
The 3% Audit Gap Became an Atomic Unit
Whatever telemetry exposes becomes an Atomic Unit: the smallest verifiable change, one named owner, binary done or not done. Chain Reactor's microloop mechanics, unchanged — and the audit gap was the one that mattered most, because 97% audit completeness fails a 100% contract exactly as badly as 40% would.
| Atomic Unit | Owner | Due | Status |
|---|---|---|---|
| AU-1 · Tolerance rule for freight-line mismatches | Technical owner | Day 34 | Done |
| AU-2 · Vendor-alias table seeded from the week-2 exception cluster | AP lead | Day 41 | Done |
| AU-3 · Confidence score written into the audit record | Technical owner | Day 48 | Done |
| AU-4 · Approving identity (agent or human) captured per run | Technical owner | Day 55 | Done |
| AU-5 · Retire the shared matching spreadsheet | AP lead | Day 62 | Deferred to Phase 3 |
AU-3 and AU-4 together closed the audit gap: the 3% of runs missing a complete trail were all cases where a human had intervened mid-flow, and the log recorded the outcome without recording who had produced it. By Day 58, audit completeness read 100%.
Reinforcement Came Off on Day 71
The threshold is crossed when reverting to the manual process would cost more friction than sustaining the agent-owned one. Four conditions, all of which must hold. Creativa stopped resolving escalations on Day 71 and watched.
No-Return Threshold confirmed: Day 84. From that date the agent-owned workflow is the system of record, with no parallel manual path running beside it.
Scored Against the Contract Signed on Day 0
No renegotiation, no adjusted baseline, no metric quietly redefined between Day 0 and Day 90. Three families cleared stretch; one cleared floor but not stretch, and is reported as such.
Cycle time landed at 0.7 days — comfortably past the 1.0-day floor, short of the 0.5-day stretch. The residual is structural rather than fixable by the agent: the ~9% of invoices that still need a buyer's answer are gated on a human replying to an email, and no amount of agent capability shortens that. It is reported as a floor-cleared result, not rounded toward a stretch it did not reach.
The full traceable chain: the contract signed before Day 1, against what the telemetry confirmed on Day 90.
The Second Workflow Had No Human on the Happy Path
Meridian approved a second AgentThrust cycle at the Day-90 review, this time at group treasury rather than inside a portfolio company. The workflow is a different shape entirely — and it is the case that tests whether the Workflow Map is a real primitive or just a diagram for simple flows.
- The trigger is a schedule, not an event: 05:30 local every business day, plus a second run at 13:00 on month-end days. A trigger can carry more than one fixed time.
- Four agents, not one: a Cash Orchestrator dispatching three workers that run in parallel — bank balances, closing FX rates, hedge exposure.
- Two collaboration modes at once: the three workers are in role cooperation (they split the work, none waits on the others); the orchestrator is leader-driven (it alone decides).
- Zero human steps when it runs clean — which is exactly why it is Tier 1 — Critical, not Tier 2.
Three things about the second cycle were different, and all three came out of the map rather than out of a new methodology:
| What changed | Why the map forced it |
|---|---|
| The fan-in became a mapped step | Three agents returning independently can disagree, and one can return nothing. The step that reconciles them needs an owner and an audit record like any other — and the map has to say whether a partial return blocks the run. |
| The Trigger field grew a deadline | An event-triggered workflow fails loudly because somebody is waiting. A scheduled one fails silently. The map carries the times and the rule: not complete by 06:15, the Treasurer hears about it. |
| Audit became per-agent attributable | One audit line for the run stops being enough once four agents contributed to it. Each agent's inputs, outputs, confidence, and timing are logged separately, so a wrong number traces to the agent that produced it — non-negotiable at Tier 1. |
The Link Was the Deliverable
Northbridge already had a working agent on Day 0. What it did not have was anything connecting that agent to a number, an owner, a permission boundary, or an audit trail. Ninety days later the same underlying capability is doing recognisably the same work — and it is now the system of record for a queue that four people used to share, with a controller's name against the number.
The bottleneck was never the model. The pilot had solved extraction eight months earlier. What took 90 days was the permissions, the escalation path, the audit trail, and one named owner — none of which a better model would have supplied.
Mapping the mess is what made it survive. The undocumented spreadsheet and the buyer-reply wait were both invisible in the official process. Had the map skipped them, the Governed Workflow Unit would have broken on its first real exception instead of clearing its floor on Day 58.
The rejected candidates protected the cycle. Vendor onboarding had the loudest internal demand and no single owner; board-pack drafting had no binary outcome. Documenting why each lost is what kept the 90 days pointed at a workflow that could actually be owned.
Missing the stretch is not a failed engagement — it's an honest one. Cycle time stopped at 0.7 days because roughly 9% of invoices wait on a human answering an email. That residual is structural, it is named in the verdict, and it is the starting point for the next cycle rather than something rounded away in this one.
One 90-day cycle, one named workflow, one number locked before the clock started. The methodology does not promise that every workflow clears its stretch target, and it does not let a pilot run indefinitely while everyone agrees it is promising. It promises a link — from a capable agent to a business outcome with an owner, a boundary, and an audit trail behind it — and a verdict at Day 90 either way.
Book a session