Book a session
Sample Deliverable — AgentThrust™ Applied

4.2 Days to 0.7: AgentThrust, Applied

One portfolio company. One vendor-invoice exception queue that four people had learned to live with. One 90-day cycle to hand that queue to an AI agent as its governed, audited owner — with a number locked before Day 1 and a verdict that had to be reported either way. Here's the full AgentThrust method — Diagnostic, Focus, Execution, Verdict — applied step by step, exactly as a real engagement runs, including the second workflow that came next and the one that was deliberately turned down.

Sponsor: Meridian Capital Partners (PE) Workflow: Vendor-invoice exception handling Volume: ≈1,400 invoices / month Cycle length: 90 days + unmetered Phase 0 Status: Fictitious · Built From Typical Patterns · Anonymized
Built directly from the live Creativa framework: AgentThrust™ →
This is a fictitious case. There is no real client and no real engagement behind this document. It was constructed from the accounts-payable dynamics, exception rates, and audit requirements that recur across mid-market finance functions, in order to demonstrate the methodology honestly on a scenario realistic enough to be useful. Every name is a fictional stand-in. Meridian Capital Partners is Creativa's standing private-equity composite — the sponsor commissioning the work. Northbridge Components is one of its portfolio companies, a mid-market industrial distributor, and the operating client. No figure anywhere in this document is, or is derived from, a real company's disclosed financials.
The Brief

The Pilot Already Worked. Nothing Was Linked to It.

  • Northbridge Components processes roughly 1,400 supplier invoices a month. About 38% fail the ERP's three-way match and drop into a manual exception queue.
  • Each exception takes 4.2 days to clear — of which 2.8 days is a buyer not replying to an email, not anyone doing work.
  • Northbridge's IT function had already run an agent pilot on invoice extraction eight months earlier. It read PDFs accurately, everyone agreed it was impressive, and it was still running in a sandbox with nothing attached to it.
  • Meridian Capital Partners flagged the queue during a quarterly value-creation review — not because the pilot had failed, but because nobody could say what it had changed.

That is the ordinary shape of the problem. The capability was never in doubt. What was missing was the link: from a working agent to a business number someone signs for, with permissions, escalation, and an audit trail attached to it.

Why a pilot doesn't become a workflow on its own A pilot proves an agent can do something. It says nothing about who owns the outcome, what the agent is allowed to touch, where a low-confidence case goes, or whether any of it would survive an audit. AgentThrust treats "the agent works" as the starting condition, not the deliverable.

What "success" had to mean here

The KPI Contract locked one distinction before Day 0: a working agent isn't the deliverable. A cycle-time number that holds after we leave is.

  • Floor: 1.0-day average cycle time on exception invoices — a credible win, with 60% of manual handoffs removed.
  • Stretch: 0.5 days, with exception volume itself falling and a complete audit trail on 100% of runs.
  • Never contracted: agent count, model choice, or "percentage of invoices touched by AI."

Everything below was built against that line.

1Diagnostic · The Candidate Shortlist

Three Candidates. One Bounded Enough to Own.

Phase 0 ran as an Executive Lab in its agentic-workflow variant — unmetered, off the clock, with Northbridge's CFO, controller, and AP lead in the room for a single day. Three candidate workflows surfaced. Each was scored on GUSTO-Winnability, and the two that lost were documented as rejected, with the reason, rather than quietly dropped.

Candidate workflowBounded?Binary outcome?Auditable?Verdict
Invoice exception handling Yes — one queue, one trigger Yes — matched or routed Yes — every run leaves evidence Selected
Vendor onboarding No — spans three functions Partly Yes Rejected
Board-pack drafting Yes No — quality is a judgement No defensible trail Rejected
Design principle Vendor onboarding was the workflow the room wanted, because it is the one that annoys everyone. It lost on a single question: who owns the outcome? Three functions each owned a piece and none owned the result — so no Governed Workflow Unit could be written for it. Enthusiasm is not a score.

Board-pack drafting failed on a different test. An agent can draft a board pack; nobody can say afterwards whether the draft was right, and there is no artifact an auditor could inspect. A workflow with no binary outcome and no defensible trail is a productivity tool, not a Governed Workflow Unit — a useful thing to buy, and the wrong thing to run a 90-day engagement against.

2Diagnostic · The KPI Contract

One Number, Signed Before the Clock Started

Phase 0 is the only unmetered phase, and it ends at a signature. Four metric families were locked — the AgentThrust library, inherited from Chain Reactor with the mechanism untouched — each with a baseline, a floor, and a stretch. Northbridge's controller was named as the single owner of the number.

Metric familyBaseline (Day 0)FloorStretchMeasured from
Cycle-time reduction4.2 days1.0 day0.5 daysPer-step cycle time on the Workflow Map
Manual handoffs eliminated4 per exception60%75%Lane crossings involving a human
Exception volume38% of invoices15%10%Cases routed to the escalation path
Audit completenessPartial — free-text ERP notes100%100%Runs with a complete, retrievable trail
The gate that nearly stopped the engagement Northbridge's CIO asked to leave the cycle-time floor open for the first month, "until we see how the agent behaves." The answer was no, and the engagement paused for nine days while the number was settled. A contract that can be adjusted mid-flight to accommodate execution friction is not a contract — it is a status report with a target attached.

The anticipated Risk Tier was set provisionally at Tier 2 — Augmented, on the expectation that a human would still sign anything above a monetary threshold. That expectation was confirmed in Phase 1, once the map existed to confirm it against.

3Focus · The Workflow Map

The Spreadsheet Nobody Had Mentioned

Day 4 of Phase 1, four hours on a shared canvas, with the AP lead and one Northbridge engineer in the room. The rule is that you map the current state exactly as it runs today — exceptions and workarounds included — before designing anything.

The documented process had four steps. The real one had seven, and the third of them was a shared spreadsheet that lived outside the ERP, maintained by one clerk, which no process document mentioned and which the earlier pilot had never seen.

FieldWhat the session captured
1 · TriggerEvent-driven: a supplier invoice arrives in the shared AP mailbox with a PO that does not reconcile.
2 · LanesThree: AP Clerk (Human), ERP & Intake (System), Matching Agent (Agent). The Agent lane was empty in the current-state map — the pilot ran beside the process, not inside it.
3 · StepsSeven, not four. Three of them workarounds, including the spreadsheet.
4 · GatewaysOne deterministic branch today (matched / not matched). In the governed design it becomes an Agentic Gateway at a $5,000 threshold, collaboration mode leader-driven.
5 · ExceptionsLow confidence or missing data routes to the AP Clerk lane regardless of invoice value. This field became the escalation path verbatim.
6 · Data FlowsExtracted lines and PO lines, both internal-confidential, never leaving the ERP boundary. One workaround was emailing invoice PDFs outside the ERP — flagged here, closed in the design.
7 · Cycle Time & Volume4.2 days end to end, ≈1,400 invoices / month, 38% exception rate. This is the KPI Contract baseline — captured here and nowhere else.
8 · Systems TouchedERP (three-way match, PO table, payment run), invoice-intake mailbox, and the spreadsheet. This list is what the stack decision waits on.
Why the messy version is the load-bearing one Had the map been drawn from the documented four-step process, the Governed Workflow Unit would have shipped without a path for the ~11% of invoices whose resolution depends on a buyer's reply. It would have looked excellent for three days and broken on its first real exception.
The full three-lane map, clickable step by step: The Workflow Map →
4Focus · Risk Tier & Governance

Governance Was Read Off the Map, Not Designed

No new oversight scale was invented. Each Agent Task was classified into one of CATO's three existing enterprise Risk Tiers, based on what the step touches — and the workflow as a whole is governed at its highest step.

Tier 1 — Critical Autonomous customer-facing or financial decision-making. Board-level oversight, real-time audit logging. Not reached here — a human still signs above the threshold, which is precisely what kept this workflow out of Tier 1.
Tier 2 — AugmentedThis workflow Internal copilot where a human curates, challenges, and decides. Two steps landed here: the line-by-line PO reconciliation, and the auto-approval below $5,000. The workflow is governed at Tier 2 because Tier 2 is its highest step.
Tier 3 — Productivity General task automation with standard data-privacy guardrails. One step landed here: header and line-item extraction — read-only, no write access anywhere. This is the step the original pilot had already solved.

The permissions specification came straight out of fields 6 and 8: read the intake mailbox and the ERP PO table; write only to the match queue. Nothing else. The escalation path came straight out of field 5, unaltered. Neither was designed in a separate workshop, because neither needed to be.

The classification that felt like paperwork The AP lead's view was that extraction is "just automation" and did not need a tier. It got one anyway. Nine weeks later, when internal audit asked how a specific October invoice had been read and by what, the Tier 3 classification and its logging requirement were the reason there was an answer.
5Focus · The Stack Decision

Nothing Was Purchased Until the Map Was Finished

Two orchestration platforms had already been demoed to Northbridge before Creativa arrived, and one had a quote attached. Both were set aside until Systems Touched was complete — because choosing the stack before the map is finished is a named anti-pattern, not a head start.

Decision layerOptions consideredChosenBecause
Orchestration Enterprise agent platform (licensed) · cross-runtime framework · lighter build on frontier-model APIs Lighter build on frontier-model APIs Northbridge has one engineer who can maintain this. Tool sophistication is gated by maturity, not by what is newest.
Interoperability Cross-system agent communication standards now, or single-agent now and add later Single-agent now One agent, one workflow. Multi-agent protocols were deferred to the second wave, where they were actually needed.
Build / license / SaaS Decided per Agent Task, not per engagement Build on the existing ERP API The ERP was already licensed and already in Systems Touched. No new platform was purchased.
Cost guardrails $1,900 / month ceiling, alert at 70% Agentic workloads have a materially less predictable cost profile than SaaS seats. Every recommendation ships with a ceiling.
Observability New tooling vs. what the client already runs Existing log stack, extended New tooling is recommended only where a genuine gap exists against the Risk Tier requirement. There wasn't one.

Phase 1 closed on Day 20 with every step on the map carrying a named owner. Two lanes had come back from the first session marked "TBD"; the session did not close until both were named. A placeholder in a lane is a red flag, not something to fill in later.

6Execution · Going Live

Real Invoices, Real Permissions, From Day 21

No sandbox. The Matching Agent went live on Day 21 against every mismatched-PO invoice arriving that week — not a cherry-picked subset, and not a shadow mode running alongside the humans. A Governed Workflow Unit that has only ever seen synthetic cases has not been tested.

Scope
Every exception invoice
No subset, no shadow run
Permissions
Exactly as mapped
Nothing broadened for convenience
Team
5-person misfit pod
Incl. a named technical owner
Cadence
Weekly, binary
Done or not done, no partial credit

The pod followed Chain Reactor's Misfit Team pattern — a small cross-functional group with explicit air cover to bypass legacy sign-off during the sprint — plus the one AgentThrust addition: a named technical owner accountable for the orchestration stack and the audit logging, present from Phase 1 onward rather than handed the work at go-live.

Week 1 was worse than the baseline, and that was expected Exception volume in week 1 was 38% — identical to baseline. The agent was routing correctly; it simply had not yet met the freight-line variances that generate most of Northbridge's mismatches. Reporting an unchanged first week honestly is what made week 4's number believable.
7Execution · Weekly Telemetry

Exception Volume Was Reported as Loudly as Cycle Time

Four metric families, every week, without softening. The dashboard was built as Northbridge's — it had to keep working after Day 90 — and the metric that got the most attention was deliberately not the flattering one.

Exception volume, week by week

Share of invoices routed to the AP Clerk lane. The contracted floor was 15%; the stretch was 10%.

38%
Wk 1
31%
Wk 2
22%
Wk 3
14%
Wk 4
11%
Wk 5
9%
Wk 6

Contracted floor — 15%, cleared in week 4

Above floor Approaching At or below floor
Cycle time
0.8 d
from 4.2 d · under floor
Handoffs removed
61%
floor 60% · cleared wk 5
Exception volume
9%
floor 15% · stretch 10%
Audit completeness
97%
3% gap — see Step 8
Why exception volume is contracted at all A fast agent with a rising exception rate is not a working workflow — it is a fast happy path with a growing human queue behind it. Contracting exception volume is what stops the engagement from optimising the metric everybody enjoys watching.
8Execution · Closing the Gaps

The 3% Audit Gap Became an Atomic Unit

Whatever telemetry exposes becomes an Atomic Unit: the smallest verifiable change, one named owner, binary done or not done. Chain Reactor's microloop mechanics, unchanged — and the audit gap was the one that mattered most, because 97% audit completeness fails a 100% contract exactly as badly as 40% would.

Atomic UnitOwnerDueStatus
AU-1 · Tolerance rule for freight-line mismatchesTechnical ownerDay 34Done
AU-2 · Vendor-alias table seeded from the week-2 exception clusterAP leadDay 41Done
AU-3 · Confidence score written into the audit recordTechnical ownerDay 48Done
AU-4 · Approving identity (agent or human) captured per runTechnical ownerDay 55Done
AU-5 · Retire the shared matching spreadsheetAP leadDay 62Deferred to Phase 3

AU-3 and AU-4 together closed the audit gap: the 3% of runs missing a complete trail were all cases where a human had intervened mid-flow, and the log recorded the outcome without recording who had produced it. By Day 58, audit completeness read 100%.

The one that was deliberately deferred AU-5 — retiring the spreadsheet — was pushed to Phase 3 on purpose. Decommissioning the manual fallback is a No-Return Threshold condition, not an execution task, and doing it early would have removed the safety net before the threshold had actually been met.
9Verdict · The No-Return Threshold

Reinforcement Came Off on Day 71

The threshold is crossed when reverting to the manual process would cost more friction than sustaining the agent-owned one. Four conditions, all of which must hold. Creativa stopped resolving escalations on Day 71 and watched.

A full weekly cycle, no Creativa intervention
Days 71–78 ran clean with the pod observing only. Two escalation cycles were resolved without us.
Exception rate at or below the contracted floor
Steady at 9% against a 15% floor, and stable across three consecutive weeks rather than one good one.
The client's own team resolves escalations unaided
The AP team handled a vendor-alias failure and a genuine duplicate-invoice case on their own. This is the hard gate — declaring the threshold crossed while Creativa is still resolving escalations quietly voids the whole engagement.
The manual fallback formally decommissioned
AU-5 executed on Day 84: the shared matching spreadsheet was archived read-only and removed from the AP workflow. A fallback left warm is a fallback the organisation drifts back to.

No-Return Threshold confirmed: Day 84. From that date the agent-owned workflow is the system of record, with no parallel manual path running beside it.

10Verdict · The Day-90 Verdict

Scored Against the Contract Signed on Day 0

No renegotiation, no adjusted baseline, no metric quietly redefined between Day 0 and Day 90. Three families cleared stretch; one cleared floor but not stretch, and is reported as such.

Metric familyBaselineFloorDeliveredVerdict
Cycle-time reduction4.2 days1.0 day0.7 daysFloor cleared
Manual handoffs eliminated4 per exception60%64%Floor cleared
Exception volume38%15%9%Stretch cleared
Audit completenessPartial100%100%Stretch cleared

Cycle time landed at 0.7 days — comfortably past the 1.0-day floor, short of the 0.5-day stretch. The residual is structural rather than fixable by the agent: the ~9% of invoices that still need a buyer's answer are gated on a human replying to an email, and no amount of agent capability shortens that. It is reported as a floor-cleared result, not rounded toward a stretch it did not reach.

Day-90 Verdict

The full traceable chain: the contract signed before Day 1, against what the telemetry confirmed on Day 90.

4.2 days → 0.7 days · 64% of handoffs removed · 38% → 9% exceptions · 100% audit
Floor 1.0 day (cleared, Day 58) · Stretch 0.5 days (not cleared) · No-Return confirmed Day 84 · Risk Tier 2 — Augmented
Northbridge Components — one queue, one agent, one number that held
What a failed verdict would have looked like There is no indefinite pilot state in AgentThrust. Had the AP team still needed Creativa to resolve escalations at Day 84, the verdict would have read not owned in production, the manual fallback would have stayed live, and the report would have said so — with the engagement stopping rather than rolling into an extension.
11What Came Next · The Scheduled Swarm

The Second Workflow Had No Human on the Happy Path

Meridian approved a second AgentThrust cycle at the Day-90 review, this time at group treasury rather than inside a portfolio company. The workflow is a different shape entirely — and it is the case that tests whether the Workflow Map is a real primitive or just a diagram for simple flows.

  • The trigger is a schedule, not an event: 05:30 local every business day, plus a second run at 13:00 on month-end days. A trigger can carry more than one fixed time.
  • Four agents, not one: a Cash Orchestrator dispatching three workers that run in parallel — bank balances, closing FX rates, hedge exposure.
  • Two collaboration modes at once: the three workers are in role cooperation (they split the work, none waits on the others); the orchestrator is leader-driven (it alone decides).
  • Zero human steps when it runs clean — which is exactly why it is Tier 1 — Critical, not Tier 2.
Agent · Orchestrator Cash Orchestrator Plans the run, fans out to three workers, decides when to stop waiting. Tier 2 at its own step.
Agent · Knowledge Balances & FX Two retrieval agents pulling 14 bank accounts and closing rates, in parallel. Tier 3, read-only.
Agent · Document Hedge exposure Recomputes exposure against policy. It computes; it does not commit. Tier 3.
Human · escalation only Group Treasurer Off the happy path. Roughly two mornings a month, when the three returns disagree.

Three things about the second cycle were different, and all three came out of the map rather than out of a new methodology:

What changedWhy the map forced it
The fan-in became a mapped step Three agents returning independently can disagree, and one can return nothing. The step that reconciles them needs an owner and an audit record like any other — and the map has to say whether a partial return blocks the run.
The Trigger field grew a deadline An event-triggered workflow fails loudly because somebody is waiting. A scheduled one fails silently. The map carries the times and the rule: not complete by 06:15, the Treasurer hears about it.
Audit became per-agent attributable One audit line for the run stops being enough once four agents contributed to it. Each agent's inputs, outputs, confidence, and timing are logged separately, so a wrong number traces to the agent that produced it — non-negotiable at Tier 1.
The same eight fields, used more carefully No new primitive was created for the swarm. The Agent lane simply split — orchestrator above, workers beneath, each a named step — and the Exceptions field absorbed the partial-return case. If the map had needed rewriting to hold this, it would not have been a portfolio-wide primitive.
Walk the swarm on the live map — the third pattern on the toggle: The Workflow Map →
Conclusions

The Link Was the Deliverable

Northbridge already had a working agent on Day 0. What it did not have was anything connecting that agent to a number, an owner, a permission boundary, or an audit trail. Ninety days later the same underlying capability is doing recognisably the same work — and it is now the system of record for a queue that four people used to share, with a controller's name against the number.

1

The bottleneck was never the model. The pilot had solved extraction eight months earlier. What took 90 days was the permissions, the escalation path, the audit trail, and one named owner — none of which a better model would have supplied.

2

Mapping the mess is what made it survive. The undocumented spreadsheet and the buyer-reply wait were both invisible in the official process. Had the map skipped them, the Governed Workflow Unit would have broken on its first real exception instead of clearing its floor on Day 58.

3

The rejected candidates protected the cycle. Vendor onboarding had the loudest internal demand and no single owner; board-pack drafting had no binary outcome. Documenting why each lost is what kept the 90 days pointed at a workflow that could actually be owned.

4

Missing the stretch is not a failed engagement — it's an honest one. Cycle time stopped at 0.7 days because roughly 9% of invoices wait on a human answering an email. That residual is structural, it is named in the verdict, and it is the starting point for the next cycle rather than something rounded away in this one.

One 90-day cycle, one named workflow, one number locked before the clock started. The methodology does not promise that every workflow clears its stretch target, and it does not let a pilot run indefinitely while everyone agrees it is promising. It promises a link — from a capable agent to a business outcome with an owner, a boundary, and an audit trail behind it — and a verdict at Day 90 either way.