Prove your agent's ROI in 6 weeks: one spreadsheet, three numbers, a kill switch
In July 2025, METR timed 16 experienced developers across 246 real tasks. With AI tools allowed, they took 19% longer to finish. Afterwards, the same developers estimated the tools had made them 20% faster. Sit with that gap for a second. Measured: slower. Felt: faster.
Operators self-report the same way. 91% of SMB leaders using AI say it boosts revenue (Salesforce, 2024). Say. Meanwhile PwC asked 4,454 CEOs this January, and 56% of large-enterprise CEOs report no significant financial benefit from AI at all. Only 12% see gains in both cost and revenue. Somewhere between "everyone feels faster" and "nobody can show the money" sits your next budget conversation.
You don't need a data team to close that gap. You need a spreadsheet, three numbers, and six weeks. Last week's post in this circle tore down the build: one governed skill. This one is the proof that it paid for itself. The full plan, runnable as written.
Week 0: pick one process and set the kill switch. One workflow, not the rollout. It qualifies if it runs at least 20 times a week, follows the same steps every time, and has one owner who can walk you through it without getting defensive. Then pick the metrics. Time per task. Error rate. Volume handled. Not eleven metrics. Three. And write the kill criteria now, before you're emotionally invested: "if time per task hasn't dropped by week 4, we stop and write up why." Only about one in four enterprise CEOs say their organization has disciplined processes for stopping underperforming initiatives (PwC, 2026). Be the one in four.
Weeks 1 and 2: baseline the old way. One row per run, logged by the person doing the work, filled in at the end of each day. Five minutes, no new software. Define on day one what counts as a completed run. Change nothing else. Without a recorded baseline, every improvement you claim later is an estimate wearing the costume of evidence.
Weeks 3 and 4: turn the agent on. Same three metrics, same people logging, same time of day. Log every failure, redo, and exception, because documented failures are what make your time-savings figure believable. A suspiciously clean win invites exactly the scrutiny you were trying to avoid.
Weeks 5 and 6: turn minutes into money. The formula, and the write-up of every assumption behind it, so anyone in finance can reproduce your number.
The entire measurement stack, one tab for before and one for after:
date | run | minutes, start to finish | error or redo? | notes
Mon 09 | contract intake #114 | 45 | no | waited 20 min on legal
Mon 09 | contract intake #115 | 52 | yes | wrong entity name, redone
The conversion:
(old minutes per task - new minutes per task)
x weekly volume
x fully loaded hourly cost = weekly savings
weekly savings x 52 = annual figure
project cost / weekly savings = payback, in weeks
Worked through at Acme SaaS (placeholder math, not a client outcome): contract intake drops from 45 to 20 minutes. 60 runs a week, $40 fully loaded. That's 25 hours back per week, $1,000 a week, $52,000 a year. If the build cost $10,000, it paid for itself in week 10. A small credible number like that beats a big unverifiable one in every budget meeting I've sat in.
And if week 4 shows the after-numbers flat? Kill it, and write up why. A killed pilot with clean documentation reads as more credible than a rollout that underperforms in silence for a year. Forcing a good story out of flat data only costs you the next budget conversation.
Agent already live and you never baselined? There's a reconstruction method (ticket timestamps, email history, system logs from the two weeks before go-live) in the full post, along with the FAQ on what a realistic first-project saving looks like: How to prove AI ROI in 6 weeks, no data team required
Now the part where you do it. Comment with three numbers from one workflow you run: minutes per task, runs per week, fully loaded hourly cost. I'll run your payback math in the replies. Rough numbers are fine. "We stopped measuring after week three" is also fine; say so. Being certain a tool is helping is not evidence that it is. Write the number down.