-
Notifications
You must be signed in to change notification settings - Fork 4
Value And The CFO
Time saved is not value realised.
A programme review, eight months after go-live. The team presents a slide. Forty-seven thousand hours saved.
The finance lead looks at it for a while.
"What did we do with those hours?"
Silence.
"Did headcount go down? Did revenue go up? Did we serve more customers?"
More silence.
"Then we saved forty-seven thousand hours of cost that we are still paying."
(Composite. The scene is assembled from a pattern that recurs, not from one meeting.)
She was right, and the slide was not dishonest. It was activity tracking wearing the clothes of a business case.
Value is realised when time saved produces a measurable outcome: cost removed, capacity added, revenue created, or risk reduced. Everything else is a leading indicator at best.
Not three grades of quality. Three different kinds of claim, each proved with different evidence, each believed by a different person.
| Tier | The claim | How it is proved |
|---|---|---|
| 1 · Efficiency | Cost came out | Roles freed multiplied by fully loaded cost. Temporary labour invoices before and after. Rework hours multiplied by rate |
| 2 · Effectiveness | We do more without spending proportionally more | Transactions per person before and after. Error or rejection rate trend. Cycle time against baseline |
| 3 · Transformation | New revenue, or a material risk avoided | Incremental recovery. Risk event cost multiplied by the reduction in probability. New product or service revenue |
Two things about this ordering.
Tier 1 is the only one most programmes ever claim, and it is the hardest one to prove. Cost out means an invoice got smaller. If no invoice got smaller, the claim is tier 2 wearing a tier 1 label.
Tier 3 is where the interesting value is and it is the tier nobody instruments for. It requires a baseline established before deployment, and by the time anyone wants the number the baseline has gone.
Before the first pound is spent on development, not after. Not because a model needs to be precise, but because building it forces the team to say out loud what they are building and why.
Cost per transaction, human. Fully loaded hourly rate multiplied by handling time in hours.
Cost per transaction, agent. Inference plus platform plus monitoring, divided by transactions processed.
Break-even. The transaction volume at which the second is below the first.
Value at scale. The difference, multiplied by annual volume.
Monitoring belongs in that second line and is routinely left out. An agent nobody watches is cheaper and it is not the thing being costed.
Then apply the rule that makes the model honest:
Use a fraction of projected volume in year one. Use less than projected accuracy. Add a quarter to infrastructure and governance costs. If the case still works under those assumptions, it is a real case.
The reason to be pessimistic on purpose is not caution. It is that the conservative model is the one that survives the second year, and a programme killed at month eighteen on economics costs more than a programme that was never funded.
Declaring value too early kills programmes. Declaring it too late loses the sponsor. A fixed cadence solves both by making the reporting boring.
Week four, at the end of shadow mode. Accuracy against the human baseline. The failure patterns found. A revised timeline if the first two require one.
Week eight, at first live volume. How much volume the agent handled. Error rate at that volume. Cost per transaction measured in the live environment rather than estimated.
Day ninety, the business review. Transactions processed. Handling time against baseline. Error rate against baseline. Hours redeployed. A preliminary return calculated from actual data.
Month six, the finance conversation. Realised savings, verified against payroll or supplier invoices. Throughput increase with supporting data. Any tier 3 evidence. A forward projection revised from actual performance.
Notice what changes between day ninety and month six. The day ninety number is calculated. The month six number is verified against something outside the programme's own reporting. That is the whole difference between a business case and a claim.
When a report says hours were redeployed, the people behind that number are owed a sentence saying where to.
A weak report says two thousand four hundred hours were saved.
A good one says two thousand four hundred hours previously spent on a task nobody valued have been redirected into work that was previously rationed, and states what more of that work became possible.
This is not presentation. It is the difference between a metric and a mandate, and it is the thing that decides whether the next agent is welcomed or resisted.
A programme that cannot say where the hours went has not finished measuring.
Derived from the AI CoE and Agent Factory Playbook, based on The Augmented Enterprise framework.
The thinking
Operating model
Frameworks
Governance
Playbooks
Value and people
Reference
In the repository
The courses
Reviewed 2026-08.