The pilots were impressive. The invoices are real. Now CFOs want to know where the value went.
Artificial intelligence is moving from experimentation into the budget. That
changes the conversation. Finance leaders no longer need another
demonstration of what AI can do; they need a credible account of what it has
delivered—and what it has truly cost.
A model can save ten minutes without saving the company ten
minutes. Until the released capacity is removed, redeployed or
converted into better output, it is potential value—not financial
return.
The meeting where the hours disappeared
The slide looks excellent. Twelve thousand employee hours saved. Faster document reviews.
Thousands of automated responses. Adoption is up and the pilot has exceeded its target. Then the
CFO asks a rather ordinary question: where, exactly, did those hours go?
Payroll has not fallen. The team has not taken on materially more work. Customer response times are
much the same. Meanwhile, licence fees, integration work and cloud usage have all appeared in the
accounts. The room becomes noticeably less enthusiastic.
This is the AI ROI reckoning. It is not a backlash against artificial intelligence, and it is not proof that the
technology has been overhyped. It is the point at which experimentation meets financial discipline. After
several years of pilots, boards are entitled to ask whether impressive demonstrations are becoming
stronger margins, new revenue, faster cash conversion or better-controlled risk.
Investment is rising faster than proof
The appetite has not disappeared. Deloitte’s Q4 2025 CFO Signals survey found that 54% of North
American CFOs considered integrating AI agents into finance a leading transformation priority for 2026,
just ahead of improving data quality, access and usability at 52%. Gartner reported in February 2026
that nearly 60% of CFOs planned to increase finance-function AI investment by at least 10% during the
year.
The realised returns tell a more restrained story. PwC’s 29th Global CEO Survey found that 56% of
chief executives had seen neither additional revenue nor lower costs from AI over the previous 12
months. Only 12% reported both. These figures do not mean that most AI projects have failed. They do
mean that widespread use and financial impact are not the same thing.
That distinction matters because corporate language has become loose. A proof of concept becomes a
use case; a use case becomes a transformation programme; a productivity estimate becomes a saving.
Each step sounds plausible, but each requires evidence the previous one did not provide.
Productivity is not profit
There is good evidence that generative AI can improve performance in specific tasks. In a large study of
customer-support agents, researchers Erik Brynjolfsson, Danielle Li and Lindsey Raymond found that
access to an AI assistant increased issues resolved per hour by nearly 14%, with much larger gains
among less experienced workers. That is meaningful operational evidence—not a marketing claim.
But task productivity does not travel automatically to the income statement. If an employee completes a
report 20 minutes faster and then spends the time clearing email, the organisation may have improved
the working day without improving profit. That can still be worthwhile. It may reduce burnout, improve
service or create room for analysis. It simply should not be booked as a cash saving.
Finance leaders need to separate four things that are too often bundled together: time saved, capacity
released, capacity redeployed and financial value realised. Time is saved when a task takes fewer
minutes. Capacity is released when enough of those minutes accumulate in the same place to change
a workflow or role. It is redeployed when the organisation deliberately uses that capacity for additional
output. Financial value appears only when the change produces lower cost, additional revenue, faster
cash or a measurable reduction in loss or risk.
The gap between the first and fourth stages is where many confident ROI percentages are born.
The denominator nobody wants to discuss
AI business cases tend to give centre stage to benefits and a supporting role to costs. The software
subscription is counted; the work required to make the software useful is not. A credible denominator
includes model and platform charges, cloud consumption, data preparation, integration, cyber security,
legal review, testing, employee training, process redesign, monitoring and the cost of human oversight.
It should also include dual running. Important processes rarely move from human to automated
overnight. For months, sometimes longer, companies pay for the existing process and the new one
while teams verify outputs and resolve exceptions. More autonomous systems add further costs:
identity controls, permission management, audit trails, escalation rules and the ability to stop an agent
before a small error becomes a large one.
None of this is an argument against investment. It is an argument for honest investment. A project
whose return survives a complete cost model is far easier to defend—and far more likely to be
supported when budgets tighten.
Why pilots are designed to look successful
Pilots usually begin with willing users, carefully selected data and visible executive attention. The
process chosen is often repetitive enough to demonstrate a quick improvement, but contained enough
to avoid the hardest integration work. That is sensible for learning. It is not a reliable forecast of
enterprise economics.
Scale introduces the awkward cases: inconsistent data, regional variations, legacy systems,
permissions, resistance from experienced employees and customers who behave differently from the
test set. Accuracy that looks acceptable in a controlled trial can become expensive when errors require
specialist review. A five-minute saving is easily erased by a fifteen-minute exception.
The solution is not to make pilots larger. It is to state what they are meant to prove. One pilot may test
technical accuracy; another may test user adoption; a third may test unit economics. Asking a short
experiment to prove all three encourages false precision.
Build an ROI chain, not an ROI headline
The strongest AI business cases can be read as a chain of evidence. Start with a stable baseline:
current volumes, labour time, error rates, cycle times, conversion rates and losses. Record the full cost
of the proposed change. Then show the operational movement and the mechanism that converts it into
money.
For accounts payable, for example, the chain might run from automated invoice matching to fewer
manual touches, then to a smaller exception queue, earlier approvals and a higher capture rate for
early-payment discounts. For collections, it might run from better account prioritisation to more
productive calls, lower days sales outstanding and reduced financing cost. For sales support, it could
run from faster proposal preparation to more proposals per employee and, ultimately, incremental gross
profit—not simply more documents produced.
Every link needs an owner. Technology may own model performance, but operations must own
workflow change and finance must validate the financial outcome. If nobody is accountable for
converting released hours into a changed cost base or additional output, the final link is wishful thinking.
Use the right clock for the use case
A common mistake is to demand the same payback period from every form of AI. A tool that
summarises internal documents should show adoption and cycle-time improvement quickly. A
forecasting system may require several planning cycles before its accuracy can be compared fairly. A
customer-facing agent needs enough volume to reveal escalation rates, satisfaction and retention. A
new AI-enabled product may need a longer investment horizon altogether.
CFOs should still set milestones, but the milestones should fit the economic logic. Early indicators can
include utilisation, completion time, quality and exception rates. Later indicators should move towards
cost per transaction, revenue per employee, working-capital impact, customer retention or loss avoided.
The danger is allowing an early indicator to remain the success measure indefinitely.
Measure against what would have happened
AI often enters a business already changing. Demand may be rising, headcount may be frozen or a
wider systems programme may be improving performance at the same time. Before-and-after
comparisons can therefore flatter the technology—or understate it.
Where practical, compare similar teams, processes or markets, introducing the tool to one group before
another. At minimum, agree a counterfactual: what would cost, service and output have looked like
without the investment? Finance should challenge heroic assumptions about wage savings, adoption
and error-free output, and run the case at conservative, expected and upside levels.
This is less glamorous than a live demonstration. It is also how a promising tool becomes a defensible
capital-allocation decision.
Manage AI as a portfolio
Not every experiment deserves to scale. That sentence is easy to support in principle and surprisingly
difficult to practise once a senior sponsor, a team and a public target are attached to a project.
A portfolio view gives the organisation permission to make different decisions. Some applications
should scale because their economics are clear. Some need a process redesign before further
technology spending. Some have strategic or risk value even when direct financial return is modest.
Others should stop. Ending a weak pilot is not failure; continuing to fund one because the launch was
celebrated is.
A useful quarterly review is short and unsentimental: actual cost against plan; operational impact
against baseline; realised financial value; unresolved risk; adoption; and the next decision—scale,
redesign, hold or stop. This is more revealing than a catalogue of use cases and far harder to game.
The workforce question hiding inside the numbers
Many business cases assume that saved time will be absorbed neatly through attrition or higher-value
work. Sometimes it will. Often, the work itself has not been redesigned, managers do not know what
should replace the automated tasks, and employees quietly maintain the old process as insurance.
There is another issue. AI often benefits less experienced workers most, yet the routine work it absorbs
is also how people learn. If junior analysts no longer build the first draft, reconcile the basic schedule or
investigate the straightforward variance, finance must create a different route to judgement. Otherwise,
today’s productivity gain may become tomorrow’s capability gap.
The best CFOs will treat workforce design as part of the return, not an HR workstream added after
implementation. They will decide which work disappears, which work expands, what controls remain
human and how employees develop when the machine handles the apprenticeship tasks.
What the board should see
Boards do not need a running commentary on every model. They need a view of exposure and value. A
clear dashboard would show total AI investment; the small number of material use cases; benefits
claimed versus independently validated; the split between cashable, capacity and risk value; major
control incidents; and decisions taken on underperforming projects.
It should also make uncertainty visible. A range with explicit assumptions is more credible than a
precise ROI figure built on speculative hours. Risk reduction should be described with evidence—fewer
fraud losses, lower error rates, quicker regulatory response—not assigned an arbitrary monetary value
simply to complete the calculation.
This is where the CFO’s role becomes especially important. Finance can bring consistency across
enthusiastic business units and sceptical boards without becoming the department that blocks every
experiment. The aim is not to make innovation prove the unknowable. It is to stop possibility being
reported as performance.
The reckoning is a sign of maturity
AI does not need to deliver an immediate headcount reduction to be valuable. It may protect revenue,
improve decisions, make scarce expertise available to more people or allow a company to grow without
adding cost at the old rate. Those are legitimate returns when they are measured honestly.
What is ending is the period in which adoption itself could stand in for achievement. The next phase will
be quieter: fewer theatrical pilots, more workflow redesign; fewer sweeping transformation claims, more
unit economics; fewer hours theoretically saved, more value that someone can point to in the operating
plan.
That is not an AI winter. It is what happens when a technology becomes important enough to be
managed like a business.