Your AI Numbers Are Going to Look Worse Before They Look Better
Budget planning, 2026
If you have not said that out loud to your CFO yet, you have a problem, and it is not a technical one.

Google's DORA team published The ROI of AI-assisted Software Development in April 2026. The finding that matters to anyone holding a budget is not about models or tools. It is about shape.
Value from AI adoption follows a J-curve. Output drops below where it was before you started. Then it climbs. DORA calls the trough the tuition cost of transformation.
Three things dig that hole
- The learning curve. Teams figuring out interfaces, workflows, and how to actually drive the thing.
- Pipeline adaptation. Code arrives faster, and testing and approval stages sized for hand-authored volume start to choke.
- The verification tax. The one nobody budgeted for: the time your engineers now spend checking AI output for hallucinations and verifying it against security and architectural standards.
That third one is the real bill. You did not remove work. You moved it from authorship to verification, and verification is the more expensive of the two because it requires reconstructing intent that was never written down.
Where leaders lose the plot
They measure at the bottom of the curve and call it failure.
Adoption is climbing. Delivery has slowed. The quarterly review lands. Somebody asks where the return is, and the honest answer is "we are in the trough, this was expected" only if you said it was expected beforehand.
Say it afterward and it sounds like an excuse. Say it beforehand and it is a plan.
Budgets get cut at the trough. That is the whole failure mode. You fund the tools, you absorb the dip, you cancel the program two months before the slope turns, and you keep every cost while forfeiting the entire return.
The part that changes where you point it
Stanford research cited in the same report: AI delivers 35 to 40% gains on simple greenfield tasks, and under 10% on complex legacy code.
Most enterprise work is the second kind. If you piloted on greenfield and projected those numbers across a fifteen-year codebase, your business case was fiction before the first sprint.
And if you run in a regulated environment, price the verification tax higher than the report does. Verification there is not only "is this correct." It is "can I evidence that it is correct to an auditor." That is a different unit of work and it does not shrink just because the code arrived faster. Who signs for agent-written code is that unit of work, spelled out.
What to do about it
- Name the dip before you spend the money. Depth and duration, in writing, to the person who controls the budget.
- Instrument verification time as its own line item, not as a mysterious slowdown in delivery metrics.
- Fix review and test capacity before you scale generation, not after the queue backs up.
- Segment your projections by task type. Greenfield and legacy are not the same investment and should never share a forecast.
None of this is an argument against AI in the lifecycle. I run agents across mine. It is an argument against promising a straight line when the research says the line dips first. It is also why the operating model puts a baseline before adoption: without one, the trough and the climb are both just stories, and the rest of the 2026 data says the climb only arrives for the organizations that redesigned the work.
The question for the leaders reading this
Has anyone in your company said the word "dip" to finance out loud?
Or is the plan to hope the chart turns before someone asks?