Your AI Numbers Are Going to Look Worse Before They Look Better

If you have not said that out loud to your CFO yet, you have a problem, and it is not a technical one.

Your AI numbers are going to look worse before they look better. A J-curve drops from the pre-AI baseline to a trough labelled the tuition cost of transformation, then climbs to higher long-term value. Three things dig that hole: the learning curve, pipeline adaptation, and the verification tax. Plan the dip. Invest through it. Capture the return.

Google's DORA team published The ROI of AI-assisted Software Development in April 2026. The finding that matters to anyone holding a budget is not about models or tools. It is about shape.

Value from AI adoption follows a J-curve. Output drops below where it was before you started. Then it climbs. DORA calls the trough the tuition cost of transformation.

Three things dig that hole

That third one is the real bill. You did not remove work. You moved it from authorship to verification, and verification is the more expensive of the two because it requires reconstructing intent that was never written down.

Where leaders lose the plot

They measure at the bottom of the curve and call it failure.

Adoption is climbing. Delivery has slowed. The quarterly review lands. Somebody asks where the return is, and the honest answer is "we are in the trough, this was expected" only if you said it was expected beforehand.

Say it afterward and it sounds like an excuse. Say it beforehand and it is a plan.

Budgets get cut at the trough. That is the whole failure mode. You fund the tools, you absorb the dip, you cancel the program two months before the slope turns, and you keep every cost while forfeiting the entire return.

The part that changes where you point it

Stanford research cited in the same report: AI delivers 35 to 40% gains on simple greenfield tasks, and under 10% on complex legacy code.

Most enterprise work is the second kind. If you piloted on greenfield and projected those numbers across a fifteen-year codebase, your business case was fiction before the first sprint.

And if you run in a regulated environment, price the verification tax higher than the report does. Verification there is not only "is this correct." It is "can I evidence that it is correct to an auditor." That is a different unit of work and it does not shrink just because the code arrived faster. Who signs for agent-written code is that unit of work, spelled out.

What to do about it

None of this is an argument against AI in the lifecycle. I run agents across mine. It is an argument against promising a straight line when the research says the line dips first. It is also why the operating model puts a baseline before adoption: without one, the trough and the climb are both just stories, and the rest of the 2026 data says the climb only arrives for the organizations that redesigned the work.

The question for the leaders reading this

Has anyone in your company said the word "dip" to finance out loud?

Or is the plan to hope the chart turns before someone asks?