Everyone quotes an AI productivity number. Almost nobody quotes the study it came from, and the studies disagree with each other in ways that turn out to be the most useful part.
Stanford’s AI Index 2026 collects them in one place (Ch4, Fig 4.4.27). Here is the honest version.
| Occupation | Application | Measured effect | Who gained most |
|---|---|---|---|
| Customer support agents | conversational assistant | +14–15% issues resolved per hour | less experienced agents (30–35%) |
| Software developers | GitHub Copilot | +26% pull requests completed | junior and less-experienced |
| Marketing teams | multimodal ad creation | +50% output per worker | human–AI teams |
| Accountants | AI-based accounting support | +55% weekly client throughput | experienced accountants |
| Authors | LLMs for content | +200% output volume | new entrants |
| Software engineers | learning new libraries | 0% (statistically insignificant) | n/a |
| Experienced open-source developers | AI assistance | −19% speed | nobody |
Two entries on that list are the reason the others are worth trusting.
The −19% study, and what happened to it
The most-cited negative result comes from Model Evaluation & Threat Research (METR), which found experienced open-source developers were 19% slower when using AI assistance, while believing they had been faster. The gap between perceived and actual help is the finding that stuck, and it deserved to.
The AI Index also reports what happened next, which almost nobody quotes: METR has not been able to replicate the result. The follow-up ran into a practical obstacle: developers had become reluctant to work without AI at all, and the team’s assessment is that developers in late 2025 were likely sped up by AI relative to the original study period (Ch4).
So the −19% is neither debunked nor confirmed. It is a real result from a specific population at a specific moment, on unfamiliar large codebases, that did not reproduce a year later. That is what most evidence in this field looks like, and pretending otherwise is how organizations end up with business cases nobody can defend.
The second inconvenient entry is the 0%: engineers using AI to learn new libraries showed no measurable speed improvement, and researchers observed what they termed “learning penalties”: heavy reliance slowing skill development over time (Shen and Tamkin 2025, in Ch4). The engineers who avoided the penalty were the ones using AI for conceptual inquiry rather than answer retrieval. Speed on the task and capability afterwards are different variables, and only one of them shows up in a quarterly metric.
Why do the numbers vary so much?
Because the tasks do. The AI Index states the pattern directly, and the sentence functions as a scoping rule:
Gains are strongest when work can be divided into well-defined, repeatable tasks with clear quality monitoring.
Read the table again through that lens and it stops looking contradictory. Support tickets are well-defined, repeatable and monitored: +14%. Pull requests are well-defined, repeatable and reviewed: +26%. Ad variants are repeatable, and quality is measured by a market: +50%. Novel work on a large unfamiliar codebase, where “correct” takes an expert an hour to establish: −19%.
Three tests, then, before you commit budget to automating something:
Is the task well-defined? Can you write down what a correct output looks like without using the word “good”?
Is it repeatable? Does it recur often enough that improvement compounds, and consistently enough that you can build evaluation around it?
Is quality monitorable? Can you tell a good output from a bad one quickly and cheaply, ideally automatically? This is the test most candidate use cases fail, and it fails silently, because a pilot with no quality signal will report success on volume.
Fail any one of the three and the project is not impossible; it is simply the harder version, and it should be budgeted, staffed, and communicated as such.
What happens at the level of the whole company
The individual-task gains are large and repeatedly measured. The organization-level gains are, so far, small and contested (Ch4, Fig 4.4.28).
| Study | Scope | Finding |
|---|---|---|
| Aldasoro et al. 2026 | 12,000 European firms | +4% labour productivity; +5.9pp for every 1% of spend on training |
| Brynjolfsson 2026 | US economy | 2.7% productivity growth in 2025 vs 1.4% decade average; framed as a “J-curve” |
| Filippucci et al. (OECD) 2025 | G7, 10-year horizon | +0.4 to +1.3pp/yr (US/UK); +0.2 to +0.8pp (Italy/Japan) |
| St. Louis Fed 2025 | US labour market | +1.1% to +1.3% |
| Yotzov et al. 2026 | 6,000 executives, US/UK/DE/AU | high adoption, minimal realized productivity impact to date |
| Penn Wharton 2025 | US economy | +0.01pp contribution to total factor productivity (negligible) |
A +26% gain on individual pull requests and a rounding error in aggregate productivity are both true at once. The usual explanation is dilution: a task improves, the process around it doesn’t, and the gain is absorbed by handoffs, review queues, and rework long before it reaches a financial statement. The organizational adoption data supports that reading: most organizations are still piloting rather than running AI in production, a gap we cover in the AI execution gap.
The one intervention with a clean multiplier attached is the least technical thing in the report: +5.9 percentage points of productivity gain for every 1% of spend that went to training. It is also, reliably, the first line cut from a proposal.
How to build a business case that survives contact with a CFO
Measure the baseline first, and write it down. This is the mistake that cannot be corrected later: once the old process is gone, so is your comparison. Cycle time, cost per unit, error rate, rework rate. Whatever you will be asked about in twelve months.
Pick one task that passes all three tests. Not a department. Not a “transformation.” One task, chosen because it is defined, repeated, and measurable.
Agree the success metric before building, with the person who will judge it. A pre-agreed number converts an argument into an observation.
Instrument quality, not just volume. Output went up is not a result. The +200% author study is instructive here: output volume tripled, and the researchers separately tracked whether quality held.
Budget the enablement. It has the best-evidenced return in the report.
Expect the J-curve. Costs land in quarter one; returns land later or not at all. A business case that promises quarter-one returns is a business case that will be judged in quarter one.
And be willing to publish the negative result internally. An organization that can say “we tried this, it produced no measurable gain, here is what we learned about why” is an organization that will eventually find the cases that work. One that only reports wins is accumulating a portfolio of pilots nobody can evaluate, which is, statistically, where most of the 88% currently are.
Our AI implementation discovery checklist covers the readiness assessment behind this, and proof-of-value pilots covers structuring the pilot itself. If you want the scoping done properly before the budget is committed, that is what our AI implementation engagements start with.
All figures cited from the Stanford HAI Artificial Intelligence Index Report 2026 (9th edition), which reports the underlying studies named above. Individual study methodologies, populations, and time periods vary; read them before relying on a single number.



