August 19, 2026

What the AI Productivity Studies Actually Say | CoreLine

+26% for developers, +55% for accountants, -19% in one developer study. Stanford's AI Index collects the evidence - and it contains a scoping rule for choosing what to automate.
date
August 19, 2026
categories
categories
Development
What the AI Productivity Studies Actually Say
table of contents

Everyone quotes an AI productivity number. Almost nobody quotes the study it came from, and the studies disagree with each other in ways that turn out to be the most useful part.

Stanford’s AI Index 2026 collects them in one place (Ch4, Fig 4.4.27). Here is the honest version.

OccupationApplicationMeasured effectWho gained most
Customer support agentsconversational assistant+14–15% issues resolved per hourless experienced agents (30–35%)
Software developersGitHub Copilot+26% pull requests completedjunior and less-experienced
Marketing teamsmultimodal ad creation+50% output per workerhuman–AI teams
AccountantsAI-based accounting support+55% weekly client throughputexperienced accountants
AuthorsLLMs for content+200% output volumenew entrants
Software engineerslearning new libraries0% (statistically insignificant)n/a
Experienced open-source developersAI assistance−19% speednobody

Two entries on that list are the reason the others are worth trusting.

The −19% study, and what happened to it

The most-cited negative result comes from Model Evaluation & Threat Research (METR), which found experienced open-source developers were 19% slower when using AI assistance, while believing they had been faster. The gap between perceived and actual help is the finding that stuck, and it deserved to.

The AI Index also reports what happened next, which almost nobody quotes: METR has not been able to replicate the result. The follow-up ran into a practical obstacle: developers had become reluctant to work without AI at all, and the team’s assessment is that developers in late 2025 were likely sped up by AI relative to the original study period (Ch4).

So the −19% is neither debunked nor confirmed. It is a real result from a specific population at a specific moment, on unfamiliar large codebases, that did not reproduce a year later. That is what most evidence in this field looks like, and pretending otherwise is how organizations end up with business cases nobody can defend.

The second inconvenient entry is the 0%: engineers using AI to learn new libraries showed no measurable speed improvement, and researchers observed what they termed “learning penalties”: heavy reliance slowing skill development over time (Shen and Tamkin 2025, in Ch4). The engineers who avoided the penalty were the ones using AI for conceptual inquiry rather than answer retrieval. Speed on the task and capability afterwards are different variables, and only one of them shows up in a quarterly metric.

Why do the numbers vary so much?

Because the tasks do. The AI Index states the pattern directly, and the sentence functions as a scoping rule:

Gains are strongest when work can be divided into well-defined, repeatable tasks with clear quality monitoring.

Read the table again through that lens and it stops looking contradictory. Support tickets are well-defined, repeatable and monitored: +14%. Pull requests are well-defined, repeatable and reviewed: +26%. Ad variants are repeatable, and quality is measured by a market: +50%. Novel work on a large unfamiliar codebase, where “correct” takes an expert an hour to establish: −19%.

Three tests, then, before you commit budget to automating something:

Is the task well-defined? Can you write down what a correct output looks like without using the word “good”?

Is it repeatable? Does it recur often enough that improvement compounds, and consistently enough that you can build evaluation around it?

Is quality monitorable? Can you tell a good output from a bad one quickly and cheaply, ideally automatically? This is the test most candidate use cases fail, and it fails silently, because a pilot with no quality signal will report success on volume.

Fail any one of the three and the project is not impossible; it is simply the harder version, and it should be budgeted, staffed, and communicated as such.

What happens at the level of the whole company

The individual-task gains are large and repeatedly measured. The organization-level gains are, so far, small and contested (Ch4, Fig 4.4.28).

StudyScopeFinding
Aldasoro et al. 202612,000 European firms+4% labour productivity; +5.9pp for every 1% of spend on training
Brynjolfsson 2026US economy2.7% productivity growth in 2025 vs 1.4% decade average; framed as a “J-curve”
Filippucci et al. (OECD) 2025G7, 10-year horizon+0.4 to +1.3pp/yr (US/UK); +0.2 to +0.8pp (Italy/Japan)
St. Louis Fed 2025US labour market+1.1% to +1.3%
Yotzov et al. 20266,000 executives, US/UK/DE/AUhigh adoption, minimal realized productivity impact to date
Penn Wharton 2025US economy+0.01pp contribution to total factor productivity (negligible)

A +26% gain on individual pull requests and a rounding error in aggregate productivity are both true at once. The usual explanation is dilution: a task improves, the process around it doesn’t, and the gain is absorbed by handoffs, review queues, and rework long before it reaches a financial statement. The organizational adoption data supports that reading: most organizations are still piloting rather than running AI in production, a gap we cover in the AI execution gap.

The one intervention with a clean multiplier attached is the least technical thing in the report: +5.9 percentage points of productivity gain for every 1% of spend that went to training. It is also, reliably, the first line cut from a proposal.

How to build a business case that survives contact with a CFO

Measure the baseline first, and write it down. This is the mistake that cannot be corrected later: once the old process is gone, so is your comparison. Cycle time, cost per unit, error rate, rework rate. Whatever you will be asked about in twelve months.

Pick one task that passes all three tests. Not a department. Not a “transformation.” One task, chosen because it is defined, repeated, and measurable.

Agree the success metric before building, with the person who will judge it. A pre-agreed number converts an argument into an observation.

Instrument quality, not just volume. Output went up is not a result. The +200% author study is instructive here: output volume tripled, and the researchers separately tracked whether quality held.

Budget the enablement. It has the best-evidenced return in the report.

Expect the J-curve. Costs land in quarter one; returns land later or not at all. A business case that promises quarter-one returns is a business case that will be judged in quarter one.

And be willing to publish the negative result internally. An organization that can say “we tried this, it produced no measurable gain, here is what we learned about why” is an organization that will eventually find the cases that work. One that only reports wins is accumulating a portfolio of pilots nobody can evaluate, which is, statistically, where most of the 88% currently are.

Our AI implementation discovery checklist covers the readiness assessment behind this, and proof-of-value pilots covers structuring the pilot itself. If you want the scoping done properly before the budget is committed, that is what our AI implementation engagements start with.

All figures cited from the Stanford HAI Artificial Intelligence Index Report 2026 (9th edition), which reports the underlying studies named above. Individual study methodologies, populations, and time periods vary; read them before relying on a single number.

need a second opinion?
Talk to our engineers about your architecture, stack, or delivery challenge.