Home/By Use Case/AI Coding Productivity Studies
By Use Case

AI Coding Productivity Studies

Does AI actually make developers more productive? It’s the most important question in AI coding — and the honest answer is messier than either the hype or the backlash suggests. The most rigorous evidence is sobering: a randomized controlled trial by METR found experienced developers were 19% slower with AI on mature codebases, even though they predicted (and still believed afterward) they were faster. Other independent studies echo the theme — speed gains offset by more bugs, output rising only a few percent, benefits concentrated among senior developers. Yet vendor studies and developer self-reports tell a far rosier story, citing 50–100% gains and 81% self-reported productivity boosts. Both can be partly true: AI clearly accelerates well-scoped, greenfield, boilerplate work, while struggling on complex changes in large existing systems — and crucially, developers feel faster even when they aren’t. Underneath it all, adoption keeps climbing while trust falls. This page lays out the measured studies, the vendor and self-report claims, and the perception gap between them.

4 visualizations 5 sources Last updated June 2026 Free to embed
Loading chart...
Loading chart...
Chart 1 · The headline paradox

Measured vs. perceived productivity

METR RCT: predicted +24%, measured -19%, still felt +20%

Loading chart...
Source: METR randomized controlled trial 2025: 16 experienced open-source developers on mature 1M+ line codebases predicted AI would make them 24% faster. They were measured 19% slower — and even after finishing, still believed they’d been ~20% faster. The gap between felt and measured productivity is the core finding.
Chart 2 · Independent studies

What rigorous research actually found

Independent studies: modest or negative

Loading chart...
Source: METR / Uplevel / Science 2025–26: independent results are modest or negative — METR ’19%; Uplevel (~800 devs) found speed gains neutralized by 41% more bugs and no PR-cycle-time improvement; a study of 30M+ GitHub commits found just +3.6% quarterly output, with experienced devs capturing nearly all of it and early-career devs showing no benefit. Rigorous ≠ rosy.
Chart 3 · Vendor & self-report claims

The much rosier other side

Vendor & self-report claims: far rosier

Loading chart...
Source: Vendor & survey data 2026: the optimistic numbers come from vendors and self-reports — GitHub/Cursor cite 50–100% gains; senior developers self-report 81% productivity increases; 74% report productivity gains; IBM reports 60% dev-time reduction for internal apps. The catch: 95% report feeling productive while measurably producing lower-quality code.
Chart 4 · Use vs. trust

Adoption climbs as confidence falls

Adoption climbs while trust falls

Loading chart...
Source: Stack Overflow / survey trend 2026: the divergence — developer favorability toward AI tools fell from 77% (2023) to 60% (2026), and only 33% trust AI code accuracy (down from 43% in 2024), even as daily use keeps climbing. A Stanford RCT even found AI users wrote less secure code while feeling more confident. Adoption and trust are moving in opposite directions.

About this data

This page compiles AI coding productivity research, contrasting independent studies (METR’s RCT, Uplevel’s ~800-developer study, a 30M+ commit analysis published in Science, a Stanford security RCT) with vendor claims (GitHub, Cursor, IBM) and developer self-reports. It is the rigorous companion to our AI coding adoption overview.

The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.

Why measured and claimed productivity diverge so sharply: they measure different things on different work. Vendor and self-report figures often reflect felt speed on favorable tasks (greenfield, boilerplate, demos); independent studies measure actual time and quality on realistic work (mature codebases, full task completion, defect rates). The METR finding — developers 19% slower but convinced they were faster — is the single most important result, because it shows self-reports can’t be trusted as productivity measurement.

Methodology notes: the METR result is a randomized controlled trial on experienced open-source developers and may not generalize to all developers, tasks, or tools; it is, however, one of the few causal (not correlational) studies. Vendor figures (50–100%) are typically from controlled demos or favorable internal metrics and should be read as best-case. Self-reports measure perception, which the same research shows diverges from measured output. “Productivity” itself is contested — speed, output volume, and quality often move in different directions, so we present multiple measures rather than one number.

Sources used on this page:

Corrections or suggestions: research@aibehaviorindex.org

For Journalists & Researchers

Use this data in your work.

Every statistic, chart, and graphic in this index is free to use and cite, with full source attribution. Can’t easily find what you need? Use our search bar to search by keyword, topic, or category.

✉️
Talk to our research team
Need a specific cut of data, an interview, or a quote? Email us — we typically respond within one business day.
research@aibehaviorindex.org →

How the data works

Every statistic shown is sourced from a publicly available study, survey, or report. We aggregate, organize, and contextualize this data — but the underlying research is conducted by the cited sources. Click any source link to access the original methodology. If you run into any issues or have a study to suggest, contact us at research@aibehaviorindex.org.