AI Coding Assistants Don't Make Teams Faster. They Move the Bottleneck.
The claim everyone repeats, and the study that complicates it
Ask most engineering leaders what AI coding assistants do to a team's workflow and you get some version of "developers ship faster." That's the pitch, and for a certain kind of work — boilerplate, unfamiliar syntax, a first draft of a function whose shape you already know — it holds up. But the most rigorous study on the subject so far found almost the opposite result for a specific, important population: developers who already know their codebase well.
In July 2025, METR ran a randomized controlled trial with 16 experienced open-source developers working in large, mature repositories they knew intimately (averaging over 22,000 GitHub stars and a million-plus lines of code). Each of 246 real tasks was randomly assigned to be done with AI assistance (mainly Cursor Pro, using Claude 3.5 and 3.7 Sonnet) or without it. The developers predicted AI would make them about 24% faster. It actually made them 19% slower. After seeing their own timing data, they still insisted AI had sped them up by about 20%.
The mechanism matters more than the headline number: developers accepted less than 44% of AI-generated suggestions, and a meaningful share of the "AI time" was spent reviewing and correcting output that didn't fit the codebase's actual constraints — constraints the model couldn't see because they lived in the developer's head, not in the prompt.
That's not an argument against AI coding tools. It's an argument that the honest version of "what changes" is more specific than "everything gets faster," and the specifics are where a buyer or an engineering lead should be looking.
Where the bottleneck actually moves
If writing code gets faster — and for a lot of tasks, it genuinely does — the bottleneck doesn't disappear. It moves downstream, to review. Faros AI's telemetry, cited in DORA's 2025 State of DevOps report, found that AI adoption correlates with developers merging 98% more pull requests that are on average 154% larger, with review times running 91% longer. Individual output went up. Organizational delivery performance stayed roughly flat. The 2025 DORA report's framing is the right one: AI acts as a multiplier on whatever engineering discipline already exists. Teams with clear review standards and small, well-scoped changes get more of those good changes, faster. Teams without that discipline get a review queue full of larger, AI-generated diffs that pass CI and then sit for days waiting for a human to actually understand them.
This is the part that doesn't show up in adoption statistics. JetBrains' 2026 survey of more than 15,000 developers found 90% using AI coding agents weekly and 68% daily, with 84% reporting they feel more productive. Feeling more productive and the team actually shipping more reliably are different claims, and the gap between them is exactly what the METR and DORA data are pointing at. Separately, 39% of organizations in that same survey population report having no way to measure AI's actual impact on their work, and another 33% are relying on developer self-report — the same self-report that, in the METR study, was off by 39 percentage points from measured reality.
What this means in practice, not in theory
A few concrete implications, not generic ones:
The task matters more than the tool. AI assistants are genuinely strong at scaffolding new projects, writing tests against a spec, translating between languages or frameworks, and producing a reasonable first draft in a part of the codebase nobody on the team has touched recently. They're weaker — and the METR data backs this up directly — when a senior developer is working inside a system they already understand well, where the model's suggestions have to be checked against institutional knowledge the model never had access to.
Review process has to change, not just scale. If AI output makes PRs larger and more frequent, running the same review process at higher volume is how delivery instability creeps in — exactly what DORA found. The fix isn't "review faster," it's smaller diffs, tighter scope per PR, and treating AI-generated code with the same scrutiny you'd give a contributor's first PR, every time, regardless of how clean it looks.
Don't let "feels faster" substitute for measurement. If a third of organizations can't say whether AI is actually helping, and the one controlled study we have shows developers misjudging their own speed by a wide margin, then internal perception is not a metric you can plan headcount or timelines around. Track cycle time and defect rates before and after adoption on comparable work, not sentiment.
For teams evaluating a vendor, this is a due-diligence question, not a trivia question. If a dev studio or in-house team tells you they "use AI heavily," the useful follow-up isn't whether — it's where. Scaffolding and tests, where the acceleration is real? Or deep changes in a codebase they're still learning, where the evidence says it can cost more time than it saves, with the added risk that nobody notices because it felt fast?
The takeaway
AI coding assistants don't make an engineering team faster by default — they change where the work and the risk sit, shifting effort from writing to reviewing and rewarding teams that already have tight scoping and review discipline. The honest question for any team adopting them isn't "how much faster are we," it's "faster at what, slower at what, and how do we actually know."
Have a system like this in mind?
Get a scoped plan ↗