Short answer: yes for the individual, not yet for most teams. In studies from 2025 and 2026, developers with AI assistants finish more tasks and merge more pull requests. Delivery data from the same period shows review queues growing, refactoring falling and stability slipping. The gain is real. Most of it gets spent before it reaches production.
The question usually arrives as a budget problem. The licences are paid for, the engineers say they're quicker, and the release calendar looks the same as it did a year ago.
So we read the research. Eight reports, all published since mid-2025, each counting something slightly different. This note goes through what each one measured, where they disagree, and what we think a team with a long-running product should take from them.
What the studies found, side by side
METR, July 2025. A randomized trial with 16 experienced open-source developers and 246 tasks. With AI tools they were 19% slower.
METR, February 2026. The rerun: 57 developers, more than 800 tasks. Returning developers came out about 18% faster and new ones 4%. Neither figure is statistically clear, and METR calls the data unreliable.
Faros AI, July 2025. Telemetry from over 10,000 developers on 1,255 teams. 21% more tasks, 98% more pull requests merged, review time up 91%. No measurable change at company level.
DORA, 2025. A survey of nearly 5,000 people. AI adoption now goes with higher delivery throughput and, still, with lower stability.
GitClear, 2026. 623 million code changes. Refactored lines fell from 13% of changes in 2023 to 3.8% in 2026, and duplicated blocks rose 81%.
Veracode, July 2026. Security tests on AI-generated code. The average pass rate is 56%; a year earlier it was 55%.
Stack Overflow, 2025 and 2026. 84% of developers used or planned to use AI tools in 2025, and 46% distrusted the accuracy of what came back. By 2026, 73% used coding assistants or agents daily.
A caveat before going further. Three of these come from vendors that sell tools for the problem they measured (Faros, GitClear and Veracode), and two are surveys of opinion. Only METR ran a controlled experiment, and its first sample was 16 people. No single report settles the question, which is why they're worth reading together.
Why did developers feel faster when they were slower?
METR's first study is the one everybody quotes, and the number people remember (19% slower) is the less interesting half. Before starting, the developers predicted AI would make them 24% faster. After finishing, having actually been slower, they still believed it had sped them up by 20%.
That gap is the finding. A team's own sense of speed is not evidence, and neither is a survey that asks people how productive they feel. More than 80% of DORA's respondents say AI has raised their productivity. They may be right. But METR's developers would have ticked the same box.
It's worth being fair to the tools here. Those 16 people worked on repositories they knew deeply (22,000+ stars, over a million lines of code) with Cursor and Claude 3.5 and 3.7 Sonnet, which were current in early 2025 and aren't now. METR says plainly that the result does not generalize to most developers.
Did the picture change in 2026?
Probably, though nobody has measured it cleanly.
METR ran the experiment again from August 2025, with 57 developers. The estimate flipped: people from the original group were about 18% faster with AI, new recruits about 4%. Both confidence intervals include zero, so neither result can be told apart from no effect.
Then the study hit a problem that says more than its numbers do. Between 30% and 50% of participants admitted they held back tasks they didn't want to do without AI. The work where AI helps most never entered the comparison. METR called its own data an unreliable signal and is redesigning the experiment — while adding that it believes developers are likely more sped up in early 2026 than a year before.
We find that detail more persuasive than any percentage. Experienced engineers, paid by the hour, declining to do certain tasks the old way.
If individuals are faster, why isn't the team shipping more?
Because writing code is one stage of several, and the others didn't speed up.
Faros AI looked at delivery telemetry instead of opinions. On teams with high AI adoption, developers completed 21% more tasks and merged 98% more pull requests. Review time rose 91%. Pull requests got 154% larger and bugs per developer went up 9%. At company level, the link between AI adoption and delivery outcomes disappeared.
Read those figures as a queue. Twice as many pull requests, each much bigger, landing on the same reviewers who already had a full day. The speed-up is produced at the keyboard and absorbed in review, testing and release.
DORA's 2025 report describes the same thing from the survey side. For the first time it found AI adoption associated with higher delivery throughput; a year earlier that relationship was negative. Stability still gets worse as adoption goes up. Their summary is that AI amplifies whatever is already there, so a good pipeline gets faster and a weak one breaks more often.
What happens to the codebase?
This part has the longest tail, and the data is thinner than we'd like: one vendor's dataset, though a large one.
GitClear tracked code changes from 2023 into 2026. The share of changed lines that were moved (their proxy for refactoring) was 21% in 2022, 13% in 2023 and 3.8% so far in 2026. Copy-paste went the other way, from 9.4% to 15.7%.
Nothing fails on the day duplicated code is merged. It costs later, when a fix has to be made in five places and somebody finds four.
Security follows a similar line. In Veracode's tests the best model reached 68%, and the average has moved one point in a year. The models are strong on SQL injection (83% pass) and weak on cross-site scripting (15%) and log injection (12%). So the failures are not random. They're predictable enough to write checks for.
And developers know it. In Stack Overflow's 2025 survey only 33% said they trust the accuracy of AI output, and 66% named the same frustration: answers that are "almost right, but not quite". Usage kept climbing anyway, which tells you how useful the tools are even at that level of trust.
What should an engineering team do about it?
Four things, in the order we'd take them.
Measure delivery before anything else
Lead time from commit to production, time waiting for review, change failure rate, rework. If those haven't moved since the licences were bought, the tools are doing their job and the pipeline isn't.
Put AI where the queue is
Most teams bring AI in for writing code first, because that's where the demos are. The Faros numbers say the wait is in review and testing. First-pass review, test generation and release checks are duller uses. They're also the ones that shorten the queue.
Cap pull request size
Nobody decided that pull requests should grow 154% — it happened because producing code got cheap. A size limit is a blunt rule and it works, since a reviewer can still read what they're approving.
Budget for refactoring
When moved lines fall from 13% to 3.8% across the industry, cleanup has stopped happening as a side effect. Schedule it, and point the assistant at it. Merging duplicates is work these tools do well once someone asks.
One more thing applies to legacy products. An assistant is only as good as what it can read. A codebase with few tests, and its history kept in people's heads, gives it little to work with, which may be part of why METR's experts on million-line repositories gained the least. The cheapest preparation is the boring kind: tests around the parts that change most often, and a short written account of how each module is supposed to behave.
Common questions
Is AI making developers worse?
There is no solid evidence on skill loss yet. What has been measured is the code: more duplication, less refactoring, flat security scores. The worry is real, though. In Stack Overflow's 2026 survey, 17% of respondents gave losing skills as their reason for avoiding AI at work.
Does AI help less on legacy code?
From what's published, yes. METR's 19% slowdown came from mature repositories with over a million lines, and the 2026 rerun pointed the other way without statistical confidence. Expect smaller gains until the codebase has tests and documentation an assistant can use.
How much faster is realistic?
Per developer, Faros measured 21% more completed tasks. Per company, nothing measurable. A fair range for a team today runs from zero to about a fifth, and where you land depends on review and release more than on the model.
Should we wait for better models?
Veracode's security pass rate moved one point in a year while the models improved at almost everything else. Waiting fixes the model. The review queue and the missing tests will still be there.
