The 19% slowdown study, revisited by the people who ran it (external link)
If you have argued about AI productivity on the internet in the last year, you have cited this study, and there is a decent chance you have cited it wrong.
The original: METR took sixteen experienced open-source developers, 246 real issues in codebases they already knew well, and randomised whether they could use AI tools. Result — they were 19% slower with the tools. The part that made it famous was the perception gap: they expected a 24% speedup going in, and after finishing slower, still believed they had been sped up by about 20%.
Go to that page today and there is a banner on it:
These results are out of date. We have released results that are current as of early 2026, in a continuation of this study.
METR published new data in February on late-2025 tools, with a much larger cohort, and the slowdown largely evaporates — the newer estimate is close enough to zero that the honest summary is "no clear effect either way", not a reversal into a speedup. Read their continuation rather than my paraphrase; the confidence intervals are the whole point and they do not survive being quoted.
Three things worth holding onto.
The headline number was always narrow. Experienced developers, on codebases they knew intimately, on early-2025 tooling. That is close to the worst case for AI assistance and it was never a claim about the general case. It got used as one anyway, by me among others.
The perception gap is the durable finding. Whether the tools make you 19% slower or 4% faster, the discovery that developers cannot tell which is happening to them survives every revision. That should terrify anyone running an engineering org on self-reported velocity, and self-reported velocity is what almost everyone is running on.
METR did the right thing. They put a warning label on their own most-cited result. That is rarer than it should be, and it is why they remain worth reading when the vendor benchmarks are not.