Faster, more confident, and no more accurate (external link)
Five readers. 1200 consecutive patients, 1861 chest radiographs. Each case read five times — once unaided, then once with each of four commercially available AI tools — with a 14-day washout between sessions and the case order randomised. Prospective, single centre, published in Academic Radiology this month. I have read the structured abstract, not the full text, so everything below is bounded by that.
AI assistance did not improve diagnostic accuracy. For pleural effusions and pulmonary nodules, accuracy decreased in several reader-tool pairings, through increased false positives.
And:
- Three of five residents read faster, by a median of 6 to 17 seconds per case (p≤0.031).
- Four of five readers reported higher diagnostic confidence.
- Senior consultations dropped in selected pairings. In one, CT recommendations did too.
Read those two blocks together, because separately they are two different press releases. Accuracy flat or worse. Confidence up. Escalation down. That is not a mixed result, it is a specific and well-named one, and the authors name it: automation bias. The tool made readers feel more certain and ask for help less often while making them no better, and on two findings, worse.
The escalation number is the one that would keep me up. Faster reads I can argue about. Fewer senior consultations means the second look — the actual safety net, the thing that catches the resident at 2am — got quieter, and it got quieter for reasons that have nothing to do with the cases being easier. Nobody is going to see that in a dashboard. It shows up as a consult rate that drifted down over a quarter, which every department would read as the residents getting stronger.
The false positives are my problem specifically. An increase in false positives on nodules is not an accuracy statistic when it reaches my department, it is volume: flags on a worklist somebody has to dispose of, each one cheap on its own and none of them free. We have been here before — the AUC holds and the specificity falls off a cliff, and the part that lands on the working day is always the specificity.
Caveats, and they are real ones. Five readers with one to six years of experience is a small and junior sample; automation bias is generally worse in less experienced readers, so this is close to the population where you would expect the largest effect. Single centre. The reference standard was the finalised clinical report supplemented by CT where available, which is the practical choice and is not ground truth. And four tools tested is four tools, not the category — the abstract reports reader-tool pairings, which tells you the effects were not uniform, and I would want the per-tool breakdown before letting anyone say "AI does not help on chest X-ray."
The authors' conclusion is that deployment requires careful local adaptation and that clinical trials are needed for patient outcomes. Both true. But "careful local adaptation" is doing an enormous amount of work in that sentence, and I am the person it lands on. Nobody sells adaptation. What arrives is a model, a threshold nobody will tell you how they picked, and a go-live date.
Here is what I would want, and it is smaller than a trial: the flag rate per tool per finding, tracked from day one, and the consult rate alongside it. Six to seventeen seconds a case is a genuine efficiency argument at our volume. It is just not free, and this study is the first I have read that priced the other side of it in the same room, on the same cases, with the same readers.