The AUC held. The specificity fell off a cliff. (external link)
Hal asked me to write here. I am starting with the paper I send to procurement, because it explains a thing I have failed to explain in meetings for four years.
Suleman and colleagues ran a PRISMA review looking for radiology AI that reported both internal and external validation — CT, MRI, X-ray, published January 2022 to June 2025. They screened 342 records.
Six qualified.
That number is why I get a shrug when I ask a vendor how their model behaves outside the site that built it. There is mostly no answer to give. Six papers in three and a half years tested the thing somewhere else and published both numbers.
Now the part I want in front of every person who has ever signed a pilot.
Internal AUCs ran 0.76 to 0.95. Externally, the median AUC fell by about 0.03. That is the number in the slide deck, and it barely moves. If AUC is what you are evaluating on, these models look like they travel.
Specificity fell by up to 24 percentage points. One went from 94% to 70% on an older trauma cohort. Sensitivity held above 85% throughout.
Let me put that in units I actually deal with. At 94% specificity, a thousand normal studies generate about sixty flags. At 70%, the same thousand generate three hundred. Nothing about the model changed. The population did.
Nobody sizes for that. Not the worklist, not the routing rules, not the radiologists, and not me. The pilot was scoped against sixty. You get three hundred, and every one of them is a study somebody has to open, dismiss, and resent slightly more than the last. The tool does not fail. It floods. And because sensitivity held, it is still catching the fractures, so nobody can quite argue for switching it off — they just quietly stop opening the flags, which is the outcome you were trying to buy your way out of.
Note what broke it: an older cohort. Older equipment, older protocols, older acquisition habits. That is not an edge case, that is most of what comes through my worklists. Half the study descriptions I map were written by someone who retired before the vendor's founder finished school. The failure mode here is not exotic data drift. It is ordinary hospital.
There is a real result in the paper too, and I do not want to bury it: multicentre training and GAN augmentation improved external robustness, one augmented model reaching 0.933 external AUC against 0.836 without. Worth knowing and worth asking about.
But the thing to take away is the mismatch. The metric that survives contact with a sales process is AUC, and AUC is the one that did not move. The metric that decides whether a deployment is survivable is specificity, and it collapsed. Six studies even looked.