01Evaluating AI

How to evaluate a radiology-AI vendor

A framework to judge any imaging AI: independent evidence, the right metrics, real-world deployment, measurements, integration, and regulatory scope.

Medically reviewed by Dr Alexandre Parpaleix, MD-PhD, CEO and co-founder of Milvue

How should you judge a vendor's claimed performance?

The first question is not "what is the score?" but "who measured it?". A figure self-reported on a private test set is not an independent, published, peer-reviewed evaluation. The gap is documented: of 100 CE-marked radiology AI products surveyed, 64 had no peer-reviewed evidence of efficacy (van Leeuwen et al., European Radiology, 2021). To structure the purchase, the ECLAIR guidelines (Omoumi et al., European Radiology, 2021) offer a question-by-question framework. The right instinct: ask for an independent study, on a population close to yours. In fracture detection, the head-to-head evaluation of three commercial solutions in the emergency department places Milvue at the top on diagnostic accuracy (90.1% versus 88.8% and 71.0%): statistically tied with the second, well ahead of the third. Reading a ranking means reading all three figures together, never the first alone (Academic Radiology, 2023).

Should you look at sensitivity, specificity or accuracy?

All three, together, never one alone. An AI can post very high sensitivity by over-calling, at the cost of degraded specificity that buries the radiologist under alerts. The reverse misses lesions. Accuracy and the real operating point therefore matter as much as the headline number. Ask for the full matrix, not an isolated metric: that is exactly what the CLAIM checklist (Mongan, Moy, Kahn, Radiology: Artificial Intelligence, 2020) requires for a result to be interpretable. As a systematic second read, Milvue AI raises the radiologist's sensitivity to adult fracture from 0.81 to 0.98 (FDA reader study, aided vs unaided); and it is accuracy, not sensitivity alone, that distinguishes it in the published head-to-head studies.

Does passing a radiology exam prove an AI is clinic-ready?

No. Passing a certification exam is a benchmark; working a shift is another. A result is only validated in a real workflow: wired into the RIS/PACS, reviewed, signed, accountable. The question to ask a vendor is not "does your model pass this test?" but "where, how many, and for how long has your solution run in routine use?". As a reference point, Milvue produces 40 million analyses a year, across 25 countries, with 1M+ TechCare Report exams already processed. The evidence that counts is not measured on a closed test set: it is measured in production.

Are measurements computed by the AI, or simply copied?

A concrete, discriminating question. Many assistants copy a number already entered elsewhere (OCR or text extraction) and drop it back into the report. That is useful, but it is not a measurement. Ask whether the value is computed from the image. TechCare Metrics produces the measurement - angle, length, index - from the image data, then writes it into the report. The difference is not cosmetic: a computed value is traceable and reproducible, a copied value is only as reliable as its source.

Is the PACS integration real or just claimed?

An AI that never reaches the workflow is never used. A vendor's "RIS/PACS integration" slide sometimes hides a fragile connector, a manual export or a second screen. The right questions: native integration or via a third-party layer? results pushed into the existing worklist or viewed separately? how many sites in production on your specific PACS? Milvue is natively integrated into the RIS/PACS of 600+ sites; it is that integration, more than the model alone, that decides whether the promised time is genuinely returned to the radiologist.

What should you check on the regulatory and data side?

A final, often decisive filter. The CE marking (and its class) covers only a precise intended-use scope: one module may be a medical device, another not. A reporting assistant or a VLM that proposes a draft for review is not a medical device, and must not be sold as one. ECLAIR details this regulatory dimension and the data question. At Milvue, the diagnostic suite is a CE-marked medical device (Class IIa); TechCare Report, subscription software, is not one; and data is held in Europe under the GDPR. Once these six filters are passed, the choice makes itself - because in the end, our AI does not replace radiologists. It makes them irreplaceable.

The Milvue solution The Milvue Suite
Explore

See Milvue AI in real conditions.