01Generative AI

Radiology VLM: the AI that drafts the report for review

What a radiology-native vision-language model changes: from flagging a finding to drafting a report the radiologist reviews. And why passing an exam is not enough.

Medically reviewed by Dr Alexandre Parpaleix, MD-PhD, CEO and co-founder of Milvue

What is a radiology VLM, and what does it change?

A VLM (vision-language model) reads the image and produces text. This is the 2026 shift: from AI that flags a finding to AI that proposes a draft report. The practical difference is clear: detection flags a finding; the VLM writes a first draft that the radiologist corrects and signs off. The judgment stays the physician's; what changes is the starting point: a page already sketched rather than a blank one.

Where does Milvue stand on the VLM?

Milvue is building a radiology-native VLM. Not a general-purpose model tweaked at the edges: a model built on radiology from the ground up. This direction was cited publicly in Microsoft's June 2026 Community Hub blog, which describes extending image-based reporting to musculoskeletal pathologies. Chest is the most heavily worked area; Milvue covers both sides of the worklist: the highest fracture-detection accuracy of the only published independent comparison, and chest expertise. The VLM shift extends that advantage: from flagging the finding to the sentence that describes it.

Is passing a certification exam the same as doing a radiologist's job?

In 2026, several models claim to pass a radiology exam. It is a feat, and it is a benchmark. But passing an exam is not working a shift: a report is only validated in a real workflow, wired into the PACS, reviewed, signed, accountable. Milvue produces 40 million analyses a year, across 25 countries, natively integrated into the RIS/PACS at 600+ sites, with 1M+ TechCare Report exams already processed. The evidence that counts is not self-reported on a test set: it is independent and published, and it is measured in routine use.

How is Milvue's VLM different from agents that draft a report?

Three differences set the product apart. First, measurements are produced, not copied: TechCare Metrics computes the angle or length from the image and writes it into the report, where an agent often just extracts a number already typed. Second, it speaks 17 languages and keeps European data residency and compliance, against models designed in English. Third, the full suite in the flow: triage and diagnostic support, measurements, then report, wired together, not one isolated module. The difference does not show in a demo: it reads in the signed report.

Is the VLM safe, and what does it not do (yet)?

Let's be precise about scope. The VLM proposes a draft for review: there is never fully autonomous report generation; the radiologist validates every line. Like the foundation models on the market, this VLM is not a medical device; certified detection and report drafting remain two separate modules with separate regulatory status. A VLM can phrase an inaccurate sentence with confidence: that is exactly why human review is not an option but the product's very architecture.

What clinical evidence backs the VLM, and when?

Milvue's discipline is constant: no figure before publication. Clinical evaluations of the VLM are under way, on chest and musculoskeletal imaging. We will share results once they are established and validated at our deployment sites, as we did for fracture detection and measurements. Meanwhile, the rule that guides this work does not change: Our AI does not replace radiologists. It makes them irreplaceable.

The Milvue solution TechCare Report
Explore

See Milvue AI in real conditions.