Radiology is where clinical AI is most real. It is the specialty with the most authorized devices, the most deployment, and — unusually for this field — one large randomized trial that measured what happens when the software is switched on in a live screening programme. It is also where the gap between what is cleared and what is evidenced is easiest to see, because both maps exist and they do not overlap neatly. This guide reads each map plainly, with every figure tied to a primary source: the FDA's own device record, two peer-reviewed audits of it, and the three dated readouts of the MASAI trial. As of July 2026.
What the FDA record shows
The closest thing the field has to a census is the FDA's AI-Enabled Medical Device List, a public register of authorized artificial-intelligence-enabled devices 1. Read by specialty, it tells one dominant story: radiology owns the list. A peer-reviewed audit of the machine-learning devices authorized in 2024 found that radiology accounted for 74.4% (125 of 168) — with cardiovascular and neurology far behind 2.
| Device panel | 2024 authorizations | Share | Source |
|---|---|---|---|
| Radiology | 125 | 74.4% | 2 |
| Cardiovascular | 11 | 6.5% | 2 |
| Neurology | 10 | 6.0% | 2 |
| Anesthesiology | 5 | 3.0% | 2 |
| Gastroenterology–Urology | 5 | 3.0% | 2 |
This is no single-year quirk. Radiology has been the largest panel every year the list has been tracked, for structural reasons: imaging produces standardized, digital, labelled data at scale, and the review pathway for image-analysis software is well worn. When people say "healthcare AI," the authorized reality is still, overwhelmingly, software that reads a scan. For the full growth curve and specialty split, see our tracker of FDA-cleared AI devices by year and specialty.
The strongest trial evidence: AI-supported screening
Counting devices tells you what reached the market. It says nothing about what happens when clinicians use them. For that, radiology has something most specialties lack: a large randomized trial. The Mammography Screening with Artificial Intelligence trial (MASAI) in Sweden randomized women in a population-based programme to AI-supported screen reading or to standard double reading by two radiologists, and it has now reported in three stages.
The interim safety analysis covered 80,033 women. AI-supported reading detected cancer at 6.1 per 1000 versus 5.1 per 1000 for standard double reading, with recall rates of 2.2% versus 2.0% and an identical 1.5% false-positive rate — while cutting the screen-reading workload by 44.3% 4. The headline there is the workload figure: the same detection, at roughly half the reading labour.
A later screening-performance analysis of 105,934 women sharpened the detection picture. AI-supported reading found 6.4 versus 5.0 cancers per 1000, a ratio of 1.29 (a 29% relative increase, p=0.0021), and did so with a higher positive predictive value (ratio 1.19) and no significant rise in recalls (recall ratio 1.08) 5. More cancers, without a proportionate flood of false alarms — the combination screening programmes want.
The third readout answers the question a screening programme actually cares about: does AI let real cancers slip through to become symptomatic interval cancers between screens? Here the trial reported 1.55 versus 1.76 interval cancers per 1000 for AI-supported versus standard reading, with specificity of 98.5% in both arms — non-inferior on the safety endpoint 6.
| MASAI readout | Women | AI-supported | Standard reading | Source |
|---|---|---|---|---|
| Interim safety (2023) | 80,033 | 6.1 / 1000 detected; workload −44.3% | 5.1 / 1000 detected | 4 |
| Screening performance (2025) | 105,934 | 6.4 / 1000 detected; PPV ratio 1.19 | 5.0 / 1000 detected | 5 |
| Interval cancer (2026) | full cohort | 1.55 / 1000 interval; spec 98.5% | 1.76 / 1000 interval; spec 98.5% | 6 |
Read the table with its design in mind. MASAI is one AI system, in one national programme, with double reading as the comparator — the standard in much of Europe, not in the United States, where single reading is common. The result is strong evidence that AI support can raise detection and cut workload without harming safety in that configuration. It does not promise that any AI system will behave the same against any comparator, and it does not yet report a mortality endpoint. The discipline for reading any such number lives in our guide on how to read an AI validation study.
Beyond screening: triage, detection, and image generation
Screening is the best-evidenced use, but it is not the most common one on the device list. A peer-reviewed taxonomy classified what 1,016 FDA-authorized AI devices actually do, and the functions cluster into a few families 3:
- Image analysis — measuring, segmenting, or characterizing a finding — is the single most common function, though its relative share has begun to decline as other uses appear.
- Triage and notification — flagging a study as likely positive so it moves up the worklist. This is the family behind time-critical uses such as suspected large-vessel-occlusion stroke or intracranial haemorrhage, where the value is speed of human review, not autonomy.
- Image generation — more than 100 authorized devices generate data, for example synthesizing or enhancing an image, rather than only measuring it 3.
Two facts from the taxonomy are worth holding onto. First, none of the authorized devices yet use large language models 3: the cleared, clinic-grade reality is quantitative imaging software, not the chatbots that dominate the headlines. Second, most of these tools are assistive — they flag, measure, or prioritize for a radiologist who remains the decision-maker. That human-in-the-loop design is both a safety property and a limit on how much a triage tool can be credited for outcomes on its own.
The transparency gap
Approval volume has outrun disclosure. In the same 2024 audit, only 29.2% of authorizations reported both sensitivity and specificity, and demographic detail on the study population was rarer still 2. A device can be authorized without a full performance-and-equity profile being published for a buyer to read.
That matters most where it is least visible: a tool validated on one scanner make, one institution's protocols, or one demographic mix can behave differently on yours. The single most important question to ask a vendor is whether the tool was externally validated — tested on data from a different site, scanner, or era than it was built on — and at what operating point, because a summary figure like AUROC hides the sensitivity and specificity you will actually run at. The Good Machine Learning Practice principles, issued jointly by the FDA, Health Canada, and the UK MHRA in October 2021, set the development expectations that frame these questions 7.
Why the workload finding travels — and where it stops
The 44.3% reduction in screen-reading labour 4 is the result most likely to shape budgets, so it is worth reading precisely. MASAI's comparator was double reading — two radiologists on every screen, the standard across much of Europe — and the AI arm replaced one of those reads for the large majority of straightforward cases while escalating the uncertain ones. That is where the workload saving comes from, and it is why the figure does not transfer unchanged to a system that already reads each screen once, as many US programmes do. The saving is real; its size is a function of the workflow it is measured against. The same caution applies to detection: the 29% relative gain 5 is a gain over double reading, so a programme with a different baseline should expect a different delta. Read the effect and the comparator together, always.
Questions that decide value in your setting
A device count and a trial result are inputs to a decision, rather than the decision itself. Whether a radiology AI tool earns its place in your reading room turns on a handful of questions the marketing rarely answers, and each maps to something in the evidence above.
- What is the cleared indication, exactly? An authorization covers a specific intended use — a body region, a modality, a finding, a role such as triage versus diagnosis. A tool cleared to flag suspected large-vessel-occlusion stroke has different obligations from one cleared to characterize a lesion. Read the indication, and read what sits outside it.
- Was it externally validated on data like yours? Because only a minority of authorizations publish a full performance profile 2, ask directly whether the tool was tested on your scanner makes, protocols, and patient mix — the external-validation question that separates a portable result from an overfit one.
- At what operating point, and with what error rates? A summary curve hides the threshold you will actually run at. Ask for the sensitivity and specificity at the deployed operating point, and for the expected alert or recall volume at your throughput — the number that governs bedside and reading-room behaviour.
- Does it change a decision or only a workflow? MASAI moved detection and workload because it altered the read itself 45. A triage tool that reorders a worklist may speed care without changing what is ultimately found — a real benefit, measured differently and credited more carefully.
- Who monitors it after go-live? Performance drifts as scanners, populations, and protocols change, so a deployment plan needs post-market monitoring and a named owner — the spirit of the Good Machine Learning Practice principles 7.
None of these questions has a universal answer, which is the point: the evidence tells you what is possible, and your setting tells you what is likely. The strongest single move an organization can make is the one MASAI models at national scale and any department can run in miniature — a local, blinded evaluation on its own images before committing. A vendor benchmark shows a tool worked somewhere; a local pilot shows whether it works here, on your equipment and your case mix, at the operating point you would actually deploy.
How to read this guide
Four cautions travel with everything above. First, cleared is not the same as proven: an authorization is a market-entry signal for an intended use, not evidence of benefit in your population — and the transparency gap means the published profile is often partial. Second, one trial is one setting: MASAI is the strongest imaging evidence in the field, and it is still a single AI system in a double-reading programme rather than a universal result. Third, triage is not diagnosis: a tool that flags a study faster has changed a workflow, not necessarily an outcome, and it depends on the radiologist who reviews the flag. Fourth, status is perishable — the device list grows monthly, new trial readouts land, and each figure here is dated to when it was pulled.
Because clearance status, reimbursement, and liability turn on exactly these distinctions, confirm the current regulatory status and evidence expectations for any specific tool with your compliance or regulatory counsel before acting; clearance is not the same as validation quality, and neither is the same as coverage. We revisit this guide on a 180-day cycle and whenever the FDA updates its list or a new imaging trial reports. For the neighboring specialties, see our guides on AI in pathology and AI in cardiology.
Sources and method
This guide draws on the FDA's AI-Enabled Medical Device List 1; two peer-reviewed audits of that record — a 2024 authorization study 2 and a taxonomy of 1,016 authorizations 3; the three dated readouts of the MASAI randomized trial 456; and the tri-regulator Good Machine Learning Practice principles 7. Every figure is tied to the primary source cited beside it, and each perishable number carries its date. Statuses and counts are current as of July 2026.