Specialties

AI in radiology: what is cleared and evidenced in 2026

Radiology is where clinical AI is most deployed and most authorized — but the cleared reality and the trial evidence are two different maps. A dated read of the FDA device record and the strongest randomized trial in imaging, each figure tied to a primary source. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Radiology is the center of gravity for cleared clinical AI: it accounted for 74.4% (125 of 168) of the machine-learning devices the FDA authorized in 2024, far ahead of cardiovascular (6.5%) and neurology (6.0%).
  • The strongest trial evidence is the MASAI randomized trial in Sweden. Its interim readout cut radiologist screen-reading workload by 44.3% while detecting cancer at a similar-or-higher rate; a later analysis of 105,934 women found 29% more cancers detected with AI support and a higher positive predictive value.
  • MASAI's safety endpoint held: interval cancers were 1.55 vs 1.76 per 1000 for AI-supported vs standard reading, with 98.5% specificity in both arms — non-inferior on the number that matters for a screening programme.
  • Most authorized radiology AI does one of three things: analyze an image, triage/flag a study, or generate image data. None of the FDA-authorized devices yet use large language models.
  • A clearance is a market-entry signal rather than a full performance profile: only 29.2% of 2024 authorizations reported both sensitivity and specificity. Confirm the evidence for any specific tool, in your own population, before you rely on it.

Radiology is where clinical AI is most real. It is the specialty with the most authorized devices, the most deployment, and — unusually for this field — one large randomized trial that measured what happens when the software is switched on in a live screening programme. It is also where the gap between what is cleared and what is evidenced is easiest to see, because both maps exist and they do not overlap neatly. This guide reads each map plainly, with every figure tied to a primary source: the FDA's own device record, two peer-reviewed audits of it, and the three dated readouts of the MASAI trial. As of July 2026.

What the FDA record shows

The closest thing the field has to a census is the FDA's AI-Enabled Medical Device List, a public register of authorized artificial-intelligence-enabled devices 1. Read by specialty, it tells one dominant story: radiology owns the list. A peer-reviewed audit of the machine-learning devices authorized in 2024 found that radiology accounted for 74.4% (125 of 168) — with cardiovascular and neurology far behind 2.

Device panel2024 authorizationsShareSource
Radiology12574.4%2
Cardiovascular116.5%2
Neurology106.0%2
Anesthesiology53.0%2
Gastroenterology–Urology53.0%2

This is no single-year quirk. Radiology has been the largest panel every year the list has been tracked, for structural reasons: imaging produces standardized, digital, labelled data at scale, and the review pathway for image-analysis software is well worn. When people say "healthcare AI," the authorized reality is still, overwhelmingly, software that reads a scan. For the full growth curve and specialty split, see our tracker of FDA-cleared AI devices by year and specialty.

The strongest trial evidence: AI-supported screening

Counting devices tells you what reached the market. It says nothing about what happens when clinicians use them. For that, radiology has something most specialties lack: a large randomized trial. The Mammography Screening with Artificial Intelligence trial (MASAI) in Sweden randomized women in a population-based programme to AI-supported screen reading or to standard double reading by two radiologists, and it has now reported in three stages.

The interim safety analysis covered 80,033 women. AI-supported reading detected cancer at 6.1 per 1000 versus 5.1 per 1000 for standard double reading, with recall rates of 2.2% versus 2.0% and an identical 1.5% false-positive rate — while cutting the screen-reading workload by 44.3% 4. The headline there is the workload figure: the same detection, at roughly half the reading labour.

A later screening-performance analysis of 105,934 women sharpened the detection picture. AI-supported reading found 6.4 versus 5.0 cancers per 1000, a ratio of 1.29 (a 29% relative increase, p=0.0021), and did so with a higher positive predictive value (ratio 1.19) and no significant rise in recalls (recall ratio 1.08) 5. More cancers, without a proportionate flood of false alarms — the combination screening programmes want.

The third readout answers the question a screening programme actually cares about: does AI let real cancers slip through to become symptomatic interval cancers between screens? Here the trial reported 1.55 versus 1.76 interval cancers per 1000 for AI-supported versus standard reading, with specificity of 98.5% in both arms — non-inferior on the safety endpoint 6.

MASAI readoutWomenAI-supportedStandard readingSource
Interim safety (2023)80,0336.1 / 1000 detected; workload −44.3%5.1 / 1000 detected4
Screening performance (2025)105,9346.4 / 1000 detected; PPV ratio 1.195.0 / 1000 detected5
Interval cancer (2026)full cohort1.55 / 1000 interval; spec 98.5%1.76 / 1000 interval; spec 98.5%6

Read the table with its design in mind. MASAI is one AI system, in one national programme, with double reading as the comparator — the standard in much of Europe, not in the United States, where single reading is common. The result is strong evidence that AI support can raise detection and cut workload without harming safety in that configuration. It does not promise that any AI system will behave the same against any comparator, and it does not yet report a mortality endpoint. The discipline for reading any such number lives in our guide on how to read an AI validation study.

Beyond screening: triage, detection, and image generation

Screening is the best-evidenced use, but it is not the most common one on the device list. A peer-reviewed taxonomy classified what 1,016 FDA-authorized AI devices actually do, and the functions cluster into a few families 3:

  • Image analysis — measuring, segmenting, or characterizing a finding — is the single most common function, though its relative share has begun to decline as other uses appear.
  • Triage and notification — flagging a study as likely positive so it moves up the worklist. This is the family behind time-critical uses such as suspected large-vessel-occlusion stroke or intracranial haemorrhage, where the value is speed of human review, not autonomy.
  • Image generation — more than 100 authorized devices generate data, for example synthesizing or enhancing an image, rather than only measuring it 3.

Two facts from the taxonomy are worth holding onto. First, none of the authorized devices yet use large language models 3: the cleared, clinic-grade reality is quantitative imaging software, not the chatbots that dominate the headlines. Second, most of these tools are assistive — they flag, measure, or prioritize for a radiologist who remains the decision-maker. That human-in-the-loop design is both a safety property and a limit on how much a triage tool can be credited for outcomes on its own.

The transparency gap

Approval volume has outrun disclosure. In the same 2024 audit, only 29.2% of authorizations reported both sensitivity and specificity, and demographic detail on the study population was rarer still 2. A device can be authorized without a full performance-and-equity profile being published for a buyer to read.

That matters most where it is least visible: a tool validated on one scanner make, one institution's protocols, or one demographic mix can behave differently on yours. The single most important question to ask a vendor is whether the tool was externally validated — tested on data from a different site, scanner, or era than it was built on — and at what operating point, because a summary figure like AUROC hides the sensitivity and specificity you will actually run at. The Good Machine Learning Practice principles, issued jointly by the FDA, Health Canada, and the UK MHRA in October 2021, set the development expectations that frame these questions 7.

Why the workload finding travels — and where it stops

The 44.3% reduction in screen-reading labour 4 is the result most likely to shape budgets, so it is worth reading precisely. MASAI's comparator was double reading — two radiologists on every screen, the standard across much of Europe — and the AI arm replaced one of those reads for the large majority of straightforward cases while escalating the uncertain ones. That is where the workload saving comes from, and it is why the figure does not transfer unchanged to a system that already reads each screen once, as many US programmes do. The saving is real; its size is a function of the workflow it is measured against. The same caution applies to detection: the 29% relative gain 5 is a gain over double reading, so a programme with a different baseline should expect a different delta. Read the effect and the comparator together, always.

Questions that decide value in your setting

A device count and a trial result are inputs to a decision, rather than the decision itself. Whether a radiology AI tool earns its place in your reading room turns on a handful of questions the marketing rarely answers, and each maps to something in the evidence above.

  • What is the cleared indication, exactly? An authorization covers a specific intended use — a body region, a modality, a finding, a role such as triage versus diagnosis. A tool cleared to flag suspected large-vessel-occlusion stroke has different obligations from one cleared to characterize a lesion. Read the indication, and read what sits outside it.
  • Was it externally validated on data like yours? Because only a minority of authorizations publish a full performance profile 2, ask directly whether the tool was tested on your scanner makes, protocols, and patient mix — the external-validation question that separates a portable result from an overfit one.
  • At what operating point, and with what error rates? A summary curve hides the threshold you will actually run at. Ask for the sensitivity and specificity at the deployed operating point, and for the expected alert or recall volume at your throughput — the number that governs bedside and reading-room behaviour.
  • Does it change a decision or only a workflow? MASAI moved detection and workload because it altered the read itself 45. A triage tool that reorders a worklist may speed care without changing what is ultimately found — a real benefit, measured differently and credited more carefully.
  • Who monitors it after go-live? Performance drifts as scanners, populations, and protocols change, so a deployment plan needs post-market monitoring and a named owner — the spirit of the Good Machine Learning Practice principles 7.

None of these questions has a universal answer, which is the point: the evidence tells you what is possible, and your setting tells you what is likely. The strongest single move an organization can make is the one MASAI models at national scale and any department can run in miniature — a local, blinded evaluation on its own images before committing. A vendor benchmark shows a tool worked somewhere; a local pilot shows whether it works here, on your equipment and your case mix, at the operating point you would actually deploy.

How to read this guide

Four cautions travel with everything above. First, cleared is not the same as proven: an authorization is a market-entry signal for an intended use, not evidence of benefit in your population — and the transparency gap means the published profile is often partial. Second, one trial is one setting: MASAI is the strongest imaging evidence in the field, and it is still a single AI system in a double-reading programme rather than a universal result. Third, triage is not diagnosis: a tool that flags a study faster has changed a workflow, not necessarily an outcome, and it depends on the radiologist who reviews the flag. Fourth, status is perishable — the device list grows monthly, new trial readouts land, and each figure here is dated to when it was pulled.

Because clearance status, reimbursement, and liability turn on exactly these distinctions, confirm the current regulatory status and evidence expectations for any specific tool with your compliance or regulatory counsel before acting; clearance is not the same as validation quality, and neither is the same as coverage. We revisit this guide on a 180-day cycle and whenever the FDA updates its list or a new imaging trial reports. For the neighboring specialties, see our guides on AI in pathology and AI in cardiology.

Sources and method

This guide draws on the FDA's AI-Enabled Medical Device List 1; two peer-reviewed audits of that record — a 2024 authorization study 2 and a taxonomy of 1,016 authorizations 3; the three dated readouts of the MASAI randomized trial 456; and the tri-regulator Good Machine Learning Practice principles 7. Every figure is tied to the primary source cited beside it, and each perishable number carries its date. Statuses and counts are current as of July 2026.

Questions & answers

  • Is AI in radiology FDA-approved?

    Many radiology AI tools are FDA-authorized, and radiology is by far the largest specialty on the FDA's AI-Enabled Medical Device List — it made up 74.4% of the machine-learning devices authorized in 2024. "Authorized" means a device cleared the FDA's review threshold for its intended use, not that it improves outcomes in every setting. Confirm the specific tool, its cleared indication, and its evidence before relying on it.

  • Does AI actually improve mammography screening?

    The strongest evidence is the MASAI randomized trial in Sweden. In an analysis of 105,934 women, AI-supported reading detected about 29% more cancers than standard double reading (6.4 vs 5.0 per 1000) without a significant rise in recalls, and it cut radiologist screen-reading workload by roughly 44%. Its safety endpoint — interval cancers — was non-inferior. One trial in one programme is strong evidence for that setting rather than a guarantee for every population or AI system.

  • Do FDA-authorized radiology AI tools use ChatGPT-style models?

    Not as of the most recent peer-reviewed taxonomy of 1,016 FDA authorizations, which found that none of the authorized devices yet rely on large language models. Most perform quantitative image analysis, triage a study, or generate image data. The cleared reality lags the chatbot conversation by a wide margin.

Sources

  1. US Food and Drug Administration. Artificial Intelligence and Machine Learning in Software as a Medical Device — the FDA AI/ML programme, which maintains the public AI-Enabled Medical Device List. Accessed July 2026. www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device
  2. Almarie B, Gonzalez-Gonzalez LF, dos Santos Barbosa LA, et al. Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration in 2024: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting. Biomedicines. 2025;13(12):3005. doi.org/10.3390/biomedicines13123005
  3. Singh R, Bapna M, Diab AR, Ruiz ES, Lotter W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. npj Digital Medicine. 2025;8:388. doi.org/10.1038/s41746-025-01800-1
  4. Lång K, Josefsson V, Larsson A-M, et al. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. Lancet Oncol. 2023;24(8):936-944. doi.org/10.1016/S1470-2045(23)00298-X
  5. Hernström V, Josefsson V, Sartor H, et al. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): a randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study. Lancet Digit Health. 2025. doi.org/10.1016/S2589-7500(24)00267-X
  6. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. The Lancet. 2026. doi.org/10.1016/S0140-6736(25)02464-X
  7. US FDA, Health Canada, and UK MHRA. Good Machine Learning Practice for Medical Device Development: Guiding Principles. October 2021. www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles