Specialties

AI in oncology: what the evidence actually shows

A guide to where artificial intelligence has earned its place across the cancer-care pathway — screening, detection, pathology — with the randomized evidence separated from the retrospective reader studies that fill vendor decks, each figure tied to its primary source. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The strongest oncology evidence is in breast screening: a randomized trial cut radiologist screen-reading workload by 44% while detecting more cancers at a similar recall rate, and its full results held up on interval cancer.
  • A real-world implementation across 461,818 screening cases found AI-supported double reading raised the cancer detection rate by 17.6% without a higher recall rate — a large but non-randomized confirmation.
  • For colonoscopy, pooled randomized evidence shows computer-aided detection raises adenoma detection (44.7% vs 36.7%), but it also increases removal of harmless polyps and does not find more advanced adenomas.
  • Most headline oncology-AI results — lung-nodule reading, prostate grading, breast detection — come from retrospective reader studies, not deployment; one famous breast-AI paper could not be independently reproduced because its code and data were withheld.
  • Across 1,016 FDA authorizations the device landscape is imaging-dominated and none use generative models, so oncology AI in practice is detection and triage with a clinician in the loop — not autonomous treatment. As of July 2026.

Oncology is where clinical AI has some of its strongest evidence and some of its most overstated claims, often about the same task. A mammography reader that saves radiologist time in a randomized trial and a lung-nodule model that "beats radiologists" in a retrospective slideshow are graded on very different scales, and the difference decides whether a result will survive contact with real patients. This guide walks the cancer-care pathway — screening, detection, pathology, and the device landscape underneath — and sorts what has randomized evidence from what has only a promising reader study. Every figure is tied to its primary source. As of July 2026.

Where the evidence actually sits

The honest map of oncology AI is less a list of capabilities than a list of evidence tiers. A tool can be technically impressive and clinically unproven at the same time, and the pathway below is arranged by how much a claim has been tested, not by how advanced the model is.

Pathway stageBest evidence tierHeadline resultThe caveat to hold
Breast screeningRandomized trial 12~44% less reading workload, more cancers detectedDetection and workload, not yet mortality
Breast screening (real world)Large implementation study 3+17.6% detection rate, non-inferior recallNon-randomized; site and reader selection
ColonoscopyMeta-analysis of randomized trials 56Adenoma detection 44.7% vs 36.7%More harmless-polyp removal; no advanced-adenoma gain
Lung-nodule readingRetrospective reader study 4AUC 94.4%; beat six readers with no prior scanNo deployment yet; single validation era
Prostate gradingDiagnostic study 7Grading comparable to pathologistsLaboratory setting, not live sign-out
Device landscapeAuthorization taxonomy 10Imaging-dominated; no generative modelsDetection and triage, not autonomous care

Read down that "caveat" column before any capability column. It is the reason two tools that both "work" can deserve very different amounts of trust.

Screening: the strongest evidence in the field

Breast screening is the clearest case where AI has been tested the hard way. In a randomized, controlled trial, women were assigned to AI-supported single reading or to standard double reading by two radiologists. AI-supported reading detected 244 screen-detected cancers versus 203 with double reading — a higher cancer detection rate at a similar recall rate — while reducing the screen-reading workload by 44.3% 1. That workload figure matters as much as the detection figure: in systems short of radiologists, halving the reading burden without losing cancers is itself the outcome.

The interim publication was a safety analysis, and the obvious question was whether the early detection signal would hold on the endpoint that matters more — interval cancers, the ones that surface symptomatically between screening rounds. The full results answered it: AI-supported screening had higher sensitivity and a non-inferior interval-cancer rate compared with double reading 2. A randomized design that survives its own harder endpoint is the strongest thing anyone can say about a clinical AI tool today.

Alongside the trial sits a different kind of evidence: real-world implementation. Across 461,818 screening cases, AI-supported double reading achieved a breast cancer detection rate of 6.7 per 1,000 versus 5.7 per 1,00017.6% higher — with a recall rate that was non-inferior to standard reading 3. That is a large number of women and a reassuring direction, but it is an observational implementation rather than a randomized comparison, so reader and site selection can flatter the result. Trial and implementation point the same way here, which is exactly the alignment you want before believing a screening claim.

Breast-screening studyDesignDetection signalRecall / workload
MASAI 12Randomized trialMore cancers; higher sensitivity44% less reading workload; non-inferior interval cancer
PRAIM 3Real-world implementation+17.6% detection rateNon-inferior recall rate

One caution travels with both: they measure detection, workload, and interval cancer, rather than cancer mortality. The pathway from "more cancers found earlier" to "fewer women die" is long and has surprised the screening field before — finding more cancer can mean catching lethal disease sooner, and it can also mean over-detecting indolent disease that would never have harmed the patient. The two look identical in a detection-rate figure and diverge only in long-term follow-up. So the right claim is "more detection at lower reading cost," and no more, until mortality and over-diagnosis data arrive.

A second caution is generalizability. Both studies ran in European screening programs with their own equipment, populations, and reading standards 13; a model that performs this well there can degrade on a different scanner, a different density distribution, or an under-represented group, which is why the screening result belongs to the systems that produced it until re-tested in yours. That is the practical face of external validation: a strong screening number is a reason to pilot locally, rather than a license to deploy sight-unseen.

Detection and diagnosis: promising, mostly retrospective

Step off screening and the evidence tier drops. The famous imaging results are overwhelmingly retrospective reader studies — a model reads stored scans and its output is compared with clinicians after the fact. These are legitimate early signals, and they are routinely presented as if they were deployment results, which they are not.

Lung-nodule reading is the archetype. A deep-learning model reached an AUC of 94.4% on low-dose CT screening cases and, when no prior scan was available for comparison, outperformed all six radiologists in the study with 11% fewer false positives and 5% fewer false negatives 4. It is a genuinely strong result — and it is a retrospective evaluation on curated cases, with the "no prior imaging" condition doing quiet work, because in practice radiologists lean heavily on comparison with earlier scans. Read the AUROC alongside the operating point the model would actually run at, and treat the reader-study framing as the ceiling rather than the expectation.

Colonoscopy is the counter-example that shows what happens when the same category is tested with randomized trials. Pooled across randomized trials, computer-aided detection raised the adenoma detection rate to 44.7% from 36.7% — a real, replicated gain in finding pre-cancerous polyps 5, consistent with individual trials showing higher detection without lengthening the procedure 6. But the same meta-analysis carries the discipline vendors omit: computer-aided detection also increased removal of non-neoplastic polyps and did not increase detection of advanced adenomas 5. More flags is not the same as more of the cancers that matter, and the extra harmless-polyp removals carry their own small costs. This is the mature version of an AI-detection story: it works, and here is precisely where it does not.

Pathology: grading at the microscope

Digital pathology is the quietest oncology-AI success. Grading a prostate biopsy — assigning a Gleason pattern — is high-stakes and famously subject to disagreement between pathologists, which makes it a natural target. A deep-learning system developed on 5,759 biopsies from 1,243 patients achieved grading comparable to pathologists 7. The value here is less about beating an expert and more about consistency: a tireless second reader that flags discordant cases and standardises grade where humans drift.

The caveat is the setting. This is a diagnostic study on assembled slides, rather than a live sign-out under time pressure with the full clinical context a pathologist carries. A tool that grades well in the study still has to prove it in the workflow, which is why deployment belongs behind a human-in-the-loop review rather than in front of it.

The device landscape: imaging, and a human in the loop

It helps to see the whole board. An independent taxonomy of 1,016 FDA authorizations found the AI device landscape dominated by imaging tasks, with none based on generative large language models 10. In oncology that shape is decisive: authorized AI clusters in radiology and pathology detection and characterisation, and the tools are regulated as software as a medical device that supports a clinician who reviews and decides. The fuller counts by year and specialty sit in our FDA-cleared AI devices tracker.

The practical reading: the oncology AI you can actually buy and deploy today helps find and measure things, and a clinician still owns the diagnosis and the treatment plan. Claims of autonomous cancer care run ahead of both the evidence and the regulatory reality.

How to read these numbers

Four cautions carry across everything above. First, design tier beats headline size: a randomized workload result at 44% is worth more than a retrospective AUC at 94%, because the first was tested against the thing that breaks most tools and the second was not. Second, retrospective reader studies flatter models — curated cases, a fixed validation era, and comparison conditions that can handicap the human arm. The discipline for reading any single paper is set out in our guide on how to read an AI validation study.

Third, reproducibility is never guaranteed. One of the most publicised breast-screening AI evaluations reported fewer false positives and false negatives than radiologists 8, but an independent commentary showed the result could not be independently reproduced because the code and data needed to check it were withheld 9. A result you cannot reproduce is a claim, not yet a finding. Fourth, detection is not mortality: nearly every number here measures finding cancer or saving reading time, and the leap to patients living longer needs its own long-horizon evidence, tracked alongside the rest of the field in our clinical AI trial results tracker.

These same patterns — strong where randomized, thin where retrospective — recur across specialties; our companion guide on AI in primary care walks the same discipline through first-contact care.

What this means for an oncology service

Frame adoption as questions, not as a shopping list.

  • What tier of evidence backs this specific tool, for this specific task? Screening reading has randomized support; most detection claims have only reader studies. Match your confidence to the tier.
  • Was it externally validated on patients like yours? A model tuned on one population and scanner can degrade elsewhere — the reason external validation is its own discipline.
  • What does it add to the workflow, and who reviews it? The deployable value today is a faster, more consistent second read, with a clinician in the loop — rather than a replacement for the diagnosis.
  • What are the downstream harms of more flags? Colonoscopy shows the cost of extra detection is extra harmless-polyp removal; every detection gain has a false-positive tail to budget for.

Because screening, diagnosis, and reimbursement decisions in oncology carry regulatory and liability weight, confirm the current authorization status and evidentiary requirements of any specific tool with your compliance or regulatory counsel before you rely on it — clearance is not the same as validation quality, and both change.

Sources and method

This guide synthesises the randomized and prospective evidence across the cancer-care pathway: the MASAI breast-screening trial and its full results 12, the PRAIM real-world implementation 3, a retrospective lung-nodule reader study 4, the pooled randomized evidence on computer-aided colonoscopy 56, a prostate-grading diagnostic study 7, a widely cited breast-AI evaluation 8 with the independent reproducibility commentary that qualifies it 9, and an independent taxonomy of the FDA device landscape 10. Every figure is drawn from the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a new randomized oncology-AI result or a longer-horizon screening follow-up is published. Dates and figures are current as of July 2026.

Questions & answers

  • Does AI improve cancer screening outcomes?

    In breast screening, the evidence is genuinely strong: a randomized trial found AI-supported reading cut radiologist workload by about 44% while detecting more cancers at a similar recall rate, and its full results were non-inferior on interval cancer. A large real-world implementation raised the detection rate by 17.6%. Both measure detection and workload; longer follow-up is still needed to show a mortality benefit.

  • Is AI used to treat cancer?

    Not autonomously. Across more than a thousand FDA-authorized AI devices, the landscape is dominated by imaging tasks and none use generative models. In oncology that means AI mostly helps detect and characterise findings — on a mammogram, a CT, a colonoscopy, a pathology slide — with a clinician reviewing and deciding. Treatment selection remains a clinician judgement.

  • Can I trust an oncology AI study that reports a high AUC?

    A high area-under-the-curve figure from a retrospective reader study is a promising signal, not proof of clinical value. Ask whether the model was externally validated on a different population, whether the comparison to clinicians was fair, and whether the result has been reproduced. One widely cited breast-AI paper could not be independently reproduced because its code and data were not shared.

Sources

  1. Lång K, Josefsson V, Larsson A-M, et al. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study. Lancet Oncology. 2023;24(8):936-944. doi.org/10.1016/S1470-2045(23)00298-X
  2. Hernström V, Josefsson V, Sartor H, et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading in the MASAI trial: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. The Lancet. 2026. doi.org/10.1016/S0140-6736(25)02464-X
  3. Eisemann N, Bunk S, Mukama T, et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening (PRAIM). Nature Medicine. 2025;31:917-924. doi.org/10.1038/s41591-024-03408-6
  4. Ardila D, Kiraly AP, Bharadwaj S, et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nature Medicine. 2019;25:954-961. doi.org/10.1038/s41591-019-0447-x
  5. Hassan C, Spadaccini M, Mori Y, et al. Real-Time Computer-Aided Detection of Colorectal Neoplasia During Colonoscopy: A Systematic Review and Meta-analysis. Annals of Internal Medicine. 2023;176(9):1209-1220. doi.org/10.7326/M22-3678
  6. Repici A, Badalamenti M, Maselli R, et al. Efficacy of Real-Time Computer-Aided Detection of Colorectal Neoplasia in a Randomized Trial. Gastroenterology. 2020;159(2):512-520.e7. doi.org/10.1053/j.gastro.2020.04.062
  7. Bulten W, Pinckaers H, van Boven H, et al. Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study. Lancet Oncology. 2020;21(2):233-241. doi.org/10.1016/S1470-2045(19)30739-9
  8. McKinney SM, Sieniek M, Godbole V, et al. International evaluation of an artificial intelligence system for breast cancer screening. Nature. 2020;577:89-94. doi.org/10.1038/s41586-019-1799-6
  9. Haibe-Kains B, Adam GA, Hosny A, et al. Transparency and reproducibility in artificial intelligence. Nature. 2020;586:E14-E16. doi.org/10.1038/s41586-020-2766-y
  10. Singh R, Bapna M, Diab AR, Ruiz ES, Lotter W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. npj Digital Medicine. 2025;8:388. doi.org/10.1038/s41746-025-01800-1