Oncology is where clinical AI has some of its strongest evidence and some of its most overstated claims, often about the same task. A mammography reader that saves radiologist time in a randomized trial and a lung-nodule model that "beats radiologists" in a retrospective slideshow are graded on very different scales, and the difference decides whether a result will survive contact with real patients. This guide walks the cancer-care pathway — screening, detection, pathology, and the device landscape underneath — and sorts what has randomized evidence from what has only a promising reader study. Every figure is tied to its primary source. As of July 2026.
Where the evidence actually sits
The honest map of oncology AI is less a list of capabilities than a list of evidence tiers. A tool can be technically impressive and clinically unproven at the same time, and the pathway below is arranged by how much a claim has been tested, not by how advanced the model is.
| Pathway stage | Best evidence tier | Headline result | The caveat to hold |
|---|---|---|---|
| Breast screening | Randomized trial 12 | ~44% less reading workload, more cancers detected | Detection and workload, not yet mortality |
| Breast screening (real world) | Large implementation study 3 | +17.6% detection rate, non-inferior recall | Non-randomized; site and reader selection |
| Colonoscopy | Meta-analysis of randomized trials 56 | Adenoma detection 44.7% vs 36.7% | More harmless-polyp removal; no advanced-adenoma gain |
| Lung-nodule reading | Retrospective reader study 4 | AUC 94.4%; beat six readers with no prior scan | No deployment yet; single validation era |
| Prostate grading | Diagnostic study 7 | Grading comparable to pathologists | Laboratory setting, not live sign-out |
| Device landscape | Authorization taxonomy 10 | Imaging-dominated; no generative models | Detection and triage, not autonomous care |
Read down that "caveat" column before any capability column. It is the reason two tools that both "work" can deserve very different amounts of trust.
Screening: the strongest evidence in the field
Breast screening is the clearest case where AI has been tested the hard way. In a randomized, controlled trial, women were assigned to AI-supported single reading or to standard double reading by two radiologists. AI-supported reading detected 244 screen-detected cancers versus 203 with double reading — a higher cancer detection rate at a similar recall rate — while reducing the screen-reading workload by 44.3% 1. That workload figure matters as much as the detection figure: in systems short of radiologists, halving the reading burden without losing cancers is itself the outcome.
The interim publication was a safety analysis, and the obvious question was whether the early detection signal would hold on the endpoint that matters more — interval cancers, the ones that surface symptomatically between screening rounds. The full results answered it: AI-supported screening had higher sensitivity and a non-inferior interval-cancer rate compared with double reading 2. A randomized design that survives its own harder endpoint is the strongest thing anyone can say about a clinical AI tool today.
Alongside the trial sits a different kind of evidence: real-world implementation. Across 461,818 screening cases, AI-supported double reading achieved a breast cancer detection rate of 6.7 per 1,000 versus 5.7 per 1,000 — 17.6% higher — with a recall rate that was non-inferior to standard reading 3. That is a large number of women and a reassuring direction, but it is an observational implementation rather than a randomized comparison, so reader and site selection can flatter the result. Trial and implementation point the same way here, which is exactly the alignment you want before believing a screening claim.
| Breast-screening study | Design | Detection signal | Recall / workload |
|---|---|---|---|
| MASAI 12 | Randomized trial | More cancers; higher sensitivity | 44% less reading workload; non-inferior interval cancer |
| PRAIM 3 | Real-world implementation | +17.6% detection rate | Non-inferior recall rate |
One caution travels with both: they measure detection, workload, and interval cancer, rather than cancer mortality. The pathway from "more cancers found earlier" to "fewer women die" is long and has surprised the screening field before — finding more cancer can mean catching lethal disease sooner, and it can also mean over-detecting indolent disease that would never have harmed the patient. The two look identical in a detection-rate figure and diverge only in long-term follow-up. So the right claim is "more detection at lower reading cost," and no more, until mortality and over-diagnosis data arrive.
A second caution is generalizability. Both studies ran in European screening programs with their own equipment, populations, and reading standards 13; a model that performs this well there can degrade on a different scanner, a different density distribution, or an under-represented group, which is why the screening result belongs to the systems that produced it until re-tested in yours. That is the practical face of external validation: a strong screening number is a reason to pilot locally, rather than a license to deploy sight-unseen.
Detection and diagnosis: promising, mostly retrospective
Step off screening and the evidence tier drops. The famous imaging results are overwhelmingly retrospective reader studies — a model reads stored scans and its output is compared with clinicians after the fact. These are legitimate early signals, and they are routinely presented as if they were deployment results, which they are not.
Lung-nodule reading is the archetype. A deep-learning model reached an AUC of 94.4% on low-dose CT screening cases and, when no prior scan was available for comparison, outperformed all six radiologists in the study with 11% fewer false positives and 5% fewer false negatives 4. It is a genuinely strong result — and it is a retrospective evaluation on curated cases, with the "no prior imaging" condition doing quiet work, because in practice radiologists lean heavily on comparison with earlier scans. Read the AUROC alongside the operating point the model would actually run at, and treat the reader-study framing as the ceiling rather than the expectation.
Colonoscopy is the counter-example that shows what happens when the same category is tested with randomized trials. Pooled across randomized trials, computer-aided detection raised the adenoma detection rate to 44.7% from 36.7% — a real, replicated gain in finding pre-cancerous polyps 5, consistent with individual trials showing higher detection without lengthening the procedure 6. But the same meta-analysis carries the discipline vendors omit: computer-aided detection also increased removal of non-neoplastic polyps and did not increase detection of advanced adenomas 5. More flags is not the same as more of the cancers that matter, and the extra harmless-polyp removals carry their own small costs. This is the mature version of an AI-detection story: it works, and here is precisely where it does not.
Pathology: grading at the microscope
Digital pathology is the quietest oncology-AI success. Grading a prostate biopsy — assigning a Gleason pattern — is high-stakes and famously subject to disagreement between pathologists, which makes it a natural target. A deep-learning system developed on 5,759 biopsies from 1,243 patients achieved grading comparable to pathologists 7. The value here is less about beating an expert and more about consistency: a tireless second reader that flags discordant cases and standardises grade where humans drift.
The caveat is the setting. This is a diagnostic study on assembled slides, rather than a live sign-out under time pressure with the full clinical context a pathologist carries. A tool that grades well in the study still has to prove it in the workflow, which is why deployment belongs behind a human-in-the-loop review rather than in front of it.
The device landscape: imaging, and a human in the loop
It helps to see the whole board. An independent taxonomy of 1,016 FDA authorizations found the AI device landscape dominated by imaging tasks, with none based on generative large language models 10. In oncology that shape is decisive: authorized AI clusters in radiology and pathology detection and characterisation, and the tools are regulated as software as a medical device that supports a clinician who reviews and decides. The fuller counts by year and specialty sit in our FDA-cleared AI devices tracker.
The practical reading: the oncology AI you can actually buy and deploy today helps find and measure things, and a clinician still owns the diagnosis and the treatment plan. Claims of autonomous cancer care run ahead of both the evidence and the regulatory reality.
How to read these numbers
Four cautions carry across everything above. First, design tier beats headline size: a randomized workload result at 44% is worth more than a retrospective AUC at 94%, because the first was tested against the thing that breaks most tools and the second was not. Second, retrospective reader studies flatter models — curated cases, a fixed validation era, and comparison conditions that can handicap the human arm. The discipline for reading any single paper is set out in our guide on how to read an AI validation study.
Third, reproducibility is never guaranteed. One of the most publicised breast-screening AI evaluations reported fewer false positives and false negatives than radiologists 8, but an independent commentary showed the result could not be independently reproduced because the code and data needed to check it were withheld 9. A result you cannot reproduce is a claim, not yet a finding. Fourth, detection is not mortality: nearly every number here measures finding cancer or saving reading time, and the leap to patients living longer needs its own long-horizon evidence, tracked alongside the rest of the field in our clinical AI trial results tracker.
These same patterns — strong where randomized, thin where retrospective — recur across specialties; our companion guide on AI in primary care walks the same discipline through first-contact care.
What this means for an oncology service
Frame adoption as questions, not as a shopping list.
- What tier of evidence backs this specific tool, for this specific task? Screening reading has randomized support; most detection claims have only reader studies. Match your confidence to the tier.
- Was it externally validated on patients like yours? A model tuned on one population and scanner can degrade elsewhere — the reason external validation is its own discipline.
- What does it add to the workflow, and who reviews it? The deployable value today is a faster, more consistent second read, with a clinician in the loop — rather than a replacement for the diagnosis.
- What are the downstream harms of more flags? Colonoscopy shows the cost of extra detection is extra harmless-polyp removal; every detection gain has a false-positive tail to budget for.
Because screening, diagnosis, and reimbursement decisions in oncology carry regulatory and liability weight, confirm the current authorization status and evidentiary requirements of any specific tool with your compliance or regulatory counsel before you rely on it — clearance is not the same as validation quality, and both change.
Sources and method
This guide synthesises the randomized and prospective evidence across the cancer-care pathway: the MASAI breast-screening trial and its full results 12, the PRAIM real-world implementation 3, a retrospective lung-nodule reader study 4, the pooled randomized evidence on computer-aided colonoscopy 56, a prostate-grading diagnostic study 7, a widely cited breast-AI evaluation 8 with the independent reproducibility commentary that qualifies it 9, and an independent taxonomy of the FDA device landscape 10. Every figure is drawn from the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a new randomized oncology-AI result or a longer-horizon screening follow-up is published. Dates and figures are current as of July 2026.