Every week brings another paper, preprint, or vendor deck reporting that a model matches or beats clinicians at some task. Most of these numbers are real. Very few of them mean what the headline implies. The difference between a result that survives contact with real patients and one that quietly fails is rarely the size of the number — it is how the study was designed, who it was tested on, and which reporting standard it was held to. This page is a working checklist you can hold against any validation study, with the base rates and the standards that tell you where to look. As of July 2026.
Why a checklist beats a headline
The evidence that headlines mislead is itself well documented. A systematic review of deep-learning studies in medical imaging found that "few prospective deep learning studies and randomised trials exist," that "most non-randomised trials are not prospective, are at high risk of bias, and deviate from existing reporting standards," and that 61 of 81 published studies still claimed the algorithm was comparable or superior to clinicians 1. A separate review of 516 AI imaging studies found that "only 6% (31 studies) performed external validation," and that none of those 31 combined the three design features — a diagnostic-cohort design, multiple institutions, and prospective data collection — that a robust test would need 2. A meta-analysis of 82 studies comparing deep learning with clinicians reached the same verdict from a different angle: an out-of-sample external validation was done in only 25, of which just 14 compared the algorithm and clinicians on the same sample, and the authors concluded that "poor reporting is prevalent" 3.
| Review | Studies examined | Externally validated |
|---|---|---|
| AI diagnostic imaging design audit 2 | 516 | 6% (31) |
| Deep learning vs clinicians meta-analysis 3 | 82 | 25 (14 same-sample) |
| AI vs clinicians reporting review 1 | 81 published | Few; most high risk of bias |
Read together, these tell you the prior you should carry into any new study: the base rate of rigorous evaluation is low, so the burden is on the paper to prove it is one of the good ones. The nine questions below are how you make it prove it.
The nine questions
Work through these in order. Each maps to something a trustworthy study is expected to report, and to the standard that defines the expectation.
| # | Ask | What a good answer looks like | Governing standard |
|---|---|---|---|
| 1 | Was it externally validated? | Tested on data from a different site, scanner, or era than it was built on | TRIPOD+AI 5 |
| 2 | What was the risk of bias? | Appraised across participants, predictors, outcome, analysis | PROBAST 6 |
| 3 | Discrimination and calibration? | Both reported, never AUROC alone | TRIPOD+AI 5 |
| 4 | Was the operating point reported? | Sensitivity and specificity at the threshold that will be used | STARD 10 |
| 5 | Prospective or retrospective; silent or live? | Timing stated; live evaluation follows silent testing | DECIDE-AI 7 |
| 6 | Who is in the sample? | Spectrum of patients resembles yours; inclusion rules stated | STARD 10 |
| 7 | What is the comparator? | A fair, adequately sized clinician or standard-of-care baseline | CONSORT-AI 8 |
| 8 | What outcome, measured how? | A clinically meaningful endpoint, defined before analysis | TRIPOD+AI 5 |
| 9 | What exactly is claimed? | The conclusion matches what was actually tested | CONSORT-AI 8 |
1. Was it externally validated?
This is the question that filters out most of the field. A model's performance on held-out data drawn from the same source it was trained on tells you how well it memorised that source, and routinely overstates how it behaves elsewhere. External validation — testing on a different site, scanner, coding practice, or time period — is where inflated results collapse. Given that only 6% of imaging studies in one large audit cleared this bar 2, a paper reporting genuine external validation has already distinguished itself. When you see only internal validation, treat every figure as an upper bound.
2. What was the risk of bias?
Risk of bias is the structured version of "where could this study have fooled itself?" PROBAST, the standard tool for prediction models, "is organized into the following 4 domains: participants, predictors, outcome, and analysis," containing "a total of 20 signaling questions" 6. You do not need to run the full instrument to use its logic: ask whether the patients were selected in a way that inflates the event rate, whether any predictor secretly encodes the outcome, whether the outcome was judged blind to the model, and whether the analysis had enough events to support the number of variables. A study at high risk of bias on any one domain can post a beautiful headline that means nothing.
3. Discrimination and calibration?
Most papers report discrimination — usually AUROC — and stop. Discrimination is only half the story. It measures whether the model ranks sicker patients above healthier ones; it says nothing about whether a predicted risk of 20% actually corresponds to a 20% event rate. That second property is calibration, and it is "often ignored" even though "poorly calibrated algorithms can be misleading and potentially harmful for clinical decision-making" 9. A model can rank well and still hand you probabilities that are systematically too high, driving over-treatment. Insist on seeing both, and read our companion guides on AUROC and calibration curves for how each is reported and where each fails.
4. Was the operating point reported?
Care happens at one threshold. An alert fires or it does not; a case is flagged or missed. So the number that governs bedside behaviour is sensitivity and specificity at the specific operating point that will be deployed — which is why STARD's "essential items for reporting diagnostic accuracy studies" include that point, rather than a summary curve alone 10. The cost of skipping this is concrete. When a widely implemented proprietary sepsis model was externally validated, it "predicted the onset of sepsis with an area under the curve of 0.63" and, at its live alerting threshold, "did not identify 1709 patients with sepsis (67%)" 4. The summary figure was mediocre; the operating-point miss rate was the number that mattered, and it was worse.
5. Prospective or retrospective; silent or live?
A retrospective study run on curated historical data is a promising start, never a finish. DECIDE-AI exists precisely to structure the next stage: it is the "reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence" 7, covering the move from silent testing — where the model runs but no one acts on it — to live use where it shapes decisions. Ask where on that path the study sits. A model evaluated only in silent mode has yet to face the thing that breaks most tools: how clinicians actually respond to its outputs.
6. Who is in the sample?
A model inherits the population it was tested on. Spectrum bias — evaluating on patients who are unusually easy to classify, such as clear positives against healthy controls — inflates accuracy that evaporates on the ambiguous middle-of-the-road patients who fill real clinics. Check the inclusion and exclusion rules, the disease prevalence in the sample, and whether the case mix resembles the patients in front of you. A dermatology model validated on biopsy-confirmed lesions from one skin-tone distribution is telling you about that distribution, and no other. Prevalence deserves particular attention: a test set enriched with obvious positives inflates accuracy figures and shifts the meaning of any threshold, so a result reported without the sample's event rate is missing a number you need to interpret every other number in the paper.
7. What is the comparator?
"Better than clinicians" is only meaningful against a fair clinician baseline. CONSORT-AI, the "reporting guidelines for clinical trial reports for interventions involving artificial intelligence" 8, formalises what a credible comparison requires. In practice the comparator is often small, unblinded, or denied the context a real clinician would have — reading a single image with no history. When the human arm is handicapped, the gap flatters the algorithm.
8. What outcome, and measured how?
An impressive result on a proxy endpoint can hide the absence of any patient benefit. A model that improves a documentation metric has not been shown to improve care; a model that flags more findings has not been shown to change management or outcomes. Ask what was measured, whether it was defined before the analysis, and whether it is the outcome you actually care about.
9. What exactly is claimed?
Finally, hold the conclusion against the design. The recurring failure — visible in the 61 of 81 studies that claimed parity with or superiority to clinicians on thin evidence 1 — is a claim that outruns what the study could support. A retrospective, internally validated, single-site model has earned the claim "promising in this setting," and no more.
How to read this checklist
Three cautions travel with it. First, no single question is disqualifying on its own; a retrospective study can be excellent groundwork, and an externally validated one can still be biased elsewhere. Weigh the answers together. Second, the checklist assesses the evidence, not the model — a well-studied model that performs modestly is more trustworthy than a dazzling one that has never been tested properly. Third, standards evolve: the specific guideline names here are current as of July 2026, and the underlying questions outlast any one revision. For regulatory-clearance and reimbursement decisions that turn on these studies, confirm the current evidentiary requirements with your compliance or regulatory counsel before acting — clearance status is not the same as validation quality.
Sources and method
This guide synthesises three peer-reviewed systematic reviews that quantify how often clinical AI studies are externally validated and at what risk of bias 123, one external-validation study used as a worked cautionary example 4, and the five reporting standards that define what a trustworthy study must disclose: TRIPOD+AI 5, PROBAST 6, DECIDE-AI 7, CONSORT-AI 8, and STARD 10, alongside the calibration literature 9. Every figure is drawn from the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a governing standard is revised. For the numbers behind the field, see our clinical AI trial results tracker; for the risk-monitoring side of deployment, see algorithmovigilance.