Guides · Evaluation
How to read the evidence behind clinical AI.
Every AI vendor arrives with numbers — an AUROC, a benchmark score, an accuracy claim. Evaluation is the discipline of asking what those numbers actually demonstrate, and it is the one skill this library treats as foundational: a clinician who can read a validation study critically can work through every other question about AI in healthcare from evidence rather than marketing.
These guides build that skill in plain language. The statistical core: what AUROC does and does not tell you, why calibration matters as much as discrimination, and how sample size and confidence intervals separate a finding from noise. The study-design layer: internal versus external validation, prospective versus retrospective evidence, subgroup performance and bias audits, and the dataset shift that quietly decays deployed models. And the applied layer: how to read a validation study and an FDA clearance summary line by line, why LLM evals are a different instrument from clinical evaluation, ten red flags of an overfit model claim, and what benchmark scores leave out. Written for clinicians, researchers, and health leaders — no machine-learning background assumed, no rigor surrendered.
13 guides in this collection
HealthBench Professional, explained: how OpenAI now measures clinician-facing AI
OpenAI's clinician benchmark grew a professional edition in 2026 — physician- authored conversations, multi-stage physician adjudication, and a headline result in which the deployed model outscored specialist-matched physicians. What the benchmark actually measures, how it was built, and what a score like that should and should never change in your decisions. As of August 2026.
Updated 23 Jul 2026LLM evals vs clinical evaluation: two different instruments
Why a benchmark score and a clinical evaluation measure different things, in different units, on different populations — and three published cases where the same models looked strong on one instrument and weak on the other. As of July 2026.
Updated 22 Jul 2026AUROC explained for clinicians
What the area under the ROC curve really measures, worked through two real published models — an AUROC of 0.991 that still forces a trade-off, and one of 0.63 that hid a 67% miss rate — and the four questions it can never answer for you. As of July 2026.
Updated 22 Jul 2026Calibration curves for clinicians
How to read a calibration plot the way it fails — each shape of curve paired with a real published model that failed that way, and the bedside decision each failure would distort. Discrimination tells you the ranking is right; calibration tells you the number is. As of July 2026.
Updated 22 Jul 2026Dataset shift and model decay
A clinical AI model is accurate at one moment, on one population. Both move. This is a working guide to dataset shift and model decay — the three ways a model's performance comes apart, the signal that warns you first, and what the evidence says actually slows it down. As of July 2026.
Updated 22 Jul 2026How to read an AI validation study
A working checklist a clinician can hold against any paper claiming an AI model works — the nine questions that decide whether a result will survive contact with real patients, each tied to the reporting standard that governs it. As of July 2026.
Updated 22 Jul 2026How to read an FDA clearance summary
A section-by-section walk of one real 510(k) summary for a cleared AI head-CT triage device — the clearance letter, the indications, the predicate, the performance table — mapped to what clearance does and does not establish. As of July 2026.
Updated 22 Jul 2026Internal vs external validation
The gap between how an AI model scores on a held-out slice of its own data and how it scores at a second hospital is where most inflated results go to die. This is a working guide to internal versus external validation — what each measures, how far performance typically falls between them, and the published figures that show it. As of July 2026.
Updated 22 Jul 2026Prospective vs retrospective evaluation
A model that looks brilliant on curated historical data has cleared the lowest bar, not the last one. This is a working guide to prospective versus retrospective evaluation — what each can prove, the documented cases where retrospective promise did not survive prospective testing, and the one where it did. As of July 2026.
Updated 22 Jul 2026Sample size and confidence intervals in AI studies
A single accuracy number with no interval around it is a guess wearing a lab coat. This is a working guide to sample size and confidence intervals in clinical AI studies — why so many are underpowered, how wide the real uncertainty is, and the published figures that prove it. As of July 2026.
Updated 22 Jul 2026Subgroup performance and bias audits
A model can post excellent overall accuracy and still fail the specific patients who most need it to work. This is a working guide to subgroup performance and bias audits — the documented failures, why they happen, and the stratified metrics that surface them before deployment. As of July 2026.
Updated 22 Jul 2026Ten red flags of an overfit model claim
Ten warning signs that a model's reported performance will not survive contact with new patients — each one anchored to a documented, published failure, and paired with the specific check that would have caught it. As of July 2026.
Updated 22 Jul 2026What benchmark scores don't tell you
A benchmark score is a real measurement of a narrow thing. This guide reads the peer-reviewed critique of medical AI benchmarks — saturation, contamination, and construct validity — and the one question to ask before a leaderboard number reaches a clinical decision. As of July 2026.