Every few weeks a new model tops a medical benchmark, and the headline writes itself: the machine "passes the boards." This page tracks what those benchmarks actually score, who currently leads, and — the part the leaderboards leave out — how far a high number sits from a clinical warrant. Every figure is dated and tied to a numbered source below. As of July 2026.
Which two benchmarks anchor the field?
Two efforts define the serious end of clinical LLM evaluation right now.
MedHELM, built on the HELM framework at Stanford's Center for Research on Foundation Models, scores models across 121 clinical tasks organised into 5 categories and 22 subcategories — a taxonomy assembled with 29 practising clinicians — using a suite of 35 benchmarks, 17 drawn from existing datasets and 18 newly built 1. Its distinguishing move is coverage: instead of a single exam, it spans diagnostic decision support, note summarisation, patient communication, and administrative work, and it includes tasks grounded in real electronic health records rather than curated question banks.
HealthBench, released by OpenAI, takes the opposite tack — depth over breadth of task type. It grades models on 5,000 multi-turn conversations against 48,562 unique rubric criteria written by 262 physicians who practise across 60 countries and 26 specialties 3. Each conversation is scored against its own rubric, so a model earns credit for the specific things a physician would want said, and loses it for the specific things a physician would flag — including under uncertainty and in emergencies.
Who currently leads?
On MedHELM's evaluation of nine frontier models, the pattern that emerged was that advanced reasoning models lead. DeepSeek R1 posted a 66% win-rate and o3-mini a 64% win-rate across the task suite, while Claude 3.5 Sonnet reached comparable results at roughly 40% lower estimated compute cost 2. The headline ranking, though, is the least durable thing on this page: it turns over with each model generation. The more stable finding is structural — reasoning-tuned models tend to do better on multi-step clinical tasks than raw scale alone predicts.
| Benchmark | What it scores | Scale | Current signal (as of July 2026) |
|---|---|---|---|
| MedHELM | 121 clinical tasks, clinician taxonomy | 35 benchmarks, 9 models | Reasoning models lead (DeepSeek R1 66%, o3-mini 64% win-rate) 12 |
| HealthBench | 5,000 graded conversations | 48,562 rubric criteria, 262 physicians | Open-ended, rubric-graded dialogue across 26 specialties 3 |
What do benchmark scores license clinically — and what don't they?
Here is the part the leaderboards omit. A systematic review of 519 studies published between January 2022 and February 2024 examined how health-care LLMs were actually being evaluated — and the shape of that evaluation is narrow 4.
- Only 5% used real patient-care data. The overwhelming majority tested models on curated questions rather than live records.
- 44.5% assessed medical-knowledge and licensing-examination questions, the single most common task, followed by diagnosis at 19.5%. Fully 84.2% of studies were question-answering; summarisation (8.9%) and open dialogue (3.3%) were rare.
- 95.4% used accuracy as the primary dimension. Fairness and bias were examined in 15.8% of studies, deployment considerations in 4.6%, and calibration and uncertainty in just 1.2%.
Read together, those numbers explain the gap between a benchmark score and a clinical warrant. A model that answers licensing-exam items correctly has demonstrated recall and reasoning on a tidy, single-answer format. That is a real capability — and a weak proxy for drafting a safe note from a rambling encounter, reconciling a contradictory record, or knowing when to say "I am not sure, escalate." The tasks that dominate the benchmarks are the tasks least like the ones that carry clinical risk.
This is why MedHELM's inclusion of real-EHR tasks matters, and why HealthBench's rubric grading of uncertainty and emergency behaviour matters: both push scoring toward the messier competencies. But even they are laboratory measures. A high score says a model is worth putting in front of a prospective, local evaluation. It does not stand in for that evaluation.
How to read these numbers
Four cautions travel with every score on this page. Leaderboard rank is a snapshot that turns over with each model release — treat any single ranking as perishable. A benchmark measures the tasks its authors chose; coverage gaps are invisible in the headline number. Exam-style performance and bedside performance are different measures, and the review above shows how rarely the second is tested. And almost all scoring optimises accuracy while leaving calibration, bias, and deployment behaviour largely unmeasured — precisely the properties that govern safety.
The practical rule: use benchmarks to shortlist and to falsify — a model that fails MedHELM's coding tasks is unlikely to surprise you in production — but require prospective, in-setting validation before a score touches a clinical decision. Our sibling clinical AI trial results tracker follows the studies that do that testing, and what benchmark scores don't tell you goes deeper on the reading frame.
Sources and method
Figures are drawn from the primary publications listed below: the MedHELM paper and its leaderboard 12, the HealthBench paper 3, and the JAMA systematic review of evaluation practice 4. We revisit this page on a ninety-day cycle and whenever a benchmark ships a new version, a frontier model posts new scores, or a peer-reviewed methodology critique lands. As of July 2026.