Statistics

LLM medical benchmark results tracker

A running read of how large language models score on the leading clinical benchmarks — MedHELM and HealthBench — set beside the peer-reviewed critique of what those scores do, and do not, tell you about bedside performance. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Two benchmarks anchor the field as of July 2026: MedHELM, which scores models across 121 clinician-defined clinical tasks, and HealthBench, which grades 5,000 conversations against 48,562 physician-written rubric criteria.
  • On MedHELM's leaderboard, advanced reasoning models lead — DeepSeek R1 (66% win-rate) and o3-mini (64%) — with Claude 3.5 Sonnet close behind at lower estimated cost.
  • The catch: a systematic review of 519 studies found only 5% used real patient-care data, while 44.5% tested medical-knowledge and licensing-exam questions. Benchmarks mostly measure exam recall, not bedside work.
  • Accuracy dominates: 95.4% of studies scored it, while calibration and uncertainty (1.2%), deployment factors (4.6%), and bias (15.8%) were rarely measured.
  • A high benchmark score licenses a hypothesis worth testing in your setting — it does not license clinical use. Prospective, local validation is still the bar.

Every few weeks a new model tops a medical benchmark, and the headline writes itself: the machine "passes the boards." This page tracks what those benchmarks actually score, who currently leads, and — the part the leaderboards leave out — how far a high number sits from a clinical warrant. Every figure is dated and tied to a numbered source below. As of July 2026.

Which two benchmarks anchor the field?

Two efforts define the serious end of clinical LLM evaluation right now.

MedHELM, built on the HELM framework at Stanford's Center for Research on Foundation Models, scores models across 121 clinical tasks organised into 5 categories and 22 subcategories — a taxonomy assembled with 29 practising clinicians — using a suite of 35 benchmarks, 17 drawn from existing datasets and 18 newly built 1. Its distinguishing move is coverage: instead of a single exam, it spans diagnostic decision support, note summarisation, patient communication, and administrative work, and it includes tasks grounded in real electronic health records rather than curated question banks.

HealthBench, released by OpenAI, takes the opposite tack — depth over breadth of task type. It grades models on 5,000 multi-turn conversations against 48,562 unique rubric criteria written by 262 physicians who practise across 60 countries and 26 specialties 3. Each conversation is scored against its own rubric, so a model earns credit for the specific things a physician would want said, and loses it for the specific things a physician would flag — including under uncertainty and in emergencies.

Who currently leads?

On MedHELM's evaluation of nine frontier models, the pattern that emerged was that advanced reasoning models lead. DeepSeek R1 posted a 66% win-rate and o3-mini a 64% win-rate across the task suite, while Claude 3.5 Sonnet reached comparable results at roughly 40% lower estimated compute cost 2. The headline ranking, though, is the least durable thing on this page: it turns over with each model generation. The more stable finding is structural — reasoning-tuned models tend to do better on multi-step clinical tasks than raw scale alone predicts.

BenchmarkWhat it scoresScaleCurrent signal (as of July 2026)
MedHELM121 clinical tasks, clinician taxonomy35 benchmarks, 9 modelsReasoning models lead (DeepSeek R1 66%, o3-mini 64% win-rate) 12
HealthBench5,000 graded conversations48,562 rubric criteria, 262 physiciansOpen-ended, rubric-graded dialogue across 26 specialties 3

What do benchmark scores license clinically — and what don't they?

Here is the part the leaderboards omit. A systematic review of 519 studies published between January 2022 and February 2024 examined how health-care LLMs were actually being evaluated — and the shape of that evaluation is narrow 4.

  • Only 5% used real patient-care data. The overwhelming majority tested models on curated questions rather than live records.
  • 44.5% assessed medical-knowledge and licensing-examination questions, the single most common task, followed by diagnosis at 19.5%. Fully 84.2% of studies were question-answering; summarisation (8.9%) and open dialogue (3.3%) were rare.
  • 95.4% used accuracy as the primary dimension. Fairness and bias were examined in 15.8% of studies, deployment considerations in 4.6%, and calibration and uncertainty in just 1.2%.

Read together, those numbers explain the gap between a benchmark score and a clinical warrant. A model that answers licensing-exam items correctly has demonstrated recall and reasoning on a tidy, single-answer format. That is a real capability — and a weak proxy for drafting a safe note from a rambling encounter, reconciling a contradictory record, or knowing when to say "I am not sure, escalate." The tasks that dominate the benchmarks are the tasks least like the ones that carry clinical risk.

This is why MedHELM's inclusion of real-EHR tasks matters, and why HealthBench's rubric grading of uncertainty and emergency behaviour matters: both push scoring toward the messier competencies. But even they are laboratory measures. A high score says a model is worth putting in front of a prospective, local evaluation. It does not stand in for that evaluation.

How to read these numbers

Four cautions travel with every score on this page. Leaderboard rank is a snapshot that turns over with each model release — treat any single ranking as perishable. A benchmark measures the tasks its authors chose; coverage gaps are invisible in the headline number. Exam-style performance and bedside performance are different measures, and the review above shows how rarely the second is tested. And almost all scoring optimises accuracy while leaving calibration, bias, and deployment behaviour largely unmeasured — precisely the properties that govern safety.

The practical rule: use benchmarks to shortlist and to falsify — a model that fails MedHELM's coding tasks is unlikely to surprise you in production — but require prospective, in-setting validation before a score touches a clinical decision. Our sibling clinical AI trial results tracker follows the studies that do that testing, and what benchmark scores don't tell you goes deeper on the reading frame.

Sources and method

Figures are drawn from the primary publications listed below: the MedHELM paper and its leaderboard 12, the HealthBench paper 3, and the JAMA systematic review of evaluation practice 4. We revisit this page on a ninety-day cycle and whenever a benchmark ships a new version, a frontier model posts new scores, or a peer-reviewed methodology critique lands. As of July 2026.

Questions & answers

  • What is the best LLM for medical tasks according to benchmarks?

    On MedHELM's leaderboard, as of July 2026, advanced reasoning models rank highest — DeepSeek R1 posted a 66% win-rate and o3-mini 64%, with Claude 3.5 Sonnet reaching comparable results at lower estimated compute cost. A leaderboard position reflects performance on defined tasks, and the ranking shifts with every new model release, so read it as a snapshot rather than a standing verdict.

  • Do medical LLM benchmarks predict real clinical performance?

    Only partly. A systematic review of 519 studies found that just 5% used real patient-care data and that 44.5% tested medical-knowledge or licensing-exam questions. Exam-style scores measure recall and reasoning on curated questions, which is a weak proxy for how a model behaves on messy records, incomplete histories, and live clinical workflows.

  • What is the difference between MedHELM and HealthBench?

    MedHELM, from Stanford's Center for Research on Foundation Models, scores models across 121 clinical tasks in a clinician-built taxonomy, including tasks drawn from real electronic health records. HealthBench, from OpenAI, grades 5,000 open-ended conversations against 48,562 rubric criteria written by 262 physicians. MedHELM emphasises task coverage; HealthBench emphasises graded, realistic dialogue.

Sources

  1. Bedi S, Liu Y, Orr-Ewing L, et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nature Medicine. 2026. doi.org/10.1038/s41591-025-04151-2
  2. Bedi S, et al. MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks. arXiv:2505.23802. 2025. arxiv.org/abs/2505.23802
  3. Arora RK, Wei J, Soskin Hicks R, et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775. 2025. arxiv.org/abs/2505.08775
  4. Bedi S, Jain SS, Bedi P, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025;333(4):319-328. doi.org/10.1001/jama.2024.21700