Benchmarks shape what the field builds, and in 2026 the most consequential new one measures something deceptively ordinary: whether a chatbot answers a clinician well. HealthBench Professional, published by OpenAI alongside its ChatGPT for Clinicians product, moves LLM evaluation from exam questions and consumer health queries to the tasks doctors actually delegate — and produced a headline in which the deployed model outscored specialist-matched physicians. This page explains how the benchmark works, where its evidence is strong, and where the claims must stop. As of August 2026.
What is HealthBench Professional?
The paper opens from usage: millions of clinicians use ChatGPT to support clinical care, while evaluations of the most common uses in model–clinician conversations remain limited 1. HealthBench Professional targets that gap with three domains drawn from real practice — care consults (reasoning through a case), writing and documentation, and medical research — each evaluated on physician-authored, multi-turn conversations scored against conversation-specific rubrics 13.
The construction details are what make it credible. Examples were curated from 15,079 candidates, with difficult cases enriched roughly 3.5-fold relative to the original pool; about a third involve deliberate adversarial testing; and three or more physicians adjudicated examples across three phases 1. The physician-comparison baseline was built generously: specialist-matched physicians wrote responses with unbounded time and web access 1. This is a benchmark designed by people who expected skeptical readers — which is exactly the posture our guide to LLM evals versus clinical evaluation recommends bringing to every leaderboard.
How does it differ from the original HealthBench?
The 2025 original was a general-population instrument: 5,000 multi-turn conversations between models and users or healthcare professionals, with 262 physicians from 60 countries writing 48,562 conversation-specific rubric criteria across themes from emergency referral to global health 2. It shipped with a physician-consensus variant and a "Hard" subset on which the best model at launch scored 32 percent 2 — and it documented steep model progress, from 16 percent for a 2023-era model to 60 percent for the o3 generation 2.
The Professional edition narrows the population (clinicians rather than laypeople), raises the difficulty deliberately, and tightens adjudication 1. Both share the rubric method: physicians specify what a good answer must contain, and a grader model scores responses against those criteria — an approach that scales far beyond what human panels could grade, while inheriting the limits of the rubrics themselves. Reference implementations of HealthBench are open-source under an MIT license 6, which matters practically: a hospital team can run the same harness against a candidate clinical LLM rather than take a vendor slide's word for it.
What did the results actually show?
Two findings carry the coverage. First, the model deployed in ChatGPT for Clinicians outperformed its own base model, other frontier models, and the physician-written responses on the benchmark's rubric scoring 1. Second, in pre-release testing reported around the launch, physician advisers reviewed roughly 7,000 conversations across clinical care, documentation and research, rating 99.6 percent of responses safe and accurate 5.
Read those numbers the way you would read any strong retrospective result — as covered in how to read an AI validation study. The comparison is between text responses, graded by rubric; it says a great deal about answer quality and calibration to clinician expectations, and something much weaker about diagnosis, treatment selection, or outcomes in deployed workflows. The distinction between benchmark performance and what it fails to capture does the heaviest lifting here: rubric grading rewards completeness and communication, an AUROC-style discrimination metric rewards ranking, and none of them measures what happens when a tired clinician half-reads a confident answer. Adversarial enrichment helps — a third of the examples probe failure modes deliberately 1 — but hallucination in clinical contexts and dataset shift after deployment remain outside any static benchmark's reach. We log frontier results, including this one, in our LLM medical benchmark tracker.
How does rubric grading actually work?
The mechanism deserves a paragraph, because it is quietly becoming the field's default. For each conversation, physicians write a rubric: a list of specific, checkable criteria a good response must satisfy — mention the red-flag symptom, advise the correct escalation, avoid the contraindicated suggestion — each carrying a positive or negative point weight. A grader model then checks a candidate response against every criterion, and the score is the weighted fraction satisfied 2. The design has two consequences worth holding onto. It scales: tens of thousands of criteria can be applied to every new model within hours, which no physician panel could do. And it inherits assumptions: the rubric encodes what its authors believed mattered at writing time, the grader model must itself judge whether prose satisfies a criterion, and both are fallible in ways the headline number hides. The original HealthBench paper validated grader agreement against physician judgment 2, which is the right response — and also a reminder that a "model-graded" benchmark is an instrument with its own error bars, a theme our guide to reading validation studies returns to often. Rubric scores also say little about calibration — whether the system's confidence tracks its correctness — which for a clinician-facing assistant is close to the whole safety question.
What is ChatGPT for Clinicians?
The product the benchmark was built to validate. Launched in April 2026 and free to verified US physicians, nurse practitioners, physician assistants and pharmacists, it packages clinical search with real-time cited answers, reusable "skills" for recurring workflows — referral letters, prior authorizations, patient instructions — continuing-education credit for clinical research queries, and an optional HIPAA business associate agreement for covered use 45. The skills list is a map of where general-purpose assistants are pressing into workflows we track elsewhere: patient-communication drafting, clinical reference lookup, and the prior-authorization pipeline. The verification gate matters legally as well as commercially — a consumer chatbot and a clinician tool sit on different sides of the state chatbot statutes in our mental-health chatbot law tracker, and covered-entity use runs through HIPAA's rules for LLMs.
What should a score like this change in your decisions?
Three calibrated conclusions travel well. First, the capability claim is real and independently checkable: the rubric harness is open, so a skeptical informatics team can reproduce the grading on its own cases 6 — a far better position than the field's usual take-our-word benchmarks. Second, the physician-comparison framing deserves its asterisks: written responses under no time pressure are a fair test of answer quality and an unfair proxy for clinical practice, where human oversight, interruption and accountability define the job. Third, local evaluation still decides: a benchmark is external validation of the general instrument, never of your deployment — build a golden set from your own consults, run the candidate model against it, and audit subgroup performance before anything touches a patient-facing workflow. Benchmarks move procurement conversations; they should never end them.
Sources and method
Benchmark construction, adjudication and results are drawn from the HealthBench Professional paper on arXiv 1 and its PDF as published by OpenAI 3; the original HealthBench's scale and scores are from the 2025 paper 2. Product details and launch facts for ChatGPT for Clinicians are from contemporaneous reporting by PYMNTS 4 and the Advisory Board's daily briefing 5, and the open-source evaluation code is the OpenAI simple-evals repository 6. Vendor-published evaluation numbers are labeled as such wherever they appear. We revisit this page twice a year, and sooner when new frontier results or independent critiques land. Current as of 1 August 2026.