Specialties

AI in primary care: exam scores versus patient outcomes

A guide to artificial intelligence at the point of first contact — decision support, conversational diagnosis, autonomous screening, and ambient documentation — reading the field's benchmark hype against the handful of large trials that measured what happened to patients. Each figure tied to its primary source. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The evidence most people cite for primary-care AI is an exam score: a large language model reached 86.5% on medical question-answering. Acing the exam and helping patients are different results.
  • The clearest test of that gap is a cluster-randomized trial in Kenyan primary care: 9,691 patients, an AI decision-support copilot, and a null patient outcome — treatment failure 2.2% vs 2.0%, adjusted odds ratio 0.77 — even though documentation improved and the tool was safe.
  • Conversational diagnostic AI that matched or beat primary-care physicians did so in a simulated, text-only exam with patient-actors, not in live clinics.
  • The one autonomous AI that makes a primary-care decision without a clinician reading the image is diabetic-retinopathy screening: 87.2% sensitivity, 90.7% specificity in its pivotal trial, and a narrow, well-bounded task.
  • Ambient scribes are the most deployable tool, but a randomized trial showed two products behave differently on the same outcome — so effects attach to specific tools, not the category. As of July 2026.

Primary care is where AI meets the widest, messiest range of patients, and where the gap between a benchmark and a bedside is largest. A model can score in the high 80s on medical exams and still fail to change what happens to the patient in front of a clinician — and in 2026 we finally have the large trials to show exactly that. This guide reads the field's exam-score progress against the handful of studies that measured patient-facing outcomes at the point of first contact: decision support, conversational diagnosis, autonomous screening, and ambient documentation. Every figure is tied to its primary source. As of July 2026.

The gap: exam scores are not patient outcomes

Start with the number everyone quotes. A large language model reached 86.5% on a benchmark of exam-style clinical questions, up from 67.6% a year earlier 4. That is a real advance in a controlled setting, and our benchmark tracker follows the record. But an exam score answers a narrow question — can the model produce the right answer when handed a clean question — and says little about the outcome that decides value in a clinic: do real patients, seen by real clinicians using the tool, end up better off.

The most disciplined test of that gap is a pragmatic cluster-randomized trial in Kenyan primary care. It enrolled 9,691 patients across 16 facilities, randomizing 103 clinical officers to an AI decision-support copilot or usual care. The primary endpoint — expert-adjudicated treatment failure within 14 days — was null: it occurred in 2.2% of AI-supported patients versus 2.0% of controls (adjusted odds ratio 0.77, 95% CI 0.55 to 1.08, P=0.13) 1. The tool was judged safe, and an independent panel rated documentation and treatment planning as higher quality in the AI arm — but the patient outcome did not move. We read this study closely in Briefing 001; it is the clearest demonstration in the field that process and outcome can decouple inside one well-run study.

That is the frame for everything that follows. The interesting question about a primary-care AI is never "how does it score," but "what did it change for patients, and how was that measured."

Decision support: what a copilot moves, and what it does not

The Kenya trial rewards a second look because its design is honest about its own limits, and those limits generalise. Randomization happened at the level of the individual clinical officer inside shared facilities, so AI-informed habits could drift into the control arm — a contamination that pushes any comparison toward null 1. The 14-day window is short, and slower benefits (a better-recognised chronic condition, an earlier risk flag) would land outside it. And treatment failure was rare, so even a trial this large stays underpowered for it.

None of that rescues an outcomes claim the trial did not earn; it explains why a clinical decision support system can improve the documented reasoning while leaving the measured outcome flat. The usable conclusion is specific: this copilot was safe and improved documentation, and that can justify a deployment when it is named as a documentation-and-reasoning benefit rather than dressed up as an outcomes result.

The setting also shapes what the result can and cannot tell you. The trial ran in a first-contact network staffed by clinical officers, where the baseline standard of care sets the ceiling on how much room any tool has to move the number — a high baseline leaves little headroom, a low one leaves more, and neither travels cleanly to a different health system 1. A copilot that adds little in one primary-care context can still add real value in another with thinner supervision or scarcer specialist backup, which is why one trial, however large, is a starting point rather than a verdict. It is also why the same tool should be re-tested where it will be used — the discipline our glossary on external validation makes concrete for every model that crosses a system boundary.

Conversational diagnosis: strong in the lab, untested in the clinic

The most striking diagnostic results in primary care come from conversational systems, and they come from the exam room of a simulation rather than a real one. In a blinded, remote objective structured clinical examination across 159 scenarios, a conversational diagnostic AI matched or exceeded primary-care physicians on diagnostic accuracy and on most conversation-quality axes, in text-based chat with trained patient-actors 2. A follow-up extended the same system to multimodal reasoning, still inside a simulated evaluation 3.

Hold two facts together. The result is genuinely impressive, and it is a text-only, actor-based exam — the diagnostic-AI equivalent of a very high test score. A clinical LLM that wins a simulated consultation has not yet faced the things that break tools in live primary care: a patient who volunteers the wrong history, a workflow that buries the suggestion, a clinician who must decide whether to trust it under time pressure. The Kenya result is what happens when a comparable class of tool meets those conditions. Treat simulated diagnostic superiority as a reason to run a real trial, not as evidence of one.

Autonomous screening: the one narrow task that works

There is exactly one place in primary care where AI makes a clinical decision without a clinician reviewing its work, and it is instructive precisely because it is narrow. Autonomous diabetic-retinopathy screening takes a retinal image in a primary-care office and returns a screening result. Its pivotal trial, across 900 patients, reported 87.2% sensitivity and 90.7% specificity for referable disease 5, and it supported the first authorization for an autonomous diagnostic AI of its kind.

Why does autonomy work here when it fails almost everywhere else? Because the task is bounded: one image type, one well-defined question, a large labelled evidence base, and a clear referral action. The lesson is not "autonomy is coming to primary care." It is that autonomy is earned task by task, where the question is narrow and the failure modes are understood — the opposite of an open-ended diagnostic conversation.

Primary-care AIDesign testedWhat it measuredReading
Decision-support copilot 1Cluster-randomized trial14-day treatment failureNull outcome; safe; better documentation
Conversational diagnosis 23Simulated OSCE with actorsDiagnostic accuracyMatched/beat physicians — in a lab
Autonomous DR screening 5Pivotal diagnostic trialSensitivity / specificity87.2% / 90.7% on a bounded task
Ambient scribe 67Randomized + deploymentDocumentation time, burnoutReal but tool-specific gains

The documentation layer: real relief, tool-specific gains

The most deployable primary-care AI does not diagnose at all — it writes the note. Ambient AI scribes listen to the encounter and draft documentation for the clinician to review. A vendor-independent randomized trial of 238 physicians found the category is not uniform: one scribe cut time-in-note by 9.5% versus control, while another showed no statistically significant change on that outcome 6. A large integrated deployment reported lower documentation burden and improved clinician experience across primary-care-heavy practices 7.

Two disciplines follow. First, effects attach to specific tools and specific workflows, so a category claim ("scribes save time") is weaker than a local measurement. Second, every draft is reviewed by a clinician before it enters the record — the human-in-the-loop step that keeps a drafting error from silently becoming part of the chart, and the reason the review time, rather than the raw capture, is the real cost to measure.

That review step is load-bearing because ambient scribes carry a distinctive failure mode: they can add detail a clinician never said and omit detail that was said, a different problem from the transcription slips of older dictation tools. A primary-care clinician seeing thirty patients has limited attention for each generated note, so the safety of the tool depends on a review that is fast enough to be sustainable and careful enough to catch a confabulated finding — a tension that a local pilot measures directly and a vendor benchmark cannot. The gains are worth pursuing; they are worth pursuing with the review discipline named, not assumed.

Screening signals that flow into the clinic

Primary care also absorbs the output of consumer AI, whether it asked to or not. A smartwatch study of 419,297 participants sent irregular-pulse notifications to 0.52% of them; among notified participants who returned an ECG patch, 84% of notifications were concordant with atrial fibrillation 8. Read from the clinic, that is a double-edged number: a plausible early signal for a serious rhythm, and a new stream of worried, notified patients arriving for confirmation — most of the cohort was young and low-risk. Consumer screening AI is now part of the primary-care workload, and planning for the false-positive tail is part of adopting it.

How to read these numbers

Four cautions carry across the category. First, the benchmark is the weakest evidence, not the strongest — an exam score is where a claim starts, and the Kenya trial is what testing it looks like. Second, simulated superiority is not clinical superiority: an actor-based OSCE removes exactly the frictions that decide whether a diagnostic tool helps in a real clinic. Third, autonomy is task-specific — it is earned where the question is narrow and the failure modes are mapped, and diabetic-retinopathy screening is the exception that proves the rule. Fourth, process gains are not outcome gains; better documentation and better-rated reasoning are worth having, and they are a different claim from "patients did better."

The general discipline for holding any one of these studies to account is in our guide on how to read an AI validation study, and the running numbers behind the field sit in the clinical AI trial results tracker. The same "strong where randomized, thin where retrospective" pattern runs through AI in oncology as well.

What this means for a primary-care practice

Frame adoption as questions, not as a search for the highest-scoring model.

  • What outcome did this tool actually change, in whom, over what horizon? Prefer the study that measured patients over the one that measured a benchmark.
  • Is the task bounded enough for the claim being made? Autonomy is credible for narrow screening, far less so for open-ended diagnosis.
  • Who reviews the output, and what does review cost? For scribes and copilots alike, the human-in-the-loop step is where safety and the real time-cost both live.
  • Can you measure it locally? A randomized trial found two scribes diverge on the same outcome — the strongest argument for a small, blinded pilot in your own clinic before committing.

Because triage, diagnosis, and reimbursement decisions in primary care carry clinical and regulatory weight, confirm the current authorization status and evidentiary requirements of any specific tool with your compliance or regulatory counsel before relying on it. Clearance status is not the same as validation quality, and both change.

Sources and method

This guide reads the field's benchmark and simulated-exam results against its patient-facing evidence: a cluster-randomized decision-support trial 1, a simulated diagnostic OSCE and its multimodal follow-up 23, the exam-style benchmark that anchors the hype 4, the pivotal trial behind autonomous diabetic-retinopathy screening 5, a vendor-independent scribe trial and a large deployment 67, and a consumer atrial-fibrillation screening study 8. Every figure is drawn from the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a primary-care AI tool is tested against a patient outcome in a new setting. Dates and figures are current as of July 2026.

Questions & answers

  • Can AI diagnose patients in primary care?

    In simulated exams, a conversational diagnostic system matched or exceeded primary-care physicians on diagnostic accuracy — but that was a text-only test with trained patient-actors, not live care. In the one large randomized trial of an AI copilot used by real clinicians on real patients, documentation improved but the 14-day patient outcome did not change. The honest reading is that diagnostic AI is promising and largely unproven at the point of care.

  • Is any primary-care AI fully autonomous?

    One narrow task is: autonomous diabetic-retinopathy screening, where the AI returns a screening result without a clinician reading the image. Its pivotal trial reported 87.2% sensitivity and 90.7% specificity. It works because the task is bounded and the failure modes are understood. Most other primary-care AI keeps a clinician in the loop by design.

  • Do ambient AI scribes help primary-care clinicians?

    They are the most deployable tool in the category, and clinicians often report real relief from documentation load. But a randomized trial found two widely used scribes behaved differently on time-in-note, so the gains depend on the specific product and how much a clinician actually uses it. Measure in your own practice rather than trusting a category claim.

Sources

  1. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine. Published online 26 June 2026. doi.org/10.1038/s41591-026-04503-6
  2. Tu T, Schaekermann M, Palepu A, et al. Towards conversational diagnostic artificial intelligence. Nature. 2025;642:442-450. doi.org/10.1038/s41586-025-08866-7
  3. Advancing conversational diagnostic AI with multimodal reasoning. Nature Medicine. 2026. doi.org/10.1038/s41591-026-04371-0
  4. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge [Med-PaLM]. Nature. 2023;620:172-180. doi.org/10.1038/s41586-023-06291-2
  5. Abràmoff MD, Lavin PT, Birch M, Shah N, Folk JC. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices. npj Digital Medicine. 2018;1:39. doi.org/10.1038/s41746-018-0040-6
  6. Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI. 2025. doi.org/10.1056/AIoa2501000
  7. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. 2024;5(3). doi.org/10.1056/CAT.23.0404
  8. Perez MV, Mahaffey KW, Hedlin H, et al. Large-Scale Assessment of a Smartwatch to Identify Atrial Fibrillation [Apple Heart Study]. New England Journal of Medicine. 2019;381:1909-1917. doi.org/10.1056/NEJMoa1901183