Primary care is where AI meets the widest, messiest range of patients, and where the gap between a benchmark and a bedside is largest. A model can score in the high 80s on medical exams and still fail to change what happens to the patient in front of a clinician — and in 2026 we finally have the large trials to show exactly that. This guide reads the field's exam-score progress against the handful of studies that measured patient-facing outcomes at the point of first contact: decision support, conversational diagnosis, autonomous screening, and ambient documentation. Every figure is tied to its primary source. As of July 2026.
The gap: exam scores are not patient outcomes
Start with the number everyone quotes. A large language model reached 86.5% on a benchmark of exam-style clinical questions, up from 67.6% a year earlier 4. That is a real advance in a controlled setting, and our benchmark tracker follows the record. But an exam score answers a narrow question — can the model produce the right answer when handed a clean question — and says little about the outcome that decides value in a clinic: do real patients, seen by real clinicians using the tool, end up better off.
The most disciplined test of that gap is a pragmatic cluster-randomized trial in Kenyan primary care. It enrolled 9,691 patients across 16 facilities, randomizing 103 clinical officers to an AI decision-support copilot or usual care. The primary endpoint — expert-adjudicated treatment failure within 14 days — was null: it occurred in 2.2% of AI-supported patients versus 2.0% of controls (adjusted odds ratio 0.77, 95% CI 0.55 to 1.08, P=0.13) 1. The tool was judged safe, and an independent panel rated documentation and treatment planning as higher quality in the AI arm — but the patient outcome did not move. We read this study closely in Briefing 001; it is the clearest demonstration in the field that process and outcome can decouple inside one well-run study.
That is the frame for everything that follows. The interesting question about a primary-care AI is never "how does it score," but "what did it change for patients, and how was that measured."
Decision support: what a copilot moves, and what it does not
The Kenya trial rewards a second look because its design is honest about its own limits, and those limits generalise. Randomization happened at the level of the individual clinical officer inside shared facilities, so AI-informed habits could drift into the control arm — a contamination that pushes any comparison toward null 1. The 14-day window is short, and slower benefits (a better-recognised chronic condition, an earlier risk flag) would land outside it. And treatment failure was rare, so even a trial this large stays underpowered for it.
None of that rescues an outcomes claim the trial did not earn; it explains why a clinical decision support system can improve the documented reasoning while leaving the measured outcome flat. The usable conclusion is specific: this copilot was safe and improved documentation, and that can justify a deployment when it is named as a documentation-and-reasoning benefit rather than dressed up as an outcomes result.
The setting also shapes what the result can and cannot tell you. The trial ran in a first-contact network staffed by clinical officers, where the baseline standard of care sets the ceiling on how much room any tool has to move the number — a high baseline leaves little headroom, a low one leaves more, and neither travels cleanly to a different health system 1. A copilot that adds little in one primary-care context can still add real value in another with thinner supervision or scarcer specialist backup, which is why one trial, however large, is a starting point rather than a verdict. It is also why the same tool should be re-tested where it will be used — the discipline our glossary on external validation makes concrete for every model that crosses a system boundary.
Conversational diagnosis: strong in the lab, untested in the clinic
The most striking diagnostic results in primary care come from conversational systems, and they come from the exam room of a simulation rather than a real one. In a blinded, remote objective structured clinical examination across 159 scenarios, a conversational diagnostic AI matched or exceeded primary-care physicians on diagnostic accuracy and on most conversation-quality axes, in text-based chat with trained patient-actors 2. A follow-up extended the same system to multimodal reasoning, still inside a simulated evaluation 3.
Hold two facts together. The result is genuinely impressive, and it is a text-only, actor-based exam — the diagnostic-AI equivalent of a very high test score. A clinical LLM that wins a simulated consultation has not yet faced the things that break tools in live primary care: a patient who volunteers the wrong history, a workflow that buries the suggestion, a clinician who must decide whether to trust it under time pressure. The Kenya result is what happens when a comparable class of tool meets those conditions. Treat simulated diagnostic superiority as a reason to run a real trial, not as evidence of one.
Autonomous screening: the one narrow task that works
There is exactly one place in primary care where AI makes a clinical decision without a clinician reviewing its work, and it is instructive precisely because it is narrow. Autonomous diabetic-retinopathy screening takes a retinal image in a primary-care office and returns a screening result. Its pivotal trial, across 900 patients, reported 87.2% sensitivity and 90.7% specificity for referable disease 5, and it supported the first authorization for an autonomous diagnostic AI of its kind.
Why does autonomy work here when it fails almost everywhere else? Because the task is bounded: one image type, one well-defined question, a large labelled evidence base, and a clear referral action. The lesson is not "autonomy is coming to primary care." It is that autonomy is earned task by task, where the question is narrow and the failure modes are understood — the opposite of an open-ended diagnostic conversation.
| Primary-care AI | Design tested | What it measured | Reading |
|---|---|---|---|
| Decision-support copilot 1 | Cluster-randomized trial | 14-day treatment failure | Null outcome; safe; better documentation |
| Conversational diagnosis 23 | Simulated OSCE with actors | Diagnostic accuracy | Matched/beat physicians — in a lab |
| Autonomous DR screening 5 | Pivotal diagnostic trial | Sensitivity / specificity | 87.2% / 90.7% on a bounded task |
| Ambient scribe 67 | Randomized + deployment | Documentation time, burnout | Real but tool-specific gains |
The documentation layer: real relief, tool-specific gains
The most deployable primary-care AI does not diagnose at all — it writes the note. Ambient AI scribes listen to the encounter and draft documentation for the clinician to review. A vendor-independent randomized trial of 238 physicians found the category is not uniform: one scribe cut time-in-note by 9.5% versus control, while another showed no statistically significant change on that outcome 6. A large integrated deployment reported lower documentation burden and improved clinician experience across primary-care-heavy practices 7.
Two disciplines follow. First, effects attach to specific tools and specific workflows, so a category claim ("scribes save time") is weaker than a local measurement. Second, every draft is reviewed by a clinician before it enters the record — the human-in-the-loop step that keeps a drafting error from silently becoming part of the chart, and the reason the review time, rather than the raw capture, is the real cost to measure.
That review step is load-bearing because ambient scribes carry a distinctive failure mode: they can add detail a clinician never said and omit detail that was said, a different problem from the transcription slips of older dictation tools. A primary-care clinician seeing thirty patients has limited attention for each generated note, so the safety of the tool depends on a review that is fast enough to be sustainable and careful enough to catch a confabulated finding — a tension that a local pilot measures directly and a vendor benchmark cannot. The gains are worth pursuing; they are worth pursuing with the review discipline named, not assumed.
Screening signals that flow into the clinic
Primary care also absorbs the output of consumer AI, whether it asked to or not. A smartwatch study of 419,297 participants sent irregular-pulse notifications to 0.52% of them; among notified participants who returned an ECG patch, 84% of notifications were concordant with atrial fibrillation 8. Read from the clinic, that is a double-edged number: a plausible early signal for a serious rhythm, and a new stream of worried, notified patients arriving for confirmation — most of the cohort was young and low-risk. Consumer screening AI is now part of the primary-care workload, and planning for the false-positive tail is part of adopting it.
How to read these numbers
Four cautions carry across the category. First, the benchmark is the weakest evidence, not the strongest — an exam score is where a claim starts, and the Kenya trial is what testing it looks like. Second, simulated superiority is not clinical superiority: an actor-based OSCE removes exactly the frictions that decide whether a diagnostic tool helps in a real clinic. Third, autonomy is task-specific — it is earned where the question is narrow and the failure modes are mapped, and diabetic-retinopathy screening is the exception that proves the rule. Fourth, process gains are not outcome gains; better documentation and better-rated reasoning are worth having, and they are a different claim from "patients did better."
The general discipline for holding any one of these studies to account is in our guide on how to read an AI validation study, and the running numbers behind the field sit in the clinical AI trial results tracker. The same "strong where randomized, thin where retrospective" pattern runs through AI in oncology as well.
What this means for a primary-care practice
Frame adoption as questions, not as a search for the highest-scoring model.
- What outcome did this tool actually change, in whom, over what horizon? Prefer the study that measured patients over the one that measured a benchmark.
- Is the task bounded enough for the claim being made? Autonomy is credible for narrow screening, far less so for open-ended diagnosis.
- Who reviews the output, and what does review cost? For scribes and copilots alike, the human-in-the-loop step is where safety and the real time-cost both live.
- Can you measure it locally? A randomized trial found two scribes diverge on the same outcome — the strongest argument for a small, blinded pilot in your own clinic before committing.
Because triage, diagnosis, and reimbursement decisions in primary care carry clinical and regulatory weight, confirm the current authorization status and evidentiary requirements of any specific tool with your compliance or regulatory counsel before relying on it. Clearance status is not the same as validation quality, and both change.
Sources and method
This guide reads the field's benchmark and simulated-exam results against its patient-facing evidence: a cluster-randomized decision-support trial 1, a simulated diagnostic OSCE and its multimodal follow-up 23, the exam-style benchmark that anchors the hype 4, the pivotal trial behind autonomous diabetic-retinopathy screening 5, a vendor-independent scribe trial and a large deployment 67, and a consumer atrial-fibrillation screening study 8. Every figure is drawn from the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a primary-care AI tool is tested against a patient outcome in a new setting. Dates and figures are current as of July 2026.