Benchmarks tell you what a model can do in a laboratory; trials tell you what a tool does in a clinic. This page tracks the notable published trials and deployment studies of clinical AI from roughly the last year, grouped by what they actually measured. Each entry carries its design, its effect, and its single most important caveat. Every figure is dated and tied to a numbered source below. As of July 2026.
Ambient documentation: real, but tool-dependent
The cleanest recent test of ambient AI scribes is a three-arm pragmatic randomized trial of 238 outpatient physicians across 14 specialties, who were assigned to one of two commercial scribes — DAX Copilot or Nabla — or to usual care. The result is a useful corrective to category-level hype: Nabla significantly reduced time-in-note versus control, while DAX showed no significant difference, and both tools improved clinician-reported experience and task load 1. The lesson is that "ambient scribe" names a category rather than a guarantee — effects attach to specific products in specific workflows.
A larger deployment study across six health systems (263 clinicians) adds the wellbeing signal: the share of clinicians reporting burnout fell from 51.9% to 38.8% thirty days after a scribe was introduced 2. That is a meaningful shift, though it comes from a pre-post design relying on self-report over a short window, so it is best read alongside the controlled time results rather than instead of them.
Imaging: the strongest positive result
The headline positive trial of the past year is MASAI, a randomized, population-based screening-accuracy study of 105,915 women comparing AI-supported mammography reading with standard double reading by radiologists. The full results are strong: sensitivity rose to 80.5% from 73.8% at an identical 98.5% specificity, the AI arm detected 29% more cancers without raising false positives, and the interval-cancer rate was non-inferior (1.55 versus 1.76 per 1,000) 3. An earlier safety analysis from the same trial had already shown AI support cut radiologists' screen-reading workload by 44% 4. MASAI is the current high-water mark for clinical AI evidence — a large randomized design with a hard outcome. The caveat is scope: it tests reading support inside an established double-reading programme in one country, so it speaks to that setting rather than to screening systems built differently.
Endoscopy: the cautionary counterweight
Against that, a multicentre observational study within the ACCEPT programme delivered the year's most sobering finding. Comparing colonoscopy quality before and after AI-assisted detection was introduced at four centres, it found that endoscopists' unassisted adenoma detection rate fell by about 6 percentage points — from roughly 28% to 22% — once they had spent time working with the AI 5. It is the first real-world evidence of AI-linked deskilling: the tool may lift performance while it is on and quietly erode unaided skill underneath. Because the design is observational and before-and-after, confounding cannot be ruled out — but the direction of the signal is the point, and it argues for monitoring baseline skill wherever a detection aid is deployed.
Screening at the edges: performance is uneven
A real-world validation of AI diabetic-retinopathy screening on 250 people in North Indian public-health clinics shows how far laboratory numbers can drift in the field. Across the algorithms tested — each validated elsewhere — sensitivity ranged from 60% to 80% and specificity from 14% to 96% on the same clinic images; the best-performing algorithm reached 78.9% sensitivity and 98.1% specificity for referable disease 6. The spread is the finding. A published accuracy figure earned on one camera and one population is a starting hypothesis for another, and this study is a reminder to validate in the setting where the tool will run.
Administrative agents: promising, unproven on outcomes
The newest frontier is agentic and administrative work, and here the evidence is early. A quality study of ChatGPT-5-drafted prior-authorization letters across 29 nephrology scenarios found strong performance on the clinical side — 89.7% showed strong reasoning, 93.1% cited valid sources, and only 3.5% contained a false statement — yet the authors concluded the drafts are not yet reliable enough to use without clinician verification, with errors clustering in codes and citations 7. This is a drafting-quality study rather than an outcome trial; whether such agents change approval times, denials, or clinician hours at scale remains untested in the peer-reviewed record.
| Study | Design | Effect | Key caveat |
|---|---|---|---|
| Ambient scribes 1 | RCT, 238 physicians, 3 arms | Nabla cut time-in-note; DAX no change | Effects are product-specific |
| Scribe deployment 2 | Pre-post, 6 systems | Burnout 51.9% → 38.8% | Self-report, short window |
| MASAI mammography 34 | RCT, 105,915 women | Sensitivity 80.5% vs 73.8%; 44% less workload | One double-reading programme |
| Colonoscopy 5 | Multicentre observational | Unassisted ADR fell ~6 points | Observational; deskilling risk |
| Retinopathy screening 6 | Real-world validation, 250 people | Specificity 14%–96% across tools | Performance setting-dependent |
| Prior-auth letters 7 | Quality study, 29 scenarios | 89.7% strong reasoning | Needs clinician review; no outcome data |
How to read these numbers
Four cautions carry across the table. Design quality varies enormously — a randomized trial with a hard outcome (MASAI) and a before-and-after audit (the colonoscopy study) support very different strengths of claim, and the effect size matters less than the design behind it. Effects are product- and setting-specific: the scribe trial shows two tools in the same category diverging. Real-world performance drifts from validation figures, sometimes dramatically. And the benefits most studies measure — time, workload, detection — are nearer than the outcomes that matter most, such as mortality or long-term harm, which few of these designs are built to see.
The practical rule mirrors our medical benchmark tracker: weigh the design before the headline, and require local, prospective validation before a result changes practice. For the framing behind that discipline, see prospective vs retrospective evaluation.
Sources and method
Figures are drawn from the primary publications listed below. We revisit this page on a ninety-day cycle and whenever a new randomized or multi-site trial, a longer-term follow-up, or a peer-reviewed outcome study of a clinical agent is published. As of July 2026.