Statistics

Clinical AI trial results tracker

A running tally of the notable trials and deployment studies of clinical AI from the last year — ambient documentation, mammography, colonoscopy, retinopathy screening, and administrative agents — each with its design, its effect, and its most important caveat. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The strongest recent evidence is modality-specific: a randomized trial found one ambient scribe cut time-in-note while another showed no significant change — the tool matters more than the category.
  • The MASAI mammography trial (105,915 women) is the headline positive result: AI-supported reading raised sensitivity to 80.5% from 73.8% at identical specificity, detected 29% more cancers, and cut reader workload by 44%.
  • A cautionary counterweight: a multicentre study found endoscopists' unassisted adenoma detection rate fell about 6 points after AI-assisted colonoscopy — the first real-world signal of deskilling.
  • Real-world screening results are uneven: across AI diabetic-retinopathy algorithms, specificity ranged from 14% to 96% on the same public-health images.
  • Administrative agents are promising but unproven on outcomes: LLM-drafted prior-authorization letters scored well on reasoning yet still require clinician verification.

Benchmarks tell you what a model can do in a laboratory; trials tell you what a tool does in a clinic. This page tracks the notable published trials and deployment studies of clinical AI from roughly the last year, grouped by what they actually measured. Each entry carries its design, its effect, and its single most important caveat. Every figure is dated and tied to a numbered source below. As of July 2026.

Ambient documentation: real, but tool-dependent

The cleanest recent test of ambient AI scribes is a three-arm pragmatic randomized trial of 238 outpatient physicians across 14 specialties, who were assigned to one of two commercial scribes — DAX Copilot or Nabla — or to usual care. The result is a useful corrective to category-level hype: Nabla significantly reduced time-in-note versus control, while DAX showed no significant difference, and both tools improved clinician-reported experience and task load 1. The lesson is that "ambient scribe" names a category rather than a guarantee — effects attach to specific products in specific workflows.

A larger deployment study across six health systems (263 clinicians) adds the wellbeing signal: the share of clinicians reporting burnout fell from 51.9% to 38.8% thirty days after a scribe was introduced 2. That is a meaningful shift, though it comes from a pre-post design relying on self-report over a short window, so it is best read alongside the controlled time results rather than instead of them.

Imaging: the strongest positive result

The headline positive trial of the past year is MASAI, a randomized, population-based screening-accuracy study of 105,915 women comparing AI-supported mammography reading with standard double reading by radiologists. The full results are strong: sensitivity rose to 80.5% from 73.8% at an identical 98.5% specificity, the AI arm detected 29% more cancers without raising false positives, and the interval-cancer rate was non-inferior (1.55 versus 1.76 per 1,000) 3. An earlier safety analysis from the same trial had already shown AI support cut radiologists' screen-reading workload by 44% 4. MASAI is the current high-water mark for clinical AI evidence — a large randomized design with a hard outcome. The caveat is scope: it tests reading support inside an established double-reading programme in one country, so it speaks to that setting rather than to screening systems built differently.

Endoscopy: the cautionary counterweight

Against that, a multicentre observational study within the ACCEPT programme delivered the year's most sobering finding. Comparing colonoscopy quality before and after AI-assisted detection was introduced at four centres, it found that endoscopists' unassisted adenoma detection rate fell by about 6 percentage points — from roughly 28% to 22% — once they had spent time working with the AI 5. It is the first real-world evidence of AI-linked deskilling: the tool may lift performance while it is on and quietly erode unaided skill underneath. Because the design is observational and before-and-after, confounding cannot be ruled out — but the direction of the signal is the point, and it argues for monitoring baseline skill wherever a detection aid is deployed.

Screening at the edges: performance is uneven

A real-world validation of AI diabetic-retinopathy screening on 250 people in North Indian public-health clinics shows how far laboratory numbers can drift in the field. Across the algorithms tested — each validated elsewhere — sensitivity ranged from 60% to 80% and specificity from 14% to 96% on the same clinic images; the best-performing algorithm reached 78.9% sensitivity and 98.1% specificity for referable disease 6. The spread is the finding. A published accuracy figure earned on one camera and one population is a starting hypothesis for another, and this study is a reminder to validate in the setting where the tool will run.

Administrative agents: promising, unproven on outcomes

The newest frontier is agentic and administrative work, and here the evidence is early. A quality study of ChatGPT-5-drafted prior-authorization letters across 29 nephrology scenarios found strong performance on the clinical side — 89.7% showed strong reasoning, 93.1% cited valid sources, and only 3.5% contained a false statement — yet the authors concluded the drafts are not yet reliable enough to use without clinician verification, with errors clustering in codes and citations 7. This is a drafting-quality study rather than an outcome trial; whether such agents change approval times, denials, or clinician hours at scale remains untested in the peer-reviewed record.

StudyDesignEffectKey caveat
Ambient scribes 1RCT, 238 physicians, 3 armsNabla cut time-in-note; DAX no changeEffects are product-specific
Scribe deployment 2Pre-post, 6 systemsBurnout 51.9% → 38.8%Self-report, short window
MASAI mammography 34RCT, 105,915 womenSensitivity 80.5% vs 73.8%; 44% less workloadOne double-reading programme
Colonoscopy 5Multicentre observationalUnassisted ADR fell ~6 pointsObservational; deskilling risk
Retinopathy screening 6Real-world validation, 250 peopleSpecificity 14%–96% across toolsPerformance setting-dependent
Prior-auth letters 7Quality study, 29 scenarios89.7% strong reasoningNeeds clinician review; no outcome data

How to read these numbers

Four cautions carry across the table. Design quality varies enormously — a randomized trial with a hard outcome (MASAI) and a before-and-after audit (the colonoscopy study) support very different strengths of claim, and the effect size matters less than the design behind it. Effects are product- and setting-specific: the scribe trial shows two tools in the same category diverging. Real-world performance drifts from validation figures, sometimes dramatically. And the benefits most studies measure — time, workload, detection — are nearer than the outcomes that matter most, such as mortality or long-term harm, which few of these designs are built to see.

The practical rule mirrors our medical benchmark tracker: weigh the design before the headline, and require local, prospective validation before a result changes practice. For the framing behind that discipline, see prospective vs retrospective evaluation.

Sources and method

Figures are drawn from the primary publications listed below. We revisit this page on a ninety-day cycle and whenever a new randomized or multi-site trial, a longer-term follow-up, or a peer-reviewed outcome study of a clinical agent is published. As of July 2026.

Questions & answers

  • Is there randomized-trial evidence that clinical AI works?

    Yes, in specific tasks. The MASAI trial randomized 105,915 women and found AI-supported mammography reading raised sensitivity and detected more cancers without more false positives. A separate randomized trial of ambient scribes found one tool cut documentation time while another did not. The evidence is real but modality- and product-specific, so a positive result for one tool does not transfer to another.

  • Can clinical AI make clinicians worse?

    One multicentre study raised that concern directly: after a period of AI-assisted colonoscopy, endoscopists' unassisted adenoma detection rate fell by roughly 6 percentage points. It is an observational finding rather than a randomized one, but it is the first real-world signal that routine AI assistance may erode unaided skill — a reason to monitor deskilling wherever a tool is deployed.

  • Do AI screening tools perform the same in the real world as in validation?

    Often no. In a real-world evaluation of AI diabetic-retinopathy screening on public-health images, specificity ranged from 14% to 96% across algorithms that had each been validated elsewhere. Real-world performance depends on the camera, the population, and the setting, which is why local validation matters more than a vendor's published figure.

Sources

  1. Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI. 2025. doi.org/10.1056/AIoa2501000
  2. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025;8(10):e2534976. doi.org/10.1001/jamanetworkopen.2025.34976
  3. Hernström V, Josefsson V, Larsson AM, et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. The Lancet. 2026. doi.org/10.1016/S0140-6736(25)02464-X
  4. Lång K, Josefsson V, Larsson AM, et al. Artificial intelligence-supported screen reading versus standard double reading in the MASAI trial: a clinical safety analysis. The Lancet Oncology. 2023;24(8):936-944. doi.org/10.1016/S1470-2045(23)00298-X
  5. Budzyń K, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology. 2025. doi.org/10.1016/S2468-1253(25)00133-5
  6. Real-World Evaluation of AI-Driven Diabetic Retinopathy Screening in Public Health Settings: Validation and Implementation Study. JMIR Medical Informatics. 2025. doi.org/10.2196/67529
  7. Aiumtrakul N, Thongprayoon C, et al. Quality assessment of large language model–generated prior authorization letters in nephrology. Frontiers in Digital Health. 2026. doi.org/10.3389/fdgth.2026.1767648