Each issue, this briefing reads one study closely and passes on what holds up. Issue 002 takes the largest study yet of the most widely purchased AI tool in care delivery: the ambient AI scribe. Deployment ran far ahead of evidence — our adoption tracker and evidence guide document that gap — and the sales pitch has always been denominated in time. This paper finally prices that pitch at scale, in minutes, from the record system's own logs 1.
What the study did
Five US academic health systems introduced ambient scribes to their clinicians between June 2023 and August 2025. The study followed 8,581 ambulatory clinicians through the rollout — 57.1% women, spread across primary care (24.4%), medical (62.4%), and surgical (13.2%) specialties — of whom 1,809 (21%) adopted a scribe, an opt-in decision at four of the five sites 1. The work was co-led by investigators at Mass General Brigham and the University of California, San Francisco 2.
The outcomes are unglamorous on purpose: total time on the electronic health record (EHR), time on documentation, time on the EHR outside scheduled hours, each normalized to 8 scheduled patient-hours, plus weekly visit volume 1 — measured from system activity data rather than from anyone's recollection of their evening.
What it found
Set the point estimates next to their intervals 1:
- Documentation time fell 16.0 minutes per 8 scheduled patient-hours (95% confidence interval, CI, 13.7 to 18.3) — about a 10% relative reduction 2.
- Total EHR time fell 13.4 minutes (95% CI 9.1 to 17.7) — about 3% relative 2.
- Visit volume rose 0.49 visits per week (95% CI 0.17 to 0.81) — half a visit per clinician.
- After-hours EHR time did not change significantly — the evenings-back claim, on this evidence, stayed a claim.
- Dose mattered. Clinicians using the scribe in at least half their visits saw roughly double the total-EHR-time reduction and triple the documentation reduction; only 32% of adopters used it that much 2.
The authors' conclusion is measured, and worth keeping verbatim in spirit: adoption was associated with modest decreases in EHR and documentation time and a modest increase in weekly visit volume 1.
Sixteen minutes is real. It is also the entire average claim — and the after-hours tail did not move.
Why this matters for healthcare
This is the study to bring to the next scribe renewal conversation. The category's marketing runs on time saved; the largest cohort in the field now puts the average at minutes per clinical day, concentrated in the minority who use the tool intensively 12. That reframes the deployment question from "which scribe" to "what utilization" — and it says the documented burden moved while the after-hours pattern, the part closest to the burnout story (see the burnout evidence guide), did not.
It also continues the theme this briefing opened with in issue 001: process measures and the outcome you were promised can decouple inside the same well-run study. There, better documentation without better patient outcomes; here, less documentation time without shorter evenings. The discipline is the same — read the endpoint the study measured, at the size it measured it. The paper: doi.org/10.1001/jama.2026.2253. The member discussion below walks the methods, the selection problem, and what a deploying leader should instrument before the pilot.