Ambient AI

AI scribes: what the evidence actually shows

Ambient scribes reached millions of uses before the first randomized trial reported a result. This guide sets the deployment scale and the vendor efficiency claims against the peer-reviewed record — how few evaluations meet real-world-evidence criteria, and what the studies that do exist actually found. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Ambient scribes reached system-wide daily use ahead of the evidence: one deployment passed 2.5 million uses in its first year, while a real-world-evidence review found only 6 of 1,450 identified studies met inclusion criteria.
  • The first randomized trial showed the effect is tool-dependent — one scribe cut documentation time 9.5% versus control, the other showed no significant change — and it framed the well-being gains as secondary findings needing confirmation.
  • Most published evidence is pre-post, single-site, self-reported, and short-window; a rapid review found the studies fell short of standardized evaluation frameworks and called for large-scale work on accuracy, safety, and cost before further expansion.
  • Accuracy is the question most effectiveness studies skip: when it is directly measured, a validated evaluation detected hallucinated content in 31% of AI notes versus 20% of clinician notes.
  • The financial case is unproven, and the rigor gap mirrors clinical AI at large — a pattern of high-risk-of-bias studies that outrun what they can support. As of July 2026.

Ambient AI scribes crossed a threshold few healthcare tools reach: system-wide, daily, at scale. One integrated group reported that its scribe passed over 2.5 million uses within a single year 1. That happened before the first randomized trial of these tools reported a result. The gap between how widely ambient scribes are used and how well they have been studied is the subject of this guide. It sets the deployment numbers and the efficiency claims against the peer-reviewed record, asks how much of that record would meet the bar we apply to any clinical technology, and reports what the strongest studies actually found. As of July 2026.

This page appraises the published evidence for ambient scribes; it is neither a product recommendation nor a clearance determination nor legal advice. Regulatory status, contract terms, and safety obligations differ by tool and setting — confirm procurement and compliance decisions with your own governance, IT, and counsel before acting.

Deployment outran the evidence

Start with the scale, because it frames everything else. The most-cited deployment — the running numbers live in our AI scribe adoption statistics — reached millions of uses in its first year 1. Vendors describe minutes saved and burnout lifted; health systems have bought in on that promise. An ambient AI scribe is now the most widely deployed generative-AI tool in clinical care.

Against that, consider what the formal evidence synthesis found. A 2025 real-world-evidence rapid review screened the literature and reported that "of the 1450 studies identified, 6 met the inclusion criteria" 2. Six. And those six were "an observational study, a case report, a peer-matched cohort study, and survey-based assessments" 2 — designs that describe and associate, rather than designs that isolate an effect. The review's own verdict was that "most studies fell short when compared to standardized frameworks" for quality, safety, and workflow evaluation, and that "before further expansion is considered, robust large-scale studies on usability, acceptance, effectiveness, patient feedback, accuracy, safety, and cost should be conducted" 2.

That is the shape of the problem: a technology in millions of encounters, resting on a published base of a handful of studies that its own reviewers judge below standard. The rest of this guide is about what that base does and does not let you conclude.

What would "real-world evidence" require?

"Real-world evidence" has become a marketing phrase, so it is worth pinning down what it means as a standard rather than a slogan. A claim that a tool improves care in practice — not in a demo, not on a curated dataset — earns trust to the degree that the study behind it did four things: measured a meaningful outcome (documentation time, burnout, note accuracy) rather than a proxy; used a fair comparator, ideally a randomized control group; ran prospectively over a realistic window; and reported enough to be appraised against a recognized standard. These are the same questions our guide on how to read an AI validation study applies to any model, and they map onto the idea of external validation: a result that holds only in the site and sample that produced it has yet to prove it travels.

Held to that bar, most of the ambient-scribe literature falls into a familiar category — the single-site, pre-versus-post, self-reported, short-window study. Those studies are useful groundwork. They are also the weakest design for the strongest claims, because a clinician who has just been given a much-hyped tool and asked how they feel a month later is answering a different question than "does this save measurable time against a comparable group who did not get it."

The base rate: what the studies actually are

The table below classifies the strongest publicly reported evidence by what kind of study it is and what it can support. Read it as a map of the field's center of gravity, current as of July 2026.

Evidence typeExample findingWhat it can support
Deployment report2.5M uses in year one 1Feasibility and reach — not effect
Pre-post self-reportBurnout 51.9% → 38.8% at 30 days 9Association; felt relief; no controlled effect
EHR log cohort~13 fewer EHR min/day; 32% used it for >½ visits 6Measured time change; partial uptake
Randomized trial9.5% time-in-note drop for one tool, none for the other 3Causal effect — tool-specific, single-center
Accuracy auditHallucinated content in 31% vs 20% of notes 5Error profile — rarely paired with effect studies
Evidence synthesis6 of 1,450 studies met inclusion 2The base rate itself: sparse, below-standard

Two things stand out. First, the deployment and self-report rows — the ones that generate the headlines — are the weakest for causal claims. Second, the rows that can support a causal or accuracy claim are recent, few, and narrow. The evidence is not absent; it is early, and it is thinner than the confidence around these tools implies.

What did the first randomized trials show?

The most important development is that randomized evidence has finally started to arrive — and it complicates the marketing story rather than confirming it.

In the first randomized clinical trial of ambient scribes, 238 outpatient physicians across 14 specialties were assigned to one of two widely used tools (DAX Copilot or Nabla) or to a usual-care control group 3. The result was mixed. Nabla users showed a 9.5% decrease in log-measured time-in-note versus control (95% CI, -17.2% to -1.8%; p=.02), while DAX users "exhibited no significant change versus control" (-1.7%; p=.66) 3. Same category of tool, same trial, opposite documentation-time result. Any claim of the form "AI scribes cut documentation time" collapses into "this tool did, in this setting, and that one did not."

On well-being, the trial found improvements — a composite Mini-Z score rose and work exhaustion fell for scribe users — but the authors were careful to classify these as secondary endpoints, writing that the well-being findings "need confirmation in larger, multicenter trials" 3. And on the safety question, reviewers rated clinically significant inaccuracies as occurring "occasionally" for both tools (around 2.7-2.8 on a 5-point scale) 3 — low, but present, and enough to justify review of every note.

A second randomized trial points the same direction with a useful nuance. In a 24-week stepped-wedge trial of 66 practitioners across two states, ambient AI "reduced health care practitioners' work exhaustion/interpersonal disengagement but did not significantly increase professional fulfillment," while documentation time "decreased without compromising diagnosis, billing compliance, or note quality" 4. Relief from exhaustion is real and worth having; it is also a narrower claim than "makes clinicians happier," and the trial's own coprimary outcome on fulfillment did not move. The randomized evidence, in short, supports a measured, tool-dependent benefit — good news, stated precisely.

Accuracy: the question most studies skip

Notice what the effectiveness studies mostly do not measure: whether the note is correct. Time saved and burnout relieved are the outcomes that get counted; the truth of the document is usually assumed. That is a strange omission for a tool whose entire job is to produce the clinical record.

When accuracy is measured directly, the picture is sobering enough to warrant attention. A validated evaluation of an ambient scribe detected hallucinated content in 31% of AI-generated notes versus 20% of clinician-written comparison notes (p=0.01) 5 — the AI notes were fuller and also more likely to contain content the encounter did not produce. The real-world-evidence review catalogued the same distinctive failure modes: hallucinations, clinically important omissions, and misattribution — including one tool that "called the patient's wife the adult female in the room" and notes "missing some details" 2. This is the mechanism behind AI hallucination in clinical contexts, and it is the reason every published deployment keeps the workflow human-in-the-loop: the clinician reads, corrects, and signs each note. A study that reports minutes saved but never checked whether the notes were accurate has measured the easy half of the question.

The financial case is unproven

The other claim doing heavy lifting in procurement is return on investment — the idea that time saved converts to enough additional throughput or revenue to pay for the tool. Here the most candid assessment came from a health-system convening rather than a single study. A 2025 report from a health-technology assessment body concluded that ambient scribes are "potentially effective at reducing clinician documentation time and cognitive load" while flagging "gaps in evidence regarding the impact of ambient scribes on productivity and financial performance" 7. Some systems reported reduced after-hours charting; others could not tell whether it had changed at all 7. Partial uptake compounds the uncertainty: the large EHR-log study found only 32% of users adopted the tool for more than half their visits 6, so even a genuine per-visit saving spreads thin across a clinician's week. The efficiency is plausible; the financial return, on the current record, is not yet demonstrated.

The rigor gap is not unique to scribes

It helps to see this pattern in context, because it is the norm for clinical AI, not an anomaly of one product category. A systematic review of deep-learning studies across medical imaging found that "most non-randomised trials are not prospective, are at high risk of bias, and deviate from existing reporting standards," and that 61 of 81 studies still claimed the algorithm was comparable or superior to clinicians on that thin footing 8. Ambient scribes inherit that culture: rapid adoption, enthusiastic claims, and an evidence base that arrives afterward and unevenly. The corrective is the same one that field needs — randomized comparisons, prospective designs, accuracy measured alongside efficiency, and claims trimmed to what the study actually tested.

How to read this

Four cautions travel with everything above. First, "the evidence is thin" is a statement about the studies, rather than a verdict that the tools fail to help — the randomized signal is real, and clinicians' reported relief is worth taking seriously even where the clock moves less. Second, effects are tool-specific: a result for one vendor's system does not transfer to another's, as the first RCT showed within a single trial 3. Third, the numbers here come from particular samples and settings; treat any single figure as a data point in an early literature rather than a settled rate. Fourth, the evidence is moving quickly — the first RCTs landed in 2025, and this page is on a 180-day review cycle precisely because the base rate should improve.

The practical posture that follows is neither hype nor dismissal. Ambient scribes have earned a serious trial in a serious workflow, with a clinician reviewing every note, accuracy audited rather than assumed, and time and burnout measured against a real baseline rather than a memory. That is what the strong studies did, and it is what separates a claim that survives from one that does not.

Sources and method

This guide synthesises the peer-reviewed and assessment-body record on ambient scribes: a real-world-evidence rapid review that quantifies how sparse and below-standard the base is 2; the first two randomized trials, one comparing two tools against control 3 and one stepped-wedge well-being trial 4; a validated accuracy evaluation 5; a large EHR-log study of time and uptake 6; a health-system assessment on the financial question 7; a one-year deployment report for scale 1; and a systematic review of clinical-AI reporting rigor for context 8. Every figure is quoted from the primary report or paper cited beside it, never from a summary of it. We revisit this page on a 180-day cycle and whenever a larger or multicenter randomized trial reports. For the running numbers, see the AI scribe adoption statistics; for the documentation-burden and burnout literature specifically, see documentation time and burnout evidence; and for the appraisal method itself, how to read an AI validation study.

Questions & answers

  • Do AI scribes actually work?

    For documentation time, the honest answer is "sometimes, depending on the tool." The first randomized trial found one scribe reduced log-measured documentation time by 9.5% versus a control group while a second tool showed no significant change. Well-being scores improved, but the trial framed those as secondary findings that need confirmation in larger, multicenter studies. The relief clinicians report is real; the measured effects are more modest and more variable than marketing suggests.

  • How strong is the evidence base for ambient AI scribes?

    Thin relative to the deployment. A 2025 real-world-evidence review identified 1,450 studies and found only 6 met inclusion criteria — an observational study, a case report, a peer-matched cohort, and surveys — and judged that most fell short of standardized evaluation frameworks. The first randomized trials appeared only in 2025, after the tools were already in millions of encounters.

  • Are AI scribe notes accurate?

    Overall error rates run low, but accuracy is rarely measured in the studies that report time savings. When it is measured directly, gaps appear: a validated evaluation detected hallucinated content in 31% of ambient notes versus 20% of clinician-written notes. This is why every published deployment keeps a clinician reviewing each note before it is signed.

Sources

  1. The Permanente Medical Group. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi.org/10.1056/CAT.25.0040
  2. Real-World Evidence Synthesis of Digital Scribes Using Ambient Listening and Generative Artificial Intelligence for Clinician Documentation Workflows: Rapid Review. JMIR AI. 2025;4:e76743. doi.org/10.2196/76743
  3. Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI. 2025;2(12). doi.org/10.1056/AIoa2501000
  4. A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. 2025. doi.org/10.1056/AIoa2500945
  5. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence. 2025;8:1691499. doi.org/10.3389/frai.2025.1691499
  6. Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes. JAMA. 2026. doi.org/10.1001/jama.2026.2253
  7. Peterson Health Technology Institute. Leading Health Systems: AI-Powered Scribes Alleviate Clinician Burnout; Financial Impact Unclear. March 2025. phti.org/announcement/ai-scribes-reduce-clinician-burnout/
  8. Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies in medical imaging. BMJ. 2020;368:m689. doi.org/10.1136/bmj.m689
  9. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025;8(10):e2534976. doi.org/10.1001/jamanetworkopen.2025.34976