Ambient AI

The Kaiser Permanente 2.5-million-encounter ambient scribe deployment, analyzed

Two NEJM Catalyst papers document the largest ambient AI scribe rollout in the published record — 7,260 physicians and more than 2.5 million encounters in a single year. This guide reads them strictly on their own terms, separating what the deployment actually measured from what it did not, so the headline number is read as adoption evidence rather than efficacy proof. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Two NEJM Catalyst papers from The Permanente Medical Group describe the largest ambient scribe deployment in the literature: 7,260 physicians used it across 2,576,627 patient encounters between October 2023 and December 2024.
  • The reported benefit is a self-derived estimate of more than 15,700 hours of documentation time saved — about 1,794 working days — alongside small voluntary surveys in which 88% of 102 physicians reported a positive impact on visit interactions.
  • Adoption concentrated in heavy users: 3,447 physicians logged at least 100 encounters each, the top third of users accounted for most uses, and age and years since graduation did not predict adoption.
  • The only note-accuracy figure is a small internal audit — a sample of 35 AI-generated transcripts scored 48 of 50 — which is a quality signal rather than a controlled error-rate study.
  • What the papers do not provide is as important as what they do: no control group, no randomization, no patient outcomes, and no note-level error rate at scale. Read 2.5 million as evidence of uptake, not of clinical effect.

"Over 2.5 million uses" is the number everyone repeats about ambient AI scribes, and it comes from a single source: two papers published in NEJM Catalyst Innovations in Care Delivery by The Permanente Medical Group, the physician group of Kaiser Permanente's Northern California region. They are the closest thing the field has to a deployment at national scale, and they are worth reading precisely because the headline is so often quoted without its context. This guide reads the two papers strictly on their own terms — every figure below is theirs — and does the one thing the headline never does: separates what the deployment measured from what it did not. The distinction decides whether "2.5 million" is evidence of adoption or evidence of effect. As of July 2026.

For the underlying tool category, see ambient AI scribe; for the note-quality studies this deployment does not supply, see our companion guide on hallucination and omission rates.

The deployment at a glance

MetricReported valueSource
Physicians who used the scribe (year 1)7,2601
Patient encounters, Oct 2023–Dec 20242,576,6271
Physicians with ≥100 encounters each3,4471
Estimated documentation time saved>15,700 hours (~1,794 working days)1
Physician survey — positive impact on visits88% of 102 respondents1
Patient survey — comfortable with the technologytwo-thirds of 118 respondents1
Note-accuracy audit35 transcripts sampled, scored 48/501
Launch — first 10 weeks (2024 paper)3,442 physicians, 303,266 encounters2

Every row below is unpacked with the caveat the paper attaches to it.

From launch to 2.5 million: the adoption curve

The story starts with the earlier paper. In October 2023 the group enabled ambient AI technology for roughly 10,000 physicians and staff across a diverse set of settings 2. Uptake was fast: within the first 10 weeks, 3,442 physicians had used the tool across 303,266 patient encounters, with 968 physicians using it in at least 100 encounters and a single physician reaching 1,210 2. That launch paper is descriptive — it reports that ambient AI "produces high-quality clinical documentation for physicians' editing" and that use was linked with reduced documentation and EHR time 2, without a control group behind either statement.

The second paper extends the window to a full year and reports the totals now in wide circulation: 7,260 physicians and 2,576,627 encounters between October 2023 and December 2024 1. The shape of that growth matters as much as its size. Usage rose roughly linearly over the year, and it concentrated: 3,447 physicians used the scribe in at least 100 encounters each, and the top third of users by volume accounted for the majority of all uses 1. Notably, standard physician characteristics such as age and years since graduation were not associated with the likelihood of adoption 1 — a finding that cuts against the assumption that scribe uptake is a generational story about younger clinicians. The group also changed scribe vendors partway through and reported that usage held steady across the transition, with prior use predicting future use 1.

What the year taught: time, satisfaction, accuracy

Three categories of result sit on top of the usage counts.

Time. The paper's efficiency headline is an estimated saving of more than 15,700 hours of documentation time across users over the year, which it frames as roughly 1,794 working days 1. This is the figure most worth handling carefully, and the next section explains why.

Satisfaction. Two voluntary surveys anchor the experience findings. Among 102 physician respondents, 88% said the scribe had a positive impact on their visit interactions, and two-thirds reported using it five or more days a week 1. Among 118 patient respondents, two-thirds said they were comfortable with the technology, 26% were neutral, and 8% reported some level of discomfort; the paper reports that patients described a positive-to- neutral impact on the quality of their visit 1.

Accuracy. Here the deployment is deliberately thin. The single accuracy figure is a small internal audit: a sample of 35 AI-generated transcripts scored 48 out of 50 on the group's rubric 1. That is a reassuring internal signal, and it is also a 35-item spot check rather than a note-level error rate across the millions of encounters. The group's response to accuracy risk was operational — it stood up an ongoing quality process rather than treating note review as a solved problem 1, which is the honest posture given how little the deployment set out to measure about errors. For the studies that do quantify note errors, and the reason their rates range so widely, see our hallucination and omission rate guide.

What the papers measured — and what they did not

The most useful thing a reader can do with these papers is grade each claim by the evidence behind it. The deployment is genuinely strong on some axes and silent on others, and conflating the two is how "2.5 million uses" gets misread as proof that scribes work.

ClaimWhat the papers provideEvidence grade
Adoption and usageHard counts drawn from the EHR across a full yearStrong
Documentation time savedA derived estimate from observational use, no control armModerate
Physician and patient experienceSmall voluntary surveys (102 and 118 respondents)Directional
Note accuracyA 35-transcript internal auditLimited
Patient clinical outcomesNot measuredAbsent
Causal effect vs. a comparison groupNo randomization, no control groupAbsent

The counts are the deployment's real contribution. Two and a half million is a reliable, EHR-derived measure of how much the tool was used, and it settles the question of whether ambient scribes can be adopted at scale in a large integrated system: they can. The adoption curve, the heavy-user concentration, and the vendor-transition resilience are all solid observational findings.

The efficiency and experience results are weaker in kind, though not in importance. A time saving estimated from an uncontrolled rollout cannot separate the tool's effect from everything else that changed over the year — training, workflow adjustments, selection of who chose to use it. The surveys are small and voluntary, which tends to over-sample enthusiasts. None of this makes the findings wrong; it makes them directional evidence to be confirmed, rather than settled causal fact.

And the absences are real. The papers report no patient clinical outcomes, no randomized or controlled comparison, and no note-level error rate across the deployment. That is a reasonable scope for an operational learnings paper — but it means the deployment cannot, on its own, answer "do these tools improve care?" or "how often is a note wrong?" Those questions belong to other study designs.

How to read this deployment

Four cautions carry over to anyone using these papers to inform their own decision.

  1. One system, generalize with care. The results come from a single, tightly integrated group with shared infrastructure, training, and an enterprise rollout 12. A small practice on different software should expect a different adoption curve and a different support burden.
  2. Adoption is not efficacy. The 2.5-million figure measures uptake. Read it as proof that scale is achievable, and look elsewhere — controlled studies and the error-rate literature — for evidence about quality and outcomes.
  3. Treat the time saving as an estimate. More than 15,700 hours is a plausible and encouraging signal 1, derived rather than measured against a control, so weigh it as a hypothesis your own setting would need to test.
  4. The accuracy question is still open here. A 35-transcript audit scoring 48/50 1 is a spot check. Keep clinician review of every note in the workflow — the human-in-the-loop control that these papers, and every other published deployment, retain.

If you are weighing a rollout of your own, use this deployment for what it proves — that adoption at scale is real and that clinicians and patients largely welcome the tool — and pair it with the controls in our implementation checklist, the field-wide numbers in the AI scribe adoption statistics, and the regulatory framing in do ambient AI scribes need FDA regulation?.

Sources and method

This analysis rests entirely on the two primary papers: the one-year learnings paper reporting the 2.5-million-encounter total 1 and the earlier launch paper reporting the first-10-week uptake 2, both published in NEJM Catalyst Innovations in Care Delivery by The Permanente Medical Group. Every figure is drawn from those papers and cited beside the claim it supports; where a number is an estimate or an internal audit rather than a controlled measurement, we say so in the text. We add no figures from outside the two papers, and we grade each claim by the design that produced it. We revisit this page on a 180-day cycle and whenever the group publishes a follow-up or a controlled study of the same deployment. Nothing here is clinical advice.

Questions & answers

  • How many encounters did Kaiser Permanente's AI scribe handle?

    The Permanente Medical Group reported 2,576,627 patient encounters assisted by ambient AI scribes between October 2023 and December 2024, used by 7,260 physicians. It is the largest ambient scribe deployment described in the peer-reviewed literature as of July 2026.

  • Did the deployment prove AI scribes save time?

    The paper reports an estimated saving of more than 15,700 hours of documentation time — about 1,794 working days — but this is a derived estimate from an observational rollout with no control group, rather than a randomized measurement. It is strong evidence of adoption and a plausible efficiency signal, weaker evidence of a causal effect.

  • What did the deployment measure about note accuracy?

    Very little, by design. The only accuracy figure is a small internal audit in which a sample of 35 AI-generated transcripts scored 48 out of 50 on the group's rubric. There is no note-level error rate across the 2.5 million encounters, which is why note-quality questions are better answered by the dedicated error-rate studies.

Sources

  1. The Permanente Medical Group. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi.org/10.1056/CAT.25.0040
  2. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. 2024;5(3). doi.org/10.1056/CAT.23.0404