"Over 2.5 million uses" is the number everyone repeats about ambient AI scribes, and it comes from a single source: two papers published in NEJM Catalyst Innovations in Care Delivery by The Permanente Medical Group, the physician group of Kaiser Permanente's Northern California region. They are the closest thing the field has to a deployment at national scale, and they are worth reading precisely because the headline is so often quoted without its context. This guide reads the two papers strictly on their own terms — every figure below is theirs — and does the one thing the headline never does: separates what the deployment measured from what it did not. The distinction decides whether "2.5 million" is evidence of adoption or evidence of effect. As of July 2026.
For the underlying tool category, see ambient AI scribe; for the note-quality studies this deployment does not supply, see our companion guide on hallucination and omission rates.
The deployment at a glance
| Metric | Reported value | Source |
|---|---|---|
| Physicians who used the scribe (year 1) | 7,260 | 1 |
| Patient encounters, Oct 2023–Dec 2024 | 2,576,627 | 1 |
| Physicians with ≥100 encounters each | 3,447 | 1 |
| Estimated documentation time saved | >15,700 hours (~1,794 working days) | 1 |
| Physician survey — positive impact on visits | 88% of 102 respondents | 1 |
| Patient survey — comfortable with the technology | two-thirds of 118 respondents | 1 |
| Note-accuracy audit | 35 transcripts sampled, scored 48/50 | 1 |
| Launch — first 10 weeks (2024 paper) | 3,442 physicians, 303,266 encounters | 2 |
Every row below is unpacked with the caveat the paper attaches to it.
From launch to 2.5 million: the adoption curve
The story starts with the earlier paper. In October 2023 the group enabled ambient AI technology for roughly 10,000 physicians and staff across a diverse set of settings 2. Uptake was fast: within the first 10 weeks, 3,442 physicians had used the tool across 303,266 patient encounters, with 968 physicians using it in at least 100 encounters and a single physician reaching 1,210 2. That launch paper is descriptive — it reports that ambient AI "produces high-quality clinical documentation for physicians' editing" and that use was linked with reduced documentation and EHR time 2, without a control group behind either statement.
The second paper extends the window to a full year and reports the totals now in wide circulation: 7,260 physicians and 2,576,627 encounters between October 2023 and December 2024 1. The shape of that growth matters as much as its size. Usage rose roughly linearly over the year, and it concentrated: 3,447 physicians used the scribe in at least 100 encounters each, and the top third of users by volume accounted for the majority of all uses 1. Notably, standard physician characteristics such as age and years since graduation were not associated with the likelihood of adoption 1 — a finding that cuts against the assumption that scribe uptake is a generational story about younger clinicians. The group also changed scribe vendors partway through and reported that usage held steady across the transition, with prior use predicting future use 1.
What the year taught: time, satisfaction, accuracy
Three categories of result sit on top of the usage counts.
Time. The paper's efficiency headline is an estimated saving of more than 15,700 hours of documentation time across users over the year, which it frames as roughly 1,794 working days 1. This is the figure most worth handling carefully, and the next section explains why.
Satisfaction. Two voluntary surveys anchor the experience findings. Among 102 physician respondents, 88% said the scribe had a positive impact on their visit interactions, and two-thirds reported using it five or more days a week 1. Among 118 patient respondents, two-thirds said they were comfortable with the technology, 26% were neutral, and 8% reported some level of discomfort; the paper reports that patients described a positive-to- neutral impact on the quality of their visit 1.
Accuracy. Here the deployment is deliberately thin. The single accuracy figure is a small internal audit: a sample of 35 AI-generated transcripts scored 48 out of 50 on the group's rubric 1. That is a reassuring internal signal, and it is also a 35-item spot check rather than a note-level error rate across the millions of encounters. The group's response to accuracy risk was operational — it stood up an ongoing quality process rather than treating note review as a solved problem 1, which is the honest posture given how little the deployment set out to measure about errors. For the studies that do quantify note errors, and the reason their rates range so widely, see our hallucination and omission rate guide.
What the papers measured — and what they did not
The most useful thing a reader can do with these papers is grade each claim by the evidence behind it. The deployment is genuinely strong on some axes and silent on others, and conflating the two is how "2.5 million uses" gets misread as proof that scribes work.
| Claim | What the papers provide | Evidence grade |
|---|---|---|
| Adoption and usage | Hard counts drawn from the EHR across a full year | Strong |
| Documentation time saved | A derived estimate from observational use, no control arm | Moderate |
| Physician and patient experience | Small voluntary surveys (102 and 118 respondents) | Directional |
| Note accuracy | A 35-transcript internal audit | Limited |
| Patient clinical outcomes | Not measured | Absent |
| Causal effect vs. a comparison group | No randomization, no control group | Absent |
The counts are the deployment's real contribution. Two and a half million is a reliable, EHR-derived measure of how much the tool was used, and it settles the question of whether ambient scribes can be adopted at scale in a large integrated system: they can. The adoption curve, the heavy-user concentration, and the vendor-transition resilience are all solid observational findings.
The efficiency and experience results are weaker in kind, though not in importance. A time saving estimated from an uncontrolled rollout cannot separate the tool's effect from everything else that changed over the year — training, workflow adjustments, selection of who chose to use it. The surveys are small and voluntary, which tends to over-sample enthusiasts. None of this makes the findings wrong; it makes them directional evidence to be confirmed, rather than settled causal fact.
And the absences are real. The papers report no patient clinical outcomes, no randomized or controlled comparison, and no note-level error rate across the deployment. That is a reasonable scope for an operational learnings paper — but it means the deployment cannot, on its own, answer "do these tools improve care?" or "how often is a note wrong?" Those questions belong to other study designs.
How to read this deployment
Four cautions carry over to anyone using these papers to inform their own decision.
- One system, generalize with care. The results come from a single, tightly integrated group with shared infrastructure, training, and an enterprise rollout 12. A small practice on different software should expect a different adoption curve and a different support burden.
- Adoption is not efficacy. The 2.5-million figure measures uptake. Read it as proof that scale is achievable, and look elsewhere — controlled studies and the error-rate literature — for evidence about quality and outcomes.
- Treat the time saving as an estimate. More than 15,700 hours is a plausible and encouraging signal 1, derived rather than measured against a control, so weigh it as a hypothesis your own setting would need to test.
- The accuracy question is still open here. A 35-transcript audit scoring 48/50 1 is a spot check. Keep clinician review of every note in the workflow — the human-in-the-loop control that these papers, and every other published deployment, retain.
If you are weighing a rollout of your own, use this deployment for what it proves — that adoption at scale is real and that clinicians and patients largely welcome the tool — and pair it with the controls in our implementation checklist, the field-wide numbers in the AI scribe adoption statistics, and the regulatory framing in do ambient AI scribes need FDA regulation?.
Sources and method
This analysis rests entirely on the two primary papers: the one-year learnings paper reporting the 2.5-million-encounter total 1 and the earlier launch paper reporting the first-10-week uptake 2, both published in NEJM Catalyst Innovations in Care Delivery by The Permanente Medical Group. Every figure is drawn from those papers and cited beside the claim it supports; where a number is an estimate or an internal audit rather than a controlled measurement, we say so in the text. We add no figures from outside the two papers, and we grade each claim by the design that produced it. We revisit this page on a 180-day cycle and whenever the group publishes a follow-up or a controlled study of the same deployment. Nothing here is clinical advice.