Ambient AI

Hallucination and omission rates in AI scribe notes: what the studies measured

Published "hallucination rates" for ambient scribe notes range from about 1.5% to 31% — a spread that looks like disagreement and is really a difference of units. This guide keys each number to what it counted, pairs every hallucination figure with its omission counterpart, and gives you one table to compare them honestly. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Published hallucination figures for AI scribe notes span roughly 1.5% to 31% — they disagree mainly because they count different units: a share of sentences, a share of notes, or errors per simulated case. Compare like with like before trusting any one number.
  • The most granular clinician-annotated measurement found a 1.47% hallucination rate and a 3.45% omission rate across 12,999 note sentences: omission is the more frequent failure, hallucination the more often harmful, with 44% of hallucinations rated major.
  • At the note level, one evaluation found at least one hallucination in 31% of ambient notes versus 20% of clinician-written notes; in a five-platform simulation the mean note error rate was 26.3%, and omissions were about three-quarters of all errors.
  • Errors carry real risk: the simulation found an average of 3.0 errors per case with potential for moderate-to-severe harm, and hallucinations cluster in the Plan section, where instructions live.
  • Rates are a moving target tied to the model and the review step — fabrication fell from 55% to 18% across two model generations, and every published deployment still keeps a clinician reviewing each note.

Ask "how often does an AI scribe hallucinate?" and the published answer is a range so wide it looks broken: from around 1.5% at one end to 31% at the other. Both numbers are real, peer-reviewed, and correct. They differ because they count different things — one measures the share of sentences that are fabricated, another the share of notes that contain at least one fabrication, a third the errors found per simulated case. Set them side by side without that key and you will conclude the field disagrees with itself. Line them up by unit and the picture is coherent, sobering, and useful. This guide supplies the key, pairs every hallucination figure with its omission counterpart — the quieter and more common failure — and ends with how to read any new number you meet. As of July 2026.

For the underlying definition of the failure mode itself, see AI hallucination in clinical contexts; for the tool category these rates describe, see ambient AI scribe.

The numbers that look like they disagree

StudyWhat it measuredUnitHallucination resultOmission result
Clinician annotation 1Errors marked in real generated notesPer sentence1.47% of sentences; 44% rated major3.45% of sentences; 16.7% rated major
Note-level evaluation 2Whether a note carries any hallucinationPer note (binary)31% of AI notes vs 20% of clinician notesReported via note "thoroughness", never as a rate
Five-platform simulation 3All errors in notes from scripted visitsPer casePart of a 26.3% mean note error rate~76.3% of all errors were omissions
Real-world audit 8Internal accuracy score on sampled notesPer sampled transcript35 transcripts scored 48/50 overallNot separately reported

Four studies, four units, one phenomenon. A "1.47%" and a "31%" are not in conflict any more than "3% of days are rainy" conflicts with "half of weeks contain a rainy day." The rest of this page walks each row and explains what its unit can and cannot tell you.

Sentence level: what is the base rate?

The most granular measurement to date comes from a clinician-in-the-loop annotation framework applied to 12,999 sentences of generated clinical documentation 1. Every sentence was reviewed against the source encounter and labelled. The result: a 1.47% hallucination rate and a 3.45% omission rate 1. Two findings inside that number matter more than the headline. First, 44% of the hallucinations were rated major, against only 16.7% of omissions — so the rarer failure is disproportionately the dangerous one 1. Second, hallucinations were most commonly (20%) located in the Plan section of the note 1, the part that carries instructions and actions for colleagues and patients, where a fabrication does the most downstream work.

A sentence-level rate is the right denominator when you want to know how "dense" errors are in the text a clinician has to review. It is the wrong one for estimating how many notes are affected — because errors are not spread evenly, a low per-sentence rate can still leave a large fraction of notes carrying at least one error. That is exactly what the next unit measures.

Note level: how often does a note carry at least one?

A separate validated evaluation scored ambient notes against clinician-written "gold" notes across 97 encounters and 194 notes, with specialists judging each on a documentation-quality instrument 2. Its headline is a note-level presence rate: hallucinations were detected in 31% of ambient notes versus 20% of gold notes (p = 0.01) 2. The same evaluation found the two note types close on overall quality — ambient notes even scored higher on thoroughness and organization — while clinician notes scored better on accuracy, succinctness, and internal consistency 2.

Read this row carefully. A 31% note-level figure and a 1.47% sentence-level figure are describing the same kind of event at different resolutions. The note-level number is the one that answers "if I pick up a note, how likely is it to contain a fabrication somewhere?" — and the answer, roughly one in three, is the number that should set your review posture. It is also the number most often quoted out of context as "AI scribes hallucinate 31% of the time," which is a misreading: it is 31% of notes, not 31% of content.

Per case in simulation: stress-testing the tools

The widest error figures come from a study that deliberately stress-tested the tools rather than sampling routine use. Researchers ran 14 scripted ambulatory cases through five ambient digital scribe platforms and counted every error 3. Across platforms, the mean clinical-note error rate was 26.3% (95% CI, 17.0%-31.0%), and omissions made up roughly three-quarters of all errors 3. The safety-relevant figure is starker: an average of 3.0 errors per case (95% CI, 0-4; range 0-21) had the potential to cause moderate-to-severe harm 3. The study also traced how transcript mistakes propagate — 19.5% of transcript errors were transmitted into the final note 3 — a reminder that the note inherits the microphone's mistakes as well as the model's.

Simulation is a magnifier rather than a census. Scripted cases can be built to include the ambiguous, information-dense moments that expose a tool, so these rates sit above what routine visits produce. That is a feature: a simulation tells you the failure surface — where and how a tool breaks — in a way a favorable real-world average never will. Use it to understand modes, not to forecast your own baseline.

Omission is the failure you were not looking for

Across every study that measured both, the same asymmetry appears: omissions outnumber hallucinations. The clinician annotation put omission at more than double the hallucination rate 1; the simulation found omissions were about 76% of all errors 3. A rapid review of the field names the pattern directly, cataloguing a scribe-specific triad of "hallucinations, critical omissions, and misattribution" 4 — the last being content assigned to the wrong speaker, a failure mode dictation never had.

Omission is harder to catch precisely because there is nothing on the page to flag. A hallucinated allergy is a wrong sentence a careful reviewer can strike; a missing allergy is an absence, visible only to a clinician who remembers the encounter well enough to notice what the note left out. This is why omission is underweighted in casual discussion of "AI errors" and why the review control for it is different — catching a fabrication means checking the note against the record, while catching an omission means checking the note against your own memory of the visit.

Why the numbers move: model generation and the review step

Any error rate you read is a snapshot of one model at one moment. Fabrication in particular falls sharply as models improve: in a controlled test of generated bibliographic citations, 55% were fabricated under one model generation and 18% under the next 6. Newer models invent less — and zero appears nowhere in the published record. The World Health Organization's guidance on large multi-modal models states the standing risk plainly: these systems "may generate outputs that are false, inaccurate, biased or incomplete" 7, which is why measurement is a recurring discipline rather than a one-time clearance.

The second moving part is the human review that sits after the model. Every published deployment pairs the tool with a human in the loop — a clinician who reads, edits, and signs each note. The largest reported deployment audited a sample of 35 of its own AI-generated transcripts and scored them 48 out of 50 on an internal rubric 8; that is an encouraging internal signal, though a small audit rather than a controlled error-rate study, and we read it in full in our analysis of that 2.5-million-encounter deployment. Meanwhile a blinded evaluation of 11 tools against human note-takers found human notes scored higher on quality across all five standardized cases 5 — a caution against assuming the review step is a formality the tools have already made redundant.

How to read these rates

Five cautions travel with any scribe error number you encounter.

  1. Check the unit first. Per-sentence, per-note, and per-case rates answer different questions and cannot be compared directly. A study reporting no unit is reporting a number you cannot use.
  2. Ask for the omission rate too. A tool advertised on a low hallucination figure may have a higher, quieter omission rate — the failure that leaves no mark on the page 13.
  3. Separate simulation from routine use. Stress-test figures like the 26.3% note error rate map the failure surface; they sit above the rates a typical clinic will see 3.
  4. Watch the denominator and the case mix. A study run on information-dense or ambiguous encounters will report higher rates than one on simple visits; the sample's difficulty is part of the result.
  5. Treat every figure as vendor- and model-specific and perishable. Rates move with the model version, the microphone setup, the specialty, and the prompt — a number measured on one product last year does not describe yours today.

For anyone evaluating a tool against these numbers, the reported rate is a starting point, never a substitute for your own governance: confirm any vendor-quoted accuracy claim against an independent measurement, and keep clinician review in the workflow regardless of how low the published rate looks. The controls that actually protect patients live in the rollout, covered in our implementation checklist for medical groups, and the question of whether these tools face any external oversight is taken up in do ambient AI scribes need FDA regulation?.

Sources and method

This guide synthesises the published error-rate studies for ambient scribe and clinical-summarisation output, each cited beside the figure it produced: a sentence-level clinician-annotation framework 1, a note-level validated evaluation against clinician notes 2, a five-platform simulated-encounter comparison 3, a rapid review of failure modes 4, a blinded multi-tool quality evaluation 5, a controlled citation-fabrication test across model generations 6, and the largest reported real-world deployment 8, read under the WHO's governance frame for large multi-modal models 7. Every percentage is quoted from its primary source, never from a summary of it, and each rate is labelled with the unit it was measured in. We revisit this page on a 180-day cycle and whenever a study re-measures these rates on a common standard. For the deployment numbers that give these rates their denominator, see the AI scribe adoption statistics; for the model class that produces the output, see clinical LLM. Nothing here is clinical advice.

Questions & answers

  • What is the hallucination rate of AI scribe notes?

    It depends on what you count. The most granular clinician-annotated study found 1.47% of sentences were hallucinated; a note-level study found at least one hallucination in 31% of notes; a five-platform simulation found a 26.3% mean note error rate. Same phenomenon, three different units — which is why a single headline number misleads more than it informs.

  • Are omissions or hallucinations more common in AI notes?

    Omissions. The clinician-annotation study measured a 3.45% omission rate against a 1.47% hallucination rate, and a five-platform simulation found omissions made up roughly three-quarters of all errors. Hallucinations are rarer but more often rated major — capable of affecting care if left uncorrected.

  • Do AI scribe errors actually reach patients?

    The controls are designed to catch them — every published deployment keeps a clinician reviewing and signing each note. But a simulation found an average of 3.0 errors per case with potential for moderate-to-severe harm, so the review step is load-bearing rather than optional. Treat the reported rates as the workload of review rather than a measure of what reaches the chart.

Sources

  1. Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8:274. doi.org/10.1038/s41746-025-01670-7
  2. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence. 2025;8:1691499. doi.org/10.3389/frai.2025.1691499
  3. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi.org/10.1016/j.mcpdig.2025.100292
  4. Real-World Evidence Synthesis of Digital Scribes Using Ambient Listening and Generative Artificial Intelligence for Clinician Documentation Workflows: Rapid Review. JMIR AI. 2025;4:e76743. doi.org/10.2196/76743
  5. Rapid Evaluation of Artificial Intelligence Technology Used for Ambient Dictation in Primary Care: Comparing the Quality of Documentation of Artificial Intelligence-Generated and Human-Produced Clinical Notes. Annals of Internal Medicine. 2026. doi.org/10.7326/annals-25-02772
  6. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports. 2023;13:14045. doi.org/10.1038/s41598-023-41032-5
  7. World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. Geneva: WHO; 18 January 2024. www.who.int/publications/i/item/9789240084759
  8. The Permanente Medical Group. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi.org/10.1056/CAT.25.0040