Glossary

AI hallucination in clinical contexts

What it means when a clinical AI system fabricates content — the measured rates, the severity splits, and the failure modes specific to healthcare. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • An AI hallucination is output a model presents as fact that has no basis in its input or in reality — in clinical settings, a fabricated symptom, dosage, finding, or citation.
  • Measured rates in clinical documentation are low but consequential: a 2025 clinician-annotation study found a 1.47% hallucination rate — and rated 44% of those hallucinations major.
  • Omission travels with hallucination: the same study measured a 3.45% omission rate, more than double the hallucination rate.
  • WHO guidance warns large multi-modal models may produce false, inaccurate, biased, or incomplete statements that could harm health decisions.
  • Clinician review of every AI-drafted output remains the standard control in published deployments, as of July 2026.

An AI hallucination is output that an artificial-intelligence system presents as fact but that has no basis in its input or in reality. In clinical contexts, it means fabricated or distorted clinical detail — a symptom, dosage, finding, or citation — delivered with the same fluent confidence as accurate content.

Why does hallucination matter in healthcare?

Fabricated content in most industries costs time. In a clinical record it can redirect care: an invented allergy, a dosage that was never stated, a negative finding recorded as positive. The World Health Organization's guidance on large multi-modal models warns that these systems may produce false, inaccurate, biased, or incomplete statements that could harm people using them to make health decisions 3. The danger is compounded by fluency — a hallucinated sentence reads exactly as smoothly as a correct one.

What do the measured rates show?

The most rigorous measurement to date comes from a 2025 clinician-annotation framework study covering 12,999 sentences of LLM-generated clinical documentation. It found a 1.47% hallucination rate and a 3.45% omission rate 1. Two details in that result matter more than the headline percentages. First, 44% of the hallucinations were rated major — judged capable of affecting patient care 1. Second, omissions outnumbered hallucinations by more than two to one: the quieter failure mode, leaving out something clinically important, occurs more often than inventing something.

Fabrication also shows a strong model-generation effect. In a controlled test of 636 bibliographic citations, 55% of GPT-3.5's citations were fabricated outright, against 18% for GPT-4 4. Improvement is real; zero is nowhere in the published record.

Where does hallucination appear today?

The setting where hallucination is most closely watched, as of July 2026, is the ambient AI scribe — software that drafts the clinical note from the recorded encounter. A real-world evidence synthesis found these tools report low overall error rates but carry a distinctive triad of failure modes: hallucination, clinically important omission, and misattribution — content assigned to the wrong speaker 2. The largest reported deployment, which passed 2.5 million uses in its first year, responded by standing up a formal quality-assurance program for note accuracy at scale 5 — the clearest signal that operators treat hallucination as a permanent operating risk rather than a launch-phase bug. The adoption numbers behind that deployment are tracked in our AI scribe adoption statistics.

Common misunderstandings

A low rate means low risk. A 1.47% sentence-level rate across millions of encounters is a large absolute number of fabricated sentences, and nearly half were rated major 1. Volume converts small rates into steady exposure.

Hallucination equals mishearing. Transcription errors distort something said; hallucinations invent something unsaid. The review controls differ — checking against the audio catches the first, while catching the second requires the clinician's own memory of the encounter 2.

Better models make review optional. Every published clinical deployment keeps a human in the loop — a clinician who reviews and signs each note before it enters the record 2. Falling rates change the workload of review, and leave the requirement in place.

Related terms

An ambient AI scribe is where clinicians most often meet hallucination in daily work. A clinical LLM is the model class that produces it. Human-in-the-loop design is the standard control. For the deployment numbers that give these failure modes their denominator, see the AI scribe adoption statistics.

Questions & answers

  • How often do clinical AI systems hallucinate?

    The best clinician-annotated measurement to date, across 12,999 sentences of LLM-generated clinical documentation, found a 1.47% hallucination rate and a 3.45% omission rate. Low percentages, but 44% of the hallucinations were rated major — capable of affecting patient care.

  • Is hallucination the same as a transcription error?

    No. A transcription error mishears something that was said. A hallucination invents content that was never said — a symptom denied, a dosage never stated, a citation that does not exist — and delivers it with full fluency, which makes it harder to catch on review.

  • Can hallucination be eliminated?

    As of July 2026, no published system claims zero hallucination. Newer models fabricate less — citation-fabrication rates fell from 55% to 18% between successive model generations in one controlled test — but every published clinical deployment still pairs the model with clinician review of each output.

Sources

  1. Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8:274. doi.org/10.1038/s41746-025-01670-7
  2. Real-World Evidence Synthesis of Digital Scribes Using Ambient Listening and Generative Artificial Intelligence for Clinician Documentation Workflows: Rapid Review. JMIR AI. 2025;4:e76743. doi.org/10.2196/76743
  3. World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. Geneva: WHO; 18 January 2024. www.who.int/publications/i/item/9789240084759
  4. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports. 2023;13:14045. doi.org/10.1038/s41598-023-41032-5
  5. The Permanente Medical Group. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi.org/10.1056/CAT.25.0040