An AI hallucination is output that an artificial-intelligence system presents as fact but that has no basis in its input or in reality. In clinical contexts, it means fabricated or distorted clinical detail — a symptom, dosage, finding, or citation — delivered with the same fluent confidence as accurate content.
Why does hallucination matter in healthcare?
Fabricated content in most industries costs time. In a clinical record it can redirect care: an invented allergy, a dosage that was never stated, a negative finding recorded as positive. The World Health Organization's guidance on large multi-modal models warns that these systems may produce false, inaccurate, biased, or incomplete statements that could harm people using them to make health decisions 3. The danger is compounded by fluency — a hallucinated sentence reads exactly as smoothly as a correct one.
What do the measured rates show?
The most rigorous measurement to date comes from a 2025 clinician-annotation framework study covering 12,999 sentences of LLM-generated clinical documentation. It found a 1.47% hallucination rate and a 3.45% omission rate 1. Two details in that result matter more than the headline percentages. First, 44% of the hallucinations were rated major — judged capable of affecting patient care 1. Second, omissions outnumbered hallucinations by more than two to one: the quieter failure mode, leaving out something clinically important, occurs more often than inventing something.
Fabrication also shows a strong model-generation effect. In a controlled test of 636 bibliographic citations, 55% of GPT-3.5's citations were fabricated outright, against 18% for GPT-4 4. Improvement is real; zero is nowhere in the published record.
Where does hallucination appear today?
The setting where hallucination is most closely watched, as of July 2026, is the ambient AI scribe — software that drafts the clinical note from the recorded encounter. A real-world evidence synthesis found these tools report low overall error rates but carry a distinctive triad of failure modes: hallucination, clinically important omission, and misattribution — content assigned to the wrong speaker 2. The largest reported deployment, which passed 2.5 million uses in its first year, responded by standing up a formal quality-assurance program for note accuracy at scale 5 — the clearest signal that operators treat hallucination as a permanent operating risk rather than a launch-phase bug. The adoption numbers behind that deployment are tracked in our AI scribe adoption statistics.
Common misunderstandings
A low rate means low risk. A 1.47% sentence-level rate across millions of encounters is a large absolute number of fabricated sentences, and nearly half were rated major 1. Volume converts small rates into steady exposure.
Hallucination equals mishearing. Transcription errors distort something said; hallucinations invent something unsaid. The review controls differ — checking against the audio catches the first, while catching the second requires the clinician's own memory of the encounter 2.
Better models make review optional. Every published clinical deployment keeps a human in the loop — a clinician who reviews and signs each note before it enters the record 2. Falling rates change the workload of review, and leave the requirement in place.
Related terms
An ambient AI scribe is where clinicians most often meet hallucination in daily work. A clinical LLM is the model class that produces it. Human-in-the-loop design is the standard control. For the deployment numbers that give these failure modes their denominator, see the AI scribe adoption statistics.