011 min
Hallucination at a Glance
- What it is — fluent, confident output that is factually wrong or unsupported by the source.
- Why it exists — models are trained to predict plausible next words, not to check facts.
- What it costs — a New York court fined a law firm $5,000 after ChatGPT invented six fake case citations.
- When it breaks — worst on specific, checkable details: names, dates, citations, quotes, and URLs.
- As of 2026 — still an open problem across every major model family, not something fine-tuning alone has solved.
021 min
The Problem Hallucination Names
Before the term "hallucination" was applied to language models, the naive expectation was that a model trained on enormous amounts of accurate text would simply reproduce accurate information back, the way a search engine returns an existing document. That expectation fails because a language model is not retrieving stored documents; researchers at the Hong Kong University of Science and Technology, in a widely cited survey, describe how models often generate text that is nonsensical, or unfaithful to the provided source input. The failure is specific and dangerous precisely because it doesn't look like a failure: a hallucinated answer is written in exactly the same fluent, confident register as a correct one, with no hedge, no uncertainty marker, and often no way to tell it apart from a true statement without checking an outside source.
The term itself was borrowed from computer vision and machine translation research, where it already described a model inserting objects or words that were never present in an image or source sentence, well before large language models existed. NLP researchers adopted the same word once generative text models began producing invented details with that same unearned confidence.
032 min
How Hallucination Works
Researchers categorize hallucinations into two kinds, walked through with the same running example. Intrinsic hallucination is generated output that contradicts the source content — asked to summarize a document that says "the first vaccine for Ebola was approved by the FDA in 2019," an intrinsically hallucinating model might produce "the first Ebola vaccine was approved in 2021," a specific, checkable contradiction of the source it was given. Extrinsic hallucination is generated output that cannot be verified from the source content at all — asked the same summarization question, a model might add a detail like a country starting clinical trials that the source document never mentions, which is not directly contradicted, just unsupported.
The mechanism traces back to what a language model is actually trained to do: predict the most statistically likely next word given everything before it, not check whether a specific factual claim is true. When a model is asked a question it has weak or no real information about — an obscure legal citation, a niche statistic, a URL — the training objective still rewards producing fluent, plausible-sounding text, so the model generates something that has the right shape for an answer (a case name, a page number, a citation format) without the process having any mechanism to check whether that specific shape corresponds to something real. The model is not lying in the human sense of knowing the truth and saying otherwise; it has no internal representation of "I don't actually know this" that reliably overrides the pressure to produce fluent output.
041 min
A Concrete Example
TruthfulQA, a benchmark of 817 questions across 38 categories including health, law, finance, and politics, was built specifically to catch this failure. Its authors crafted questions that some humans would answer falsely due to a false belief or misconception, to test whether models repeat popular misconceptions rather than the truth. Testing GPT-3, GPT-Neo/J, GPT-2, and a T5-based model, they found the best model was truthful on 58% of questions, while human performance was 94%. The gap is the measurable size of the hallucination problem on questions specifically designed to tempt a model toward a fluent, popular, wrong answer over an accurate one — and the paper's authors noted the largest models were generally the least truthful, the opposite pattern from most other NLP benchmarks, where bigger models do better.
051 min
What This Means for Your Work
For an engineer shipping an AI feature, hallucination changes what counts as an acceptable design: any feature that presents model output as a fact rather than a suggestion — a citation, a phone number, a code API that may not exist — needs either a retrieval step that grounds the answer in a real source, or a visible signal to the user that the output is unverified. Treating a model's confident tone as evidence of accuracy is exactly the mistake that produces a shipped feature that quietly fabricates details nobody catches until a user does.
For a founder or PM deciding where to deploy a model, hallucination changes which use cases are safe to launch without a human check: a brainstorming tool where a wrong idea costs nothing tolerates a much higher hallucination rate than a legal-research tool, a medical-symptom checker, or anything that generates a citation a user might submit somewhere official.
For a data or eval team, the practical decision is what an eval set actually measures: an eval built only from easy, well-known questions will not catch a model's tendency to hallucinate on the specific, checkable details — names, dates, exact figures — where the failure is worst and least visible to a casual reviewer.
061 min
What Hallucination Costs
The clearest documented cost of a hallucination reaching production is legal, not computational. In a 2023 case in the Southern District of New York, a plaintiff's lawyers used ChatGPT to generate a legal motion, which contained numerous fake legal cases involving fictitious airlines with fabricated quotations, submitted as real precedent opposing a motion to dismiss. Judge P. Kevin Castel dismissed the underlying personal injury case against Avianca and ordered the plaintiff's attorneys to pay a $5,000 fine, in a decision that has since been frequently cited by courts in cases where usage of generative AI during the course of proceedings leads to the creation and citation of nonexistent caselaw.
A single hallucinated fact in a high-visibility setting can also cost real money quickly: after Google's Bard chatbot claimed in a February 2023 promotional demo that the James Webb Space Telescope took the very first pictures of a planet outside our own solar system, a claim that contradicted the well-documented 2004 discovery of exoplanet 2M1207b, about 10 percent of Alphabet's market value, some $120 billion, was wiped out that week.
071 min
What Hallucination Does Not Solve
Hallucination is not solved by asking a model to "be careful" or "only state facts it is sure of," because the model has no reliable internal signal distinguishing a memorized fact from a plausible-sounding guess; both are generated by the same next-word-prediction process. Detecting it in production generally means checking model output against an independent source, not trusting the model's own confidence.
Two approaches reduce the rate rather than eliminate it. Retrieval-augmented generation gives the model real source documents to draw from at answer time, turning an open-ended "recall this fact from training" task into a narrower "summarize what's actually in front of you" task, which measurably reduces, though does not eliminate, both intrinsic and extrinsic hallucination. The second is citation verification as a separate step after generation: checking that a claimed URL resolves, that a claimed quote actually appears on the cited page, and that a claimed study exists under the name and year given, before any of it reaches a user — treating the model's own citation as a claim to verify, never as a fact already checked.
Neither approach is a complete fix as of 2026: retrieval reduces the rate of fabricated facts but can still be misread, and citation verification catches a bad citation only if the verification step itself is actually run before publishing, not treated as optional.
081 min
Hallucination vs. Nearby Concepts
The nearest neighbor is retrieval-augmented generation, and the relationship is closer to cause-and-fix than to confusable near-synonyms: RAG is one of the main techniques used specifically to reduce hallucination, by grounding a model's answer in retrieved documents rather than relying entirely on what it memorized during training. A model can still hallucinate with RAG in place, either by misreading the retrieved document (an intrinsic hallucination against the retrieved source instead of against training data) or by adding an unsupported detail the retrieved document never mentioned, so RAG lowers the rate without ever fully closing the gap.
A second neighbor is plain model error, such as a large language model doing arithmetic wrong or misreading an instruction. The distinguishing fact is confidence and shape: an arithmetic error is usually a computation mistake the model would likely get right with a different approach, while a hallucination is a specific, fluent, checkable factual claim, a name, a date, a URL, a quote, presented with no internal signal that it might be fabricated.
091 min
A Second Case
The Avianca court case and the Bard demo show hallucination in two very different failure shapes. The Avianca case was intrinsic: the model was asked to find real cases and instead invented ones that did not exist at all, wholesale fabrication with no real source behind it. The Bard incident was closer to a confidently wrong recall of a real fact: the model conflated "first image captured by JWST specifically" with "first image of an exoplanet, period," producing a false claim that sounded exactly like the true, more modest claim it should have made instead.
The two cases also differ in how the error was caught. The Avianca fabrications were caught only because opposing counsel could not find the cited cases in any legal database, a slow, manual verification process. The Bard error was caught almost immediately because it was checkable against a well-known, widely reported fact in a scientific field with many watchful outside experts, showing that how quickly a hallucination is caught depends heavily on how easy the specific claim is to check, not on how confidently it was stated.
101 min
Where the Evidence Is Contested
Some researchers push back on treating every unsupported or unverifiable output as a failure. The same survey that defines extrinsic hallucination notes that such hallucination is not always erroneous because it could be from factually correct external information, meaning a model adding a true detail that simply wasn't in the specific source document it was given is technically an extrinsic hallucination by definition, even though the added information is accurate. Under a strict definition, a model that correctly adds outside knowledge and a model that fabricates a fake citation are both labeled the same way, which several researchers argue conflates a useful behavior with a dangerous one.
A separate disagreement concerns whether hallucination should be treated as a bug to be eliminated or as an unavoidable statistical consequence of how these models generate text at all. TruthfulQA's own finding that larger models were generally less truthful than smaller ones on its benchmark complicated the assumption that hallucination is simply a capability gap that scaling up model size would eventually close on its own.
?6 questions
Questions people ask
Why do language models hallucinate?
Is hallucination the same as a language model making a mistake?
What are the risks or limits of relying on model output without checking it?
Does retrieval-augmented generation eliminate hallucination?
How much does hallucination cost in practice?
Do bigger, more capable models hallucinate less?
§4 sources
Sources
Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12), 1–38.
Lin, S., Hilton, J., & Evans, O. (2021). TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958.
Mata v. Avianca, Inc. — Wikipedia, for the court's findings and sanction.
Alphabet shares dumped after Google AI chatbot flub — The Register, for the Bard JWST incident and its market impact.





