AI Hallucinations: Why Models Make Things Up, and Which Controls Actually Reduce It
An AI hallucination is output from a language model that is fluent, confident and false. It is not a malfunction. It is the normal operation of a system that produces the most plausible answer and has no built-in signal separating plausible from true. That framing matters because it changes what executives should ask for. "Keep a human in the loop" is the standard advice and it is incomplete: it does not say which errors the human is looking for, how many outputs they can check, or what happens at volume when they stop looking carefully. The evidence points to a more useful approach. Hallucinations come in three distinct kinds, each caught by different controls; those controls differ by an order of magnitude in cost; and the right combination for a use case follows from two questions, how expensive a wrong answer is and whether anyone checks the output before it matters.

Key Takeaways
- Fluency and truth come from the same mechanism. A language model generates the most plausible continuation of its input, so a well-formed false answer and a well-formed true one are produced the same way, and the model has no internal alarm that goes off when it is guessing.
- There are three kinds of hallucination: fabrication of facts, faithfulness failures where the model contradicts a source it was given, and reasoning errors. Each needs a different control, and a program that treats them as one problem will under-invest in at least one.
- The controls that do the most work are structural rather than instructional: grounding answers in retrieved sources with citations, constraining output to a schema or closed set, allowing and rewarding abstention, and validating outputs against a system of record.
- Three popular fixes work less well than claimed: telling the model not to hallucinate, upgrading to a bigger or newer model on its own, and taking model confidence at face value.
- A hallucination rate is only meaningful for a specific use case. Leaderboard numbers measure someone else's task. Measure on your own questions, by hallucination type, and track the abstention rate alongside the error rate.

Why Models Make Things Up
The mechanism in plain language
A large language model is trained to predict the next piece of text given everything before it. Asked a question, it produces the continuation that its training makes most plausible. When the answer is well represented in what it learned, or supplied in the prompt, the most plausible continuation is usually the true one. When it is not, the model still produces the most plausible continuation, and that continuation is shaped like an answer: the right format, the right tone, a specific name or number or citation.
This is why hallucinations are so persuasive. The machinery that makes the model fluent on questions it can answer makes it equally fluent on questions it cannot, and there is no internal flag that says "this part is a guess".
What makes it worse
Three conditions reliably raise the rate.
Missing context. A question about an internal policy, a customer account or a recent event that the model has never seen invites a plausible reconstruction. The model does not know what it does not know.
Ambiguous context. Two retrieved documents that disagree, or a question that could refer to either of two products, give the model room to blend them into an answer that matches neither.
False premises. A question such as "which clause of the contract allows early termination without notice" presupposes that such a clause exists. Models are strongly inclined to answer the question as asked rather than challenge its premise, and they will often locate or invent a clause.
Why training does not remove it
Research from OpenAI published in September 2025, Why Language Models Hallucinate by Kalai, Nachum, Vempala and Zhang, argues that standard training and evaluation actively reward guessing. Benchmarks that grade only on accuracy give zero credit for "I don't know" and some chance of credit for a guess, so a model tuned to score well learns to guess, like a student facing a multiple-choice exam with no penalty for wrong answers. The authors' proposed fix is to change how evaluations are scored so that abstaining is not punished. The practical implication for enterprises is direct: the same incentive exists in every internal evaluation that counts only correct answers.
Why it matters beyond accuracy
Two public cases show where the cost lands. In Moffatt v. Air Canada, decided by British Columbia's Civil Resolution Tribunal in February 2024, the airline was held liable for negligent misrepresentation after its website chatbot described a bereavement refund policy that did not exist. In Mata v. Avianca in 2023, a US federal court sanctioned lawyers who filed a brief citing cases that a chatbot had invented. In both, the organization that deployed the output carried the consequence, not the model provider.

The Three Kinds of Hallucination
The word covers three failures that look similar on the page and are caught in different ways. Separating them is the first step to choosing controls.
Fabrication is a stated fact that is false and was never in any source: a nonexistent court case, a made-up product feature, an invented statistic with a plausible decimal place.
Faithfulness failure is output that contradicts or goes beyond the source the model was given. The source document says a fee is waived for the first year; the summary says it is waived permanently. This is the dominant failure in retrieval-based systems and document processing, because the right information was present and the model still got it wrong.
Reasoning error is a wrong conclusion drawn from correct facts: an arithmetic slip, a policy rule applied to the wrong case, a correct date from which the wrong deadline is computed.
| Type | What it looks like | Example | Control that catches it |
|---|---|---|---|
| Fabrication | A specific, confident fact with no source | A citation to a case that does not exist | Grounding with citations; validation against a system of record; abstention |
| Faithfulness failure | Output that contradicts or exceeds the provided source | "Waived for the first year" summarized as "waived" | Verbatim citation checks; source-entailment checks; structured extraction with source spans |
| Reasoning error | Correct inputs, wrong conclusion | Right contract date, wrong notice deadline | Deterministic calculation outside the model; cross-field validation; sampling consistency |
| Premise acceptance | Answering a question built on a false assumption | Finding a termination clause that is not there | Instructions and evaluations that reward challenging the premise; abstention |
The fourth row is a special case of fabrication, listed separately because the fix is partly question design.
The Controls, Ranked by Cost and What They Fix
Controls are listed roughly from cheapest to most expensive to operate. The ordering matters because the cheap structural controls reduce the volume that reaches the expensive ones.
1. Constrain the output
Where the answer belongs to a known set, such as a category, a yes or no, a product code, or a field in a schema, force the model to return one of the allowed values. Schema-constrained output removes an entire class of invention, because the model cannot produce a product code that is not on the list. The mechanics are covered in structured outputs for LLMs.
Cost: very low. Fixes: fabricated values in closed domains, malformed output. Does not fix: choosing the wrong allowed value.
2. Allow and reward abstention
Give the model an explicit, legitimate way to say the answer is not available, and treat that response as a success in evaluation when the source genuinely lacks the answer. In a schema, this means a "not found" value for every field. In a chat assistant, it means an instruction and an evaluation set that both treat "the policy documents provided do not cover this" as correct behavior.
Cost: low, plus a lower answer rate that product teams must accept. Fixes: a large share of fabrication and premise acceptance. Does not fix: faithfulness failures on sources that do contain an answer.
3. Ground answers in retrieved sources, with citations
Retrieval-augmented generation supplies the model with relevant documents at question time and instructs it to answer from them, as described in retrieval-augmented generation. The critical addition is requiring a citation for every claim, ideally a verbatim quote with a document reference, and then checking mechanically that the quoted text exists in the cited document.
Cost: moderate, since retrieval quality has to be engineered and maintained. Fixes: most fabrication on topics the documents cover. Does not fix: faithfulness failures, poor retrieval, or questions the documents do not cover. Grounding is not a cure on its own: a Stanford RegLab and HAI study, Hallucination-Free?, published as a preprint in 2024 and peer-reviewed in 2025, found that leading retrieval-based legal research tools still hallucinated on between 17 and 33 percent of queries, while performing better than a general-purpose model.
4. Validate against a system of record
When the output contains a fact the organization already holds, check it. An account balance, a policy limit, a customer's plan, an invoice total, a date of birth. The model drafts and a deterministic check confirms, and any mismatch blocks or flags the output. Calculations belong outside the model entirely: have it extract the inputs and let code compute the deadline, the fee or the total.
Cost: moderate, mostly integration work. Fixes: fabrication and reasoning errors on anything verifiable. Does not fix: claims with no system of record behind them. This is the control that makes document extraction safe enough for finance workflows, as set out in document AI extraction.
5. Check consistency across samples
Ask the same question several times and compare the answers. A model that knows the answer tends to give the same one each time; a model that is guessing tends to vary. A June 2024 paper in Nature by Farquhar, Kossen, Kuhn and Gal, Detecting hallucinations in large language models using semantic entropy, formalized this by clustering answers by meaning rather than wording, and showed it detects a class of arbitrary fabrications the authors call confabulations without needing the correct answer.
Cost: high per query, since it multiplies model calls. Fixes: flags likely guesses, usefully as a trigger for review. Does not fix: errors the model makes consistently, which are exactly the ones a model "believes".
6. Human review, sized to error cost
Human review is the most expensive control and the least reliable at volume, because reviewers who see correct output most of the time learn to approve it. It works when it is targeted: routed by the cheaper controls above to the outputs that failed a check, sampled on the ones that passed, and staffed against the expected exception rate rather than total volume. Reviewers should know which hallucination type they are looking for, since checking a citation and checking a calculation are different tasks.
Cost: highest. Fixes: whatever the reviewer is equipped to catch. Does not fix: errors in outputs the reviewer never sees, or approves on autopilot.
| Control | Relative cost | Main type it reduces | Main blind spot |
|---|---|---|---|
| Constrained output | Very low | Fabricated values in closed domains | Wrong choice among allowed values |
| Abstention | Low, plus lower answer rate | Fabrication, premise acceptance | Misreading a source that has the answer |
| Grounding with verified citations | Moderate | Fabrication on covered topics | Retrieval misses, faithfulness failures |
| System-of-record validation | Moderate | Fabrication and reasoning on verifiable facts | Claims with nothing to check against |
| Sampling consistency | High per query | Guesses that vary between samples | Consistent errors |
| Targeted human review | Highest | Anything flagged and in scope for the reviewer | Unflagged output, reviewer fatigue |
What Works Less Well Than Claimed
A "do not hallucinate" instruction. Telling a model to be accurate or to avoid making things up does not give it knowledge it lacks or a signal it does not have. Instructions help when they grant permission to abstain and define what abstaining looks like. On their own they change tone more than truth.
A bigger or newer model on its own. Newer models are often better on average, and not reliably better on this. OpenAI's own April 2025 system card for o3 and o4-mini reported hallucination rates on its PersonQA evaluation of 33 percent for o3 and 48 percent for o4-mini, against 16 percent for the earlier o1, attributing part of the increase to the newer models making more claims overall. A model change is a reason to re-run the evaluation, not a reason to skip the controls.
Confidence scores at face value. A model asked how confident it is will produce a plausible confidence level, which is subject to the same problem as any other generated answer. Token probabilities carry more signal, but they conflate uncertainty about wording with uncertainty about facts, which is the problem semantic-entropy methods were designed to address. Confidence is useful as one routing input, calibrated against measured accuracy on the organization's own data, and misleading as a guarantee.
Retrieval as a complete fix. Grounding reduces fabrication and introduces its own failure: confidently wrong answers built from the wrong retrieved passage, as the Stanford legal study showed.
How to Measure a Hallucination Rate for Your Own Use Case
A published hallucination rate describes a specific model on a specific benchmark. It does not describe a claims assistant answering questions about the organization's own policies, and it can be misleading in either direction. The measurement has to be local.
Build the test set from real traffic. Collect a few hundred questions that users actually ask, including ones the source material does not answer and ones with false premises. The unanswerable questions are essential: without them, abstention can never score as correct.
Label by type, not just right or wrong. Record whether each error is a fabrication, a faithfulness failure or a reasoning error. The distribution tells you which control to invest in next. A system with mostly faithfulness failures needs better citation checking, not a better retriever.
Track three numbers together. The hallucination rate on answered questions, the abstention rate, and the share of abstentions that were correct. A system can drive its hallucination rate toward zero by refusing everything, so the error rate is only meaningful alongside how often it answers.
Re-run on every change. A model upgrade, a prompt change or a new document collection changes the rate. The evaluation is a regression test, not a launch gate that runs once. The broader discipline is covered in AI evaluations for executives.
Sample production. Offline tests drift from real use. A small, continuous sample of production outputs reviewed by type keeps the measured rate honest.
A Decision Framework: Which Controls a Use Case Needs
The minimum controls follow from two properties: the cost of a wrong answer, and whether a qualified person checks the output before it has an effect. Runtime guardrails that enforce these controls in production are covered in AI guardrails for LLM safety.
| Use case | Cost of a wrong answer | Who checks before it matters | Minimum controls |
|---|---|---|---|
| Internal brainstorming, first drafts | Low | The author, always | Abstention permitted; no further controls required |
| Internal knowledge assistant (policies, how-to) | Medium: wasted time, wrong internal action | The user, sometimes | Grounding with verified citations; abstention; local evaluation |
| Agent-assist drafts for support staff | Medium to high: wrong statement to a customer | The agent, before sending | Grounding with citations; system-of-record validation for account facts; agent sees sources |
| Customer-facing assistant | High: liability, as Moffatt v. Air Canada showed | Nobody before the customer reads it | Constrained scope; grounding; abstention to a human handoff; validation; production sampling |
| Document extraction into financial systems | High: wrong payment or record | Only flagged items | Schema constraints; source spans; cross-field and system-of-record validation; review queue |
| Legal, medical, credit or regulatory output | Very high, often irreversible | A qualified professional, every time | All of the above, plus mandatory expert review of every output with sources shown |
The position, and its tradeoff
The highest-return change most enterprises can make is not a model or a vendor. It is making abstention a first-class, measured behavior: a legitimate "the sources do not answer this" in every schema and every assistant, and evaluations that score a correct abstention as a success. Combined with validation against systems of record for any fact the organization already holds, that converts a large share of confident fabrications into visible gaps.
The tradeoff is that the system answers less often, and users and product teams perceive it as less capable. A customer assistant that says "I can't confirm that, here is how to reach someone who can" on a fifth of questions will score worse on satisfaction surveys in the short run than one that answers everything. For low-stakes internal use, that trade may not be worth making. For anything customer-facing or financial, an assistant that knows when to stop is worth more than one that is always fluent, and the liability cases are the reason.
Frequently Asked Questions
Why do large language models hallucinate?
A language model generates the most plausible continuation of its input, and it produces fluent true answers and fluent false ones by the same mechanism, with no reliable internal signal that separates them. Missing context, ambiguous sources and questions with false premises raise the rate. Research from OpenAI in 2025 also argues that standard training and evaluation reward guessing over admitting uncertainty, because benchmarks give no credit for "I don't know".
Can AI hallucinations be eliminated completely?
Not with current language models. They can be reduced substantially and, more importantly, made visible: constraining outputs, allowing abstention, grounding answers in cited sources and validating facts against systems of record turn many fabrications into flagged gaps. The remaining rate has to be measured for each use case and matched with review proportional to the cost of an error.
Does retrieval-augmented generation stop hallucinations?
It reduces fabrication on topics the retrieved documents cover, and it does not stop hallucinations entirely. Retrieval can surface the wrong passage, and models can still misstate a source they were given. A Stanford study found leading retrieval-based legal research tools hallucinated on 17 to 33 percent of test queries. Requiring verbatim citations and checking them mechanically closes much of the remaining gap.
Are newer, bigger models less likely to hallucinate?
Not reliably. Newer models are often more capable overall, but hallucination can rise as well as fall: OpenAI's own published evaluation in 2025 found two of its newer reasoning models hallucinated more often than an earlier one on a person-facts test. Every model change should be treated as a reason to re-run the organization's own evaluation.
How do you measure an AI hallucination rate?
Build a test set from real user questions, including questions the sources cannot answer and questions with false premises. Label each error as fabrication, faithfulness failure or reasoning error, and track the hallucination rate on answered questions together with the abstention rate, since refusing everything would drive errors to zero. Re-run the set on every model, prompt or data change, and sample production outputs continuously.
The Bottom Line
Hallucination is not a defect waiting for a patch. It is what a system built to produce plausible text does when it lacks the information to produce true text, and it will remain a property of language models for the foreseeable future. That makes it an engineering and governance problem with known tools rather than a mystery to be managed with caution.
The organizations that deploy these systems well stop asking whether a model hallucinates and start asking which kind of hallucination their use case is exposed to, which structural controls catch it, and who checks what is left. They let the model say it does not know, they check its facts against what they already hold, and they measure the result on their own questions. The ones that rely on a stronger model and a reminder to be accurate are relying on the model's fluency, which is exactly the property that makes its errors hard to see.