An AI hallucination is output from an AI model, most often a large language model (LLM), that sounds confident and plausible but is false, invented or not supported by the information the model was given. Examples include a made-up statistic, a citation to a document that doesn’t exist, a policy the company never had or a summary that adds details absent from the original. The term is borrowed loosely from psychology; the model isn’t perceiving anything, it is producing likely-sounding text.
At a glance
- Hallucinations are a known property of generative AI, not a rare bug.
- They are hard to spot because wrong answers are written just as fluently as right ones.
- Grounding answers in trusted sources, narrowing the task and human review reduce the rate; no current technique eliminates it.
- The risk depends on the use: a wrong draft email is minor, a wrong answer to a customer or in a contract is not.
- The organization using the output generally carries the consequences.
What problem it solves
Hallucination is a risk, not a solution; the term names a failure that buyers need to plan for. Generative AI tools work by predicting plausible text, not by looking facts up. Most of the time the result is accurate enough to be useful. Sometimes, especially when the model lacks the right information or the question is unusual, it fills the gap with something invented, and it does so without signaling doubt.
For a business, the problem shows up when AI output is trusted without checking: a chatbot promises a refund policy that doesn’t exist, an AI assistant drafts a report with fabricated figures, or an AI agent acts on a wrong assumption. Understanding hallucination helps buyers decide where AI can be used freely, where it needs review and where it shouldn’t be used yet.
How it works
Why it happens. A language model generates text one piece at a time, choosing what is statistically likely given everything before it. It has no built-in check that a statement is true. Hallucinations are more likely when:
- the question is about facts the model wasn’t trained on or wasn’t given, such as recent events or your internal policies;
- the question is ambiguous or assumes something false;
- the model is asked for specifics such as numbers, names, quotes or citations;
- the documents it was given are incomplete, conflicting or out of date.
How it is reduced. Common measures, usually combined:
- Grounding: retrieving trusted documents at question time and instructing the model to answer only from them, with citations.
- Narrow scope: limiting the tool to a defined set of tasks and topics, and letting it say “I don’t know”.
- Verification: automated checks against source data, and human review for anything high-stakes.
- Model and settings choice: some models and configurations are more reliable for factual tasks; test them on your own questions.
- Design: showing sources so users can check, and making clear to users that answers may be wrong.
When it matters for buyers
- When deploying customer-facing AI. Chatbots and virtual agents that state prices, policies or eligibility need tight grounding, limits and escalation to people. A chatbot that only answers from approved content is easier to control than an open-ended one.
- When AI output feeds decisions or documents. Financial, legal, medical, HR and contract work need human verification.
- When rolling out AI assistants to staff. Training people to check output is as important as the license.
- When agents take actions. A hallucinated fact becomes an incorrect action.
- When setting AI governance. Policies should say where review is mandatory. Our artificial intelligence overview covers evaluation approaches.
Questions to ask vendors
- How do you ground answers in our content, and does the tool show its sources?
- What does the tool do when it doesn’t know or can’t find an answer?
- How do you measure accuracy and hallucination rates, and can we test with our own questions?
- Can we restrict the topics it answers and the commitments it can make?
- How quickly do updates to our source content reach the tool’s answers?
- What liability do you accept for incorrect output?
How it differs from prompt injection
Prompt injection is an attack: someone deliberately feeds an AI model instructions to make it misbehave. A hallucination needs no attacker; the model produces wrong output on its own. The controls overlap, such as limiting what the AI can do and reviewing high-impact output, but prompt injection is handled as a security threat, while hallucination is a quality and accuracy issue. Both can lead to the same visible result, an AI saying or doing something wrong, so incident reviews should check for each.
