What Is an AI Hallucination?

Also called: AI confabulation

Related problems: AI tools giving confident answers that turn out to be wrong; Chatbot telling customers about policies we don't have; Made-up sources and citations in AI-drafted documents; Not knowing whether we can trust AI output enough to use it

An AI hallucination is output from an AI model, most often a large language model (LLM), that sounds confident and plausible but is false, invented or not supported by the information the model was given. Examples include a made-up statistic, a citation to a document that doesn’t exist, a policy the company never had or a summary that adds details absent from the original. The term is borrowed loosely from psychology; the model isn’t perceiving anything, it is producing likely-sounding text.

At a glance

  • Hallucinations are a known property of generative AI, not a rare bug.
  • They are hard to spot because wrong answers are written just as fluently as right ones.
  • Grounding answers in trusted sources, narrowing the task and human review reduce the rate; no current technique eliminates it.
  • The risk depends on the use: a wrong draft email is minor, a wrong answer to a customer or in a contract is not.
  • The organization using the output generally carries the consequences.

What problem it solves

Hallucination is a risk, not a solution; the term names a failure that buyers need to plan for. Generative AI tools work by predicting plausible text, not by looking facts up. Most of the time the result is accurate enough to be useful. Sometimes, especially when the model lacks the right information or the question is unusual, it fills the gap with something invented, and it does so without signaling doubt.

For a business, the problem shows up when AI output is trusted without checking: a chatbot promises a refund policy that doesn’t exist, an AI assistant drafts a report with fabricated figures, or an AI agent acts on a wrong assumption. Understanding hallucination helps buyers decide where AI can be used freely, where it needs review and where it shouldn’t be used yet.

How it works

Why it happens. A language model generates text one piece at a time, choosing what is statistically likely given everything before it. It has no built-in check that a statement is true. Hallucinations are more likely when:

  • the question is about facts the model wasn’t trained on or wasn’t given, such as recent events or your internal policies;
  • the question is ambiguous or assumes something false;
  • the model is asked for specifics such as numbers, names, quotes or citations;
  • the documents it was given are incomplete, conflicting or out of date.

How it is reduced. Common measures, usually combined:

  • Grounding: retrieving trusted documents at question time and instructing the model to answer only from them, with citations.
  • Narrow scope: limiting the tool to a defined set of tasks and topics, and letting it say “I don’t know”.
  • Verification: automated checks against source data, and human review for anything high-stakes.
  • Model and settings choice: some models and configurations are more reliable for factual tasks; test them on your own questions.
  • Design: showing sources so users can check, and making clear to users that answers may be wrong.

When it matters for buyers

  • When deploying customer-facing AI. Chatbots and virtual agents that state prices, policies or eligibility need tight grounding, limits and escalation to people. A chatbot that only answers from approved content is easier to control than an open-ended one.
  • When AI output feeds decisions or documents. Financial, legal, medical, HR and contract work need human verification.
  • When rolling out AI assistants to staff. Training people to check output is as important as the license.
  • When agents take actions. A hallucinated fact becomes an incorrect action.
  • When setting AI governance. Policies should say where review is mandatory. Our artificial intelligence overview covers evaluation approaches.

Questions to ask vendors

  • How do you ground answers in our content, and does the tool show its sources?
  • What does the tool do when it doesn’t know or can’t find an answer?
  • How do you measure accuracy and hallucination rates, and can we test with our own questions?
  • Can we restrict the topics it answers and the commitments it can make?
  • How quickly do updates to our source content reach the tool’s answers?
  • What liability do you accept for incorrect output?

How it differs from prompt injection

Prompt injection is an attack: someone deliberately feeds an AI model instructions to make it misbehave. A hallucination needs no attacker; the model produces wrong output on its own. The controls overlap, such as limiting what the AI can do and reviewing high-impact output, but prompt injection is handled as a security threat, while hallucination is a quality and accuracy issue. Both can lead to the same visible result, an AI saying or doing something wrong, so incident reviews should check for each.

Frequently Asked Questions

Why do AI models hallucinate?
Generative models produce the most likely-looking continuation of text based on patterns, not by checking facts. When they lack the right information, or the question is ambiguous, they can still produce a fluent answer. Gaps or errors in training data and in the documents they are given add to the problem.
Can hallucinations be eliminated?
Not with current techniques. They can be reduced, often substantially, by grounding answers in trusted documents, narrowing the task, requiring citations and adding human review, but any generative model can still produce wrong output.
What is grounding?
Grounding means giving the model trusted source material, such as your policies or product documents, and instructing it to answer from that material. A common approach retrieves relevant documents at the time of the question. It reduces hallucinations but doesn't guarantee accuracy, especially if the sources are wrong or out of date.
Who is responsible if an AI chatbot gives a customer wrong information?
Generally the organization that deployed it, though this depends on the facts and the jurisdiction. Treat customer-facing AI output as your own statements, limit what it may promise, and check liability with counsel and in vendor contracts.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.