A hallucination is output that sounds fluent and confident but is false: an invented statistic, a non-existent citation, an API function that doesn't exist.
Why It Happens
Language models generate text that is likely given their training and the prompt; they don't look facts up or verify them. When the model lacks the information, it still produces plausible-sounding text. Training that rewards confident, complete answers can make this worse.
When It's Most Likely
- Specific facts: exact figures, dates, names, quotes.
- Citations, URLs and references.
- Niche or recent topics beyond the training data.
- Details about you or your organisation that weren't provided.
- Long outputs with many details.
Reducing It
- Ground the model: provide the source documents and instruct it to answer only from them, quoting what it relies on.
- Give it permission to not know: "If the answer isn't in the documents, say so."
- Ask narrower questions.
- Use retrieval (RAG) for knowledge-heavy applications.
- Use tools such as search or calculators for facts and arithmetic.
Catching It
- Verify important facts against primary sources.
- Check that cited passages actually exist and say what's claimed.
- Run generated code and tests.
- Use automated checks — for example, a second model grading whether answers are supported by the source.
Set Expectations
Treat outputs as drafts. The person publishing or acting on them remains responsible for accuracy.