Advanced Prompt EngineeringReducing hallucinations and defending against prompt injection · Lesson 10 of 17

Reducing hallucinations: grounding, citations and 'I don't know'

Article · 13 min · 8 min lecture

Video lecture

Reducing hallucinations: grounding, citations and 'I don't know'

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Reducing hallucinations

  • Why models hallucinate
  • Grounding, abstention, citations
  • Verifying quotes in code
  • Measuring properly

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why models hallucinate

Language models generate the most plausible continuation of text. When the true answer is in their context or strongly represented in training, plausible and true coincide. When it is not, the model can still produce something fluent and confident, because fluency is what it is optimised for. A hallucination is not a bug in one prompt; it is the default behaviour when the model lacks information and has no permission to say so.

That gives us three levers: give the model the information, give it permission to abstain, and make claims checkable.

Lever 1: grounding

Grounding means asking the model to answer from supplied sources rather than memory.

Answer using only the information in <sources>. If the sources do not
contain the answer, say "The provided documents don't cover this" and
state what information would be needed.

Grounding works best combined with good retrieval (so the right passages are actually present) and clear labelling of sources. Be aware of a subtle failure: models may blend source content with background knowledge. If strict grounding matters (legal, medical, financial contexts), say so explicitly and verify.

Lever 2: permission to say "I don't know"

Models often guess because the prompt implies an answer must exist. Explicitly allowing abstention reduces fabrication:

  • "If you are not confident, say so and explain what is uncertain."
  • "It is better to answer 'insufficient information' than to guess."
  • Provide a structured way to abstain: a confidence field with defined levels, or an answerable: false flag.

Pair this with product design: an abstention should route somewhere useful (a human, a search, a follow-up question), otherwise users learn to hate the "I don't know" answer and teams quietly remove it.

Lever 3: citations and quote-first answers

Requiring citations makes claims auditable and discourages unsupported statements:

First, extract the quotes from the documents most relevant to the
question into <quotes>, each with its document index.
Then answer in <answer>, citing quote numbers like [1][3] after each claim.
If no quote supports a claim, do not make the claim.

Then verify in code: every quote must appear verbatim (or near-verbatim, allowing whitespace differences) in the source. If a quote does not match, retract the claim or flag it. Some APIs offer built-in citation features that return source spans; where available they reduce the parsing work, but still spot-check them.

Other effective techniques

  • Narrow the task. "Summarise the risks mentioned in section 4" hallucinates less than "What are the risks of this company?"
  • Ask for specific formats for uncertain facts. "Give dates only if stated; otherwise write 'not stated'."
  • Use tools for facts. For arithmetic, dates, current prices or live data, give the model a calculator, code execution or a search tool rather than relying on memory.
  • Lower creative pressure. Instructions like "be comprehensive" or "give at least ten examples" push the model to pad with invented items. Ask for "up to ten, only those supported."
  • Check consistency. For high-stakes answers, generate several times; if answers disagree substantially, treat the result as uncertain.

Worked example: a policy assistant

An HR team builds an internal assistant for leave policy across offices in London, Dubai and Karachi. Early tests show the assistant confidently answering questions about local public holidays not in the documents. Fixes:

  1. Retrieval limited to the official policy documents, each labelled with office and effective date.
  2. Quote-first answers with citations, verified in code.
  3. An explicit rule: "If the policy for the employee's office is not in the sources, say so and suggest contacting HR."
  4. Office is required as input; the assistant asks for it if missing.

Unsupported answers drop sharply in their test set, and the remaining errors are mostly retrieval misses, which are easier to fix than a model inventing policy.

Measuring hallucination

You cannot manage what you do not measure. Build a test set that includes:

  • answerable questions (the answer is in the sources)
  • unanswerable questions (it is not)
  • trick questions with false premises ("Why did we cancel the Dubai office?")

Track: correct answers, correct abstentions, wrong answers and wrong abstentions. A system that never abstains will fail the unanswerable set; one that abstains constantly is useless. You are tuning the balance.

Hands-on: grounded answers with verified quotes

import json, os, re, anthropic
client = anthropic.Anthropic()

ANSWER_SCHEMA = {
    "type": "object",
    "properties": {
        "answerable": {"type": "boolean"},
        "quotes": {"type": "array", "items": {"type": "object", "properties": {
            "doc": {"type": "integer"}, "text": {"type": "string"}},
            "required": ["doc", "text"], "additionalProperties": False}},
        "answer": {"type": "string"},
    },
    "required": ["answerable", "quotes", "answer"], "additionalProperties": False,
}

def norm(s):
    return re.sub(r"\s+", " ", s).strip().lower()

def grounded_answer(question, docs):
    labelled = "\n".join(f'<document index="{i}">{d}</document>' for i, d in enumerate(docs))
    resp = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=2048,
        output_config={"format": {"type": "json_schema", "schema": ANSWER_SCHEMA}},
        messages=[{"role": "user", "content": f"""<documents>{labelled}</documents>
Answer using only the documents. First copy the supporting quotes exactly.
If the documents do not contain the answer, set answerable to false.
<question>{question}</question>"""}],
    )
    out = json.loads(next(b.text for b in resp.content if b.type == "text"))
    bad = [q for q in out["quotes"] if norm(q["text"]) not in norm(docs[q["doc"]])]
    if bad:
        out["answerable"] = False
        out["answer"] = "Could not verify supporting quotes; routing to a human."
    return out

Several APIs also offer built-in citations. With Claude, you can pass documents as content blocks with citations enabled and receive answers split into text blocks that carry cited spans (character ranges or PDF page numbers); note that this feature cannot be combined with schema-constrained output in the same request. Web search tools in the major APIs likewise return source URLs. Built-in citations reduce parsing work, but you should still spot-check them.

Grounding with live data

For questions about current facts (prices, regulations, availability), connect a search or database tool instead of relying on training data. Tell the model which sources are authoritative ("official government sites and our product database outrank blogs") and require the source URL or record ID in the answer.

Running your hallucination test set

Score four outcomes per case: correct answer, correct abstention, wrong answer, wrong abstention. Report them separately by tag. A change that reduces wrong answers but doubles wrong abstentions may be acceptable for a legal assistant and unacceptable for a sales assistant; the right balance depends on the cost of each error type in your context.

Going further

Groundedness can be checked automatically with a second model call that labels each sentence of the answer as supported, partially supported or unsupported by the cited passages. Treat this checker as a classifier with its own error rate and validate it against human judgements before relying on it.

Key takeaways

  • Hallucination is the default when the model lacks information and lacks permission to abstain.
  • Ground answers in labelled sources, allow and route 'I don't know', and require quote-backed citations.
  • Verify quotes in code, narrow tasks, use tools for facts, and avoid prompts that pressure padding.
  • Measure with answerable, unanswerable and false-premise questions to balance answers and abstentions.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which instruction is most likely to INCREASE fabricated items?
  2. Why include false-premise questions in a hallucination test set?
  3. After a model returns citations with quotes, what should your system do?

Put it into practice

Build a 15-question test for a document assistant: 7 answerable, 5 unanswerable, 3 false-premise. Score your current prompt, then add grounding, abstention and quote-first rules and score again.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.