Advanced Prompt EngineeringReducing hallucinations and defending against prompt injection · Lesson 10 of 17
Reducing hallucinations: grounding, citations and 'I don't know'
Video lecture
Reducing hallucinations: grounding, citations and 'I don't know'
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Reducing hallucinations
An HR assistant for offices in London, Dubai and Karachi confidently told an employee about a public holiday that was not in any policy document. It was not broken. It was doing exactly what language models do when they lack information and lack permission to say so. In this lecture you will learn why models hallucinate, the three levers that reduce it, grounding, abstention and citations, how to verify quotes in code, and how to measure hallucination properly.
0:34 Why it happens
Language models generate the most plausible continuation of text. When the true answer is in their context or strongly represented in training, plausible and true coincide. When it is not, the model can still produce something fluent and confident, because fluency is what it is optimised for. So hallucination is not a bug in one prompt. It is the default behaviour when the model lacks information and has no permission to abstain. That gives us three levers: give the model the information, give it permission to say I don't know, and make its claims checkable.
1:15 Lever 1: grounding
Lever one is grounding: answer using only the information in the sources provided. If the sources do not contain the answer, say so and state what information would be needed. Grounding works best with good retrieval, so the right passages are actually present, and with clear source labels. Beware of blending, where the model mixes source content with background knowledge. If strict grounding matters, in legal, medical or financial contexts, say so explicitly and verify. For current facts, connect a search or database tool and say which sources are authoritative.
1:54 Lever 2: permission to abstain
Lever two is permission to abstain. Models often guess because the prompt implies an answer must exist. Explicitly allow uncertainty: it is better to answer insufficient information than to guess. Better still, give a structured way to abstain, such as an answerable true or false field, or a confidence field with defined levels. And pair it with product design. An abstention must route somewhere useful, to a human, a search or a follow-up question. Otherwise users learn to hate the I don't know answer, and teams quietly remove it.
2:33 Lever 3: quote first, then verify
Lever three is citations and quote-first answers. First, extract the quotes most relevant to the question, each with its document index. Then answer, citing quote numbers after each claim. If no quote supports a claim, do not make the claim. Then verify in code: every quote must appear verbatim, allowing for whitespace differences, in the source. In the lesson's code, a schema-enforced answer contains an answerable flag, the quotes and the answer; code normalises whitespace and checks each quote against its document. If any quote fails, the answer is withheld and routed to a human.
3:14 Built-in citations
Several APIs also offer built-in citations. With Claude, you can pass documents with citations enabled and receive answers whose text blocks carry cited spans, such as character ranges or PDF page numbers. Note that this feature cannot be combined with schema-constrained output in the same request. Web search tools in the major APIs likewise return source URLs. Built-in citations reduce parsing work, but spot-check them anyway.
3:43 More techniques
Other techniques help too. Narrow the task: summarise the risks in section four hallucinates less than what are the risks of this company. Ask for specific formats for uncertain facts, like give dates only if stated, otherwise write not stated. Use tools for arithmetic and live data. Lower creative pressure: be comprehensive or give at least ten examples pushes the model to pad with invented items, so ask for up to ten, only those supported. And for high-stakes answers, generate several times; if answers disagree substantially, treat the result as uncertain.
4:23 Worked example: HR policy assistant
Back to the HR assistant. The fixes: retrieval limited to official policy documents, each labelled with office and effective date. Quote-first answers with citations verified in code. An explicit rule: if the policy for the employee's office is not in the sources, say so and suggest contacting HR. And office is now a required input; the assistant asks for it if missing. Unsupported answers drop sharply in their test set, and the remaining errors are mostly retrieval misses, which are much easier to fix than a model inventing policy.
5:02 Measuring hallucination
You cannot manage what you do not measure. Build a test set with three kinds of questions: answerable, where the answer is in the sources; unanswerable, where it is not; and false-premise questions, like why did we close the Dubai office, when you did not. Score four outcomes: correct answers, correct abstentions, wrong answers and wrong abstentions, and report them separately by tag. A system that never abstains fails the unanswerable set. One that abstains constantly is useless. You are tuning a balance, and the right balance depends on the cost of each error in your context.
5:44 Example 1: the town hall question
A simple worked example. Ask a model: what did our CEO say about pricing in last Tuesday's town hall? With no transcript, a model may produce a plausible quote. Now paste the transcript and add three instructions. Answer only from the transcript. Quote the exact words first. If the transcript does not cover pricing, say so. The model finds that pricing was not discussed at all, and says exactly that. The honest answer, not discussed, is only possible because you gave it the source and permission to say no.
6:23 Example 2: plan-features assistant (illustrative)
Now a business scenario, with illustrative numbers. A telecom reseller in Karachi runs a sales assistant answering questions about about forty business plans. In testing on one hundred and twenty questions, it invented plan features for about fifteen percent of questions about plans that did not include them. The team applied the three levers. Retrieval limited to the current plan sheet, each chunk labelled with plan name and effective date. A schema with an answerable flag and quotes. And code that verifies every quote exists in the plan sheet. They also added twenty unanswerable and ten false-premise questions to the test set. Wrong answers fell to around two percent, correct abstentions rose, and the few remaining errors were mostly retrieval pulling the wrong plan, which they fixed with better labels. Illustrative numbers.
7:20 Recap
To recap. Hallucination is the default when information and permission are missing. Ground answers in labelled sources, allow and route abstentions, require quote-first citations, and verify quotes in code. Narrow tasks, use tools for facts, and avoid prompts that pressure padding. Measure with answerable, unanswerable and false-premise questions. Try this now: build a fifteen-question test for a document assistant, seven answerable, five unanswerable and three false-premise, score your current prompt, then add grounding, abstention and quote-first rules and score it again. Next: prompt injection.
Why models hallucinate
Language models generate the most plausible continuation of text. When the true answer is in their context or strongly represented in training, plausible and true coincide. When it is not, the model can still produce something fluent and confident, because fluency is what it is optimised for. A hallucination is not a bug in one prompt; it is the default behaviour when the model lacks information and has no permission to say so.
That gives us three levers: give the model the information, give it permission to abstain, and make claims checkable.
Lever 1: grounding
Grounding means asking the model to answer from supplied sources rather than memory.
Answer using only the information in <sources>. If the sources do not
contain the answer, say "The provided documents don't cover this" and
state what information would be needed.Grounding works best combined with good retrieval (so the right passages are actually present) and clear labelling of sources. Be aware of a subtle failure: models may blend source content with background knowledge. If strict grounding matters (legal, medical, financial contexts), say so explicitly and verify.
Lever 2: permission to say "I don't know"
Models often guess because the prompt implies an answer must exist. Explicitly allowing abstention reduces fabrication:
- "If you are not confident, say so and explain what is uncertain."
- "It is better to answer 'insufficient information' than to guess."
- Provide a structured way to abstain: a
confidencefield with defined levels, or ananswerable: falseflag.
Pair this with product design: an abstention should route somewhere useful (a human, a search, a follow-up question), otherwise users learn to hate the "I don't know" answer and teams quietly remove it.
Lever 3: citations and quote-first answers
Requiring citations makes claims auditable and discourages unsupported statements:
First, extract the quotes from the documents most relevant to the
question into <quotes>, each with its document index.
Then answer in <answer>, citing quote numbers like [1][3] after each claim.
If no quote supports a claim, do not make the claim.Then verify in code: every quote must appear verbatim (or near-verbatim, allowing whitespace differences) in the source. If a quote does not match, retract the claim or flag it. Some APIs offer built-in citation features that return source spans; where available they reduce the parsing work, but still spot-check them.
Other effective techniques
- Narrow the task. "Summarise the risks mentioned in section 4" hallucinates less than "What are the risks of this company?"
- Ask for specific formats for uncertain facts. "Give dates only if stated; otherwise write 'not stated'."
- Use tools for facts. For arithmetic, dates, current prices or live data, give the model a calculator, code execution or a search tool rather than relying on memory.
- Lower creative pressure. Instructions like "be comprehensive" or "give at least ten examples" push the model to pad with invented items. Ask for "up to ten, only those supported."
- Check consistency. For high-stakes answers, generate several times; if answers disagree substantially, treat the result as uncertain.
Worked example: a policy assistant
An HR team builds an internal assistant for leave policy across offices in London, Dubai and Karachi. Early tests show the assistant confidently answering questions about local public holidays not in the documents. Fixes:
- Retrieval limited to the official policy documents, each labelled with office and effective date.
- Quote-first answers with citations, verified in code.
- An explicit rule: "If the policy for the employee's office is not in the sources, say so and suggest contacting HR."
- Office is required as input; the assistant asks for it if missing.
Unsupported answers drop sharply in their test set, and the remaining errors are mostly retrieval misses, which are easier to fix than a model inventing policy.
Measuring hallucination
You cannot manage what you do not measure. Build a test set that includes:
- answerable questions (the answer is in the sources)
- unanswerable questions (it is not)
- trick questions with false premises ("Why did we cancel the Dubai office?")
Track: correct answers, correct abstentions, wrong answers and wrong abstentions. A system that never abstains will fail the unanswerable set; one that abstains constantly is useless. You are tuning the balance.
Hands-on: grounded answers with verified quotes
import json, os, re, anthropic
client = anthropic.Anthropic()
ANSWER_SCHEMA = {
"type": "object",
"properties": {
"answerable": {"type": "boolean"},
"quotes": {"type": "array", "items": {"type": "object", "properties": {
"doc": {"type": "integer"}, "text": {"type": "string"}},
"required": ["doc", "text"], "additionalProperties": False}},
"answer": {"type": "string"},
},
"required": ["answerable", "quotes", "answer"], "additionalProperties": False,
}
def norm(s):
return re.sub(r"\s+", " ", s).strip().lower()
def grounded_answer(question, docs):
labelled = "\n".join(f'<document index="{i}">{d}</document>' for i, d in enumerate(docs))
resp = client.messages.create(
model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=2048,
output_config={"format": {"type": "json_schema", "schema": ANSWER_SCHEMA}},
messages=[{"role": "user", "content": f"""<documents>{labelled}</documents>
Answer using only the documents. First copy the supporting quotes exactly.
If the documents do not contain the answer, set answerable to false.
<question>{question}</question>"""}],
)
out = json.loads(next(b.text for b in resp.content if b.type == "text"))
bad = [q for q in out["quotes"] if norm(q["text"]) not in norm(docs[q["doc"]])]
if bad:
out["answerable"] = False
out["answer"] = "Could not verify supporting quotes; routing to a human."
return outSeveral APIs also offer built-in citations. With Claude, you can pass documents as content blocks with citations enabled and receive answers split into text blocks that carry cited spans (character ranges or PDF page numbers); note that this feature cannot be combined with schema-constrained output in the same request. Web search tools in the major APIs likewise return source URLs. Built-in citations reduce parsing work, but you should still spot-check them.
Grounding with live data
For questions about current facts (prices, regulations, availability), connect a search or database tool instead of relying on training data. Tell the model which sources are authoritative ("official government sites and our product database outrank blogs") and require the source URL or record ID in the answer.
Running your hallucination test set
Score four outcomes per case: correct answer, correct abstention, wrong answer, wrong abstention. Report them separately by tag. A change that reduces wrong answers but doubles wrong abstentions may be acceptable for a legal assistant and unacceptable for a sales assistant; the right balance depends on the cost of each error type in your context.
Going further
Groundedness can be checked automatically with a second model call that labels each sentence of the answer as supported, partially supported or unsupported by the cited passages. Treat this checker as a classifier with its own error rate and validate it against human judgements before relying on it.
Key takeaways
- Hallucination is the default when the model lacks information and lacks permission to abstain.
- Ground answers in labelled sources, allow and route 'I don't know', and require quote-backed citations.
- Verify quotes in code, narrow tasks, use tools for facts, and avoid prompts that pressure padding.
- Measure with answerable, unanswerable and false-premise questions to balance answers and abstentions.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a 15-question test for a document assistant: 7 answerable, 5 unanswerable, 3 false-premise. Score your current prompt, then add grounding, abstention and quote-first rules and score again.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.