Evaluating and Monitoring LLM ApplicationsEvaluating RAG pipelines and agents · Lesson 7 of 16
Evaluating RAG: retrieval, generation and abstention
Video lecture
Evaluating RAG: retrieval, generation and abstention
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Evaluating RAG
When a RAG assistant gives a wrong answer, there are two very different culprits. Either retrieval never found the right information, or the model had it and misused it. If you only score final answers, you are guessing which one to fix. In this lecture you will learn to evaluate retrieval and generation separately, the metrics for each, and a simple two by two that tells you where to invest.
0:30 Analogy: researcher and writer
Here is an analogy for evaluating RAG in parts. Think of a research assistant and a writer working together. The assistant fetches books from the library; the writer reads them and writes the report. If the report is wrong, you need to know whether the assistant brought the wrong books or the writer misread the right ones. Firing the writer when the assistant is the problem fixes nothing. Stage-level metrics tell you who to coach.
1:03 Stage-by-stage
Picture the pipeline: a question, perhaps a query rewrite, retrieval of the top results, maybe a reranker, then generation with the retrieved context. Each stage can be evaluated. For retrieval you need, for each test question, the identifiers of the chunks or documents that are genuinely relevant. Experts can label these, or you can bootstrap with a judge and then review. Do not let the retriever you are testing label its own results. That is circular.
1:36 Retrieval metrics
Five retrieval metrics are worth knowing. Recall at k: of the relevant items, how many made the top k? Precision at k: of the top k, how many are relevant? Hit rate: did at least one relevant item appear? Mean reciprocal rank: how high is the first relevant item? And nDCG, which rewards putting the most relevant items at the top when relevance is graded. For RAG, recall usually matters most, because the model cannot use what it never saw.
2:11 Generation metrics
On the generation side, a vocabulary popularized by the Ragas framework is now common. Faithfulness: are the answer's claims supported by the retrieved context? Answer relevance: does it actually address the question? Context precision: is the context focused? Context recall: does the context contain everything the reference answer needs? And answer correctness against a reference or key points.
2:36 Abstention
Do not forget abstention. When the answer is not in your knowledge base, does the system say so, or does it confidently guess? Build a set of unanswerable questions, like asking about a policy that does not exist, and measure how often the system correctly declines. Many teams discover their assistant never says I don't know, which is a trust problem waiting to happen.
3:04 The 2×2
Now the diagnostic two by two. Rows: was relevant context retrieved or not. Columns: was the answer correct or wrong. Relevant context and correct answer: working. Relevant context but wrong answer: a generation problem, like the prompt, the model, or conflicting documents. Missing context but wrong answer: a retrieval problem, like chunking, embeddings, filters or query rewriting. Missing context yet correct: the model used general knowledge, so check your groundedness policy.
3:35 Case: three-country HR policies
Here is a real-world shaped case. An HR assistant for a company with offices in Lahore, Dubai and London answered from policy documents. End-to-end correctness was disappointing. Stage metrics showed the cause: recall on leave questions was low, while faithfulness was high. The three countries' policies were nearly identical text, and chunks had no country metadata. Adding a country filter from the user profile and putting section headings into chunks fixed retrieval, and correctness followed. No prompt change needed.
4:09 Record everything
To make this repeatable, store every test result with the question, slice tags, relevant and retrieved identifiers, all metrics, the answer, and the configuration: embedder, chunking strategy, reranker and prompt version. The lesson shows a JSON record like this. With configuration recorded, you can compare chunking strategies or rerankers directly. And change one component at a time, so you know what caused the improvement.
4:37 Example: one question, five results
A simple example of computing recall. A question has two relevant chunks, A and B. Your retriever returns five results: C, A, D, E, F. Recall at five: one of the two relevant chunks is present, so fifty percent. Precision at five: one of five results is relevant, twenty percent. Reciprocal rank: the first relevant chunk is at position two, so one half. Now you know the model was missing chunk B entirely, so any answer that needed B was doomed before generation started.
5:14 Common mistakes
Common mistakes when evaluating RAG. Reporting only one end-to-end number. Letting the retriever you are testing decide which chunks are relevant. Forgetting unanswerable questions, so abstention is never tested. And changing the chunker, the embedder and the prompt all in the same week, so nobody knows which change helped. Try this now: pick ten failed answers from your system and place each one in the two by two. The pattern will tell you where to spend next sprint.
5:48 Deeper: testing abstention
One level deeper on abstention. Build ten unanswerable questions for the HR assistant, like asking about a sabbatical policy that does not exist. A good system says it cannot find that policy and suggests contacting HR. Measure the share of correct abstentions. In the first run, the assistant invented a sabbatical policy for four of them, which became a blocking test.
6:15 Watch me do it: diagnosing HR RAG
Watch me do it on the three-country HR assistant. First, labels. For thirty leave questions, an HR specialist marks the relevant chunk IDs, for example the UAE policy's leave section two for a Dubai employee's annual leave question. Second, I run retrieval only and compute metrics with the functions from the lesson. Recall at five averages around sixty percent; mean reciprocal rank is low. Third, I print one failing example. The Dubai question retrieved, in order, the UK leave section, the Pakistan leave section and then the UAE section at position three, plus two unrelated chunks. So recall at five was one, but for other questions the UAE chunk never made the top five at all. Fourth, generation metrics on the same set: faithfulness is high, the answers faithfully summarize whatever arrived. That confirms the diagnosis. Fifth, the two by two. Of twelve wrong answers, ten sit in the retrieval quadrant, relevant context missing. Two sit in the generation quadrant, where the right chunk arrived but the answer mixed in a UK detail. Sixth, one change at a time. I add a country metadata filter taken from the user's profile and rerun: recall at five rises sharply, and eight of the ten retrieval failures now answer correctly. Only then do I look at the two generation failures, with a prompt tweak, measured separately.
7:52 Recap
Recap. Evaluate retrieval and generation separately, then end to end. Prioritize recall at k for retrieval. Measure faithfulness, relevance, context quality, correctness and abstention for generation. Use the two by two to locate failures, and record configuration with every result. Your next step: label relevant chunks for thirty questions, run the retrieval metrics code from the lesson, and place each failure in the two by two.
Evaluate the pipeline in parts
A retrieval-augmented generation (RAG) system fails in two places: retrieval (the right information never reached the model) and generation (the model had the right information but misused it). If you only score final answers, you cannot tell which part to fix. Evaluate each stage separately, then end to end.
question -> [query rewrite] -> [retrieve top-k] -> [rerank] -> [generate with context] -> answer
^ eval: rewrite quality ^ eval: recall/precision/rank ^ eval: faithfulness, relevance, correctnessRetrieval metrics
You need, for each test question, the IDs of relevant chunks or documents (labeled by experts, or bootstrapped with a judge and then reviewed).
| Metric | Question it answers |
|---|---|
| Recall@k | Of the relevant items, what share appear in the top k? (Did we find it at all?) |
| Precision@k | Of the top k, what share are relevant? (How much noise did we add?) |
| Hit rate@k | Did at least one relevant item appear in the top k? |
| MRR (mean reciprocal rank) | How high is the first relevant item, on average? |
| nDCG@k | Are the most relevant items ranked highest, with graded relevance? |
For RAG, recall@k usually matters most: the generator cannot use what was never retrieved. Precision matters for cost and for avoiding distracting context.
# retrieval_metrics.py
import math
def recall_at_k(retrieved: list[str], relevant: set[str], k: int) -> float:
return len(set(retrieved[:k]) & relevant) / len(relevant) if relevant else 1.0
def precision_at_k(retrieved, relevant, k):
top = retrieved[:k]
return len([d for d in top if d in relevant]) / len(top) if top else 0.0
def mrr(retrieved, relevant):
for i, d in enumerate(retrieved, start=1):
if d in relevant:
return 1.0 / i
return 0.0
def ndcg_at_k(retrieved, grades: dict[str, int], k: int) -> float:
dcg = sum((2 ** grades.get(d, 0) - 1) / math.log2(i + 1) for i, d in enumerate(retrieved[:k], start=1))
ideal = sorted(grades.values(), reverse=True)[:k]
idcg = sum((2 ** g - 1) / math.log2(i + 1) for i, g in enumerate(ideal, start=1))
return dcg / idcg if idcg else 0.0Generation metrics
Frameworks such as Ragas popularized a vocabulary that is now common:
- Faithfulness / groundedness: are answer claims supported by the retrieved context? (Module 2.)
- Answer relevance: does the answer address the question asked?
- Context precision: is the retrieved context focused on what is needed?
- Context recall: does the retrieved context contain everything needed for the reference answer?
- Answer correctness: does it match the reference or key points?
Also measure abstention: when the answer is not in the knowledge base, does the system say so instead of guessing? Build a set of unanswerable questions for this.
Diagnosing with a 2×2
| Answer correct | Answer wrong | |
|---|---|---|
| Relevant context retrieved | Working | Generation problem (prompt, model, context too long, conflicting docs) |
| Relevant context missing | Lucky or using general knowledge (check groundedness policy) | Retrieval problem (chunking, embeddings, filters, query rewriting) |
Tagging each failed test with its quadrant tells you where to invest.
Worked example: tuning a policy assistant
An HR assistant for a company with offices in Lahore, Dubai and London answered questions from policy PDFs. End-to-end correctness was 74% (illustrative). Stage metrics showed recall@5 of only 61% on questions about leave, while faithfulness was high. The problem was retrieval: policies for three countries were near-identical text, and chunks lacked country metadata. Adding a country filter from the user profile and including section headings in chunks raised recall@5 substantially, and end-to-end correctness followed. No prompt change was needed.
Hands-on: a RAG eval record
Store every test result with enough detail to diagnose:
{
"id": "hr-leave-017",
"question": "How many days of annual leave do I get in the Dubai office?",
"slice": {"country": "AE", "topic": "leave", "answerable": true},
"relevant_ids": ["ae-policy#leave-2"],
"retrieved_ids": ["gb-policy#leave-1", "ae-policy#leave-2", "pk-policy#leave-3"],
"metrics": {"recall@5": 1.0, "precision@5": 0.33, "mrr": 0.5,
"faithfulness": 1.0, "correct": true, "abstained": false},
"answer": "...",
"config": {"embedder": "v3", "chunking": "heading-aware-800", "reranker": "on", "prompt": "hr-v12"}
}With the configuration recorded, you can compare chunking strategies, embedders and rerankers directly.
Pitfalls
- Only end-to-end scores, leaving you guessing which stage broke.
- Relevance labels generated by the retriever you are testing (circular). Review them.
- No unanswerable questions, so you never measure abstention.
- Changing several components at once; change one, measure, then the next.
How to measure success
Stage-level metrics for every release, a 2×2 diagnosis of failures, and improvements traced to specific component changes.
Key takeaways
- Evaluate retrieval and generation separately, then end to end.
- Recall@k usually matters most for RAG; also track precision, MRR and nDCG.
- Measure faithfulness, answer relevance, context precision and recall, correctness and abstention.
- Use the context-vs-answer 2×2 to locate failures and record configuration with every result.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Label relevant chunks for 30 questions, compute recall@5, precision@5 and MRR with the code provided, and classify every failure in the 2×2.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.