---
title: "Evaluating RAG: retrieval, generation and abstention"
description: "Evaluate the pipeline in parts A retrieval-augmented generation (RAG) system fails in two places: retrieval (the right information never reached the…"
url: https://optimizeall.com/learn/llm-evals-and-observability/evaluating-rag
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Evaluating RAG pipelines and agents · lesson 7 of 16 · 15 min

# Evaluating RAG: retrieval, generation and abstention

## Evaluate the pipeline in parts

A retrieval-augmented generation (RAG) system fails in two places: **retrieval** (the right information never reached the model) and **generation** (the model had the right information but misused it). If you only score final answers, you cannot tell which part to fix. Evaluate each stage separately, then end to end.

```text
question -> [query rewrite] -> [retrieve top-k] -> [rerank] -> [generate with context] -> answer
               ^ eval: rewrite quality   ^ eval: recall/precision/rank   ^ eval: faithfulness, relevance, correctness
```

## Retrieval metrics

You need, for each test question, the IDs of **relevant** chunks or documents (labeled by experts, or bootstrapped with a judge and then reviewed).

| Metric | Question it answers |
|---|---|
| **Recall@k** | Of the relevant items, what share appear in the top k? (Did we find it at all?) |
| **Precision@k** | Of the top k, what share are relevant? (How much noise did we add?) |
| **Hit rate@k** | Did at least one relevant item appear in the top k? |
| **MRR** (mean reciprocal rank) | How high is the first relevant item, on average? |
| **nDCG@k** | Are the most relevant items ranked highest, with graded relevance? |

For RAG, **recall@k** usually matters most: the generator cannot use what was never retrieved. Precision matters for cost and for avoiding distracting context.

```python
# retrieval_metrics.py
import math

def recall_at_k(retrieved: list[str], relevant: set[str], k: int) -> float:
    return len(set(retrieved[:k]) & relevant) / len(relevant) if relevant else 1.0

def precision_at_k(retrieved, relevant, k):
    top = retrieved[:k]
    return len([d for d in top if d in relevant]) / len(top) if top else 0.0

def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, start=1):
        if d in relevant:
            return 1.0 / i
    return 0.0

def ndcg_at_k(retrieved, grades: dict[str, int], k: int) -> float:
    dcg = sum((2 ** grades.get(d, 0) - 1) / math.log2(i + 1) for i, d in enumerate(retrieved[:k], start=1))
    ideal = sorted(grades.values(), reverse=True)[:k]
    idcg = sum((2 ** g - 1) / math.log2(i + 1) for i, g in enumerate(ideal, start=1))
    return dcg / idcg if idcg else 0.0
```

## Generation metrics

Frameworks such as **Ragas** popularized a vocabulary that is now common:

- **Faithfulness / groundedness:** are answer claims supported by the retrieved context? (Module 2.)
- **Answer relevance:** does the answer address the question asked?
- **Context precision:** is the retrieved context focused on what is needed?
- **Context recall:** does the retrieved context contain everything needed for the reference answer?
- **Answer correctness:** does it match the reference or key points?

Also measure **abstention**: when the answer is not in the knowledge base, does the system say so instead of guessing? Build a set of unanswerable questions for this.

## Diagnosing with a 2×2

| | Answer correct | Answer wrong |
|---|---|---|
| **Relevant context retrieved** | Working | Generation problem (prompt, model, context too long, conflicting docs) |
| **Relevant context missing** | Lucky or using general knowledge (check groundedness policy) | Retrieval problem (chunking, embeddings, filters, query rewriting) |

Tagging each failed test with its quadrant tells you where to invest.

## Worked example: tuning a policy assistant

An HR assistant for a company with offices in Lahore, Dubai and London answered questions from policy PDFs. End-to-end correctness was 74% (illustrative). Stage metrics showed recall@5 of only 61% on questions about leave, while faithfulness was high. The problem was retrieval: policies for three countries were near-identical text, and chunks lacked country metadata. Adding a country filter from the user profile and including section headings in chunks raised recall@5 substantially, and end-to-end correctness followed. No prompt change was needed.

## Hands-on: a RAG eval record

Store every test result with enough detail to diagnose:

```json
{
  "id": "hr-leave-017",
  "question": "How many days of annual leave do I get in the Dubai office?",
  "slice": {"country": "AE", "topic": "leave", "answerable": true},
  "relevant_ids": ["ae-policy#leave-2"],
  "retrieved_ids": ["gb-policy#leave-1", "ae-policy#leave-2", "pk-policy#leave-3"],
  "metrics": {"recall@5": 1.0, "precision@5": 0.33, "mrr": 0.5,
              "faithfulness": 1.0, "correct": true, "abstained": false},
  "answer": "...",
  "config": {"embedder": "v3", "chunking": "heading-aware-800", "reranker": "on", "prompt": "hr-v12"}
}
```

With the configuration recorded, you can compare chunking strategies, embedders and rerankers directly.

## Pitfalls

- **Only end-to-end scores**, leaving you guessing which stage broke.
- **Relevance labels generated by the retriever you are testing** (circular). Review them.
- **No unanswerable questions**, so you never measure abstention.
- **Changing several components at once**; change one, measure, then the next.

## How to measure success

Stage-level metrics for every release, a 2×2 diagnosis of failures, and improvements traced to specific component changes.

## Video lecture: Evaluating RAG: retrieval, generation and abstention

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Evaluating RAG
2. Analogy: researcher and writer
3. Stage-by-stage
4. Retrieval metrics
5. Generation metrics
6. Abstention
7. The 2×2
8. Case: three-country HR policies
9. Record everything
10. Example: one question, five results
11. Common mistakes
12. Deeper: testing abstention
13. Watch me do it: diagnosing HR RAG
14. Recap

## Lecture transcript

### Evaluating RAG

When a RAG assistant gives a wrong answer, there are two very different culprits. Either retrieval never found the right information, or the model had it and misused it. If you only score final answers, you are guessing which one to fix. In this lecture you will learn to evaluate retrieval and generation separately, the metrics for each, and a simple two by two that tells you where to invest.

### Analogy: researcher and writer

Here is an analogy for evaluating RAG in parts. Think of a research assistant and a writer working together. The assistant fetches books from the library; the writer reads them and writes the report. If the report is wrong, you need to know whether the assistant brought the wrong books or the writer misread the right ones. Firing the writer when the assistant is the problem fixes nothing. Stage-level metrics tell you who to coach.

### Stage-by-stage

Picture the pipeline: a question, perhaps a query rewrite, retrieval of the top results, maybe a reranker, then generation with the retrieved context. Each stage can be evaluated. For retrieval you need, for each test question, the identifiers of the chunks or documents that are genuinely relevant. Experts can label these, or you can bootstrap with a judge and then review. Do not let the retriever you are testing label its own results. That is circular.

### Retrieval metrics

Five retrieval metrics are worth knowing. Recall at k: of the relevant items, how many made the top k? Precision at k: of the top k, how many are relevant? Hit rate: did at least one relevant item appear? Mean reciprocal rank: how high is the first relevant item? And nDCG, which rewards putting the most relevant items at the top when relevance is graded. For RAG, recall usually matters most, because the model cannot use what it never saw.

### Generation metrics

On the generation side, a vocabulary popularized by the Ragas framework is now common. Faithfulness: are the answer's claims supported by the retrieved context? Answer relevance: does it actually address the question? Context precision: is the context focused? Context recall: does the context contain everything the reference answer needs? And answer correctness against a reference or key points.

### Abstention

Do not forget abstention. When the answer is not in your knowledge base, does the system say so, or does it confidently guess? Build a set of unanswerable questions, like asking about a policy that does not exist, and measure how often the system correctly declines. Many teams discover their assistant never says I don't know, which is a trust problem waiting to happen.

### The 2×2

Now the diagnostic two by two. Rows: was relevant context retrieved or not. Columns: was the answer correct or wrong. Relevant context and correct answer: working. Relevant context but wrong answer: a generation problem, like the prompt, the model, or conflicting documents. Missing context but wrong answer: a retrieval problem, like chunking, embeddings, filters or query rewriting. Missing context yet correct: the model used general knowledge, so check your groundedness policy.

### Case: three-country HR policies

Here is a real-world shaped case. An HR assistant for a company with offices in Lahore, Dubai and London answered from policy documents. End-to-end correctness was disappointing. Stage metrics showed the cause: recall on leave questions was low, while faithfulness was high. The three countries' policies were nearly identical text, and chunks had no country metadata. Adding a country filter from the user profile and putting section headings into chunks fixed retrieval, and correctness followed. No prompt change needed.

### Record everything

To make this repeatable, store every test result with the question, slice tags, relevant and retrieved identifiers, all metrics, the answer, and the configuration: embedder, chunking strategy, reranker and prompt version. The lesson shows a JSON record like this. With configuration recorded, you can compare chunking strategies or rerankers directly. And change one component at a time, so you know what caused the improvement.

### Example: one question, five results

A simple example of computing recall. A question has two relevant chunks, A and B. Your retriever returns five results: C, A, D, E, F. Recall at five: one of the two relevant chunks is present, so fifty percent. Precision at five: one of five results is relevant, twenty percent. Reciprocal rank: the first relevant chunk is at position two, so one half. Now you know the model was missing chunk B entirely, so any answer that needed B was doomed before generation started.

### Common mistakes

Common mistakes when evaluating RAG. Reporting only one end-to-end number. Letting the retriever you are testing decide which chunks are relevant. Forgetting unanswerable questions, so abstention is never tested. And changing the chunker, the embedder and the prompt all in the same week, so nobody knows which change helped. Try this now: pick ten failed answers from your system and place each one in the two by two. The pattern will tell you where to spend next sprint.

### Deeper: testing abstention

One level deeper on abstention. Build ten unanswerable questions for the HR assistant, like asking about a sabbatical policy that does not exist. A good system says it cannot find that policy and suggests contacting HR. Measure the share of correct abstentions. In the first run, the assistant invented a sabbatical policy for four of them, which became a blocking test.

### Watch me do it: diagnosing HR RAG

Watch me do it on the three-country HR assistant. First, labels. For thirty leave questions, an HR specialist marks the relevant chunk IDs, for example the UAE policy's leave section two for a Dubai employee's annual leave question. Second, I run retrieval only and compute metrics with the functions from the lesson. Recall at five averages around sixty percent; mean reciprocal rank is low. Third, I print one failing example. The Dubai question retrieved, in order, the UK leave section, the Pakistan leave section and then the UAE section at position three, plus two unrelated chunks. So recall at five was one, but for other questions the UAE chunk never made the top five at all. Fourth, generation metrics on the same set: faithfulness is high, the answers faithfully summarize whatever arrived. That confirms the diagnosis. Fifth, the two by two. Of twelve wrong answers, ten sit in the retrieval quadrant, relevant context missing. Two sit in the generation quadrant, where the right chunk arrived but the answer mixed in a UK detail. Sixth, one change at a time. I add a country metadata filter taken from the user's profile and rerun: recall at five rises sharply, and eight of the ten retrieval failures now answer correctly. Only then do I look at the two generation failures, with a prompt tweak, measured separately.

### Recap

Recap. Evaluate retrieval and generation separately, then end to end. Prioritize recall at k for retrieval. Measure faithfulness, relevance, context quality, correctness and abstention for generation. Use the two by two to locate failures, and record configuration with every result. Your next step: label relevant chunks for thirty questions, run the retrieval metrics code from the lesson, and place each failure in the two by two.

## Key takeaways

- Evaluate retrieval and generation separately, then end to end.
- Recall@k usually matters most for RAG; also track precision, MRR and nDCG.
- Measure faithfulness, answer relevance, context precision and recall, correctness and abstention.
- Use the context-vs-answer 2×2 to locate failures and record configuration with every result.

## Try it

Label relevant chunks for 30 questions, compute recall@5, precision@5 and MRR with the code provided, and classify every failure in the 2×2.

- [Previous: Groundedness, safety, privacy and fairness metrics](https://optimizeall.com/learn/llm-evals-and-observability/groundedness-safety-and-quality)
- [Next: Evaluating agents: outcomes, trajectories and efficiency](https://optimizeall.com/learn/llm-evals-and-observability/evaluating-agents-and-trajectories)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
