---
title: "Evaluating RAG systems | Optimize All Academy"
description: "Evaluate the parts, not just the whole When a RAG answer is wrong, the cause is usually one of: 1. Retrieval failure: the right passage never reached the…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/evaluating-rag
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Retrieval-augmented generation (RAG) · lesson 4 of 20 · 12 min

# Evaluating RAG systems

## Evaluate the parts, not just the whole

When a RAG answer is wrong, the cause is usually one of:

1. **Retrieval failure:** the right passage never reached the model.
2. **Generation failure:** the passage was there, but the model ignored, misread or embellished it.
3. **Source problem:** the documents themselves are wrong, outdated or contradictory.
4. **Question problem:** the question is ambiguous or out of scope.

End-to-end scores alone cannot tell these apart. Evaluate retrieval and generation separately, and tag failures by cause.

## Building the test set

For each test question, record:

```json
{
  "id": "hr-031",
  "question": "How many days do I have to submit a travel claim in Dubai?",
  "relevant_chunk_ids": ["hr-policy-2026#s4.2"],
  "reference_answer": "14 days from the travel date.",
  "tags": ["dubai", "expenses", "single-hop"]
}
```

Include:

- **Single-hop** questions (one passage answers).
- **Multi-hop** questions (the answer combines passages).
- **Unanswerable** questions (nothing in the corpus answers; correct behaviour is to abstain).
- **Conflicting-source** questions (old vs new policy).
- **Real user questions**, with typos and informal phrasing, sampled from logs where privacy rules allow.

You can use a model to draft candidate questions from your documents to bootstrap the set, but have humans review them; synthetic questions tend to reuse document wording and make retrieval look easier than it is.

## Retrieval metrics

- **Recall at k:** the share of questions where at least one relevant chunk is in the top k. The most important retrieval metric, because the model cannot use what it never sees.
- **Precision at k:** the share of the top k that is relevant. Low precision means noisy context.
- **Mean reciprocal rank (MRR):** rewards putting the relevant chunk near the top.

These need the labelled relevant chunk IDs, which is why labelling is worth the effort.

## Generation metrics

- **Faithfulness (groundedness):** is every claim supported by the retrieved passages?
- **Answer correctness:** does the answer match the reference?
- **Answer relevance:** does it address the question asked?
- **Citation accuracy:** do citations point to passages that actually support the claim?
- **Abstention quality:** does it abstain on unanswerable questions and answer the answerable ones?

These are usually scored with a mix of code checks (citation IDs exist, quotes match) and LLM-as-judge prompts with rubrics, calibrated against human grading on a sample. Open-source evaluation frameworks provide ready-made versions of these metrics; understand what each metric's prompt actually checks before trusting its numbers.

## A diagnostic table

Run the full set and fill a table like this (illustrative):

| Tag | Recall@5 | Faithfulness | Correctness | Main failure |
|---|---|---|---|---|
| single-hop | high | high | high | - |
| multi-hop | medium | high | low | second passage not retrieved |
| unanswerable | n/a | n/a | low | answers instead of abstaining |
| dubai | low | high | low | office filter missing |

Now you know where to work: multi-hop retrieval, abstention instructions, and metadata filtering, rather than tweaking the answer prompt blindly.

## Online evaluation

Offline tests are necessary but not sufficient. In production, track:

- user feedback (thumbs, follow-up "that's wrong" messages, escalations)
- abstention rate and "no results" rate
- citation click-through, if shown
- sampled human review of real conversations, stratified by topic

Watch for drift: new documents, new products and new question types appear constantly.

## Worked example

An e-commerce support RAG shows high end-to-end satisfaction but a spike in refunds complaints. Tagged evaluation reveals the returns policy was updated but the old version was still indexed and often ranked higher. The generation step was faithful, faithfully citing the outdated source. The fix was index hygiene (remove superseded versions) and a date-based precedence rule, not a prompt change. Without separating retrieval, generation and source quality, the team would have kept rewriting the prompt.

## Failure modes in evaluation

- Test questions written by people who know the documents, so they mirror the wording.
- No unanswerable questions, so a system that never abstains scores well.
- Judges that reward fluent answers without checking support.
- Evaluating once at launch and never again.

## Hands-on: a retrieval and faithfulness harness

Start with retrieval, because it needs no model at all: just your labelled test set and your search function.

```python
import json, statistics

def recall_at_k(results, relevant, k):
    return any(r in relevant for r in results[:k])

def reciprocal_rank(results, relevant):
    for i, r in enumerate(results, start=1):
        if r in relevant:
            return 1 / i
    return 0.0

def eval_retrieval(test_set, search, k=5):
    rows = []
    for case in test_set:
        if not case["relevant_chunk_ids"]:      # unanswerable: skip retrieval metrics
            continue
        got = search(case["question"])           # returns a ranked list of chunk IDs
        rows.append({"id": case["id"], "tags": case["tags"],
                     "hit": recall_at_k(got, case["relevant_chunk_ids"], k),
                     "rr": reciprocal_rank(got, case["relevant_chunk_ids"])})
    by_tag = {}
    for r in rows:
        for t in r["tags"]:
            by_tag.setdefault(t, []).append(r["hit"])
    print(f"recall@{k}:", statistics.mean(r["hit"] for r in rows),
          "MRR:", round(statistics.mean(r["rr"] for r in rows), 3))
    for t, hits in sorted(by_tag.items()):
        print(f"  {t:15} recall@{k}={statistics.mean(hits):.2f} (n={len(hits)})")
    return rows

test_set = [json.loads(line) for line in open("rag_tests.jsonl", encoding="utf-8")]
```

Then grade generation with a rubric-based judge. Keep the judge's job narrow and binary: it is far easier to calibrate against humans than a 1 to 10 score.

```text
You are grading an answer for FAITHFULNESS.
<passages>{{PASSAGES}}</passages>
<answer>{{ANSWER}}</answer>
List each factual claim in the answer. For each, write SUPPORTED or UNSUPPORTED
based only on the passages. Then output a final line: VERDICT: PASS if every claim
is supported, otherwise VERDICT: FAIL.
```

Calibrate before you trust it: have two people grade 30 to 50 answers, compare with the judge, and read every disagreement. Adjust the rubric until agreement is high, and re-check whenever you change the judge model.

## Wiring it into releases

Save each run's metrics by tag with the versions that produced them (index build, embedding model, reranker, prompt, answer model). A simple CI rule such as "block the release if recall@5 on any tag drops by more than a few points, or faithfulness falls below the agreed threshold" turns evaluation from a launch ritual into a safety net. The flagship course **Evaluating and Monitoring LLM Applications** goes much deeper into judges, datasets, tracing and online evaluation.

## Going further

Automate evaluation in your deployment pipeline: every change to chunking, embeddings, retrieval parameters, reranker or prompt runs the suite and reports metrics by tag, with thresholds that block a release if recall or faithfulness drops.

## Video lecture: Evaluating RAG systems

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Evaluating RAG in parts
2. Analogy: diagnosing a car
3. Four failure causes
4. Build the test set
5. Simple example: 5 questions
6. Retrieval metrics
7. Generation metrics
8. Read results by tag
9. Business example (illustrative)
10. Hands-on + operations
11. Common mistakes
12. Is your evaluation good?
13. Watch me do it: the evaluation harness
14. Recap
15. Try this now (30 minutes)

## Lecture transcript

### Evaluating RAG in parts

Your RAG assistant gave a customer the wrong returns window. Whose fault was it? The search, the model, the documents, or the question? If you can't answer that in minutes, you'll spend weeks rewriting prompts that were never the problem. In this lesson you'll learn to evaluate RAG in parts, build a test set that actually catches failures, and run a harness that tells you exactly where to work.

### Analogy: diagnosing a car

Why does evaluating in parts matter? Think of a car that won't start. You wouldn't replace the engine before checking the battery and the fuel. Yet that's what teams do when they rewrite prompts every time an AI answer is wrong. Evaluation is your diagnostic kit. Recall tells you whether the fuel reached the engine, meaning whether the right passage reached the model. Faithfulness tells you whether the engine used it properly.

### Four failure causes

There are four usual culprits. A retrieval failure means the right passage never reached the model. A generation failure means it was there, but the model ignored it, misread it, or added something. A source problem means the documents themselves are wrong, outdated or contradictory. And a question problem means it was ambiguous or out of scope. An end-to-end score blends all four together, so evaluate retrieval and generation separately, and tag every failure with its cause.

### Build the test set

Your test set is the foundation. For each question, record the question, the IDs of the chunks that contain the answer, a reference answer, and tags. Include single-hop questions answered by one passage, multi-hop questions that combine passages, conflicting-source questions where old and new policies disagree, and, critically, unanswerable questions where the right behaviour is to say I don't know. Add real user questions with typos and informal phrasing. You can use a model to draft candidate questions, but review them, because synthetic questions tend to copy document wording and flatter your retrieval.

### Simple example: 5 questions

Here's a simple example with five test questions. For four of them, the right chunk appears in the top five results. For one, it doesn't. That's recall at five of four out of five, or eighty percent. Now look at the four answers where the right chunk was present. Three are fully supported by the passages; one adds a delivery time that isn't in any source. That's a faithfulness failure, not a retrieval failure, and it needs a different fix.

### Retrieval metrics

For retrieval, the key metric is recall at k: for what share of questions does at least one relevant chunk appear in the top k? The model can't use what it never sees, so this comes first. Precision at k tells you how noisy the context is, and mean reciprocal rank rewards putting the right chunk near the top. None of these need a language model. They just need your labels and your search function, which makes them cheap to run on every change.

### Generation metrics

For generation, measure faithfulness, meaning every claim is supported by the passages; correctness against the reference; relevance to the question; citation accuracy; and abstention quality, which means declining unanswerable questions while still answering the answerable ones. Use code for objective checks, like whether cited IDs exist and quotes match. Use an LLM judge with a narrow, binary rubric for the rest, and calibrate it against human grading on a sample before you trust its numbers.

### Read results by tag

Now the payoff: a diagnostic table by tag. Maybe single-hop is strong, multi-hop has weak recall because the second passage never arrives, unanswerable questions get answered anyway, and everything tagged Dubai fails because the office filter is missing. Each row points at a different fix. Here's a real-world style example. An online shop's assistant kept quoting an old returns window. Faithfulness was perfect. It faithfully cited an outdated document that was still indexed. The fix was index hygiene, not the prompt.

### Business example (illustrative)

A deeper business example, illustrative. A Karachi fintech's support assistant scored well on average, yet complaints kept arriving about card limits. Tagging the test set by topic showed card-limit questions had recall at five of only forty percent, because limits lived in a PDF table that was never chunked with its headers. Fixing that one document moved the tag to ninety percent and the complaints stopped. Without tags, that problem was hidden inside a healthy-looking average.

### Hands-on + operations

The hands-on section gives you a small Python harness that computes recall at k and mean reciprocal rank by tag, plus a faithfulness judge prompt. Run it before you change anything, so you have a baseline. Then connect it to your release process: when chunking, embeddings, the reranker or the prompt changes, the suite runs and blocks the release if key metrics drop. Offline tests aren't the end, though. In production, watch feedback, abstention rates and sampled human reviews, because new documents and new questions arrive every week.

### Common mistakes

Common evaluation mistakes. Writing test questions from the documents themselves, so they mirror the wording. Leaving out unanswerable questions, so a system that never says I don't know scores well. Trusting an LLM judge you've never compared with human grading. Reporting one average score instead of results by tag. And evaluating once at launch, then never again, while your documents and users keep changing.

### Is your evaluation good?

How will you know your evaluation itself is good? Your test set includes every tag you care about, including unanswerable and multi-hop questions. Your judge agrees with human graders on a sample most of the time, and you've read every disagreement. And your metrics predict reality: when offline scores improve, user feedback and escalations improve too. If they don't move together, your test set is missing something real users do.

### Watch me do it: the evaluation harness

Watch me do it. I open the harness. First, recall at k checks whether any retrieved ID is in the relevant set. Reciprocal rank returns one over the position of the first relevant chunk. Next, eval retrieval loops over the test set, skipping unanswerable cases, calls my search function, and stores hit and reciprocal rank with the case's tags. Then I group hits by tag and print recall at five and mean reciprocal rank overall, then per tag. I load twenty cases from the JSON lines file and run it. The output shows recall at five of point eight, and a per-tag list where multi-hop sits at point five. Next, I paste one multi-hop answer into the faithfulness judge prompt. It lists each claim as supported or unsupported and ends with verdict fail, because one claim cites a second document we never retrieved. That tells me the fix is retrieval, not the prompt.

### Recap

To recap: separate retrieval, generation, source and question failures. Build a test set with hard cases and unanswerable questions. Measure recall first, then faithfulness and abstention, with calibrated judges. Read results by tag and fix the biggest cause. Your next step is to write twenty test questions for one knowledge base, including four unanswerable and four multi-hop, and measure recall at five before touching anything else. When you're ready for more depth, the Evaluating and Monitoring LLM Applications course takes this much further.

### Try this now (30 minutes)

Try this now. Create a simple sheet with twenty questions for one knowledge base: twelve normal ones, four that need two documents, and four that nothing in your documents answers. For each answerable question, write the ID of the chunk that contains the answer. Then run your search and count how many land in the top five. That number, measured before you change anything, is your baseline, and every improvement from now on can be proven against it.

## Key takeaways

- Separate retrieval, generation, source and question failures; end-to-end scores hide causes.
- Test sets need single-hop, multi-hop, unanswerable, conflicting-source and real user questions.
- Measure recall at k, precision and MRR for retrieval; faithfulness, correctness, citation accuracy and abstention for generation.
- Combine offline suites with online signals, and re-run on every pipeline change.

## Try it

Create 20 test questions for a knowledge base (including 4 unanswerable and 4 multi-hop) with relevant chunk IDs. Measure recall at 5 before changing anything else.

- [Previous: The RAG pipeline: retrieval, reranking and citations](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/rag-pipeline)
- [Next: Advanced RAG patterns: agentic, graph, visual and long-context](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/advanced-rag-patterns)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
