Latest AI Techniques: RAG, Tool Use, Agents & MCPRetrieval-augmented generation (RAG) · Lesson 4 of 20

Evaluating RAG systems

Article · 12 min · 9 min lecture

Video lecture

Evaluating RAG systems

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Evaluating RAG in parts

  • Four causes of a wrong answer
  • Test sets that catch real failures
  • Metrics that point to the fix

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Evaluate the parts, not just the whole

When a RAG answer is wrong, the cause is usually one of:

  1. Retrieval failure: the right passage never reached the model.
  2. Generation failure: the passage was there, but the model ignored, misread or embellished it.
  3. Source problem: the documents themselves are wrong, outdated or contradictory.
  4. Question problem: the question is ambiguous or out of scope.

End-to-end scores alone cannot tell these apart. Evaluate retrieval and generation separately, and tag failures by cause.

Building the test set

For each test question, record:

{
  "id": "hr-031",
  "question": "How many days do I have to submit a travel claim in Dubai?",
  "relevant_chunk_ids": ["hr-policy-2026#s4.2"],
  "reference_answer": "14 days from the travel date.",
  "tags": ["dubai", "expenses", "single-hop"]
}

Include:

  • Single-hop questions (one passage answers).
  • Multi-hop questions (the answer combines passages).
  • Unanswerable questions (nothing in the corpus answers; correct behaviour is to abstain).
  • Conflicting-source questions (old vs new policy).
  • Real user questions, with typos and informal phrasing, sampled from logs where privacy rules allow.

You can use a model to draft candidate questions from your documents to bootstrap the set, but have humans review them; synthetic questions tend to reuse document wording and make retrieval look easier than it is.

Retrieval metrics

  • Recall at k: the share of questions where at least one relevant chunk is in the top k. The most important retrieval metric, because the model cannot use what it never sees.
  • Precision at k: the share of the top k that is relevant. Low precision means noisy context.
  • Mean reciprocal rank (MRR): rewards putting the relevant chunk near the top.

These need the labelled relevant chunk IDs, which is why labelling is worth the effort.

Generation metrics

  • Faithfulness (groundedness): is every claim supported by the retrieved passages?
  • Answer correctness: does the answer match the reference?
  • Answer relevance: does it address the question asked?
  • Citation accuracy: do citations point to passages that actually support the claim?
  • Abstention quality: does it abstain on unanswerable questions and answer the answerable ones?

These are usually scored with a mix of code checks (citation IDs exist, quotes match) and LLM-as-judge prompts with rubrics, calibrated against human grading on a sample. Open-source evaluation frameworks provide ready-made versions of these metrics; understand what each metric's prompt actually checks before trusting its numbers.

A diagnostic table

Run the full set and fill a table like this (illustrative):

TagRecall@5FaithfulnessCorrectnessMain failure
single-hophighhighhigh-
multi-hopmediumhighlowsecond passage not retrieved
unanswerablen/an/alowanswers instead of abstaining
dubailowhighlowoffice filter missing

Now you know where to work: multi-hop retrieval, abstention instructions, and metadata filtering, rather than tweaking the answer prompt blindly.

Online evaluation

Offline tests are necessary but not sufficient. In production, track:

  • user feedback (thumbs, follow-up "that's wrong" messages, escalations)
  • abstention rate and "no results" rate
  • citation click-through, if shown
  • sampled human review of real conversations, stratified by topic

Watch for drift: new documents, new products and new question types appear constantly.

Worked example

An e-commerce support RAG shows high end-to-end satisfaction but a spike in refunds complaints. Tagged evaluation reveals the returns policy was updated but the old version was still indexed and often ranked higher. The generation step was faithful, faithfully citing the outdated source. The fix was index hygiene (remove superseded versions) and a date-based precedence rule, not a prompt change. Without separating retrieval, generation and source quality, the team would have kept rewriting the prompt.

Failure modes in evaluation

  • Test questions written by people who know the documents, so they mirror the wording.
  • No unanswerable questions, so a system that never abstains scores well.
  • Judges that reward fluent answers without checking support.
  • Evaluating once at launch and never again.

Hands-on: a retrieval and faithfulness harness

Start with retrieval, because it needs no model at all: just your labelled test set and your search function.

import json, statistics

def recall_at_k(results, relevant, k):
    return any(r in relevant for r in results[:k])

def reciprocal_rank(results, relevant):
    for i, r in enumerate(results, start=1):
        if r in relevant:
            return 1 / i
    return 0.0

def eval_retrieval(test_set, search, k=5):
    rows = []
    for case in test_set:
        if not case["relevant_chunk_ids"]:      # unanswerable: skip retrieval metrics
            continue
        got = search(case["question"])           # returns a ranked list of chunk IDs
        rows.append({"id": case["id"], "tags": case["tags"],
                     "hit": recall_at_k(got, case["relevant_chunk_ids"], k),
                     "rr": reciprocal_rank(got, case["relevant_chunk_ids"])})
    by_tag = {}
    for r in rows:
        for t in r["tags"]:
            by_tag.setdefault(t, []).append(r["hit"])
    print(f"recall@{k}:", statistics.mean(r["hit"] for r in rows),
          "MRR:", round(statistics.mean(r["rr"] for r in rows), 3))
    for t, hits in sorted(by_tag.items()):
        print(f"  {t:15} recall@{k}={statistics.mean(hits):.2f} (n={len(hits)})")
    return rows

test_set = [json.loads(line) for line in open("rag_tests.jsonl", encoding="utf-8")]

Then grade generation with a rubric-based judge. Keep the judge's job narrow and binary: it is far easier to calibrate against humans than a 1 to 10 score.

You are grading an answer for FAITHFULNESS.
<passages>{{PASSAGES}}</passages>
<answer>{{ANSWER}}</answer>
List each factual claim in the answer. For each, write SUPPORTED or UNSUPPORTED
based only on the passages. Then output a final line: VERDICT: PASS if every claim
is supported, otherwise VERDICT: FAIL.

Calibrate before you trust it: have two people grade 30 to 50 answers, compare with the judge, and read every disagreement. Adjust the rubric until agreement is high, and re-check whenever you change the judge model.

Wiring it into releases

Save each run's metrics by tag with the versions that produced them (index build, embedding model, reranker, prompt, answer model). A simple CI rule such as "block the release if recall@5 on any tag drops by more than a few points, or faithfulness falls below the agreed threshold" turns evaluation from a launch ritual into a safety net. The flagship course Evaluating and Monitoring LLM Applications goes much deeper into judges, datasets, tracing and online evaluation.

Going further

Automate evaluation in your deployment pipeline: every change to chunking, embeddings, retrieval parameters, reranker or prompt runs the suite and reports metrics by tag, with thresholds that block a release if recall or faithfulness drops.

Key takeaways

  • Separate retrieval, generation, source and question failures; end-to-end scores hide causes.
  • Test sets need single-hop, multi-hop, unanswerable, conflicting-source and real user questions.
  • Measure recall at k, precision and MRR for retrieval; faithfulness, correctness, citation accuracy and abstention for generation.
  • Combine offline suites with online signals, and re-run on every pipeline change.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. The right passage is retrieved but the answer adds a claim not in it. Which component failed?
  2. Why include unanswerable questions in a RAG test set?
  3. A model faithfully cites an outdated policy that is still indexed. What is the right fix?

Put it into practice

Create 20 test questions for a knowledge base (including 4 unanswerable and 4 multi-hop) with relevant chunk IDs. Measure recall at 5 before changing anything else.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.