Evaluating and Monitoring LLM ApplicationsEvaluating RAG pipelines and agents · Lesson 7 of 16

Evaluating RAG: retrieval, generation and abstention

Article · 15 min · 8 min lecture

Video lecture

Evaluating RAG: retrieval, generation and abstention

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Evaluating RAG

  • Two failure points
  • Retrieval metrics
  • Generation metrics
  • The 2×2 diagnosis

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Evaluate the pipeline in parts

A retrieval-augmented generation (RAG) system fails in two places: retrieval (the right information never reached the model) and generation (the model had the right information but misused it). If you only score final answers, you cannot tell which part to fix. Evaluate each stage separately, then end to end.

question -> [query rewrite] -> [retrieve top-k] -> [rerank] -> [generate with context] -> answer
               ^ eval: rewrite quality   ^ eval: recall/precision/rank   ^ eval: faithfulness, relevance, correctness

Retrieval metrics

You need, for each test question, the IDs of relevant chunks or documents (labeled by experts, or bootstrapped with a judge and then reviewed).

MetricQuestion it answers
Recall@kOf the relevant items, what share appear in the top k? (Did we find it at all?)
Precision@kOf the top k, what share are relevant? (How much noise did we add?)
Hit rate@kDid at least one relevant item appear in the top k?
MRR (mean reciprocal rank)How high is the first relevant item, on average?
nDCG@kAre the most relevant items ranked highest, with graded relevance?

For RAG, recall@k usually matters most: the generator cannot use what was never retrieved. Precision matters for cost and for avoiding distracting context.

# retrieval_metrics.py
import math

def recall_at_k(retrieved: list[str], relevant: set[str], k: int) -> float:
    return len(set(retrieved[:k]) & relevant) / len(relevant) if relevant else 1.0

def precision_at_k(retrieved, relevant, k):
    top = retrieved[:k]
    return len([d for d in top if d in relevant]) / len(top) if top else 0.0

def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, start=1):
        if d in relevant:
            return 1.0 / i
    return 0.0

def ndcg_at_k(retrieved, grades: dict[str, int], k: int) -> float:
    dcg = sum((2 ** grades.get(d, 0) - 1) / math.log2(i + 1) for i, d in enumerate(retrieved[:k], start=1))
    ideal = sorted(grades.values(), reverse=True)[:k]
    idcg = sum((2 ** g - 1) / math.log2(i + 1) for i, g in enumerate(ideal, start=1))
    return dcg / idcg if idcg else 0.0

Generation metrics

Frameworks such as Ragas popularized a vocabulary that is now common:

  • Faithfulness / groundedness: are answer claims supported by the retrieved context? (Module 2.)
  • Answer relevance: does the answer address the question asked?
  • Context precision: is the retrieved context focused on what is needed?
  • Context recall: does the retrieved context contain everything needed for the reference answer?
  • Answer correctness: does it match the reference or key points?

Also measure abstention: when the answer is not in the knowledge base, does the system say so instead of guessing? Build a set of unanswerable questions for this.

Diagnosing with a 2×2

Answer correctAnswer wrong
Relevant context retrievedWorkingGeneration problem (prompt, model, context too long, conflicting docs)
Relevant context missingLucky or using general knowledge (check groundedness policy)Retrieval problem (chunking, embeddings, filters, query rewriting)

Tagging each failed test with its quadrant tells you where to invest.

Worked example: tuning a policy assistant

An HR assistant for a company with offices in Lahore, Dubai and London answered questions from policy PDFs. End-to-end correctness was 74% (illustrative). Stage metrics showed recall@5 of only 61% on questions about leave, while faithfulness was high. The problem was retrieval: policies for three countries were near-identical text, and chunks lacked country metadata. Adding a country filter from the user profile and including section headings in chunks raised recall@5 substantially, and end-to-end correctness followed. No prompt change was needed.

Hands-on: a RAG eval record

Store every test result with enough detail to diagnose:

{
  "id": "hr-leave-017",
  "question": "How many days of annual leave do I get in the Dubai office?",
  "slice": {"country": "AE", "topic": "leave", "answerable": true},
  "relevant_ids": ["ae-policy#leave-2"],
  "retrieved_ids": ["gb-policy#leave-1", "ae-policy#leave-2", "pk-policy#leave-3"],
  "metrics": {"recall@5": 1.0, "precision@5": 0.33, "mrr": 0.5,
              "faithfulness": 1.0, "correct": true, "abstained": false},
  "answer": "...",
  "config": {"embedder": "v3", "chunking": "heading-aware-800", "reranker": "on", "prompt": "hr-v12"}
}

With the configuration recorded, you can compare chunking strategies, embedders and rerankers directly.

Pitfalls

  • Only end-to-end scores, leaving you guessing which stage broke.
  • Relevance labels generated by the retriever you are testing (circular). Review them.
  • No unanswerable questions, so you never measure abstention.
  • Changing several components at once; change one, measure, then the next.

How to measure success

Stage-level metrics for every release, a 2×2 diagnosis of failures, and improvements traced to specific component changes.

Key takeaways

  • Evaluate retrieval and generation separately, then end to end.
  • Recall@k usually matters most for RAG; also track precision, MRR and nDCG.
  • Measure faithfulness, answer relevance, context precision and recall, correctness and abstention.
  • Use the context-vs-answer 2×2 to locate failures and record configuration with every result.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Faithfulness is high but end-to-end correctness is low, and recall@5 is poor. Where should you focus?
  2. Why build a set of unanswerable questions?

Put it into practice

Label relevant chunks for 30 questions, compute recall@5, precision@5 and MRR with the code provided, and classify every failure in the 2×2.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.