Latest AI Techniques: RAG, Tool Use, Agents & MCPRetrieval-augmented generation (RAG) · Lesson 3 of 20

The RAG pipeline: retrieval, reranking and citations

Article · 13 min · 9 min lecture

Video lecture

The RAG pipeline: retrieval, reranking and citations

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Retrieval-augmented generation (RAG)

  • Look up your sources at question time
  • Answer from them, with citations
  • Six steps, six failure points

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What RAG is and why it exists

Retrieval-augmented generation (RAG) retrieves relevant information from your own sources at question time and gives it to the model as context for its answer. It addresses three limits of models on their own: they do not know your private data, their training knowledge has a cut-off date, and they may hallucinate when unsure. RAG does not eliminate hallucination, but grounding answers in retrieved sources, with citations, makes errors less frequent and much easier to detect.

The pipeline

User question
  -> 1. Query understanding / rewriting
  -> 2. Retrieval (hybrid search + metadata filters)   -> top 20-50 candidates
  -> 3. Reranking                                       -> top 3-8 passages
  -> 4. Prompt assembly (passages + instructions)
  -> 5. Generation with citations
  -> 6. Post-checks (citation verification, abstention handling)

Step 1: query understanding

User questions are often short, ambiguous or conversational ("what about for contractors?"). Improvements:

  • Contextualise follow-ups: rewrite using the conversation into a standalone query ("What is the travel reimbursement deadline for contractors?").
  • Query expansion: generate alternative phrasings or sub-questions and retrieve for each.
  • Hypothetical answer embedding: have a model draft a plausible answer and embed that, since answers often resemble documents more than questions do. Useful, but test it; it can pull retrieval toward the model's assumptions.
  • Routing: decide which index or tool to use (policies vs product docs vs a database).

Step 2: retrieval

Use hybrid search with metadata filters (permissions, region, product, date). Retrieve generously at this stage (for example 20 to 50 candidates), because the next step will be more selective.

Step 3: reranking

A reranker (often a cross-encoder model, or an LLM prompted to score relevance) looks at the query and each candidate together and scores relevance more accurately than vector similarity alone. It is slower per item, which is why it runs only on the candidate set. Reranking is frequently one of the highest-impact upgrades to a basic RAG system.

Step 4: prompt assembly

<sources>
  <source id="S1" title="Travel Policy 2026 s4.2" date="2026-01-01">...</source>
  <source id="S2" title="Contractor Handbook s2.1" date="2025-09-10">...</source>
</sources>

<instructions>
Answer using only the sources. Cite source ids in square brackets after
each claim, e.g. [S1]. If sources conflict, prefer the most recent official
policy and mention the conflict. If the sources do not contain the answer,
say so and suggest who to contact.
</instructions>

<question>{{standalone_question}}</question>

Order passages by relevance, include titles and dates, and keep total context lean: more passages are not always better, since irrelevant text can distract the model.

Step 5: generation with citations

Ask for claim-level citations. Several model APIs now offer native citation features that return exact source spans (the hands-on section below uses one); these remove fragile bracket-parsing and make verification straightforward. Either way, citations are for the user's trust and for your verification.

Step 6: post-checks

  • Verify every cited ID exists and, ideally, that quoted text appears in that source.
  • Detect abstentions and route them usefully (search suggestion, human handoff).
  • Log question, retrieved IDs, reranker scores, answer and citations for evaluation.

Worked example: an internal policy assistant

A regional firm with offices in London, Dubai and Karachi builds an HR assistant. Early problem: answers cite the London policy to Karachi staff. Fix: office metadata filter from the user's profile, plus an instruction to state which office the policy applies to. Second problem: follow-up questions ("and for part-timers?") retrieve nothing useful. Fix: rewrite follow-ups into standalone queries. Third: correct passages retrieved at rank 15 never reach the model. Fix: a reranker over the top 40. Each fix is targeted at a measured failure, not guessed.

When RAG is the wrong tool

  • The answer requires aggregating over structured data ("total sales by region last quarter"): use a database query (text-to-SQL or a tool), not passage retrieval.
  • The whole corpus is small enough to fit comfortably in context and changes rarely: sometimes simply including it (with caching) is simpler and better.
  • The task needs the model to learn a style or skill rather than facts: examples or fine-tuning may fit better.

Hands-on: a minimal cited RAG answerer with the Claude API

This example takes passages from your retriever (for instance the hybrid search from module 1, plus a reranker) and asks Claude to answer with native citations. Citations come back as structured data pointing at the exact text used, so you can verify them in code rather than parsing brackets.

pip install anthropic sentence-transformers
export ANTHROPIC_API_KEY=...        # never hard-code keys
export ANTHROPIC_MODEL=claude-opus-5  # check the current model list in the docs
import os
import anthropic
from sentence_transformers import CrossEncoder

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")  # open-source cross-encoder reranker

def rerank(question, candidates, top_n=5):
    scores = reranker.predict([(question, c["text"]) for c in candidates])
    ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
    return [c for _, c in ranked[:top_n]]

def answer(question, candidates):
    passages = rerank(question, candidates)
    content = [
        {"type": "document",
         "source": {"type": "text", "media_type": "text/plain", "data": p["text"]},
         "title": f'{p["title"]} (updated {p["date"]})',
         "citations": {"enabled": True}}
        for p in passages
    ]
    content.append({"type": "text", "text":
        "Answer the question using only the documents. If they conflict, prefer the most "
        "recent official policy and mention the conflict. If they do not contain the answer, "
        "say so and suggest contacting HR.\n\nQuestion: " + question})
    try:
        resp = client.messages.create(model=MODEL, max_tokens=2000,
                                      messages=[{"role": "user", "content": content}])
    except anthropic.RateLimitError:
        return {"answer": None, "error": "rate_limited_retry_later"}
    except anthropic.APIStatusError as e:
        return {"answer": None, "error": f"api_error_{e.status_code}"}
    parts, cites = [], []
    for block in resp.content:
        if block.type == "text":
            parts.append(block.text)
            for c in (block.citations or []):
                cites.append({"doc": c.document_title, "quote": c.cited_text})
    return {"answer": "".join(parts), "citations": cites, "passages": [p["id"] for p in passages]}

Things to check when you run it:

  • Every claim should carry a citation. Log answers where text blocks have no citations; they are candidates for unsupported claims.
  • Abstention works. Ask something the passages do not cover and confirm the answer says so.
  • Log the passage IDs you sent, so evaluation (next lesson) can separate retrieval failures from generation failures.

Other providers offer comparable patterns (file search tools, grounding with citations); the architecture stays the same: retrieve generously, rerank, send a lean set of passages with titles and dates, and verify citations.

Going further

Advanced variants include graph-based retrieval (building entity-relationship graphs over documents for multi-hop questions), agentic RAG (the model decides when and what to search, iteratively), and multi-vector representations. Each adds complexity; adopt them when your evaluation shows the simpler pipeline failing on a class of questions they are designed for.

Key takeaways

  • RAG retrieves from your sources at question time to ground answers and enable citations.
  • Pipeline: query rewriting, hybrid retrieval with filters, reranking, lean prompt assembly, cited generation, post-checks.
  • Reranking a generous candidate set is often one of the highest-impact upgrades.
  • Use databases for aggregation questions, and consider full-context when the corpus is small.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A user asks 'and for contractors?' after a policy question, and retrieval fails. What is the best fix?
  2. Why retrieve 30-50 candidates before reranking to a handful?
  3. 'What were total sales by region last quarter?' is best answered by:

Put it into practice

Sketch the six-step pipeline for a knowledge base you know. For each step, write the one failure you expect most and how you would detect it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.