Latest AI Techniques: RAG, Tool Use, Agents & MCPRetrieval-augmented generation (RAG) · Lesson 3 of 20
The RAG pipeline: retrieval, reranking and citations
Video lecture
The RAG pipeline: retrieval, reranking and citations
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Retrieval-augmented generation (RAG)
Ask a general AI model about your company's travel policy and it will do one of two things. It will admit it doesn't know, or worse, it will confidently invent something plausible. Retrieval-augmented generation, or RAG, fixes this by looking things up in your own sources at question time and handing them to the model, with instructions to answer from those sources and cite them. In this lesson you'll walk through the six steps of a production RAG pipeline, and see where each one typically breaks.
0:37 Analogy: the open-book exam
Here's an analogy that makes RAG click. Imagine an open-book exam. A student with no book has to rely on memory and may bluff. A student with the book, and good bookmarks, can look up the answer and quote the page. RAG turns your AI into the second student. Retrieval is finding the right pages, reranking is choosing the best bookmarks, and citations are quoting the page number so the examiner, your user, can check.
1:10 Why RAG?
RAG exists because models have three limits. They don't know your private data. Their knowledge stops at a training cut-off. And when they're unsure, they may make things up. Grounding answers in retrieved passages doesn't eliminate that last problem, but it makes mistakes rarer and, crucially, easier to check, because every claim points to a source a human can open.
1:36 Steps 1–2
Step one is understanding the question. Real users ask follow-ups like, and what about contractors? That means nothing to a search engine on its own, so you rewrite it into a standalone query using the conversation. You might also expand it into alternative phrasings, or route it to the right index, policies versus product docs versus a database. Step two is retrieval: hybrid search with metadata filters for permissions, region and date, pulling a generous set of candidates, say twenty to fifty.
2:12 Steps 3–4
Step three is reranking, and it's often the biggest single upgrade you can make. A reranker, usually a cross-encoder model, reads the question and each candidate together and scores relevance much more accurately than vector similarity. It's slower per item, which is why you only rerank the candidate set, then keep the best three to eight passages. Step four is prompt assembly: order passages by relevance, include titles and dates, and tell the model what to do when sources conflict or don't contain the answer. More text is not better. Irrelevant passages distract.
2:52 Steps 5–6
Step five is generation with citations. Ask for claim-level citations. Several model APIs, including Claude's, now return citations natively as structured data pointing at the exact text used, which beats parsing square brackets out of prose. Step six is post-checks. Verify every citation points at a real passage, detect when the model abstained and route that usefully, perhaps to a human or a search page, and log the question, the passage IDs and the answer so you can evaluate later.
3:27 Simple example: 'How long do refunds take?'
First, a simple example. A customer asks, how long do refunds take? The retriever finds the refunds section of your policy and an old FAQ. The reranker ranks the current policy first. The prompt includes both with dates and says prefer the most recent. The model answers, refunds are processed within five working days of receiving the item, and cites the policy. Four steps, one clear answer, and a citation the customer can click.
3:59 Worked example: a three-office HR assistant
Let's make it concrete. A firm with offices in London, Dubai and Karachi builds an HR assistant. First problem: Karachi staff get London's policy. Fix: filter by office from the user's profile. Second: follow-up questions retrieve nothing. Fix: rewrite follow-ups into standalone queries. Third: the right passage sits at rank fifteen and never reaches the model. Fix: rerank the top forty. Notice the pattern. Each fix targets a measured failure. Nobody guessed.
4:30 Business example (illustrative)
Let's put numbers on the HR assistant, labelled illustrative. Before the fixes, HR received around a hundred and twenty policy questions a week by email, and the assistant answered correctly on about sixty percent of a test set. After the office filter, follow-up rewriting and reranking, test accuracy rose to about eighty-five percent, and HR saw email questions fall by roughly half. The remaining questions were mostly about individual cases, which the assistant correctly routed to a person.
5:04 When RAG is the wrong tool
And know when RAG is the wrong tool. Questions like total sales by region last quarter need a database query, not passage retrieval. If your whole corpus is small and stable, it may be simpler to put all of it in the prompt with caching. And if you want the model to learn a style rather than facts, examples or fine-tuning fit better. In the lesson you'll find a hands-on script that reranks passages with an open-source cross-encoder and asks Claude for an answer with native citations, returning the quotes it relied on.
5:44 Common mistakes
Now the common mistakes. Sending far too many passages, so the model is distracted by noise. Skipping query rewriting, so follow-ups fail. Relying on vector search alone for product codes and names. Asking for citations but never checking them. And treating RAG as the answer to every question, including ones that really need a database query. If you remember one thing, make it this: measure which step fails before you change anything.
6:15 How you'll know it's working
How will you know the pipeline works? Three signals. Answers carry valid citations that actually support the claims. Users stop asking the same question twice. And your abstentions happen on genuinely unanswerable questions, not on ones your documents cover. Log the passage IDs you send with every answer, so when a user flags a bad answer you can tell in a minute which step failed.
6:43 Watch me do it: the cited answerer
Watch me do it. I open the cited answerer. First, rerank: the cross-encoder scores each question and passage pair, and I keep the top five. Next, I build the content list: one document block per passage, with the text as the source, a title that includes the updated date, and citations enabled. Then I append the instruction text: answer only from the documents, prefer the most recent policy, and say so if the answer isn't there. I call messages dot create inside a try block, catching rate-limit and API errors separately. Finally, I loop over the response blocks. For every text block I collect the text, and for each citation I record the document title and the cited text. I run it with the question, how many days do I have to submit a travel claim in Dubai? The answer says fourteen days, and the citation shows the exact sentence from the Dubai travel policy.
7:50 Recap
To recap: RAG retrieves from your sources at question time. The pipeline is rewrite, retrieve generously, rerank, assemble a lean prompt, generate with citations, and check. Fix measured failures one at a time. Your next step is to sketch these six steps for a knowledge base you know, and write down the most likely failure at each. In the next lesson, you'll learn how to evaluate RAG properly, so you know which step to fix.
8:23 Try this now (15 minutes)
Try this now. Pick one knowledge base you know well. Write down the six pipeline steps on a single page, and next to each, the failure you think is most likely and how you'd detect it. Then run one real question through the hands-on script with three to five passages from that knowledge base, and check every citation it returns against the source text. Fifteen minutes, and you'll have a map of where to focus next.
What RAG is and why it exists
Retrieval-augmented generation (RAG) retrieves relevant information from your own sources at question time and gives it to the model as context for its answer. It addresses three limits of models on their own: they do not know your private data, their training knowledge has a cut-off date, and they may hallucinate when unsure. RAG does not eliminate hallucination, but grounding answers in retrieved sources, with citations, makes errors less frequent and much easier to detect.
The pipeline
User question
-> 1. Query understanding / rewriting
-> 2. Retrieval (hybrid search + metadata filters) -> top 20-50 candidates
-> 3. Reranking -> top 3-8 passages
-> 4. Prompt assembly (passages + instructions)
-> 5. Generation with citations
-> 6. Post-checks (citation verification, abstention handling)Step 1: query understanding
User questions are often short, ambiguous or conversational ("what about for contractors?"). Improvements:
- Contextualise follow-ups: rewrite using the conversation into a standalone query ("What is the travel reimbursement deadline for contractors?").
- Query expansion: generate alternative phrasings or sub-questions and retrieve for each.
- Hypothetical answer embedding: have a model draft a plausible answer and embed that, since answers often resemble documents more than questions do. Useful, but test it; it can pull retrieval toward the model's assumptions.
- Routing: decide which index or tool to use (policies vs product docs vs a database).
Step 2: retrieval
Use hybrid search with metadata filters (permissions, region, product, date). Retrieve generously at this stage (for example 20 to 50 candidates), because the next step will be more selective.
Step 3: reranking
A reranker (often a cross-encoder model, or an LLM prompted to score relevance) looks at the query and each candidate together and scores relevance more accurately than vector similarity alone. It is slower per item, which is why it runs only on the candidate set. Reranking is frequently one of the highest-impact upgrades to a basic RAG system.
Step 4: prompt assembly
<sources>
<source id="S1" title="Travel Policy 2026 s4.2" date="2026-01-01">...</source>
<source id="S2" title="Contractor Handbook s2.1" date="2025-09-10">...</source>
</sources>
<instructions>
Answer using only the sources. Cite source ids in square brackets after
each claim, e.g. [S1]. If sources conflict, prefer the most recent official
policy and mention the conflict. If the sources do not contain the answer,
say so and suggest who to contact.
</instructions>
<question>{{standalone_question}}</question>Order passages by relevance, include titles and dates, and keep total context lean: more passages are not always better, since irrelevant text can distract the model.
Step 5: generation with citations
Ask for claim-level citations. Several model APIs now offer native citation features that return exact source spans (the hands-on section below uses one); these remove fragile bracket-parsing and make verification straightforward. Either way, citations are for the user's trust and for your verification.
Step 6: post-checks
- Verify every cited ID exists and, ideally, that quoted text appears in that source.
- Detect abstentions and route them usefully (search suggestion, human handoff).
- Log question, retrieved IDs, reranker scores, answer and citations for evaluation.
Worked example: an internal policy assistant
A regional firm with offices in London, Dubai and Karachi builds an HR assistant. Early problem: answers cite the London policy to Karachi staff. Fix: office metadata filter from the user's profile, plus an instruction to state which office the policy applies to. Second problem: follow-up questions ("and for part-timers?") retrieve nothing useful. Fix: rewrite follow-ups into standalone queries. Third: correct passages retrieved at rank 15 never reach the model. Fix: a reranker over the top 40. Each fix is targeted at a measured failure, not guessed.
When RAG is the wrong tool
- The answer requires aggregating over structured data ("total sales by region last quarter"): use a database query (text-to-SQL or a tool), not passage retrieval.
- The whole corpus is small enough to fit comfortably in context and changes rarely: sometimes simply including it (with caching) is simpler and better.
- The task needs the model to learn a style or skill rather than facts: examples or fine-tuning may fit better.
Hands-on: a minimal cited RAG answerer with the Claude API
This example takes passages from your retriever (for instance the hybrid search from module 1, plus a reranker) and asks Claude to answer with native citations. Citations come back as structured data pointing at the exact text used, so you can verify them in code rather than parsing brackets.
pip install anthropic sentence-transformers
export ANTHROPIC_API_KEY=... # never hard-code keys
export ANTHROPIC_MODEL=claude-opus-5 # check the current model list in the docsimport os
import anthropic
from sentence_transformers import CrossEncoder
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
reranker = CrossEncoder("BAAI/bge-reranker-v2-m3") # open-source cross-encoder reranker
def rerank(question, candidates, top_n=5):
scores = reranker.predict([(question, c["text"]) for c in candidates])
ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
return [c for _, c in ranked[:top_n]]
def answer(question, candidates):
passages = rerank(question, candidates)
content = [
{"type": "document",
"source": {"type": "text", "media_type": "text/plain", "data": p["text"]},
"title": f'{p["title"]} (updated {p["date"]})',
"citations": {"enabled": True}}
for p in passages
]
content.append({"type": "text", "text":
"Answer the question using only the documents. If they conflict, prefer the most "
"recent official policy and mention the conflict. If they do not contain the answer, "
"say so and suggest contacting HR.\n\nQuestion: " + question})
try:
resp = client.messages.create(model=MODEL, max_tokens=2000,
messages=[{"role": "user", "content": content}])
except anthropic.RateLimitError:
return {"answer": None, "error": "rate_limited_retry_later"}
except anthropic.APIStatusError as e:
return {"answer": None, "error": f"api_error_{e.status_code}"}
parts, cites = [], []
for block in resp.content:
if block.type == "text":
parts.append(block.text)
for c in (block.citations or []):
cites.append({"doc": c.document_title, "quote": c.cited_text})
return {"answer": "".join(parts), "citations": cites, "passages": [p["id"] for p in passages]}Things to check when you run it:
- Every claim should carry a citation. Log answers where text blocks have no citations; they are candidates for unsupported claims.
- Abstention works. Ask something the passages do not cover and confirm the answer says so.
- Log the passage IDs you sent, so evaluation (next lesson) can separate retrieval failures from generation failures.
Other providers offer comparable patterns (file search tools, grounding with citations); the architecture stays the same: retrieve generously, rerank, send a lean set of passages with titles and dates, and verify citations.
Going further
Advanced variants include graph-based retrieval (building entity-relationship graphs over documents for multi-hop questions), agentic RAG (the model decides when and what to search, iteratively), and multi-vector representations. Each adds complexity; adopt them when your evaluation shows the simpler pipeline failing on a class of questions they are designed for.
Key takeaways
- RAG retrieves from your sources at question time to ground answers and enable citations.
- Pipeline: query rewriting, hybrid retrieval with filters, reranking, lean prompt assembly, cited generation, post-checks.
- Reranking a generous candidate set is often one of the highest-impact upgrades.
- Use databases for aggregation questions, and consider full-context when the corpus is small.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Sketch the six-step pipeline for a knowledge base you know. For each step, write the one failure you expect most and how you would detect it.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.