Latest AI Techniques: RAG, Tool Use, Agents & MCPRetrieval-augmented generation (RAG) · Lesson 5 of 20
Advanced RAG patterns: agentic, graph, visual and long-context
Video lecture
Advanced RAG patterns: agentic, graph, visual and long-context
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Advanced RAG patterns
Your RAG system is good now. Single-fact questions come back right, with citations. But a few kinds of question keep failing, however much you tune. Questions that need two documents. Questions about themes across everything. Questions about a chart on slide fourteen. In this lesson you'll meet the advanced patterns built for exactly those failures, and a decision table that stops you adopting shiny techniques you don't need.
0:30 Key idea: diagnose, then treat
Here's the key idea for this lesson, with an analogy. A good doctor doesn't prescribe the strongest medicine for every symptom. They diagnose first, then choose the lightest treatment that works. Advanced RAG patterns are strong medicine. They cost more, add side effects like latency and complexity, and only help specific conditions. Your evaluation tags are the diagnosis. Without them, you're prescribing blind.
0:57 The golden rule
Start with the golden rule. Adopt an advanced pattern only when your evaluation shows a class of questions failing that the pattern is designed to fix. Each one adds cost, latency and moving parts. If you skipped the evaluation lesson, go back, because without tagged failures you'll be guessing, and advanced RAG guessed at is expensive RAG that isn't any better.
1:24 Simple example: plan comparison
A simple example first. Someone asks: how do the Pro and Business plans differ on storage and support? A single search might find the Pro plan page but not the Business one. Query decomposition splits it into two searches: Pro plan storage and support, and Business plan storage and support. Each finds its page, the results are fused, and the model compares them side by side. One extra model call, and a whole class of comparison questions starts working.
1:58 Multi-hop options
Pattern one: agentic RAG. Retrieval becomes a tool the model can call repeatedly, with different queries, filters or indexes, until it has enough. Which UAE clients renewed after the price change, and did any complain? First search contracts, then search tickets for those accounts, then answer with citations to both. It's powerful for multi-step questions, but it's slower and more variable, so give it step limits and well-described search tools. Pattern two is its cheaper cousin: query decomposition. One call splits the question into sub-queries, your code searches each in parallel, fuses the rankings and reranks. Predictable, cacheable, and often enough.
2:42 Indexing-time patterns
Pattern three, contextual retrieval, you met in the chunking lesson: add situating context to each chunk at indexing time. No extra latency when users ask questions, which makes it great value. Pattern four is graph-based retrieval, often called GraphRAG. It extracts entities and relationships into a knowledge graph, sometimes with summaries of related clusters. It's designed for global questions, like what themes appear across five hundred customer interviews, and relationship questions, like which suppliers share a parent company. But indexing is expensive and extraction errors are hard to spot, so try clustering and summarising embeddings first.
3:24 Visual and long-context
Pattern five handles visual documents. Slide decks, brochures and scanned invoices lose tables, charts and layout when flattened to text. You can have a vision model describe each page at indexing time and index that description, or use visual document retrieval models that embed the page image itself, then send the retrieved pages to a vision-capable model to answer. Pattern six is the opposite of clever retrieval: for a small, stable corpus like a handbook, put the whole thing in the prompt and use prompt caching, so repeated questions reuse the processed text cheaply. Very long context windows make this practical, but it doesn't suit large, fast-changing or permission-sensitive content.
4:12 Worked example: agency knowledge assistant
Here's how a marketing agency in Riyadh and London used this. Their evaluation showed single-hop questions were fine, but questions about results hidden in charts failed, comparison questions failed, and what objections come up most in lost pitches failed. They added vision page descriptions for decks, then query decomposition, then a monthly cluster-and-summarise job for lost-pitch themes. They trialled a full graph pipeline, found clustering answered their themes questions nearly as well for far less indexing cost, and stopped there.
4:47 Business example (illustrative)
Let's add numbers to the agency example, illustrative. Their comparison and multi-hop questions made up about a fifth of real queries, and recall at five on that tag was around forty-five percent. Query decomposition lifted it to about seventy percent, at the cost of one extra model call per complex question. Because only a fifth of questions triggered it, total cost rose by only a few percent. That's the kind of trade-off you can defend to a finance lead.
5:21 Hands-on in the lesson
The lesson includes a decision table mapping failing question classes to the pattern to try first, and a hands-on query decomposition script using structured outputs to produce sub-queries, keeping the original question, fusing rankings with reciprocal rank fusion, and handing the result to your reranker. Measure it on your multi-hop tag only, and keep it only if recall rises there without hurting simple questions.
5:49 Common mistakes
Common mistakes with advanced patterns. Adopting a knowledge graph because it sounds sophisticated, then discovering the indexing bill. Letting agentic RAG loop without step limits. Testing a new pattern only on the questions it was designed for, and missing that it made simple questions worse. Using long context for content that changes daily or needs per-user permissions. And forgetting that every pattern still depends on clean, current documents underneath.
6:19 Did the pattern earn its place?
How will you know a pattern earned its place? It lifts the metric on the tag it targets, by a margin you decided in advance. It doesn't hurt your simple questions. Its extra cost and latency are acceptable for that route. And six weeks later, it's still helping, because the question mix hasn't shifted away from what it was built for. If any of these fail, remove it. Simpler is a feature.
6:50 Watch me do it: decomposition
Watch me do it. First, I define a Plan model with one field, a list of sub-queries. Next, decompose calls messages dot parse with that model and a prompt that says break the question into one to four standalone queries and keep product names and codes exactly. I cap the result at four. Then rrf fuses rankings the same way as before. Multi query retrieve builds the query list, always including the original question first, runs my search function on each, fuses the rankings and keeps the top thirty for the reranker. I test it with, how do the Pro and Business plans differ on storage and support? The printed sub-queries are Pro plan storage limits, Pro plan support, Business plan storage limits and Business plan support. Both plan pages now appear in the fused top five. Finally, I re-run only the multi-hop tag of my evaluation set and compare against the baseline.
7:57 Recap
To recap: advanced RAG patterns are targeted fixes. Agentic RAG and decomposition for multi-hop. Contextual retrieval for messy chunks. Graphs, or cheaper clustering, for themes and relationships. Visual retrieval for decks and scans. Full context with caching for small, stable corpora. Your next step is to pick your weakest evaluation tag, choose one pattern from the table, predict the improvement, and design the experiment to prove it.
8:26 Try this now (under an hour)
Try this now. Open your evaluation results and find your weakest tag. Write one sentence describing the failing question type, pick one pattern from the decision table, and write your prediction, for example: decomposition will raise multi-hop recall at five from forty to sixty percent without hurting single-hop. Then run the experiment on that tag and on your simple questions. Whether you're right or wrong, you'll learn something real in under an hour.
Beyond the basic pipeline
The six-step pipeline from the previous lessons handles most business questions. But some question types keep failing no matter how you tune chunk sizes or prompts: questions that need several hops across documents, questions about "themes across everything", questions over slide decks and scanned forms, and questions where the right search depends on what the first search found. This lesson surveys the advanced patterns built for those failures, and, just as importantly, how to decide whether you need them.
The rule of thumb: adopt an advanced pattern only when your evaluation shows a class of questions failing that the pattern is designed for. Each one adds cost, latency and moving parts.
Pattern 1: Agentic RAG
In classic RAG, your code retrieves once, then the model answers. In agentic RAG, retrieval becomes a tool the model can call as many times as it needs, with different queries, filters or indexes, deciding when it has enough.
User: "Which of our UAE clients renewed after the 2026 price change, and did any complain?"
Model -> search_contracts(query="renewal", region="UAE", after="2026-03-01")
Model -> search_tickets(query="price complaint", accounts=[...ids from first result...])
Model -> answer with citations to bothIt shines for multi-step questions and when the right query depends on earlier results. Costs: more model calls, more latency and more variability. Guard it with a step limit, a budget and the tool design principles from module 3 (each search tool should say what it covers and what it does not).
Pattern 2: Query decomposition and multi-query
A cheaper cousin of agentic RAG: one model call breaks a complex question into sub-questions or alternative phrasings, your code retrieves for each (in parallel), merges with rank fusion and reranks. It is predictable, easy to cache and often enough for comparison questions ("How do the Pro and Business plans differ on storage and support?").
Pattern 3: Contextual retrieval and enriched indexing
Covered in the chunking lesson: prepend model-written or structural context to each chunk before embedding and keyword indexing. It is an indexing-time change, so it adds no latency at question time, which makes it one of the best value upgrades for messy corpora.
Pattern 4: Graph-based retrieval (GraphRAG)
Graph approaches extract entities (people, products, companies, clauses) and relationships from your documents into a knowledge graph, sometimes with community summaries for clusters of related entities. At question time, the system retrieves via the graph as well as via text similarity.
- Good for: "global" questions across a corpus ("What are the main themes in 500 customer interviews?") and multi-hop relationship questions ("Which suppliers share a parent company with our top three vendors?").
- Costs: expensive indexing (many model calls), graph maintenance as documents change, and extraction errors that are hard to see.
- Alternative first: for many "themes" questions, clustering embeddings and summarising clusters is cheaper and nearly as useful.
Pattern 5: Visual and multimodal retrieval
Slide decks, brochures, invoices and scanned reports lose much of their meaning when flattened to OCR text: tables break, charts vanish, layout disappears. Two options:
- Describe, then index: use a vision-capable model to write a text description of each page or figure, and index that text (plus the OCR).
- Embed the page image: visual document retrieval models (the ColPali family and similar multi-vector approaches) embed page images directly and match text queries against them. Then send the retrieved page images to a vision-capable model to answer.
Test both on your documents; visual retrieval often wins on dense, visual pages, while text pipelines remain cheaper for plain documents.
Pattern 6: Long context plus caching instead of retrieval
Frontier models now accept very long inputs (several vendors offer context windows around a million tokens at the time of writing; check current limits). For a small, stable corpus, such as a 60-page policy handbook, you can put the whole thing in the prompt and use prompt caching so repeated questions reuse the processed prefix at a much lower cost. This removes chunking and retrieval failures entirely for that corpus. It stops working well when the corpus is large, changes often, needs per-user permissions, or when "lost in the middle" effects appear; many teams combine both: long context for the core handbook, retrieval for everything else.
Choosing: a decision table
| Failing question class (from your eval tags) | Try first | Then consider |
|---|---|---|
| Paraphrase and synonym misses | Hybrid search, better embeddings | Contextual retrieval |
| Multi-hop across documents | Query decomposition | Agentic RAG, graph retrieval |
| "Themes across everything" | Cluster-and-summarise | GraphRAG with community summaries |
| Tables, charts, scanned pages | Vision descriptions at indexing | Visual document retrieval |
| Small, stable core corpus | Full context + prompt caching | Hybrid of both |
| Answer depends on live data | Tools / database queries | Agentic RAG with tools |
Worked example: a B2B agency knowledge assistant
A marketing agency in Riyadh and London indexes 3 years of proposals, case studies and campaign reports (many as slide decks). Evaluation tags show: single-hop questions fine; "which case studies show ROAS improvements for fashion clients in KSA" failing (multi-hop plus data inside charts); and "what objections come up most in lost pitches" failing (a themes question).
Their fixes, in order: (1) vision-model page descriptions at indexing time for decks, which lifted recall on chart-heavy questions; (2) query decomposition for comparison and multi-hop questions; (3) a monthly cluster-and-summarise job over lost-pitch notes, stored as its own "themes" index. They trialled a full graph pipeline but found the cluster approach answered their themes questions nearly as well at a fraction of the indexing cost, so they stopped there. Every step was justified by a tag in their evaluation table.
Hands-on: query decomposition with rank fusion
import json, os
import anthropic
from pydantic import BaseModel
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
class Plan(BaseModel):
sub_queries: list[str]
def decompose(question: str) -> list[str]:
resp = client.messages.parse(
model=MODEL, max_tokens=1000,
messages=[{"role": "user", "content":
"Break this question into 1-4 standalone search queries that together cover it. "
"Keep product names, codes and dates exactly as written.\n\nQuestion: " + question}],
output_format=Plan,
)
return resp.parsed_output.sub_queries[:4]
def rrf(rankings, k=60):
fused = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking):
fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
return sorted(fused, key=fused.get, reverse=True)
def multi_query_retrieve(question, search, top_n=30):
queries = [question] + decompose(question) # always keep the original
rankings = [search(q) for q in queries] # your hybrid search -> ranked chunk IDs
return rrf(rankings)[:top_n] # then rerank and answer as beforeMeasure it against your baseline on the multi-hop tag only. If recall rises there without hurting single-hop questions, keep it; if not, remove it.
Key takeaways
- Adopt advanced RAG patterns only when evaluation shows a question class failing that they target.
- Agentic RAG and query decomposition handle multi-hop questions; decomposition is cheaper and more predictable.
- Graph retrieval suits global and relationship questions but has high indexing cost; try cluster-and-summarise first.
- Visual retrieval and long context with prompt caching are strong options for visual documents and small stable corpora.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take your RAG evaluation tags from the previous lesson. For the weakest tag, pick one pattern from the decision table, state the expected improvement, and design the experiment that would prove it.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.