Latest AI Techniques: RAG, Tool Use, Agents & MCPRetrieval-augmented generation (RAG) · Lesson 5 of 20

Advanced RAG patterns: agentic, graph, visual and long-context

Article · 14 min · 9 min lecture

Video lecture

Advanced RAG patterns: agentic, graph, visual and long-context

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Advanced RAG patterns

  • Multi-hop, themes, visuals, long context
  • Match the pattern to the failure
  • Only when evaluation says so

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Beyond the basic pipeline

The six-step pipeline from the previous lessons handles most business questions. But some question types keep failing no matter how you tune chunk sizes or prompts: questions that need several hops across documents, questions about "themes across everything", questions over slide decks and scanned forms, and questions where the right search depends on what the first search found. This lesson surveys the advanced patterns built for those failures, and, just as importantly, how to decide whether you need them.

The rule of thumb: adopt an advanced pattern only when your evaluation shows a class of questions failing that the pattern is designed for. Each one adds cost, latency and moving parts.

Pattern 1: Agentic RAG

In classic RAG, your code retrieves once, then the model answers. In agentic RAG, retrieval becomes a tool the model can call as many times as it needs, with different queries, filters or indexes, deciding when it has enough.

User: "Which of our UAE clients renewed after the 2026 price change, and did any complain?"
Model -> search_contracts(query="renewal", region="UAE", after="2026-03-01")
Model -> search_tickets(query="price complaint", accounts=[...ids from first result...])
Model -> answer with citations to both

It shines for multi-step questions and when the right query depends on earlier results. Costs: more model calls, more latency and more variability. Guard it with a step limit, a budget and the tool design principles from module 3 (each search tool should say what it covers and what it does not).

Pattern 2: Query decomposition and multi-query

A cheaper cousin of agentic RAG: one model call breaks a complex question into sub-questions or alternative phrasings, your code retrieves for each (in parallel), merges with rank fusion and reranks. It is predictable, easy to cache and often enough for comparison questions ("How do the Pro and Business plans differ on storage and support?").

Pattern 3: Contextual retrieval and enriched indexing

Covered in the chunking lesson: prepend model-written or structural context to each chunk before embedding and keyword indexing. It is an indexing-time change, so it adds no latency at question time, which makes it one of the best value upgrades for messy corpora.

Pattern 4: Graph-based retrieval (GraphRAG)

Graph approaches extract entities (people, products, companies, clauses) and relationships from your documents into a knowledge graph, sometimes with community summaries for clusters of related entities. At question time, the system retrieves via the graph as well as via text similarity.

  • Good for: "global" questions across a corpus ("What are the main themes in 500 customer interviews?") and multi-hop relationship questions ("Which suppliers share a parent company with our top three vendors?").
  • Costs: expensive indexing (many model calls), graph maintenance as documents change, and extraction errors that are hard to see.
  • Alternative first: for many "themes" questions, clustering embeddings and summarising clusters is cheaper and nearly as useful.

Pattern 5: Visual and multimodal retrieval

Slide decks, brochures, invoices and scanned reports lose much of their meaning when flattened to OCR text: tables break, charts vanish, layout disappears. Two options:

  1. Describe, then index: use a vision-capable model to write a text description of each page or figure, and index that text (plus the OCR).
  2. Embed the page image: visual document retrieval models (the ColPali family and similar multi-vector approaches) embed page images directly and match text queries against them. Then send the retrieved page images to a vision-capable model to answer.

Test both on your documents; visual retrieval often wins on dense, visual pages, while text pipelines remain cheaper for plain documents.

Pattern 6: Long context plus caching instead of retrieval

Frontier models now accept very long inputs (several vendors offer context windows around a million tokens at the time of writing; check current limits). For a small, stable corpus, such as a 60-page policy handbook, you can put the whole thing in the prompt and use prompt caching so repeated questions reuse the processed prefix at a much lower cost. This removes chunking and retrieval failures entirely for that corpus. It stops working well when the corpus is large, changes often, needs per-user permissions, or when "lost in the middle" effects appear; many teams combine both: long context for the core handbook, retrieval for everything else.

Choosing: a decision table

Failing question class (from your eval tags)Try firstThen consider
Paraphrase and synonym missesHybrid search, better embeddingsContextual retrieval
Multi-hop across documentsQuery decompositionAgentic RAG, graph retrieval
"Themes across everything"Cluster-and-summariseGraphRAG with community summaries
Tables, charts, scanned pagesVision descriptions at indexingVisual document retrieval
Small, stable core corpusFull context + prompt cachingHybrid of both
Answer depends on live dataTools / database queriesAgentic RAG with tools

Worked example: a B2B agency knowledge assistant

A marketing agency in Riyadh and London indexes 3 years of proposals, case studies and campaign reports (many as slide decks). Evaluation tags show: single-hop questions fine; "which case studies show ROAS improvements for fashion clients in KSA" failing (multi-hop plus data inside charts); and "what objections come up most in lost pitches" failing (a themes question).

Their fixes, in order: (1) vision-model page descriptions at indexing time for decks, which lifted recall on chart-heavy questions; (2) query decomposition for comparison and multi-hop questions; (3) a monthly cluster-and-summarise job over lost-pitch notes, stored as its own "themes" index. They trialled a full graph pipeline but found the cluster approach answered their themes questions nearly as well at a fraction of the indexing cost, so they stopped there. Every step was justified by a tag in their evaluation table.

Hands-on: query decomposition with rank fusion

import json, os
import anthropic
from pydantic import BaseModel

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

class Plan(BaseModel):
    sub_queries: list[str]

def decompose(question: str) -> list[str]:
    resp = client.messages.parse(
        model=MODEL, max_tokens=1000,
        messages=[{"role": "user", "content":
            "Break this question into 1-4 standalone search queries that together cover it. "
            "Keep product names, codes and dates exactly as written.\n\nQuestion: " + question}],
        output_format=Plan,
    )
    return resp.parsed_output.sub_queries[:4]

def rrf(rankings, k=60):
    fused = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
    return sorted(fused, key=fused.get, reverse=True)

def multi_query_retrieve(question, search, top_n=30):
    queries = [question] + decompose(question)          # always keep the original
    rankings = [search(q) for q in queries]            # your hybrid search -> ranked chunk IDs
    return rrf(rankings)[:top_n]                       # then rerank and answer as before

Measure it against your baseline on the multi-hop tag only. If recall rises there without hurting single-hop questions, keep it; if not, remove it.

Key takeaways

  • Adopt advanced RAG patterns only when evaluation shows a question class failing that they target.
  • Agentic RAG and query decomposition handle multi-hop questions; decomposition is cheaper and more predictable.
  • Graph retrieval suits global and relationship questions but has high indexing cost; try cluster-and-summarise first.
  • Visual retrieval and long context with prompt caching are strong options for visual documents and small stable corpora.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Evaluation shows single-hop questions pass but questions combining two documents fail. What should you try first?
  2. A 60-page handbook changes twice a year and every employee may read it. What is a strong, simple design?
  3. Which question type is GraphRAG with community summaries mainly designed for?

Put it into practice

Take your RAG evaluation tags from the previous lesson. For the weakest tag, pick one pattern from the decision table, state the expected improvement, and design the experiment that would prove it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.