---
title: "Advanced RAG patterns: agentic, graph, visual and…"
description: "Beyond the basic pipeline The six-step pipeline from the previous lessons handles most business questions. But some question types keep failing no matter…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/advanced-rag-patterns
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Retrieval-augmented generation (RAG) · lesson 5 of 20 · 14 min

# Advanced RAG patterns: agentic, graph, visual and long-context

## Beyond the basic pipeline

The six-step pipeline from the previous lessons handles most business questions. But some question types keep failing no matter how you tune chunk sizes or prompts: questions that need several hops across documents, questions about "themes across everything", questions over slide decks and scanned forms, and questions where the right search depends on what the first search found. This lesson surveys the advanced patterns built for those failures, and, just as importantly, how to decide whether you need them.

The rule of thumb: **adopt an advanced pattern only when your evaluation shows a class of questions failing that the pattern is designed for.** Each one adds cost, latency and moving parts.

## Pattern 1: Agentic RAG

In classic RAG, your code retrieves once, then the model answers. In **agentic RAG**, retrieval becomes a tool the model can call as many times as it needs, with different queries, filters or indexes, deciding when it has enough.

```text
User: "Which of our UAE clients renewed after the 2026 price change, and did any complain?"
Model -> search_contracts(query="renewal", region="UAE", after="2026-03-01")
Model -> search_tickets(query="price complaint", accounts=[...ids from first result...])
Model -> answer with citations to both
```

It shines for multi-step questions and when the right query depends on earlier results. Costs: more model calls, more latency and more variability. Guard it with a step limit, a budget and the tool design principles from module 3 (each search tool should say what it covers and what it does not).

## Pattern 2: Query decomposition and multi-query

A cheaper cousin of agentic RAG: one model call breaks a complex question into sub-questions or alternative phrasings, your code retrieves for each (in parallel), merges with rank fusion and reranks. It is predictable, easy to cache and often enough for comparison questions ("How do the Pro and Business plans differ on storage and support?").

## Pattern 3: Contextual retrieval and enriched indexing

Covered in the chunking lesson: prepend model-written or structural context to each chunk before embedding and keyword indexing. It is an indexing-time change, so it adds no latency at question time, which makes it one of the best value upgrades for messy corpora.

## Pattern 4: Graph-based retrieval (GraphRAG)

Graph approaches extract entities (people, products, companies, clauses) and relationships from your documents into a knowledge graph, sometimes with community summaries for clusters of related entities. At question time, the system retrieves via the graph as well as via text similarity.

- **Good for:** "global" questions across a corpus ("What are the main themes in 500 customer interviews?") and multi-hop relationship questions ("Which suppliers share a parent company with our top three vendors?").
- **Costs:** expensive indexing (many model calls), graph maintenance as documents change, and extraction errors that are hard to see.
- **Alternative first:** for many "themes" questions, clustering embeddings and summarising clusters is cheaper and nearly as useful.

## Pattern 5: Visual and multimodal retrieval

Slide decks, brochures, invoices and scanned reports lose much of their meaning when flattened to OCR text: tables break, charts vanish, layout disappears. Two options:

1. **Describe, then index:** use a vision-capable model to write a text description of each page or figure, and index that text (plus the OCR).
2. **Embed the page image:** visual document retrieval models (the ColPali family and similar multi-vector approaches) embed page images directly and match text queries against them. Then send the retrieved page images to a vision-capable model to answer.

Test both on your documents; visual retrieval often wins on dense, visual pages, while text pipelines remain cheaper for plain documents.

## Pattern 6: Long context plus caching instead of retrieval

Frontier models now accept very long inputs (several vendors offer context windows around a million tokens at the time of writing; check current limits). For a small, stable corpus, such as a 60-page policy handbook, you can put the whole thing in the prompt and use **prompt caching** so repeated questions reuse the processed prefix at a much lower cost. This removes chunking and retrieval failures entirely for that corpus. It stops working well when the corpus is large, changes often, needs per-user permissions, or when "lost in the middle" effects appear; many teams combine both: long context for the core handbook, retrieval for everything else.

## Choosing: a decision table

| Failing question class (from your eval tags) | Try first | Then consider |
|---|---|---|
| Paraphrase and synonym misses | Hybrid search, better embeddings | Contextual retrieval |
| Multi-hop across documents | Query decomposition | Agentic RAG, graph retrieval |
| "Themes across everything" | Cluster-and-summarise | GraphRAG with community summaries |
| Tables, charts, scanned pages | Vision descriptions at indexing | Visual document retrieval |
| Small, stable core corpus | Full context + prompt caching | Hybrid of both |
| Answer depends on live data | Tools / database queries | Agentic RAG with tools |

## Worked example: a B2B agency knowledge assistant

A marketing agency in Riyadh and London indexes 3 years of proposals, case studies and campaign reports (many as slide decks). Evaluation tags show: single-hop questions fine; "which case studies show ROAS improvements for fashion clients in KSA" failing (multi-hop plus data inside charts); and "what objections come up most in lost pitches" failing (a themes question).

Their fixes, in order: (1) vision-model page descriptions at indexing time for decks, which lifted recall on chart-heavy questions; (2) query decomposition for comparison and multi-hop questions; (3) a monthly cluster-and-summarise job over lost-pitch notes, stored as its own "themes" index. They trialled a full graph pipeline but found the cluster approach answered their themes questions nearly as well at a fraction of the indexing cost, so they stopped there. Every step was justified by a tag in their evaluation table.

## Hands-on: query decomposition with rank fusion

```python
import json, os
import anthropic
from pydantic import BaseModel

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

class Plan(BaseModel):
    sub_queries: list[str]

def decompose(question: str) -> list[str]:
    resp = client.messages.parse(
        model=MODEL, max_tokens=1000,
        messages=[{"role": "user", "content":
            "Break this question into 1-4 standalone search queries that together cover it. "
            "Keep product names, codes and dates exactly as written.\n\nQuestion: " + question}],
        output_format=Plan,
    )
    return resp.parsed_output.sub_queries[:4]

def rrf(rankings, k=60):
    fused = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
    return sorted(fused, key=fused.get, reverse=True)

def multi_query_retrieve(question, search, top_n=30):
    queries = [question] + decompose(question)          # always keep the original
    rankings = [search(q) for q in queries]            # your hybrid search -> ranked chunk IDs
    return rrf(rankings)[:top_n]                       # then rerank and answer as before
```

Measure it against your baseline on the multi-hop tag only. If recall rises there without hurting single-hop questions, keep it; if not, remove it.

## Video lecture: Advanced RAG patterns: agentic, graph, visual and long-context

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Advanced RAG patterns
2. Key idea: diagnose, then treat
3. The golden rule
4. Simple example: plan comparison
5. Multi-hop options
6. Indexing-time patterns
7. Visual and long-context
8. Worked example: agency knowledge assistant
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. Did the pattern earn its place?
13. Watch me do it: decomposition
14. Recap
15. Try this now (under an hour)

## Lecture transcript

### Advanced RAG patterns

Your RAG system is good now. Single-fact questions come back right, with citations. But a few kinds of question keep failing, however much you tune. Questions that need two documents. Questions about themes across everything. Questions about a chart on slide fourteen. In this lesson you'll meet the advanced patterns built for exactly those failures, and a decision table that stops you adopting shiny techniques you don't need.

### Key idea: diagnose, then treat

Here's the key idea for this lesson, with an analogy. A good doctor doesn't prescribe the strongest medicine for every symptom. They diagnose first, then choose the lightest treatment that works. Advanced RAG patterns are strong medicine. They cost more, add side effects like latency and complexity, and only help specific conditions. Your evaluation tags are the diagnosis. Without them, you're prescribing blind.

### The golden rule

Start with the golden rule. Adopt an advanced pattern only when your evaluation shows a class of questions failing that the pattern is designed to fix. Each one adds cost, latency and moving parts. If you skipped the evaluation lesson, go back, because without tagged failures you'll be guessing, and advanced RAG guessed at is expensive RAG that isn't any better.

### Simple example: plan comparison

A simple example first. Someone asks: how do the Pro and Business plans differ on storage and support? A single search might find the Pro plan page but not the Business one. Query decomposition splits it into two searches: Pro plan storage and support, and Business plan storage and support. Each finds its page, the results are fused, and the model compares them side by side. One extra model call, and a whole class of comparison questions starts working.

### Multi-hop options

Pattern one: agentic RAG. Retrieval becomes a tool the model can call repeatedly, with different queries, filters or indexes, until it has enough. Which UAE clients renewed after the price change, and did any complain? First search contracts, then search tickets for those accounts, then answer with citations to both. It's powerful for multi-step questions, but it's slower and more variable, so give it step limits and well-described search tools. Pattern two is its cheaper cousin: query decomposition. One call splits the question into sub-queries, your code searches each in parallel, fuses the rankings and reranks. Predictable, cacheable, and often enough.

### Indexing-time patterns

Pattern three, contextual retrieval, you met in the chunking lesson: add situating context to each chunk at indexing time. No extra latency when users ask questions, which makes it great value. Pattern four is graph-based retrieval, often called GraphRAG. It extracts entities and relationships into a knowledge graph, sometimes with summaries of related clusters. It's designed for global questions, like what themes appear across five hundred customer interviews, and relationship questions, like which suppliers share a parent company. But indexing is expensive and extraction errors are hard to spot, so try clustering and summarising embeddings first.

### Visual and long-context

Pattern five handles visual documents. Slide decks, brochures and scanned invoices lose tables, charts and layout when flattened to text. You can have a vision model describe each page at indexing time and index that description, or use visual document retrieval models that embed the page image itself, then send the retrieved pages to a vision-capable model to answer. Pattern six is the opposite of clever retrieval: for a small, stable corpus like a handbook, put the whole thing in the prompt and use prompt caching, so repeated questions reuse the processed text cheaply. Very long context windows make this practical, but it doesn't suit large, fast-changing or permission-sensitive content.

### Worked example: agency knowledge assistant

Here's how a marketing agency in Riyadh and London used this. Their evaluation showed single-hop questions were fine, but questions about results hidden in charts failed, comparison questions failed, and what objections come up most in lost pitches failed. They added vision page descriptions for decks, then query decomposition, then a monthly cluster-and-summarise job for lost-pitch themes. They trialled a full graph pipeline, found clustering answered their themes questions nearly as well for far less indexing cost, and stopped there.

### Business example (illustrative)

Let's add numbers to the agency example, illustrative. Their comparison and multi-hop questions made up about a fifth of real queries, and recall at five on that tag was around forty-five percent. Query decomposition lifted it to about seventy percent, at the cost of one extra model call per complex question. Because only a fifth of questions triggered it, total cost rose by only a few percent. That's the kind of trade-off you can defend to a finance lead.

### Hands-on in the lesson

The lesson includes a decision table mapping failing question classes to the pattern to try first, and a hands-on query decomposition script using structured outputs to produce sub-queries, keeping the original question, fusing rankings with reciprocal rank fusion, and handing the result to your reranker. Measure it on your multi-hop tag only, and keep it only if recall rises there without hurting simple questions.

### Common mistakes

Common mistakes with advanced patterns. Adopting a knowledge graph because it sounds sophisticated, then discovering the indexing bill. Letting agentic RAG loop without step limits. Testing a new pattern only on the questions it was designed for, and missing that it made simple questions worse. Using long context for content that changes daily or needs per-user permissions. And forgetting that every pattern still depends on clean, current documents underneath.

### Did the pattern earn its place?

How will you know a pattern earned its place? It lifts the metric on the tag it targets, by a margin you decided in advance. It doesn't hurt your simple questions. Its extra cost and latency are acceptable for that route. And six weeks later, it's still helping, because the question mix hasn't shifted away from what it was built for. If any of these fail, remove it. Simpler is a feature.

### Watch me do it: decomposition

Watch me do it. First, I define a Plan model with one field, a list of sub-queries. Next, decompose calls messages dot parse with that model and a prompt that says break the question into one to four standalone queries and keep product names and codes exactly. I cap the result at four. Then rrf fuses rankings the same way as before. Multi query retrieve builds the query list, always including the original question first, runs my search function on each, fuses the rankings and keeps the top thirty for the reranker. I test it with, how do the Pro and Business plans differ on storage and support? The printed sub-queries are Pro plan storage limits, Pro plan support, Business plan storage limits and Business plan support. Both plan pages now appear in the fused top five. Finally, I re-run only the multi-hop tag of my evaluation set and compare against the baseline.

### Recap

To recap: advanced RAG patterns are targeted fixes. Agentic RAG and decomposition for multi-hop. Contextual retrieval for messy chunks. Graphs, or cheaper clustering, for themes and relationships. Visual retrieval for decks and scans. Full context with caching for small, stable corpora. Your next step is to pick your weakest evaluation tag, choose one pattern from the table, predict the improvement, and design the experiment to prove it.

### Try this now (under an hour)

Try this now. Open your evaluation results and find your weakest tag. Write one sentence describing the failing question type, pick one pattern from the decision table, and write your prediction, for example: decomposition will raise multi-hop recall at five from forty to sixty percent without hurting single-hop. Then run the experiment on that tag and on your simple questions. Whether you're right or wrong, you'll learn something real in under an hour.

## Key takeaways

- Adopt advanced RAG patterns only when evaluation shows a question class failing that they target.
- Agentic RAG and query decomposition handle multi-hop questions; decomposition is cheaper and more predictable.
- Graph retrieval suits global and relationship questions but has high indexing cost; try cluster-and-summarise first.
- Visual retrieval and long context with prompt caching are strong options for visual documents and small stable corpora.

## Try it

Take your RAG evaluation tags from the previous lesson. For the weakest tag, pick one pattern from the decision table, state the expected improvement, and design the experiment that would prove it.

- [Previous: Evaluating RAG systems](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/evaluating-rag)
- [Next: Function calling fundamentals](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/function-calling-fundamentals)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
