Latest AI Techniques: RAG, Tool Use, Agents & MCPEmbeddings, semantic search and indexing · Lesson 2 of 20
Chunking and indexing strategies
Video lecture
Chunking and indexing strategies
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Chunking decides what can be found
Imagine you tear a policy manual into strips and hand one strip to a colleague. The strip says: it must be submitted within fourteen days. What must? Travel claims? Complaints? Your colleague has no idea, and neither does your AI system. That's chunking, and it quietly decides what your AI can and can't find. In this lesson you'll learn the trade-offs, the strategies that work, and a hands-on way to give every chunk the context it needs.
0:33 Why it matters
Why does this matter so much? Because chunking happens once, quietly, at indexing time, and then shapes every answer your system ever gives. Think of it like cutting a pizza. Cut it into tiny crumbs and nobody gets a proper bite. Leave it whole and nobody can share it. The right cut depends on the pizza and the people eating it. Documents are the same: the right chunk size depends on the document's structure and the questions people ask.
1:07 The size trade-off
Here's the trade-off. Small chunks, a sentence or a short paragraph, are precise. Each vector represents one idea. But they lose context, like our mysterious it. Large chunks, several pages, keep the context but blur the meaning. The vector becomes an average of many topics, so it matches no specific question strongly, and you stuff lots of irrelevant text into the model's prompt. There's no universal best size. A sensible starting point is a few hundred words with a little overlap, then you tune on your own test questions.
1:46 Pick a strategy per document type
Match the strategy to the document. For well-structured content like policies, manuals and web pages, split on the structure: headings, sections, list items, table rows. For long unstructured text, like call transcripts, semantic chunking splits where the topic shifts. And some documents need their own rules. Split code by function, spreadsheets by row groups with headers repeated, email threads by message, and FAQs by question and answer pair. A fixed-size splitter with overlap is a fine baseline, but structure-aware splitting usually wins on business documents.
2:23 Simple example: a returns page
Here's a simple example. A returns policy page has three headings: eligibility, how to return, and refunds. A fixed-size splitter might cut straight through the middle of the refunds section, so the refund amount ends up in one chunk and the refund timing in another. Split on the headings instead, and each chunk answers one question cleanly: can I return this, how do I send it back, and when do I get my money? Same page, far better retrieval.
2:57 Make chunks self-describing
Now the single most useful trick: make every chunk self-describing. Before you embed a chunk, prefix it with where it came from. Employee Handbook 2026, Expenses, Travel reimbursement. Suddenly our it-must-be-submitted strip is obviously about travel claims, and a question about travel expenses finds it. You embed the header plus the text, but you still show users the clean text with a citation. Store metadata too, such as office, effective date and access level, so you can filter.
3:31 Contextual retrieval
For messy documents with weak headings, you can go further with what Anthropic published as contextual retrieval. At indexing time, you show a model the whole document and one chunk, and ask it to write one or two sentences situating that chunk. You embed that context with the chunk, and index it for keyword search too. It costs extra processing, so use prompt caching to avoid paying for the whole document over and over, and measure recall before and after to prove it's worth it.
4:08 More patterns and hygiene
A few more patterns. Small-to-big retrieval matches on small precise chunks, then sends the model the larger parent section, so you get precision and context. Tables must keep their header row, or be converted to sentences. Scanned PDFs need good text extraction, and it's worth checking a sample by eye before indexing. And keep your index clean: remove superseded versions, date everything, re-index when content changes, and record permissions at index time. Looking ahead, late chunking and image-based page retrieval are maturing fast, and are worth benchmarking on slide decks and scanned forms.
4:48 Business example (illustrative)
Here's a deeper business example, with illustrative numbers. A UK insurance broker indexes policy wordings from twelve insurers. With fixed thousand-word chunks, their test set of sixty questions found the right clause in the top five only about half the time, because exclusions and definitions were mixed together. Splitting on clause headings, adding the insurer and policy name as a header, and keeping definition tables whole raised that to around four in five, without changing the embedding model at all.
5:23 Hands-on in the lesson
In the lesson you'll find a short Python function that splits Markdown on headings, keeps chunks under a word limit with overlap, and adds the section path as a header. You'll also find the prompt pattern for contextual retrieval. Try both on one real document: a policy, a price list, or a product manual. Then ask five real questions and check which version puts the right chunk in the top three.
5:54 Common mistakes
Common mistakes to avoid. Indexing boilerplate like cookie banners, menus and footers, which then match everything weakly. Splitting tables without their header row, so a chunk says fifty, eighty, one hundred and twenty with no idea what those numbers are. Keeping every old version of a document, so retrieval finds contradictions. Never checking OCR output from scanned PDFs. And tuning chunk size by gut feel instead of re-running your recall test after each change.
6:26 How you'll know it's working
How will you know your chunking works? Measure recall at k before and after each change, on the same questions. Watch the average length of passages you send to the model, which should stay lean. And read ten retrieved chunks at random: could a stranger understand each one on its own? If not, add context headers. Chunking is an experiment you run, not a setting you pick once.
6:56 Watch me do it: chunk_markdown()
Watch me do it. I paste the handbook snippet into the chunk markdown function. First, the regular expression splits the text on Markdown headings and keeps the heading text. Next, for each heading, I work out its level from the number of hash marks and trim the section path, so Travel reimbursement sits under Expenses, under Employee Handbook. Then the body is split into windows of two hundred and fifty words, with forty words of overlap. For each window, I build the context header by joining the document title and path with arrows, and I store two versions: embed text, which is header plus body, and text, which is the clean body for display. I run it, and the output reads Employee Handbook twenty twenty-six, four Expenses, four point two Travel reimbursement, followed by the sentence about fourteen days. Finally, I feed embed text into the hybrid search from the last lesson and re-run my five questions.
8:04 Recap
To recap: chunking decides what can be found. Small chunks are precise, large chunks carry context, and you tune the size with evaluation. Split along structure where you can, make every chunk self-describing, consider contextual retrieval for messy documents, and keep your index hygienic. Your next step is to treat chunking as an experiment, not a guess. In the next module, we'll assemble all of this into a full retrieval-augmented generation pipeline.
8:35 Try this now (25 minutes)
Try this now. Take one real document you'd want an AI assistant to know, a policy, a price list or a product guide. Chunk it two ways: a fixed size of about two hundred and fifty words, and by headings with the section path added as a header. Write five questions a customer might ask, and check which version puts the answering chunk in the top three. Keep the winning approach and the five questions: they're the start of your evaluation set.
Why chunking matters
Documents are usually too long to embed and retrieve as a whole. You split them into chunks, embed each chunk, and retrieve the best chunks for a query. Chunking quietly decides what your system can find: if the answer is split across two chunks, or buried in a chunk dominated by unrelated text, retrieval will struggle no matter how good your model is.
The chunk size trade-off
- Small chunks (a sentence or short paragraph) are precise: the vector represents one idea. But they may lack context ("It must be submitted within 14 days": what is "it"?).
- Large chunks (several pages) carry context but dilute meaning: the vector averages many topics, so similarity to a specific question drops, and you send more irrelevant text to the model.
There is no universal best size. Start with chunks of a few hundred words with some overlap, then tune on your evaluation set.
Chunking strategies
- Fixed-size with overlap. Split every N tokens with an overlap (for example 10 to 20 percent) so sentences at boundaries appear in both chunks. Simple and a reasonable baseline.
- Structure-aware. Split on headings, sections, paragraphs, list items, table rows or slide boundaries. Usually better for well-structured documents like policies, manuals and web pages.
- Semantic chunking. Split where the topic shifts, detected by embedding similarity between adjacent sentences. Useful for long, unstructured text such as transcripts.
- Document-specific. Code split by function; spreadsheets by row groups with headers repeated; emails by message; FAQs by question-answer pair.
Enriching chunks with context
A powerful, widely used technique is to add context to each chunk before embedding:
{
"chunk_id": "hr-policy-2026#s4.2",
"text": "Requests must be submitted within 14 days of the travel date.",
"context_header": "Employee Handbook 2026 > Section 4 Expenses > 4.2 Travel reimbursement",
"metadata": {"office": "Dubai", "effective": "2026-01-01", "access": "all-staff"}
}Embedding the header together with the text ("Travel reimbursement: requests must be submitted within 14 days...") makes the chunk findable by questions about travel expenses. Some teams go further and use a model to write a one or two sentence summary situating each chunk within its document, prepended before embedding. This costs extra processing at indexing time but can noticeably improve retrieval on ambiguous chunks.
Parent-child (small-to-big) retrieval
Retrieve with small, precise chunks, but send the model the larger parent section they belong to. You get precise matching and enough surrounding context for a correct answer.
Handling tables, images and PDFs
- Tables: keep header rows with every chunk, or convert rows to sentences ("Plan: Pro; Price: ...; Users: ..."). A table split mid-way without headers is meaningless.
- Scanned PDFs: need text extraction (OCR) or a vision-capable model; check extraction quality before indexing, since garbage in means garbage retrieved.
- Charts and diagrams: generate text descriptions with a vision model, or store the image and use a multimodal embedding model.
Index hygiene
- Deduplicate near-identical documents (old and new versions of a policy), or at least mark the current one.
- Version and date everything, and remove or down-rank superseded content.
- Re-index on change. Connect indexing to your content system so edits propagate.
- Respect permissions at index time. Record who can access each chunk; filter at query time.
Worked example: a product manual
A manufacturer indexes 40 product manuals. Version 1 uses fixed 1,000-token chunks; answers often mix instructions from different models. Version 2 splits by section, prefixes each chunk with product model and section path, stores the model number as metadata, and filters by the model the user selected. Retrieval accuracy on their test questions improves substantially, and the remaining failures are mostly questions that span several sections, which parent-child retrieval then addresses.
Failure modes
- Chunks without headers ("See table below").
- Boilerplate (cookie banners, navigation) indexed as content.
- Overlapping versions of the same document contradicting each other.
- A single giant chunk for a whole FAQ page, so every question matches it weakly.
Hands-on: structure-aware chunking with context headers
The script below splits a Markdown or exported web page on its headings, keeps each chunk under a size limit, and prefixes every chunk with its document title and section path before embedding. It is dependency-free, so you can drop it into any pipeline.
import re
def chunk_markdown(doc_title: str, text: str, max_words: int = 250, overlap_words: int = 40):
"""Split on headings, then on size; prepend a context header to every chunk."""
chunks, path = [], []
sections = re.split(r"(?m)^(#{1,6})\s+(.*)$", text)
# re.split returns: [preamble, hashes, heading, body, hashes, heading, body, ...]
blocks = [("", sections[0])] + [
(sections[i] + " " + sections[i + 1], sections[i + 2]) for i in range(1, len(sections) - 2, 3)
]
for heading, body in blocks:
if heading:
level = heading.count("#")
path = path[: level - 1] + [heading.lstrip("# ").strip()]
words = body.split()
start = 0
while start < len(words):
piece = " ".join(words[start : start + max_words])
header = " > ".join([doc_title] + path)
chunks.append({"id": f"{doc_title}#{len(chunks)}", "context_header": header,
"embed_text": f"{header}\n{piece}", "text": piece})
if start + max_words >= len(words):
break
start += max_words - overlap_words
return [c for c in chunks if c["text"].strip()]
policy = """# Employee Handbook 2026
## 4 Expenses
### 4.2 Travel reimbursement
Requests must be submitted within 14 days of the travel date. Attach receipts.
"""
for c in chunk_markdown("Employee Handbook 2026", policy):
print(c["context_header"], "|", c["text"][:60])Embed embed_text (header plus text), but show users text with a citation to context_header. Store metadata (office, effective date, access level) alongside.
Contextual retrieval: a model-written header
When headings are not enough (transcripts, long contracts), you can ask a model to write a one or two sentence "situating" note for each chunk at indexing time, a technique Anthropic published as contextual retrieval. Prompt pattern:
<document>{{WHOLE_DOCUMENT}}</document>
Here is a chunk from the document:
<chunk>{{CHUNK}}</chunk>
Write one or two sentences that situate this chunk within the overall document,
to improve search retrieval of the chunk. Answer only with the context.Because the whole document is repeated for every chunk, use your provider's prompt caching so the document prefix is paid for once per document rather than once per chunk. Then embed "context + chunk" and, ideally, index the same enriched text in BM25 too. As always, measure recall before and after on your own evaluation set; it is an extra indexing cost you should justify with numbers.
Late chunking and visual documents (outlook)
Two newer ideas are worth knowing. Late chunking embeds a long passage with a long-context embedding model first and splits the token embeddings afterwards, so each chunk's vector "remembers" its surroundings. Visual document retrieval embeds page images directly (the ColPali family of approaches), which can beat OCR pipelines on slide decks, forms and scanned reports full of tables and charts. Both are maturing quickly; treat them as options to benchmark, not defaults.
Going further
Treat chunking as a hyperparameter. Build a retrieval evaluation set (next module), then compare chunk sizes, overlap, structure-aware splitting and context enrichment. Measure recall at k: the share of questions where a chunk containing the answer appears in the top k results. Small, measured changes here often beat switching models.
Key takeaways
- Chunking decides what can be found: too small loses context, too large dilutes meaning.
- Prefer structure-aware chunking for structured documents and tune size with evaluation.
- Enrich chunks with headers, document context and metadata; consider small-to-big retrieval.
- Keep tables with headers, check extraction quality, deduplicate versions and re-index on change.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take one document you would index. Chunk it two ways (fixed-size and by structure with context headers) and test five questions: which approach retrieves the right chunk more often?
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.