Latest AI Techniques: RAG, Tool Use, Agents & MCPEmbeddings, semantic search and indexing · Lesson 2 of 20

Chunking and indexing strategies

Article · 12 min · 9 min lecture

Video lecture

Chunking and indexing strategies

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Chunking decides what can be found

  • The chunk size trade-off
  • Strategies by document type
  • Context headers and contextual retrieval

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why chunking matters

Documents are usually too long to embed and retrieve as a whole. You split them into chunks, embed each chunk, and retrieve the best chunks for a query. Chunking quietly decides what your system can find: if the answer is split across two chunks, or buried in a chunk dominated by unrelated text, retrieval will struggle no matter how good your model is.

The chunk size trade-off

  • Small chunks (a sentence or short paragraph) are precise: the vector represents one idea. But they may lack context ("It must be submitted within 14 days": what is "it"?).
  • Large chunks (several pages) carry context but dilute meaning: the vector averages many topics, so similarity to a specific question drops, and you send more irrelevant text to the model.

There is no universal best size. Start with chunks of a few hundred words with some overlap, then tune on your evaluation set.

Chunking strategies

  1. Fixed-size with overlap. Split every N tokens with an overlap (for example 10 to 20 percent) so sentences at boundaries appear in both chunks. Simple and a reasonable baseline.
  2. Structure-aware. Split on headings, sections, paragraphs, list items, table rows or slide boundaries. Usually better for well-structured documents like policies, manuals and web pages.
  3. Semantic chunking. Split where the topic shifts, detected by embedding similarity between adjacent sentences. Useful for long, unstructured text such as transcripts.
  4. Document-specific. Code split by function; spreadsheets by row groups with headers repeated; emails by message; FAQs by question-answer pair.

Enriching chunks with context

A powerful, widely used technique is to add context to each chunk before embedding:

{
  "chunk_id": "hr-policy-2026#s4.2",
  "text": "Requests must be submitted within 14 days of the travel date.",
  "context_header": "Employee Handbook 2026 > Section 4 Expenses > 4.2 Travel reimbursement",
  "metadata": {"office": "Dubai", "effective": "2026-01-01", "access": "all-staff"}
}

Embedding the header together with the text ("Travel reimbursement: requests must be submitted within 14 days...") makes the chunk findable by questions about travel expenses. Some teams go further and use a model to write a one or two sentence summary situating each chunk within its document, prepended before embedding. This costs extra processing at indexing time but can noticeably improve retrieval on ambiguous chunks.

Parent-child (small-to-big) retrieval

Retrieve with small, precise chunks, but send the model the larger parent section they belong to. You get precise matching and enough surrounding context for a correct answer.

Handling tables, images and PDFs

  • Tables: keep header rows with every chunk, or convert rows to sentences ("Plan: Pro; Price: ...; Users: ..."). A table split mid-way without headers is meaningless.
  • Scanned PDFs: need text extraction (OCR) or a vision-capable model; check extraction quality before indexing, since garbage in means garbage retrieved.
  • Charts and diagrams: generate text descriptions with a vision model, or store the image and use a multimodal embedding model.

Index hygiene

  • Deduplicate near-identical documents (old and new versions of a policy), or at least mark the current one.
  • Version and date everything, and remove or down-rank superseded content.
  • Re-index on change. Connect indexing to your content system so edits propagate.
  • Respect permissions at index time. Record who can access each chunk; filter at query time.

Worked example: a product manual

A manufacturer indexes 40 product manuals. Version 1 uses fixed 1,000-token chunks; answers often mix instructions from different models. Version 2 splits by section, prefixes each chunk with product model and section path, stores the model number as metadata, and filters by the model the user selected. Retrieval accuracy on their test questions improves substantially, and the remaining failures are mostly questions that span several sections, which parent-child retrieval then addresses.

Failure modes

  • Chunks without headers ("See table below").
  • Boilerplate (cookie banners, navigation) indexed as content.
  • Overlapping versions of the same document contradicting each other.
  • A single giant chunk for a whole FAQ page, so every question matches it weakly.

Hands-on: structure-aware chunking with context headers

The script below splits a Markdown or exported web page on its headings, keeps each chunk under a size limit, and prefixes every chunk with its document title and section path before embedding. It is dependency-free, so you can drop it into any pipeline.

import re

def chunk_markdown(doc_title: str, text: str, max_words: int = 250, overlap_words: int = 40):
    """Split on headings, then on size; prepend a context header to every chunk."""
    chunks, path = [], []
    sections = re.split(r"(?m)^(#{1,6})\s+(.*)$", text)
    # re.split returns: [preamble, hashes, heading, body, hashes, heading, body, ...]
    blocks = [("", sections[0])] + [
        (sections[i] + " " + sections[i + 1], sections[i + 2]) for i in range(1, len(sections) - 2, 3)
    ]
    for heading, body in blocks:
        if heading:
            level = heading.count("#")
            path = path[: level - 1] + [heading.lstrip("# ").strip()]
        words = body.split()
        start = 0
        while start < len(words):
            piece = " ".join(words[start : start + max_words])
            header = " > ".join([doc_title] + path)
            chunks.append({"id": f"{doc_title}#{len(chunks)}", "context_header": header,
                           "embed_text": f"{header}\n{piece}", "text": piece})
            if start + max_words >= len(words):
                break
            start += max_words - overlap_words
    return [c for c in chunks if c["text"].strip()]

policy = """# Employee Handbook 2026
## 4 Expenses
### 4.2 Travel reimbursement
Requests must be submitted within 14 days of the travel date. Attach receipts.
"""
for c in chunk_markdown("Employee Handbook 2026", policy):
    print(c["context_header"], "|", c["text"][:60])

Embed embed_text (header plus text), but show users text with a citation to context_header. Store metadata (office, effective date, access level) alongside.

Contextual retrieval: a model-written header

When headings are not enough (transcripts, long contracts), you can ask a model to write a one or two sentence "situating" note for each chunk at indexing time, a technique Anthropic published as contextual retrieval. Prompt pattern:

<document>{{WHOLE_DOCUMENT}}</document>
Here is a chunk from the document:
<chunk>{{CHUNK}}</chunk>
Write one or two sentences that situate this chunk within the overall document,
to improve search retrieval of the chunk. Answer only with the context.

Because the whole document is repeated for every chunk, use your provider's prompt caching so the document prefix is paid for once per document rather than once per chunk. Then embed "context + chunk" and, ideally, index the same enriched text in BM25 too. As always, measure recall before and after on your own evaluation set; it is an extra indexing cost you should justify with numbers.

Late chunking and visual documents (outlook)

Two newer ideas are worth knowing. Late chunking embeds a long passage with a long-context embedding model first and splits the token embeddings afterwards, so each chunk's vector "remembers" its surroundings. Visual document retrieval embeds page images directly (the ColPali family of approaches), which can beat OCR pipelines on slide decks, forms and scanned reports full of tables and charts. Both are maturing quickly; treat them as options to benchmark, not defaults.

Going further

Treat chunking as a hyperparameter. Build a retrieval evaluation set (next module), then compare chunk sizes, overlap, structure-aware splitting and context enrichment. Measure recall at k: the share of questions where a chunk containing the answer appears in the top k results. Small, measured changes here often beat switching models.

Key takeaways

  • Chunking decides what can be found: too small loses context, too large dilutes meaning.
  • Prefer structure-aware chunking for structured documents and tune size with evaluation.
  • Enrich chunks with headers, document context and metadata; consider small-to-big retrieval.
  • Keep tables with headers, check extraction quality, deduplicate versions and re-index on change.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A chunk reads 'It must be submitted within 14 days' and is rarely retrieved for expense questions. What is the best fix?
  2. What is small-to-big (parent-child) retrieval?
  3. Which metric measures whether the right chunk appears in the top results?

Put it into practice

Take one document you would index. Chunk it two ways (fixed-size and by structure with context headers) and test five questions: which approach retrieves the right chunk more often?

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.