All concepts
Chunking Strategies
Split documents into retrieval units that are small enough to focus but large enough to make sense.
RAG & Retrieval · Intermediate · ~8 min
In plain English
Deciding where to cut your documents. Cut in the wrong place and half a sentence gets retrieved without the half that gave it meaning.
Why it's worth your time
This is where most RAG systems actually fail, weeks before anyone notices — and it's fixed at index time, not query time.
If you remember three things
- The chunk is the unit of retrieval AND the unit of reasoning
- Cut on structure, not character counts
- Overlap saves facts that straddle a boundary
Overview
How you slice documents into retrieval units. Chunks must be small enough to be focused embeddings yet large enough to stay self-contained. Split on structure (headings, paragraphs, tokens, or semantic breaks), add overlap to preserve cross-boundary context, and attach metadata like source, section, and access tags.
How it works
- Start: Long Document Raw documents are often too long to embed or retrieve as one unit.
- Long Document -> Chunk Boundaries Split by headings, paragraphs, tokens, semantic similarity, or tables.
- Chunk Boundaries -> Overlap Overlap preserves context that crosses boundaries.
- Overlap -> Metadata Attach source IDs, section titles, page numbers, timestamps, and ACLs.
- Metadata -> Retrieval Eval Measure whether chunks retrieve the evidence needed by downstream answers.
In an interview
Chunking is splitting documents into the units you embed and retrieve. Too large and the embedding is diffuse and retrieval imprecise; too small and a chunk loses the context needed to answer. I split on natural boundaries — headings, paragraphs, or semantic shifts — add token overlap, and tag each chunk with metadata.
Production defaults
- Size
- 600–1000 tokens prose; one function for code; one row-group for tables
- Overlap
- 10–20%
- Enrich
- prepend the document title and section path to every chunk — the cheapest retrieval win there is
- Small-to-big
- index small chunks for matching, return the parent section for context
What breaks
- Right document, wrong part — Chunks too big — the embedding averaged several topics into mush. Cut size and index a per-chunk summary.
- Tables and code come back scrambled — A prose splitter ran over structured content. Route by content type.