All concepts

Chunking Strategies

Split documents into retrieval units that are small enough to focus but large enough to make sense.

RAG & Retrieval · Intermediate · ~8 min

In plain English

Deciding where to cut your documents. Cut in the wrong place and half a sentence gets retrieved without the half that gave it meaning.

Why it's worth your time

This is where most RAG systems actually fail, weeks before anyone notices — and it's fixed at index time, not query time.

If you remember three things

  • The chunk is the unit of retrieval AND the unit of reasoning
  • Cut on structure, not character counts
  • Overlap saves facts that straddle a boundary

Overview

How you slice documents into retrieval units. Chunks must be small enough to be focused embeddings yet large enough to stay self-contained. Split on structure (headings, paragraphs, tokens, or semantic breaks), add overlap to preserve cross-boundary context, and attach metadata like source, section, and access tags.

How it works

  1. Start: Long Document Raw documents are often too long to embed or retrieve as one unit.
  2. Long Document -> Chunk Boundaries Split by headings, paragraphs, tokens, semantic similarity, or tables.
  3. Chunk Boundaries -> Overlap Overlap preserves context that crosses boundaries.
  4. Overlap -> Metadata Attach source IDs, section titles, page numbers, timestamps, and ACLs.
  5. Metadata -> Retrieval Eval Measure whether chunks retrieve the evidence needed by downstream answers.

In an interview

Chunking is splitting documents into the units you embed and retrieve. Too large and the embedding is diffuse and retrieval imprecise; too small and a chunk loses the context needed to answer. I split on natural boundaries — headings, paragraphs, or semantic shifts — add token overlap, and tag each chunk with metadata.

Production defaults

Size
600–1000 tokens prose; one function for code; one row-group for tables
Overlap
10–20%
Enrich
prepend the document title and section path to every chunk — the cheapest retrieval win there is
Small-to-big
index small chunks for matching, return the parent section for context

What breaks

  • Right document, wrong part — Chunks too big — the embedding averaged several topics into mush. Cut size and index a per-chunk summary.
  • Tables and code come back scrambled — A prose splitter ran over structured content. Route by content type.

Watch it explained

Chunking Methods for RAG Explained: Overlapped vs Semantic vs Late Chunking (2026) — TecAdRise, 7:35

Related