All concepts

Contextual Retrieval

Prepend an LLM-written sentence of surrounding context to every chunk before you embed it — the cheapest large retrieval win there is.

Advanced RAG · Advanced · ~6 min

In plain English

A page torn from a report says 'revenue grew 3%'. Whose revenue? Which quarter? Write that on the top of the page before you file it, and it becomes findable.

Why it's worth your time

Anthropic measured roughly a third fewer retrieval failures from this alone, and about two thirds with BM25 and a reranker on top. Little else in RAG returns that much for that little code.

If you remember three things

  • Chunking severs identity; this restores it at index time
  • Contextualise for BOTH the vector index and BM25
  • Prompt caching over the document is what makes the cost trivial

Overview

Chunking destroys context. A paragraph that says 'revenue grew 3% over the previous quarter' is unretrievable for 'ACME Q2 2024 growth', because the chunk never names the company, the quarter, or the year — the document did, forty paragraphs earlier. Contextual retrieval fixes it at index time: for each chunk, ask a cheap model to write one or two sentences situating that chunk in its document, prepend them, and embed the combined text. Do the same for the BM25 index. Anthropic's published numbers put the retrieval failure-rate reduction at roughly 35% for contextual embeddings, ~49% combined with contextual BM25, and ~67% with a reranker on top. Prompt caching over the document makes the indexing cost small enough to be irrelevant.

In an interview

Contextual retrieval rewrites each chunk at index time by prepending an LLM-generated sentence that situates it in its parent document — which entity, which period, which section. Because the embedding and the BM25 posting list now contain the identifying terms the chunk itself omitted, queries that name the entity actually match. It's the highest-return-per-line change available in RAG, and prompt caching keeps the indexing cost trivial.

Production defaults

Model
the cheapest capable model (Haiku-class), 50-100 tokens of context per chunk
Caching
cache the document prefix and generate all its chunks against it — per-chunk cost lands well under $0.001
Stack
contextual embeddings + contextual BM25, fused with RRF, then a reranker on the top 150

What breaks

  • No improvement — You generated the context from the chunk alone. The prompt must include the whole document — the missing information is by definition not in the chunk.
  • Index is inconsistent — The context prompt changed mid-run. Version the prompt with the index and treat a change as a full re-index.

Watch it explained

Advanced RAG techniques for developers — Google Cloud Tech, 8:17

Related