All concepts

RAPTOR Hierarchical Index

Cluster chunks, summarise each cluster, cluster the summaries, and index every level — so one retriever answers both detail and 'what is this about' questions.

Advanced RAG · Advanced · ~6 min

In plain English

Chapter summaries, then a summary of the summaries, all filed in the same drawer as the pages. Ask a detail question and you get a page; ask what the book is about and you get the summary.

Why it's worth your time

It is the only clean answer to 'what are the themes across these documents' — a question flat top-k retrieval structurally cannot answer.

If you remember three things

  • Cluster → summarise → repeat, until one node remains
  • Index EVERY level in one collection (collapsed tree)
  • Soft clustering so a multi-topic chunk reaches both summaries

Overview

Flat chunk retrieval can only answer questions whose evidence fits in a chunk. Ask 'what were the main themes across these fifty reports' and top-k returns fifty unrelated fragments, none of which contains the answer. RAPTOR builds a tree instead: embed the leaf chunks, cluster them softly with a Gaussian mixture over a UMAP projection, summarise each cluster with an LLM, then embed those summaries and repeat until one node remains. Every node at every level goes into the same index. A specific question naturally matches a leaf; an abstract one matches a summary node that already aggregates the evidence. The collapsed-tree variant — search all levels at once — outperforms tree traversal and is what most implementations use.

In an interview

RAPTOR builds a summarisation tree over your corpus: cluster the chunks, summarise each cluster, then cluster and summarise the summaries, recursively. Every node at every level is indexed together, so a detailed question retrieves a leaf and a thematic question retrieves a summary node whose text already aggregates the underlying evidence. It's the standard fix for questions no single chunk can answer.

Production defaults

Clustering
UMAP to ~10 dims, Gaussian mixture, soft-assign at p > 0.10, target ~8 chunks per cluster
Retrieval
collapsed tree — all levels, one index, ordinary top-k. It beat tree traversal in the paper and is far simpler to run
Rebuild
per collection on a schedule. Incremental summary invalidation is the genuinely hard part — scope it small

What breaks

  • Built the tree, saw no gain — You're only retrieving leaves. Index the summary nodes in the same collection or the construction is wasted.
  • Summaries drift from the source — Too many levels over too little content. Cap the height and tag nodes with their level so you can inspect what answered.

Watch it explained

RAG From Scratch: Part 13 (RAPTOR) — LangChain, 7:40

Related