Cluster chunks, summarise each cluster, cluster the summaries, and index every level — so one retriever answers both detail and 'what is this about' questions.
Chapter summaries, then a summary of the summaries, all filed in the same drawer as the pages. Ask a detail question and you get a page; ask what the book is about and you get the summary.
It is the only clean answer to 'what are the themes across these documents' — a question flat top-k retrieval structurally cannot answer.
Flat chunk retrieval can only answer questions whose evidence fits in a chunk. Ask 'what were the main themes across these fifty reports' and top-k returns fifty unrelated fragments, none of which contains the answer. RAPTOR builds a tree instead: embed the leaf chunks, cluster them softly with a Gaussian mixture over a UMAP projection, summarise each cluster with an LLM, then embed those summaries and repeat until one node remains. Every node at every level goes into the same index. A specific question naturally matches a leaf; an abstract one matches a summary node that already aggregates the evidence. The collapsed-tree variant — search all levels at once — outperforms tree traversal and is what most implementations use.
RAPTOR builds a summarisation tree over your corpus: cluster the chunks, summarise each cluster, then cluster and summarise the summaries, recursively. Every node at every level is indexed together, so a detailed question retrieves a leaf and a thematic question retrieves a summary node whose text already aggregates the underlying evidence. It's the standard fix for questions no single chunk can answer.
RAG From Scratch: Part 13 (RAPTOR) — LangChain, 7:40