All concepts

HyDE & Query Expansion

Search with a hallucinated answer instead of the question — a fake document is closer to the real one than a query ever is.

Advanced RAG · Advanced · ~6 min

In plain English

Don't describe the book you're after — sketch a page of it. Handing the librarian a page that looks like what you want works better than describing it.

Why it's worth your time

It fixes the query–document asymmetry that dense retrieval papers over, and it costs one cheap LLM call.

If you remember three things

  • A hypothetical answer embeds nearer real answers than a question does
  • Wrong facts are harmless — you never show the text to anyone
  • RRF fuses ranked lists without any score calibration

Overview

Queries and documents are different kinds of text. A question is short, interrogative, and vague; a document is long, declarative, and specific. Embedding both into one space and comparing them is an asymmetry the model has to paper over. HyDE — Hypothetical Document Embeddings — sidesteps it: ask an LLM to write the answer it imagines, then embed that hypothetical document and search with it. The invented facts don't matter, because you never show it to the user; the embedding only needs to land in the right neighbourhood. The sibling technique is multi-query expansion: generate several rewordings, retrieve for each, and fuse the rankings with Reciprocal Rank Fusion. Both cost one extra LLM call and both are strongest exactly where dense retrieval is weakest — short, ambiguous, jargon-light queries.

In an interview

HyDE has an LLM write a hypothetical answer to the query, then searches with that document's embedding instead of the query's. Because documents match documents better than questions match documents, this lands in the right region even when the query is short or vague — and the hallucinated details are harmless since the text is never shown. Multi-query expansion is the sibling trick: several rewordings, retrieved separately and fused with Reciprocal Rank Fusion.

Production defaults

Generation
a small fast model, ~150 tokens, temperature 0.7. Quality of prose barely matters; topical density does
Fusion
RRF with k ≈ 60 across the original query, HyDE, and 3-4 rewordings
Caching
cache by normalised query — head queries repeat far more than teams expect

What breaks

  • Retrieval got worse — The hypothetical answer confidently drifted to another topic. Fuse it WITH the original query rather than replacing it.
  • p95 latency jumped — You added a serial LLM hop. Run HyDE and the plain query search in parallel, and cache aggressively.

Watch it explained

RAG from scratch: Part 9 (Query Translation -- HyDE) — LangChain, 4:47

Related