All concepts
Hybrid Retrieval + Reranking
Combine keyword (BM25) and vector search, then rerank with a cross-encoder for precision.
RAG & Retrieval · Advanced · ~8 min
In plain English
Vector search understands meaning but misses exact strings like error codes. Keyword search is the opposite. Run both and merge the two lists.
Why it's worth your time
It's an afternoon of work and reliably the biggest single quality jump a RAG system gets.
If you remember three things
- Semantic search finds paraphrases; lexical search finds identifiers
- Reciprocal Rank Fusion merges without needing comparable scores
- Run the two searches in parallel, not in sequence
Overview
Pure vector search misses exact keywords, IDs, and rare terms; pure keyword search misses paraphrase. Hybrid retrieval fuses both (e.g. with reciprocal rank fusion), then a cross-encoder reranker rescoring the merged candidates gives production-grade precision.
How it works
- Keyword search (BM25) The query runs lexical BM25 search — great for exact terms, names, and IDs vectors miss.
- Vector search too In parallel, dense vector search catches paraphrase and meaning that keywords miss.
- Fuse both lists Reciprocal rank fusion merges the two result lists — you get the best recall from each.
- Both feed the fusion Keyword and vector hits combine into one high-recall candidate set.
- Rerank precisely A cross-encoder reads query + chunk together and rescoring them sharpens the final order.
- Top-k to the LLM The precise top-k context goes to generation — the reliable production recipe for RAG retrieval.
In an interview
Hybrid retrieval runs BM25 keyword search and dense vector search together and fuses their results, then a cross-encoder reranks the top candidates. Keyword search nails exact terms and IDs; vectors catch paraphrase; reranking sharpens the final order. It's the reliable production recipe for RAG retrieval.
Production defaults
- Candidates
- top-50 from each side
- Fusion
- RRF, score = Σ 1/(60 + rank). No normalization needed — that's why it's the default
- Weighting
- start 50/50, then tune. Identifier-heavy corpora want more lexical weight
What breaks
- Exact IDs still not found — BM25 is tokenizing your identifiers into fragments. Fix the analyzer so codes stay whole.
- Fusion made results worse — You normalized and added raw scores from two different scales. Use rank-based fusion instead.