All concepts
Cross-Encoder Reranking
Retrieve wide and cheap with a bi-encoder, then rerank narrow and precise with a cross-encoder.
RAG & Retrieval · Advanced · ~8 min
In plain English
First pass compares pre-computed summaries of each document to the query — fast but shallow. Second pass actually reads the query and document together, for the shortlist only.
Why it's worth your time
It's the highest quality-per-hour change in retrieval, and the shape (cheap-and-broad, then expensive-and-accurate) is a pattern you'll reuse everywhere.
If you remember three things
- Bi-encoder: two independent vectors, comparable in advance
- Cross-encoder: one pass over the pair, far more accurate, can't be pre-computed
- Latency is linear in depth; quality saturates
Overview
Bi-encoders compare precomputed query and document vectors — fast, but coarse. A cross-encoder reads the query and a candidate together and outputs a precise relevance score, at the cost of a forward pass per pair. So you retrieve ~100 candidates cheaply, then rerank the shortlist expensively.
How it works
- Retrieve wide & cheap A bi-encoder compares precomputed query/doc vectors — fast, so you fetch ~100 candidates.
- Coarse candidate list These are recall-oriented: the right doc is probably in here, but the ordering is rough.
- Cross-encoder rescoring A cross-encoder reads each (query, document) pair jointly and outputs a precise relevance score.
- Reorder by relevance Candidates are sorted by the new scores — the truly relevant docs rise to the top.
- Keep the top-k Hand the top few to the LLM. Retrieve wide and cheap; rerank narrow and expensive.
In an interview
Retrieval uses a bi-encoder: query and docs are embedded separately and compared by similarity — cheap and recall-oriented. Reranking uses a cross-encoder that reads (query, doc) jointly for a precise relevance score, but it's O(N) forward passes, so you only run it on the top ~50–100 candidates. Retrieve wide and cheap, rerank narrow and precise.
Production defaults
- Depth
- rerank 25–50 candidates, keep 3–5. The knee is almost always in that range
- Budget
- roughly 50–150 ms for 50 passages on a GPU. Measure before promising
- Prerequisite
- the gold passage must be in the candidate set. A reranker cannot promote what it never saw
What breaks
- Reranking made results worse — Domain mismatch in the reranker, or the right passage wasn't in the candidates. Check recall@50 first.
- Latency doubled — Depth set to 100+ 'to be safe'. Plot quality against depth and pick the knee.