All concepts

Cross-Encoder Reranking

Retrieve wide and cheap with a bi-encoder, then rerank narrow and precise with a cross-encoder.

RAG & Retrieval · Advanced · ~8 min

In plain English

First pass compares pre-computed summaries of each document to the query — fast but shallow. Second pass actually reads the query and document together, for the shortlist only.

Why it's worth your time

It's the highest quality-per-hour change in retrieval, and the shape (cheap-and-broad, then expensive-and-accurate) is a pattern you'll reuse everywhere.

If you remember three things

  • Bi-encoder: two independent vectors, comparable in advance
  • Cross-encoder: one pass over the pair, far more accurate, can't be pre-computed
  • Latency is linear in depth; quality saturates

Overview

Bi-encoders compare precomputed query and document vectors — fast, but coarse. A cross-encoder reads the query and a candidate together and outputs a precise relevance score, at the cost of a forward pass per pair. So you retrieve ~100 candidates cheaply, then rerank the shortlist expensively.

How it works

  1. Retrieve wide & cheap A bi-encoder compares precomputed query/doc vectors — fast, so you fetch ~100 candidates.
  2. Coarse candidate list These are recall-oriented: the right doc is probably in here, but the ordering is rough.
  3. Cross-encoder rescoring A cross-encoder reads each (query, document) pair jointly and outputs a precise relevance score.
  4. Reorder by relevance Candidates are sorted by the new scores — the truly relevant docs rise to the top.
  5. Keep the top-k Hand the top few to the LLM. Retrieve wide and cheap; rerank narrow and expensive.

In an interview

Retrieval uses a bi-encoder: query and docs are embedded separately and compared by similarity — cheap and recall-oriented. Reranking uses a cross-encoder that reads (query, doc) jointly for a precise relevance score, but it's O(N) forward passes, so you only run it on the top ~50–100 candidates. Retrieve wide and cheap, rerank narrow and precise.

Production defaults

Depth
rerank 25–50 candidates, keep 3–5. The knee is almost always in that range
Budget
roughly 50–150 ms for 50 passages on a GPU. Measure before promising
Prerequisite
the gold passage must be in the candidate set. A reranker cannot promote what it never saw

What breaks

  • Reranking made results worse — Domain mismatch in the reranker, or the right passage wasn't in the candidates. Check recall@50 first.
  • Latency doubled — Depth set to 100+ 'to be safe'. Plot quality against depth and pick the knee.

Watch it explained

Semantic Reranker for better Search Results in 10 minutes — Ambarish Ganguly Academy, 9:52

Related