All concepts

Late Interaction (ColBERT)

Keep one vector per token instead of one per document, and score with MaxSim — cross-encoder quality at nearly bi-encoder speed.

Advanced Embeddings · Advanced · ~7 min

In plain English

Instead of summarising a book in one sentence and comparing summaries, you keep every sentence and, for each thing you're looking for, find the single best matching line.

Why it's worth your time

It recovers the exact-term precision that single-vector dense retrieval throws away, without a per-candidate model pass.

If you remember three things

  • One vector per token, not per document
  • MaxSim: best doc token per query token, then sum
  • Storage is the price — compression is mandatory at scale

Overview

A bi-encoder squashes a whole document into one vector, which is fast but throws away term-level detail. A cross-encoder reads query and document together, which is accurate but must run the model once per candidate. Late interaction sits between them: encode every token of the document into its own small vector offline, encode the query's tokens at query time, and score with MaxSim — for each query token, take its best match among the document's tokens, then sum. All the expensive encoding still happens offline; the only query-time work is a lot of cheap dot products. ColBERTv2 and PLAID cut the storage cost with residual compression, and the pattern has since been ported to images (ColPali) and to multi-vector support in Vespa, Qdrant, and Weaviate.

In an interview

Late interaction stores one vector per token rather than one per document, and scores a pair with MaxSim: each query token takes its best-matching document token, and you sum those maxima. The document encoding is precomputed, so query time is just dot products — you get much of a cross-encoder's term-level precision at a fraction of its cost. The trade is storage: a document is now dozens of vectors instead of one.

Production defaults

Storage budget
tokens × dim × bytes. ColBERTv2 residual compression cuts it ~6-10×; assume you need it above a few million docs
Where it fits
first-stage reranker over BM25 + dense top-1000, or the retriever itself in a multi-vector store
Engine
Vespa, Qdrant, or Weaviate multi-vector — emulating MaxSim in app code is an order of magnitude slower

What breaks

  • Index size exploded — You stored uncompressed token vectors. Turn on residual/centroid compression before scaling the corpus, not after.
  • No faster than a cross-encoder — You're re-encoding documents at query time. Document encoding must be offline — that's the entire premise.

Watch it explained

ColBERT Passage Retrieval: Late Interaction in PyTorch — Professor Py: Information Retrieval with Python, 7:22

Related