Keep one vector per token instead of one per document, and score with MaxSim — cross-encoder quality at nearly bi-encoder speed.
Instead of summarising a book in one sentence and comparing summaries, you keep every sentence and, for each thing you're looking for, find the single best matching line.
It recovers the exact-term precision that single-vector dense retrieval throws away, without a per-candidate model pass.
A bi-encoder squashes a whole document into one vector, which is fast but throws away term-level detail. A cross-encoder reads query and document together, which is accurate but must run the model once per candidate. Late interaction sits between them: encode every token of the document into its own small vector offline, encode the query's tokens at query time, and score with MaxSim — for each query token, take its best match among the document's tokens, then sum. All the expensive encoding still happens offline; the only query-time work is a lot of cheap dot products. ColBERTv2 and PLAID cut the storage cost with residual compression, and the pattern has since been ported to images (ColPali) and to multi-vector support in Vespa, Qdrant, and Weaviate.
Late interaction stores one vector per token rather than one per document, and scores a pair with MaxSim: each query token takes its best-matching document token, and you sum those maxima. The document encoding is precomputed, so query time is just dot products — you get much of a cross-encoder's term-level precision at a fraction of its cost. The trade is storage: a document is now dozens of vectors instead of one.
ColBERT Passage Retrieval: Late Interaction in PyTorch — Professor Py: Information Retrieval with Python, 7:22