All concepts

Binary Quantization + Rescoring

Keep one bit per dimension, search with XOR and popcount, then rescore the survivors — 32× less memory for a couple of points of recall.

Advanced Vector Search · Advanced · ~6 min

In plain English

Keeping only whether each measurement was above or below average. Sounds reckless — but with a thousand measurements, the pattern of ups and downs still identifies the thing.

Why it's worth your time

32× less memory and a distance function that is two CPU instructions, for a couple of points of recall you get back by rescoring.

If you remember three things

  • One sign bit per dimension; distance is XOR + popcount
  • Oversample then rescore — never serve binary results directly
  • Needs high dimensions and cosine-trained, centred embeddings

Overview

Binary quantization is the most aggressive compression that still works: each dimension becomes a single bit, sign only. A 1024-dim float32 vector drops from 4096 bytes to 128. Distance becomes Hamming distance — an XOR and a popcount, both single CPU instructions across whole words — so scanning is not just smaller but dramatically faster. On its own that costs real recall, so the pattern is always oversample-and-rescore: retrieve 4-10× more candidates than you need using binary distance, then rescore those with the full-precision (or int8) vectors and keep the true top-k. On modern high-dimensional embeddings this recovers 95-99% of full-precision recall, which is why Qdrant, Weaviate, Milvus, and pgvector all shipped it.

In an interview

Binary quantization stores one sign bit per dimension, cutting memory 32× and turning distance into XOR plus popcount. Used alone it loses accuracy, so you oversample — retrieve maybe eight times more candidates than you need — and then rescore those with full-precision vectors. On high-dimensional modern embeddings that recovers nearly all the recall, which makes it the cheapest big win in vector search right now.

Production defaults

Oversampling
start at 4× the k you want, raise to 8-10× if recall@10 is short. It is the main dial
Dimensions
≥768 for it to hold up. Below ~256 dimensions the sign pattern isn't enough
Rescore store
int8 or float on SSD is fine — you only fetch a few hundred vectors per query

What breaks

  • Recall collapsed — Either no rescoring stage, or the embedding isn't zero-centred. Check the mean of each dimension before blaming the method.
  • No memory saved — The rescore vectors are still in the same hot RAM. Move them to a colder tier — the candidate set is tiny.

Watch it explained

Vector Quantization Techniques | Qdrant Multi-Vector Search — Qdrant Vector Search, 3:33

Related