Keep one bit per dimension, search with XOR and popcount, then rescore the survivors — 32× less memory for a couple of points of recall.
Keeping only whether each measurement was above or below average. Sounds reckless — but with a thousand measurements, the pattern of ups and downs still identifies the thing.
32× less memory and a distance function that is two CPU instructions, for a couple of points of recall you get back by rescoring.
Binary quantization is the most aggressive compression that still works: each dimension becomes a single bit, sign only. A 1024-dim float32 vector drops from 4096 bytes to 128. Distance becomes Hamming distance — an XOR and a popcount, both single CPU instructions across whole words — so scanning is not just smaller but dramatically faster. On its own that costs real recall, so the pattern is always oversample-and-rescore: retrieve 4-10× more candidates than you need using binary distance, then rescore those with the full-precision (or int8) vectors and keep the true top-k. On modern high-dimensional embeddings this recovers 95-99% of full-precision recall, which is why Qdrant, Weaviate, Milvus, and pgvector all shipped it.
Binary quantization stores one sign bit per dimension, cutting memory 32× and turning distance into XOR plus popcount. Used alone it loses accuracy, so you oversample — retrieve maybe eight times more candidates than you need — and then rescore those with full-precision vectors. On high-dimensional modern embeddings that recovers nearly all the recall, which makes it the cheapest big win in vector search right now.
Vector Quantization Techniques | Qdrant Multi-Vector Search — Qdrant Vector Search, 3:33