Put the graph on SSD, keep a compressed copy in RAM to steer, and serve a billion vectors from one machine.
A library too big for your desk. You keep an index card for every book in a drawer you can reach, and walk to the shelf only for the handful you actually need.
RAM costs about a hundred times what NVMe does per byte, and at a billion vectors that ratio decides whether the system is one machine or forty.
HNSW assumes the graph and the vectors live in RAM, which caps a single node at tens of millions of vectors. DiskANN removes that assumption. It builds a Vamana graph — a single flat layer with a controlled long-range degree, designed so a greedy search converges in few hops, because each hop is now an SSD read. The full-precision vectors and adjacency lists live on disk; RAM holds only PQ-compressed vectors used to decide which neighbour to visit next. A search does a few dozen random reads instead of a few million, and the candidates it returns are rescored against the exact vectors it fetched along the way. FreshDiskANN adds in-place updates so the index isn't rebuild-only.
DiskANN serves billion-scale vector search from a single machine by putting the graph and full vectors on NVMe and keeping only PQ-compressed vectors in RAM to guide traversal. Its Vamana graph is built to converge in few hops, because every hop is a disk read. Candidates are rescored with the exact vectors fetched during the walk, so the compression used for navigation doesn't determine final accuracy.
Research talk: Approximate nearest neighbor search systems at scale — Microsoft Research, 9:33