All concepts

How a Vector Database Works

Embed, index, and search: the anatomy of a vector database for RAG.

RAG & Retrieval · Intermediate · ~8 min

In plain English

A database where the query is a point in space and the answer is 'what's near it'. It finds close neighbours without checking every item.

Why it's worth your time

Every AI feature that needs similarity at latency runs on one, and the index choices you make are recall choices.

If you remember three things

  • Approximate nearest neighbour trades a little accuracy for a lot of speed
  • Filtering during the search beats filtering after it
  • Recall is a tunable knob you own, not a fixed property

Overview

A vector database stores embeddings and serves nearest-neighbour queries with metadata filtering. Documents are chunked and embedded, vectors go into an ANN index (HNSW/IVF), and at query time the same encoder embeds the query, ANN search finds the closest vectors, filters constrain by metadata, and the top-k payloads return.

How it works

  1. Ingest & embed Each document or chunk is turned into a vector by an embedding model.
  2. Build the index Vectors go into an ANN index (HNSW/IVF-PQ) so search is sub-linear, not brute force.
  3. Embed the query At query time the same encoder embeds the query into the same vector space.
  4. Search + filter ANN search finds close vectors; metadata filters (tenant, date, ACL) constrain the results.
  5. Return payloads Return the top-k ids with their stored payload/metadata for the app to use.

In an interview

A vector DB embeds documents into vectors, indexes them with an ANN structure like HNSW, and at query time embeds the query, runs approximate nearest-neighbour search, applies metadata filters, and returns the top-k with their payloads. It's the retrieval backbone of RAG.

Production defaults

Index
HNSW: M=16, efConstruction=200, efSearch=100 as a starting point
Distance
cosine on normalized text embeddings. Must match between index and query
Filtering
native pre-filtering, or over-fetch 5–10× before filtering
Sizing
~4 bytes × dims × vectors for floats, then ~1.5× for the graph

What breaks

  • Filtered queries return almost nothing — Post-filtering: ANN found 50, the filter removed 48. Use pre-filtering or over-fetch.
  • Recall quietly degrades as the corpus grows — efSearch that was fine at 100k isn't at 10M. Track recall@k against an exact-search sample on a schedule.

Watch it explained

Vector Databases simply explained! (Embeddings & Indexes) — AssemblyAI, 4:23

Related