All concepts
Embedding Vector Space
Map text to vectors where semantic similarity becomes geometric closeness.
NLP & Embeddings · Beginner · ~8 min
In plain English
A map where meaning is location. Things that mean similar things end up near each other, so 'find similar' becomes 'find nearby'.
Why it's worth your time
It's the foundation of search, recommendations, clustering, and every RAG system you will ever build.
If you remember three things
- Direction carries meaning; cosine similarity compares direction
- Embeddings from different models are not comparable
- The space reflects the training data, biases included
Overview
Embeddings map tokens, words, or documents to dense vectors so that similar meanings land near each other. Cosine similarity then measures semantic relatedness. Embeddings power semantic search, RAG retrieval, clustering, and recommendations.
How it works
- Words become vectors Each word or document is mapped to a point in a high-dimensional space (shown here in 2D).
- Similar meanings cluster The model is trained so related items — 'king', 'queen', 'prince' — land close together.
- Cosine similarity = closeness The angle between two vectors measures how related they are: small angle → high similarity.
- A query finds its neighbors Embed a query and its nearest neighbors are the most semantically relevant items.
- This powers semantic search & RAG Retrieval just returns the closest vectors — the backbone of semantic search and RAG.
In an interview
Embeddings turn text into dense vectors where distance reflects meaning, so semantic similarity becomes cosine similarity. They're the backbone of semantic search, RAG retrieval, clustering, and recommendations.
Production defaults
- Similarity
- cosine on normalized vectors. Normalize once at index time, not per query
- Dimensions
- 384–1024 covers most needs. Larger costs memory and rarely repays it
- Versioning
- changing the embedding model means re-indexing everything. Store the model version with every vector
What breaks
- Search quality collapsed after a model upgrade — Mixed embedding versions in one index. They live in different spaces — re-index the whole corpus.
- Everything looks 0.8 similar — Cosine on unnormalized vectors, or a model where all embeddings cluster. Normalize, and judge by RANK not absolute score.