All concepts

Matryoshka Embeddings

Train one embedding whose first 64 dimensions are already a usable embedding — so you can truncate for speed and only pay full width when it matters.

Advanced Embeddings · Advanced · ~6 min

In plain English

A set of nesting dolls. Open the big one and there's a smaller one inside that is still recognisably the same doll — so you can carry the small one when you're travelling light.

Why it's worth your time

It turns embedding width from a training decision into a serving dial, and vector-search cost is almost entirely a memory number.

If you remember three things

  • Every prefix of the vector is itself a valid embedding
  • Search truncated over everything, rescore full over hundreds
  • Re-normalise after truncating — the norm changes

Overview

Matryoshka Representation Learning trains a single embedding so that every nested prefix — the first 64, 128, 256, 512 dimensions — is independently a good representation. The loss is applied at each truncation level during training, which forces the model to pack the most discriminative information into the earliest dimensions. At serving time you get a dial nobody had before: search the whole corpus with 64-dim vectors (8× less memory, 8× faster distance math), then re-score the top few hundred with the full 1536 dimensions. Accuracy lands within a point or two of full-width search at a fraction of the cost. OpenAI's text-embedding-3 and Nomic's models ship this by default, which is why `dimensions=256` is a valid API parameter rather than a lossy hack.

In an interview

Matryoshka embeddings are trained so that each nested prefix of the vector is itself a valid embedding. That means you can truncate to 256 dimensions for a cheap first-pass search over the whole corpus, then re-rank the survivors with the full vector. You get most of the accuracy of full-width search at a fraction of the memory and latency, and it's a serving-time dial rather than a retraining decision.

Production defaults

Start at
256 dimensions for the index, full width for the rescore. Measure recall@10 at 64/128/256/512 on your corpus before committing
Rescore depth
300-500 candidates. Cheap, because it's a dot product over a few hundred vectors
Storage
full vectors on object storage or a column store; only the truncated index needs RAM

What breaks

  • Truncation destroyed accuracy — The model wasn't trained with MRL. Only Matryoshka-trained models survive truncation — check the model card before assuming.
  • Scores drift after truncation — You didn't re-normalise. Cosine assumes unit vectors and truncation changes the norm.

Watch it explained

Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings — Google for Developers, 4:13

Related