Train one embedding whose first 64 dimensions are already a usable embedding — so you can truncate for speed and only pay full width when it matters.
A set of nesting dolls. Open the big one and there's a smaller one inside that is still recognisably the same doll — so you can carry the small one when you're travelling light.
It turns embedding width from a training decision into a serving dial, and vector-search cost is almost entirely a memory number.
Matryoshka Representation Learning trains a single embedding so that every nested prefix — the first 64, 128, 256, 512 dimensions — is independently a good representation. The loss is applied at each truncation level during training, which forces the model to pack the most discriminative information into the earliest dimensions. At serving time you get a dial nobody had before: search the whole corpus with 64-dim vectors (8× less memory, 8× faster distance math), then re-score the top few hundred with the full 1536 dimensions. Accuracy lands within a point or two of full-width search at a fraction of the cost. OpenAI's text-embedding-3 and Nomic's models ship this by default, which is why `dimensions=256` is a valid API parameter rather than a lossy hack.
Matryoshka embeddings are trained so that each nested prefix of the vector is itself a valid embedding. That means you can truncate to 256 dimensions for a cheap first-pass search over the whole corpus, then re-rank the survivors with the full vector. You get most of the accuracy of full-width search at a fraction of the memory and latency, and it's a serving-time dial rather than a retraining decision.
Introducing EmbeddingGemma: The Best-in-Class Open Model for On-Device Embeddings — Google for Developers, 4:13