All concepts

Learned Sparse Retrieval (SPLADE)

Let the language model decide the term weights — a sparse vector over the vocabulary that expands documents with terms they never contained.

Advanced Embeddings · Advanced · ~6 min

In plain English

Keyword search where the keywords are chosen by someone who read the document and understood it — including words the document never actually used.

Why it's worth your time

It closes most of the semantic gap while staying on inverted-index infrastructure you already run and can debug by reading it.

If you remember three things

  • Sparse vector over the vocabulary, weights learned not counted
  • FLOPS regulariser is what keeps it sparse and servable
  • Fully interpretable: print the terms and their weights

Overview

BM25 is sparse and interpretable but literal: it can only match words that are actually present. Dense embeddings generalise but lose exact terms. Learned sparse retrieval keeps the sparse, invertible-index-friendly shape of BM25 and lets a transformer fill in the weights. SPLADE projects each token onto the full vocabulary through the MLM head, takes a log-saturated max over positions, and applies FLOPS regularisation so the result stays sparse — a few hundred non-zero terms out of 30k. The output is a bag of weighted terms that includes words the document never used, so a document about 'reimbursement' picks up weight on 'refund'. Because it is still a sparse vector, it serves on the same inverted index infrastructure you already run.

In an interview

SPLADE produces a sparse vector over the vocabulary where a transformer, not a term-frequency formula, sets the weights. Because the model projects through its masked-language-model head, documents get weight on related terms they never contained — expansion for free. It keeps BM25's inverted-index serving story while closing most of the semantic gap, which is why it's a strong hybrid partner for dense retrieval.

Production defaults

Sparsity
aim for 100-300 non-zero terms per document. More effectiveness, worse p95 — that's the whole trade
Efficient config
expand documents at index time, keep the query side lean, when latency is tight
Fusion
RRF with a dense retriever; they fail on different queries, which is what makes fusion worth the second index

What breaks

  • Query latency doubled — Expansion touches more posting lists. Raise the FLOPS coefficient, or expand documents only and leave queries as plain BM25.
  • Index bloated — The regulariser is too weak and the 'sparse' vectors are filling in. Sparsity is a tuned property, not a given.

Watch it explained

SPLADE: Sparse Lexical Models for Efficient Search Ranking — IR with PUGGY, 6:35

Related