All concepts

Fine-tuning Embeddings

Contrastive training on your own query–document pairs, with mined hard negatives — usually the single biggest retrieval win available.

Advanced Embeddings · Advanced · ~7 min

In plain English

Teaching the search engine your company's dialect. Right now it thinks 'churn' is about butter, and every question your users ask lands in the wrong part of the map.

Why it's worth your time

A few thousand real query–document pairs usually beat months of chunking and prompt tuning. It is the highest-leverage retrieval work available.

If you remember three things

  • Contrastive: pull the true pair together, push the rest apart
  • Hard negatives are where the gain is; in-batch negatives plateau
  • A mined negative that is secretly correct teaches the opposite lesson

Overview

An off-the-shelf embedding model knows general English, not that in your product 'churn' means a subscription event and not butter. Fine-tuning fixes the geometry: pull genuine query–document pairs together, push everything else apart. The signal that matters is the negatives. In-batch negatives are nearly free but too easy, so the model plateaus; mined hard negatives — documents the current retriever ranks highly but that are wrong — are what actually move recall. The training loop is InfoNCE with a temperature, the evaluation is recall@k and nDCG on a held-out golden set, and the trap is false negatives: a mined 'negative' that is actually a correct answer teaches the model precisely the wrong thing.

In an interview

Fine-tuning an embedding model on your own query–document pairs reshapes the vector space around your domain's meaning of words. The training objective is contrastive — pull the true pair together, push others apart — and the quality of the result is dominated by the negatives you mine. Hard negatives, filtered so they aren't secretly correct, are where the recall gain comes from; in-batch negatives alone plateau quickly.

Production defaults

Data
1-5k query–document pairs is enough to see a real move. Click logs, resolved tickets, or LLM-generated queries filtered by a cross-encoder
Negatives
4-8 mined hard negatives per query, filtered with a cross-encoder to drop false negatives
Training
MultipleNegativesRankingLoss, batch 64+ (batch size IS the negative count), 1-3 epochs, warmup ~10%

What breaks

  • Recall barely moved — Random negatives only. Mine hard negatives with the current retriever — that's the signal that changes the geometry.
  • Recall got worse on some queries — False negatives in the mined set. Filter with a cross-encoder and drop anything scoring above threshold.
  • Great offline, no better in prod — You evaluated on the training distribution. Hold out a golden query set before you start.

Watch it explained

Fine tuning Embeddings Model — Genpakt, 7:07

Related