All concepts
Two Tower
Embed users and items separately so nearest-neighbor search can retrieve recommendations quickly.
Applied ML · Advanced · ~8 min
In plain English
Train two encoders — one for users, one for items — so that a user's vector lands near the items they'd like. Item vectors can be computed in advance.
Why it's worth your time
It's how recommendation retrieval works at scale: all the expensive work happens offline, and serving is a nearest-neighbour lookup.
If you remember three things
- Towers never see each other except through the dot product
- Item embeddings are pre-computed and indexed
- Serving is one user encode plus an ANN search
Overview
A retrieval architecture that embeds users and items with separate encoders into a shared vector space, trained so compatible pairs land close together. Because item vectors are precomputed, serving reduces to fast approximate nearest-neighbor search over the user vector — the workhorse of retrieval-stage recommenders.
How it works
- Start: User Features A user tower encodes profile, history, and context.
- User Features -> Item Features An item tower encodes product, video, article, or company features.
- Item Features -> Shared Space Train with contrastive or ranking losses so compatible user-item pairs are close.
- Shared Space -> ANN Search At serving, embed the user and find nearest item vectors.
- ANN Search -> Recommendations Two-tower models are the workhorse for retrieval-stage recommenders.
In an interview
Two-tower models learn separate encoders for users and items that map into one embedding space, trained with a contrastive or ranking loss so matching pairs have high similarity. Item embeddings are precomputed and indexed, so retrieval is just an ANN lookup on the user vector — fast enough for candidate generation over millions of items.
Production defaults
- Negatives
- in-batch negatives plus hard negatives. Random-only negatives make the task too easy to learn from
- Refresh
- re-embed items on change; re-train on a cadence matched to catalogue churn
- Serving
- ANN index over item vectors — same infrastructure as any vector search
- Then rerank
- the two-tower model is a retrieval stage, not a final ranker
What breaks
- Recommends only popular items — Popularity bias from sampled negatives. Correct with logQ / sampling-bias correction.
- Cold-start items never surface — No interactions means no signal. Add content features to the item tower.