All concepts
CLIP Alignment
Train image and text encoders so matching captions land near matching images.
Deep Learning · Advanced · ~8 min
In plain English
Train two encoders — one for pictures, one for captions — until a photo of a dog and the words 'a dog' land in the same spot on the same map.
Why it's worth your time
It's the reason you can search images with words, and the backbone of nearly every text-to-image system.
If you remember three things
- Contrastive loss: matched pairs pulled together, mismatched pushed apart
- One shared embedding space for both modalities
- Zero-shot classification falls out for free
Overview
CLIP trains an image encoder and a text encoder jointly so matching image-caption pairs land close in a shared embedding space. A contrastive loss pulls true pairs together and pushes mismatches apart, enabling zero-shot classification and image-text retrieval via cosine similarity.
How it works
- Start: Image An image encoder converts pixels into a dense vector.
- Image -> Caption A text encoder converts the caption into the same embedding space.
- Caption -> Contrastive Loss Matched image-text pairs are pulled together; mismatched pairs are pushed apart.
- Contrastive Loss -> Shared Space Images and text become comparable with cosine similarity.
- Shared Space -> Zero-shot Search A text prompt can retrieve images or classify unseen labels without task-specific training.
In an interview
CLIP jointly trains image and text encoders on hundreds of millions of image-caption pairs with a contrastive objective, so a caption's embedding sits near its image's embedding. At inference you compare a text prompt to images by cosine similarity, giving zero-shot classification and retrieval without task-specific training.
Production defaults
- Batch size
- the negatives come from within the batch, so large batches matter more here than almost anywhere else
- Temperature
- learned, initialized around 0.07
- Use it for
- retrieval and zero-shot labelling. For fine-grained classification, fine-tune a head on top
What breaks
- Zero-shot accuracy is poor on your domain — Prompt phrasing matters enormously. Ensemble several templates ('a photo of a {}', 'a blurry photo of a {}').
- Retrieval returns the same few images — Embedding collapse or unnormalized vectors. Normalize before cosine similarity.