All concepts

CLIP Alignment

Train image and text encoders so matching captions land near matching images.

Deep Learning · Advanced · ~8 min

In plain English

Train two encoders — one for pictures, one for captions — until a photo of a dog and the words 'a dog' land in the same spot on the same map.

Why it's worth your time

It's the reason you can search images with words, and the backbone of nearly every text-to-image system.

If you remember three things

  • Contrastive loss: matched pairs pulled together, mismatched pushed apart
  • One shared embedding space for both modalities
  • Zero-shot classification falls out for free

Overview

CLIP trains an image encoder and a text encoder jointly so matching image-caption pairs land close in a shared embedding space. A contrastive loss pulls true pairs together and pushes mismatches apart, enabling zero-shot classification and image-text retrieval via cosine similarity.

How it works

  1. Start: Image An image encoder converts pixels into a dense vector.
  2. Image -> Caption A text encoder converts the caption into the same embedding space.
  3. Caption -> Contrastive Loss Matched image-text pairs are pulled together; mismatched pairs are pushed apart.
  4. Contrastive Loss -> Shared Space Images and text become comparable with cosine similarity.
  5. Shared Space -> Zero-shot Search A text prompt can retrieve images or classify unseen labels without task-specific training.

In an interview

CLIP jointly trains image and text encoders on hundreds of millions of image-caption pairs with a contrastive objective, so a caption's embedding sits near its image's embedding. At inference you compare a text prompt to images by cosine similarity, giving zero-shot classification and retrieval without task-specific training.

Production defaults

Batch size
the negatives come from within the batch, so large batches matter more here than almost anywhere else
Temperature
learned, initialized around 0.07
Use it for
retrieval and zero-shot labelling. For fine-grained classification, fine-tune a head on top

What breaks

  • Zero-shot accuracy is poor on your domain — Prompt phrasing matters enormously. Ensemble several templates ('a photo of a {}', 'a blurry photo of a {}').
  • Retrieval returns the same few images — Embedding collapse or unnormalized vectors. Normalize before cosine similarity.

Watch it explained

CLIP by OpenAI The Most Powerful AI for Image & Text Understanding (Deep Dive Explained — NobleX Infinity Labs®️, 4:31

Related