All concepts
Vision Transformer (ViT)
Treat an image as a sequence of patches and run a transformer over them.
Deep Learning · Advanced · ~8 min
In plain English
Chop the image into a grid of small tiles, treat each tile as a word, and run the same machinery that reads sentences.
Why it's worth your time
It showed that the convolution's built-in assumptions aren't necessary if you have enough data — and unified vision and language under one architecture.
If you remember three things
- Patches become tokens; position embeddings restore the layout
- No built-in translation invariance, so it needs more data than a CNN
- Attention is global from layer one
Overview
ViT applies the transformer architecture to images by splitting a picture into fixed-size patches, embedding each patch as a token, and running self-attention over the sequence. Every patch can attend to every other, giving a global receptive field from the first layer instead of the local one of CNNs.
How it works
- Start: Image Start with an image tensor rather than a sentence.
- Image -> Patchify Split the image into fixed-size patches, like visual tokens.
- Patchify -> Patch Embeddings Flatten and project each patch into a vector, then add positional information.
- Patch Embeddings -> Transformer Self-attention lets every patch attend to every other patch.
- Transformer -> Class Token A classification head reads the final class token or pooled representation.
In an interview
A Vision Transformer treats an image as a sequence of patches — typically 16x16 — flattens and linearly projects each into a token, adds positional embeddings, and feeds them to a standard transformer encoder. A class token's final representation drives classification. With enough pretraining data it matches or beats CNNs, since self-attention gives global context immediately.
Production defaults
- Patch size
- 16×16 for 224px inputs — that's 196 tokens
- Data
- pretrain on something large or start from pretrained weights. From scratch on a small dataset, a CNN wins
- Fine-tuning
- at a higher resolution than pretraining usually helps; interpolate the position embeddings
What breaks
- Badly beaten by a ResNet on your dataset — Too little data. ViTs lack the CNN's spatial prior and have to learn it — use pretrained weights.
- Memory blows up at high resolution — Attention is quadratic in token count, and tokens scale with area. Use a windowed variant (Swin) or larger patches.