All concepts

Vision Transformer (ViT)

Treat an image as a sequence of patches and run a transformer over them.

Deep Learning · Advanced · ~8 min

In plain English

Chop the image into a grid of small tiles, treat each tile as a word, and run the same machinery that reads sentences.

Why it's worth your time

It showed that the convolution's built-in assumptions aren't necessary if you have enough data — and unified vision and language under one architecture.

If you remember three things

  • Patches become tokens; position embeddings restore the layout
  • No built-in translation invariance, so it needs more data than a CNN
  • Attention is global from layer one

Overview

ViT applies the transformer architecture to images by splitting a picture into fixed-size patches, embedding each patch as a token, and running self-attention over the sequence. Every patch can attend to every other, giving a global receptive field from the first layer instead of the local one of CNNs.

How it works

  1. Start: Image Start with an image tensor rather than a sentence.
  2. Image -> Patchify Split the image into fixed-size patches, like visual tokens.
  3. Patchify -> Patch Embeddings Flatten and project each patch into a vector, then add positional information.
  4. Patch Embeddings -> Transformer Self-attention lets every patch attend to every other patch.
  5. Transformer -> Class Token A classification head reads the final class token or pooled representation.

In an interview

A Vision Transformer treats an image as a sequence of patches — typically 16x16 — flattens and linearly projects each into a token, adds positional embeddings, and feeds them to a standard transformer encoder. A class token's final representation drives classification. With enough pretraining data it matches or beats CNNs, since self-attention gives global context immediately.

Production defaults

Patch size
16×16 for 224px inputs — that's 196 tokens
Data
pretrain on something large or start from pretrained weights. From scratch on a small dataset, a CNN wins
Fine-tuning
at a higher resolution than pretraining usually helps; interpolate the position embeddings

What breaks

  • Badly beaten by a ResNet on your dataset — Too little data. ViTs lack the CNN's spatial prior and have to learn it — use pretrained weights.
  • Memory blows up at high resolution — Attention is quadratic in token count, and tokens scale with area. Use a windowed variant (Swin) or larger patches.

Watch it explained

An image is worth 16x16 words: ViT | Vision Transformer explained — AI Coffee Break with Letitia, 5:25

Related