All concepts

CNN Convolution

A small kernel slides over the image, computing dot products to detect local patterns.

Deep Learning · Intermediate · ~8 min

In plain English

A small stencil slid across the picture. At each position it checks 'does this little patch look like an edge?' and writes down how strongly.

Why it's worth your time

Convolution encodes the assumption that a cat is a cat wherever it appears — which is why it needs far less data than a dense network for images.

If you remember three things

  • Weight sharing: the same filter runs everywhere
  • Early layers learn edges, deep layers learn objects
  • Stride and pooling shrink the map; padding preserves size

Overview

Convolution applies a small learnable kernel across an image, computing a dot product at each position to produce a feature map. Weight sharing and locality make CNNs parameter-efficient and translation-equivariant; stacking conv+pool layers builds from edges to objects.

How it works

  1. The image An image is a grid of pixel values (here a single channel).
  2. The kernel A small learnable filter (e.g. 3×3) that detects a local pattern — here a vertical edge.
  3. Slide & multiply-add Place the kernel over a patch (its receptive field) and compute the dot product.
  4. Feature map Each result fills one cell of the output. Bright cells = strong response to the pattern.
  5. Pooling & depth Pooling downsamples; stacking conv layers composes edges → textures → objects.

In an interview

A convolution slides a small learnable kernel over the image, computing a dot product at each position to build a feature map — detecting local patterns like edges. Weight sharing makes it parameter-efficient and translation-equivariant; stacking conv and pooling layers composes low-level features into high-level ones.

Production defaults

Kernel
3×3, stacked. Two 3×3 layers see as much as one 5×5 with fewer parameters
Channels
double the channels each time you halve the spatial size
Start from
a pretrained backbone and fine-tune. Training a vision model from scratch is rarely the right call

What breaks

  • Overfits immediately on a small dataset — Augment (flip, crop, colour jitter) and fine-tune a pretrained backbone instead of training from scratch.
  • Works on clean images, fails on real photos — Your training set lacks the real variation. Match augmentation to the actual capture conditions.

Watch it explained

What are Convolutional Neural Networks (CNNs)? — IBM Technology, 6:20

Related