All concepts
Activation Functions
The non-linear bend that lets stacked layers model complex functions — sigmoid, tanh, ReLU, GELU.
Deep Learning · Beginner · ~8 min
In plain English
The decision about whether a neuron passes its signal on, and how strongly. Without one, the whole network is just one big multiplication.
Why it's worth your time
It's a one-token architectural choice that decides whether your network trains, and it's the single most common 'why is my model dead' cause.
If you remember three things
- ReLU: fast, sparse, can die at zero
- GELU/SiLU: smooth, the transformer default
- Sigmoid/softmax belong at the output, not in the hidden stack
Overview
Without a non-linearity, any depth of linear layers collapses to a single linear map. Activations add the bend. Sigmoid/tanh saturate (vanishing gradients); ReLU is the cheap workhorse; GELU is the smooth default in transformers.
How it works
- Why non-linearity Stacking linear layers stays linear. We need a non-linear function between them to gain expressive power.
- Sigmoid Squashes to (0,1). Interpretable as a probability, but saturates in the tails → tiny gradients (vanishing gradients).
- Tanh Like sigmoid but zero-centered (−1,1), which helps optimization. Still saturates.
- ReLU max(0,x): cheap, no saturation for x>0, trains fast. Can 'die' (stuck at 0) — leaky variants fix that.
- GELU A smooth, probabilistic gate — the default activation inside modern transformers.
In an interview
Activations are the non-linearity between linear layers — without them a deep net collapses to one linear layer. Sigmoid and tanh saturate and cause vanishing gradients; ReLU is the cheap default that avoids saturation for positive inputs; GELU is the smooth activation used in transformers.
Production defaults
- Hidden layers
- ReLU for convnets, GELU for transformers. SiLU/SwiGLU in modern LLM feed-forward blocks
- Output
- sigmoid for binary, softmax for multiclass, none for regression
- Dying ReLU
- if many units sit at zero forever, use LeakyReLU or lower the learning rate
What breaks
- A large fraction of units output zero always — Dead ReLUs from too-large updates. Lower the LR, or switch to LeakyReLU/GELU.
- Sigmoid in a deep hidden stack won't train — Saturating derivative — this is the vanishing gradient problem in its original form.