All concepts

Activation Functions

The non-linear bend that lets stacked layers model complex functions — sigmoid, tanh, ReLU, GELU.

Deep Learning · Beginner · ~8 min

In plain English

The decision about whether a neuron passes its signal on, and how strongly. Without one, the whole network is just one big multiplication.

Why it's worth your time

It's a one-token architectural choice that decides whether your network trains, and it's the single most common 'why is my model dead' cause.

If you remember three things

  • ReLU: fast, sparse, can die at zero
  • GELU/SiLU: smooth, the transformer default
  • Sigmoid/softmax belong at the output, not in the hidden stack

Overview

Without a non-linearity, any depth of linear layers collapses to a single linear map. Activations add the bend. Sigmoid/tanh saturate (vanishing gradients); ReLU is the cheap workhorse; GELU is the smooth default in transformers.

How it works

  1. Why non-linearity Stacking linear layers stays linear. We need a non-linear function between them to gain expressive power.
  2. Sigmoid Squashes to (0,1). Interpretable as a probability, but saturates in the tails → tiny gradients (vanishing gradients).
  3. Tanh Like sigmoid but zero-centered (−1,1), which helps optimization. Still saturates.
  4. ReLU max(0,x): cheap, no saturation for x>0, trains fast. Can 'die' (stuck at 0) — leaky variants fix that.
  5. GELU A smooth, probabilistic gate — the default activation inside modern transformers.

In an interview

Activations are the non-linearity between linear layers — without them a deep net collapses to one linear layer. Sigmoid and tanh saturate and cause vanishing gradients; ReLU is the cheap default that avoids saturation for positive inputs; GELU is the smooth activation used in transformers.

Production defaults

Hidden layers
ReLU for convnets, GELU for transformers. SiLU/SwiGLU in modern LLM feed-forward blocks
Output
sigmoid for binary, softmax for multiclass, none for regression
Dying ReLU
if many units sit at zero forever, use LeakyReLU or lower the learning rate

What breaks

  • A large fraction of units output zero always — Dead ReLUs from too-large updates. Lower the LR, or switch to LeakyReLU/GELU.
  • Sigmoid in a deep hidden stack won't train — Saturating derivative — this is the vanishing gradient problem in its original form.

Watch it explained

Activation Functions In Neural Networks Explained | Deep Learning Tutorial — AssemblyAI, 6:43

Related