All concepts

Neural Network Forward Pass

Data flows layer by layer: linear combination, non-linear activation, repeat — producing a prediction.

Deep Learning · Beginner · ~9 min

In plain English

A stack of adjustable dimmer switches. Numbers go in the front, each layer mixes and re-weights them, and by the end the mixture that comes out is the answer.

Why it's worth your time

Every model in the second half of this site — transformers, diffusion, agents — is this same object with a different wiring diagram.

If you remember three things

  • A layer is a matrix multiply plus a bias plus a non-linearity
  • Without the non-linearity, a hundred layers collapse into one
  • Width learns more features; depth learns more composition

Overview

A feed-forward network stacks layers of neurons. Each layer computes a weighted sum of its inputs, adds a bias, and applies a non-linear activation. Stacking non-linear layers lets the network approximate complex functions. The forward pass is just this data flow from input to output.

How it works

  1. Input features The input vector x enters the network — one value per input neuron.
  2. Weighted sum Each neuron computes z = Σ wᵢxᵢ + b: a linear combination of its inputs plus a bias.
  3. Activation A non-linearity like ReLU or sigmoid transforms z. Without it, stacked layers collapse to one linear map.
  4. Hidden layer The activations become inputs to the next layer. Deeper layers compose features into higher-level ones.
  5. Output The final layer produces the prediction — softmax for classes, linear for regression.

In an interview

A neural network's forward pass alternates linear layers (weighted sum + bias) with non-linear activations. Stacking non-linear layers lets it approximate complex functions. Without the non-linearity, depth would collapse into a single linear model.

Production defaults

Start
2–3 hidden layers, width 64–256, ReLU/GELU. Go deeper only when the shallow one is bias-limited
Batch
32–256. Larger batches need a proportionally larger learning rate
Normalize
batch norm for vision, layer norm for sequences. It's a training-stability tool, not an accuracy trick

What breaks

  • Loss won't move at all — Overfit a batch of 8 examples first. If it can't reach zero there, the bug is in the data pipeline or the loss, not the model.
  • Beaten by gradient boosting on tabular data — Usually the correct outcome. Neural nets win on text, images and audio; trees still win on tables.

Watch it explained

Neural Networks Explained in 5 minutes — IBM Technology, 4:31

Related