All concepts
Neural Network Forward Pass
Data flows layer by layer: linear combination, non-linear activation, repeat — producing a prediction.
Deep Learning · Beginner · ~9 min
In plain English
A stack of adjustable dimmer switches. Numbers go in the front, each layer mixes and re-weights them, and by the end the mixture that comes out is the answer.
Why it's worth your time
Every model in the second half of this site — transformers, diffusion, agents — is this same object with a different wiring diagram.
If you remember three things
- A layer is a matrix multiply plus a bias plus a non-linearity
- Without the non-linearity, a hundred layers collapse into one
- Width learns more features; depth learns more composition
Overview
A feed-forward network stacks layers of neurons. Each layer computes a weighted sum of its inputs, adds a bias, and applies a non-linear activation. Stacking non-linear layers lets the network approximate complex functions. The forward pass is just this data flow from input to output.
How it works
- Input features The input vector x enters the network — one value per input neuron.
- Weighted sum Each neuron computes z = Σ wᵢxᵢ + b: a linear combination of its inputs plus a bias.
- Activation A non-linearity like ReLU or sigmoid transforms z. Without it, stacked layers collapse to one linear map.
- Hidden layer The activations become inputs to the next layer. Deeper layers compose features into higher-level ones.
- Output The final layer produces the prediction — softmax for classes, linear for regression.
In an interview
A neural network's forward pass alternates linear layers (weighted sum + bias) with non-linear activations. Stacking non-linear layers lets it approximate complex functions. Without the non-linearity, depth would collapse into a single linear model.
Production defaults
- Start
- 2–3 hidden layers, width 64–256, ReLU/GELU. Go deeper only when the shallow one is bias-limited
- Batch
- 32–256. Larger batches need a proportionally larger learning rate
- Normalize
- batch norm for vision, layer norm for sequences. It's a training-stability tool, not an accuracy trick
What breaks
- Loss won't move at all — Overfit a batch of 8 examples first. If it can't reach zero there, the bug is in the data pipeline or the loss, not the model.
- Beaten by gradient boosting on tabular data — Usually the correct outcome. Neural nets win on text, images and audio; trees still win on tables.