All concepts
Recurrent Neural Network
Process sequences one step at a time while carrying hidden state forward.
Deep Learning · Intermediate · ~8 min
In plain English
Reading a sentence one word at a time while keeping a running note of what you've understood so far, and updating that note at every word.
Why it's worth your time
It's the architecture transformers replaced, and knowing why they replaced it is the cleanest way to understand what attention buys you.
If you remember three things
- State carries information forward one step at a time
- Sequential by nature — it cannot be parallelized across time
- LSTM/GRU gates exist to keep the gradient alive over long spans
Overview
A recurrent neural network processes a sequence one step at a time, updating a hidden state that carries information forward. The same weights apply at every timestep via h_t = f(W_x x_t + W_h h_{t-1}), so the hidden state acts as a running summary of everything seen so far.
How it works
- Start: Token 1 The first item enters with an initial hidden state.
- Token 1 -> Hidden State The RNN updates memory using the current input and previous state.
- Hidden State -> Next Token The same cell repeats across time with shared weights.
- Next Token -> Sequence Memory Hidden state summarizes earlier tokens, though long dependencies are difficult.
- Sequence Memory -> Output RNNs power simple sequence models; LSTMs/GRUs add gates to preserve information longer.
In an interview
An RNN handles sequences by maintaining a hidden state updated at each step: h_t = f(W_x x_t + W_h h_{t-1}), reusing the same weights across time. That recurrence models order and variable length, but backpropagation through many steps causes vanishing or exploding gradients, so plain RNNs struggle with long-range dependencies.
Production defaults
- Use today
- rarely for text; still reasonable for short, strictly-ordered sensor or time-series data
- If you must
- GRU over vanilla RNN, gradient clipping at 1.0, and truncated backprop through time
- Otherwise
- a small transformer will usually be both faster to train and more accurate
What breaks
- Forgets anything more than ~20 steps back — The vanishing gradient over time. Gates help; attention solves it properly.
- Training is far slower than a transformer — Inherent — the recurrence forbids parallelism across the sequence. That's the whole reason for the switch.