All concepts

Recurrent Neural Network

Process sequences one step at a time while carrying hidden state forward.

Deep Learning · Intermediate · ~8 min

In plain English

Reading a sentence one word at a time while keeping a running note of what you've understood so far, and updating that note at every word.

Why it's worth your time

It's the architecture transformers replaced, and knowing why they replaced it is the cleanest way to understand what attention buys you.

If you remember three things

  • State carries information forward one step at a time
  • Sequential by nature — it cannot be parallelized across time
  • LSTM/GRU gates exist to keep the gradient alive over long spans

Overview

A recurrent neural network processes a sequence one step at a time, updating a hidden state that carries information forward. The same weights apply at every timestep via h_t = f(W_x x_t + W_h h_{t-1}), so the hidden state acts as a running summary of everything seen so far.

How it works

  1. Start: Token 1 The first item enters with an initial hidden state.
  2. Token 1 -> Hidden State The RNN updates memory using the current input and previous state.
  3. Hidden State -> Next Token The same cell repeats across time with shared weights.
  4. Next Token -> Sequence Memory Hidden state summarizes earlier tokens, though long dependencies are difficult.
  5. Sequence Memory -> Output RNNs power simple sequence models; LSTMs/GRUs add gates to preserve information longer.

In an interview

An RNN handles sequences by maintaining a hidden state updated at each step: h_t = f(W_x x_t + W_h h_{t-1}), reusing the same weights across time. That recurrence models order and variable length, but backpropagation through many steps causes vanishing or exploding gradients, so plain RNNs struggle with long-range dependencies.

Production defaults

Use today
rarely for text; still reasonable for short, strictly-ordered sensor or time-series data
If you must
GRU over vanilla RNN, gradient clipping at 1.0, and truncated backprop through time
Otherwise
a small transformer will usually be both faster to train and more accurate

What breaks

  • Forgets anything more than ~20 steps back — The vanishing gradient over time. Gates help; attention solves it properly.
  • Training is far slower than a transformer — Inherent — the recurrence forbids parallelism across the sequence. That's the whole reason for the switch.

Watch it explained

The Power of Recurrent Neural Networks (RNN) — IBM Technology, 7:46

Related