Level N3 · Foundations · Classic networks · runs in your browser

Recurrent networks

How can a network read a sequence one step at a time?

Side trip · best after level 10 · Probability and sampling

An MLP takes a fixed number of inputs. A sentence does not have a fixed length, and the order of its words matters. A recurrent network (RNN) reads a sequence one item at a time. It carries a short list of numbers, the hidden state h, from one step to the next. The hidden state is how the network remembers what it has read so far.

1. One step at a time

At every step the network mixes the new input with the hidden state, then passes the result through tanh, which keeps it between −1 and 1:

ht=tanh⁡(wx xt+wh ht−1+b)h_t = \tanh(w_x \, x_t + w_h \, h_{t-1} + b)

Below, the hidden state is a single number. The sequence is 1, 0, 0, 0, and h starts at 0. Press Next step to run it.

A recurrent network, one step at a time

Press Next step. Change the weights or the inputs x and watch h change at every step.

h₀=0× w_xx=1h1?× w_h× w_xx=0h2…× w_h× w_xx=0h3…× w_h× w_xx=0h4…

z = w_x·x_t + w_h·h_(t−1) + b, then h_t = tanh(z)

tx_tzh_tsize of h
11.00×1 + 0.50×0.0000 + 0.00 = 1.0000 ?
2……
3……
4……
step 1 of 4: h1 = ?. The same w_x, w_h and b are used at every step; only h carries anything forward.
h above 0 h below 0 deeper color: bigger |h|
Number wₓ = 1, wₕ = 0.5, b = 0, h₀ = 0 and x₁ = 1. Use tanh(1) ≈ 0.76. What is h₁ = tanh(wₓ·x₁ + wₕ·h₀ + b)? (2 decimals)
🔒 Answer the question above to unlock

Starting at step 2, the input is 0. The only thing that carries the 1 forward is wh ht−1w_h \, h_{t-1}.

Number wₓ = 1, wₕ = 0.5, b = 0. h₁ = 0.76 and x₂ = 0. Use tanh(0.38) ≈ 0.36. What is h₂ = tanh(wₓ·x₂ + wₕ·h₁ + b)? (2 decimals)
🔒 Answer the question above to unlock

Each step keeps about half of what was there, so the 1 from step 1 gets smaller and smaller: 0.76, 0.36, 0.18, 0.09.

Number This RNN has a one-number hidden state: hₜ = tanh(wₓ·xₜ + wₕ·hₜ₋₁ + b). It reads a sequence of 1,000 inputs. How many learned numbers (weights and biases) does it have?
I got stuck here Is there a separate network for every step?

No. Drawings often show the network copied once per step, side by side. That drawing is called unrolling, and it helps to see the flow. But every copy uses the same wxw_x, whw_h and bb. A sequence of 4 steps and a sequence of 4,000 steps use the same 3 weights in this lab. With vectors, it is the same 3 matrices.

2. Training through time

🔒 Answer the question above to unlock

To train the network, the loss at the last step has to tell step 1 what to change. Backpropagation walks back through the unrolled steps. At every step it multiplies by the same kind of factor: whw_h times the slope of tanh there, which is 1−h21 - h^2.

∂h4∂h1=(wh(1−h22))(wh(1−h32))(wh(1−h42))\frac{\partial h_4}{\partial h_1} = \big(w_h (1 - h_2^2)\big)\big(w_h (1 - h_3^2)\big)\big(w_h (1 - h_4^2)\big)
Choosewₕ = 0.5. Going back one step multiplies the gradient by wₕ·(1 − h²). Ignoring tanh, three steps back would give 0.5³ = 0.125. With tanh, ∂h₄/∂h₁ is…

The next question leaves tanh out, so every factor is just whw_h.

ChooseNo activation, wₕ = 1.5. After 30 steps the gradient is a product of 29 factors of 1.5. About how big is it?

How much does step 1 still matter at step t?

Drag w_h and switch the activation. The line is the gradient d h_t / d h_1 over 30 steps.

vanishingexplodingd h_t / d h_1 (log scale)1e61e31e01e-31e-61e-91e-12t=1t=10t=20t=30
at t = 30, d h_30 / d h_1 = 1.55e-9: the first input has almost no say, so training does not see that it matters
input 1 at step 1, then 0; w_x = 1, b = 0 each step multiplies by w_h × (slope of the activation); tanh’s slope is at most 1 vanishing or exploding

Each factor is whw_h times a tanh slope, and the slope 1−h21 - h^2 is always between 0 and 1.

  • If ∣wh∣|w_h| is below 1, every factor is below 1, and a long product shrinks toward 0. This is the vanishing gradient: the first steps get almost no training signal.
  • If ∣wh∣|w_h| is well above 1, a factor can stay above 1 and the product grows very large: an exploding gradient. tanh does not prevent that. It only helps where h is near ±1, because there the slope is close to 0.

Either way, a product of many factors almost never stays near 1. The explode question above left tanh out to make the arithmetic easy, but the gradient grows the same way with tanh whenever wh(1−h2)w_h (1 - h^2) stays above 1.

Go deeper An exploding gradient is easy to fix, a vanishing one is not

For an exploding gradient there is a simple fix: if its length is above some limit, scale it down to the limit. That is gradient clipping, and demo.py uses it (clip_grad_norm_(…, 1.0)).

A vanishing gradient has no such fix. Making a number like 1e-9 bigger also makes the noise from every other step. The information about step 1 is simply lost on the way back. The fix has to change the network itself. That is level N4.

Here is the same chain drawn as a network: the sequence 1, 0, 0, 0 with wx=1w_x = 1, wh=0.5w_h = 0.5, b=0b = 0, as in the lab of section 1.

One cell, used four times

Press Next for the forward steps, then keep going to send the gradient back.

x1h1x2h2x3h3x4h4w_xw_xw_hw_xw_hw_xw_h
Press Next to start. · 3 parameters in all
Layers (8)

Look at the color: every step is the same cell with the same three weights. Going back, the bright area shrinks at every step.

positive valuenegative value A disc’s fill is its value: the stronger the color, the further from 0. the same weights, used again forward (this step) backward: gradient “∂ −0.5” = the gradient ∂L/∂(that value). Select a layer to see its numbers.

3. Remember the first symbol

🔒 Answer the question above to unlock

Here is a task where the network must remember. A sequence starts with A or B, then k − 1 random extra symbols (c, d or e). At the end the network must say which symbol came first. Guessing gives 50%.

ChooseAn RNN with a 32-number hidden state learns “remember the first symbol” perfectly for k = 20. Now k = 40: same rule, 1,500 training steps. What accuracy will it reach?

Remember the first symbol: real training runs

Pick a sequence length k. Try the longer ones.

Loading runs…

…
RNN 1,500 steps of 64 sequences, a 32-number hidden state, accuracy on 1,000 new sequences

Up to k = 20, the RNN learns the task perfectly. At k = 40 it stays at guessing, even though the rule is just as simple.

I got stuck here Why not train longer, or with a bigger learning rate?

The training signal that should reach step 1 is a product of 39 factors, each well below 1. It is far smaller than the noise coming from the random symbols near the end. More steps or a bigger learning rate amplify that noise just as much. The network never “hears” that step 1 mattered.

Try it

In the step lab, set whw_h to 1 and run all 4 steps. The 1 now shrinks much more slowly. Then set whw_h to −1. What happens to the sign of h at each step?

4. Write it yourself

With vectors, x is a list of D numbers and h is a list of H numbers. As in every level since level 1, vectors are rows and a layer is “row @ matrix”:

ht=tanh⁡(xt Wx+ht−1 Wh+b)h_t = \tanh(x_t \, W_x + h_{t-1} \, W_h + b)

So WxW_x is (D, H): it turns D input numbers into H. WhW_h is (H, H), and b has H numbers. In code that is x @ Wx + h @ Wh + b. With a one-number hidden state, xtWxx_t W_x is just wxxtw_x x_t, the formula from section 1.

Write the one line inside the loop.

CodeWrite the update inside the loop: h becomes tanh of (x times Wₓ, plus h times Wₕ, plus b).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

rnn keeps only the last h. Often you need the hidden state after every step, for example to read the whole sequence again later (level N5 does that). Collect them in a list: start with hs = [], and after each update call hs.append(h). At the end, np.stack(hs) turns a list of L rows of H numbers into one array of shape (L, H).

CodeWrite rnn_all: run the same loop, but keep every hidden state, so the result has one row per step, shape (L, H). Replace ____ with as many lines as you need.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Run an RNN by hand: mix the input and the old hidden state, then take tanh.
  • Explain why the gradient vanishes or explodes over many steps: it is a product of one factor per step.
  • Write an RNN loop in NumPy that keeps every hidden state, shape (L, H).

Keep in mind

  • , with (D, H), (H, H), b (H,)
  • Every step uses the same , and : the number of weights does not depend on the length
  • One step back multiplies the gradient by ; many such factors shrink to 0 or grow very large

Common mistakes

  • Counting new weights for every step: the steps share one set of weights.
  • Adding the factors instead of multiplying them: 29 factors of 1.5 give about 128,000, not 45.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars