Level 6 · Foundations · runs in your browser

Backpropagation

How does the gradient flow back through the network, layer by layer?

Boss level

Write the backward pass for a 2→6→1 network and prove it right with a gradient check. Then train the network until all four XOR points are correct, in your browser.

In level 2 you could write the gradient by hand: one formula, two parameters. A network has thousands of parameters spread over many layers. Backpropagation is a way to find all their gradients in one pass, using a single rule again and again. That rule is the chain rule (the math page shows it with numbers).

Here is the smallest network that still has layers: one input, one neuron, one loss.

z=w⋅x+b→a=sigmoid(z)→L=(a−y)2z = w \cdot x + b \quad\to\quad a = \text{sigmoid}(z) \quad\to\quad L = (a - y)^2

With x=2x = 2, w=0.5w = 0.5, b=−1b = -1, y=1y = 1: z=0z = 0, so a=sigmoid(0)=0.5a = \text{sigmoid}(0) = 0.5 and L=(0.5−1)2=0.25L = (0.5 - 1)^2 = 0.25. We want ∂L/∂w\partial L / \partial w.

1. The one rule

Each box knows only its own local gradient: how its output moves when its input moves. Sigmoid knows ∂a/∂z=a(1−a)\partial a / \partial z = a(1 - a). The square knows ∂L/∂a=2(a−y)\partial L / \partial a = 2(a - y). To get from LL back to ww, multiply the local gradients along the path. That is the chain rule.

Press “Step back” one step at a time. Each new gradient shows as “?” until you compute it in the question below.

Send the gradient back from L, one link at a time

Press “Step back” to move the gradient one box further back. The numbers in w, x, b and the target y can be edited.

∂L/∂w ·
∂L/∂x ·
∂L/∂b ·
z = w·x + b 0 ∂L/∂z ·
a = σ(z) 0.5 ∂L/∂a ·
L = (a − y)² 0.25 ∂L/∂L ·
  1. Forward: z = 0.5·2 + -1 = 0, a = σ(0) = 0.5, L = (0.5 − 1)² = 0.25
step 0 / 4
Forward pass done. Press “Step back” to start the gradient at L.
0.5 forward value ∂L/∂· gradient, flowing back where the gradient is now

Check ∂L/∂w with a small change

Change w by ε both ways and compare the slope with the chain-rule answer. Move ε to change the size of the change.

(L(w + ε) − L(w − ε)) / 2ε = ?
chain rule ∂L/∂w = ?
difference = 1.89e-11
Number a = 0.5 and y = 1. What is the local gradient ∂L/∂a = 2(a − y)?
🔒 Answer the question above to unlock

The next step needs your arithmetic. The lab stops at a question mark.

Number A neuron computes a = sigmoid(z), and the loss is L = (a − y)². Going back, ∂L/∂a = −1 and ∂a/∂z = a(1 − a) = 0.5 × 0.5 = 0.25. Use the chain rule: what is ∂L/∂z?
🔒 Answer the question above to unlock

One more multiplication reaches ww.

Number A neuron computes z = w·x + b with x = 2. Going back, ∂L/∂z = −0.25, and the local gradient is ∂z/∂w = x = 2. What is ∂L/∂w?

Notice what ∂L/∂z\partial L / \partial z did. Once you had it, ww and bb each cost one more multiplication. Going backward means every gradient is computed once and reused by all the earlier layers. That is why it is fast.

I got stuck here Why go backward? Couldn’t I start at w and go forward?

You could, for one parameter. Starting at ww, you would carry “how much does z move per unit of w” forward to L. But then for bb you would start again from the beginning, and for every other weight again. A network with a million weights would need a million forward passes.

Going backward, you start from the one thing everyone shares, the loss, and each box passes one number to the boxes before it. One backward pass gives every gradient, for roughly the cost of one or two forward passes.

🔒 Answer the question above to unlock

2. How do you know the gradient is right?

Go back to the definition: move ww up by a tiny amount ε\varepsilon, then down by the same amount, and see how much the loss changes.

∂L∂w≈L(w+ε)−L(w−ε)2ε\frac{\partial L}{\partial w} \approx \frac{L(w + \varepsilon) - L(w - \varepsilon)}{2\varepsilon}

This numeric gradient needs no calculus at all, just two forward passes. It is too slow to train with (two passes per weight), but it is a perfect checker. If it agrees with your backprop, your backprop is right. This is called a gradient check.

The lab writes tiny numbers the way computers do: 1e−5 means 1×10−5=0.000011 \times 10^{-5} = 0.00001, and 2e−11 means 2×10−112 \times 10^{-11}. This “e” has nothing to do with e≈2.72e \approx 2.72 from level 3.

Number Check on an easy function: L(w) = w². At w = 3 with ε = 0.1, what is (L(w + ε) − L(w − ε)) / (2ε)?
Predict firstIn the gradient check, slide ε down to 1e−12. Compared with ε = 1e−5, the difference from the chain-rule answer will be…
🔒 Answer the question above to unlock

3. From numbers to matrices

A real layer does the same thing for many neurons and many examples at once. Each scalar becomes a matrix. In the rest of this level, dz means ∂L/∂z\partial L/\partial z: the gradient of the loss at a layer’s output before its activation, one number per example and per neuron. In general it is not the error (prediction − truth) of level 2: at a hidden layer it is whatever the chain rule gives. Only at the output, with sigmoid and cross-entropy, does it come out as (p−y)/n(p - y)/n (the Deeper box below shows why). Three rules move it around:

  1. gradient of a layer’s weights: dW = X.T @ dz, where X is the layer’s input and X.T its transpose (level 1). “The layer’s input” means what goes into that layer: for the hidden layer it is the data X, but for the output layer it is the hidden layer’s output (the boss below calls it a1), not X.
  2. gradient of a layer’s bias: db = dz.sum(axis=0), one number per neuron. A bias is like a weight whose input is always 1, so its gradient is the plain sum of dz over the examples. (.sum(axis=0) adds down the rows, as in level 1, section 7.)
  3. one layer back: dz_prev = (dz @ W.T) * f'(z_prev). The @ sends the gradient back through the weights. The * is element-wise: it multiplies each number by the activation’s slope at that point.

Here are the same rules on a small network with 2 inputs, 3 hidden neurons and 1 output. It uses sigmoid and the squared loss, as in section 1. Hidden neuron 1 is section 1’s neuron: x1=2x_1 = 2, w=0.5w = 0.5, b=−1b = -1. The output also lands on z=0z = 0, so the backward pass starts with the same −1-1 and −0.25-0.25. Step forward, then keep pressing Next to go back.

Forward with numbers, then the gradient flows back along the same lines

Press Next to step. Select a layer to see its numbers.

xhaLW1W2(a − y)²
Press Next to run the forward pass with x = [2, 1]. · 13 parameters in all
Layers (4)

Look at the lines into h when you step back: each one turns into ∂L/∂w, and its width now shows the gradient, not the weight.

positive weight or valuenegative weight or value A disc’s fill is its value: the stronger the color, the further from 0. forward (this step) backward: gradient “∂ −0.5” = the gradient ∂L/∂(that value). Thicker line = larger |w|. Select a layer to see its numbers.

Where does X.T @ dz come from? Look at one example first. For one input row xx and one output gradient, the weight gradient is x⋅dzx \cdot dz, exactly as in section 1 (∂L/∂w=x⋅∂L/∂z\partial L/\partial w = x \cdot \partial L/\partial z). With several examples, the loss is a sum over them, so the gradients add up. Two examples, two inputs, one output:

exampleinput xdzx × dz (each input times dz)
0[1, 2]0.5[0.5, 1.0]
1[3, 1]−1?

dW is the sum of the last column over both examples. X.T @ dz computes exactly that sum in one multiplication.

Number Two examples: X = [[1, 2], [3, 1]] and dz = [[0.5], [−1]]. What is dW[0][0], the first entry of X.T @ dz?
Number Two examples have dz = [[0.5], [−1]]. The layer has one output neuron with one bias. What is db = dz.sum(axis=0)?

Our network for XOR is XX (4, 2) → W1W_1 (2, 6) → tanh → W2W_2 (6, 1) → sigmoid.

Shape X is (4, 2). The gradient at the hidden layer, dz1, is (4, 6). What is the shape of dW1 = X.T @ dz1?

Rule 3 takes the gradient one layer back. Try it on the XOR network, then on two numbers by hand.

Shape Rule 3 for the XOR network: dz2 is (4, 1), one number per example, and W2 is (6, 1). What is the shape of dz2 @ W2.T?
Number One example, a hidden layer of 2 tanh neurons, one output. dz2 = [[0.5]], W2 = [[2], [−1]] (so W2.T = [[2, −1]]), and the hidden outputs are a1 = [[0.5, 0]]. The slope of tanh at a neuron is 1 − a², where a is that neuron’s output. Use rule 3 from section 3. What is the first number of dz1?
I got stuck here Why is there a transpose in dW = Xᵀ @ dz?

Let the shapes decide. XX is (4, 2): 4 examples, 2 inputs. The gradient at the hidden layer, dz1, is (4, 6). The gradient must have the same shape as W1W_1, which is (2, 6). The only way to multiply a (4, 2) and a (4, 6) into a (2, 6) is XTX^T (2, 4) @ dz1 (4, 6).

The 4 that disappears is the examples: the transpose adds up each weight’s gradient over all 4 examples. When a shape doesn’t line up, write the shapes down and look for the one arrangement that works.

Go deeper Why the output gradient is simply p − y

The output uses sigmoid and the loss is the cross-entropy from level 3: L=−[yln⁡p+(1−y)ln⁡(1−p)]L = -[y \ln p + (1-y)\ln(1-p)]. Chain the two local gradients:

∂L∂p=p−yp(1−p)∂p∂z=p(1−p)\frac{\partial L}{\partial p} = \frac{p - y}{p(1-p)} \qquad \frac{\partial p}{\partial z} = p(1-p)

Multiply them and p(1−p)p(1-p) cancels: ∂L/∂z=p−y\partial L / \partial z = p - y. That is why the output gradient in the code below is just p - y (divided by 4, because the loss is the mean over 4 examples). It is also why sigmoid and cross-entropy are almost always used together.

🔒 Answer the question above to unlock

4. Boss: write the backward pass

The forward pass is written. You write the whole backward pass: start at the output with dz2 = (p − y) / n (the Deeper box above shows why), then use the three rules to get all four gradients. The local gradient of tanh is 1−tanh⁡(z)21 - \tanh(z)^2, and you already have tanh⁡(z)\tanh(z): it is a1. In Python the square is a1 ** 2 (** is a power, as in level 2; ^ is not). The hidden tests run a gradient check on every gradient your backward returns, with ε=10−5\varepsilon = 10^{-5} (written 1e-5).

CodeWrite the whole backward pass: every gradient the training step needs, from dz2 at the output down to dW1 and db1. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Your gradients match the numeric ones. Now train the network. Each step moves every weight against its gradient, exactly as in level 2. So that this cell doesn’t show the boss’s answer, it doesn’t call your backward: it gets the gradients the slow way, by changing each number a little (the check from section 2). Your gradient check showed that those are the same numbers your backward gives, only much slower to compute.

CodeWrite the update, then run 1000 steps of gradient descent on XOR.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

The lab below runs the same network, written in JavaScript instead of NumPy. Press Train and watch it learn.

A 2 → 6 → 1 network learns XOR

Press Train and watch the background bend around the four points. Then turn the activation off and try again.

0110x1 →x2 ↑
loss0steps →Train to draw the loss
loss while training (last 200 records)
inputtargetoutput
[0, 0]00.000
[0, 1]10.000
[1, 0]10.000
[1, 1]00.000
steps 0 · loss (cross-entropy) — · ✗ not learned yet
background: how sure the network is of 1 target 1 target 0 a wrong answer
Try it

In the backprop lab at the top, set the target to y = 0 and step back again. Which gradients change sign? Then set w = 5. Why does ∂L/∂w almost vanish? Look at a(1 − a) when a is close to 1.

Optional side trip: level U3 builds a small program that does this backward pass for you, the way PyTorch does, in about 50 lines of Python. Level 7 continues the main line.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Follow the chain rule backward through a small network by hand, multiplying the local gradients.
  • Check any gradient with the numeric estimate .
  • Write the backward pass of a two-layer network in NumPy and train it on XOR.

Keep in mind

  • Chain rule:
  • dW = X.T @ dz, the same shape as W
  • db = dz.sum(axis=0): one number per neuron, summed over the examples
  • dz_prev = (dz @ W.T) * f'(z_prev); for tanh,
  • Sigmoid with cross-entropy: the output dz = (p - y) / n

Common mistakes

  • Writing db = dz without summing over the examples, so its shape is wrong.
  • Using W @ dz or dz @ W where the shapes call for a transpose: write the shapes down first.
Side trips after this levelOptional; the next level does not need them.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars