Now the output layer, then the shapes.
Multilayer perceptron
What do you gain by stacking layers of neurons?
A short test to skip this level
Solve these 5 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
In level 3 one neuron drew one straight line. In level 4, XOR needed a second neuron and a bend between the layers. That small network has a name: a multilayer perceptron, or MLP. It is just layers of neurons, each layer feeding the next, with an activation function in between.
This level asks what you get from more neurons per layer (width) and more layers (depth). First, the forward pass written the way every library writes it: as matrices.
1. One hidden layer, as matrices
Put the inputs in a matrix with one row per example. A hidden layer of neurons is one matrix multiplication, a bias, and an activation. The output layer is one more:
As in level 4, is the matrix product X @ W1 in code.
Column of holds neuron ‘s two weights, so is (2, n). holds the hidden values: one row per example,
one column per hidden neuron.
Try one example by hand, with ReLU as the activation:
2. A harder pattern: two spirals
XOR had four points. Here are 200, in two arms that wind around each other. No straight line separates them, and neither do two or three. Before you press anything:
The lab below starts with 1 hidden layer of 8 tanh units. Press Train. The Map view shows the network’s answer as a background color, one color for each arm; switch to 3D to see the same answer as a height over the plane.
A network learns two spirals
Pick the size of the network, then press Train. The background shows the network’s answer at every point.
Now change the width. Set 2 units per layer and train again. Then try 32. With 2 units the network can only bend the plane twice, and it fails on most of the spiral. With enough units the background starts to follow the arms.
I got stuck here Why does it find a different answer every time I press New random start?
Every run starts from different random weights, and training only ever walks downhill from where it starts. So two runs can end in different valleys: different boundaries, sometimes one that fits and one that gets stuck. That randomness is normal. Real training fixes the random seed when it needs to repeat a run exactly.
I got stuck here How does the lab know which way to move all those weights?
It computes the gradient of the loss for every weight, then takes a small step against it, exactly the gradient descent from level 2, with a smarter step size you will meet in level 7. How it gets a gradient for a weight buried inside a hidden layer is the chain rule, applied layer by layer. That is level 6.
3. How big is it?
Every line between two neurons is one weight, and every neuron after the input has one bias.
Now add a layer.
The parameter count in the lab is now unlocked. Compare your answers with it.
Here is the same idea in 3D. It starts with the example from section 1 (, ReLU). Then pick a width and a depth and count the lines: each one is a weight.
Every line is one weight
Pick a width and a depth. Press Next to run one example through it.
Layers (3)
Look at the lines between two hidden layers: width × width of them. Width 6 has four times as many as width 3.
positive weight or valuenegative weight or value A disc’s fill is its value: the stronger the color, the further from 0. forward (this step) Thicker line = larger |w|. Select a layer to see its numbers.
4. What each hidden unit does
Check show each first-layer unit’s line. Every unit in the first layer is a neuron from level 3: it draws one straight line and reports which side a point is on. The layer after it combines those reports. Many lines, combined with bends in between, can trace almost any shape: here, a spiral.
Go deeper Width or depth?
A wide single hidden layer can, in principle, approximate any reasonable function: add enough units and their lines can split the plane into as many pieces as you like. But “enough” can be enormous.
Depth reuses work. A second layer does not see points, it sees the first layer’s answers, so it can combine pieces that the first layer has already found. That is why the same number of parameters often does more when it is spread over two or three layers than when it is all in one. Try it below.
Depth has a cost: the gradient has to travel back through every layer, and level 4 showed how it can get smaller at every layer. Levels 6 and 7 deal with that.
Compare 1 layer of 32 units with 3 layers of 8 units. Which has more parameters? Which one draws the spiral better after 1000 steps? Then switch to ReLU: the boundary is made of straight pieces instead of curves. Why?
5. Write it yourself
Write the hidden layer of a one-hidden-layer MLP with tanh. Everything is a matrix, so one line handles
a whole batch of examples. In NumPy, tanh is np.tanh: np.tanh(np.array([0., 1.])) gives [0., 0.76], one tanh per element.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Now add depth. A second hidden layer is the same line again, with as its input instead of .
Use ReLU this time: in NumPy it is np.maximum(0, …), as in level 4. This is the 2 → 2 → 2 → 1 version
of the network in section 1:
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
- Compute an MLP’s forward pass by hand for one example: hidden layer, activation, output.
- Give the shape of every layer’s output for a batch of examples.
- Count the weights and biases of a network from its layer sizes.
Keep in mind
- ,
(N, d_in) @ (d_in, n)→(N, n): one row per example, one column per hidden unit- Parameters per layer: inputs × outputs weights + one bias per output
- Each first-layer unit draws one line; later layers combine those lines
Common mistakes
- Counting only the lines (weights) and forgetting one bias per neuron.
- Adding the bias after the activation instead of before it:
act(X @ W1 + b1).
Press ? for keyboard shortcuts