Level 2 · Foundations · runs in your browser

Gradient descent

How does a machine learn from its mistakes?

A model is a formula with adjustable numbers called parameters (or weights). Learning means changing the parameters until the formula’s answers match the data. This level uses the smallest model there is, a straight line with two parameters:

y^=w⋅x+b\hat y = w \cdot x + b

y^\hat y (read “y hat”) is the line’s prediction for an input xx. Here “·” between two single numbers is ordinary multiplication. (In level 1 it joined two lists; with single numbers it is just ×.)

The data is three points: (1, 2), (2, 4), (3, 6). You can see the answer is w = 2, b = 0. The machine can’t. It starts from a guess, w = 0.5 and b = 0, and has to reach the answer step by step.

1. Measure the mistake

For each point, the error is prediction minus truth. Square each error so that misses in either direction count as bad, then take the mean. That one number is the loss:

L=mean((wx+b−y)2)L = \text{mean}\big((w x + b - y)^2\big)

Gradient descent on three points

Move w and b by hand, or take steps. The line, the errors and the loss surface move together.

02468123x →y
The data, the line y = w·x + b, and each error as a dashed line.
Drag to turn · click, then scroll to zoom

Your first click or tap selects the view. After that, a click or tap on the surface moves the point there.

Height is ln(1 + loss) for every (w, b): low is good, and the big losses are made much smaller so the bottom stays visible.
xyw·x + berrorerror²2·error·x2·error
120.5-1.5???
241-3???
361.5-4.5???
meanloss = ?grad_w = ?grad_b = ?
The step buttons unlock once you compute one step by hand below.
you: (w, b) now the lowest loss, w = 2, b = 0 the path of steps an error (prediction − y)
Number Start at w = 0.5, b = 0. The predictions are 0.5, 1, 1.5 and the targets are 2, 4, 6. What is the mean squared error loss?
🔒 Answer the question above to unlock

2. Which way is downhill?

The loss is a surface over (w, b): the right panel of the lab draws it. In 3D the height is not the loss itself but a smaller number made from it: more loss is still higher, but the big values are made much smaller, so the low ground stays visible. (For the curious: the height is ln⁡(1+loss)\ln(1 + \text{loss}), and ln is on the math page.) To lower the loss, you need to know: if you make w a little bigger, does the loss go up or down, and how fast? That rate is the gradient. In plain words: how much L changes for each tiny step of w. It is written ∂L/∂w\partial L/\partial w and read “the partial derivative of L with respect to w”. “Partial” means only w moves and everything else stays fixed. ∂L/∂b\partial L/\partial b is the same with only b moving. For a function with only one input, the slope has a short name: f′(x)f'(x), read “f prime of x”. It is the slope of ff at xx. (New to derivatives? The math page shows one with numbers.) For this loss the two gradients are:

∂L∂w=mean(2⋅error⋅x)∂L∂b=mean(2⋅error)\frac{\partial L}{\partial w} = \text{mean}(2 \cdot \text{error} \cdot x) \qquad \frac{\partial L}{\partial b} = \text{mean}(2 \cdot \text{error})

Compute it on paper. Once you answer, the table under the lab shows every term so you can check.

Number A line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6) with the loss L = mean((w·x + b − y)²). At w = 0.5, b = 0, what is ∂L/∂w?
🔒 Answer the question above to unlock

3. Take a step

Move each parameter a small amount against its gradient. The size of “a small amount” is the learning rate:

w←w−lr⋅∂L∂wb←b−lr⋅∂L∂bw \leftarrow w - \text{lr} \cdot \frac{\partial L}{\partial w} \qquad b \leftarrow b - \text{lr} \cdot \frac{\partial L}{\partial b}

Compute the new w by hand with the learning rate at 0.05. Then the step buttons in the lab unlock, and “Take one step” checks you.

Number w = 0.5 and its gradient is −14. With learning rate 0.05, what is w after one step?
🔒 Answer the question above to unlock
I got stuck here The gradient of w is −14. Why is it negative, and what does that tell me?

Negative means: if w goes up, the loss goes down. Every prediction is too low right now (0.5, 1, 1.5 against 2, 4, 6), so a steeper line helps.

The size, 14, says how sensitive the loss is to w right now. It is large because we are far from the answer. Near the bottom of the valley the gradient shrinks toward 0, and the steps get smaller by themselves, with no change to the learning rate.

The loss dropped from 10.5 to about 2.117 in one step. Press “10 steps” a few times and watch the path move down the surface toward the lowest point, w = 2, b = 0.

ChooseA line ŷ = w·x + b is fit to the points (1, 2), (2, 4), (3, 6). At w = 0.5, b = 0 the loss is 10.5, ∂L/∂w = −14 and ∂L/∂b = −6. The best line has w = 2, b = 0. You take one step on both w and b with learning rate 0.2 instead of 0.05. What happens to the loss?
Go deeper Where the gradient formula comes from

Take one point. Its loss is (wx+b−y)2(wx + b - y)^2. Call the inside e=wx+b−ye = wx + b - y, the error.

The loss is e2e^2, and e2e^2 changes at rate 2e2e when ee changes. ee changes at rate xx when ww changes (because ww is multiplied by xx), and at rate 11 when bb changes.

Multiply the two rates: ∂L/∂w=2e⋅x\partial L/\partial w = 2e \cdot x and ∂L/∂b=2e⋅1\partial L/\partial b = 2e \cdot 1. Average over the points because the loss is an average. That is the whole formula.

“Multiply the rates along the way” is called the chain rule. Level 6 is that idea applied to a whole network.

4. Write it yourself

Write the gradient of w. x and y are NumPy arrays, so np.mean averages over the points.

One piece of Python you will need in every later level: a power is written **. So the loss is np.mean(error ** 2), where error ** 2 squares every element. Don’t write ^ for a power: in Python it means something else, and on decimal numbers it gives an error. (More on the math page.)

CodeWrite the gradient of w.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

5. Why against the gradient?

I got stuck here Why do we subtract the gradient? Wouldn’t adding it also change the parameters?

The gradient points in the direction in which the loss grows fastest. We want the loss to shrink, so we go the opposite way. Adding it would climb the surface, and the loss would grow every step.

Try it

Move to a far corner of the loss surface, for example w = −1, b = 3 (select a point on the 3D surface, or drag in the Map view), and take steps from there. Then set the learning rate to 0.15 and do it again. Which path goes back and forth across the valley, and why?

6. The whole loop

Put the three steps together: predict, compute both gradients, step both parameters. Repeat. This is the loop every model in this course trains with; only the formula for the gradients gets bigger.

CodeWrite the body of the training loop: one gradient-descent step on w and b. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

One last check: the same procedure on different data.

Number The points are now (1, 3), (2, 5), (3, 7). Start at w = 0.5, b = 0 with learning rate 0.05 and take one step of gradient descent (by hand, or with descend from the code box). What is b after that one step?

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute a mean squared error loss by hand for a few points.
  • Compute and for a line and take one step against them.
  • Write the whole training loop: predict, compute the gradients, step, repeat.

Keep in mind

  • ,
  • : step against the gradient
  • Too large a learning rate jumps past the minimum and the loss grows

Common mistakes

  • Adding the gradient instead of subtracting it: minus a negative gradient is plus.
  • Forgetting the x in the gradient of w, or summing where the loss takes the mean.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars