Level D1 · Foundations · Diffusion · runs in your browser

Adding and removing noise

If you mix a picture with noise step by step, can you undo it?

Side trip · best after level 10 · Probability and sampling

A Transformer writes one word at a time. Diffusion models make pictures in a different way. They start from pure noise and remove a little of it, many times, until a picture is left.

To learn how to remove noise, the model first watches the opposite: clean data being mixed with noise. That direction is easy. It takes one line of math.

What you need: Foundations (levels 1–10). D1 and D2 use nothing else. D3 also uses attention and Transformer blocks, so take D3 after level 16.

Our “picture” is a cloud of 2D points shaped like a spiral. Every point is two numbers, so you can see everything.

1. Adding noise

Each step mixes a bit of random noise into every point. After t steps, a point x0 has become:

xt=αˉt  x0+1−αˉt  εx_t = \sqrt{\bar\alpha_t}\; x_0 + \sqrt{1-\bar\alpha_t}\; \varepsilon

As ᾱ_t falls toward 0, the signal’s weight √ᾱ_t falls toward 0 and the noise’s weight √(1 − ᾱ_t) rises toward 1. The lab below uses 50 steps (call that number T = 50). Before you drag anything, ask yourself: what does the last step look like?

ChooseNoise is mixed into a spiral of 2D points with xₜ = √ᾱₜ x0 + √(1 − ᾱₜ) eps. On a 50-step schedule, alpha-bar at t = 50 is 0.0045. What does the spiral look like there?

Now drag t and watch the spiral turn into noise. The highlighted point is the one we follow.

The forward process: a spiral turning into noise

Drag t. Every point is mixed with its own fixed noise, so it moves smoothly.

ᾱ (ab)1.0000
x0[0.293, 0.146]
eps[-0.583, 0.641]
x_t = 1.000 × x0 + 0.000 × eps
x_t[0.293, 0.146]
t = 0: nothing added yet, every point is still on the spiral.
the cloud at step t where the point started the point we follow, and its path

Where does alpha-bar come from? Each step has a small number β (beta), the amount of noise it adds. That step keeps α = 1 − β of the signal. After several steps, multiply what each step kept.

Try it on a small example with only 3 steps: β = 0.1, 0.2, 0.5. So α = 0.9, 0.8, 0.5.

Number With β = 0.1, 0.2, 0.5, what is alpha-bar after 3 steps?
🔒 Answer the question above to unlock

Why the square roots?

Alpha-bar is 0.36, yet the formula multiplies the signal by √0.36 = 0.6. Both are right, because they measure different things. In level 10 you measured how spread out numbers are with the std. Squaring the std gives the variance, and variance is what alpha-bar counts: 36% of the signal’s variance is left. Multiplying a number by 0.6 multiplies its variance by 0.6² = 0.36.

The noise gets the rest. Its weight is √(1 − 0.36) = √0.64 = 0.8, so its variance is 0.8² = 0.64.

Number The signal is multiplied by 0.6 and the noise by 0.8. What is 0.6² + 0.8²?

That sum is the reason for the square roots. The two weights, squared, always add up to 1, at every step. So if the clean data has variance 1, the noisy data has variance 1 at every step too: the cloud only changes from “spiral” into “noise”, and never grows or shrinks. Real models scale their data to variance 1 for this reason. Our spiral has a variance of about 0.6, so its cloud grows a little as it becomes noise with variance 1.

Now compute it for one point. Take x0 = [1, 2] and the noise eps = [0.5, -1]. At t = 3, √0.36 = 0.6 and √(1 − 0.36) = 0.8. After you answer, switch the lab to Hand example · 3 steps and drag to t = 3 to watch the point move.

Number x0 = [1, 2], eps = [0.5, −1], alpha-bar = 0.36. What is the first number of x₃?
I got stuck here Why jump straight to step t? Don’t you have to add the noise one step at a time?

You can add it one step at a time and get the same kind of result. Adding two independent normal noises gives one bigger normal noise, so all t steps combine into one formula. That shortcut matters for training: to make a training example at step 37, you don’t have to run 37 steps. Compute alpha-bar once and jump there.

Write the shortcut yourself. It works on one point or on a whole array of points at once.

CodeWrite add_noise. ab is alpha-bar. It should work for one point or for an array of points.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Now put both parts together: start from the list of betas and return the noisy point at step t. Three pieces of NumPy do it:

  • betas[:t] takes the first t betas (positions 0 to t − 1), so step t uses steps 1 to t.
  • 1 - betas[:t] turns them into the α values, what each step keeps.
  • np.prod multiplies all the numbers of an array: np.prod(np.array([0.9, 0.8, 0.5])) is 0.36.
CodeWrite noisy_at: from the clean point x0, the noise eps, the list of betas and a step t (counting from 1), compute alpha-bar for step t, then return the noisy point xₜ.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Adding noise is the easy half. Next comes the half that makes generation possible: removing it.

🔒 Answer the question above to unlock

2. Running it backward

Here is the key idea. If you knew exactly which noise eps was added, you could undo the formula and get x0 back:

x0=xt−1−αˉt  εαˉtx_0 = \frac{x_t - \sqrt{1-\bar\alpha_t}\;\varepsilon}{\sqrt{\bar\alpha_t}}

Try it on a different point. At t = 3 (so 0.6 and 0.8 again), a point sits at x_3 = [0.7, 1.0], and you are told its noise was eps = [0.5, 0.5].

Number At t = 3 (alpha-bar 0.36), x₃ = [0.7, 1.0] and the true noise was eps = [0.5, 0.5]. What is the second number of x0?

When generating, nobody tells you the noise. So a network is trained to guess it. The network sees a noisy point and the step number t, and outputs its guess of the noise. The loss is the mean squared error between the guess and the true noise.

Number The simplest network always guesses that the noise is [0, 0]. For one point the true noise is [0.5, −1]. What is the loss, mean((guess − eps)²)?
Shape One training batch has 128 noisy points. Each input row is the point’s x and y, plus 12 numbers that encode the step t. What shape goes into the network?
Go deeper Why the network needs to know t

Take the same noisy position at two different steps:

  • At step 2 the point has hardly moved, so almost all of it is signal. The right noise guess is small.
  • At step 48 the point is almost pure noise. The right guess is close to the point itself.

The same position needs a different answer at different steps.

So the network gets t as 12 extra numbers: sin and cos of t/T (T = 50, the number of steps) at 6 different frequencies (how fast each one repeats). A single number would work too, but those 12 make it easy for the network to treat nearby steps alike and far-apart steps differently. Transformers use the same idea for word positions (level 16).

3. Training the noise-guessing network

Next you train this network in your browser. Each training step:

  1. takes 128 points from the spiral,
  2. picks a random t for each and adds the matching noise,
  3. asks the network to guess that noise and changes its weights a little to reduce the loss.

The simplest guess, always [0, 0], scored 0.625 on our one point. Averaged over many random noises it scores 1.0, because each noise number has variance 1, so the average of ε² is 1. Training should do better than that. How much better?

Predict firstGuess before you train: if training goes perfectly, where does the loss stop?

Now press Train. This is real gradient descent with Adam, the same loop you met in levels 2 and 7.

Train a noise-guessing network in your browser

Press Train. Then drag t to see how well it guesses at each amount of noise.

step 0 of 4000 · loss (average of the last 50 steps) —
noisy points at step t the network’s guess of where each started: (x_t − √(1 − ᾱ)·guess) / √ᾱ always guessing 0
Try it

Train all 4,000 steps, then drag “look at step t”. At t = 5 the guessed points sit right on the spiral. At t = 45 they no longer follow the spiral; they land in an unclear cloud. Why can’t the network do better from so much noise? Is it a bad network, or is the question impossible?

One piece is left: turning a noise guess into a guess of the clean point. It is the undo formula from section 2, with the network’s guess in place of the true noise. The guess is written ε̂ (“eps-hat”, eps_hat in code): a hat on a letter marks an estimate of it.

CodeWrite guess_clean: given a noisy point, the network’s noise guess, and alpha-bar, return the guess of x0.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

You now have both halves. Forward: one line that adds noise. Backward: a network that guesses the noise. One guess from heavy noise gives an answer that is not sharp, as you saw. In level D2 you take many small steps instead of one big one, and the spiral comes back sharp.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute alpha-bar from a list of betas: multiply what each step keeps, α = 1 − β.
  • Compute the noisy point at any step t in one line, without running the steps in between.
  • Undo the formula with a noise guess to get a guess of the clean point.

Keep in mind

  • with ; in code np.prod(1 - betas[:t])
  • The two weights squared add up to 1, so the variance stays the same:
  • The network sees the noisy point and t and guesses the noise; its loss is mean((eps_hat - eps)**2)

Common mistakes

  • Using alpha-bar itself as the weight: the signal gets (0.8), not (0.64).
  • Multiplying the betas (the noise added) instead of the alphas (the signal kept).

Press ? for keyboard shortcuts

Reading mode · every part open, no stars