Level 17 · Theory · runs in your browser

The whole Transformer

How does a model produce text one word at a time?

You have every part now: embeddings, attention, masks, the FFN, residuals and LayerNorm, and level 16 put them together into a decoder block. In this level you stack those blocks into a whole model and watch it work on a real task.

The task: read a number aloud. The input is a string of digits, the output is English words. The model sees both as one sequence: the digits, then =, then the words, then <eos> (“end”).

4 2 7 = four hundred twenty seven <eos>

The model only ever does one thing: guess the next token. Give it 4 2 7 = and it continues the sequence, one word at a time, until it writes <eos>. The part you give it is the prompt; the part it writes is the answer.

There are 1,000 numbers, from 0 to 999. The model trains on 900 of them. The other 100 are held out (the model never trains on them), so the only way to get them right is to learn the rule, not to remember the answers.

Number Level 17’s model reads 905 as one sequence: the digits, then =, then the words, then <eos>. One token per digit and one per word. How many tokens is that?
🔒 Answer the question above to unlock

1. The whole machine

In order, from input to output: a token table turns each id into a vector. Then the position numbers of level 16 are added. Then come a stack of decoder blocks, one final LayerNorm, and an output layer that gives one score per token. Select a block to see its input, its output, and how many parameters it holds. The shapes are for one sequence of 8 tokens, the longest input this task has.

The whole model, block by block

Pick a block to see its input, its output and how many parameters it holds. Change the sizes below.

One stack: reads “427 =”, writes the words
Decoder block × 2 (1, 8, 64) → (1, 8, 64)

Each block: LayerNorm, masked self-attention (4 heads of size 16; each position sees only itself and earlier ones), added back to x; then LayerNorm, the FFN (64 → 128 → 64), added back. One block has 33,280 parameters: attention 16,448, FFN 16,576, two LayerNorms 256.

parameters: 66,560 (92.3% of all)

total parameters: 72,106 · d_k per head = 16
fraction of all parameters not a layer (no parameters)
Shape A batch holds 64 numbers. The model input x is (64, 8): 8 tokens per number, padded. There are 42 tokens in the vocabulary and d_model = 64. What shape are the logits?
🔒 Answer the question above to unlock
Choosed_model stays 64 and you change heads from 4 to 8. The parameter count…

Most of the parameters are in the blocks, and inside each block most of them are in two places: the attention matrices and the FFN. You can count one FFN by hand.

Number d_model = 64 and d_ff = 128. How many parameters does one FFN have? It is Linear(64 → 128), ReLU, Linear(128 → 64). Count weights and biases.
Go deeper Where the 72,106 parameters are

At the default size (d_model 64, 4 heads, 2 blocks, d_ff 128):

partparameters
token table: 42 × 642,688
one attention block: W_Q, W_K, W_V (64 × 64 each) + W_O (64 × 64 + 64)16,448
one FFN16,576
one block: attention + FFN + 2 LayerNorms (128 each)33,280
final LayerNorm128
output layer: 64 × 42 + 422,730
total: 2 blocks + token table + final LayerNorm + output72,106

The position numbers are added, not learned, so they have no parameters. The number of heads does not appear anywhere in this table. It only decides how the 64 columns are cut into slices. The 42 tokens are <pad>, <eos>, the ten digits, =, and 29 words: zero to nineteen, the eight tens words, and “hundred”.

🔒 Answer the question above to unlock

2. Under the microscope

The diagram shows the blocks. Now look inside them. Below, a tiny copy of this model (1 block, 1 head, d_model = 6, d_ff = 12) reads 8 9 = and writes the next word. Every size is different from every other size, so each number shows you which axis it is. Each step shows the tensor’s shape with named axes, its numbers, and the shape of the same tensor in the full model on this page.

The whole model under a microscope

Step through one input, from token ids to the next token. Point at or tap any number to read where it sits.

Loading the recorded run…

Point at or tap a number to see its full index. ← → also move between steps.
positive negative −∞ blocked by the mask batch positions widths heads vocabulary

Here is the block of the same tiny model as a picture, with the numbers from the microscope. The line underneath is the residual path: each sublayer reads a normalized copy of x and adds its result to x.

Norm first, then add to the residual path

Press Next to follow “8 9 =” through one decoder block. Select a block for its numbers.

Drag to turn · click, then scroll to zoom
Press Next to start.
Layers (8)

Look at where the residual path starts: before LN1. Each LayerNorm gives its output only to its own sublayer, so x itself is never normalized until the final LayerNorm.

forward (this step) Select a layer to see its numbers.

I got stuck here Why is the embedding multiplied by a number before the position is added?

Look at steps 3 and 4 in the microscope. Every position number is between −1 and 1, so the factor decides how strong the token is compared with its position.

This model’s token table starts with numbers of size about 1/√d_model. Multiplying by √d_model makes them about size 1, the same as the position numbers, so that one doesn’t hide the other. In the full model, a token’s row has 64 numbers of size 1/8, so its length is about √64 × 1/8 = 1. After × √64 = 8 it is about 8 long, and a position’s row is about 5.7.

It matters. Suppose the table starts at size 1 instead. A token’s row is then about √64 × 1 × 8 = 64 long, 11 times longer than its position’s row, and the model can hardly see where a token is. It reads 82 as “eight hundred twenty two”. Over three training runs it gets 55% to 79% of the unseen numbers right, while the model on this page gets 99% to 100%. In level 21 the position table is learned too, and both tables start at the same size, so that model needs no factor.

With several heads, attention scores have the shape (B, h, L_q, L_k): batch, heads, then the length of Q and the length of K. In a decoder-only model Q and K come from the same sequence.

Shape Same batch: x is (64, 8), and the model has 4 heads. What shape are one block’s attention scores?
🔒 Answer the question above to unlock

3. One word at a time

Level 14 called it inference: running a trained model to get an answer. It is the whole forward pass, from token ids to logits, with every weight fixed. Training runs the same forward pass, then the backward pass, and then changes the weights.

ChooseA trained model reads a number aloud (inference). Which of these changes while it runs?

The model never writes the whole answer at once. It writes one word, then runs again:

  1. Start with the prompt: the digits and =.
  2. Run the model over everything so far. The output has one row per token.
  3. Only the last row matters. The output layer turns its 64 numbers into 42 scores, one per token. These raw scores are usually called logits.
  4. Pick the highest score (this is called greedy decoding), and add that token to the end.
  5. Repeat until the model picks <eos>.

Below is a real model, trained on this page’s task. Press Next step and watch each stage. The bars above the tokens show where the last row’s attention looked while it chose the next word.

Greedy decoding, one word at a time

Pick a number for the trained model to read aloud, then press Next step to watch it choose each word.

Loading the trained model’s recording…

…
positive number negative number the word it picks the right word, when it picks wrong
Number Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the model at this step?
🔒 Answer the question above to unlock

Step 4 is the only new code. NumPy’s np.argmax gives the position of the largest value, not the value itself: np.argmax(np.array([0.2, 1.5, -0.3])) is 1.

CodeWrite one step of greedy decoding. h is the last row of the model’s output (after the final LayerNorm), W and b are the output layer. Return the id of the next word.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Now the whole loop from the list above. The model returns one row of scores per token, so scores[-1] is the last row. ids.append(nxt) adds the picked id to the end, and break leaves the loop once the model picks <eos>. A fixed limit on new tokens stops a model that never picks <eos>.

CodeWrite the whole greedy loop. Start from the prompt (the digits and =). Each step runs the model on everything so far, takes the last row of the scores, appends the top id, and stops after the model picks eos (keep eos in the list). Stop after max_new new tokens even without eos.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

One epoch is one pass over all 900 training numbers. The lab has the same model saved twice: after 2 epochs and after 60.

Try it

Switch the model to after 2 epochs and read 14 aloud. It says “six hundred twenty nine”. It has learned the shape of a long answer: a digit, “hundred”, a tens word, a digit. It has not learned which digit goes where, or even how long the number is.

Then switch back to after 60 epochs and read 896. Watch the bars: for “eight” the last row looks at the 8, for “ninety” mostly at the 9, and for “six” at the 6. Nobody told it which digit goes with which word. This model gets all 100 unseen numbers right.

🔒 Answer the question above to unlock

4. How it learned

While training, the model does not run one word at a time. It gets the whole sequence and makes every guess in one pass. The input x is the sequence without its last token, and the target y is the sequence without its first token:

427:    x   4  2  7  =     four     hundred  twenty  seven
        y   2  7  =  four  hundred  twenty   seven   <eos>

The causal mask stops each position from seeing the token it has to guess. Training on the correct previous tokens like this, instead of on the model’s own guesses, is called teacher forcing.

I got stuck here Why are x and y the same tokens, shifted by one?

Because position t has to guess the next token. Write them one under the other: under = the answer is “four”; under “four” the answer is “hundred”. y is x moved one step to the left, with <eos> at the end.

If you fed the same row as both input and target, position t could copy its own input, and the loss would drop to zero while the model learns nothing.

Not every guess is worth learning. Under the 4 the target is 2, and under the 2 it is 7: the model would be guessing the digits of the question, and those are random. Only the targets that are words or <eos> are the answer. Every target before them, = included, is part of the question. The loss counts the answer targets and skips the rest, by setting the skipped targets to PAD.

Number For 427, x = 4 2 7 = four hundred twenty seven and y = 2 7 = four hundred twenty seven <eos>: 8 targets. The loss counts only targets that are part of the answer (the words and <eos>). How many of the 8 count?
🔒 Answer the question above to unlock

Training repeats four steps on batches of 64 numbers:

for x, y in batches:              # 1. a batch; skipped targets are PAD
    logits = model(x)             # 2. forward: (64, L, 42)
    # 3. loss, ignore_index=PAD
    loss = loss_fn(logits.reshape(-1, 42), y.reshape(-1))
    # 4. backward, then update
    opt.zero_grad(); loss.backward(); opt.step()

Short numbers have short sequences. In a batch with the numbers 0 to 63, the longest x is 5 tokens (for example 2 3 = twenty three), so the logits are (64, 5, 42) and there are 320 targets. Only 167 of them count: 118 are digits or =, and 35 are PAD after a short sequence.

Shape For the loss, the (64, 5, 42) logits are flattened so that every position is one row. What shape do they become?
I got stuck here Why reshape to (B × L, V) at all?

The loss function wants a simple list of guesses: one row of scores per guess, and one correct id per guess. It does not care which sequence or which position a guess came from. Every position in every sequence is one independent guess, so (64, 5, 42) is stored as 320 rows of 42.

ChooseYou delete the causal mask and train again. What happens?
Go deeper A debugging habit: the copy task

When a model learns nothing, the problem can be the task, or it can be a bug in the code. What would you check first?

ChooseYou build a new model for a new task. After 10 epochs it gets 0% right, and the training loss stopped falling long ago. What do you try first?

The copy task is 1 2 3 = 1 2 3. Any correct model learns it in about a minute of training. So if yours can’t, the bug is in the mask, the shift by one or the loss.

Last step. Write the loss that skips the PAD targets. In NumPy, an array of True and False picks entries: np.array([5, 6, 7])[np.array([True, False, True])] is array([5, 7]). The browser gives you softmax(x), and it works on each row of x separately.

CodeWrite the whole loss that skips PAD, starting from the scores. logits is (N, V), targets is (N,), and the argument pad is the id of PAD (0 by default). Return the mean loss over the rows whose target is not PAD. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Recommended next: side trip N7. If you want to read the original 2017 Transformer design, now is a good time. N7 builds the two-stack encoder–decoder on this same task. Level 18 continues the main line.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Follow the shapes through a whole decoder-only model, from token ids to (B, L, V) logits.
  • Write the greedy loop: run the model, take the last row, append the top id, stop at <eos>.
  • Build x and y by a shift of one, and compute a loss that counts only the answer.

Keep in mind

  • One sequence: 4 2 7 = four hundred twenty seven <eos>; the prompt is the start of the sequence
  • logits: (B, L, d_model) → (B, L, V); for the loss, reshape(-1, V)
  • x = seq[:-1], y = seq[1:]: position t guesses token t + 1
  • Greedy step: np.argmax(scores[-1]), the last row only

Common mistakes

  • Averaging the loss over the skipped rows too: (nll * keep).mean() still divides by every row.
  • Counting only the words when you count rows: the digits and = are in the same sequence.
Side trips after this levelOptional; the next level does not need them.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars