Solve these 6 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Shape Batch 2. The decoder input has 3 tokens (<bos> plus 2 target words), and the source has 6 words. What shape are one head’s cross-attention weights?
Number Cross-attention by hand: one query q = [√2 · ln 3, 0], dₖ = 2, keys k₀ = [1, 0] and k₁ = [0, 1] from memory, values v₀ = [4, 0] and v₁ = [0, 8]. What is the output’s first number?
Shaped_model = 64, a batch of 32 numbers, source digits padded to 3 tokens, a decoder input of 5 tokens. What shape is memory, the encoder’s output?
Shape The source has 5 tokens. The decoder input is <bos> plus 5 target words. What shape is the causal mask in the decoder’s self-attention (for one sentence)?
Number Reading 42 aloud, in a batch where longer answers need 5 words. The decoder input is <bos> forty two PAD PAD: 5 positions. The decoder’s self-attention mask blocks a cell if the causal mask blocks it or if its column is a PAD word. How many of the 25 cells are blocked?
CodeBuild the three masks of one example. src holds the digit ids, tgt_in the decoder input ids, and pad is the id of PAD. Return (enc, dec, cross), True where a cell is blocked: enc is (Ls, Ls), dec is (Lt, Lt), cross is (Lt, Ls). Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
This level prepares you to read the original 2017 Transformer design. Its diagram has two stacks, and this level builds both.
Level 17’s model is one stack: the digits, then =, then the words, all in one sequence, with masked self-attention
everywhere. The first Transformer, from 2017, was built differently. It was made for translation: read a
source sentence in one language, write a target sentence in another. So it has two stacks, an encoder that reads
and a decoder that writes. (On this page, “target” means the sentence to write. It is not the training target y.)
This level builds that design and runs it on level 17’s task, reading a number aloud, with the same 900 training numbers
and the same 100 held-out ones:
"427" → four hundred twenty seven
If you did side trip N5, you solved this kind of task there with two LSTMs and attention between them. This model
keeps that plan (an encoder, a decoder, and attention from one to the other) and replaces the LSTMs with attention
layers. N5 is optional; you don’t need it for this level.
demo.py prints every number on this page. To run it yourself, download it into a
folder of its own and run python demo.py there (setup). It needs PyTorch and takes about 20 seconds.
1. Two stacks
The two stacks split the work:
the encoder reads the whole source at once: every source token may look at every other source token;
the decoder writes the target one word at a time. It looks at the words written so far, and at the source.
An encoder layer is self-attention, then the FFN. A decoder layer has three sublayers: masked self-attention over
the words written so far, cross-attention to read the source, and the FFN.
Cross-attention is ordinary attention, with one difference: its Q comes from one sequence and its K and V from another.
The encoder’s final output is called memory: one vector per source token. It is what the decoder reads.
ChooseIn the decoder’s cross-attention, where do K and V come from?
🔒 Answer the question above to unlock
The 2017 design also puts the LayerNorm in a different place from level 17. Each sublayer is wrapped as
LayerNorm(x + sublayer(x)), “add & norm”: add the residual first, then normalize. That is the wrapping of level 16,
section 3. It is called post-norm. Level 17 normalizes what goes into each sublayer instead (pre-norm), and
level 19 shows why that trains deep stacks better.
One encoder layer and one decoder layer
Select a block to see what it does and the shapes it works on. The sliders change the sizes.
encoder layer
↓
↓
↓
↓
memory (1, 5, 8) → to the decoder’s cross-attention
decoder layer
↓
↓
↓
↓
↓
↓
cross-attention
Every target word looks at the source sentence. This is how the decoder reads the input.
input
(B, L, d_model) = (1, 4, 8) and memory (1, 5, 8)
output
(B, L, d_model) = (1, 4, 8)
Q, K, V
Q comes from the decoder. K and V come from memory, the encoder’s output.
shape
weights per head (B, L_tgt, L_src) = (1, 4, 5): one row per target word, one column per source word
parameters
256
Remove it: Without it, the decoder never sees the source: it can only write a sentence, not translate one.
Inside “cross-attention”, tensor by tensor
tensorshapewhat happens
1Q per head(B, heads, L_tgt, d_k) = (1, 2, 4, 4)from the decoder, d_model split into heads
2K, V from memory(B, heads, L_src, d_k) = (1, 2, 5, 4)from the encoder’s output
3weights(B, heads, L_tgt, L_src) = (1, 2, 4, 5)one row per query word, one column per key word
4out(B, L_tgt, d_model) = (1, 4, 8)heads side by side again, then W_O
It is the output of the last encoder layer: one vector per source token, shape (B, L_src, d_model).
It is computed once per source. Every decoder layer reads the same memory through its cross-attention,
using it for K and V. Nothing is stored or trained in it. It is a name for “the encoder’s answer”.
It is not the hidden state of an RNN (level N3). On this page, “memory” alone means the encoder’s output.
The “GPU memory” of U6 and level 20 is something else.
Q decides the rows of the attention weights and K the columns, as in level 14. In cross-attention, Q comes from the
decoder input and K from memory.
The decoder input always begins with a start token, written <bos> (“beginning of sequence”). It gives the
decoder a first input before it has written any word. Section 4 shows how the decoder uses it.
Shape Batch 2. The decoder input has 3 tokens (<bos> plus 2 target words), and the source has 6 words. What shape are one head’s cross-attention weights?
🔒 Answer the question above to unlock
Now compute one cross-attention by hand. The steps are those of level 14: one score per key, softmax, then a weighted
average of the values. Only the place each table comes from is new.
Number One decoder token reads a memory of 2 source tokens, with dₖ = 2. Its query is q = [√2 · ln 3, 0], about [1.55, 0]. From memory, the keys are k₀ = [1, 0] and k₁ = [0, 1], and the values are v₀ = [4, 0] and v₁ = [0, 8]. The output is a vector of 2 numbers. What is its first number?
🔒 Answer the question above to unlock
2. The whole machine
The encoder reads the digits. The decoder writes the words. While it writes, it reads the encoder’s output (memory)
through cross-attention. Select a block to see its input, its output, and how many parameters it holds. The shapes are for one
number (“427”) and a decoder input of 5 tokens: the start token <bos> and up to 4 words.
The whole model, block by block
Pick a block to see its input, its output and how many parameters it holds. Change the sizes below.
Encoder: reads “427”
Decoder: writes the words
Decoder layer × 2(1, 5, 64) → (1, 5, 64)
Each layer has three parts: masked self-attention (each word sees only earlier words), cross-attention (Q from the words, K and V from memory), and the FFN. One layer has 49,856 parameters.
parameters: 99,712 (58.2% of all)
total parameters: 171,232 · d_k per head = 16
fraction of all parametersnot a layer (no parameters)
Shape A model reads numbers aloud, with d_model = 64. A batch holds 32 numbers. The source digits are padded to 3 tokens, and the decoder input has 5 tokens. What shape is memory, the encoder’s output?
🔒 Answer the question above to unlock
ChooseAt the default size, a decoder layer has 49,856 parameters and an encoder layer 33,280. The difference is 16,576, the same as one FFN. What does the decoder layer have that the encoder layer doesn’t?
🔒 Answer the question above to unlock
Go deeper Where the 171,232 parameters are
At the default size (d_model 64, 4 heads, 2 layers, d_ff 128):
part
parameters
one attention block: WQ, WK, WV (64 × 64 each) + WO (64 × 64 + 64)
The position numbers are added, not learned, so they have no parameters.
The number of heads does not appear anywhere in this table. It only decides how the 64 columns are cut into slices,
so 8 heads instead of 4 still give 171,232.
The encoder and the decoder have 2 layers each here, but the two numbers are independent and can differ.
There are two embedding tables: one for the 13 source tokens (digits and 3 special tokens) and one for the 32 target words.
🔒 Answer the question above to unlock
3. Under the microscope
Below, a tiny copy of this model (1 layer, 1 head, d_model = 6, d_ff = 12) reads “89” and writes the next word.
All the sizes are different, so each axis can be recognized by its size. Each step shows the tensor’s
shape with named axes, its numbers, and the shape of the same tensor in the full model on this page.
The whole model under a microscope
Step through one input, from token ids to the next token. Point at or tap any number to read where it sits.
Loading the recorded run…
Point at or tap a number to see its full index. ← → also move between steps.
positivenegative−∞ blocked by the maskbatchpositionswidthsheadsvocabulary
Here is the encoder layer of the same tiny model as a picture, with the numbers from the microscope. The line
under the blocks is the residual path. In this 2017 layer, a LayerNorm follows each add.
In the 2017 layer, the residual path is normalized after every add
Press Next to follow “89” through one encoder layer. Select a block for its numbers.
Drag to turn · click, then scroll to zoom
Press Next to start.
Layers (7)
Look at the residual path under the blocks: after each add, a LayerNorm rescales the sum. This is post-norm. Compare it with level 17’s block, where x is never normalized until the end.
forward (this step)Select a layer to see its numbers.
I got stuck here Why is the embedding multiplied by a factor before the position is added?
Look at steps 3 and 4 in the microscope. Every position number is between −1 and 1. The factor decides how strong
the token is compared with its position.
This model’s token tables start with numbers of size about 1/√dmodel, as in level 17. Multiplying by
√dmodel makes them about size 1, the same as the position numbers. So neither one is much bigger than the
other. With d_model = 64, the factor is √64 = 8. As in level 17, a token’s row is then about 8 long, and a position’s
row about 5.7.
The 2017 design is usually explained the same way. There, the output layer shares its weights with the embedding table.
That is one reason for a table with small numbers.
It matters. Suppose the table starts at size 1 instead. A token’s row is then about 64 long, 11 times longer than its
position’s row, as in level 17. Here is the fraction of unseen numbers each model gets right:
level 17’s one-stack model, after 60 epochs: 55% to 79% with the size-1 table, 99% to 100% with the small table;
this two-stack model, after 20 epochs: 71% with the size-1 table, 98% with the small table;
this two-stack model, after 60 epochs: 99% with the size-1 table, 100% with the small table.
So with the size-1 table, this model needs more epochs to reach the same result.
With several heads, attention scores have the shape (B, h, L_q, L_k): batch, heads, then the length of Q and the length of K.
Shape A model reads numbers aloud. A batch holds 64 numbers, the source digits are padded to 3 tokens, the decoder input has 5 tokens (words), and there are 4 heads. What shape are the cross-attention scores?
🔒 Answer the question above to unlock
4. One word at a time, with a source
The decoder writes the way level 17’s model does: run, take the last row, pick the top word, append it, run again.
One thing is new. Level 17’s model starts from the digits and =, which are already in its sequence. This decoder’s
input holds only words, so it starts from the start token <bos>. The digits reach it only through
cross-attention, and the encoder reads them once, before the first word.
Press Next step and watch each stage. The bars under the digits show which digit cross-attention looked at for each word.
Greedy decoding, one word at a time
Pick a number for the trained model to read aloud, then press Next step to watch it choose each word.
Loading the trained model’s recording…
…
positive numbernegative numberthe word it picksthe right word, when it picks wrong
Number Reading 742 aloud. The model has already written “seven hundred forty”. How many rows go into the decoder at this step?
🔒 Answer the question above to unlock
Number A decoder with 2 layers reads 742 aloud: 4 words, one per step. The model keeps (caches) every result that cannot change between steps, and reuses it. In one decoder layer, how many times is memory multiplied by W_K for the whole answer?
🔒 Answer the question above to unlock
This decoder uses two caches. The cross-attention cache holds memory’s keys and values: one row per source token,
made once. The self-attention cache gets one new row at every step. Level 20 counts what such
a cache costs in memory and time.
Try it
Switch the model to after 2 epochs and read 896 aloud. It says “six hundred eighty”.
It has learned the shape of an answer (a digit, “hundred”, a tens word) but not which digit goes where.
Look at the cross-attention bars. They are spread out, almost even.
Then switch back to after 60 epochs and read 896 again. Now each word looks at its own digit: “eight” puts all its
weight on the 8, and “six” puts 0.92 on the 6. The trained model gets all 100 unseen numbers right.
🔒 Answer the question above to unlock
5. Three attentions, three masks
In training, the decoder gets the whole correct answer at once, shifted by one:
427: tgt_in <bos> four hundred twenty seven tgt_out four hundred twenty seven <eos>
I got stuck here Why are the decoder’s input and its targets the same words, shifted by one?
Because position t has to guess the next word. Write them one under the other: under <bos> the answer is “four”;
under “four” the answer is “hundred”. tgt_out is tgt_in moved one step to the left,
with <eos> at the end. Level 17 does the same shift on its one sequence; here the shift only covers the words.
The model uses three attentions, and each one has its own mask. A mask always has the shape of the scores it blocks,
(length of Q, length of K). L_src is the source length and L_tgt the decoder input length, <bos> included.
Every head uses the same mask:
attention
Q from
K and V from
mask
encoder self-attention
the digits
the digits
blocks PAD digits only
decoder self-attention
tgt_in
tgt_in
causal, and blocks PAD words
cross-attention
tgt_in
memory
blocks PAD digits only
Their scores have these shapes:
encoder self-attention: (B, h, L_src, L_src);
decoder self-attention: (B, h, L_tgt, L_tgt);
cross-attention: (B, h, L_tgt, L_src).
Here are the three masks for one short example, side by side. Rows are the tokens that Q comes from; columns are the
tokens that K comes from.
encoder self-attention
(3, 3), 3 blocked
1
5
PAD
1
0
0
1
5
0
0
1
PAD
0
0
1
rows (Q): the digits columns (K): the digits
decoder self-attention
(4, 4), 9 blocked: 6 future, 3 PAD
<bos>
fifteen
PAD
PAD
<bos>
0
1
1
1
fifteen
0
0
1
1
PAD
0
0
1
1
PAD
0
0
1
1
rows (Q): the decoder input columns (K): the decoder input
cross-attention
(4, 3), 4 blocked
1
5
PAD
<bos>
0
0
1
fifteen
0
0
1
PAD
0
0
1
PAD
0
0
1
rows (Q): the decoder input columns (K): memory (the digits)
blocked: futureblocked: PAD1 means True: blocked. A blocked score becomes −∞, so its weight is 0.
Reading 15 aloud: the source “15” is padded to 3 digits, and the decoder input is<bos> fifteen PAD PAD. A cell that is both future and PAD is shown as future.
Shape The source has 5 tokens. The decoder input is <bos> plus 5 target words. What shape is the causal mask in the decoder’s self-attention (for one sentence)?
🔒 Answer the question above to unlock
I got stuck here Which length decides the size of the causal mask?
The causal mask is used only in the decoder’s self-attention, where both Q and K come from the decoder input:
<bos> plus the 5 target words, 6 tokens. The source length never enters a causal mask.
I got stuck here Why don’t the encoder and cross-attention use a causal mask?
The causal mask stops a position from seeing the word it has to guess. Only the decoder guesses words, and only the
decoder’s own input contains them. The digits are the question, not the answer: every word may read the whole number,
and every digit may read every other digit.
Number Reading 42 aloud, in a batch where longer answers need 5 words. The decoder input is <bos> forty two PAD PAD: 5 positions. The decoder’s self-attention mask blocks a cell if the causal mask blocks it or if its column is a PAD word. How many of the 25 cells are blocked?
🔒 Answer the question above to unlock
Now build all three masks for one example. True means blocked, as in level 15. Level 15’s padding mask had one entry per
column; to give it the shape of the scores, stretch that row over every query row. np.broadcast_to(row, (n, len(row)))
does it: np.broadcast_to(np.array([False, True]), (3, 2)) is a (3, 2) table whose three rows are all [False, True].
| combines two True/False tables cell by cell: a cell is True if either one is True. For the causal part, recall
level 15’s np.triu.
CodeBuild the three masks of one example. src holds the digit ids, tgt_in the decoder input ids, and pad is the id of PAD. Return (enc, dec, cross), True where a cell is blocked: enc is (Ls, Ls), dec is (Lt, Lt), cross is (Lt, Ls). Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
🔒 Answer the question above to unlock
Now use the cross mask inside cross-attention. Write one head, with the PAD source tokens blocked.
CodeWrite one cross-attention head. y holds the decoder’s rows and memory the encoder’s output; src_pad is True for a PAD source token. Return (out, A): the output and the weights. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Go deeper Label smoothing, and why the loss never reaches 0
This model uses label_smoothing=0.1, as the 2017 design did. The target for each position is not “100% the right
word”, but “90% the right word, and the remaining 10% spread over all 32 words”. It stops the model from becoming
completely certain, which tends to help it on unseen inputs.
Because of this, even a perfect model can’t reach a loss of 0. The lowest possible loss is the entropy of the
smoothed target. The right word gets 0.9 + 0.1 / 32 = 0.903125, and each of the other 31 words gets
0.1 / 32 = 0.003125. So the lowest loss is −(0.903125 ln 0.903125 + 31 × 0.003125 ln 0.003125) ≈ 0.651. This run ends at a loss of 0.652, only 0.001 above it, while
getting 100% of the unseen numbers right. Unseen accuracy over the run: 0% after epoch 1, 89% after 10, 98% after 20,
98% after 40, 100% after 60.
🔒 Answer the question above to unlock
6. Reading the original 2017 design
You now know every part of the original 2017 Transformer. Its diagram and its text use these names.
The right column says where this course teaches each one. You may not have done some of those levels yet.
name in the 2017 design
what it is
where this course teaches it
Input Embedding, Output Embedding
the token tables: one for the source, one for the target
The original sizes: d_model = 512, h = 8 heads, so dₖ = 512 / 8 = 64; d_ff = 2048 = 4 × d_model; 6 layers on each side.
Today’s chat models keep only the decoder stack, as level 17 does: the “source” is simply the start of the sequence.
The encoder–decoder design is still used where the input and the output are clearly two different things, for example
translation or speech to text.
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
Name the sublayers of an encoder layer and a decoder layer, and say where cross-attention takes Q, K and V from.
Give the shape of every attention’s scores and mask in an encoder–decoder, <bos> included.
Build the three masks of one example in NumPy.
Keep in mind
Cross-attention: Q from the decoder, K and V from memory; scores (B, h, L_tgt, L_src)
tgt_in = <bos> + answer, tgt_out = answer + <eos>: position t guesses word t + 1
Only the decoder’s self-attention is causal; encoder self-attention and cross-attention block PAD source tokens only
The 2017 placement wraps each sublayer as LayerNorm(x + sublayer(x)) (post-norm)
Common mistakes
Making the decoder’s causal mask from the source length: it follows the decoder input, <bos> included.
Swapping the axes of cross-attention: rows are the decoder’s tokens (Q), columns the source tokens (K).