1. The whole machine
In order, from input to output: a token table turns each id into a vector. Then the position numbers of level 16 are added. Then come a stack of decoder blocks, one final LayerNorm, and an output layer that gives one score per token. Select a block to see its input, its output, and how many parameters it holds. The shapes are for one sequence of 8 tokens, the longest input this task has.
The whole model, block by block
Pick a block to see its input, its output and how many parameters it holds. Change the sizes below.
Each block: LayerNorm, masked self-attention (4 heads of size 16; each position sees only itself and earlier ones), added back to x; then LayerNorm, the FFN (64 → 128 → 64), added back. One block has 33,280 parameters: attention 16,448, FFN 16,576, two LayerNorms 256.
parameters: 66,560 (92.3% of all)
d_model = 64. What shape are the logits?d_model stays 64 and you change heads from 4 to 8. The parameter count…Most of the parameters are in the blocks, and inside each block most of them are in two places: the attention matrices and the FFN. You can count one FFN by hand.
d_model = 64 and d_ff = 128. How many parameters does one FFN have? It is Linear(64 → 128), ReLU, Linear(128 → 64). Count weights and biases.Go deeper Where the 72,106 parameters are
At the default size (d_model 64, 4 heads, 2 blocks, d_ff 128):
| part | parameters |
|---|---|
| token table: 42 × 64 | 2,688 |
one attention block: W_Q, W_K, W_V (64 × 64 each) + W_O (64 × 64 + 64) | 16,448 |
| one FFN | 16,576 |
| one block: attention + FFN + 2 LayerNorms (128 each) | 33,280 |
| final LayerNorm | 128 |
| output layer: 64 × 42 + 42 | 2,730 |
| total: 2 blocks + token table + final LayerNorm + output | 72,106 |
The position numbers are added, not learned, so they have no parameters.
The number of heads does not appear anywhere in this table. It only decides how the 64 columns are cut into slices.
The 42 tokens are <pad>, <eos>, the ten digits, =, and 29 words: zero to nineteen, the eight tens words, and “hundred”.