paddingExtra placeholder tokens (often written <PAD>) added to the end of short sequences so every sequence in a batch has the same length. Masks and the loss ignore them.
Taught in level 15: Masks and heads → padding (convolution)also: zero paddingExtra border cells (usually zeros) added around an input so the filter can sit on the edges and the output keeps its size.
Taught in level N1: Convolutions → padding maskA mask that blocks every PAD token (the extra tokens that make short sequences as long as the longest one), so no word looks at padding.
Taught in level 15: Masks and heads → parameterAny number the model learns during training, such as a weight or a bias. Model size is usually counted in parameters.
Taught in level 5: Multilayer perceptron → patchA small square cut from a picture. A Transformer for images treats each patch as one token.
Taught in level D3: Latents and DiT → perplexitye to the power of the average cross-entropy. A perplexity of 4 means the model is as unsure as if it were choosing among 4 equally likely words.
Taught in level 10: Probability and sampling → plateaualso: plateausA part of training where the loss or accuracy stays flat for a while before it starts to improve again.
Taught in level 21: Write your own GPT → poolingShrinking a feature map by keeping one number (often the largest) from each small block.
Taught in level N1: Convolutions → positional encodingNumbers added to each token’s embedding that say where it is in the sequence. Attention alone does not know the word order.
Taught in level 16: Transformer parts → post-normThe 2017 order: add the sublayer’s output to x, then normalize, LayerNorm(x + f(x)). Deep post-norm stacks need a long warmup to train.
Taught in level 19: Modern blocks → pre-normPutting the normalization before each sublayer instead of after it, so nothing rescales the residual path and deep stacks train reliably.
Taught in level 16: Transformer parts → precisionHow exact a float format is: how many numbers it has between one power of 2 and the next. More mantissa bits give more precision.
Taught in level U2: Numbers in a computer → prefillReading the whole prompt in one forward pass before the answer starts. The weights are read once for all prompt tokens. So for a long prompt, arithmetic limits it, not memory: on the example GPU of level 20, longer than about 100 tokens. It also fills the KV cache.
Taught in level 20: Inference cost → promptThe tokens you give the model at the start. The model continues them one token at a time, and the tokens it writes are the answer.
Taught in level 17: The whole Transformer →