2. The shape rule
The shape rule comes from that one rule. To multiply a row of A with a column of B number by number,
the row and the column must have the same length. In NumPy, and everywhere in this course, @ means matrix multiplication. So:
The two inner numbers must match, and they disappear. The two outer numbers are what is left. When the inner numbers match, we say the shapes line up.
The flip in that answer has a name: the transpose, written B.T. Row i of B becomes column i of B.T,
so a (3, 2) matrix turns into (2, 3). You will use it in level 6 and again in attention.
I got stuck here Why not just multiply cell by cell, like adding two matrices?
That operation exists too. In NumPy it is A * B. It needs both matrices to have the same shape,
and it is used in places like masks. But it never mixes information between positions:
cell (0,0) only ever sees cell (0,0).
Matrix multiplication does mix. Every output cell combines a whole row with a whole column. That mixing is exactly what a neural network layer needs. Each output is a weighted sum of all the inputs.
3. Write it yourself
Write the multiplication with three loops, without @ or np.dot. The innermost loop runs over the inner size,
which is the dimension that disappears.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs