Level 13 · Theory · runs in your browser

Embeddings and similarity

How does a token id become a list of numbers that means something?

Level 12 turned text into token ids. For hello, cut into characters, the vocabulary is e, h, l, o (ids 0 to 3), and the word becomes [1, 0, 2, 2, 3].

A network can’t use those ids as they are. It would treat them as amounts: l (2) would be “twice” h (1). So each id is replaced by a short list of numbers, a vector, looked up in a table. This is called embedding.

1. The embedding table

The table has one row per id. Embedding a token means taking its row. Here each row has 2 numbers; real models use hundreds or thousands. So a sentence of L tokens becomes an array of shape (L, d): one row per token.

Shape The table E has shape (4, 2): 4 ids, 2 numbers each. You look up the 5 ids of “hello”, [1, 0, 2, 2, 3]. What is the shape of the result?
Choose“hello” has two l’s, at positions 2 and 3 (counting from 0). After the lookup, are their two vectors the same?
🔒 Answer the question above to unlock

From token to vector: a table lookup

Pick a token to see its one-hot row times the table. Drag a point (or focus it and use the arrow keys) to change its row.

Token ids (5 tokens). Pick one.

Embedding table E, shape (4, 2)

idcharone-hotE row
0e0[-1.4, 0.1]
1h1[1.5, -1.2]
2l0[-0.1, -1.5]
3o0[-1.3, 1.5]
-2-2-1-11122number 1 →↑ number 2ehlo
“h” → id 1 → one-hot [0, 1, 0, 0] @ E = row 1 = [1.5, -1.2]

“Take row 2” can also be written as a matrix multiplication. Make a one-hot row: all zeros, except a 1 at position 2. Multiply it by the table, and every row except row 2 is multiplied by 0. Select a token in the lab and look at the one-hot column.

CodeEmbed by multiplying one-hot rows by the table.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

I got stuck here Why not give the ids to the network as numbers? h = 1, l = 2 are already numbers.

Because the network would treat them as amounts. It would “think” l (2) is twice h (1), and that o (3) is somewhere past l. None of that is true. The ids are only names, in alphabetical order.

A vector of several numbers per token, which the network is free to learn, avoids this. Two tokens can become close or far apart for real reasons, and there are many directions to be close in.

Go deeper Where the numbers in the table come from

With the BPE tokens of level 12, the table has tens of thousands of rows, and each row has hundreds to thousands of numbers. The idea is the same as here: an id picks a row.

The rows start as random numbers. They are weights, exactly like the weights in level 3, and training moves them with gradient descent. Tokens used in similar ways get similar rows, because that helps the model predict the next token. Nobody writes the meaning in by hand. Section 3 shows why this happens.

🔒 Answer the question above to unlock

2. Comparing two vectors: the dot product

If similar tokens get similar vectors, you need a number for “how similar”. The simplest is the dot product: multiply the two vectors number by number and add up. You did exactly this in level 1: it is one cell of a matrix multiplication.

Dot product: direction and length

Drag an arrow tip (or focus it and use the arrow keys); tips move in steps of 0.5. Pick a pair to see its angle.

x →↑ ycatdogcar
wordvectorlength
cat[2, 1]2.24
dog[1.5, 1.5]2.12
car[-1, 2]2.24
pairdot productcos(angle)
2×1.5 + 1×1.5 = 4.50.949
2×-1 + 1×2 = 00.000
1.5×-1 + 1.5×2 = 1.50.316
Pick a pair in the table to see its angle.
positive dot product: roughly the same direction 0: at a right angle negative: roughly opposite
Number What is [2, 1] · [1, 3]?
Choosecat = [2, 1]. You turn car so it points the opposite way from cat. What happens to cat·car?
Try it

cat·car starts at exactly 0: cat and car are at a right angle. Now make dog·car exactly 0 without touching dog. Then drag dog farther from the center, in the same direction. Its dot product with cat grows, but cos(angle) stays the same. The dot product mixes two things: direction (the angle) and length.

Go deeper Dot product, length, and angle

For two vectors aa and bb:

a⋅b=∣a∣ ∣b∣cos⁡θa \cdot b = |a|\,|b|\cos\theta

cos⁡θ\cos\theta is 1 when they point the same way, 0 at a right angle, and −1 when they point opposite ways. Dividing the dot product by both lengths gives just the angle part, called cosine similarity.

Attention, in the next level, uses the plain dot product. Length matters there: a longer vector gets bigger scores.

More than two numbers

Real embeddings have hundreds of numbers per token, and the dot product works the same way: multiply number by number and add up. With three numbers you can still see the vectors. Turn the picture and pick two words. The three axes are simply the 1st, 2nd and 3rd number of each vector. They have no names: training decides what they mean.

Word vectors in 3D

Pick two words. Drag the picture to turn it (select it first to zoom).

Drag to turn · click, then scroll to zoom

First word

Second word

Dot product with cat, biggest first

  1. lion6
  2. dog4
  3. bus1
  4. car-1
  5. truck-4
cat · dog = 2×2 + 1×0 + 0×1 = 4 = shadow 1.79 × length of cat 2.24 cos(angle) = 0.80
first word second word its shadow on the first word’s line The small arc marks the angle between the two words; its cosine and the shadow’s length are in the box above. Gray arrows are the other words; one that sits close to a picked word is not named, so pick it to see which it is. The 1st, 2nd and 3rd numbers of each word are its x, y and z. Animals point one way and vehicles another: same direction gives a big dot product, a right angle gives 0 (no shadow), opposite directions go negative (the shadow falls behind the start).
🔒 Answer the question above to unlock

3. Where do the vectors come from?

The table starts as random numbers, and training moves them. But what pushes “cat” and “dog” together? The idea is short: the meaning of a word comes from the words around it. Cat and dog appear in the same places (“the … sleeps”, “my … is hungry”), so a model that has to predict a word’s neighbors does best if it gives them similar vectors.

You can see this with nothing but counting. Take four tiny sentences: “cat eats fish”, “dog eats meat”, “cat drinks milk”, “car needs fuel”. For each word, count how often each of eats, drinks and needs sits right next to it (one word before or after).

Number Sentences: “cat eats fish”, “dog eats meat”, “cat drinks milk”, “car needs fuel”. How many times does one of eats, drinks, needs sit right next to “cat” (one word before or after)?
🔒 Answer the question above to unlock

So cat = [1, 1, 0] over (eats, drinks, needs). The same count gives dog = [1, 0, 0] and car = [0, 0, 1]. These counts are already vectors, and the dot product from section 2 compares them.

Number Over (eats, drinks, needs): cat = [1, 1, 0], dog = [1, 0, 0], car = [0, 0, 1]. What is cat · dog?
🔒 Answer the question above to unlock
ChooseTake thousands of sentences like “cat eats fish”, “dog eats meat” and “car needs fuel”. cat and dog keep appearing in the same places, and car in different ones. Whose count vector is finally closest to cat’s?
🔒 Answer the question above to unlock

Our own rules wrote 3,850 short sentences about animals, foods and vehicles. Count every word within two words of every other, and each word gets a long row of counts. Comparing those rows (with cosine similarity) and drawing them in 2D places the words like this:

Words placed by their neighbors

Pick a word. Lines join it to the 3 words whose rows of counts are most alike.

Loading the vectors…

Similarities appear when the vectors have loaded.

Nobody said which words are animals. The three groups appear because each group is used in its own kind of sentence. Inside a group the words are close but not equal: each word also has a few sentences of its own (“the dog barks at night”).

Try it

Pick “fish”. Its most frequent neighbors mix two uses: “people” (people eat fish) next to “swims”, “drinks” and “water”. Our rules sometimes use fish as an animal. Its point sits between the foods and the animals, closest to the foods. One vector has to serve both uses, which is the problem of the next section.

Here is the counting step in code. C[i, j] += 1 adds 1 to row i, column j of the table.

CodeCount neighbors: for every word, add 1 for each word right before and right after it.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper Counting versus predicting

This level counts neighbors and then compresses the counts (the compression step is a matrix factorization; it keeps the directions in which the rows differ most). The more famous method predicts instead: a tiny network gets a word and is trained to give high probability to the words around it, and its embedding table is the result. Both finish with similar vectors, because both are driven by the same thing: which words share neighbors.

A language model (level 11, and the GPT of level 21) does the same without being asked to. Its embedding table is trained only to help predict the next token, and words used in similar ways get similar vectors.

🔒 Answer the question above to unlock

4. One vector per word is not enough

A lookup table always returns the same row for the same token, whatever the sentence. That is a problem for a word like apple: a fruit in one sentence, a company in another.

Here the four numbers have readable meanings: fruit, tech, sweet, action. The table gives apple = [1, 1, 0, 0]: one on the fruit number and one on the tech number, because the same row has to serve both meanings. Take the two sentences “sweet apple” and “apple releases phone”.

ChooseThe word “apple” is looked up in the same table in both sentences. Is its vector the same in “sweet apple” and in “apple releases phone”?
🔒 Answer the question above to unlock

Same word, two sentences

Check the box to average every word in the sentence with equal weights, a simple first version of attention.

sweet apple

fruittechsweetaction
sweet1010
apple1100
apple →1100

apple releases phone

fruittechsweetaction
apple1100
releases0101
phone0100
apple →1100
With a lookup table, “apple” gets exactly the same vector in both sentences: fruit 1, tech 1.

Check the box to mix each sentence’s words together with equal weights. In the table, sweet = [1, 0, 1, 0], releases = [0, 1, 0, 1], phone = [0, 1, 0, 0].

Number Average the three vectors of “apple releases phone” (apple [1, 1, 0, 0], releases [0, 1, 0, 1], phone [0, 1, 0, 0]). What is the fruit number, the first entry, of the result? (3 decimals)

Now write the step in code, with one change: the words don’t have to count equally. Each word gets a weight, the weights add up to 1, and the mix is the weighted sum of the rows. Equal weights (1/2 and 1/2, or 1/3 each) give the plain average from the lab. Indexing a table with a list of ids takes those rows, in that order:

T = np.array([[1, 2], [3, 4], [5, 6]])
T[[2, 0]]                 # [[5, 6], [1, 2]]   rows 2 and 0, shape (2, 2)
0.75 * T[2] + 0.25 * T[0] # [4, 5]   weights 0.75 and 0.25: 0.75·[5, 6] + 0.25·[1, 2]

Writing one term per word only works for a fixed number of words. Level 1 has a shorter way: a vector of weights @ a matrix gives exactly this weighted sum of its rows, for any number of rows.

CodeWrite mix: look up the rows of ids in the table E, then mix them with the weights w (one weight per id, adding up to 1) into one vector of 4 numbers. For example, ids [1, 0] with w = [0.75, 0.25] is 0.75 × sweet + 0.25 × apple. Two lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Averaging with equal weights is too simple: every neighbor counts the same, even unrelated words. The next level fixes exactly this. Each word computes how much to take from each neighbor, using the dot product you just learned.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Turn a list of token ids into an (L, d) array by taking rows of the embedding table.
  • Compute a dot product by hand and read its sign: same way, right angle, opposite way.
  • Count neighbors to give each word a row, and average a sentence’s vectors with equal weights.

Keep in mind

  • Embedding = taking row id of E; onehot @ E gives the same rows
  • n ids looked up in E of shape (V, d) → (n, d)
  • : [2, 1] · [1, 3] = 5
  • E[ids].mean(axis=0): an equal-weight mix, shape (d,)
  • Words used in the same places get similar vectors

Common mistakes

  • Expecting the table to give a word a different vector in a different sentence: the lookup always returns the same row.
  • Putting the table first in E @ onehot: the one-hot rows (L, 4) go first.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars