Now check yourself in the lab, and type anything else you like.
A character tokenizer
Type any text. Point at, tap or tab to a character to see every place it is used.
Vocabulary: 4 distinct characters, sorted. The small number is the id.
Encoded: 5 tokens, each character above its id.
The vocabulary gives two dictionaries: stoi (string to id) and itos (id to string).
Code that builds them uses two Python tools:
list(enumerate(["e", "h", "l"])) # [(0, 'e'), (1, 'h'), (2, 'l')] each item with its position
{c: i for i, c in enumerate(["e", "h"])} # {'e': 0, 'h': 1} a dictionary built in one lineThat one-line form is called a dictionary comprehension. The same form with [ ] builds a list (a list comprehension).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Encoding looks up every character. Decoding goes back. Nothing is lost on the way.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
I got stuck here Why sort the characters? Does it matter which character gets which id?
It doesn’t matter at all. Any fixed order works, as long as encoding and decoding use the same one. Sorting is an easy way to get the same order every time you run the code.
The ids are names, not amounts: “l” being 2 doesn’t make it twice “h”. Level 13 shows how the model avoids reading them as amounts.