LLM by Hand

Run it on your computer

Most levels run in your browser and need nothing installed. A few parts run on your own computer: the bosses that train real models with PyTorch, and a few optional runs. Set this up once; it takes about ten minutes.

1. Get the files

Each part that runs on your computer links the files it needs, for example boss.py. Make one course folder, and inside it one folder for each level. Download that level’s files into its folder, and run every command for that level in that folder. The table in step 6 lists every file.

Open a terminal in that folder. A terminal is a window where you type commands. On macOS it is the app Terminal (Applications → Utilities). On Windows use PowerShell (from the Start menu). In the terminal,cd followed by the folder’s path moves you into the folder.

2. Install Python

You need Python 3.10 or newer. Check what you have:

python3 --version

macOS

Install from python.org, or with Homebrew: brew install python.

Windows

Install from python.org and check the box “Add python.exe to PATH”. Then use py where this page says python3.

Linux

Use your package manager, for example sudo apt install python3 python3-venv.

3. Make a virtual environment and install the packages

A virtual environment is a private folder of packages for this course, so nothing conflicts with the rest of your computer. Make it once, in a folder that holds your level folders (for example llm-by-hand/):

python3 -m venv .venv
# on Windows use:  .venv\Scripts\activate
source .venv/bin/activate
pip install numpy torch

When the environment is active, your prompt starts with (.venv). In a new terminal, run thesource line again before working on the course.

4. Run a script

With the environment active, go into the level’s folder and run the file there. For example, for level N4:

cd lstm          # the folder that holds boss.py
python boss.py

The scripts write their own files (a trained model, downloaded digits) into that folder or into a cache in your home folder.

5. CPU, NVIDIA GPU or Apple GPU

Everything works on a plain CPU. Some scripts use a faster device when they find one: "cuda" on a computer with an NVIDIA graphics card, "mps" on a Mac with an Apple chip (M1 or newer), otherwise"cpu". You can check what PyTorch sees with:

python -c "import torch; print(torch.cuda.is_available(), torch.backends.mps.is_available())"

6. Every file you can run, and how long it takes

Times on a recent laptop; a CPU without a GPU can be 2–3× slower.

levelfilestime
9 From NumPy to PyTorchfirst_run.pyabout 1 s
U5 Debugging a modelbroken_train.py, check.pya few seconds
N4 LSTM and GRUboss.pya few minutes
N5 Seq2seq and the first attentionboss.pya few minutes
N6 Autoencoders and VAEsdemo.pyabout 12 s (downloads the digits once)
D3 Latents and DiTboss.pyabout 2–3 min (downloads the digits once)
21 Write your own GPTskeleton.py, skeleton_hints.py, check.py, check_weights.jsonyour model trains in about 2 min

7. If something goes wrong