Level U2 · Foundations · Under the hood · runs in your browser

Numbers in a computer

Why does a loss suddenly become NaN, and why does softmax subtract the largest score?

Side trip · best after level 3 · Neurons and loss

Sections 6 and 7 use softmax and cross-entropy, which level 10 teaches. This page explains the part it needs.

On paper a number can be as big or as precise as you like. In a computer each number gets a fixed number of bits: 32, 16, sometimes 8. That limit explains several training bugs that seem to have no cause:

  1. a loss that suddenly becomes nan (“not a number”);
  2. softmax code that subtracts the largest score before exp;
  3. weights that stop changing although the gradient is not 0.

1. A number in bits

A float is stored like a number in scientific notation, but with powers of 2 instead of powers of 10 (the math page has a reminder). For example, 6=1.5×226 = 1.5 \times 2^2. The bits hold three parts:

value=(−1)s×(1+mant/2M)×2 exp−bias\text{value} = (-1)^{s} \times (1 + \text{mant}/2^{M}) \times 2^{\,\text{exp} - \text{bias}}

Three formats matter for deep learning:

formatbitsexponent bitsmantissa bits MMbiaslargest number
float3232823127about 3.4×10383.4 \times 10^{38}
float16165101565504
bfloat161687127about 3.4×10383.4 \times 10^{38}

In NumPy these are np.float32 and np.float16; NumPy’s default np.float64 uses 64 bits. NumPy has no bfloat16. PyTorch has all three: torch.float32, torch.float16, torch.bfloat16.

Type a number into the lab, or tap a bit to flip it. The line under the bits computes the value with the formula above.

What the bits store

Type a number or pick one, then tap a bit to flip it.

sign
exponent (5 bits) = 11
mantissa (10 bits) = 614
(−1)0 × (1 + 614/1024) × 211 − 15 = 0.0999755859375
largest65504gap after 10.000977
smallest normal6.10e−5numbers in [1, 2)1024
0.1 in float16 → 0.0999755859375, off by 2.4e−5
sign: 0 is +, 1 is − exponent: which power of 2 mantissa: where between that power and the next
Number A float16 number has the bits 0 10001 1000000000: sign 0, exponent 10001 (that is 17), mantissa 1000000000 (that is 512). float16 has 10 mantissa bits and bias 15. What number is it?
🔒 Answer the question above to unlock

2. Range and precision

The two 16-bit formats split their bits differently, and that is the whole difference between them:

  • The exponent bits set the range: how big and how small a number can be. bfloat16 keeps float32’s 8 exponent bits, so it has float32’s range. float16 has only 5, so it stops at 65504.
  • The mantissa bits set the precision: how many numbers there are between one power of 2 and the next.
Number bfloat16 has 7 mantissa bits. How many different bfloat16 numbers are there from 1 up to 2 (1 included, 2 not included)?
🔒 Answer the question above to unlock

The same 128 mantissa patterns are used between every two powers of 2: from 1 to 2, from 2 to 4, from 256 to 512. So the gap between neighbors grows with the size of the number. Precision is relative: about 2 to 3 correct digits for bfloat16, wherever the number is.

Number In bfloat16 there are 128 numbers from 1 up to 2, so neighbors are 1/128 apart. The numbers from 256 up to 512 use the same 128 mantissa patterns, scaled by 256. How far apart are neighbors there?
🔒 Answer the question above to unlock

A result that falls between two neighbors is rounded to the nearer one. That is all a computer can store.

ChooseIn bfloat16, neighbors from 256 to 512 are 2 apart. What is 256 + 0.75, stored in bfloat16?

Most numbers are not stored exactly, even simple ones. 0.1 is not a fraction with a power of 2 at the bottom (the denominator), so each format stores the nearest number it has: 0.100000001490 in float32, 0.099975585938 in float16 and 0.100097656250 in bfloat16.

Try it

Press 0.1 in the lab, then switch between the three formats and compare how far each one is from 0.1. Then press 256.75 in bfloat16 and in float16: float16 has more mantissa bits, so it keeps the 0.75.

The same rounding happens in training. A weight gets a small change at every step, and in a 16-bit format a small change to a weight near 1 can be lost completely in the rounding.

Number A weight w = 1 is stored in bfloat16, where the next number above 1 is 1 + 1/128 ≈ 1.0078. An update adds 0.001. What is w after the update, stored in bfloat16?
Go deeper Exactly halfway (ties), and the numbers below the smallest normal one

Ties. When a result lies exactly halfway between two neighbors, it goes to the neighbor whose last mantissa bit is 0 (“round half to even”). So 256 + 1 in bfloat16 is 256, not 258. Always rounding halves upward would make long sums slowly grow too big; this rule rounds up and down equally often.

Subnormals. When the exponent bits are all 0, the format drops the “1 +” and stores mant/2M×21−bias\text{mant}/2^{M} \times 2^{1 - \text{bias}}. These tiny numbers fill the space between 0 and the smallest normal number. In float16 the smallest one is 2−24≈6×10−82^{-24} \approx 6 \times 10^{-8}. A smaller number rounds to whichever of 0 and 2−242^{-24} is nearer, so anything below about 3×10−83 \times 10^{-8} becomes 0. When the exponent bits are all 1, the pattern means inf (mantissa 0) or nan (anything else).

🔒 Answer the question above to unlock

3. Overflow and underflow

When a result is bigger than the largest number of its format, it overflows and becomes inf (infinity). NumPy gives no error, only a warning that is easy to miss.

ChooseThe largest float16 is 65504. What is np.float16(300) * np.float16(300)?

exp is where this happens most often in a network, because it grows so fast: each +1 to zz multiplies eze^z by about 2.72 (the math page has a reminder about ee).

Number e¹¹ ≈ 59 874 and e¹² ≈ 162 755. The largest float16 is 65504. What is the largest whole number z for which np.exp(np.float16(z)) is not inf?
🔒 Answer the question above to unlock

The other side is underflow: a result too close to 0 becomes exactly 0. In float32, e−1000e^{-1000} is 0. In float16, any number below about 3×10−83 \times 10^{-8} rounds to 0. Small gradients in float16 can underflow like this and disappear.

Underflow becomes dangerous at the next step. Level 3’s loss takes −ln⁡p-\ln p. If pp underflowed to 0, then −ln⁡0=-\ln 0 = inf.

4. How nan appears, and how it spreads

nan means “not a number”. It comes from a few operations that have no sensible answer:

operationresult
inf - infnan
0 * infnan
inf / inf, 0 / 0nan
np.log(-1)nan

So an inf from overflow, or a 0 from underflow, is often one step away from a nan.

Choosew = [0.5, nan, −1] and x = [1, 2, 3]. What is w @ x?
I got stuck here My loss is a number for 300 steps and then nan. How do I find where it starts?

Look for the first inf, not the first nan: the nan is usually one step later. Print the largest absolute value of the scores before every exp and every softmax, and of the gradients (np.isnan(x).any() and np.isinf(x).any() check a whole array). The usual causes, in order:

  1. a learning rate so high that the weights grow without limit (level 2);
  2. exp of big scores without subtracting the max (section 6);
  3. np.log of a probability that underflowed to 0 (section 7);
  4. a division by a sum or a std that can be 0. Add a small ε to it (level 16’s LayerNorm does this).
🔒 Answer the question above to unlock

5. Why big models train in 16 bits anyway

A 16-bit number takes half the memory of a float32. A model with 7 billion weights takes 28 GB in float32 and 14 GB in bfloat16, and the graphics card moves half as many bytes. Graphics cards also multiply 16-bit matrices several times faster. So large models use mixed precision: each operation runs in the format that is safe for it.

  • The big matrix multiplications run in bfloat16 (or float16). They are most of the work.
  • A master copy of the weights stays in float32. The update is added there, so a change of 0.001 to a weight of 1 is not rounded away, as it was in section 2.
  • Sums of many numbers, softmax and the loss run in float32.

bfloat16 was made for this. It keeps float32’s range, so scores and gradients rarely overflow or underflow. It has less precision, so the float32 master copy and float32 sums keep the precision that matters.

Go deeper float16 needs loss scaling

float16 has more precision than bfloat16 but far less range. Gradients of 10−810^{-8} are common, and float16 stores them as 0. The fix is loss scaling: multiply the loss by a big number, such as 1024, before backward(). Every gradient is then 1024 times bigger and stays above float16’s smallest number (10−8×1024≈10−510^{-8} \times 1024 \approx 10^{-5} is stored). Divide the gradients by 1024 again, in float32, before the update. If any gradient became inf, skip that step and use a smaller scale. PyTorch’s torch.amp does all of this for you. With bfloat16, most training needs no loss scaling at all.

6. Softmax subtracts the largest score

You don’t need level 10 for this section: everything it uses is defined here. (Level 10 teaches softmax properly, and why it is used.) Softmax turns a list of scores into probabilities that add up to 1. Take ee to the power of each score, then divide each result by their sum:

pi=ezi∑jezjp_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

A model uses it to choose among many answers. Here two scores are enough. With z=[1,2]z = [1, 2]: e1≈2.72e^1 \approx 2.72 and e2≈7.39e^2 \approx 7.39, their sum is 10.1110.11, so softmax gives [2.72/10.11, 7.39/10.11]=[0.27,0.73][2.72/10.11,\ 7.39/10.11] = [0.27, 0.73]. Now try big scores.

ChooseScores z = [1000, 1001] in float32, where e⁸⁹ is already inf. What does the plain formula e^z / sum(e^z) give?
🔒 Answer the question above to unlock

The fix uses one fact. Adding the same number cc to every score does not change softmax:

ezi+c∑jezj+c=ezi⋅ec∑jezj⋅ec=ezi∑jezj\frac{e^{z_i + c}}{\sum_j e^{z_j + c}} = \frac{e^{z_i} \cdot e^{c}}{\sum_j e^{z_j} \cdot e^{c}} = \frac{e^{z_i}}{\sum_j e^{z_j}}

The factor ece^c is in the top and in the bottom of the fraction, so it cancels. Choose c=−max⁡(z)c = -\max(z): subtract the largest score. Then the largest score becomes 0 and its e0=1e^0 = 1. No eze^{z} can overflow, and the sum is at least 1, so it can’t be 0 either.

The lab computes both ways, rounding every step to the format you choose.

Softmax two ways

Pick two scores and a format. Watch where inf and nan appear.

score 0score 1sum
directly from ez
z10001001
ezinfinfinf
pnannan
subtract the largest score first (1001)
z − max−10
ez − max0.36811.37
p??
straight: p = [nan, nan] · subtract max first: answer the question below to see p
inf, nan too big for float32, or a result of inf / inf 0.73 the probabilities that softmax should give
Number Scores z = [1000, 1001]. Subtract the largest score first: [−1, 0]. Use e⁻¹ ≈ 0.37 and e⁰ = 1. What is the softmax probability of score 1 (the 1001), to two decimals?
🔒 Answer the question above to unlock
Try it

Press [−1000, −999]: now every eze^z underflows to 0, and 0 / 0 is nan. Then switch to float16 and press [10, 12]: float16 overflows at a score of about 11. Subtracting the max fixes every case.

In NumPy, z.max(axis=-1, keepdims=True) gives the largest score of each row, with shape (B, 1), so that it lines up with z of shape (B, V) (level 1’s broadcasting).

CodeWrite the line that makes softmax safe for big scores: subtract each row’s largest score before exp.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

7. Log-sum-exp: the loss without inf

Level 3’s loss for a yes/no answer was −ln⁡p-\ln p, with pp the probability of the right answer. With softmax over many answers the loss is the same, −ln⁡pt-\ln p_t for the target answer tt (its probability from softmax). It is called cross-entropy; level 10 uses it to train a model that picks a word. Here you only need this one formula.

Computed in two steps, first pp and then −ln⁡p-\ln p, it fails when pp underflows to 0. Write ptp_t out instead, and use ln⁡(a/b)=ln⁡a−ln⁡b\ln(a/b) = \ln a - \ln b and ln⁡ez=z\ln e^{z} = z (both on the math page):

−ln⁡pt=−ln⁡ezt∑jezj=ln⁡∑jezj  −  zt-\ln p_t = -\ln \frac{e^{z_t}}{\sum_j e^{z_j}} = \ln \sum_j e^{z_j} \;-\; z_t

The first part is the log-sum-exp of the scores. It has the same overflow problem as softmax, and the same fix. Subtract the largest score mm from every score:

ln⁡∑jezj=m+ln⁡∑jezj−m\ln \sum_j e^{z_j} = m + \ln \sum_j e^{z_j - m}
Number Use ln 2 ≈ 0.69. What is ln(e¹⁰⁰⁰ + e¹⁰⁰⁰), to two decimals? (A computer gets inf when it computes this directly, but you can do it by hand.)
🔒 Answer the question above to unlock

Now write it. The template picks each row’s target score for you: z[np.arange(B), t] takes, from row 0, the score at index t[0]; from row 1, the score at index t[1]; and so on. You write lse, one log-sum-exp per row. m[:, 0] turns the (B, 1) maximum into shape (B,).

CodeWrite cross_entropy from raw scores with log-sum-exp, so it never sees inf. z has shape (B, V); t holds each row’s target index. Compute lse, one log-sum-exp per row, shape (B,).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper This is what PyTorch’s cross_entropy does

torch.nn.functional.cross_entropy(z, t) takes the raw scores, not probabilities, for exactly this reason: inside, it computes log-sum-exp with the max subtracted first, and subtracts the target score. torch.logsumexp and torch.log_softmax exist for the same reason. If you ever write torch.log(torch.softmax(z, -1)), replace it with torch.log_softmax(z, -1): the same numbers when they are small, and no inf when they are big.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Read a float from its bits: sign, exponent and mantissa.
  • Say when a format overflows to inf, underflows to 0, or loses a small update.
  • Compute softmax and cross-entropy without inf or nan, with the max trick and log-sum-exp.

Keep in mind

  • value =
  • float32: 8 + 23 bits · float16: 5 + 10 (largest 65504) · bfloat16: 8 + 7 (range of float32, but only 7 mantissa bits)
  • : subtract the max before exp
  • , and cross-entropy = log-sum-exp −
  • inf − inf, 0 × inf, inf / inf and 0 / 0 give nan; every result computed from a nan is nan

Common mistakes

  • Computing np.log(softmax(z)) in two steps: when p underflows to 0 the loss is inf.
  • Keeping the weights themselves in bfloat16: small updates are rounded to 0 and training stops.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars