# 4. Activation functions

> Why does a network need a non-linear step, and which one should it use?

LLM by Hand · Foundations · runs in your browser · interactive page: https://llm.liko.page/learn/activations/

In level 3 one neuron drew one straight line, and no line can separate XOR. On XOR the neuron's best
cross-entropy loss was ln 2 ≈ 0.6931, the cost of a coin flip. This level stacks neurons, finds out what must sit
between them, and compares the four functions used there most often.

## 1. Stacking neurons: does that fix it?

If one line is not enough, use several neurons side by side (a **layer**) and pass their outputs to
another neuron. First try it without sigmoid, with plain linear layers.

Take a point $x = [2, 1]$ and two layers:

$$
W_1 = \begin{bmatrix} 1 & 2 \\ 0 & 1 \end{bmatrix} \qquad W_2 = \begin{bmatrix} 1 \\ 1 \end{bmatrix}
$$

The first layer computes $h = x W_1$. The second computes $h W_2$.
(In math, $x W_1$ with nothing between the two letters means the matrix product `x @ W1` in code.)

**Question.** x = [2, 1], W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. First h = x @ W1, then out = h @ W2. What is out?

*Answer it on the page to check your work.*

Matrix multiplication lets you group the two weight matrices first: $(x W_1) W_2 = x (W_1 W_2)$.

**Question.** W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. W1 @ W2 has shape (2, 1). What is its top entry?

*Answer it on the page to check your work.*

So the two layers are really one layer with weights $W_1 W_2$. However many linear layers you stack,
the result is still a single line. Extra layers add nothing unless something nonlinear sits **between** them.

**Deeper: What the nonlinearity actually does**

Put a nonlinear function between the layers, for example $h = \tanh(x W_1 + b_1)$. tanh is sigmoid stretched to
the range −1 to 1: $\tanh(x) = 2\,\text{sigmoid}(2x) - 1$. Now $(\tanh(\cdot)) W_2$
can't be rewritten as one matrix, because tanh does not distribute over addition: $\tanh(a + b) \ne \tanh(a) + \tanh(b)$.

Each hidden neuron still draws one line, then tanh bends its output into "this side" (about +1) or
"that side" (about −1). The output neuron combines those answers. Two lines plus a combination can make
the XOR pattern: "between the two lines" is 1, "outside" is 0. That is the pattern you'll see the network
find in the lab below.

Any nonlinear function works in principle. The last part of this page compares the four you will meet most.

Here is a network with 6 hidden neurons, learning XOR for real in your browser. Its hidden neurons use tanh
(sigmoid stretched to −1…1, see the box above). Press Train.

*[Interactive lab: Mlp — open the page to use it]*

**Predict.** Guess before you try it: if you uncheck “activation on (tanh)” and press Train, what happens? Two layers, 6 hidden neurons, but no tanh.

A. It learns XOR, just more slowly
B. It gets 3 out of 4, like one neuron
C. It stops learning and answers 0.5 for every point

*Answer it on the page to check your work.*

With the activation back on, the network gets all four points right.

**Try it**

Press “New random weights” a few times with the activation on. The network finds a different boundary each time, but always
gets all four points right. What patterns do you see? Does it ever use one line?

Each of the 6 hidden neurons in the lab uses tanh. Is tanh the best choice? The next sections compare it with three other functions.

## 2. Four activation functions

The function between the layers is called the **activation function**. Here are the four you will meet. f′(x), “f prime of x”, is the
slope of f at x (level 2).

| name | f(x) | slope f′(x) |
|---|---|---|
| sigmoid | $\sigma(x) = 1 / (1 + e^{-x})$ | $\sigma(x)\,(1 - \sigma(x))$ |
| tanh | $\tanh(x)$ | $1 - \tanh(x)^2$ |
| ReLU | $\max(0, x)$ | 1 if $x > 0$, else 0 |
| GELU | $x \cdot \Phi(x)$ | $\Phi(x) + x\,\varphi(x)$ |

$\Phi(x)$ is the chance that a random number from the bell curve (mean 0, standard deviation 1; level 10 shows it)
lands below $x$, and $\varphi$ is the bell curve itself.
You don't need to memorize them. What matters is the **slope**, because training is driven by slopes:
level 6 shows that the gradient passes backwards through every layer, and at each layer it is multiplied
by that layer's f′.

**Question.** σ(0) = 0.5. Using f′(x) = σ(x)·(1 − σ(x)), what is the slope of sigmoid at x = 0?

*Answer it on the page to check your work.*

ReLU is even simpler. Its slope has only two values.

**Question.** ReLU(x) = max(0, x). What is its slope at x = −2?

*Answer it on the page to check your work.*

Now chain the slopes. Every layer multiplies the gradient by its own f′. (The weights multiply it too. Here we
ignore them; level 7 shows how the starting weight scale changes the picture.)

**Predict.** A gradient flows back through 10 sigmoid layers. Each layer multiplies it by that layer’s σ′, which is at most 0.25. Start with a gradient of 1 and ignore the weights. What reaches the first layer, at most?

A. About 0.25
B. About 0.06
C. Less than one millionth

*Answer it on the page to check your work.*

**Question.** Five sigmoid layers, ignoring the weights, best case: the gradient is multiplied by 0.25 five times. 0.25⁵ = 1 / N. What is N?

*Answer it on the page to check your work.*

## 3. See all four at once

Drag on any panel. The solid line is f and the dashed line is its slope. The table at the bottom multiplies
the slope through n layers.

*[Interactive lab: Activation — open the page to use it]*

Three things to see:

- **Sigmoid and tanh become flat at both ends.** Move x to 3: tanh′(3) is about 0.0099. A neuron whose input
  is that far from 0 passes almost no gradient back. This is called **saturation**. Stack many such layers and the
  gradient for the early layers shrinks toward zero: the **vanishing gradient**.
- **ReLU never flattens on the right.** Its slope is exactly 1 for every positive input, so gradients pass
  through unchanged. That is why it made deep networks trainable.
- **ReLU is completely flat on the left.** A neuron whose input is negative for every example gets zero gradient
  and never changes again.

**If you are stuck: If ReLU’s slope is 0 on the left, how does a neuron that went negative ever come back?**

It doesn't, through that path. If a neuron's input $z$ is negative for every training example, its slope is 0
for every example, so none of its incoming weights get any gradient. Nothing moves it back. This is a
**dead ReLU**. With a large learning rate a big update can push many neurons there at once, and a large part of the
network stops learning.

Other neurons can still change what feeds into it, so it can sometimes start learning again. But this does not happen reliably.
The usual ways to prevent it are a sensible starting scale for the weights (level 7), a learning rate that isn't too large,
and variants that keep a small slope on the left, such as Leaky ReLU ($0.01x$ for $x < 0$). GELU also has a small
negative side near 0, though for very negative inputs its slope gets close to 0 too.

**Question.** GELU(x) = x · Φ(x), and Φ(−1) = 0.1587. What is GELU(−1)? (Four decimals.)

*Answer it on the page to check your work.*

**Deeper: Why GELU replaced ReLU in many Transformers**

GELU looks like ReLU from far away: about 0 for very negative inputs, about $x$ for large positive ones.
Up close it differs in two ways that help training:

1. **It is smooth.** ReLU has a corner at 0, so its slope jumps from 0 to 1. GELU's slope changes
   gradually, which makes the loss surface smoother for the optimizer.
2. **It is not dead on the left.** Between about −3 and 0 it dips slightly below zero and still has a slope,
   so a neuron whose input becomes negative keeps getting gradient and can come back.

You can read GELU as "multiply $x$ by the chance that it should pass": large $x$ passes almost fully,
very negative $x$ is blocked, and values near 0 pass partly. The small Transformer in level 16 keeps ReLU because it is simpler to compute
by hand. Many large models use GELU there instead, and level 19 shows the gated version that most models use today.

## 4. Write the slope yourself

Backpropagation (level 6) needs each activation's slope as a function. Two NumPy tools make this one line.
A comparison works on every element of an array at once, and `.astype(float)` turns True/False into 1.0/0.0:

```
x = np.array([-1.0, 0.0, 2.0])
x > 0                  # array([False, False,  True])
(x < 1).astype(float)  # array([1., 1., 0.])
```

A plain `if x > 0:` does not work on a whole array, because Python cannot decide if the whole array is "true".
Write ReLU's slope for a whole array.

**Code question.** Write the slope of ReLU for every element of an array x (1 where x > 0, else 0).

Fill in the blank (`____`):

```python
def relu(x):
    return np.maximum(0, x)

def relu_grad(x):
    return ____

x = np.array([-2.0, -0.5, 0.0, 0.5, 3.0])
print("x        ", x)
print("relu     ", relu(x))
print("relu_grad", relu_grad(x))
```

*Answer it on the page to check your work.*

Now the sigmoid's slope from the table in section 2: $\sigma(x)\,(1 - \sigma(x))$. In code, $\sigma(x)$ is
`1 / (1 + np.exp(-x))`, as in the neuron of level 3, and `np.exp` works on every element of an array.
Compute $\sigma(x)$ once, give it a name, then use it twice.

**Code question.** Write the slope of sigmoid for every element of an array x: first σ(x), then σ(x)·(1 − σ(x)). Two lines.

Fill in the blank (`____`):

```python
def sigmoid_grad(x):
    ____

x = np.array([-2.0, 0.0, 3.0])
print("sigmoid_grad", sigmoid_grad(x).round(4))
print("five layers at x = 0:", sigmoid_grad(np.array([0.0]))[0] ** 5)
```

*Answer it on the page to check your work.*

**Try it**

In the lab, set x = 0.5 and n = 20. Which functions still pass a usable gradient after 20 layers?
Now set x = −0.5. What happens to ReLU, and what happens to GELU?

## You can now

- Show that two layers with no activation between them merge into one layer.
- Read the slope of sigmoid, tanh, ReLU and GELU, and compute them in NumPy for a whole array.
- Explain why the gradient vanishes through many sigmoid layers.
