Inside the model · LayerNorm & Residual

LayerNorm & Residual

pre-LN · ln1 of block 1

shown
toy scale
features
8 GPT-2 768
LayerNorms
25 2 per block + lnf
params each
16 GPT-2 1,536
ε
1e-5

All steps

  1. Residual stream

    The residual stream runs straight through every block: each sub-layer reads a normalised copy and adds its result back, so information is only ever added, and that direct path keeps training stable. GPT-2 normalises before each sub-layer (pre-LN); the 2017 Transformer normalised after each add. Click attn or mlp to open them.

    GPT-2: pre-LN, 12 blocks

  2. Subtract the mean

    LayerNorm works on one token at a time, over its features. First the mean of the token's features is subtracted, so they centre on 0. (BatchNorm, common in vision, normalises each feature across a batch instead; a token's result would then depend on the other sequences in the batch, so Transformers use LayerNorm.)

    μ over 768 features

  3. Divide by the std

    Then everything is divided by the standard deviation, so every token's features have the same spread no matter how large the stream has grown. That keeps each sub-layer's input in a range it was trained on.

    σ = √(var + ε)

  4. Scale and shift

    Finally each feature is scaled by γ and shifted by β, both learned. The result X is what the attention layer receives.

    X = γ ⊙ x̂ + β

Code

class Block(nn.Module):
    def forward(self, x):
        x = x + self.attn(self.ln_1(x))   # read a normalised copy, add back
        x = x + self.mlp(self.ln_2(x))
        return x

# what ln_1 computes for each token's 768 features (F.layer_norm):
mu = x.mean(-1, keepdim=True)
var = x.var(-1, keepdim=True, unbiased=False)
x_hat = (x - mu) / torch.sqrt(var + 1e-5)
y = self.weight * x_hat + self.bias      # γ ⊙ x̂ + β

Go deeper