Inside the model · LayerNorm & Residual
- shown
- toy scale
- features
- 8 GPT-2 768
- LayerNorms
- 25 2 per block + lnf
- params each
- 16 GPT-2 1,536
- ε
- 1e-5
All steps
Residual stream
The residual stream runs straight through every block: each sub-layer reads a normalised copy and adds its result back, so information is only ever added, and that direct path keeps training stable. GPT-2 normalises before each sub-layer (pre-LN); the 2017 Transformer normalised after each add. Click attn or mlp to open them.
GPT-2: pre-LN, 12 blocksSubtract the mean
LayerNorm works on one token at a time, over its features. First the mean of the token's features is subtracted, so they centre on 0. (BatchNorm, common in vision, normalises each feature across a batch instead; a token's result would then depend on the other sequences in the batch, so Transformers use LayerNorm.)
μ over 768 featuresDivide by the std
Then everything is divided by the standard deviation, so every token's features have the same spread no matter how large the stream has grown. That keeps each sub-layer's input in a range it was trained on.
σ = √(var + ε)Scale and shift
Finally each feature is scaled by γ and shifted by β, both learned. The result X is what the attention layer receives.
X = γ ⊙ x̂ + β
Code
class Block(nn.Module):
def forward(self, x):
x = x + self.attn(self.ln_1(x)) # read a normalised copy, add back
x = x + self.mlp(self.ln_2(x))
return x
# what ln_1 computes for each token's 768 features (F.layer_norm):
mu = x.mean(-1, keepdim=True)
var = x.var(-1, keepdim=True, unbiased=False)
x_hat = (x - mu) / torch.sqrt(var + 1e-5)
y = self.weight * x_hat + self.bias # γ ⊙ x̂ + β