Architectures · Decoder-only
- compared
- LLaMA 3 8B vs GPT-2 small
- layers
- 32 GPT-2 12
- dmodel
- 4,096 GPT-2 768
- heads
- 32 q · 8 kv GPT-2 12
- dff
- 14,336 GPT-2 3,072
- vocab
- 128,256 GPT-2 50,257
- context
- 8,192 GPT-2 1,024
All steps
GPT-2 vs LLaMA
LLaMA keeps GPT-2's block: pre-norm, residual adds, causal attention. Four parts change, marked in the lower row; click one to jump to it. Also different: no bias terms anywhere, an untied output matrix (not WEᵀ), RMSNorm as the final norm, and a 128K-token vocabulary.
4 changes · same blockRotary positions
GPT-2 adds a learned position vector once at the input. LLaMA instead rotates each pair of query and key numbers by an angle that grows with position, inside every attention layer. Use the q, k and shift controls below to test it.
θⱼ = base^(−2j / dhead)RMSNorm
LayerNorm centres each token and scales it to unit spread. RMSNorm only rescales by the root mean square: simpler, slightly faster, and just as stable.
x / √(mean(x²) + ε)SwiGLU MLP
GPT-2's MLP widens, applies GELU and narrows. LLaMA's runs two projections side by side and lets one gate the other, a pattern called SwiGLU.
(SiLU(x·Wgate) ⊙ x·Wup) · WdownGrouped-query attention
While generating, past tokens’ keys and values are kept in a KV cache, which grows with every token and layer. GPT-2 gives every head its own keys and values; LLaMA 3 shares each key/value head among 4 query heads, so the cache is 4× smaller.
32 q heads · 8 kv heads
Code
# RMSNorm
x = x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + eps) * self.weight
# RoPE, inside attention: rotate each pair of q and k by position × θ
q, k = apply_rotary_pos_emb(q, k, cos, sin)
# GQA: 8 key/value heads serve 32 query heads
k, v = repeat_kv(k, 4), repeat_kv(v, 4)
# SwiGLU MLP
y = self.down_proj(F.silu(self.gate_proj(x)) * self.up_proj(x))
# the block, shaped like GPT-2's
h = x + self.self_attn(self.input_layernorm(x))
out = h + self.mlp(self.post_attention_layernorm(h))