Architectures · Decoder-only
- compared
- DeepSeek-V3 vs GPT-2 small
- layers
- 61 GPT-2 12
- dmodel
- 7,168 GPT-2 768
- attention
- MLA · 128 heads GPT-2 MHA · 12
- experts
- 1 shared + 8 of 256 GPT-2 1 dense MLP
- params
- 671B · 37B active GPT-2 124M
- context
- 128K GPT-2 1,024
All steps
GPT-2 vs DeepSeek-V3
DeepSeek-V3 changes both halves of GPT-2’s block. Attention becomes multi-head latent attention (MLA), which caches one small latent per token instead of every head’s keys and values. The MLP becomes DeepSeekMoE: one shared expert plus 256 small routed experts, 8 per token. Click a label to jump there.
61 layers · 671B total · 37B per tokenMulti-head latent attention
MLA squeezes each token’s vector into a latent c of 512 numbers (2 here) and caches only that. Every head’s keys and values are rebuilt from c by up-projections. Position goes into a separate 64-number RoPE key shared by all heads, which is cached too. Hover the cells.
c = h · WDKV · K = c · WUK · V = c · WUVKV cache per token
Per token and layer, full multi-head attention would cache 2 × 128 heads × 128 numbers; MLA caches 512 + 64. Over 61 layers that is 3.8 MiB against 69 KiB per token, 57 times less, which is what makes a 128K-token context affordable.
(512 + 64) × 61 layers × 2 bytesShared and routed experts
DeepSeekMoE splits the experts into many small ones: 256 routed experts, 8 chosen per token, plus one shared expert every token uses. Scores are sigmoids, normalised over the chosen 8. More, smaller experts give far more combinations to specialise.
1 shared + top 8 of 256 routedBalancing with a bias
To keep experts evenly used without an extra loss term, each expert has a bias that is added to its score only when choosing experts. After each step, an overloaded expert’s bias goes down and an idle one’s goes up. Shown with 16 toy experts and 64 tokens.
choose by si + bi · weight by si
Code
# MLA: cache one small latent per token instead of every head's K and V
c_kv = kv_a_proj(h) # (tokens, 512) cached
k_rope = rope(k_rope_proj(h)) # (tokens, 64) cached, shared by all heads
k_nope, v = kv_b_proj(c_kv).split(...) # each head's keys and values, rebuilt
k = torch.cat([k_nope, k_rope.expand(heads)], dim=-1)
# DeepSeekMoE: one shared expert plus the top 8 of 256 routed experts
s = torch.sigmoid(gate(x)) # (tokens, 256) affinities
idx = torch.topk(s + bias, 8).indices # the bias only affects the choice
w = s.gather(-1, idx); w = w / w.sum(-1, keepdim=True)
y = shared_expert(x) + sum(w[..., j] * experts[idx[..., j]](x) for j in range(8))
# after each step: nudge each expert's bias toward an even load
bias += gamma * torch.sign(load.mean() - load)