Architectures · Decoder-only

DeepSeek

DeepSeek-V3 · latent attention and fine-grained experts

compared
DeepSeek-V3 vs GPT-2 small
layers
61 GPT-2 12
dmodel
7,168 GPT-2 768
attention
MLA · 128 heads GPT-2 MHA · 12
experts
1 shared + 8 of 256 GPT-2 1 dense MLP
params
671B · 37B active GPT-2 124M
context
128K GPT-2 1,024

All steps

  1. GPT-2 vs DeepSeek-V3

    DeepSeek-V3 changes both halves of GPT-2’s block. Attention becomes multi-head latent attention (MLA), which caches one small latent per token instead of every head’s keys and values. The MLP becomes DeepSeekMoE: one shared expert plus 256 small routed experts, 8 per token. Click a label to jump there.

    61 layers · 671B total · 37B per token

  2. Multi-head latent attention

    MLA squeezes each token’s vector into a latent c of 512 numbers (2 here) and caches only that. Every head’s keys and values are rebuilt from c by up-projections. Position goes into a separate 64-number RoPE key shared by all heads, which is cached too. Hover the cells.

    c = h · WDKV · K = c · WUK · V = c · WUV

  3. KV cache per token

    Per token and layer, full multi-head attention would cache 2 × 128 heads × 128 numbers; MLA caches 512 + 64. Over 61 layers that is 3.8 MiB against 69 KiB per token, 57 times less, which is what makes a 128K-token context affordable.

    (512 + 64) × 61 layers × 2 bytes

  4. Shared and routed experts

    DeepSeekMoE splits the experts into many small ones: 256 routed experts, 8 chosen per token, plus one shared expert every token uses. Scores are sigmoids, normalised over the chosen 8. More, smaller experts give far more combinations to specialise.

    1 shared + top 8 of 256 routed

  5. Balancing with a bias

    To keep experts evenly used without an extra loss term, each expert has a bias that is added to its score only when choosing experts. After each step, an overloaded expert’s bias goes down and an idle one’s goes up. Shown with 16 toy experts and 64 tokens.

    choose by si + bi · weight by si

Code

# MLA: cache one small latent per token instead of every head's K and V
c_kv = kv_a_proj(h)                          # (tokens, 512)  cached
k_rope = rope(k_rope_proj(h))                # (tokens, 64)   cached, shared by all heads
k_nope, v = kv_b_proj(c_kv).split(...)       # each head's keys and values, rebuilt
k = torch.cat([k_nope, k_rope.expand(heads)], dim=-1)
# DeepSeekMoE: one shared expert plus the top 8 of 256 routed experts
s = torch.sigmoid(gate(x))                   # (tokens, 256) affinities
idx = torch.topk(s + bias, 8).indices        # the bias only affects the choice
w = s.gather(-1, idx); w = w / w.sum(-1, keepdim=True)
y = shared_expert(x) + sum(w[..., j] * experts[idx[..., j]](x) for j in range(8))
# after each step: nudge each expert's bias toward an even load
bias += gamma * torch.sign(load.mean() - load)

Go deeper