Architectures · Compare

Architecture diff

any two models, read from their own checkpoints

compared
GPT-2 small vs DeepSeek-V3
layers
12 · 61
dmodel
768 · 7,168
params
124M · 671B
per token
124M · 37.6B
context
1,024 · 160K

All steps

  1. What differs

    GPT-2 small and DeepSeek-V3, row by row; the rows that differ are bright. Pick any two models above. Every number comes from the model’s own config and checkpoint on Hugging Face.

    9 of 9 rows differ

  2. Layer by layer

    One column per layer: its attention above, its MLP below. GPT-2 small: 12 full. DeepSeek-V3: 61 latent.

    12 layers · 61 layers

  3. Heads and the cache per token

    Query heads above, the key and value heads they read below. Grouped-query attention shares one key/value head among several query heads; latent attention caches one small vector per token instead; a linear layer keeps a fixed state.

    per token: 36 KiB · 68.6 KiB

  4. The MLP: dense or experts

    A dense MLP sends every token through one wide network. A mixture of experts stores many smaller ones and sends each token to a few, plus any shared expert every token uses. Lit cells are one token’s experts.

    width per token: 3,072 · 18,432

  5. Where the parameters are

    GPT-2 small stores 124M parameters and a token runs through 124M; DeepSeek-V3 stores 671B and runs 37.6B. Routed experts a token is not sent to (hatched) sit in memory unused.

    124M · 37.6B per token

  6. KV cache against context

    The cache for one sequence, in 16-bit, up to each model’s context length. At 1,024 tokens GPT-2 small holds 36 MiB and DeepSeek-V3 holds 68.6 MiB. Sliding-window and linear layers stop growing; full and latent layers grow with every token.

    Σ layers min(n, window) × per token, or a state

Code

# every number on this page, from the model’s own files on the Hub
cfg = json.load(open(hf_hub_download(repo, 'config.json')))
n = struct.unpack('<Q', get_range(url, 0, 7))[0]           # a safetensors header's length
header = json.loads(get_range(url, 8, 7 + n))              # every tensor: name, dtype, shape
params = sum(prod(t["shape"]) for name, t in header.items() if not is_scale(name))
idle = routed_experts * (1 - top_k / n_experts)           # stored, not used by this token
active = params - idle
per_token = 2 * kv_heads * head_dim                        # a full layer; MLA: latent + rope key
kv(n) = sum(min(n, window) * per_token or state for layer in layers)

Go deeper