Architectures · Compare
- compared
- GPT-2 small vs DeepSeek-V3
- layers
- 12 · 61
- dmodel
- 768 · 7,168
- params
- 124M · 671B
- per token
- 124M · 37.6B
- context
- 1,024 · 160K
All steps
What differs
GPT-2 small and DeepSeek-V3, row by row; the rows that differ are bright. Pick any two models above. Every number comes from the model’s own config and checkpoint on Hugging Face.
9 of 9 rows differLayer by layer
One column per layer: its attention above, its MLP below. GPT-2 small: 12 full. DeepSeek-V3: 61 latent.
12 layers · 61 layersHeads and the cache per token
Query heads above, the key and value heads they read below. Grouped-query attention shares one key/value head among several query heads; latent attention caches one small vector per token instead; a linear layer keeps a fixed state.
per token: 36 KiB · 68.6 KiBThe MLP: dense or experts
A dense MLP sends every token through one wide network. A mixture of experts stores many smaller ones and sends each token to a few, plus any shared expert every token uses. Lit cells are one token’s experts.
width per token: 3,072 · 18,432Where the parameters are
GPT-2 small stores 124M parameters and a token runs through 124M; DeepSeek-V3 stores 671B and runs 37.6B. Routed experts a token is not sent to (hatched) sit in memory unused.
124M · 37.6B per tokenKV cache against context
The cache for one sequence, in 16-bit, up to each model’s context length. At 1,024 tokens GPT-2 small holds 36 MiB and DeepSeek-V3 holds 68.6 MiB. Sliding-window and linear layers stop growing; full and latent layers grow with every token.
Σ layers min(n, window) × per token, or a state
Code
# every number on this page, from the model’s own files on the Hub
cfg = json.load(open(hf_hub_download(repo, 'config.json')))
n = struct.unpack('<Q', get_range(url, 0, 7))[0] # a safetensors header's length
header = json.loads(get_range(url, 8, 7 + n)) # every tensor: name, dtype, shape
params = sum(prod(t["shape"]) for name, t in header.items() if not is_scale(name))
idle = routed_experts * (1 - top_k / n_experts) # stored, not used by this token
active = params - idle
per_token = 2 * kv_heads * head_dim # a full layer; MLA: latent + rope key
kv(n) = sum(min(n, window) * per_token or state for layer in layers)Go deeper
- Ainslie et al. 2023, GQA: Training Generalized Multi-Query Transformer Models
- DeepSeek-AI 2024, DeepSeek-V2 (multi-head latent attention)
- Jiang et al. 2023, Mistral 7B (sliding-window attention)
- OpenAI 2025, gpt-oss-120b & gpt-oss-20b model card
- Hugging Face, reading a safetensors header (how these checkpoints were read)