Start here
What is a Transformer?
The model on this site is GPT-2 small (OpenAI, 2019, 124M parameters), a Transformer. It cuts text into tokens, turns each token into a vector of 768 numbers, and passes the vectors through 12 blocks. In each block, attention lets every token read from the tokens before it, and an MLP then works on each token alone. The last token's vector is finally turned into a probability for every possible next token.
Decoder-only means it reads left to right: a token never sees the ones after it. Generating text is just this, repeated: pick a token, append it, run again. LLaMA and most chat models keep the same design with a few changes (see Architectures).
Where the 124M numbers live
- MLPs, 12 × 4.72M 56.7M 46%
- token embeddings WE (also the output) 38.6M 31%
- attention, 12 × 2.36M 28.3M 23%
- positions WP 786K 0.63%
- LayerNorms 38K 0.03%
Running it costs about 2 floating-point operations per weight per token, roughly 250M FLOPs for each new token, plus attention’s share, which grows with the context. Bigger models keep the same parts, only wider and deeper: GPT-2 XL has 1.5B weights, LLaMA 3 has 8B and 70B.
The path
How to read the pictures
- Each position keeps its own colour on every page (colour means position, not word; after seven positions the colours repeat).
- A lane is one token’s vector flowing through the model; its colour blends as it takes in other tokens.
- A glass plate is a layer the lanes pass through.
- In a matrix, a filled cell is positive, an outlined cell is negative, and a hatched cell is masked out. In the Serving pictures of memory, hatching marks space that is reserved but empty.
- In a matrix product C = A · B, row i of A (left) meets column j of B (above) at cell (i, j) of C. The drawings fill cells one at a time so you can follow them; the hardware computes every cell at once.
- ⊕ adds a layer’s output back onto the lane (the residual stream).
- Parts marked ↗ open a detail view. Hover or tap any matrix cell to see its formula; click to pin it.
Space play or pause · ← → previous or next step. Each step pauses at its end so there is time to read; switch to Auto in the controls to play straight through.
Where do the numbers come from?
Every weight (124M of them) was set by training: GPT-2 read about 40 GB of web text, predicted each next token, and after every batch nudged all its weights so the actual next token got a little more probability. Nothing in the model was written by hand. Even the tokenizer was learned, by counting which pairs of symbols appear together most often. The Next-token loss page shows the signal it learns from.
Real or toy?
The Forward pass and Unembed pages show the numbers of a real GPT-2 small run, computed offline. The detail views that animate every matrix product use a toy model (8 numbers per token instead of 768) so each cell fits on screen; they say shown: toy and give GPT-2's real sizes alongside.