Architectures · Origin

Transformer (2017)

Vaswani et al., “Attention Is All You Need” · what GPT-2 changed

compared
base 2017 vs GPT-2 small
layers
6 enc + 6 dec GPT-2 12
dmodel
512 GPT-2 768
heads
8 GPT-2 12
dff
2,048 GPT-2 3,072
vocab
~37,000 shared GPT-2 50,257
params
65M GPT-2 124M

All steps

  1. Before: one step at a time

    Before 2017, translation models were recurrent neural networks (RNNs) that read a sentence one token at a time: each step needs the one before, so a GPU cannot run them side by side. Self-attention links every pair of positions in one matrix product.

    RNN: n steps · attention: 1

  2. GPT-2 vs the 2017 Transformer

    GPT-2 is one stack. The 2017 Transformer has two: an encoder that reads the source sentence and a decoder that writes the translation, reading the encoder’s output through cross-attention. It also normalises after each residual add, uses fixed sinusoids for positions and ReLU in the MLP. Click a label to jump to that change.

    6 encoder + 6 decoder layers · dmodel 512

  3. Translating a sentence

    The encoder reads the English sentence once. The decoder then writes German one token at a time, like GPT-2, and each pass reads the encoder’s output. In training, the correct translation is fed in (teacher forcing) and the causal mask hides the future, so all positions run in one pass.

    encoder once · decoder once per token

  4. Three kinds of attention

    There are three attentions, and they all compute softmax(Q·Kᵀ / √dk) · V. The encoder sees the whole source, the decoder sees only earlier target tokens, and cross-attention lets each target token see the whole source. GPT-2 has only the middle kind.

    source × source · target × target · target × source

  5. Cross-attention

    In cross-attention the queries come from the decoder and the keys and values from the encoder output, so the score matrix is target × source and is not square. Each row looks for the source word it needs next: ‘gesehen’ goes back to ‘seen’, although German moves it to the end. Hover the cells.

    Q 7 × dk · Kᵀ dk × 6 → 7 × 6

  6. Post-LN vs pre-LN

    The 2017 model normalises after each residual add (post-LN), which puts LayerNorm on the stream’s main path; deep post-LN models need a long learning-rate warmup (4,000 steps in the paper). GPT-2 normalises the copy each sub-layer reads (pre-LN), so the main path only adds.

    LN(x + f(x)) → x + f(LN(x))

  7. Sinusoidal positions

    Attention ignores order, so both models add a position vector to each token. The 2017 model computes it from sine and cosine waves, with no parameters and for any position; GPT-2 learns a table of 1,024 rows. Moving k positions turns each sin/cos pair by a fixed angle.

    PE(pos, 2i) = sin(pos / 10000^(2i/dmodel))

Code

# positions: fixed sine and cosine waves, added to embeddings scaled by √d_model
pe[:, 0::2] = torch.sin(pos * 10000 ** (-i2 / d_model))
pe[:, 1::2] = torch.cos(pos * 10000 ** (-i2 / d_model))
x = embed(src) * math.sqrt(d_model) + pe[:len(src)]
# encoder layer, 6 times: LayerNorm after each residual add (post-LN)
x = norm1(x + self_attn(x, x, x))                     # no mask
x = norm2(x + ffn(x))                                 # ReLU
# decoder layer, 6 times
y = norm1(y + self_attn(y, y, y, mask=causal))
y = norm2(y + cross_attn(q=y, k=memory, v=memory))    # memory = encoder output
y = norm3(y + ffn(y))
# GPT-2, for comparison: pre-LN
h = x + attn(ln_1(x)); out = h + mlp(ln_2(h))

Go deeper