Architectures · Encoder
- compared
- BERT-base vs GPT-2 small
- layers
- 12 GPT-2 12
- dmodel
- 768 GPT-2 768
- heads
- 12 GPT-2 12
- dff
- 3,072 GPT-2 3,072
- vocab
- 30,522 WordPiece GPT-2 50,257
- params
- 110M GPT-2 124M
All steps
Same shape, both directions
BERT-base has exactly GPT-2 small’s shape: 12 layers, 768 wide, 12 heads. What differs is the direction: there is no causal mask, so every token sees the whole sentence, and it is trained to fill in hidden words instead of predicting the next one. Like the 2017 Transformer, it normalises after each residual add.
12 × 768 · 110M · encoder-onlyLooking both ways
Two real heads. GPT-2’s clearest previous-token head can only look back. BERT’s layer 3 head 1 is its mirror image: nearly all of its attention goes to the next token, which a causal mask would forbid. Hover the cells.
causal triangle · full squareWhat goes in
BERT reads one or two sentences at a time: [CLS] first and [SEP] after each sentence. Each input vector is the sum of a WordPiece token embedding, a segment embedding (sentence A or B) and a learned position embedding. WordPiece splits rare words: mailman → mail ##man.
token + segment + positionFill in the blank
In training, 15% of the tokens are chosen and hidden (mostly as [MASK]), and BERT predicts them from both sides: masked language modelling. These are real predictions from BERT-base, next to GPT-2’s guesses from only the words before the blank. Use the arrows below to switch sentences.
predict [MASK] from both sidesReading, not writing
BERT is built to read, not to write: with no left-to-right order it has no natural way to generate text. Instead a small classifier is trained on top, on the final [CLS] vector for a label per sentence (such as sentiment) or on each token’s vector for tagging.
label = softmax(hCLS · W)
Code
# BERT: GPT-2 small's shape, no causal mask, LayerNorm after each add
x = layer_norm(tok_emb[ids] + seg_emb[segments] + pos_emb[positions])
for layer in layers: # 12 × (768 wide, 12 heads, 3,072)
x = layer_norm(x + self_attn(x)) # every token sees every token
x = layer_norm(x + mlp(x))
# pretraining: predict hidden tokens from both sides
logits = mlm_head(x[masked_positions]) # 30,522 WordPiece scores
loss = F.cross_entropy(logits, original_ids)
# using it: a classifier on the [CLS] vector
probs = F.softmax(x[:, 0] @ W_cls, dim=-1)