Architectures · Encoder

BERT

BERT-base · the same shape as GPT-2, reading both ways

compared
BERT-base vs GPT-2 small
layers
12 GPT-2 12
dmodel
768 GPT-2 768
heads
12 GPT-2 12
dff
3,072 GPT-2 3,072
vocab
30,522 WordPiece GPT-2 50,257
params
110M GPT-2 124M

All steps

  1. Same shape, both directions

    BERT-base has exactly GPT-2 small’s shape: 12 layers, 768 wide, 12 heads. What differs is the direction: there is no causal mask, so every token sees the whole sentence, and it is trained to fill in hidden words instead of predicting the next one. Like the 2017 Transformer, it normalises after each residual add.

    12 × 768 · 110M · encoder-only

  2. Looking both ways

    Two real heads. GPT-2’s clearest previous-token head can only look back. BERT’s layer 3 head 1 is its mirror image: nearly all of its attention goes to the next token, which a causal mask would forbid. Hover the cells.

    causal triangle · full square

  3. What goes in

    BERT reads one or two sentences at a time: [CLS] first and [SEP] after each sentence. Each input vector is the sum of a WordPiece token embedding, a segment embedding (sentence A or B) and a learned position embedding. WordPiece splits rare words: mailman → mail ##man.

    token + segment + position

  4. Fill in the blank

    In training, 15% of the tokens are chosen and hidden (mostly as [MASK]), and BERT predicts them from both sides: masked language modelling. These are real predictions from BERT-base, next to GPT-2’s guesses from only the words before the blank. Use the arrows below to switch sentences.

    predict [MASK] from both sides

  5. Reading, not writing

    BERT is built to read, not to write: with no left-to-right order it has no natural way to generate text. Instead a small classifier is trained on top, on the final [CLS] vector for a label per sentence (such as sentiment) or on each token’s vector for tagging.

    label = softmax(hCLS · W)

Code

# BERT: GPT-2 small's shape, no causal mask, LayerNorm after each add
x = layer_norm(tok_emb[ids] + seg_emb[segments] + pos_emb[positions])
for layer in layers:                          # 12 × (768 wide, 12 heads, 3,072)
    x = layer_norm(x + self_attn(x))            # every token sees every token
    x = layer_norm(x + mlp(x))
# pretraining: predict hidden tokens from both sides
logits = mlm_head(x[masked_positions])        # 30,522 WordPiece scores
loss = F.cross_entropy(logits, original_ids)
# using it: a classifier on the [CLS] vector
probs = F.softmax(x[:, 0] @ W_cls, dim=-1)

Go deeper