Architectures · Encoder–decoder
- compared
- T5-small vs GPT-2 small
- layers
- 6 enc + 6 dec GPT-2 12
- dmodel
- 512 GPT-2 768
- heads
- 8 GPT-2 12
- dff
- 2,048 · ReLU GPT-2 3,072 · GELU
- vocab
- 32,128 SentencePiece GPT-2 50,257
- params
- 60M GPT-2 124M
All steps
Text in, text out
T5 treats every task the same way: text in, text out. A short prefix names the task, and the answer is always generated as text, even a label (“negative”) or a score (“3.6”). These are real outputs of T5-small.
prefix: input → output textSpan corruption
T5 is pretrained by span corruption: about 15% of the tokens, in spans of 3 on average, are dropped and replaced by sentinel tokens, and the decoder writes out only what was dropped, each span after its sentinel. The last lines are T5-small’s real output.
input with sentinels → the missing spansGPT-2 vs T5
T5 is an encoder–decoder like the 2017 Transformer, with changes that became standard: LayerNorm before each sub-layer and scale-only (no mean, no β: an RMSNorm), no bias terms anywhere, and position as a learned bias on the attention scores instead of added sinusoids. Click a label to jump.
6 + 6 layers · dmodel 512 · 60MRelative position buckets
Instead of a vector per position, T5 sorts each query–key pair by distance into one of 32 buckets: one per distance up to 7, then wider and wider, and everything from 91 on shares the last. The encoder keeps separate buckets for keys before and after the query; the decoder only looks back.
32 buckets · exact below 8 · log-spaced to 128Learned position biases
Each bucket has one learned number per head, added to the attention score of every pair in that bucket and shared by all layers: T5’s relative position bias. These are T5-small’s real biases: some encoder heads lean toward earlier tokens, some toward later ones, and one strongly avoids its own position. Hover the cells.
score = q · k + b[bucket(j − i), head]
Code
# every task is a string in and a string out
ids = tok("translate English to German: I have seen the cat.").input_ids
out = model.generate(ids) # "Ich habe die Katze gesehen."
# pretraining: drop spans, mark them with sentinels, write them back
inp = "Thank you <extra_id_0> me to your party <extra_id_1> week."
tgt = "<extra_id_0> for inviting <extra_id_1> last <extra_id_2>"
# the block: norm first, scale only (RMSNorm), no biases anywhere
x = x + self_attn(rms_norm(x), position_bias)
x = x + ffn(rms_norm(x)) # ReLU in T5 v1.0
# one learned number per (bucket, head), added to every score
position_bias = rel_bias[bucket(key_pos - query_pos)]
scores = q @ k.transpose(-1, -2) + position_bias # no ÷ √d in T5