Architectures · Encoder–decoder

T5

T5-small · every task as text in, text out

compared
T5-small vs GPT-2 small
layers
6 enc + 6 dec GPT-2 12
dmodel
512 GPT-2 768
heads
8 GPT-2 12
dff
2,048 · ReLU GPT-2 3,072 · GELU
vocab
32,128 SentencePiece GPT-2 50,257
params
60M GPT-2 124M

All steps

  1. Text in, text out

    T5 treats every task the same way: text in, text out. A short prefix names the task, and the answer is always generated as text, even a label (“negative”) or a score (“3.6”). These are real outputs of T5-small.

    prefix: input → output text

  2. Span corruption

    T5 is pretrained by span corruption: about 15% of the tokens, in spans of 3 on average, are dropped and replaced by sentinel tokens, and the decoder writes out only what was dropped, each span after its sentinel. The last lines are T5-small’s real output.

    input with sentinels → the missing spans

  3. GPT-2 vs T5

    T5 is an encoder–decoder like the 2017 Transformer, with changes that became standard: LayerNorm before each sub-layer and scale-only (no mean, no β: an RMSNorm), no bias terms anywhere, and position as a learned bias on the attention scores instead of added sinusoids. Click a label to jump.

    6 + 6 layers · dmodel 512 · 60M

  4. Relative position buckets

    Instead of a vector per position, T5 sorts each query–key pair by distance into one of 32 buckets: one per distance up to 7, then wider and wider, and everything from 91 on shares the last. The encoder keeps separate buckets for keys before and after the query; the decoder only looks back.

    32 buckets · exact below 8 · log-spaced to 128

  5. Learned position biases

    Each bucket has one learned number per head, added to the attention score of every pair in that bucket and shared by all layers: T5’s relative position bias. These are T5-small’s real biases: some encoder heads lean toward earlier tokens, some toward later ones, and one strongly avoids its own position. Hover the cells.

    score = q · k + b[bucket(j − i), head]

Code

# every task is a string in and a string out
ids = tok("translate English to German: I have seen the cat.").input_ids
out = model.generate(ids)          # "Ich habe die Katze gesehen."
# pretraining: drop spans, mark them with sentinels, write them back
inp = "Thank you <extra_id_0> me to your party <extra_id_1> week."
tgt = "<extra_id_0> for inviting <extra_id_1> last <extra_id_2>"
# the block: norm first, scale only (RMSNorm), no biases anywhere
x = x + self_attn(rms_norm(x), position_bias)
x = x + ffn(rms_norm(x))                        # ReLU in T5 v1.0
# one learned number per (bucket, head), added to every score
position_bias = rel_bias[bucket(key_pos - query_pos)]
scores = q @ k.transpose(-1, -2) + position_bias  # no ÷ √d in T5

Go deeper