架构演进 · 编码器–解码器

T5

T5-small · 每个任务都是文本进、文本出

对比
T5-small vs GPT-2 small
层
6 编码 + 6 解码 GPT-2 12
dmodel
512 GPT-2 768
头
8 GPT-2 12
dff
2,048 · ReLU GPT-2 3,072 · GELU
词表
32,128 SentencePiece GPT-2 50,257
参数
60M GPT-2 124M

全部步骤

  1. 文本进,文本出

    T5 用同一种方式对待每个任务:文本进,文本出。一个短前缀指明任务,答案总是以文本形式生成,哪怕是一个标签(“negative”)或一个分数(“3.6”)。这些是 T5-small 的真实输出。

    前缀:输入 → 输出文本

  2. 片段破坏

    T5 用片段破坏来预训练:大约 15% 的词元以平均长度 3 的片段被丢弃,换成哨兵词元,解码器只写出被丢弃的内容,每个片段跟在它的哨兵后面。最后几行是 T5-small 的真实输出。

    带哨兵的输入 → 缺失的片段

  3. GPT-2 vs T5

    T5 和 2017 年的 Transformer 一样是编码器–解码器,但做了几处后来成为标准的改动:LayerNorm 放在每个子层之前,且只缩放(没有均值、没有 β:即 RMSNorm),任何地方都没有偏置项,位置是注意力分数上学习得到的偏置,而不是加上去的正弦波。点击标签跳过去。

    6 + 6 层 · dmodel 512 · 60M

  4. 相对位置桶

    T5 不为每个位置准备一个向量,而是按距离把每一对查询–键分到 32 个桶之一:7 以内每个距离一个桶,之后越来越宽,从 91 起全部共用最后一个桶。编码器为查询之前和之后的键分别设桶;解码器只往回看。

    32 个桶 · 8 以内精确 · 到 128 为对数间隔

  5. 学习得到的位置偏置

    每个桶在每个头上有一个学习得到的数,加到该桶中每一对的注意力分数上,并由所有层共享:这就是 T5 的相对位置偏置。这些是 T5-small 的真实偏置:有些编码器头偏向之前的词元,有些偏向之后的,还有一个强烈回避自己的位置。把鼠标停在格子上看看。

    score = q · k + b[bucket(j − i), head]

代码

# every task is a string in and a string out
ids = tok("translate English to German: I have seen the cat.").input_ids
out = model.generate(ids)          # "Ich habe die Katze gesehen."
# pretraining: drop spans, mark them with sentinels, write them back
inp = "Thank you <extra_id_0> me to your party <extra_id_1> week."
tgt = "<extra_id_0> for inviting <extra_id_1> last <extra_id_2>"
# the block: norm first, scale only (RMSNorm), no biases anywhere
x = x + self_attn(rms_norm(x), position_bias)
x = x + ffn(rms_norm(x))                        # ReLU in T5 v1.0
# one learned number per (bucket, head), added to every score
position_bias = rel_bias[bucket(key_pos - query_pos)]
scores = q @ k.transpose(-1, -2) + position_bias  # no ÷ √d in T5

延伸阅读