智能体 · 提示
- 模型
- GPT-2 small 权重固定
- 任务
- 国家 → 首都
- 示例
- 0 到 3
- 复制头
- 第 9 层,头 7
- 指令微调
- Qwen3-1.7B
全部步骤
没有示例
GPT-2 small 只拿到 “Egypt:”,别无其他。没有任何东西说明后面该接什么,所以它的猜测很泛:换行,“The”。“ Cairo” 只得到 <0.1%。
p( Cairo | “Egypt:” ) = <0.1%示例决定任务
同一个模型,在问题前加上「country: capital」形式的示例。一个示例就够了:“ Cairo” 跃升到 58%。权重没有变化;示例只是放在上下文里。在 12 个国家上,一个示例让正确答案有 10 次(共 12 次)排到第一。
示例在上下文中,权重固定提示词就是程序
问题不变,换掉示例,模型就计算出不同的函数:首都、语言或大洲。每一个都是 GPT-2 看过两个示例后的真实首选。这就是上下文学习:提示词充当了程序。
同样的权重 · 同样的问题 · 三个任务会复制的头
部分原理:在第 9 层、头 7,最后一个 “:” 把 97% 的注意力放在之前冒号后面的词上,也就是答案上。它寻找这个词元之前出现时后面接的内容,并从那里复制。做这件事的头叫归纳头。
第 9 层,头 7:来自最后一个 “:” 的注意力示例也会误导
复制也带来偏向。三个示例中有两个 “Europe” 标签时,GPT-2 给 Egypt 的答案是 “ Europe”;“ Africa” 降到 11%。在首都任务里,最后一个示例(Rome)把 “ Rome” 和 “ Milan” 拉了进来。更多示例在这里也没用:用 1 到 4 个示例时,12 个中分别对了 10、9、9、8 个。
多数偏向与近因偏向用指令,而不是示例
只给指令时,GPT-2 会继续写出像它训练数据那样的文本。Qwen3-1.7B 经过了指令微调:在对话模板中的指令和回答上进一步训练过。它不需要任何示例就回答 “Cairo”。示例仍有助于确定格式,这里三个示例要 12 个词元,每次调用都要付出。
指令就是程序
代码
examples = [('France', 'Paris'), ('Japan', 'Tokyo'), ('Italy', 'Rome')]
prompt = ''.join(f'{c}: {a}\n' for c, a in examples) + 'Egypt:'
ids = tokenizer(prompt, return_tensors="pt").input_ids # no training: the weights stay as they are
probs = model(ids).logits[0, -1].softmax(-1) # the next token after "Egypt:"
probs[tokenizer(' Cairo').input_ids[0]] # 0.1% with no examples, 58% with one
# an instruction-tuned model reads its chat template instead of examples
text = tokenizer.apply_chat_template([{'role': 'user', 'content': question}],
add_generation_prompt=True, enable_thinking=False, tokenize=False)延伸阅读
- Brown et al. 2020, Language Models are Few-Shot Learners
- Olsson et al. 2022, In-context Learning and Induction Heads
- Zhao et al. 2021, Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Ouyang et al. 2022, Training language models to follow instructions (InstructGPT)
- Qwen Team 2025, Qwen3 Technical Report