Agents · Prompting
- model
- GPT-2 small weights fixed
- task
- country → capital
- examples
- 0 to 3
- copying head
- layer 9, head 7
- instruction-tuned
- Qwen3-1.7B
All steps
No examples
GPT-2 small is given “Egypt:” and nothing else. Nothing says what should follow, so its guesses are generic: a newline, “The”. “ Cairo” gets <0.1%.
p( Cairo | “Egypt:” ) = <0.1%Examples set the task
The same model, with examples of country: capital lines before the question. One example is enough: “ Cairo” jumps to 58%. No weight changes; the examples only sit in the context. Over 12 countries, one example takes the right answer to first place 10 times out of 12.
examples in the context, weights fixedThe prompt is the program
Keep the question, change the examples, and the model computes a different function: capital, language or continent. Each is GPT-2’s real top guess after two examples. This is in-context learning: the prompt acts as the program.
same weights · same question · three tasksA head that copies
Part of how: in layer 9, head 7, the last “:” puts 97% of its attention on the words that followed the earlier colons, the answers. It looks for what came after this token before and copies from there. Heads that do this are called induction heads.
layer 9, head 7: attention from the last “:”Examples also mislead
Copying also brings bias. With two “Europe” labels among three examples, GPT-2 says “ Europe” for Egypt; “ Africa” drops to 11%. With capitals, the last example (Rome) pulls in “ Rome” and “ Milan”. More examples do not help here: 10, 9, 9, 8 of 12 right with 1 to 4.
majority and recency biasInstructions, not examples
Given only the instruction, GPT-2 keeps writing text like its training data. Qwen3-1.7B was instruction-tuned: trained further on instructions and answers in a chat template. It answers “Cairo” with no examples. Examples still help to pin down a format, at 12 tokens for three here, paid on every call.
the instruction is the program
Code
examples = [('France', 'Paris'), ('Japan', 'Tokyo'), ('Italy', 'Rome')]
prompt = ''.join(f'{c}: {a}\n' for c, a in examples) + 'Egypt:'
ids = tokenizer(prompt, return_tensors="pt").input_ids # no training: the weights stay as they are
probs = model(ids).logits[0, -1].softmax(-1) # the next token after "Egypt:"
probs[tokenizer(' Cairo').input_ids[0]] # 0.1% with no examples, 58% with one
# an instruction-tuned model reads its chat template instead of examples
text = tokenizer.apply_chat_template([{'role': 'user', 'content': question}],
add_generation_prompt=True, enable_thinking=False, tokenize=False)Go deeper
- Brown et al. 2020, Language Models are Few-Shot Learners
- Olsson et al. 2022, In-context Learning and Induction Heads
- Zhao et al. 2021, Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Ouyang et al. 2022, Training language models to follow instructions (InstructGPT)
- Qwen Team 2025, Qwen3 Technical Report