Training · After pretraining
- before
- Qwen3-1.7B-Base
- after
- Qwen3-1.7B Qwen’s post-training
- example
- 43 tokens 20 trained
All steps
What the model sees
A model never sees chat bubbles. The chat template turns the conversation into one token sequence, with special tokens around each turn: <|im_start|>, the role, the text, <|im_end|>. SFT trains on many such sequences; this one has 43 tokens, 20 of them the answer.
43 tokens · 20 in the answerGraded on the answer only
Training reads the whole sequence and, at every position, asks the model for the next token, exactly as in pretraining. Only the answer’s guesses are graded: their −log p is added up into the loss, while the prompt and template are masked. Before any tuning, Qwen3-1.7B-Base pays 2.47 nats per answer token.
loss = mean −log p over the 20 answer tokensWhat it learns to do
What that training does, over many thousands of such examples: the same prompt, before and after. The base model treats the chat as a transcript to continue and writes the speaker label itself; the tuned model just answers. (Qwen’s real post-training used far more than SFT.)
base: continue the text · tuned: answer itSure of its own words
Training on answers also makes a model very sure of its own phrasing. On the answer written for this page, the tuned model gives most tokens a probability near 1, but where the page’s wording leaves its own it drops to nearly 0. Per token it pays 0.12 nats on its own answer and 4.12 on this one.
post-training sharpens the distributionTokens it never learned
Masked tokens get no gradient, so nothing holds the model’s guesses there in place. After post-training most of these tokens became more likely, but a few fell off a cliff: Qwen3 now gives “assistant” after <|im_start|> a probability below 10⁻¹⁵. It never has to predict it: the program writes the template.
no loss, no training signal
Code
ids = tokenizer.apply_chat_template([{'role': 'user', 'content': q}, {'role': 'assistant', 'content': a}])
labels = ids.clone()
labels[:n_prompt] = -100 # prompt and template: masked
logits = model(ids[:-1]).logits
loss = F.cross_entropy(logits, labels[1:], ignore_index=-100) # the answer only
loss.backward(); opt.step()