Training · After pretraining

SFT

instruction tuning: the loss, on answers only

before
Qwen3-1.7B-Base
after
Qwen3-1.7B Qwen’s post-training
example
43 tokens 20 trained

All steps

  1. What the model sees

    A model never sees chat bubbles. The chat template turns the conversation into one token sequence, with special tokens around each turn: <|im_start|>, the role, the text, <|im_end|>. SFT trains on many such sequences; this one has 43 tokens, 20 of them the answer.

    43 tokens · 20 in the answer

  2. Graded on the answer only

    Training reads the whole sequence and, at every position, asks the model for the next token, exactly as in pretraining. Only the answer’s guesses are graded: their −log p is added up into the loss, while the prompt and template are masked. Before any tuning, Qwen3-1.7B-Base pays 2.47 nats per answer token.

    loss = mean −log p over the 20 answer tokens

  3. What it learns to do

    What that training does, over many thousands of such examples: the same prompt, before and after. The base model treats the chat as a transcript to continue and writes the speaker label itself; the tuned model just answers. (Qwen’s real post-training used far more than SFT.)

    base: continue the text · tuned: answer it

  4. Sure of its own words

    Training on answers also makes a model very sure of its own phrasing. On the answer written for this page, the tuned model gives most tokens a probability near 1, but where the page’s wording leaves its own it drops to nearly 0. Per token it pays 0.12 nats on its own answer and 4.12 on this one.

    post-training sharpens the distribution

  5. Tokens it never learned

    Masked tokens get no gradient, so nothing holds the model’s guesses there in place. After post-training most of these tokens became more likely, but a few fell off a cliff: Qwen3 now gives “assistant” after <|im_start|> a probability below 10⁻¹⁵. It never has to predict it: the program writes the template.

    no loss, no training signal

Code

ids = tokenizer.apply_chat_template([{'role': 'user', 'content': q}, {'role': 'assistant', 'content': a}])
labels = ids.clone()
labels[:n_prompt] = -100                     # prompt and template: masked
logits = model(ids[:-1]).logits
loss = F.cross_entropy(logits, labels[1:], ignore_index=-100)   # the answer only
loss.backward(); opt.step()

Go deeper