Training

Learning the tokenizer

byte-pair encoding: count pairs, merge, repeat

corpus
this site’s glossary 15,307 bytes
merges
300 GPT-2 50,000
start
256 byte symbols

All steps

  1. Start from bytes

    A tokenizer is learned from text before the model is. Here the text is this site’s glossary: 15,307 bytes in 3,216 pre-split words. Training starts with one symbol per byte: 256 in all, and nothing unknown.

    15,307 bytes · vocabulary 256

  2. Count every adjacent pair

    Count how often each pair of neighbouring symbols occurs, inside words only and weighted by how often each word appears. The most frequent pair here is Ġ + t, 426 times.

    count pairs within words

  3. Merge the most frequent, repeat

    Replace that pair everywhere with one new symbol, add it to the vocabulary, and count again. Each merge is a rule, kept in order. Watch the sample words fall from bytes into larger pieces as 300 merges are learned.

    300 merges · vocabulary 556

  4. Vocabulary against length

    More merges mean fewer tokens per word, so a context holds more text, but every new token needs a row in the embedding matrix and in the output layer. GPT-2 stopped at 50,000 merges; LLaMA 3 has about 128,000 tokens.

    fewer tokens per word vs a bigger vocabulary

  5. GPT-2’s merges

    GPT-2’s first merges, from its real merges.txt (learned on 40 GB of web text), start much like ours: a space joined to a common letter, “h e”, “i n”. Encoding a new word replays the merges in this order.

    GPT-2: 50,000 merges, 50,257 tokens

Code

words = Counter(regex.findall(PAT, text))              # pre-split, counted
vocab = {w: tuple(w.encode()) for w in words}          # every word as bytes
merges = []
for _ in range(num_merges):
    pairs = Counter()
    for w, n in words.items():
        for a, b in zip(vocab[w], vocab[w][1:]): pairs[a, b] += n
    best = max(pairs, key=pairs.get)
    merges.append(best)                                # the rule, in order
    vocab = {w: merge(s, best) for w, s in vocab.items()}

Go deeper