Training
- corpus
- this site’s glossary 15,307 bytes
- merges
- 300 GPT-2 50,000
- start
- 256 byte symbols
All steps
Start from bytes
A tokenizer is learned from text before the model is. Here the text is this site’s glossary: 15,307 bytes in 3,216 pre-split words. Training starts with one symbol per byte: 256 in all, and nothing unknown.
15,307 bytes · vocabulary 256Count every adjacent pair
Count how often each pair of neighbouring symbols occurs, inside words only and weighted by how often each word appears. The most frequent pair here is Ġ + t, 426 times.
count pairs within wordsMerge the most frequent, repeat
Replace that pair everywhere with one new symbol, add it to the vocabulary, and count again. Each merge is a rule, kept in order. Watch the sample words fall from bytes into larger pieces as 300 merges are learned.
300 merges · vocabulary 556Vocabulary against length
More merges mean fewer tokens per word, so a context holds more text, but every new token needs a row in the embedding matrix and in the output layer. GPT-2 stopped at 50,000 merges; LLaMA 3 has about 128,000 tokens.
fewer tokens per word vs a bigger vocabularyGPT-2’s merges
GPT-2’s first merges, from its real merges.txt (learned on 40 GB of web text), start much like ours: a space joined to a common letter, “h e”, “i n”. Encoding a new word replays the merges in this order.
GPT-2: 50,000 merges, 50,257 tokens
Code
words = Counter(regex.findall(PAT, text)) # pre-split, counted
vocab = {w: tuple(w.encode()) for w in words} # every word as bytes
merges = []
for _ in range(num_merges):
pairs = Counter()
for w, n in words.items():
for a, b in zip(vocab[w], vocab[w][1:]): pairs[a, b] += n
best = max(pairs, key=pairs.get)
merges.append(best) # the rule, in order
vocab = {w: merge(s, best) for w, s in vocab.items()}