Reference

Glossary

terms of art, in plain words

Active parameters
The weights one token actually runs through. In a mixture of experts that is far fewer than the model stores: the experts a token is not sent to sit in memory unused.
adaLN-Zero
DiT’s conditioning: the timestep and class set each block’s LayerNorm scale and shift and a gate on each sub-layer’s output; the gates start at zero, so blocks start as the identity.
Adam
The standard optimizer for Transformers: it keeps a running mean of each weight’s gradient (m) and of its square (v), and steps by m / √v, so every weight moves at a similar pace. AdamW adds weight decay separately.
Agent
A program that runs a language model in a loop: the model chooses actions such as tool calls, the program carries them out and feeds the results back, until the task is done.
Attention
Lets each position mix in information from itself and earlier positions, weighted by how well its query matches their keys.
Attention head
One of several attention computations run side by side, each on its own slice of the vector (64 of 768 numbers in GPT-2), free to track a different relation.
Attention sink
Many heads park spare attention on the first token, whose value adds almost nothing; the overview colours treat that attention as a no-op. gpt-oss gives each head a learned sink instead: an extra score that can take attention without any token.
Backpropagation
Computing the gradient of the loss for every weight at once, by applying the chain rule layer by layer from the output back to the input.
Base model
A model after pretraining only: it continues text in the style of its training data and has not been tuned to follow instructions.
BatchNorm
Normalises each feature across the examples in a batch (common in vision). Transformers use LayerNorm instead, which normalises each token on its own.
BPE
Byte-pair encoding: a tokenizer that starts from bytes and repeatedly merges the most frequent adjacent pair seen in training text. GPT-2 learned 50,000 merges this way.
Causal mask
Hides later positions from each position, so the model cannot peek at the tokens it is learning to predict.
Chain rule
If y depends on x and the loss depends on y, the loss’s slope with respect to x is its slope with respect to y times y’s slope with respect to x.
Chat template
The text format a chat model was tuned on: special tokens such as <|im_start|> open each message and name who wrote it, and the model’s turn begins where the template stops.
Chunked prefill
Splitting a long prompt’s prefill into pieces that ride along with other requests’ decode steps, so no single step becomes slow.
Classifier-free guidance
At sampling time, running a diffusion model with and without the condition and moving further in the direction the condition adds, for images that match it more strongly.
Compute-bound and memory-bound
Whether a GPU step is limited by how fast it can do arithmetic or by how fast it can read memory; a step doing few operations per byte read (like decoding one token) waits on memory.
Compute-optimal
The split of a fixed training budget between model size and training tokens that gives the lowest loss; about 20 tokens per parameter by the Chinchilla fit.
Constrained decoding
Masking, at each step, every token that would break a required format (a JSON grammar, a list of tool names) before the next token is chosen, so the output always fits the format.
Continuous batching
Rescheduling a serving batch after every decode step: finished requests leave at once and waiting ones take their place, instead of waiting for the whole batch to finish.
Contrastive learning
Training by comparison: matching pairs (an image and its caption) are pulled together and every mismatched pair in the batch is pushed apart.
Copy-on-write
Sharing memory between users until one of them writes to it; only then is a private copy made.
Cosine similarity
How closely two vectors point the same way: their dot product divided by both lengths, from −1 to 1. For vectors of length 1 it is just the dot product.
Cross-attention
Attention whose queries come from one sequence and keys and values from another: in a translation model, the decoder (target) reads the encoder’s output (source).
Decoder-only
A Transformer that reads left to right under a causal mask and predicts the next token. GPT-2 and LLaMA are decoder-only.
Diffusion model
A model trained to remove noise from data; generating means starting from pure noise and denoising step by step.
DPO
Direct preference optimization: tuning a model on pairs of preferred and rejected answers with one loss that raises the preferred answer’s probability relative to a frozen reference model and lowers the other’s.
Embedding
A learned vector of numbers (768 in GPT-2) that stands for a token or a position; tokens used in similar ways end up with similar vectors.
Embedding model
A model that turns a whole text into one vector, trained so that texts with related meanings get vectors pointing in similar directions.
Encoder and decoder
In the 2017 Transformer, the encoder reads the whole source sentence at once; the decoder writes the output one token at a time, under a causal mask, reading the encoder’s output through cross-attention.
Encoder-only
A Transformer with no causal mask that reads a whole text at once and outputs a vector per token, for understanding rather than generating text. BERT is encoder-only.
Few-shot prompting
Putting a few worked examples of a task in the prompt before the real question, so the model continues the pattern.
Fine-tuning
Training a pretrained model a little further, on a smaller dataset for one task, usually with a small new output layer on top.
Finite difference
Estimating a slope by moving an input a tiny step each way and dividing the change in output by the step; slow, but a good check.
FlashAttention
An exact attention kernel that works through tiles of queries, keys and values in the GPU’s on-chip SRAM, with an online softmax, so the N × N score matrix is never written to HBM.
FLOPs
Floating-point operations. One multiply-add counts as two.
GELU
A smooth nonlinearity: large positive inputs pass almost unchanged, small ones are damped and negative ones are pushed close to 0.
GEMM
General matrix multiply, C = A · B: the operation almost all of a Transformer’s compute goes into.
GQA
Grouped-query attention: several query heads share one key/value head, which shrinks the KV cache.
Gradient
For every weight, how fast the loss changes as that weight changes. Training moves each weight a little against its gradient.
Greedy decoding
Always picking the most likely next token.
Handoff
Passing a piece of work, or the whole conversation, from one agent to another, usually by a tool call; the receiving agent sees only what is passed to it.
In-context learning
A model doing a task it is shown in its prompt, by examples or instructions, with no change to its weights.
Induction head
A head that finds an earlier copy of the current token and attends to the token that came after it, so a repeated pattern can be continued.
Instruction tuning
Fine-tuning a pretrained model on instructions paired with good answers, written in a chat template, so that it answers requests instead of continuing the text.
JSON Schema
A standard way to describe the shape of JSON data: its fields, their types and which are required. Tool definitions use it for their parameters.
KV cache
While generating, the keys and values of past tokens are stored so each new token only computes its own. It grows with every token and every layer.
LayerNorm
Rescales each token’s vector to mean 0 and spread 1, then applies a learned scale (γ) and shift (β) per feature.
Learning rate
How far each training step moves the weights against the gradient (η). It is usually warmed up from zero, then decayed.
Linear attention
An attention layer that folds the past into a fixed-size state instead of keeping every key and value, so its memory does not grow with the text; Qwen3-Next and Kimi Linear mix it with full attention.
Logit lens
Reading a middle layer through the final LayerNorm and the unembedding, to see what the model would predict if it stopped there.
Logits
Raw, unnormalised scores, one per vocabulary token; softmax turns them into probabilities.
LoRA
Low-rank adaptation: fine-tuning by keeping the weights frozen and learning a small change to them as the product of two thin matrices.
Masked language modelling
BERT’s training task: some tokens are hidden and the model predicts them from the words on both sides.
Mixture of experts (MoE)
A layer with several expert MLPs and a router; each token runs through only the few experts the router picks, so the model can store many more parameters than it uses per token.
MLP
Multi-layer perceptron: two matrix products with a nonlinearity (GELU) between them, applied to each token on its own.
Momentum
Stepping along a running sum of past gradients instead of the latest one alone, so oscillations cancel and consistent directions build up speed.
Multi-head latent attention (MLA)
DeepSeek’s attention: each token’s keys and values are compressed into one small latent vector, which is all the KV cache stores; every head’s keys and values are rebuilt from it.
Multi-token prediction
Training a model to also predict tokens further ahead (the one after next), as an extra, denser training signal.
Online softmax
Computing a softmax in chunks with a running maximum and a running sum, rescaling earlier terms when the maximum grows; the result equals the ordinary softmax.
Orchestrator
In a multi-agent system, the agent that splits a task, hands the pieces to worker agents and combines what they report.
PagedAttention
vLLM’s way of storing the KV cache: in fixed-size blocks handed out on demand and found through a per-request block table, like virtual-memory pages.
Patch embedding
ViT’s replacement for a vocabulary lookup: each 16 × 16 image patch is flattened and multiplied by one shared matrix, the same as a convolution with a 16 × 16 kernel and stride 16.
Perplexity
e raised to the average next-token loss: roughly, how many tokens the model is choosing between at each step. Lower is better.
Post-LN and pre-LN
Where LayerNorm sits: after each residual add (post-LN, the 2017 Transformer) or on the copy each sub-layer reads (pre-LN, GPT-2 and later), which trains more stably.
Prefill and decode
The two phases of generation: prefill runs the whole prompt in one pass and fills the KV cache; each decode step then runs a single new token.
Prefix caching
Keeping the KV cache of a prompt’s beginning between requests, so a long shared start such as a system prompt is not processed again.
QK-Norm
Normalising each head’s queries and keys before their dot product, which keeps attention scores from growing too large during training.
Quantization
Storing numbers in fewer bits, for example weights as 8- or 4-bit integers with a scale, to make a model smaller and faster to serve at a small cost in accuracy.
Query, key, value
Three vectors made from each token for attention: what it is looking for, what it offers, and what it passes on when chosen.
RAG
Retrieval-augmented generation: find the passages most related to a question in a document collection and put them in the prompt, so the model answers from them.
ReAct
A prompt format for agents (Yao et al., 2022): the model alternates a written Thought, an Action that names a tool, and an Observation holding the tool’s result.
Relative position bias
T5’s way of giving attention a sense of order: each query–key pair is sorted by distance into a bucket, and a learned number per bucket and head is added to its score.
Residual stream
The per-token vector that runs through the whole model. Every layer reads it and adds its output back instead of replacing it.
Reward model
A model trained on preference pairs to give each answer a score, higher for the answers people preferred; RLHF then tunes the language model to raise that score.
RLHF
Reinforcement learning from human feedback: train a reward model on people’s preferences between answers, then tune the language model with reinforcement learning (usually PPO) to write answers it scores highly.
RMSNorm
LayerNorm without the mean: divides by the root mean square, then applies a learned scale.
RNN
Recurrent neural network: reads a sequence one token at a time, updating a hidden state h[t] = f(h[t−1], x[t]). Each step waits for the previous one, so the steps cannot run in parallel.
RoPE
Rotary position embedding: query and key pairs are rotated by angles that grow with position, so the position part of a score depends only on the offset between tokens.
Router
In a mixture of experts, a small matrix that scores every expert for a token; the top-scoring experts run, weighted by a softmax over their scores.
Scaling law
The observation that a model’s loss falls smoothly and predictably, as a power law, as its parameters, training data and compute grow.
Selective scan
Mamba’s state-space step, in which the step size Δ and the matrices B and C depend on the current token, so the model chooses per token what to write into its state and what to keep.
Sentinel token
In T5’s span corruption, a placeholder token (<extra_id_0>, <extra_id_1>, …) that marks where a span was dropped; the target lists each sentinel followed by its span.
SFT
Supervised fine-tuning: training a pretrained model further on example conversations, with the next-token loss counted only on the answers.
Sliding window attention
Attention that only looks back a fixed number of tokens, the window; such a layer keeps at most that many keys and values, however long the text.
Softmax
Turns a list of scores into positive weights that sum to 1: the exp of each score divided by the sum of all the exps.
Span corruption
T5’s pretraining task: spans of the input are replaced by sentinel tokens, and the model writes out only the dropped spans.
Speculative decoding
Generating with a small draft model that guesses several tokens and a large model that checks them all in one pass, keeping the agreeing prefix; the output is the same as the large model’s.
SRAM and HBM
A GPU’s two main memories: SRAM is small and very fast, on the chip beside the arithmetic units; HBM is large (tens of GB) and slower, next to the chip.
State-space model (SSM)
A sequence layer that carries a fixed-size state from token to token, ht = Ā ht−1 + B̄ xt, and reads its output from it, yt = C ht. Work grows linearly with length.
Stop string
Text that ends generation as soon as the model writes it; an agent loop uses one to take control back before the model makes up a tool’s result.
SwiGLU
An MLP in which one projection, passed through SiLU, gates another element by element.
Teacher forcing
In training, the decoder is fed the correct previous tokens rather than its own guesses, so every position can be trained at once under the causal mask.
Temperature
Logits are divided by T before softmax: below 1 sharpens the distribution, above 1 flattens it, and T → 0 becomes greedy.
Text to text
T5’s framing: every task, from translation to classification, is a string in and a string out, named by a prefix such as “summarize:”.
Token
A piece of text the model reads as one unit (a word, part of a word, a byte or punctuation), each with an id in the vocabulary.
Tool call
Text a model writes to ask the surrounding program to run a function, such as JSON with a name and arguments; the program runs it and puts the result back in the context.
Top-k and top-p
Sampling that first keeps only the k likeliest tokens, or the smallest set whose probabilities add up to p, then renormalises.
Training
Adjusting every weight, over many steps, so the model gives a higher probability to the actual next token across a large body of text.
Vision Transformer (ViT)
A Transformer encoder whose tokens are image patches plus a [CLS] token; the [CLS] output is classified.
Vocabulary
The fixed set of tokens a model knows, each with an id and its own row in the embedding matrix; set by the tokenizer before training.
Weight decay
Shrinking every weight by a small fraction each training step, which keeps weights small unless the loss needs them large.
WordPiece
BERT’s tokenizer: words are split greedily into the longest pieces in its vocabulary; ## marks a piece that continues a word.
x · W and W x
Two ways to write the same product. Papers often write W x (a column vector); code and this site write x · W (x @ W), with one row per token.
Zero-shot
Doing a task without any training examples for it, such as classifying images by comparing them with a caption written for each class.