Reference
- Active parameters
- The weights one token actually runs through. In a mixture of experts that is far fewer than the model stores: the experts a token is not sent to sit in memory unused.
- adaLN-Zero
- DiT’s conditioning: the timestep and class set each block’s LayerNorm scale and shift and a gate on each sub-layer’s output; the gates start at zero, so blocks start as the identity.
- Adam
- The standard optimizer for Transformers: it keeps a running mean of each weight’s gradient (m) and of its square (v), and steps by m / √v, so every weight moves at a similar pace. AdamW adds weight decay separately.
- Agent
- A program that runs a language model in a loop: the model chooses actions such as tool calls, the program carries them out and feeds the results back, until the task is done.
- Attention
- Lets each position mix in information from itself and earlier positions, weighted by how well its query matches their keys.
- Attention head
- One of several attention computations run side by side, each on its own slice of the vector (64 of 768 numbers in GPT-2), free to track a different relation.
- Attention sink
- Many heads park spare attention on the first token, whose value adds almost nothing; the overview colours treat that attention as a no-op. gpt-oss gives each head a learned sink instead: an extra score that can take attention without any token.
- Backpropagation
- Computing the gradient of the loss for every weight at once, by applying the chain rule layer by layer from the output back to the input.
- Base model
- A model after pretraining only: it continues text in the style of its training data and has not been tuned to follow instructions.
- BatchNorm
- Normalises each feature across the examples in a batch (common in vision). Transformers use LayerNorm instead, which normalises each token on its own.
- BPE
- Byte-pair encoding: a tokenizer that starts from bytes and repeatedly merges the most frequent adjacent pair seen in training text. GPT-2 learned 50,000 merges this way.
- Causal mask
- Hides later positions from each position, so the model cannot peek at the tokens it is learning to predict.
- Chain rule
- If y depends on x and the loss depends on y, the loss’s slope with respect to x is its slope with respect to y times y’s slope with respect to x.
- Chat template
- The text format a chat model was tuned on: special tokens such as <|im_start|> open each message and name who wrote it, and the model’s turn begins where the template stops.
- Chunked prefill
- Splitting a long prompt’s prefill into pieces that ride along with other requests’ decode steps, so no single step becomes slow.
- Classifier-free guidance
- At sampling time, running a diffusion model with and without the condition and moving further in the direction the condition adds, for images that match it more strongly.
- Compute-bound and memory-bound
- Whether a GPU step is limited by how fast it can do arithmetic or by how fast it can read memory; a step doing few operations per byte read (like decoding one token) waits on memory.
- Compute-optimal
- The split of a fixed training budget between model size and training tokens that gives the lowest loss; about 20 tokens per parameter by the Chinchilla fit.
- Constrained decoding
- Masking, at each step, every token that would break a required format (a JSON grammar, a list of tool names) before the next token is chosen, so the output always fits the format.
- Continuous batching
- Rescheduling a serving batch after every decode step: finished requests leave at once and waiting ones take their place, instead of waiting for the whole batch to finish.
- Contrastive learning
- Training by comparison: matching pairs (an image and its caption) are pulled together and every mismatched pair in the batch is pushed apart.
- Copy-on-write
- Sharing memory between users until one of them writes to it; only then is a private copy made.
- Cosine similarity
- How closely two vectors point the same way: their dot product divided by both lengths, from −1 to 1. For vectors of length 1 it is just the dot product.
- Cross-attention
- Attention whose queries come from one sequence and keys and values from another: in a translation model, the decoder (target) reads the encoder’s output (source).
- Decoder-only
- A Transformer that reads left to right under a causal mask and predicts the next token. GPT-2 and LLaMA are decoder-only.
- Diffusion model
- A model trained to remove noise from data; generating means starting from pure noise and denoising step by step.
- DPO
- Direct preference optimization: tuning a model on pairs of preferred and rejected answers with one loss that raises the preferred answer’s probability relative to a frozen reference model and lowers the other’s.
- Embedding
- A learned vector of numbers (768 in GPT-2) that stands for a token or a position; tokens used in similar ways end up with similar vectors.
- Embedding model
- A model that turns a whole text into one vector, trained so that texts with related meanings get vectors pointing in similar directions.
- Encoder and decoder
- In the 2017 Transformer, the encoder reads the whole source sentence at once; the decoder writes the output one token at a time, under a causal mask, reading the encoder’s output through cross-attention.
- Encoder-only
- A Transformer with no causal mask that reads a whole text at once and outputs a vector per token, for understanding rather than generating text. BERT is encoder-only.
- Few-shot prompting
- Putting a few worked examples of a task in the prompt before the real question, so the model continues the pattern.
- Fine-tuning
- Training a pretrained model a little further, on a smaller dataset for one task, usually with a small new output layer on top.
- Finite difference
- Estimating a slope by moving an input a tiny step each way and dividing the change in output by the step; slow, but a good check.
- FlashAttention
- An exact attention kernel that works through tiles of queries, keys and values in the GPU’s on-chip SRAM, with an online softmax, so the N × N score matrix is never written to HBM.
- FLOPs
- Floating-point operations. One multiply-add counts as two.
- GELU
- A smooth nonlinearity: large positive inputs pass almost unchanged, small ones are damped and negative ones are pushed close to 0.
- GEMM
- General matrix multiply, C = A · B: the operation almost all of a Transformer’s compute goes into.
- GQA
- Grouped-query attention: several query heads share one key/value head, which shrinks the KV cache.
- Gradient
- For every weight, how fast the loss changes as that weight changes. Training moves each weight a little against its gradient.
- Greedy decoding
- Always picking the most likely next token.
- Handoff
- Passing a piece of work, or the whole conversation, from one agent to another, usually by a tool call; the receiving agent sees only what is passed to it.
- In-context learning
- A model doing a task it is shown in its prompt, by examples or instructions, with no change to its weights.
- Induction head
- A head that finds an earlier copy of the current token and attends to the token that came after it, so a repeated pattern can be continued.
- Instruction tuning
- Fine-tuning a pretrained model on instructions paired with good answers, written in a chat template, so that it answers requests instead of continuing the text.
- JSON Schema
- A standard way to describe the shape of JSON data: its fields, their types and which are required. Tool definitions use it for their parameters.
- KV cache
- While generating, the keys and values of past tokens are stored so each new token only computes its own. It grows with every token and every layer.
- LayerNorm
- Rescales each token’s vector to mean 0 and spread 1, then applies a learned scale (γ) and shift (β) per feature.
- Learning rate
- How far each training step moves the weights against the gradient (η). It is usually warmed up from zero, then decayed.
- Linear attention
- An attention layer that folds the past into a fixed-size state instead of keeping every key and value, so its memory does not grow with the text; Qwen3-Next and Kimi Linear mix it with full attention.
- Logit lens
- Reading a middle layer through the final LayerNorm and the unembedding, to see what the model would predict if it stopped there.
- Logits
- Raw, unnormalised scores, one per vocabulary token; softmax turns them into probabilities.
- LoRA
- Low-rank adaptation: fine-tuning by keeping the weights frozen and learning a small change to them as the product of two thin matrices.
- Masked language modelling
- BERT’s training task: some tokens are hidden and the model predicts them from the words on both sides.
- Mixture of experts (MoE)
- A layer with several expert MLPs and a router; each token runs through only the few experts the router picks, so the model can store many more parameters than it uses per token.
- MLP
- Multi-layer perceptron: two matrix products with a nonlinearity (GELU) between them, applied to each token on its own.
- Momentum
- Stepping along a running sum of past gradients instead of the latest one alone, so oscillations cancel and consistent directions build up speed.
- Multi-head latent attention (MLA)
- DeepSeek’s attention: each token’s keys and values are compressed into one small latent vector, which is all the KV cache stores; every head’s keys and values are rebuilt from it.
- Multi-token prediction
- Training a model to also predict tokens further ahead (the one after next), as an extra, denser training signal.
- Online softmax
- Computing a softmax in chunks with a running maximum and a running sum, rescaling earlier terms when the maximum grows; the result equals the ordinary softmax.
- Orchestrator
- In a multi-agent system, the agent that splits a task, hands the pieces to worker agents and combines what they report.
- PagedAttention
- vLLM’s way of storing the KV cache: in fixed-size blocks handed out on demand and found through a per-request block table, like virtual-memory pages.
- Patch embedding
- ViT’s replacement for a vocabulary lookup: each 16 × 16 image patch is flattened and multiplied by one shared matrix, the same as a convolution with a 16 × 16 kernel and stride 16.
- Perplexity
- e raised to the average next-token loss: roughly, how many tokens the model is choosing between at each step. Lower is better.
- Post-LN and pre-LN
- Where LayerNorm sits: after each residual add (post-LN, the 2017 Transformer) or on the copy each sub-layer reads (pre-LN, GPT-2 and later), which trains more stably.
- Prefill and decode
- The two phases of generation: prefill runs the whole prompt in one pass and fills the KV cache; each decode step then runs a single new token.
- Prefix caching
- Keeping the KV cache of a prompt’s beginning between requests, so a long shared start such as a system prompt is not processed again.
- QK-Norm
- Normalising each head’s queries and keys before their dot product, which keeps attention scores from growing too large during training.
- Quantization
- Storing numbers in fewer bits, for example weights as 8- or 4-bit integers with a scale, to make a model smaller and faster to serve at a small cost in accuracy.
- Query, key, value
- Three vectors made from each token for attention: what it is looking for, what it offers, and what it passes on when chosen.
- RAG
- Retrieval-augmented generation: find the passages most related to a question in a document collection and put them in the prompt, so the model answers from them.
- ReAct
- A prompt format for agents (Yao et al., 2022): the model alternates a written Thought, an Action that names a tool, and an Observation holding the tool’s result.
- Relative position bias
- T5’s way of giving attention a sense of order: each query–key pair is sorted by distance into a bucket, and a learned number per bucket and head is added to its score.
- Residual stream
- The per-token vector that runs through the whole model. Every layer reads it and adds its output back instead of replacing it.
- Reward model
- A model trained on preference pairs to give each answer a score, higher for the answers people preferred; RLHF then tunes the language model to raise that score.
- RLHF
- Reinforcement learning from human feedback: train a reward model on people’s preferences between answers, then tune the language model with reinforcement learning (usually PPO) to write answers it scores highly.
- RMSNorm
- LayerNorm without the mean: divides by the root mean square, then applies a learned scale.
- RNN
- Recurrent neural network: reads a sequence one token at a time, updating a hidden state h[t] = f(h[t−1], x[t]). Each step waits for the previous one, so the steps cannot run in parallel.
- RoPE
- Rotary position embedding: query and key pairs are rotated by angles that grow with position, so the position part of a score depends only on the offset between tokens.
- Router
- In a mixture of experts, a small matrix that scores every expert for a token; the top-scoring experts run, weighted by a softmax over their scores.
- Scaling law
- The observation that a model’s loss falls smoothly and predictably, as a power law, as its parameters, training data and compute grow.
- Selective scan
- Mamba’s state-space step, in which the step size Δ and the matrices B and C depend on the current token, so the model chooses per token what to write into its state and what to keep.
- Sentinel token
- In T5’s span corruption, a placeholder token (<extra_id_0>, <extra_id_1>, …) that marks where a span was dropped; the target lists each sentinel followed by its span.
- SFT
- Supervised fine-tuning: training a pretrained model further on example conversations, with the next-token loss counted only on the answers.
- Sliding window attention
- Attention that only looks back a fixed number of tokens, the window; such a layer keeps at most that many keys and values, however long the text.
- Softmax
- Turns a list of scores into positive weights that sum to 1: the exp of each score divided by the sum of all the exps.
- Span corruption
- T5’s pretraining task: spans of the input are replaced by sentinel tokens, and the model writes out only the dropped spans.
- Speculative decoding
- Generating with a small draft model that guesses several tokens and a large model that checks them all in one pass, keeping the agreeing prefix; the output is the same as the large model’s.
- SRAM and HBM
- A GPU’s two main memories: SRAM is small and very fast, on the chip beside the arithmetic units; HBM is large (tens of GB) and slower, next to the chip.
- State-space model (SSM)
- A sequence layer that carries a fixed-size state from token to token, ht = Ā ht−1 + B̄ xt, and reads its output from it, yt = C ht. Work grows linearly with length.
- Stop string
- Text that ends generation as soon as the model writes it; an agent loop uses one to take control back before the model makes up a tool’s result.
- SwiGLU
- An MLP in which one projection, passed through SiLU, gates another element by element.
- Teacher forcing
- In training, the decoder is fed the correct previous tokens rather than its own guesses, so every position can be trained at once under the causal mask.
- Temperature
- Logits are divided by T before softmax: below 1 sharpens the distribution, above 1 flattens it, and T → 0 becomes greedy.
- Text to text
- T5’s framing: every task, from translation to classification, is a string in and a string out, named by a prefix such as “summarize:”.
- Token
- A piece of text the model reads as one unit (a word, part of a word, a byte or punctuation), each with an id in the vocabulary.
- Tool call
- Text a model writes to ask the surrounding program to run a function, such as JSON with a name and arguments; the program runs it and puts the result back in the context.
- Top-k and top-p
- Sampling that first keeps only the k likeliest tokens, or the smallest set whose probabilities add up to p, then renormalises.
- Training
- Adjusting every weight, over many steps, so the model gives a higher probability to the actual next token across a large body of text.
- Vision Transformer (ViT)
- A Transformer encoder whose tokens are image patches plus a [CLS] token; the [CLS] output is classified.
- Vocabulary
- The fixed set of tokens a model knows, each with an id and its own row in the embedding matrix; set by the tokenizer before training.
- Weight decay
- Shrinking every weight by a small fraction each training step, which keeps weights small unless the loss needs them large.
- WordPiece
- BERT’s tokenizer: words are split greedily into the longest pieces in its vocabulary; ## marks a piece that continues a word.
- x · W and W x
- Two ways to write the same product. Papers often write W x (a column vector); code and this site write x · W (x @ W), with one row per token.
- Zero-shot
- Doing a task without any training examples for it, such as classifying images by comparing them with a caption written for each class.