Warm-up
- shown
- small examples
- takes
- ≈ 30 s
All steps
Dot product
A dot product multiplies two vectors number by number and adds the results. It works as a similarity score: large when the vectors point the same way.
a · b = Σ aₖ bₖMatrix product
A matrix product is a grid of dot products: cell (i, j) takes row i of A and column j of B. Every page draws it this way, with A on the left and B above.
[3 × 4] · [4 × 3] → [3 × 3]Softmax
Softmax turns any list of scores into positive weights that sum to 1, keeping their order. Attention and next-token prediction both use it.
exp(sₖ) / Σ exp(s)One-hot lookup
A one-hot row has a single 1. Multiplied by a matrix, it simply picks out one row: the idea behind the embedding lookup.
onehot · M = one row
Code
a @ b # dot product: (a * b).sum()
C = A @ B # (3, 4) @ (4, 3) → (3, 3); C[i, j] = A[i] @ B[:, j]
p = torch.softmax(s, dim=-1) # exp(s) / exp(s).sum(), rows sum to 1
F.one_hot(i, 4).float() @ M # == M[i]