Warm-up

Foundations

the four pieces of math the rest uses

shown
small examples
takes
≈ 30 s

All steps

  1. Dot product

    A dot product multiplies two vectors number by number and adds the results. It works as a similarity score: large when the vectors point the same way.

    a · b = Σ aₖ bₖ

  2. Matrix product

    A matrix product is a grid of dot products: cell (i, j) takes row i of A and column j of B. Every page draws it this way, with A on the left and B above.

    [3 × 4] · [4 × 3] → [3 × 3]

  3. Softmax

    Softmax turns any list of scores into positive weights that sum to 1, keeping their order. Attention and next-token prediction both use it.

    exp(sₖ) / Σ exp(s)

  4. One-hot lookup

    A one-hot row has a single 1. Multiplied by a matrix, it simply picks out one row: the idea behind the embedding lookup.

    onehot · M = one row

Code

a @ b                        # dot product: (a * b).sum()
C = A @ B                    # (3, 4) @ (4, 3) → (3, 3); C[i, j] = A[i] @ B[:, j]
p = torch.softmax(s, dim=-1) # exp(s) / exp(s).sum(), rows sum to 1
F.one_hot(i, 4).float() @ M  # == M[i]

Go deeper