What You'll Learn
- How queries, keys, and values work together to let tokens exchange information
- Why attention scores are divided by the square root of the key dimension
- How scaling prevents softmax from saturating and preserves useful gradient flow
- Why transformers split attention across multiple heads instead of using one large attention operation
- What researchers have discovered about the specialized roles learned by attention heads
- How induction heads learn a match-and-copy algorithm that helps explain in-context learning
What You'll Learn
- How attention behaves like a soft version of predicate logic
- How top-k and top-p sampling use ideas from set theory
- How Boolean logic determines which tokens can attend to each other
- Why chain-of-thought reasoning resembles the structure of a proof
- How positional encoding uses periodic patterns to represent position
- How discrete math helps explain why common LLM techniques work
![]()