2 "Softmax" Posts

Inside Attention, Part 1: The Mechanism

Attention is the engine. The rest of the transformer architecture stabilizes it, organizes it, and makes deep training possible.
What You'll Learn
  • How queries, keys, and values work together to let tokens exchange information
  • Why attention scores are divided by the square root of the key dimension
  • How scaling prevents softmax from saturating and preserves useful gradient flow
  • Why transformers split attention across multiple heads instead of using one large attention operation
  • What researchers have discovered about the specialized roles learned by attention heads
  • How induction heads learn a match-and-copy algorithm that helps explain in-context learning

Temperature and Top-P: The Creativity Knobs

How sampling parameters shape AI personality
What You'll Learn
  • How an LLM turns raw token scores into probabilities before choosing what comes next
  • How temperature reshapes a probability distribution to make outputs more predictable or more varied
  • How top-p sampling limits which tokens the model is allowed to consider
  • Why top-p adapts to model confidence differently from a fixed top-k cutoff
  • How temperature and top-p interact when both are applied to the same distribution
  • How to choose sampling settings for factual, structured, professional, and creative tasks