What You'll Learn
- How an LLM turns raw token scores into probabilities before choosing what comes next
- How temperature reshapes a probability distribution to make outputs more predictable or more varied
- How top-p sampling limits which tokens the model is allowed to consider
- Why top-p adapts to model confidence differently from a fixed top-k cutoff
- How temperature and top-p interact when both are applied to the same distribution
- How to choose sampling settings for factual, structured, professional, and creative tasks
What You'll Learn
- Why language models break text into tokens instead of reading whole words
- How subword tokenization balances vocabulary size with the amount of text a model must process
- How Byte Pair Encoding (BPE) learns useful token boundaries from patterns in training data
- Why the same sentence can use very different numbers of tokens across languages, code, and rare words
- How tokenization can cause surprising failures in spelling, letter counting, and unusual inputs
- Why token counts affect context limits, processing efficiency, and the cost of using an LLM
What You'll Learn
- Why an LLM’s context window is not the same thing as memory
- How chat applications create continuity even though the underlying model is stateless
- How attention lets earlier parts of a conversation influence the next token
- Why longer conversations become increasingly expensive for a transformer to process
- Why information can become harder to use even while it remains inside the context window
- How truncation, summarization, retrieval, and KV caching help manage long conversations
![]()