08 · LLMs · 5 min read · Interactive · updated
What does the temperature parameter do in an LLM, and what are top-k and top-p?
In short
Temperature divides the logits before the softmax: low values sharpen the distribution, high ones flatten it. Top-k and top-p cut off the unlikely tail.
What it is
Sampling is the way the next token is chosen from the probability distribution a language model returns. The temperature T is the number by which the logits are divided before the softmax function: T < 1 amplifies the differences and makes the model stick to the most probable tokens, while T > 1 smooths them out and increases diversity. Top-k and top-p (nucleus sampling) additionally restrict sampling to the most probable candidates.
The model itself does not choose words — it outputs a distribution. Whether the text will be predictable and dry or creative and chaotic is decided largely by the decoding strategy, i.e. by what we do with that distribution.
That is why the same model can sound like a cautious civil servant one moment and a poet the next: often the only difference is a few sampling parameters.
Mechanism — why it works this way
Temperature. The softmax with temperature is p_i = e^(z_i / T) / Σ e^(z_j / T). As T falls towards zero, the largest logit dominates more and more, and the distribution converges to choosing a single token (greedy decoding). As T rises, the differences between logits shrink and the distribution approaches uniform. The ranking of tokens never changes — only how strongly we favour the front-runners. The name comes from statistical physics, where the identical formula (the Boltzmann distribution) describes the states of particles at a given temperature.
Why not always choose the best token? Intuition suggests the most probable token is the "best" one. Holtzman et al. (2020) showed, however, that greedy decoding and beam search produce repetitive, looping text, noticeably less diverse than text written by humans. People do not always choose the most obvious word. On the other hand, pure sampling from the full distribution sometimes lands in the long tail of thousands of improbable tokens — and a single absurd token can derail the rest of the text, because it stays in the context.
Top-k. Fan et al. (2018) proposed sampling only from the k most probable tokens (after renormalizing). The drawback: a fixed k does not suit every situation. After "The capital of Poland is" practically one token makes sense, while after "I like to eat" hundreds do.
Top-p (nucleus sampling). Holtzman et al. (2020) propose taking the smallest set of tokens whose combined probability exceeds p (e.g. 0.9). When the model is confident, the nucleus has one or two tokens; when it is uncertain, dozens. The cut-off adapts to the shape of the distribution.
What temperature does not do. A low temperature does not make the model smarter or more truthful — it picks what the model considers most probable, and that can be wrong. It does, however, reduce random mistakes caused by sampling weak tokens, which is why it is usually lowered for data extraction and code.
By example
After "For dinner I'll have", the model gave these logits: " soup" 2.0, " dumplings" 1.0, " a sandwich" 0.5, " a rock" −1.0. At T = 1 the softmax gives 0.609, 0.224, 0.136, 0.030. At T = 0.5 the distribution sharpens: 0.842, 0.114, 0.042, 0.002 — "a rock" practically disappears. At T = 2 it flattens: 0.434, 0.263, 0.205, 0.097 — the absurd "a rock" now comes up in almost one draw in ten. The entropy of the distribution rises from 0.78 bits (T = 0.5) through 1.46 to 1.82 bits (T = 2).
Now top-p = 0.9 at T = 1. The cumulative probabilities are 0.609, 0.834, 0.970. We cross the 0.9 threshold only at the third token, so "soup", "dumplings" and "a sandwich" remain, and after renormalization they have 0.629, 0.231, 0.140. "A rock" is cut off completely. Top-k = 2 would keep only the first two tokens, with probabilities of 0.731 and 0.269.
In practice
- In
transformers:model.generate(do_sample=True, temperature=0.7, top_p=0.9, top_k=50);do_sample=Falsegives greedy decoding. - Typical settings: 0–0.3 for extraction, classification and code; 0.7–1.0 for free-form writing and brainstorming.
- Usually change one parameter at a time — temperature and top-p reinforce each other, making the effect hard to judge afterwards.
- A repetition penalty (
repetition_penalty) lowers the logits of tokens already used; it helps with loops, but too large a value breaks grammar. - For evaluation and reproducible tests, set a fixed random seed or temperature 0 — otherwise differences between prompt versions may just be sampling noise.
Frequently asked questions
- Does temperature 0 always give the same answer?
- In theory yes, because the token with the highest logit is chosen. In practice, parallel computation on GPUs and request batching can change the logits minutely and, in the case of ties, produce a different answer.
- Does a higher temperature increase creativity?
- It increases diversity, which we often perceive as creativity. Beyond a certain level, though, what mainly grows is the number of errors and inconsistencies, because the model more often picks tokens it itself rated as weak.
- Which is better: top-k or top-p?
- Top-p adapts better to the model's confidence, which is why it is used more often. In practice the two are combined: top-k as a hard safety limit, top-p as the main cut-off.
Sources
- Holtzman A. et al., 2020, "The Curious Case of Neural Text Degeneration", ICLR 2020.
- Fan A., Lewis M., Dauphin Y., 2018, "Hierarchical Neural Story Generation", ACL 2018.
- Hinton G., Vinyals O., Dean J., 2015, "Distilling the Knowledge in a Neural Network", arXiv:1503.02531 (softmax with temperature).
- Jurafsky D., Martin J. H., "Speech and Language Processing", 3rd ed. (online draft), chapter on large language models (sampling).
- Hugging Face Transformers documentation, "Generation strategies", https://huggingface.co/docs/transformers/generation_strategies