Temperature, Top-p, Top-k, and Min-p: How the LLM Sampler Actually Works

Every LLM application has two knobs that get blamed for everything: temperature and top-p. When outputs are boring, someone raises temperature. When outputs are unhinged, someone lowers top-p. Both rituals miss the point, because temperature and top-p solve different problems, and top-p has a subtle flaw that a newer parameter — min-p — fixes more cleanly. Understanding what each knob actually does to the token distribution is the difference between tuning a sampler and performing cargo cult inference.

This post walks through the full sampler chain as an inference engine actually executes it: temperature reshaping, top-k truncation, top-p (nucleus) filtering, min-p filtering, and the final multinomial draw. We’ll look at the math each stage applies, when the stages interact badly, and what sensible configurations look like for real workloads — extraction, code generation, and creative writing all want different chains.

The raw distribution comes out of the logits

The model ends every forward pass with a vector of raw scores — logits — one per vocabulary entry. Those are unnormalized real numbers like 23.1, 18.7, −4.2. The first sampler stage converts them to a probability distribution with a softmax: exponentiate each logit and divide by the sum. After softmax, every token has a probability in [0, 1] and all probabilities sum to 1. Everything that follows operates on this distribution.

The key mental model: sampling is a pipeline of filters and transforms applied in a fixed order. Each stage can remove candidates or reshape their probabilities, and the stages are not commutative — temperature before truncation produces different results than truncation before temperature. Inference engines like vLLM and Hugging Face Transformers fix the order for you: temperature first, then truncation filters, then the draw. That order matters when you reason about why two settings that “should” be equivalent are not.

Temperature: reshaping before filtering

Temperature doesn’t remove any tokens. It divides every logit by a scalar T and re-softmaxes. With T = 1 you get the model’s native distribution. With T below 1, differences between logits get amplified: the gap between 23.1 and 18.7 becomes a much larger probability gap, mass concentrates on the top tokens, and sampling converges toward greedy argmax selection as T approaches 0. With T above 1, logits flatten toward uniform and unlikely tokens gain real probability mass.

Two practical consequences follow. First, temperature is the only knob that changes the relative shape of the distribution before filtering, so it determines how much material the truncation stages have to work with. Second, T = 0 (or engines interpreting temperature as 0) switches off sampling entirely — the draw becomes argmax, and top-p and top-k become irrelevant because there is nothing left to sample. If you’re debugging nondeterminism, check temperature first.

Top-k: the bluntest truncation

Top-k keeps only the k highest-probability tokens and zeroes the rest, renormalizing the survivors. It’s the oldest trick and the least adaptive: k = 50 means 50 candidates whether the model is confident the next token is “the” (one token holds essentially all the mass) or standing at a genuine fork in a story (twenty tokens are all plausible). In the confident case, top-k keeps 49 near-zero candidates that add noise; in the fork case, it amputates legitimate options at position 51.

Top-k is still useful as a cheap upper bound on candidate set size, and some model vendors recommend it in fixed combinations — for example, the Qwen3 model cards specify temperature 0.6, top-p 0.95, top-k 20 for thinking mode. But as a standalone diversity control it’s been superseded, because the two distributions it handles badly are precisely the interesting ones.

Top-p (nucleus): adaptive, but confidence-blind

Top-p sorts tokens by probability, walks down the list accumulating mass, and cuts off once the cumulative sum exceeds p. Keep everything up to that cut, zero the rest, renormalize. With p = 0.9 on a confident distribution, maybe three tokens survive; on an uncertain one, hundreds do. That adaptivity is why nucleus sampling (introduced in the paper “The Curious Case of Neural Text Degeneration”) displaced top-k as the default truncation method.

The flaw: top-p is calibrated to the absolute probability scale, which shifts with model confidence. When the model is confident — say the top token has 90% probability — a p = 0.9 nucleus might contain only that token plus one or two others. Fine. But when the model is uncertain — top token at 8% — the same p = 0.9 nucleus now sweeps in dozens of junk tokens with probabilities in the 0.1–2% range, because that’s what it takes to fill the 90% budget. The cut line moves with the shape of the distribution, but it doesn’t track the top token. The result is that at high temperature, top-p either over-truncates (confident steps) or under-truncates (uncertain steps), and you can’t fix both with one value of p.

Min-p: truncation relative to the top token

Min-p changes the reference point. Instead of asking “how much cumulative probability should I keep?”, it asks “how probable must a token be, relative to the most probable token?” A min-p value of 0.1 means: any token whose probability is at least 10% of the top token’s probability survives; everything else is dropped, and the survivors are renormalized.

This single change makes the truncation dynamic in exactly the way top-p isn’t. On a confident step where the top token holds 90%, the threshold is 9% — junk at 0.5% never enters. On an uncertain step where the top token holds 8%, the threshold drops to 0.8% — genuinely plausible alternatives survive even though their absolute probabilities look tiny. The candidate set scales with model confidence, not with an absolute mass budget.

The method was validated in “Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs” (arXiv:2407.01082), which showed that min-p combined with higher temperatures produces text that is simultaneously more coherent and more diverse than temperature-plus-top-p baselines — breaking what had looked like a fixed trade-off. The intuition is that at high temperature the distribution flattens, top-p floods the candidate set with low-quality tokens, and the resulting text degrades; min-p prunes exactly those tokens while letting the flattening increase diversity among the survivors. The paper found optimal quality at temperatures around 1.5 with min-p filtering, far above the 0.7–1.0 range where top-p setups saturate.

Min-p is supported across the major serving stacks: it’s a first-class SamplingParams field in vLLM (defaulting to 0, i.e. disabled), a min_p generation parameter in Transformers’ GenerationConfig, and an accepted parameter in OpenAI-compatible servers like llama.cpp and LM Studio. If your serving layer doesn’t expose it, you can emulate the effect with a logits processor, but the native parameter is simpler.

How the chain composes in practice

Putting it together, a modern sampler in a typical engine executes, in order:

  • Temperature scaling — divide logits by T, softmax.
  • Top-k truncation — cap the candidate list (optional).
  • Top-p truncation — cut by cumulative mass (optional).
  • Min-p truncation — cut by relative-to-top probability (optional).
  • Multinomial draw — sample one token from what’s left.

Stages stack, and that’s where misconfigurations hide. A common mistake is pairing a low temperature with aggressive truncation: temperature 0.2 already concentrates mass on one or two tokens, and then min-p 0.1 amputates the runner-up, so your “sampling” is greedy with extra steps and occasional inexplicable jumps. Another is setting top-p and min-p to fight each other — top-p 0.9 lets in a token with 3% absolute probability on an uncertain step, min-p 0.2 then kills it. The filters are applied sequentially, so the most restrictive one wins; there’s no negotiation.

A rules of thumb table by workload:

Structured extraction, classification, JSON: temperature 0–0.3, no sampling filters that matter, constrain with grammar/JSON schema if the engine supports it. Determinism and validity dominate; diversity is a defect.

Code generation: temperature 0.2–0.6, top-p 0.9–0.95 or min-p 0.05–0.1. You want enough sampling to escape local repetition in long generations but not enough to hallucinate identifiers. Vendor recommendations for coding-capable models cluster here.

Creative writing, brainstorming, agent exploration: this is min-p’s home turf. Temperature 1.0–1.5 with min-p 0.05–0.1 gives broader exploration than top-p can tolerate at the same quality level. If you must stay on top-p, keep temperature ≤ 1.0 and accept less diversity.

Reasoning / thinking modes: follow the model card. Reasoning-tuned models are typically calibrated for moderate temperature (0.6) with top-p 0.95 and a top-k cap, and they degrade measurably when you “improve” the settings — the vendor calibrated these against their RL training distribution.

Common pitfalls

  • Stacking defaults blindly. Many servers default to temperature 1.0, top-p 1.0, top-k −1, min-p 0 — pure multinomial sampling, which produces repetition loops and incoherence on small models. “It repeats itself” is often just unset parameters.
  • Changing two knobs at once. Temperature and truncation interact. If you tune top-p while temperature is non-default, you’re fitting two variables with one observation.
  • Trusting temperature after a fine-tune. Sampler settings that were right for the base model are frequently wrong for a fine-tune with a narrower output distribution. Re-validate.
  • Forgetting seeds. With temperature above 0, identical requests can produce different outputs. If you need reproducibility for tests or evals, pass a fixed seed alongside the sampling parameters — most engines support it, though results are only deterministic per-engine-version, not across hardware or quantizations.
  • Using sampling to fix a model problem. If the model doesn’t know the answer, no sampler setting rescues it. Truncation filters change which plausible token you get; they don’t create knowledge.

Wrapping up

Temperature, top-k, top-p, and min-p are four different operations on the same distribution: reshape, cap by count, cap by mass, cap by relative confidence. Once you read sampler configs as pipelines instead of magic numbers, most “tuning” questions answer themselves. Start with the workload table above, change one parameter at a time, and prefer min-p over aggressive top-p when you need diversity without incoherence — the era of treating the sampler as an untouchable black box is over, because the inference engines have exposed every stage of it.

Leave a Reply

Your email address will not be published. Required fields are marked *