← Math in Practice
Concept Math behind AI

Exponentials, softmax, and temperature

Follow candidate scores into a distribution, then check what that distribution can tell you.

A writing assistant starts choosing unusually repetitive next words after a decoding setting changes. The top candidate still looks reasonable, so the team is unsure whether the model's belief changed or only its sampling behavior. To investigate, follow the raw scores through the calculation that turns them into a probability distribution.

The judgment to keep

Softmax converts finite scores into positive probabilities by exponentiating and normalizing. Temperature scales score differences before that conversion: lower values concentrate mass while higher values spread it. The resulting confidence is a property of the distribution, not proof that the chosen token or answer is correct.

TypeScriptGo Logits · exponentials · softmax · temperature · log-sum-exp · numerical stability
01 / Read the ranking review

The candidate with the largest score is easy to name; its probability takes another step.

A language model assigns next-token logits to the words or token fragments that could follow. For this small example, three candidates have logits [2.4, 1.1, 0.2]. A logit is an unnormalized score: adding the same constant to every logit leaves their relative ranking unchanged, and the values do not have to lie between zero and one or sum to anything.

The question from the incident is empirical: did the model's scores change, did the temperature change, or did another decoding control alter token choice? Capture the logits and decoding configuration for the same prompt before changing settings. This lesson isolates the softmax stage so that each effect can be checked.

Case file / Next-token selectionWhy did the assistant become less predictable?
Observed
Generated text became more varied after a deployment.
Candidate scores
Three illustrative logits: 2.4, 1.1, and 0.2.
Competing explanations
Changed logits, temperature, top-k/top-p filtering, or random sampling.
First check
Compare captured logits and decoding configuration on the same prompt.
02 / Turn scores into weights

The exponential makes every weight positive and amplifies score differences.

For a logit zᵢ, exponentiation gives the unnormalized weight eᶻⁱ. Because exponentials are positive, the weights can be normalized. Their ratio makes the role of differences visible: for two candidates, eᶻ¹ / eᶻ² = e⁽ᶻ¹⁻ᶻ²⁾. A difference of 1.3 means a weight ratio of e¹·³ ≈ 3.67, no matter whether the logits are 2.4 and 1.1 or both shifted upward by 100.

With the three example logits and temperature T = 1, subtract the maximum first: [2.4, 1.1, 0.2] − 2.4 = [0, −1.3, −2.2]. The stable weights are [1, e⁻¹·³, e⁻²·²] ≈ [1, 0.2725, 0.1108]. The denominator is their sum, approximately 1.3833.

01 / Largest logit2.4

Subtract it from each score without changing score differences.

02 / Exponential weights1 · 0.2725 · 0.1108

Positive values preserve which candidate ranked higher.

03 / DenominatorΣ weights ≈ 1.3833

Every weight will be divided by this shared total.

03 / Normalize with softmax

Dividing each weight by the total produces probabilities that sum to one.

At temperature T > 0, softmax is pᵢ = exp(zᵢ/T) / Σⱼ exp(zⱼ/T). For this example at T = 1, the probabilities are approximately [0.723, 0.197, 0.080]. The first candidate has the largest modeled share of next-token probability mass; the values sum to one apart from display rounding.

Adding the same constant c to every logit does not change the result because exp((zᵢ+c)/T) contains a shared factor exp(c/T), which cancels between numerator and denominator. This is useful diagnostically: large-looking absolute logits alone do not tell us the distribution's sharpness; relative gaps and temperature do.

Hand-checkable calculation; displayed values rounded to four decimals.
CandidateLogitWeight after max subtractionProbability
A2.4exp(0) = 10.7229
B1.1exp(−1.3) ≈ 0.27250.1970
C0.2exp(−2.2) ≈ 0.11080.0801
Total—1.38331.0000*

*Each displayed probability is rounded independently; the unrounded probabilities sum to one.

04 / Change temperature

Temperature changes the spread of probability, not the ordering of finite logits.

Dividing by T scales each difference from the maximum. At T < 1, gaps grow before exponentiation, so the top candidate gets a larger share. At T > 1, gaps shrink, so probability mass spreads toward the other candidates. As positive T approaches zero, the distribution concentrates on the maximum-logit candidate; as T grows very large, a finite candidate set approaches a uniform distribution.

Temperature does not reorder candidates when logits are finite and T is positive. But it can change which token a random sampler emits. Top-k or nucleus (top-p) filtering can then remove candidates and renormalize a different set. To diagnose a generation change, compare these settings separately rather than attributing all variety to temperature alone.

05 / Keep the arithmetic stable

Compute relative exponentials so a large shared score cannot overflow.

A direct calculation can try to evaluate exp(1000), which exceeds ordinary floating-point range even though the final probabilities are well-defined. Let m = max(z). Then exp(zᵢ/T) / Σexp(zⱼ/T) = exp((zᵢ−m)/T) / Σexp((zⱼ−m)/T) because the same factor exp(−m/T) cancels. Every exponent is now non-positive and the largest is zero, so at least one weight is exactly one.

The corresponding log denominator uses log-sum-exp: log Σexp(zᵢ/T) = m/T + log Σexp((zᵢ−m)/T). This is the stable route for log-probabilities and log-likelihoods too. The next lesson, Log-probabilities and information theory, uses this connection to explain surprisal, cross-entropy, and perplexity.

06 / Inspect the distribution

Move temperature while holding candidate scores fixed.

These controls change only the temperature. The logits remain [2.4, 1.1, 0.2], so any visible change comes from rescaling the same score gaps and applying softmax. The bars represent next-token probability mass for this illustrative candidate set.

Probability total: 1.0000. The largest logit remains Candidate 1; sampling randomness and filtering may still affect the emitted token.
Candidate 1logit 2.472.3%
Candidate 2logit 1.119.7%
Candidate 3logit 0.28.0%

A temperature change alters the distribution, not the candidate scores or their rank. This lab has no random sampling step.

07 / Practice in code

Keep score transformation and probability normalization explicit.

The examples validate finite logits and a positive finite temperature, subtract the maximum before exponentiating, and return probabilities. The TypeScript companion also shows a log-sum-exp helper. They do not implement token filtering, random sampling, or model inference.

Compare the same softmax in TypeScript and Go.

Both examples use the same three logits and compare temperatures 0.5, 1, and 2.

TypeScriptCalculate stable softmax probabilities
softmax.ts
export type Probability = { label: string; logit: number; probability: number };

/** Softmax for finite logits. Subtracting the maximum preserves probabilities and avoids overflow. */
export function softmax(logits: readonly number[], temperature = 1): Probability[] {
	if (logits.length === 0) throw new RangeError('logits must not be empty');
	if (!Number.isFinite(temperature) || temperature <= 0) {
		throw new RangeError('temperature must be finite and greater than zero');
	}
	if (logits.some((logit) => !Number.isFinite(logit))) {
		throw new TypeError('every logit must be finite');
	}

	const maximum = logits.reduce((current, logit) => Math.max(current, logit), -Infinity);
	const weights = logits.map((logit) => Math.exp((logit - maximum) / temperature));
	const total = weights.reduce((sum, weight) => sum + weight, 0);
	return logits.map((logit, index) => ({
		label: `Candidate ${index + 1}`,
		logit,
		probability: weights[index] / total
	}));
}

/** log(sum(exp(logits / temperature))) using the same max-subtraction trick. */
export function logSumExp(logits: readonly number[], temperature = 1): number {
	if (logits.length === 0) throw new RangeError('logits must not be empty');
	if (!Number.isFinite(temperature) || temperature <= 0) {
		throw new RangeError('temperature must be finite and greater than zero');
	}
	if (logits.some((logit) => !Number.isFinite(logit))) {
		throw new TypeError('every logit must be finite');
	}
	const maximum = logits.reduce((current, logit) => Math.max(current, logit), -Infinity);
	const scaledSum = logits.reduce(
		(sum, logit) => sum + Math.exp((logit - maximum) / temperature),
		0
	);
	return maximum / temperature + Math.log(scaledSum);
}

export const rankingLogits = [2.4, 1.1, 0.2] as const;
GoCalculate stable softmax probabilities
softmax.go
package main

import (
	"errors"
	"fmt"
	"math"
)

type Probability struct {
	Logit       float64
	Probability float64
}

// Softmax subtracts the maximum logit before exponentiating to avoid overflow.
func Softmax(logits []float64, temperature float64) ([]Probability, error) {
	if len(logits) == 0 {
		return nil, errors.New("logits must not be empty")
	}
	if math.IsNaN(temperature) || math.IsInf(temperature, 0) || temperature <= 0 {
		return nil, errors.New("temperature must be finite and greater than zero")
	}
	maximum := logits[0]
	for _, logit := range logits {
		if math.IsNaN(logit) || math.IsInf(logit, 0) {
			return nil, errors.New("every logit must be finite")
		}
		if logit > maximum {
			maximum = logit
		}
	}

	weights := make([]float64, len(logits))
	total := 0.0
	for i, logit := range logits {
		weights[i] = math.Exp((logit - maximum) / temperature)
		total += weights[i]
	}
	probabilities := make([]Probability, len(logits))
	for i, logit := range logits {
		probabilities[i] = Probability{Logit: logit, Probability: weights[i] / total}
	}
	return probabilities, nil
}

func main() {
	logits := []float64{2.4, 1.1, 0.2}
	for _, temperature := range []float64{0.5, 1, 2} {
		probabilities, err := Softmax(logits, temperature)
		if err != nil {
			panic(err)
		}
		fmt.Printf("temperature %.1f:", temperature)
		for _, item := range probabilities {
			fmt.Printf(" %.4f", item.Probability)
		}
		fmt.Println()
	}
}
08 / Check what the result means

A normalized distribution is internally consistent; it can still be wrong about the world.

Softmax always assigns probabilities over the candidate set it receives. That property does not show that the set contains the right answer, that model probabilities match real frequencies, or that the model's text is truthful. A model can be sharply concentrated on a wrong next token. Correctness is an outcome to measure; calibration asks whether predictions assigned a given probability are correct at approximately that rate over an appropriate population.

To investigate the repetitive-output incident, reproduce a fixed prompt set, record raw logits and all decoding controls, then change one control at a time. Compare token diversity and answer-level evaluation against reviewed outcomes. If logits themselves changed, inspect model version, prompt construction, tokenization, and input data. If only temperature changed, expect a changed sampling distribution while preserving score rank. Document filtering because it changes the candidate set seen by a sampler.

Transfer question: suppose every logit has 1,000 added. Before running code, predict whether probabilities change and whether naive exponentiation remains safe. The distribution is invariant to the shift; the direct exponential calculation becomes more overflow-prone.

References