← Math in Practice
Concept Math behind AI

Log-probabilities and information theory

Read what a model's probability scores say about observed tokens, and where that evidence stops.

A model update reduces perplexity on a benchmark, yet reviewers still find confident factual errors in generated answers. The evaluation number changed, but it measures token prediction, not truth directly. To interpret the report, trace how probability becomes log loss, how losses average into cross-entropy, and what a comparison between distributions can establish.

The judgment to keep

Log probability makes products of token probabilities additive. In base 2, the negative log of a probability is surprisal in bits. Average surprisal on observed outcomes is cross-entropy; perplexity is that average converted back to an effective number of choices. These are measures of predictive fit on a specified population, not direct measures of factual correctness.

TypeScriptGo Log probability · surprisal · entropy · cross-entropy · perplexity · KL divergence
01 / Read the evaluation report

A lower perplexity tells us something specific about the observed tokens.

A language model predicts one token at a time. For each position, it assigns probabilities to possible next tokens; the actual token in the held-out text contributes a score. An evaluator may report average negative log probability or perplexity across many positions. Those values can help compare predictive fit when the tokenizer, text population, masking rules, and aggregation method are held consistent.

They do not answer whether a generated paragraph is correct, useful, safe, or appropriate. A model may predict common connective words well while inventing a key fact. First inspect the metric definition and evaluation sample, then decide what separate answer-level evidence is needed for the product claim.

Case file / Model release reviewPerplexity fell; factual review did not improve.
Metric population
Held-out next-token examples from one fixed benchmark.
Reported change
Lower average token prediction loss after a model update.
Observed disagreement
Reviewers still find plausible-sounding factual errors.
Next diagnosis
Check token evaluation comparability and measure reviewed answer correctness separately.
02 / Measure one outcome

A probability becomes an additive score when we take its logarithm.

If an event has probability p, its surprisal is −log₂(p) bits. An unlikely event carries more information when it occurs: probability 1/2 gives 1 bit, and probability 1/8 gives 3 bits. If the observed next token is “cache” and a model assigned it probability 0.2, its negative log probability is −log₂(0.2) ≈ 2.3219 bits.

The logarithm turns a product into a sum: log(ab) = log(a) + log(b). If a model assigns successive tokens probabilities p₁, …, pₙ and we use the chain rule's conditional probabilities, the sequence log probability is Σ log(pᵢ). Summing negative log probabilities is easier to accumulate and compare than multiplying many tiny values that can underflow.

01 / Observed token“cache”

One outcome from the illustrative three-token vocabulary.

02 / Assigned probabilityQ(cache) = 0.2

The model's conditional next-token probability for this context.

03 / Surprisal−log₂(0.2) ≈ 2.3219 bits

Contribution from this single observed token.

03 / Describe uncertainty

Entropy is the average surprisal if outcomes follow a stated distribution.

Use a small reference distribution P = [0.7, 0.2, 0.1] over the outcomes “deploy,” “cache,” and “queue.” Its entropy is H(P) = −Σ pᵢ log₂(pᵢ). Substituting gives −(0.7 log₂ 0.7 + 0.2 log₂ 0.2 + 0.1 log₂ 0.1) ≈ 1.1568 bits/outcome. This is the average surprisal under P, not a statement that each event carries 1.1568 bits.

For three equally likely outcomes, entropy would be log₂(3) ≈ 1.585 bits. The skew in P lowers the average: “deploy” is common and unsurprising, while “queue” is less common and contributes more when it occurs. Entropy depends on the distribution and the defined outcome vocabulary. In language, context changes that distribution from token to token.

Reference probabilities; bit contributions rounded to four decimals.
OutcomeP(outcome)−p log₂(p)
deploy0.70.3602 bits
cache0.20.4644 bits
queue0.10.3322 bits
Total entropy1.01.1568 bits/outcome
04 / Score model predictions

Cross-entropy asks how surprising reference outcomes are under the model.

Let a model distribution be Q = [0.6, 0.3, 0.1] over the same outcomes. Cross-entropy from target/reference P to model Q is H(P,Q) = −Σ pᵢ log₂(qᵢ). For this example: −(0.7 log₂ 0.6 + 0.2 log₂ 0.3 + 0.1 log₂ 0.1) ≈ 1.1955 bits/outcome.

The model assigns too little mass to “deploy” relative to P and more to “cache,” so it pays a cross-entropy penalty. The identity H(P,Q) = H(P) + Dₖₗ(P ∥ Q) separates irreducible reference uncertainty from the model's distribution mismatch. Here the excess is about 1.1955 − 1.1568 = 0.0387 bits/outcome.

01 / Reference uncertaintyH(P) ≈ 1.1568 bits

Expected surprisal when scoring with P itself.

02 / Model scoreH(P,Q) ≈ 1.1955 bits

Expected code length when outcomes follow P but are scored with Q.

03 / Distribution mismatchDₖₗ(P ∥ Q) ≈ 0.0387 bits

Cross-entropy minus reference entropy, in this direction.

05 / Translate to perplexity

Exponentiating average loss expresses it as an effective number of choices.

For cross-entropy measured in bits per token, perplexity is 2ᴴ. Here, 2¹·¹⁹⁵⁵ ≈ 2.2902 effective choices per outcome. This is a compact way to express average token uncertainty under this scoring setup; it is not a literal count of candidates the model considered and not the probability that a complete answer is right.

If using nats instead, perplexity is eᴴ. The bases must match: do not exponentiate a bit-valued loss with e. For text evaluation, report which tokens count, how losses are averaged, and which tokenizer defines a token. Different tokenization changes the unit of “per token,” so perplexities across different tokenizers are not directly comparable.

06 / Compare two distributions

KL divergence measures the extra expected log loss from using Q when P describes outcomes.

The Kullback–Leibler divergence in bits is Dₖₗ(P ∥ Q) = Σ pᵢ log₂(pᵢ/qᵢ). With our numbers, the terms are 0.7 log₂(0.7/0.6) + 0.2 log₂(0.2/0.3) + 0.1 log₂(0.1/0.1) ≈ 0.0387 bits. It is non-negative for valid distributions and equals zero when P and Q match.

Direction matters: Dₖₗ(P ∥ Q) generally differs from Dₖₗ(Q ∥ P). It is not a symmetric distance. If Pᵢ > 0 while Qᵢ = 0, the divergence is infinite: the model says an outcome that occurs under P is impossible. In finite datasets, smoothing or model support choices matter, so inspect them before interpreting a very large score.

Same three outcomes, different mass assignments.
OutcomeReference PModel QP log₂(P/Q)
deploy0.70.6+0.1557 bits
cache0.20.3−0.1170 bits
queue0.10.10 bits
Total KLDₖₗ(P ∥ Q)≈ 0.0387 bits/outcome
07 / Inspect an observed token

Choose which held-out token occurred and see its individual surprisal.

The model distribution stays fixed at Q = [0.6, 0.3, 0.1]. Select an observed token to calculate its negative log probability. The entropy, cross-entropy, KL, and perplexity cards summarize the complete illustrative distributions and do not change with this one-token choice.

Single-token surprisal1.7370 bits−log₂(0.3) for “cache”
Reference entropy1.1568 bitsH(P)
Cross-entropy1.1955 bitsH(P,Q)
KL divergence0.0387 bitsDₖₗ(P ∥ Q)
Perplexity2.29022 raised to cross-entropy in bits

The values use three illustrative outcomes, and the displayed precision is rounded. A real evaluation averages over a defined set of token positions.

08 / Practice in code

Name the base and handle impossible events deliberately.

The examples calculate entropy, cross-entropy, KL divergence, and perplexity using base-two logarithms. They validate normalized distributions, ignore zero-probability terms when the target mass is zero, and return infinity when the model assigns zero probability to an observed outcome with positive target mass. A production evaluator must additionally define masks, tokenization, sample weighting, and aggregation.

Compare the information measures in TypeScript and Go.

Both snippets use the same reference and model distributions and report base-two units.

TypeScriptCalculate entropy and distribution mismatch
information.ts
export type Distribution = readonly number[];

function validate(distribution: Distribution, name: string): void {
	if (distribution.length === 0) throw new RangeError(`${name} must not be empty`);
	if (distribution.some((p) => !Number.isFinite(p) || p < 0 || p > 1)) {
		throw new RangeError(`${name} probabilities must be finite and between zero and one`);
	}
	const total = distribution.reduce((sum, p) => sum + p, 0);
	if (Math.abs(total - 1) > 1e-9) throw new RangeError(`${name} probabilities must sum to one`);
}

/** Shannon entropy measured in bits per outcome. */
export function entropyBits(probabilities: Distribution): number {
	validate(probabilities, 'probabilities');
	return -probabilities.reduce((sum, p) => sum + (p === 0 ? 0 : p * Math.log2(p)), 0);
}

/** Cross-entropy in bits per outcome; infinity if target mass meets a zero model probability. */
export function crossEntropyBits(target: Distribution, model: Distribution): number {
	validate(target, 'target');
	validate(model, 'model');
	if (target.length !== model.length)
		throw new RangeError('target and model must have equal length');
	return target.reduce((sum, p, i) => {
		if (p === 0) return sum;
		if (model[i] === 0) return Number.POSITIVE_INFINITY;
		return sum - p * Math.log2(model[i]);
	}, 0);
}

/** KL(target || model) in bits per outcome. */
export function klBits(target: Distribution, model: Distribution): number {
	validate(target, 'target');
	validate(model, 'model');
	if (target.length !== model.length)
		throw new RangeError('target and model must have equal length');
	return target.reduce((sum, p, i) => {
		if (p === 0) return sum;
		if (model[i] === 0) return Number.POSITIVE_INFINITY;
		return sum + p * Math.log2(p / model[i]);
	}, 0);
}

export function perplexityFromBits(crossEntropy: number): number {
	if (Number.isNaN(crossEntropy) || crossEntropy < 0) {
		throw new RangeError('crossEntropy must be non-negative');
	}
	return 2 ** crossEntropy;
}

export const referenceDistribution = [0.7, 0.2, 0.1] as const;
export const modelDistribution = [0.6, 0.3, 0.1] as const;
GoCalculate entropy and distribution mismatch
information.go
package main

import (
	"errors"
	"fmt"
	"math"
)

func validate(probabilities []float64) error {
	if len(probabilities) == 0 {
		return errors.New("distribution must not be empty")
	}
	total := 0.0
	for _, p := range probabilities {
		if math.IsNaN(p) || math.IsInf(p, 0) || p < 0 || p > 1 {
			return errors.New("probabilities must be finite and between zero and one")
		}
		total += p
	}
	if math.Abs(total-1) > 1e-9 {
		return errors.New("probabilities must sum to one")
	}
	return nil
}

// EntropyBits returns Shannon entropy in bits per outcome; zero terms contribute zero.
func EntropyBits(probabilities []float64) (float64, error) {
	if err := validate(probabilities); err != nil {
		return 0, err
	}
	entropy := 0.0
	for _, p := range probabilities {
		if p > 0 {
			entropy -= p * math.Log2(p)
		}
	}
	return entropy, nil
}

// CrossEntropyBits returns target cross-entropy in bits per outcome.
func CrossEntropyBits(target, model []float64) (float64, error) {
	if err := validate(target); err != nil {
		return 0, fmt.Errorf("target: %w", err)
	}
	if err := validate(model); err != nil {
		return 0, fmt.Errorf("model: %w", err)
	}
	if len(target) != len(model) {
		return 0, errors.New("target and model must have equal length")
	}
	crossEntropy := 0.0
	for i, p := range target {
		if p == 0 {
			continue
		}
		if model[i] == 0 {
			return math.Inf(1), nil
		}
		crossEntropy -= p * math.Log2(model[i])
	}
	return crossEntropy, nil
}

// KLDivergenceBits returns KL(target || model) in bits per outcome.
func KLDivergenceBits(target, model []float64) (float64, error) {
	if err := validate(target); err != nil {
		return 0, fmt.Errorf("target: %w", err)
	}
	if err := validate(model); err != nil {
		return 0, fmt.Errorf("model: %w", err)
	}
	if len(target) != len(model) {
		return 0, errors.New("target and model must have equal length")
	}
	divergence := 0.0
	for i, p := range target {
		if p == 0 {
			continue
		}
		if model[i] == 0 {
			return math.Inf(1), nil
		}
		divergence += p * math.Log2(p/model[i])
	}
	return divergence, nil
}

func main() {
	target := []float64{0.7, 0.2, 0.1}
	model := []float64{0.6, 0.3, 0.1}
	entropy, err := EntropyBits(target)
	if err != nil {
		panic(err)
	}
	crossEntropy, err := CrossEntropyBits(target, model)
	if err != nil {
		panic(err)
	}
	divergence, err := KLDivergenceBits(target, model)
	if err != nil {
		panic(err)
	}
	fmt.Printf("entropy=%.4f bits, cross-entropy=%.4f bits, KL=%.4f bits, perplexity=%.4f\n",
		entropy, crossEntropy, divergence, math.Pow(2, crossEntropy))
}
09 / Diagnose the metric

Model confidence describes assigned probability, not whether a claim is true.

Cross-entropy is sensitive to probability assigned to observed targets. A single low-probability target can contribute a large loss; changing tokenization, masks, truncation, or which examples are averaged changes the measured population. Before reading a reported change as a model improvement or regression, verify the evaluation pipeline and compare identical examples with identical scoring rules.

For the opening release review, reproduce the benchmark, compare per-example and per-token loss contributions, inspect the tokenizer and data split, and check for contamination between training and evaluation. Then run reviewed answer-level tasks that assess factual correctness and the product behavior users need. Calibration is another question: across examples assigned similar confidence, do outcomes occur at similar frequencies? A lower average loss does not answer that question by itself.

Transfer question: if the model distribution exactly matches P, predict entropy, cross-entropy, KL, and perplexity. Then suppose the tokenizer changes but the raw text does not. Which comparisons remain valid without re-running the evaluation under a shared scoring convention?

References