A lower perplexity tells us something specific about the observed tokens.
A language model predicts one token at a time. For each position, it assigns probabilities to possible next tokens; the actual token in the held-out text contributes a score. An evaluator may report average negative log probability or perplexity across many positions. Those values can help compare predictive fit when the tokenizer, text population, masking rules, and aggregation method are held consistent.
They do not answer whether a generated paragraph is correct, useful, safe, or appropriate. A model may predict common connective words well while inventing a key fact. First inspect the metric definition and evaluation sample, then decide what separate answer-level evidence is needed for the product claim.
- Metric population
- Held-out next-token examples from one fixed benchmark.
- Reported change
- Lower average token prediction loss after a model update.
- Observed disagreement
- Reviewers still find plausible-sounding factual errors.
- Next diagnosis
- Check token evaluation comparability and measure reviewed answer correctness separately.
A probability becomes an additive score when we take its logarithm.
If an event has probability p, its surprisal is −log₂(p) bits.
An unlikely event carries more information when it occurs: probability 1/2 gives 1 bit, and probability 1/8 gives 3 bits. If the observed next token is
“cache” and a model assigned it probability 0.2, its negative log probability
is −log₂(0.2) ≈ 2.3219 bits.
The logarithm turns a product into a sum: log(ab) = log(a) + log(b). If a
model assigns successive tokens probabilities p₁, …, pₙ and we use the chain
rule's conditional probabilities, the sequence log probability is Σ log(pᵢ).
Summing negative log probabilities is easier to accumulate and compare than multiplying
many tiny values that can underflow.
One outcome from the illustrative three-token vocabulary.
The model's conditional next-token probability for this context.
Contribution from this single observed token.
Entropy is the average surprisal if outcomes follow a stated distribution.
Use a small reference distribution P = [0.7, 0.2, 0.1] over the outcomes
“deploy,” “cache,” and “queue.” Its entropy is H(P) = −Σ pᵢ log₂(pᵢ). Substituting gives −(0.7 log₂ 0.7 + 0.2 log₂ 0.2 + 0.1 log₂ 0.1) ≈ 1.1568 bits/outcome. This is
the average surprisal under P, not a statement that each event carries 1.1568 bits.
For three equally likely outcomes, entropy would be log₂(3) ≈ 1.585 bits. The
skew in P lowers the average: “deploy” is common and unsurprising, while “queue” is less
common and contributes more when it occurs. Entropy depends on the distribution and the
defined outcome vocabulary. In language, context changes that distribution from token to
token.
| Outcome | P(outcome) | −p log₂(p) |
|---|---|---|
| deploy | 0.7 | 0.3602 bits |
| cache | 0.2 | 0.4644 bits |
| queue | 0.1 | 0.3322 bits |
| Total entropy | 1.0 | 1.1568 bits/outcome |
Cross-entropy asks how surprising reference outcomes are under the model.
Let a model distribution be Q = [0.6, 0.3, 0.1] over the same outcomes.
Cross-entropy from target/reference P to model Q is H(P,Q) = −Σ pᵢ log₂(qᵢ). For this example: −(0.7 log₂ 0.6 + 0.2 log₂ 0.3 + 0.1 log₂ 0.1) ≈ 1.1955 bits/outcome.
The model assigns too little mass to “deploy” relative to P and more to “cache,” so it
pays a cross-entropy penalty. The identity H(P,Q) = H(P) + Dₖₗ(P ∥ Q) separates irreducible reference uncertainty from the model's distribution mismatch. Here the
excess is about 1.1955 − 1.1568 = 0.0387 bits/outcome.
Expected surprisal when scoring with P itself.
Expected code length when outcomes follow P but are scored with Q.
Cross-entropy minus reference entropy, in this direction.
Exponentiating average loss expresses it as an effective number of choices.
For cross-entropy measured in bits per token, perplexity is 2ᴴ. Here, 2¹·¹⁹⁵⁵ ≈ 2.2902 effective choices per outcome. This is
a compact way to express average token uncertainty under this scoring setup; it is not a literal
count of candidates the model considered and not the probability that a complete answer is right.
If using nats instead, perplexity is eᴴ. The bases must match: do not
exponentiate a bit-valued loss with e. For text evaluation, report which
tokens count, how losses are averaged, and which tokenizer defines a token. Different
tokenization changes the unit of “per token,” so perplexities across different tokenizers
are not directly comparable.
KL divergence measures the extra expected log loss from using Q when P describes outcomes.
The Kullback–Leibler divergence in bits is Dₖₗ(P ∥ Q) = Σ pᵢ log₂(pᵢ/qᵢ). With our numbers, the terms are 0.7 log₂(0.7/0.6) + 0.2 log₂(0.2/0.3) + 0.1 log₂(0.1/0.1) ≈ 0.0387 bits. It
is non-negative for valid distributions and equals zero when P and Q match.
Direction matters: Dₖₗ(P ∥ Q) generally differs from Dₖₗ(Q ∥ P). It is not a symmetric distance. If Pᵢ > 0 while Qᵢ = 0, the divergence is infinite: the model says an outcome that occurs
under P is impossible. In finite datasets, smoothing or model support choices matter, so
inspect them before interpreting a very large score.
| Outcome | Reference P | Model Q | P log₂(P/Q) |
|---|---|---|---|
| deploy | 0.7 | 0.6 | +0.1557 bits |
| cache | 0.2 | 0.3 | −0.1170 bits |
| queue | 0.1 | 0.1 | 0 bits |
| Total KL | Dₖₗ(P ∥ Q) | ≈ 0.0387 bits/outcome | |
Choose which held-out token occurred and see its individual surprisal.
The model distribution stays fixed at Q = [0.6, 0.3, 0.1]. Select an observed
token to calculate its negative log probability. The entropy, cross-entropy, KL, and
perplexity cards summarize the complete illustrative distributions and do not change with
this one-token choice.
The values use three illustrative outcomes, and the displayed precision is rounded. A real evaluation averages over a defined set of token positions.
Name the base and handle impossible events deliberately.
The examples calculate entropy, cross-entropy, KL divergence, and perplexity using base-two logarithms. They validate normalized distributions, ignore zero-probability terms when the target mass is zero, and return infinity when the model assigns zero probability to an observed outcome with positive target mass. A production evaluator must additionally define masks, tokenization, sample weighting, and aggregation.
Both snippets use the same reference and model distributions and report base-two units.
export type Distribution = readonly number[];
function validate(distribution: Distribution, name: string): void {
if (distribution.length === 0) throw new RangeError(`${name} must not be empty`);
if (distribution.some((p) => !Number.isFinite(p) || p < 0 || p > 1)) {
throw new RangeError(`${name} probabilities must be finite and between zero and one`);
}
const total = distribution.reduce((sum, p) => sum + p, 0);
if (Math.abs(total - 1) > 1e-9) throw new RangeError(`${name} probabilities must sum to one`);
}
/** Shannon entropy measured in bits per outcome. */
export function entropyBits(probabilities: Distribution): number {
validate(probabilities, 'probabilities');
return -probabilities.reduce((sum, p) => sum + (p === 0 ? 0 : p * Math.log2(p)), 0);
}
/** Cross-entropy in bits per outcome; infinity if target mass meets a zero model probability. */
export function crossEntropyBits(target: Distribution, model: Distribution): number {
validate(target, 'target');
validate(model, 'model');
if (target.length !== model.length)
throw new RangeError('target and model must have equal length');
return target.reduce((sum, p, i) => {
if (p === 0) return sum;
if (model[i] === 0) return Number.POSITIVE_INFINITY;
return sum - p * Math.log2(model[i]);
}, 0);
}
/** KL(target || model) in bits per outcome. */
export function klBits(target: Distribution, model: Distribution): number {
validate(target, 'target');
validate(model, 'model');
if (target.length !== model.length)
throw new RangeError('target and model must have equal length');
return target.reduce((sum, p, i) => {
if (p === 0) return sum;
if (model[i] === 0) return Number.POSITIVE_INFINITY;
return sum + p * Math.log2(p / model[i]);
}, 0);
}
export function perplexityFromBits(crossEntropy: number): number {
if (Number.isNaN(crossEntropy) || crossEntropy < 0) {
throw new RangeError('crossEntropy must be non-negative');
}
return 2 ** crossEntropy;
}
export const referenceDistribution = [0.7, 0.2, 0.1] as const;
export const modelDistribution = [0.6, 0.3, 0.1] as const;
package main
import (
"errors"
"fmt"
"math"
)
func validate(probabilities []float64) error {
if len(probabilities) == 0 {
return errors.New("distribution must not be empty")
}
total := 0.0
for _, p := range probabilities {
if math.IsNaN(p) || math.IsInf(p, 0) || p < 0 || p > 1 {
return errors.New("probabilities must be finite and between zero and one")
}
total += p
}
if math.Abs(total-1) > 1e-9 {
return errors.New("probabilities must sum to one")
}
return nil
}
// EntropyBits returns Shannon entropy in bits per outcome; zero terms contribute zero.
func EntropyBits(probabilities []float64) (float64, error) {
if err := validate(probabilities); err != nil {
return 0, err
}
entropy := 0.0
for _, p := range probabilities {
if p > 0 {
entropy -= p * math.Log2(p)
}
}
return entropy, nil
}
// CrossEntropyBits returns target cross-entropy in bits per outcome.
func CrossEntropyBits(target, model []float64) (float64, error) {
if err := validate(target); err != nil {
return 0, fmt.Errorf("target: %w", err)
}
if err := validate(model); err != nil {
return 0, fmt.Errorf("model: %w", err)
}
if len(target) != len(model) {
return 0, errors.New("target and model must have equal length")
}
crossEntropy := 0.0
for i, p := range target {
if p == 0 {
continue
}
if model[i] == 0 {
return math.Inf(1), nil
}
crossEntropy -= p * math.Log2(model[i])
}
return crossEntropy, nil
}
// KLDivergenceBits returns KL(target || model) in bits per outcome.
func KLDivergenceBits(target, model []float64) (float64, error) {
if err := validate(target); err != nil {
return 0, fmt.Errorf("target: %w", err)
}
if err := validate(model); err != nil {
return 0, fmt.Errorf("model: %w", err)
}
if len(target) != len(model) {
return 0, errors.New("target and model must have equal length")
}
divergence := 0.0
for i, p := range target {
if p == 0 {
continue
}
if model[i] == 0 {
return math.Inf(1), nil
}
divergence += p * math.Log2(p/model[i])
}
return divergence, nil
}
func main() {
target := []float64{0.7, 0.2, 0.1}
model := []float64{0.6, 0.3, 0.1}
entropy, err := EntropyBits(target)
if err != nil {
panic(err)
}
crossEntropy, err := CrossEntropyBits(target, model)
if err != nil {
panic(err)
}
divergence, err := KLDivergenceBits(target, model)
if err != nil {
panic(err)
}
fmt.Printf("entropy=%.4f bits, cross-entropy=%.4f bits, KL=%.4f bits, perplexity=%.4f\n",
entropy, crossEntropy, divergence, math.Pow(2, crossEntropy))
}
Model confidence describes assigned probability, not whether a claim is true.
Cross-entropy is sensitive to probability assigned to observed targets. A single low-probability target can contribute a large loss; changing tokenization, masks, truncation, or which examples are averaged changes the measured population. Before reading a reported change as a model improvement or regression, verify the evaluation pipeline and compare identical examples with identical scoring rules.
For the opening release review, reproduce the benchmark, compare per-example and per-token loss contributions, inspect the tokenizer and data split, and check for contamination between training and evaluation. Then run reviewed answer-level tasks that assess factual correctness and the product behavior users need. Calibration is another question: across examples assigned similar confidence, do outcomes occur at similar frequencies? A lower average loss does not answer that question by itself.
References
- Claude Shannon, “A Mathematical Theory of Communication” — information, entropy, and communication.
- Jurafsky and Martin, Speech and Language Processing, Chapter 3 draft — language-model probability and perplexity concepts.
- Google Machine Learning Crash Course: Log loss — negative log probability as a loss for observed labels.