← Math in Practice
Concept Estimates, intervals, and visual evidence

Error bars and uncertainty communication

A chart can make uncertainty visible only when the reader knows what the bars encode.

A release chart shows that the share of failed requests rose from 2.4% to 4.8%. Each point has a vertical line, but the dashboard caption only says “error bars.” Those bars could show run-to-run spread, uncertainty in an estimate, or a chosen range. The picture has no useful statistical meaning until its construction is named.

The judgment to keep

Pair every interval with its estimate, denominator, population, time window, and method. Error bars describe one defined source of variation or uncertainty; they do not explain causes or make two groups automatically comparable.

TypeScriptGo Proportion estimates · Wilson score interval · chart labels · comparison limits
01 / Read the chart contract

“Error bar” is a drawing, not a statistical definition.

An interval drawn around a point can represent a range of observed values, standard deviation, standard error, a confidence interval, a prediction interval, or a credible interval. These answer different questions. A standard deviation describes spread among individual observations. A confidence interval describes uncertainty in an estimated parameter under a sampling model. A prediction interval concerns a future observation or outcome.

Suppose a service team counts failed requests during two 30-minute release windows. In each window, “failure” means a completed eligible request returned a 5xx or timed out. The old build had 24 failures among 1,000 eligible requests (2.4%); the candidate had 48 among 1,000 (4.8%). The interval here will be a 95% Wilson score interval for each binomial failure proportion. These are authored examples, not real telemetry.

The chart contract also names the target population: eligible requests served in that region during the specified windows. A percentage without this denominator can hide load changes, filtered requests, or changes in the failure definition.

Case file / Release comparisonTwo rates, both with denominators.
Old build
24 failures / 1,000 eligible requests = 2.4%
Candidate
48 failures / 1,000 eligible requests = 4.8%
Measurement unit
Request-level binary failure outcome
Interval method
95% Wilson score interval for each proportion
02 / Build a proportion interval

For a request failure rate, count events and eligible trials.

For a binomial outcome, each eligible request contributes one trial, and each failure contributes one success in the counted event. The sample estimate is p = x/n. Here the old rate is 24/1,000 = 0.024 = 2.4%. The point estimate is exact arithmetic for the sample; uncertainty arises when we use it to say something about a broader request population.

This lesson uses the Wilson score interval, a useful interval for a binomial proportion that behaves better than the simple estimate ± a normal standard error in small samples and near 0 or 1. With x=24, n=1,000, and z=1.96, it is approximately 1.61% to 3.56%. For 48/1,000, it is approximately 3.64% to 6.32%. The 1.96 value is the approximate standard-normal critical value for a two-sided 95% interval.

Old · 24 of 1,0002.40% · about [1.61%, 3.56%]
Candidate · 48 of 1,0004.80% · about [3.64%, 6.32%]
03 / Write the caption

Put the interpretation beside the graphic.

A useful caption can be short while still being auditable: “Request-level 5xx-or-timeout share among eligible requests, per 30-minute window; points are sample proportions and bars are two-sided 95% Wilson score intervals; old 24/1,000, candidate 48/1,000.” Then specify how requests were sampled or whether the telemetry is a complete census of the window.

Do not round the point and endpoints so aggressively that the difference disappears or precision is exaggerated. Keep enough precision to reproduce the calculation, but avoid implying that the interval endpoints are known with infinite exactness. If the graph is printed without its caption or denominator, it becomes much easier to misread.

Caption draftFailure share among eligible requests
Point
Sample proportion per build
Bars
Two-sided 95% Wilson score interval
Denominator
Eligible completed requests, n=1,000 per build
Window
Separate 30-minute windows; confirm traffic and region match
04 / Compare two estimates

Separate intervals on a chart are not the comparison itself.

The visual intervals nearly touch in this illustration. That can attract attention, but “the bars overlap” is not a universal test of whether two rates differ. Separate confidence intervals address uncertainty for each group; a direct comparison needs an interval or test for the difference, using a model appropriate to the data and experiment.

Even a well-calculated interval cannot decide whether a change matters operationally. A rise from 2.4% to 4.8% is a 2.4 percentage-point increase and a 100% relative increase, but impact also depends on total traffic, severity, retry behavior, and the release's user-facing objective. Establish a practical threshold before staring at the plot.

Diagnostic pause: if candidate traffic doubled and the failure count doubled too, what would happen to the failure share? Which would you put in a release decision: count, rate, or both?
05 / Challenge the sample

A million requests can still act like only a few independent observations.

Requests in one deployment window may share host load, a dependency incident, a customer cohort, or a single routing change. If failures cluster by host, zone, or minute, a calculation that treats every request as independent can make the interval too narrow. Use the experimental unit and sampling design: perhaps compare deployment blocks, zones, or randomized cohorts rather than pretending the individual request count captures all uncertainty.

Also check whether “eligible” changed, missing telemetry could be related to failure, retries create multiple counted attempts per user operation, or the dashboard selects only successful traces. A proportion describes the outcome as defined and observed. It does not tell you why requests failed, nor does a sample interval contain uncertainty about a misspecified population or biased collection pipeline.

06 / Practice in code

Make the event definition and denominator explicit.

The helper accepts integer event and trial counts and returns percentages. It rejects impossible counts and uses the Wilson score formula. It does not inspect request eligibility, dependence, region, time windows, or causal assignment; those belong in the measurement design. Keep rates as proportions during calculation and convert to percentage only for display.

Compare the same interval calculation in TypeScript and Go.

Both versions use counts, validate the denominator, and label the interval method.

TypeScriptFailure proportion · Wilson score interval
interval.ts
export type Interval = { estimate: number; low: number; high: number; count: number };

/** Wilson score interval for a binomial proportion, shown as percentages. */
export function wilson(successes: number, trials: number, z = 1.96): Interval {
	if (
		!Number.isInteger(successes) ||
		!Number.isInteger(trials) ||
		trials < 1 ||
		successes < 0 ||
		successes > trials
	) {
		throw new Error('use integer counts with 0 <= successes <= trials and trials >= 1');
	}
	const p = successes / trials;
	const z2 = z * z;
	const denominator = 1 + z2 / trials;
	const center = (p + z2 / (2 * trials)) / denominator;
	const half = (z / denominator) * Math.sqrt((p * (1 - p)) / trials + z2 / (4 * trials * trials));
	return {
		estimate: p * 100,
		low: Math.max(0, center - half) * 100,
		high: Math.min(1, center + half) * 100,
		count: trials
	};
}
GoFailure proportion · Wilson score interval
interval.go
package interval

import (
	"errors"
	"math"
)

type Interval struct {
	Estimate, Low, High float64
	Count               int
}

// Wilson returns a score interval for a binomial proportion, expressed in percentages.
func Wilson(successes, trials int, z float64) (Interval, error) {
	if trials < 1 || successes < 0 || successes > trials || z <= 0 || math.IsNaN(z) || math.IsInf(z, 0) {
		return Interval{}, errors.New("require trials >= 1, 0 <= successes <= trials, and finite z > 0")
	}
	p := float64(successes) / float64(trials)
	z2 := z * z
	denom := 1 + z2/float64(trials)
	center := (p + z2/(2*float64(trials))) / denom
	half := (z / denom) * math.Sqrt(p*(1-p)/float64(trials)+z2/(4*float64(trials*trials)))
	return Interval{p * 100, math.Max(0, center-half) * 100, math.Min(1, center+half) * 100, trials}, nil
}
07 / Make the next measurement

Use the graph to decide what evidence is missing.

For the meaning and construction of confidence intervals, consult NIST/SEMATECH's confidence interval overview and intervals for a binomial proportion. This lesson uses the Wilson score interval and emphasizes the assumptions behind its interpretation.