← Math in Practice
Concept Change, uncertainty, and evidence

Mean, median, variance, and percentiles

Averages describe a center; they do not describe the whole tail.

A release dashboard says the service’s mean response time is still 195 milliseconds. That sounds stable—but a mean can stay the same while a group of users starts waiting two seconds. Compare what the measurements say before deciding whether the release is healthy.

The judgment to keep

Choose a statistic to answer a specific question about a measured population. Keep the request count, time window, and limits of that summary beside the number.

TypeScriptGo Latency samples · p95 · release comparison
01 / Start with the dashboard

These two windows have the same mean and different users.

Imagine each row is one window of 100 requests. In the steady window, every request takes 195 ms. In the second, 95 requests take 100 ms and five take 2,000 ms. Both means are 195 ms: the fast majority offsets the slow requests when all 100 values are added together.

100 requests per window · nearest-rank percentiles · population variance
WindowMeanMedianp95p99Variance
Steady: 100 × 195 ms195 ms195 ms195 ms195 ms0 ms²
Long tail: 95 × 100 ms, 5 × 2,000 ms195 ms100 ms100 ms2,000 ms171,475 ms²

The mean is the sum divided by the count, so both windows average to 195 ms. The median is the middle observation after sorting: it shows the typical request shifted from 195 ms to 100 ms. The table also shows why the mean alone cannot tell you whether the slow requests matter to users.

02 / Measure the spread

Spread matters when requests do not behave alike.

The range reports only the minimum and maximum. Variance uses every observation: subtract the mean from each value, square each difference, then average those squared differences. Squaring keeps fast and slow deviations from canceling and gives distant values more influence.

For the 100-request long-tail window above, population variance is 171,475 ms². Standard deviation is its square root, about 414 ms, which brings the spread back to latency units. The steady window has zero variance because every request takes the same time.

Variance is useful when comparing how dispersed two sets of measurements are, but squared units are hard to explain as a user-facing target. Standard deviation is still sensitive to extreme values and does not say what share of requests crossed a deadline. Use percentiles when the question is about a latency threshold.

03 / Read the tail

p95 is a threshold for most requests, not a report of the worst ones.

For nearest rank, sort n observations, calculate the one-based position ceil(0.95 × n), and select that value. With 100 observations, p95 is item 95. In the long-tail example, items 1 through 95 are 100 ms, so p95 is 100 ms; p99 reaches the 2,000 ms requests. A p95 tells you where the 95% threshold falls. It does not summarize the remaining five percent.

Nearest rank is one convention. Some tools interpolate between observations, so compare systems using the same definition. Do not average two reported p95 values to get a combined p95; use the underlying observations or a mergeable distribution summary whose approximation limits are known.

04 / Change the measurements

Move the tail and watch each summary respond.

Start with the same-mean example. The two windows both average 195 ms, but their p99 and variance tell a different story. Change one or two current measurements, then compare the median, p95, p99, and maximum. Try the small-sample example too: with only five readings, nearest-rank p95 selects the maximum, so one request controls the result.

Change the evidence

Compare two latency windows

Local calculation · nothing is saved

Edit either list or paste up to 100 non-negative measurements in milliseconds. Separate values with commas, spaces, or new lines. Results update as you type.

Comparing 100 baseline and 100 current measurements. Mean: 195 → 195 ms. p95: 195 → 100 ms. p99: 195 → 2,000 ms.

Sorted request times · bar height uses the same scale in both windows

Summary statistics · nearest-rank percentiles
MeasureBaselineCurrent
Count100100
Mean195 ms195 ms
Median (p50)195 ms100 ms
p95195 ms100 ms
p99195 ms2,000 ms
Population variance0 ms²171,475 ms²
Standard deviation0 ms414.1 ms
Maximum195 ms2,000 ms

A mean can stay flat while the tail changes. Read the count and distribution with each summary; this small sample calculator does not establish an SLO or explain why a request was slow.

05 / Make a release decision

Compare like with like, then investigate the change.

Compare the same request population, units, time window, and percentile convention. Relative change is (current − baseline) / baseline × 100. A release policy can use a stated threshold to flag a change for investigation—for example, a p95 increase greater than 15%.

A flag is a decision rule, not proof that the release caused the change. Check traffic mix, sample size, errors, and other changes in the same window. Use the threshold to decide when a person should look closer, not as a substitute for understanding what users experienced.

06 / Practice in code

Calculate p95 and use it in a release check.

Work through nearest-rank p95, a fuller latency summary, and a threshold comparison. Each ticket uses the same percentile rule in TypeScript and Go.

Math in Practice practice 8 min

Find the nearest-rank p95

This is an experiment with ticket-style exercises, giving beginners a feel for how tasks may be described in the workplace. Leave feedback

Checking your sign-in status. Your lesson remains available while we check.

TypeScript Go

Math in Practice practice 8 min

This is an experiment with ticket-style exercises, giving beginners a feel for how tasks may be described in the workplace. Leave feedback

Find the nearest-rank p95

TYPESCRIPT

Work item OBS-104

Find the nearest-rank p95

Implementation exercise Ready

Context

A service dashboard needs a p95 from one window of request latencies. The samples arrive unsorted, and the dashboard must use one stated percentile convention so two implementations agree.

Acceptance criteria
  1. AC-1Sort a copy of the samples so caller order is unchanged.
  2. AC-2Use nearest rank: rank = ceil(0.95 × sample count), with rank counted from one.
  3. AC-3Return the sample at that rank, including the maximum when the sample is small.
Notes
  • All inputs are non-empty latency measurements in the same unit.
  • Nearest rank is one explicit convention; interpolated percentiles may produce a different number.
  • A p95 is a threshold statistic, not the average request time.

Copy the ticket to research the problem in your own notes or AI tool. Your code stays here.

Your implementation

Edit the function in the editor. Run the visible checks as often as you like; your code stays in this tab.

Checks cover

  • Unsorted samples use nearest rank
  • Twenty samples select rank nineteen
  • A small sample can place p95 at the maximum