The center can stay calm while the slowest users wait longer.
Imagine the same checkout endpoint, route, and region before and after a release. In two illustrative windows of 10,000 eligible requests, the median is 120 ms in both. The first window's p99 is 450 ms; the second's is 1,800 ms. These are teaching values, not production measurements. They describe a change worth investigating; they do not establish that the release caused it.
Start by asking how the requests divide up. Did a new database lookup affect only a route or cohort? Did retries, a cold cache, a downstream slowdown, or a burst create a few long waits? Did trace sampling or aggregation change? “The tail got slower” is an observation to explain, not yet a remedy.
- Population
- Eligible checkout requests in matched 10-minute windows.
- Sample count
- 10,000 per window, illustrative and assumed fully retained.
- Median
- 120 ms before and after the release.
- p99
- 450 ms before; 1,800 ms after, using the same percentile convention.
Latency is often asymmetric because requests take different paths.
A distribution records how often measurements fall at different values. For latency, values cannot drop below zero, but they can stretch far upward: most requests may finish quickly while a small group waits on a lock, a remote service, disk, garbage collection, a retry, or a queue. Mixing request types with different costs can create the same right-skewed shape even when no single request type has an extreme tail.
A long right tail is an operational description: comparatively few requests take much longer than the bulk. In statistics, heavy-tailed has more specific meanings about how tail probability decays. A handful of slow observations is not enough to prove a power law or any particular distribution family. Do not assume latency is normal, log-normal, or heavy-tailed without evidence and a stated model.
Cache hits and short queries cluster near the center.
Large payloads or expensive routes form another cohort.
Queues, locks, and downstream waits stretch some calls.
Retries or coincident slow dependencies compound a request.
A percentile locates an observation; it does not explain the observations above it.
Sort n durations from smallest to largest. Under the nearest-rank convention,
the pth percentile is the value at one-based position ceil((p / 100) × n). With 10,000 requests, p99 is at position 9,900: about
99% of observations are at or below that value and the upper one percent are at or above
the boundary, subject to ties. Another tool may interpolate between values, so check its
convention before comparing results.
The mean answers a different question: total elapsed request time divided by request count. It is useful for aggregate demand and comparisons, but distant values pull it upward. The median (p50) is the midpoint of ordered observations and is less moved by a few extremes. Neither statistic is “the real latency.” A percentile is not the maximum, and p99 is not a guarantee that 99% of future requests will meet that value.
Before comparing two tails, make sure they describe comparable requests.
Averages and percentiles inherit the choices in the telemetry pipeline. A dashboard might use server processing time while a client measures end-to-end wait. A trace collector might sample one request class more heavily, drop overloaded periods, or merge distributions into sketches. Those choices can shift the visible tail. Record the service boundary, units, time window, request count, inclusion rules, and aggregation method beside the statistic.
Compare like with like first: same endpoint or named cohort, comparable request mix, equal window duration, same percentile implementation, and similar measurement coverage. Then split by region, route, payload class, downstream dependency, and retry outcome. A change in the aggregate can come from a change in request composition, a change inside one cohort, or both.
Tail latency is a symptom to investigate; average demand still matters for capacity.
A slow request can occupy a worker, connection, or lock longer, keeping resources busy while other requests arrive. If arrivals bunch up or the service is near its sustainable processing limit, a queue can grow and add waiting time to later requests. That feedback makes latency especially sensitive near saturation. But a request's end-to-end latency is not automatically its CPU service time: time spent waiting on a dependency or network may consume a different resource. Measure arrival rate, resource service demand, concurrency, queue depth, and utilization for the resource you are sizing.
The mean helps estimate aggregate work when paired with throughput and resource cost. A tail percentile helps assess whether a slow fraction threatens a user-facing objective. Neither alone gives safe capacity. Use load tests and production evidence across the expected workload, including bursts and dependency behavior, and keep headroom for variability. A p99 crossing a latency objective may justify a page or rollout pause under an agreed policy; it does not by itself say whether the right fix is more capacity, less contention, fewer retries, or isolation of a costly route.
See how the same slow share looks at different sample sizes.
This small discrete model creates a sample where most requests take 100 ms and a chosen share takes longer. Change the sample size and tail share. Watch the mean, p99, and maximum respond. It is a hand-checkable model of order statistics, not a realistic simulation of a service or a prediction about your traffic.
At this sample size the p99 rank still lands among the 990 fast requests. The slow observations are real in this model, but p99 does not include every request above its threshold. The maximum reveals one endpoint, not how often it occurs.
Summarize observations without hiding the conventions.
The TypeScript and Go examples use the same deterministic observations: 990 requests at 100 ms and 10 at 2,000 ms. They report the mean, nearest-rank p99, maximum, sample count, and number at or above a 1,000 ms threshold. The values are synthetic so every result can be checked by hand.
Both examples use nearest rank and preserve the count behind the reported tail.
export type LatencySummary = {
count: number;
meanMs: number;
p99Ms: number;
maxMs: number;
aboveThreshold: number;
};
/** Nearest-rank quantile summary for one complete, retained sample. */
export function summarizeLatency(observationsMs: number[], thresholdMs: number): LatencySummary {
if (observationsMs.length === 0) {
throw new Error('at least one observation is required');
}
if (!Number.isFinite(thresholdMs) || thresholdMs < 0) {
throw new Error('thresholdMs must be finite and non-negative');
}
for (const value of observationsMs) {
if (!Number.isFinite(value) || value < 0) {
throw new Error('latencies must be finite and non-negative');
}
}
const sorted = [...observationsMs].sort((a, b) => a - b);
const totalMs = sorted.reduce((sum, value) => sum + value, 0);
const p99Index = Math.ceil(0.99 * sorted.length) - 1;
return {
count: sorted.length,
meanMs: totalMs / sorted.length,
p99Ms: sorted[p99Index],
maxMs: sorted[sorted.length - 1],
aboveThreshold: sorted.filter((value) => value >= thresholdMs).length
};
}
const observations = [...Array<number>(990).fill(100), ...Array<number>(10).fill(2000)];
const summary = summarizeLatency(observations, 1000);
console.log(summary);
// { count: 1000, meanMs: 119, p99Ms: 100, maxMs: 2000, aboveThreshold: 10 }
package main
import (
"errors"
"fmt"
"math"
"sort"
)
type LatencySummary struct {
Count int
MeanMs float64
P99Ms float64
MaxMs float64
AboveThreshold int
}
// SummarizeLatency uses the nearest-rank p99 convention on a complete sample.
func SummarizeLatency(observationsMs []float64, thresholdMs float64) (LatencySummary, error) {
if len(observationsMs) == 0 {
return LatencySummary{}, errors.New("at least one observation is required")
}
if math.IsNaN(thresholdMs) || math.IsInf(thresholdMs, 0) || thresholdMs < 0 {
return LatencySummary{}, errors.New("thresholdMs must be finite and non-negative")
}
values := append([]float64(nil), observationsMs...)
var total float64
aboveThreshold := 0
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return LatencySummary{}, errors.New("latencies must be finite and non-negative")
}
total += value
if value >= thresholdMs {
aboveThreshold++
}
}
sort.Float64s(values)
p99Index := int(math.Ceil(0.99*float64(len(values)))) - 1
return LatencySummary{
Count: len(values),
MeanMs: total / float64(len(values)),
P99Ms: values[p99Index],
MaxMs: values[len(values)-1],
AboveThreshold: aboveThreshold,
}, nil
}
func main() {
observations := make([]float64, 0, 1000)
for i := 0; i < 990; i++ {
observations = append(observations, 100)
}
for i := 0; i < 10; i++ {
observations = append(observations, 2000)
}
summary, err := SummarizeLatency(observations, 1000)
if err != nil {
panic(err)
}
fmt.Printf("count=%d mean=%.1fms p99=%.0fms max=%.0fms above=%.0fms:%d\n",
summary.Count, summary.MeanMs, summary.P99Ms, summary.MaxMs, 1000.0, summary.AboveThreshold)
// count=1000 mean=119.0ms p99=100ms max=2000ms above=1000ms:10
}
Make the next query capable of separating plausible causes.
For the checkout example, first verify equivalent windows and retained counts. Then compare latency distributions by route, region, payload, and downstream call. Align those cohorts with queue depth, pool waits, retries, resource utilization, and release changes. If one cohort accounts for the tail, investigate its path. If all cohorts shift together, inspect shared resources and dependencies. If only the telemetry method changed, reproduce the comparison on consistent observations.
A useful incident note separates observation from inference: “In matched 10-minute windows of 10,000 eligible requests, median latency stayed at 120 ms while nearest-rank p99 moved from 450 ms to 1,800 ms. The cause is unknown. We are checking route cohorts, queue waits, and trace coverage before changing concurrency.” That statement is useful because the evidence and the unknown are both explicit.