The average of group averages treats every group as equally large.
During a 15-minute window, an interactive API handles 10,000 requests with a mean latency of 100 ms. A batch endpoint handles 1,000 requests with a mean of 500 ms. Assume every request is measured once, all durations use the same clock and definition, and these means are exact for the illustrative dataset.
The unweighted mean of group means is (100 + 500) / 2 = 300 ms. It answers a
different question: what is the average of the two cohort averages if both cohorts receive
equal weight? It does not describe the average request among the 11,000 requests observed.
When people say “overall average,” ask which objects are being averaged, what each denominator is, and whether the groups partition the same target window. Averages with no counts beside them are difficult to audit and easy to combine incorrectly.
- Interactive API
- 10,000 requests · mean 100 ms
- Batch endpoint
- 1,000 requests · mean 500 ms
- Unweighted group mean
- (100 + 500) / 2 = 300 ms; wrong target for all requests
- Combined population
- 11,000 requests in one stated 15-minute window
Multiply each mean by its count, add, and divide by the total count.
The request-weighted mean is Σ(nᵢ × x̄ᵢ) / Σnᵢ. Here the interactive requests
contribute an estimated total of 10,000 × 100 ms = 1,000,000 request·ms;
batch contributes 1,000 × 500 ms = 500,000 request·ms. Together that is
1,500,000 request·ms across 11,000 requests.
So the combined mean is 1,500,000 / 11,000 ≈ 136.36 ms/request. The units
reduce back to milliseconds per request. A sanity check: the answer must lie between 100
and 500 ms; because most requests came from the lower-latency API, it should be much
closer to 100 ms.
This works because a mean multiplied by its count recovers the sum of observations (subject to rounding). If the dashboard stores rounded means, the reconstructed sum is approximate. Prefer raw sums and counts when designing aggregations.
A weighted mean formula does not turn p95 values into a global p95.
A percentile is a position in an ordered distribution, not an additive total. The global 95th percentile depends on how values from all cohorts interleave. Group p95s omit most of that ordering information, so weighting or averaging the two p95 values cannot generally recover the p95 of the combined population.
For a tiny nearest-rank illustration, cohort A has observed latencies [10, 20, 30, 40, 50] ms, so its p95 is 50 ms; cohort B has [100, 200, 300, 400, 500] ms, p95 500
ms. Their p95 average is 275 ms. Pooling the ten raw observations gives sorted values
ending … 400, 500; at nearest-rank rank ceil(0.95 × 10) = 10,
the combined p95 is 500 ms. This small example exposes the loss, though production systems
use far more observations and may use different interpolation conventions.
To combine percentiles, merge raw values or merge a representation designed for quantile queries, such as compatible histograms or a mergeable sketch. Approximate structures trade memory and precision in documented ways; their buckets, units, and configuration must remain compatible.
- Cohort A p95
- 50 ms
- Cohort B p95
- 500 ms
- Average of the two p95s
- (50 + 500) / 2 = 275 ms
- Pooled p95
- Rank 10 of 10 combined values = 500 ms
Different questions need different mergeable evidence.
For a global arithmetic mean, retain sum and count. For a rate, retain event count and eligible-trial count, then divide. For a percentage, carry both numerator and denominator; do not average percentages from differently sized groups without weights. For a median or p95, retain raw values or a compatible distribution summary, not just the scalar percentile. For variance, the mean and count alone are insufficient: retain a variance-aware summary or mergeable moments.
Aggregation boundaries matter. If one request can appear in multiple endpoint cohorts, summing the counts double-counts the target population. If one region reports request-weighted latency and another reports user-weighted latency, the metrics do not combine into one interpretable quantity even if both are labeled “average latency.”
| Question | Useful retained evidence | Common trap |
|---|---|---|
| Mean of requests | Sum of values + request count | Equal-weighting group means |
| Failure rate | Failure count + eligible request count | Averaging cohort percentages |
| p95 latency | Raw values, compatible histogram, or mergeable sketch | Averaging group p95s |
| Variance | Count plus mergeable moments or raw observations | Using only mean and count |
Correct arithmetic cannot reconcile mismatched populations.
Before combining, verify the same unit (milliseconds, not a mixture of seconds and milliseconds), same outcome definition, same eligibility rules, same measurement window, compatible clock and sampling behavior, and mutually understood group membership. Ask whether the counts include retries, whether missing observations are excluded, and whether a high-volume cohort is overrepresented because it truly had more requests or because one pipeline sampled it differently.
A request-weighted global mean describes the average observed request. It gives large tenants and busy routes greater influence. If the question is “what does the typical tenant experience?”, calculate a tenant-level distribution or explicitly use equal tenant weights. Those are different estimands. Choose the weighting unit from the decision, not from whichever denominator is easiest to query.
Keep denominators in the data shape.
The code combines cohort means by count and implements a nearest-rank percentile from raw observations to make the information requirement concrete. The percentile helper copies before sorting and validates its input. It deliberately does not accept group percentiles as inputs: there is no general correct aggregation function for those scalar values alone.
Weighted means consume cohort counts; percentiles consume observations.
export type Cohort = { name: string; count: number; mean: number };
/** Combines cohort means while preserving each cohort's denominator. */
export function weightedMean(cohorts: Cohort[]): number {
if (
!cohorts.length ||
cohorts.some((c) => !Number.isSafeInteger(c.count) || c.count < 0 || !Number.isFinite(c.mean))
) {
throw new Error('provide cohorts with finite means and non-negative integer counts');
}
const total = cohorts.reduce((sum, c) => sum + c.count, 0);
if (total === 0) throw new Error('total count must be positive');
return cohorts.reduce((sum, c) => sum + c.count * c.mean, 0) / total;
}
/** Percentile of raw observations using nearest-rank convention (one-based rank). */
export function percentile(values: number[], p: number): number {
if (
!values.length ||
!Number.isFinite(p) ||
p <= 0 ||
p > 1 ||
values.some((v) => !Number.isFinite(v))
) {
throw new Error('provide finite observations and percentile p in (0, 1]');
}
const sorted = [...values].sort((a, b) => a - b);
return sorted[Math.ceil(p * sorted.length) - 1];
}
package aggregate
import (
"errors"
"math"
"sort"
)
type Cohort struct {
Name string
Count int
Mean float64
}
// WeightedMean combines cohort means while preserving their denominators.
func WeightedMean(cohorts []Cohort) (float64, error) {
if len(cohorts) == 0 {
return 0, errors.New("at least one cohort is required")
}
total, weighted := 0, 0.0
for _, c := range cohorts {
if c.Count < 0 || math.IsNaN(c.Mean) || math.IsInf(c.Mean, 0) {
return 0, errors.New("counts must be non-negative and means finite")
}
total += c.Count
weighted += float64(c.Count) * c.Mean
}
if total == 0 {
return 0, errors.New("total count must be positive")
}
return weighted / float64(total), nil
}
// Percentile returns the nearest-rank percentile from raw observations, with p in (0, 1].
func Percentile(values []float64, p float64) (float64, error) {
if len(values) == 0 || p <= 0 || p > 1 || math.IsNaN(p) || math.IsInf(p, 0) {
return 0, errors.New("require observations and p in (0, 1]")
}
ordered := append([]float64(nil), values...)
for _, v := range ordered {
if math.IsNaN(v) || math.IsInf(v, 0) {
return 0, errors.New("observations must be finite")
}
}
sort.Float64s(ordered)
index := int(math.Ceil(p*float64(len(ordered)))) - 1
return ordered[index], nil
}
Preserve the evidence needed by the next reader.
For a practical discussion of percentiles and distribution summaries, see the project's companion lesson on mean, median, variance, and percentiles. This lesson's examples use nearest-rank for clarity; production telemetry tools should document their quantile and histogram semantics.