“5.5% of requests were slow” is incomplete without a population.
A team compares traces from an interactive API and a batch export worker. For this lesson, “slow” means a request took at least 500 ms. The target is one illustrative 100,000-request production window: 90,000 interactive requests and 10,000 batch requests. These figures are constructed for teaching; they are not measurements from a real service.
The benchmark deliberately retains 1,000 requests from each workload so engineers can
compare them with equal detail. Among those 2,000 retained requests, 10 interactive and
100 batch requests meet the slow threshold. Pooling gives (10 + 100) ÷ (1,000 + 1,000) = 110 ÷ 2,000 = 5.5%.
That arithmetic is correct for the retained sample. The question is whether the sample composition matches the population named in the claim. Here production is 90% interactive and 10% batch, while the sample is 50% of each. The batch workload has been deliberately overrepresented.
- Target population
- 100,000 requests: 90,000 interactive, 10,000 batch.
- Observed sample
- 2,000 retained traces: 1,000 from each cohort.
- Measured event
- Request duration is at least 500 ms.
- Observed slow
- 10 interactive + 100 batch = 110 / 2,000 = 5.5%.
Selection bias starts with the path from population to records.
The target population is the set of requests the conclusion is meant to describe. The observed sample is the subset the benchmark or telemetry pipeline retained. A selection mechanism is the rule that determines which population members can appear and how likely each is to be retained.
Here the benchmark quota takes the same number of records from each workload. That is a reasonable design for comparing interactive and batch behavior. It is not a miniature copy of the production mix, so its unweighted pooled rate estimates the benchmark mix, not the production mix. This is a design issue, not evidence that the benchmark was careless.
Now ask how the 1,000 records within each cohort were retained. Were requests randomly sampled, were the first 1,000 kept, were only completed traces retained, or did a tail-sampling rule favor slow requests? Cohort balancing corrects the between-cohort mix only if the within-cohort rate estimates are useful for their groups.
| Part | In this example | Diagnostic question |
|---|---|---|
| Target | All 100,000 requests in the stated window | Which users, workload classes, regions, and time window does the claim cover? |
| Observed | 2,000 retained trace records | Which requests are missing, and are missingness and retention recorded? |
| Selection | Fixed quota of 1,000 per workload | What were the inclusion rules inside each workload? |
Keep the cohort rates, then combine them with the population shares.
First calculate each cohort's slow rate from its sample. Interactive: 10 ÷ 1,000 = 0.01 = 1%. Batch: 100 ÷ 1,000 = 0.10 = 10%. These are still sample-based estimates,
and their quality depends on how traces were retained within each cohort.
Next compute each workload's share of the target population. Interactive: 90,000 ÷ 100,000 = 0.90. Batch: 10,000 ÷ 100,000 = 0.10. A post-stratified estimate multiplies each cohort rate by its target share and
adds the pieces:
(0.90 × 1%) + (0.10 × 10%) = 0.9% + 1.0% = 1.9%
The interactive workload contributes an estimated 900 slow requests in 90,000; batch contributes 1,000 in 10,000. That gives 1,900 / 100,000 = 1.9%. The higher batch rate still matters: its 10% slow rate contributes just over half the estimated slow requests despite representing only a tenth of traffic.
| Workload | Population share | Sample slow rate | Weighted contribution |
|---|---|---|---|
| Interactive | 90,000 / 100,000 = 90% | 10 / 1,000 = 1% | 90% × 1% = 0.9 percentage points |
| Batch | 10,000 / 100,000 = 10% | 100 / 1,000 = 10% | 10% × 10% = 1.0 percentage point |
| Combined | 100% | Unweighted sample: 5.5% | Post-stratified estimate: 1.9% |
A representative sample gives the same answer here, in expectation.
Imagine a second sample selected to mirror the target mix: 9,000 interactive requests and
1,000 batch requests. If the within-cohort rates match the first sample, it would contain
an estimated 90 slow interactive requests (9,000 × 1%) and 100 slow batch
requests (1,000 × 10%). The combined rate is (90 + 100) ÷ 10,000 = 1.9%.
A production-shaped random sample estimates the mix directly; stratified weighting estimates it from group rates and known shares. Neither produces exactly 1.9% every time. Actual random samples vary, and the cohort rates themselves carry sampling uncertainty. The equality here is an illustrative expected result based on the same cohort-rate inputs.
Weighting fixes a known mix mismatch; it cannot repair every selection effect.
The calculation assumes that the sampled requests within interactive and batch are informative about all requests in those cohorts. If the tracing system keeps only slow batch traces, then 100 / 1,000 is not a valid batch rate; multiplying it by the correct 10% share does not fix that within-cohort selection.
Cohorts may also be too broad. Batch requests could include exports of very different sizes; a region or client version could have a distinct rate. A shift in workload mix after a release can make yesterday's target shares stale. Missing traces, failed requests omitted from the collector, or changed event definitions can alter the measured population too.
If every eligible request has a known inclusion probability, sampling methods can sometimes account for that unequal selection using inverse-probability weights. The simple cohort formula here is not a universal correction: it only uses known cohort shares and assumes representative within-group rates. Tail-based selection requires the selection probabilities or a separate unbiased measurement path; unknown inclusion probabilities leave the population rate unidentified from these counts alone.
Make the raw and weighted answers hard to confuse.
Both examples calculate the raw pooled sample rate and the population-weighted estimate
from the same cohort counts. The weighted rate is Σ(population share × cohort sample rate). They reject empty populations, empty samples, non-integer or negative counts, and slow
counts larger than their sample. Inputs are rates and shares from an illustrative design;
the function cannot determine if the samples are unbiased or if the population shares are
current.
The same counts produce the 5.5% pooled sample rate and the 1.9% estimate under the target mix.
type Cohort = {
name: string;
populationRequests: number;
sampledRequests: number;
slowRequests: number;
};
type Estimate = {
naiveSampleRate: number;
weightedRate: number;
cohorts: Array<{ name: string; sampleRate: number; populationShare: number }>;
};
function estimateSlowRate(cohorts: Cohort[]): Estimate {
if (cohorts.length === 0) throw new Error('at least one cohort is required');
const populationTotal = cohorts.reduce((sum, cohort) => sum + cohort.populationRequests, 0);
const sampleTotal = cohorts.reduce((sum, cohort) => sum + cohort.sampledRequests, 0);
const slowTotal = cohorts.reduce((sum, cohort) => sum + cohort.slowRequests, 0);
if (populationTotal <= 0 || sampleTotal <= 0)
throw new Error('population and sample must be non-empty');
let weightedRate = 0;
const details = cohorts.map((cohort) => {
if (
!Number.isInteger(cohort.populationRequests) ||
!Number.isInteger(cohort.sampledRequests) ||
!Number.isInteger(cohort.slowRequests) ||
cohort.populationRequests < 0 ||
cohort.sampledRequests <= 0 ||
cohort.slowRequests < 0 ||
cohort.slowRequests > cohort.sampledRequests
) {
throw new Error(`invalid counts for ${cohort.name}`);
}
const sampleRate = cohort.slowRequests / cohort.sampledRequests;
const populationShare = cohort.populationRequests / populationTotal;
weightedRate += populationShare * sampleRate;
return { name: cohort.name, sampleRate, populationShare };
});
return { naiveSampleRate: slowTotal / sampleTotal, weightedRate, cohorts: details };
}
const example = estimateSlowRate([
{ name: 'Interactive', populationRequests: 90_000, sampledRequests: 1_000, slowRequests: 10 },
{ name: 'Batch', populationRequests: 10_000, sampledRequests: 1_000, slowRequests: 100 }
]);
console.log(`Unweighted sample: ${(example.naiveSampleRate * 100).toFixed(1)}%`);
console.log(`Population-weighted estimate: ${(example.weightedRate * 100).toFixed(1)}%`);
package main
import (
"fmt"
)
type Cohort struct {
Name string
PopulationRequests int64
SampledRequests int64
SlowRequests int64
}
type Estimate struct {
NaiveSampleRate float64
WeightedRate float64
}
func estimateSlowRate(cohorts []Cohort) (Estimate, error) {
if len(cohorts) == 0 {
return Estimate{}, fmt.Errorf("at least one cohort is required")
}
var populationTotal, sampleTotal, slowTotal int64
var weightedRate float64
for _, cohort := range cohorts {
if cohort.PopulationRequests < 0 || cohort.SampledRequests <= 0 ||
cohort.SlowRequests < 0 || cohort.SlowRequests > cohort.SampledRequests {
return Estimate{}, fmt.Errorf("invalid counts for %s", cohort.Name)
}
populationTotal += cohort.PopulationRequests
sampleTotal += cohort.SampledRequests
slowTotal += cohort.SlowRequests
}
if populationTotal <= 0 || sampleTotal <= 0 {
return Estimate{}, fmt.Errorf("population and sample must be non-empty")
}
for _, cohort := range cohorts {
sampleRate := float64(cohort.SlowRequests) / float64(cohort.SampledRequests)
populationShare := float64(cohort.PopulationRequests) / float64(populationTotal)
weightedRate += populationShare * sampleRate
}
return Estimate{NaiveSampleRate: float64(slowTotal) / float64(sampleTotal), WeightedRate: weightedRate}, nil
}
func main() {
estimate, err := estimateSlowRate([]Cohort{
{Name: "Interactive", PopulationRequests: 90_000, SampledRequests: 1_000, SlowRequests: 10},
{Name: "Batch", PopulationRequests: 10_000, SampledRequests: 1_000, SlowRequests: 100},
})
if err != nil {
panic(err)
}
fmt.Printf("Unweighted sample: %.1f%%\n", estimate.NaiveSampleRate*100)
fmt.Printf("Population-weighted estimate: %.1f%%\n", estimate.WeightedRate*100)
}
Choose a sample that can answer the question you actually have.
If the question is “which workload is slow?”, keep the balanced sample and compare cohort distributions. If it is “what fraction of all production requests crossed 500 ms?”, use a production-shaped sample or estimate each cohort and weight by current production shares. If it is “did the release cause this?”, neither percentage answers causation: compare before and after under matching workload mix, and investigate other changes that moved with the release.
A useful report says what was observed and how: “In this illustrative benchmark, 5.5% of the balanced retained sample was at least 500 ms. Using the 90/10 target workload mix and the observed within-cohort rates gives a 1.9% estimate. This assumes the retained requests represent each cohort; we have not established that from the pooled count.” In real work, replace every illustrative quantity with a queryable dataset and document missingness.
For the distinction between a population and a sample, see NIST's discussion of data and sampling. For why representative or random selection matters when drawing conclusions, see NIST's sampling discussion.