← Math in Practice
Concept Samples, populations, and evidence

Sampling and selection bias

A measurement can be accurate about its sample and still mislead about the system.

A benchmark report says that 5.5% of requests were slow. Before changing a timeout or declaring a regression, ask which requests could enter the report. A carefully balanced sample can reveal how two workloads differ while giving the wrong combined rate for production traffic.

The judgment to keep

Name the target population, the observed sample, and how records were selected. If the sample mix differs from the population, calculate cohort rates and weight them by the target mix—then check whether the sample is representative inside each cohort.

TypeScriptGo Target population · selection mechanism · cohort weighting
01 / Read the claim

“5.5% of requests were slow” is incomplete without a population.

A team compares traces from an interactive API and a batch export worker. For this lesson, “slow” means a request took at least 500 ms. The target is one illustrative 100,000-request production window: 90,000 interactive requests and 10,000 batch requests. These figures are constructed for teaching; they are not measurements from a real service.

The benchmark deliberately retains 1,000 requests from each workload so engineers can compare them with equal detail. Among those 2,000 retained requests, 10 interactive and 100 batch requests meet the slow threshold. Pooling gives (10 + 100) ÷ (1,000 + 1,000) = 110 ÷ 2,000 = 5.5%.

That arithmetic is correct for the retained sample. The question is whether the sample composition matches the population named in the claim. Here production is 90% interactive and 10% batch, while the sample is 50% of each. The batch workload has been deliberately overrepresented.

Case file / API and batch traces The report has two different mixes.
Target population
100,000 requests: 90,000 interactive, 10,000 batch.
Observed sample
2,000 retained traces: 1,000 from each cohort.
Measured event
Request duration is at least 500 ms.
Observed slow
10 interactive + 100 batch = 110 / 2,000 = 5.5%.
Diagnostic pause: the sample tells us the balanced benchmark's event rate. What facts would you need before saying it estimates production?
02 / Inspect the selection

Selection bias starts with the path from population to records.

The target population is the set of requests the conclusion is meant to describe. The observed sample is the subset the benchmark or telemetry pipeline retained. A selection mechanism is the rule that determines which population members can appear and how likely each is to be retained.

Here the benchmark quota takes the same number of records from each workload. That is a reasonable design for comparing interactive and batch behavior. It is not a miniature copy of the production mix, so its unweighted pooled rate estimates the benchmark mix, not the production mix. This is a design issue, not evidence that the benchmark was careless.

Now ask how the 1,000 records within each cohort were retained. Were requests randomly sampled, were the first 1,000 kept, were only completed traces retained, or did a tail-sampling rule favor slow requests? Cohort balancing corrects the between-cohort mix only if the within-cohort rate estimates are useful for their groups.

Questions to ask before moving from a sample to a population
PartIn this exampleDiagnostic question
TargetAll 100,000 requests in the stated windowWhich users, workload classes, regions, and time window does the claim cover?
Observed2,000 retained trace recordsWhich requests are missing, and are missingness and retention recorded?
SelectionFixed quota of 1,000 per workloadWhat were the inclusion rules inside each workload?
03 / Rebuild the estimate

Keep the cohort rates, then combine them with the population shares.

First calculate each cohort's slow rate from its sample. Interactive: 10 ÷ 1,000 = 0.01 = 1%. Batch: 100 ÷ 1,000 = 0.10 = 10%. These are still sample-based estimates, and their quality depends on how traces were retained within each cohort.

Next compute each workload's share of the target population. Interactive: 90,000 ÷ 100,000 = 0.90. Batch: 10,000 ÷ 100,000 = 0.10. A post-stratified estimate multiplies each cohort rate by its target share and adds the pieces:

(0.90 × 1%) + (0.10 × 10%) = 0.9% + 1.0% = 1.9%

The interactive workload contributes an estimated 900 slow requests in 90,000; batch contributes 1,000 in 10,000. That gives 1,900 / 100,000 = 1.9%. The higher batch rate still matters: its 10% slow rate contributes just over half the estimated slow requests despite representing only a tenth of traffic.

Illustrative estimates · duration ≥ 500 ms
WorkloadPopulation shareSample slow rateWeighted contribution
Interactive90,000 / 100,000 = 90%10 / 1,000 = 1%90% × 1% = 0.9 percentage points
Batch10,000 / 100,000 = 10%100 / 1,000 = 10%10% × 10% = 1.0 percentage point
Combined100%Unweighted sample: 5.5%Post-stratified estimate: 1.9%
04 / Match the target mix

A representative sample gives the same answer here, in expectation.

Imagine a second sample selected to mirror the target mix: 9,000 interactive requests and 1,000 batch requests. If the within-cohort rates match the first sample, it would contain an estimated 90 slow interactive requests (9,000 × 1%) and 100 slow batch requests (1,000 × 10%). The combined rate is (90 + 100) ÷ 10,000 = 1.9%.

A production-shaped random sample estimates the mix directly; stratified weighting estimates it from group rates and known shares. Neither produces exactly 1.9% every time. Actual random samples vary, and the cohort rates themselves carry sampling uncertainty. The equality here is an illustrative expected result based on the same cohort-rate inputs.

05 / Test the assumptions

Weighting fixes a known mix mismatch; it cannot repair every selection effect.

The calculation assumes that the sampled requests within interactive and batch are informative about all requests in those cohorts. If the tracing system keeps only slow batch traces, then 100 / 1,000 is not a valid batch rate; multiplying it by the correct 10% share does not fix that within-cohort selection.

Cohorts may also be too broad. Batch requests could include exports of very different sizes; a region or client version could have a distinct rate. A shift in workload mix after a release can make yesterday's target shares stale. Missing traces, failed requests omitted from the collector, or changed event definitions can alter the measured population too.

If every eligible request has a known inclusion probability, sampling methods can sometimes account for that unequal selection using inverse-probability weights. The simple cohort formula here is not a universal correction: it only uses known cohort shares and assumes representative within-group rates. Tail-based selection requires the selection probabilities or a separate unbiased measurement path; unknown inclusion probabilities leave the population rate unidentified from these counts alone.

Question the estimate: are requests missing or selected differently within either workload? Could a region, payload size, or failure state change both the chance of being sampled and the chance of being slow?
06 / Practice in code

Make the raw and weighted answers hard to confuse.

Both examples calculate the raw pooled sample rate and the population-weighted estimate from the same cohort counts. The weighted rate is Σ(population share × cohort sample rate). They reject empty populations, empty samples, non-integer or negative counts, and slow counts larger than their sample. Inputs are rates and shares from an illustrative design; the function cannot determine if the samples are unbiased or if the population shares are current.

Compare selection-aware estimation in TypeScript and Go.

The same counts produce the 5.5% pooled sample rate and the 1.9% estimate under the target mix.

TypeScriptPooled and post-stratified slow-request estimates
estimate.ts
type Cohort = {
	name: string;
	populationRequests: number;
	sampledRequests: number;
	slowRequests: number;
};

type Estimate = {
	naiveSampleRate: number;
	weightedRate: number;
	cohorts: Array<{ name: string; sampleRate: number; populationShare: number }>;
};

function estimateSlowRate(cohorts: Cohort[]): Estimate {
	if (cohorts.length === 0) throw new Error('at least one cohort is required');
	const populationTotal = cohorts.reduce((sum, cohort) => sum + cohort.populationRequests, 0);
	const sampleTotal = cohorts.reduce((sum, cohort) => sum + cohort.sampledRequests, 0);
	const slowTotal = cohorts.reduce((sum, cohort) => sum + cohort.slowRequests, 0);
	if (populationTotal <= 0 || sampleTotal <= 0)
		throw new Error('population and sample must be non-empty');

	let weightedRate = 0;
	const details = cohorts.map((cohort) => {
		if (
			!Number.isInteger(cohort.populationRequests) ||
			!Number.isInteger(cohort.sampledRequests) ||
			!Number.isInteger(cohort.slowRequests) ||
			cohort.populationRequests < 0 ||
			cohort.sampledRequests <= 0 ||
			cohort.slowRequests < 0 ||
			cohort.slowRequests > cohort.sampledRequests
		) {
			throw new Error(`invalid counts for ${cohort.name}`);
		}
		const sampleRate = cohort.slowRequests / cohort.sampledRequests;
		const populationShare = cohort.populationRequests / populationTotal;
		weightedRate += populationShare * sampleRate;
		return { name: cohort.name, sampleRate, populationShare };
	});

	return { naiveSampleRate: slowTotal / sampleTotal, weightedRate, cohorts: details };
}

const example = estimateSlowRate([
	{ name: 'Interactive', populationRequests: 90_000, sampledRequests: 1_000, slowRequests: 10 },
	{ name: 'Batch', populationRequests: 10_000, sampledRequests: 1_000, slowRequests: 100 }
]);

console.log(`Unweighted sample: ${(example.naiveSampleRate * 100).toFixed(1)}%`);
console.log(`Population-weighted estimate: ${(example.weightedRate * 100).toFixed(1)}%`);
GoPooled and post-stratified slow-request estimates
estimate.go
package main

import (
	"fmt"
)

type Cohort struct {
	Name               string
	PopulationRequests int64
	SampledRequests    int64
	SlowRequests       int64
}

type Estimate struct {
	NaiveSampleRate float64
	WeightedRate    float64
}

func estimateSlowRate(cohorts []Cohort) (Estimate, error) {
	if len(cohorts) == 0 {
		return Estimate{}, fmt.Errorf("at least one cohort is required")
	}
	var populationTotal, sampleTotal, slowTotal int64
	var weightedRate float64
	for _, cohort := range cohorts {
		if cohort.PopulationRequests < 0 || cohort.SampledRequests <= 0 ||
			cohort.SlowRequests < 0 || cohort.SlowRequests > cohort.SampledRequests {
			return Estimate{}, fmt.Errorf("invalid counts for %s", cohort.Name)
		}
		populationTotal += cohort.PopulationRequests
		sampleTotal += cohort.SampledRequests
		slowTotal += cohort.SlowRequests
	}
	if populationTotal <= 0 || sampleTotal <= 0 {
		return Estimate{}, fmt.Errorf("population and sample must be non-empty")
	}
	for _, cohort := range cohorts {
		sampleRate := float64(cohort.SlowRequests) / float64(cohort.SampledRequests)
		populationShare := float64(cohort.PopulationRequests) / float64(populationTotal)
		weightedRate += populationShare * sampleRate
	}
	return Estimate{NaiveSampleRate: float64(slowTotal) / float64(sampleTotal), WeightedRate: weightedRate}, nil
}

func main() {
	estimate, err := estimateSlowRate([]Cohort{
		{Name: "Interactive", PopulationRequests: 90_000, SampledRequests: 1_000, SlowRequests: 10},
		{Name: "Batch", PopulationRequests: 10_000, SampledRequests: 1_000, SlowRequests: 100},
	})
	if err != nil {
		panic(err)
	}
	fmt.Printf("Unweighted sample: %.1f%%\n", estimate.NaiveSampleRate*100)
	fmt.Printf("Population-weighted estimate: %.1f%%\n", estimate.WeightedRate*100)
}
07 / Make the next measurement

Choose a sample that can answer the question you actually have.

If the question is “which workload is slow?”, keep the balanced sample and compare cohort distributions. If it is “what fraction of all production requests crossed 500 ms?”, use a production-shaped sample or estimate each cohort and weight by current production shares. If it is “did the release cause this?”, neither percentage answers causation: compare before and after under matching workload mix, and investigate other changes that moved with the release.

A useful report says what was observed and how: “In this illustrative benchmark, 5.5% of the balanced retained sample was at least 500 ms. Using the 90/10 target workload mix and the observed within-cohort rates gives a 1.9% estimate. This assumes the retained requests represent each cohort; we have not established that from the pooled count.” In real work, replace every illustrative quantity with a queryable dataset and document missingness.

For the distinction between a population and a sample, see NIST's discussion of data and sampling. For why representative or random selection matters when drawing conclusions, see NIST's sampling discussion.

Rule to keep: state the target, sample, and selection mechanism before reporting a rate. Weighting can restore a known group mix; it cannot recover information the sample never captured.