← Math in Practice
Concept Math in engineering decisions

Capacity, headroom, and saturation

A percentage is useful only when demand and capacity describe the same workload.

An API team sees p95 latency rise from 290 ms to 1.2 s during a rollout. The cluster dashboard shows 68% average CPU. Meanwhile, the report endpoint has changed the work each request performs, and the waiting queue is growing. “We have 32% CPU left” does not yet answer whether this API has room for more requests.

The judgment to keep

Compare offered demand with capacity measured for the same workload and service objective. Then check whether accepted work is accumulating and which resource or dependency limits completion.

TypeScriptGo Demand · measured capacity · utilization · headroom · queue growth
01 / Read the latency report

The cluster looks comfortable while one service falls behind.

For this illustrative incident, eight API replicas receive about 820 accepted requests each second and complete about 760 requests each second over a five-minute window. The queue grows while p95 latency reaches 1.2 seconds. A prior load test sustained 100 requests/s per replica for this request mix while keeping p95 at or below 250 ms and errors below 0.1%.

The numbers are a teaching scenario, not production measurements. “Accepted” means work enters this service’s queue; completions leave the same boundary. Count retry attempts in offered demand, and track them separately so retries do not masquerade as new customer work.

Separate three observations: how much work arrives, how much leaves, and whether the agreed latency/error objective still holds.
02 / Define the quantities

Demand and capacity need the same unit and the same service boundary.

Demand is work offered to or accepted by the boundary you are studying, here requests per second entering the API queue. Be explicit about which one: attempted ingress can include requests rejected before queueing, while accepted arrivals are the right quantity for the queue balance. Retries are work too, even when they do not represent a new user action.

Capacity is the sustainable rate this configured service can handle while meeting its stated objective for a defined workload. The illustrative test says one replica sustained 100 requests/s for the current request mix with p95 at or below 250 ms and errors below 0.1%. That is a measured operating point, not a universal maximum. Change the payloads, cache hit rate, database, concurrency, or latency target and the estimate may change.

Utilization here is a planning ratio: offered demand divided by measured capacity. A resource gauge such as CPU utilization is a different ratio. Planning utilization can exceed 100% when offered work exceeds the measured rate; that means a deficit in this model, not a CPU gauge that somehow reads 102.5%.

Headroom is capacity minus demand. Positive headroom is room under the chosen measurement; zero means demand equals it; negative headroom is a deficit. Keep the units, requests/s, so the sign and magnitude remain interpretable.

Planning utilizationU = demand / capacity
Signed headroomH = capacity − demand
03 / Work the incident numbers

The offered rate is slightly above the tested service rate.

Illustrative five-minute incident and capacity estimate
QuantityCalculationResultWhat it says
Estimated capacity100 requests/s/replica × 8 replicas800 requests/sAssumes replicas scale near-linearly for this mix
Planning utilization820 requests/s ÷ 800 requests/s1.025 = 102.5%Offered accepted work is above the estimate
Signed headroom800 − 820 requests/s−20 requests/sA 20 requests/s deficit in this model
Queue change rate820 arrivals/s − 760 completions/s+60 requests/sAccepted work accumulates if rates persist

Over five minutes, the simple queue balance gives 60 requests/s × 300 s = 18,000 additional queued requests. This is a projection from constant rates, not a forecast: queues can hit limits, shed work, trigger timeouts, or cause retries. The arithmetic is a useful consistency check against the measured queue slope.

04 / Question the capacity estimate

Eight times one replica is a hypothesis about scaling, too.

Multiplying 100 requests/s/replica × 8 replicas assumes those replicas contribute independent capacity. Shared databases, locks, caches, network limits, or load imbalance can break that assumption. A benchmark that saturates one replica against an isolated fake dependency may measure a different system from production.

Capacity is also tied to a workload and a service objective. A tiny cached read and a large report-generation request each count as one request, but they do not consume equal work. When the mix changes, requests/s can hide a demand shift. Segment by route or use a workload-specific cost unit, and preserve p95, error rate, queue age, and rejected work alongside the aggregate rate.

Averages hide bursts and uneven distribution. A five-minute average below capacity can still miss a short spike or one hot replica. Sample shorter windows, inspect per-replica histograms, and compare the queue’s actual slope. A capacity number without its load shape and time window can create false confidence.

A useful capacity statement: “With configuration X and workload mix Y, we measured Z accepted requests/s while meeting latency objective A and error objective B.”
05 / Change the assumptions

Adjust demand and replicas, then check whether the story still fits.

Capacity and headroom labU = D / C   ·   H = C − D

Capacity is per replica for the workload and objective described above; linear scaling is an assumption.

Estimated aggregate capacity800 requests/s100 requests/s/replica × 8 replicas
Planning utilization102.5%820 offered ÷ 800 capacity
Signed headroom-20 requests/sNegative means offered work exceeds the estimate
Simple queue change60 requests/sAbout 18000 requests over five minutes if rates persist

The queue estimate assumes accepted arrivals and completions share a boundary and that drops/cancellations are accounted for separately. Zero capacity has no finite utilization ratio.

Try increasing demand to 900 requests/s while holding eight replicas: the planning ratio rises to 112.5%, and estimated headroom falls to −100 requests/s. Now double replicas. The lab shows a lower aggregate ratio only because it assumes each additional replica brings another 100 requests/s of capacity. In a real system, verify that assumption against shared dependencies and a representative load test.

06 / Practice in code

Keep the three calculations explicit.

The helpers return the planning utilization, signed headroom, and queue rate change. They validate finite, non-negative rates and refuse a zero capacity denominator. They cannot confirm that two callers measured the same workload, service objective, interval, or boundary; those assumptions belong in the metric contract and benchmark notes.

Compare the same calculation in TypeScript and Go.

The examples retain request-rate units in names and label the ratio as a planning measure.

TypeScriptCapacity, utilization, and headroom · units retained
capacity.ts · utilization and headroom
export type CapacityAssessment = {
	utilization: number;
	headroomPerSecond: number;
	backlogChangePerSecond: number;
	status: 'within-measured-capacity' | 'at-capacity' | 'over-capacity';
};

/** Compare offered demand and completed work with capacity for the same workload. */
export function assessCapacity(
	offeredRequestsPerSecond: number,
	completedRequestsPerSecond: number,
	sustainableRequestsPerSecond: number
): CapacityAssessment {
	for (const [name, value] of [
		['offeredRequestsPerSecond', offeredRequestsPerSecond],
		['completedRequestsPerSecond', completedRequestsPerSecond],
		['sustainableRequestsPerSecond', sustainableRequestsPerSecond]
	] as const) {
		if (!Number.isFinite(value) || value < 0) {
			throw new Error(`${name} must be finite and non-negative`);
		}
	}
	if (sustainableRequestsPerSecond === 0) {
		throw new Error('sustainableRequestsPerSecond must be positive');
	}

	const utilization = offeredRequestsPerSecond / sustainableRequestsPerSecond;
	return {
		utilization,
		headroomPerSecond: sustainableRequestsPerSecond - offeredRequestsPerSecond,
		backlogChangePerSecond: offeredRequestsPerSecond - completedRequestsPerSecond,
		status:
			utilization > 1
				? 'over-capacity'
				: utilization === 1
					? 'at-capacity'
					: 'within-measured-capacity'
	};
}
GoCapacity, utilization, and headroom · units retained
capacity.go · utilization and headroom
package mathpractice

import (
	"errors"
	"math"
)

type CapacityAssessment struct {
	Utilization            float64
	HeadroomPerSecond      float64
	BacklogChangePerSecond float64
	Status                 string
}

// AssessCapacity compares offered demand and completed work with capacity
// measured for the same workload and service objective.
func AssessCapacity(offeredRequestsPerSecond, completedRequestsPerSecond, sustainableRequestsPerSecond float64) (CapacityAssessment, error) {
	values := []struct {
		name  string
		value float64
	}{
		{"offeredRequestsPerSecond", offeredRequestsPerSecond},
		{"completedRequestsPerSecond", completedRequestsPerSecond},
		{"sustainableRequestsPerSecond", sustainableRequestsPerSecond},
	}
	for _, input := range values {
		if math.IsNaN(input.value) || math.IsInf(input.value, 0) || input.value < 0 {
			return CapacityAssessment{}, errors.New(input.name + " must be finite and non-negative")
		}
	}
	if sustainableRequestsPerSecond == 0 {
		return CapacityAssessment{}, errors.New("sustainableRequestsPerSecond must be positive")
	}

	utilization := offeredRequestsPerSecond / sustainableRequestsPerSecond
	status := "within-measured-capacity"
	if utilization > 1 {
		status = "over-capacity"
	} else if utilization == 1 {
		status = "at-capacity"
	}
	return CapacityAssessment{
		Utilization:            utilization,
		HeadroomPerSecond:      sustainableRequestsPerSecond - offeredRequestsPerSecond,
		BacklogChangePerSecond: offeredRequestsPerSecond - completedRequestsPerSecond,
		Status:                 status,
	}, nil
}
07 / Choose the next measurement

A ratio can locate a mismatch; evidence locates the constraint.

Keep the model concise: U = D/C compares offered demand to measured capacity, H = C−D is the signed room left, and queue change ≈ arrivals−completions estimates accumulation over a short window. Each equation answers a different question, and none identifies the bottleneck without measurements of the system.

For queue accounting, see John D. C. Little, “A Proof for the Queuing Formula: L = λW”, Operations Research 9(3), 1961. Little’s Law relates long-run averages in a stable system; the queue-slope calculation here is a separate, short-window balance.