“There are 240 requests” needs a service boundary and a time window.
For this exercise, suppose a service dashboard reports an average of 240 requests in the system, and the service completes 80 requests per second over the same steady five-minute observation window. The count includes requests waiting and requests being processed. These are illustrative measurements, not a production benchmark.
A queue length is an amount of work at a moment or averaged over time. Throughput is completed work per unit time. Latency is elapsed time for one request; when averaged across the same request population and interval, it can be related to the first two quantities.
Little’s Law relates three long-run averages.
Little’s Law is usually written L = λW: L is
the average number of items in a system, λ is the effective arrival rate, and W is the average time each item spends inside that same system. In a stable system, effective
arrivals and completions balance over the long run, so a matching completion-throughput measurement
can stand in for λ.
Rearrange it to answer the alert’s question: W = L / λ. The units check the
result: requests ÷ (requests/second) = seconds. The request boundary must
match. If L covers the application queue but λ counts database completions,
you are dividing unlike quantities.
It is an accounting relationship for a stable system over a representative period, not a causal explanation. The equation does not tell you whether time was spent waiting for a worker, a lock, a network call, or a database.
Keep the units through the division.
| Quantity | Value | Unit | Meaning |
|---|---|---|---|
| Average population, L | 240 | requests | Waiting plus active requests in this service boundary |
| Throughput, λ | 80 | requests/second | Completed requests over the matching period |
| Average time, W = L/λ | 240 / 80 = 3 | seconds | Average time in the system, under the stated assumptions |
A useful independent check: at 80 completions each second, three seconds of average time corresponds to 240 requests in the system. The arithmetic agrees in both directions. Rounding the displayed averages can make the equality approximate rather than exact.
When arrivals outrun completions, the stock accumulates.
Let λin mean the rate of arriving requests and μout the rate of
completed requests. In a simplified fluid model, backlog changes at approximately λin − μout requests per second while those rates persist. If arrivals are 105 requests/s
and completions are 100 requests/s, the backlog grows by about 5 requests each second.
This difference is a short-window diagnostic, not a stable Little’s Law estimate. Real services have bursts, variable service times, limits, cancellations, and feedback such as retries. If arrivals later fall below completions, the queue may drain. If they remain above completions, waiting work accumulates and some requests may time out before completion.
Change one average and see the implied time move.
Use matching averages for one stable system boundary. Values are illustrative.
A zero throughput has no finite implied time. The backlog calculation is a simplified rate difference, not a queueing simulation or forecast.
Try doubling the average population while holding throughput fixed: the implied average time doubles. Then double both population and throughput: the implied time returns to its starting value. In the separate imbalance example, equal arrival and completion rates imply zero net change; if completions exceed arrivals, a negative result represents a draining queue in this simplified model.
Make the boundary and invalid rates visible.
The helper computes W = L/λ and rejects a non-positive throughput. It cannot verify
that a caller used the same service boundary or a stable observation window; those belong in
the surrounding metric contract.
Both examples preserve units in names and validate their numeric inputs.
export function averageTimeInSystemSeconds(
averageRequestsInSystem: number,
throughputPerSecond: number
): number {
if (!Number.isFinite(averageRequestsInSystem) || averageRequestsInSystem < 0) {
throw new Error('averageRequestsInSystem must be non-negative');
}
if (!Number.isFinite(throughputPerSecond) || throughputPerSecond <= 0) {
throw new Error('throughputPerSecond must be positive');
}
return averageRequestsInSystem / throughputPerSecond;
}
export function projectedBacklogChangePerSecond(
arrivalRatePerSecond: number,
completionRatePerSecond: number
): number {
if (!Number.isFinite(arrivalRatePerSecond) || arrivalRatePerSecond < 0) {
throw new Error('arrivalRatePerSecond must be non-negative');
}
if (!Number.isFinite(completionRatePerSecond) || completionRatePerSecond < 0) {
throw new Error('completionRatePerSecond must be non-negative');
}
return arrivalRatePerSecond - completionRatePerSecond;
}
package mathpractice
import (
"errors"
"math"
)
func AverageTimeInSystemSeconds(averageRequestsInSystem, throughputPerSecond float64) (float64, error) {
if math.IsNaN(averageRequestsInSystem) || math.IsInf(averageRequestsInSystem, 0) || averageRequestsInSystem < 0 {
return 0, errors.New("averageRequestsInSystem must be non-negative")
}
if math.IsNaN(throughputPerSecond) || math.IsInf(throughputPerSecond, 0) || throughputPerSecond <= 0 {
return 0, errors.New("throughputPerSecond must be positive")
}
return averageRequestsInSystem / throughputPerSecond, nil
}
func ProjectedBacklogChangePerSecond(arrivalRatePerSecond, completionRatePerSecond float64) (float64, error) {
if math.IsNaN(arrivalRatePerSecond) || math.IsInf(arrivalRatePerSecond, 0) || arrivalRatePerSecond < 0 {
return 0, errors.New("arrivalRatePerSecond must be non-negative")
}
if math.IsNaN(completionRatePerSecond) || math.IsInf(completionRatePerSecond, 0) || completionRatePerSecond < 0 {
return 0, errors.New("completionRatePerSecond must be non-negative")
}
return arrivalRatePerSecond - completionRatePerSecond, nil
}
The law relates averages; diagnosis needs evidence about the system.
Keep the relationship and its limit together: L = λW connects average work in the
system, throughput, and average time for a stable system. It does not turn a growing queue into
a stable one or replace tail-latency evidence.
Further reading: John D. C. Little, “A Proof for the Queuing Formula: L = λW”, Operations Research 9(3), 1961; Leonard Kleinrock, Queueing Systems, Volume 1: Theory.