A service path fails when any required part is unavailable.
In this illustrative release review, the team has monthly time-based availability estimates of 99.95% for the checkout API, 99.90% for identity, and 99.95% for the database. The path needs all three to serve a checkout. The numbers are authored examples, not production observations.
Availability describes the fraction of a defined interval when a specified service boundary met its availability rule. The rule might be “health checks pass” or “eligible requests meet the service objective”; those are not interchangeable. Here the illustrative figures are monthly time fractions over one 30-day calendar month, or 43,200 minutes. They are not probabilities that a particular request succeeds.
- API
- 99.95% time availability during the stated month.
- Identity
- 99.90% time availability during the same interval.
- Database
- 99.95% time availability during the same interval.
- Question
- What can multiplication estimate, and what assumption would make it wrong?
For a series path, every required component must be up.
If a request requires components A, B, and C, the path is available only when all three
are available at the same time. If their states are independent at the time scale being
measured, then Apath = AA × AB × AC. This is a series system: failure of any required component breaks the path.
Substitute fractions, not whole-number percentages: 0.9995 × 0.9990 × 0.9995 = 0.99800124975, or 99.800124975%. For a 30-day interval, expected unavailable time under the
stationary model is (1 − 0.99800124975) × 43,200 min ≈ 86.35 min. This is a
model-based conversion of a time fraction; it does not say when the downtime occurs.
Multiplying measured marginal availabilities does not generally give the actual joint availability. It gives an estimate when independence is credible. If component states are correlated, use aligned incident timelines to measure when the whole path was unavailable, or build a model that represents the shared causes.
Either replica can serve the work, if either can fail on its own.
For two equivalent replicas, the redundant group is down only when both replicas are down.
If their failures are independent and each has availability A, then Aparallel = 1 − (1 − A)². With two 99.9% replicas, that is 1 − (0.001 × 0.001) = 0.999999, or 99.9999% for the replica group.
For n independent interchangeable replicas, estimate 1 − (1 − A)ⁿ. “Interchangeable” matters: replicas must have compatible data,
capacity, routing, and permissions, and the load balancer must direct work to a healthy
one. The equation does not include detection delay, failover errors, overload of the
survivor, or recovery time unless those effects are represented in the input measurements.
Two copies do little for an outage that reaches both copies.
Suppose the two API replicas each measure 99.9% availability, but both sit behind one
99.9%-available regional network dependency. The independence-only replica calculation
says 99.9999%. When the network dependency is down, though, neither replica is reachable.
If its state is independent of replica failures, include it in series: 0.999 × 0.999999 = 0.998999001, or about 99.8999% for this simplified path.
The shared dependency now dominates; adding more replicas in the same failure domain cannot raise path availability above that dependency’s 99.9% availability. If the network failures also correlate with replica failures, even that product is not justified. A measured joint availability or a dependency model with common causes is needed.
Look for shared regions and zones, power and network paths, DNS, identity and secrets, storage, control planes, deploy pipelines, and human actions. Redundancy is a topology and operating property as well as a replica count.
Put the shared dependency back into the estimate.
Change the values to compare a three-service series path with a redundant replica group. The window is held at 30 days (43,200 minutes). The calculator treats the entered percentages as time availability estimates; it cannot decide whether the definitions, evidence, or independence assumptions are valid.
Estimated unavailable time: 86.35 minutes per 30-day window.
Assumes independent replica states and equivalent capacity.
Formula: shared availability × independent replica-group availability.
These are time-fraction estimates, not per-request success probabilities. For simplicity, the shared dependency is assumed independent of replica failures; real common causes can violate that assumption. Confirm availability rules and failure domains against incident timelines and end-to-end SLI data.
Make the assumptions visible at the function boundary.
The helpers take availability percentages and an explicit window in minutes. The series helper multiplies component fractions; the redundant helper combines interchangeable replica availability with an independent shared dependency. Input validation catches values outside 0–100% and invalid replica counts, but code cannot validate whether the operational assumptions are true.
Both examples preserve the time unit and document the independence assumptions.
export type AvailabilityEstimate = {
availabilityPercent: number;
downtimeMinutes: number;
};
function validateAvailability(value: number, label: string): void {
if (!Number.isFinite(value) || value < 0 || value > 100) {
throw new Error(`${label} must be between 0 and 100 percent`);
}
}
/** Estimates a series path where every component must be available.
* This product model assumes component states are independent in the stated window.
*/
export function estimateSeriesAvailability(
componentAvailabilityPercent: number[],
windowMinutes: number
): AvailabilityEstimate {
if (componentAvailabilityPercent.length === 0) {
throw new Error('at least one component is required');
}
if (!Number.isFinite(windowMinutes) || windowMinutes < 0) {
throw new Error('windowMinutes must be finite and non-negative');
}
const availability = componentAvailabilityPercent.reduce((product, component, index) => {
validateAvailability(component, `component ${index + 1}`);
return product * (component / 100);
}, 1);
return {
availabilityPercent: availability * 100,
downtimeMinutes: (1 - availability) * windowMinutes
};
}
/** Estimates interchangeable replicas plus one independent shared dependency.
* The replica failures must be independent conditional on the shared dependency being up.
*/
export function estimateRedundantAvailability(
replicaAvailabilityPercent: number,
replicaCount: number,
sharedDependencyAvailabilityPercent: number
): number {
validateAvailability(replicaAvailabilityPercent, 'replica availability');
validateAvailability(sharedDependencyAvailabilityPercent, 'shared dependency availability');
if (!Number.isSafeInteger(replicaCount) || replicaCount < 1) {
throw new Error('replicaCount must be a positive safe integer');
}
const replicaAvailability = replicaAvailabilityPercent / 100;
const sharedAvailability = sharedDependencyAvailabilityPercent / 100;
const allReplicasUnavailable = (1 - replicaAvailability) ** replicaCount;
return sharedAvailability * (1 - allReplicasUnavailable) * 100;
}
package availability
import (
"errors"
"fmt"
"math"
)
type Estimate struct {
AvailabilityPercent float64
DowntimeMinutes float64
}
func validateAvailability(value float64, label string) error {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 100 {
return fmt.Errorf("%s must be between 0 and 100 percent", label)
}
return nil
}
// Series estimates a path where every component must be available.
// The product model assumes independent component states in the stated window.
func Series(componentsPercent []float64, windowMinutes float64) (Estimate, error) {
if len(componentsPercent) == 0 {
return Estimate{}, errors.New("at least one component is required")
}
if math.IsNaN(windowMinutes) || math.IsInf(windowMinutes, 0) || windowMinutes < 0 {
return Estimate{}, errors.New("windowMinutes must be finite and non-negative")
}
availability := 1.0
for index, component := range componentsPercent {
if err := validateAvailability(component, fmt.Sprintf("component %d", index+1)); err != nil {
return Estimate{}, err
}
availability *= component / 100
}
return Estimate{
AvailabilityPercent: availability * 100,
DowntimeMinutes: (1 - availability) * windowMinutes,
}, nil
}
// Redundant estimates interchangeable replicas plus one independent shared dependency.
// Replica failures must be independent conditional on the shared dependency being up.
func Redundant(replicaPercent float64, replicaCount int, sharedDependencyPercent float64) (float64, error) {
if err := validateAvailability(replicaPercent, "replica availability"); err != nil {
return 0, err
}
if err := validateAvailability(sharedDependencyPercent, "shared dependency availability"); err != nil {
return 0, err
}
if replicaCount < 1 {
return 0, errors.New("replicaCount must be at least 1")
}
replicaAvailability := replicaPercent / 100
sharedAvailability := sharedDependencyPercent / 100
allReplicasUnavailable := math.Pow(1-replicaAvailability, float64(replicaCount))
return sharedAvailability * (1 - allReplicasUnavailable) * 100, nil
}
Use the estimate to ask better questions of the real system.
Availability and error budget math is a way to reason about tolerated service risk, not a substitute for observing user outcomes. See the Google SRE Book chapters on embracing risk and the availability table for the relationship between SLOs and time budgets.
For the architectural limit on simple redundancy, see the Google SRE Book’s discussion of production environments and failure domains. The lesson’s formulas are simplified teaching models; a production claim should be based on a defined SLI and evidence from the actual system.