The page is evidence. The incident is still a question.
The checkout monitor evaluates one service-region window at a time. An event is present when more than 5% of checkout requests in a five-minute window return a user-visible error. The monitor emits a positive alert when its score crosses a fixed threshold. The on-call engineer sees one positive alert at 14:07.
A positive alert could indicate the defined degradation, or it could be a false positive. A missed alert is also possible. We need the event’s base rate, the detector’s performance when the event is present, and its performance when the event is absent. For this worked example, those rates come from an illustrative set of 10,000 labeled, comparable windows—not from a live service.
- Event
- >5% user-visible checkout errors in one service-region five-minute window.
- Population
- 10,000 comparable labeled service-region five-minute windows.
- Base rate
- 1% of windows contain the defined degradation.
- Question
- Given a positive alert, how likely is the event in that window?
Three rates describe three different slices of the evidence.
Let E mean that the defined degradation is present in a window, and let A mean the detector alerts for that window. The base rate is P(E): the fraction of all comparable windows where the event is present. Here
it is 100 / 10,000 = 1%.
Sensitivity, also called the true-positive rate, is P(A | E): among only the affected windows, the fraction that alert. It is 90 / 100 = 90%. The false-positive rate is P(A | not E): among only unaffected
windows, the fraction that nevertheless alert. It is 99 / 9,900 = 1%. Neither
rate is the fraction of alerts that are real.
The reverse conditional, P(E | A), is often called the positive predictive
value (PPV) or precision. It is what the page asks for: among positive alerts, how many
correspond to the event? The denominator is now all positive alerts, including true and
false positives.
100 degraded windows among all 10,000 windows.
90 alerts among the 100 degraded windows.
99 alerts among the 9,900 unaffected windows.
Make the rare event visible before applying the formula.
Start with 10,000 comparable windows. At a 1% base rate, 10,000 × 0.01 = 100 windows contain the event; 10,000 − 100 = 9,900 do not. Sensitivity of 90%
means the detector alerts on 100 × 0.90 = 90 affected windows. It misses the
other 10. The 1% false-positive rate means it also alerts on 9,900 × 0.01 = 99 unaffected windows.
So the detector emits 90 + 99 = 189 positive alerts in this cohort. Ninety correspond
to the event and 99 do not. There are 9,801 true negatives and 10 false negatives. The counts
here are whole because the chosen rates and cohort happen to produce whole values; estimated
rates in another cohort can produce fractional expected counts.
| Window condition | Alert | No alert | Total |
|---|---|---|---|
| Degradation present | 90 true positives | 10 false negatives | 100 |
| Degradation absent | 99 false positives | 9,801 true negatives | 9,900 |
| Total | 189 positive alerts | 9,811 no-alert windows | 10,000 |
90 + 99 = 189. If
the denominator is only 100 affected windows, you are calculating sensitivity, not the
probability that an alert is true.Only 90 of the 189 alerts come from affected windows.
Bayes’ rule combines the event’s base rate with how the detector behaves in each
condition: P(E | A) = P(A | E)P(E) / P(A). A positive alert can arise in two ways: the
event is present and the detector catches it, or the event is absent and the detector
fires by mistake. Therefore P(A) = P(A | E)P(E) + P(A | not E)P(not E).
Substitute the rates: (0.90 × 0.01) / ((0.90 × 0.01) + (0.01 × 0.99)). The
numerator is 0.009. The denominator is 0.009 + 0.0099 = 0.0189.
So P(E | A) = 0.009 / 0.0189 ≈ 0.476, or about 47.6%. That
matches the count: 90 / 189 ≈ 47.6% of positive alerts are true positives in this
modeled cohort.
The positive alert raises the probability from 1% before observing the alert to about 47.6% after it. This is a substantial update, but it still leaves the event less likely than not under these assumptions. A posterior is a conditional probability, not a verdict about cause or a command to ignore the page.
9 expected true alerts per 1,000 comparable windows.
9.9 expected false alerts per 1,000 windows.
Positive-alert probability under this model.
Event probability given a positive alert.
Look for evidence that separates the plausible explanations.
A positive alert is compatible with a real checkout degradation, a benign shift that confuses the detector, or a telemetry/data-quality fault. “Rollback now” and “ignore it” are both premature conclusions from this alert alone. First check an independent user-outcome measure for the same five-minute window, then compare the detector’s inputs and alert behavior by region, release, and request volume.
If raw checkout errors rise in the same windows and requests share a new release or dependency failure, the event gains support; that still does not uniquely identify the cause. If the detector score rises while independent errors stay flat, inspect feature freshness, missing data, threshold changes, and traffic mix. If one region alone changes, aggregate rates may be hiding a localized condition. Keep these observations separate from their explanations in the incident notes.
| Possible explanation | What it predicts | Next evidence to inspect |
|---|---|---|
| Checkout degradation is present | Independent user-visible error counts rise in the same cohort and window. | Request outcomes by region, release, status, and total request denominator. |
| Traffic or feature distribution shifted | Alert inputs change with traffic mix, while the event’s independent definition may not. | Raw feature values, request volume, threshold/configuration changes, and score by cohort. |
| Telemetry is delayed or malformed | Alert timing or score disagrees with source timestamps and event counts. | Ingestion lag, missing-sample rate, timestamp distribution, and raw event records. |
Let the calculation carry its inputs and denominator.
These functions take the base rate, sensitivity, and false-positive rate as probabilities
between zero and one, plus a cohort size. They return the expected number of true and
false alerts and their ratio. With the example inputs—0.01, 0.90, 0.01, and 10_000—the function returns 90 true positives, 99
false positives, and a posterior near 0.47619.
The resulting counts are expectations from rates, not necessarily observed counts; fractional counts can occur for other inputs. The function does not validate whether the rates were measured on representative, correctly labeled windows. That is an empirical question outside the arithmetic.
Both versions make the alert denominator explicit and reject invalid rates.
export type AlertEstimate = {
truePositive: number;
falsePositive: number;
positiveAlerts: number;
probabilityEventGivenAlert: number;
};
/**
* Estimates P(event | alert) from a base rate, sensitivity, and false-positive rate.
* Inputs are probabilities from 0 to 1. This is a model, not a calibration check.
*/
export function estimateAlertPosterior(
baseRate: number,
sensitivity: number,
falsePositiveRate: number,
windowCount: number
): AlertEstimate {
for (const [name, value] of [
['baseRate', baseRate],
['sensitivity', sensitivity],
['falsePositiveRate', falsePositiveRate]
] as const) {
if (!Number.isFinite(value) || value < 0 || value > 1) {
throw new Error(`${name} must be between 0 and 1`);
}
}
if (!Number.isSafeInteger(windowCount) || windowCount < 0) {
throw new Error('windowCount must be a non-negative safe integer');
}
const affectedWindows = windowCount * baseRate;
const unaffectedWindows = windowCount - affectedWindows;
const truePositive = affectedWindows * sensitivity;
const falsePositive = unaffectedWindows * falsePositiveRate;
const positiveAlerts = truePositive + falsePositive;
if (positiveAlerts === 0) {
throw new Error('posterior is undefined when the model predicts no positive alerts');
}
return {
truePositive,
falsePositive,
positiveAlerts,
probabilityEventGivenAlert: truePositive / positiveAlerts
};
}
package mathpractice
import (
"errors"
"math"
)
type AlertEstimate struct {
TruePositive float64
FalsePositive float64
PositiveAlerts float64
ProbabilityEventGivenAlert float64
}
// EstimateAlertPosterior estimates P(event | alert) from a base rate,
// sensitivity, and false-positive rate. Inputs are probabilities in [0, 1].
// This is a model, not a calibration check.
func EstimateAlertPosterior(baseRate, sensitivity, falsePositiveRate float64, windowCount int64) (AlertEstimate, error) {
for _, value := range []float64{baseRate, sensitivity, falsePositiveRate} {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 1 {
return AlertEstimate{}, errors.New("rates must be finite probabilities between 0 and 1")
}
}
if windowCount < 0 || windowCount > (1<<53)-1 {
return AlertEstimate{}, errors.New("window count must be between 0 and 2^53-1")
}
count := float64(windowCount)
affectedWindows := count * baseRate
unaffectedWindows := count - affectedWindows
estimate := AlertEstimate{
TruePositive: affectedWindows * sensitivity,
FalsePositive: unaffectedWindows * falsePositiveRate,
}
estimate.PositiveAlerts = estimate.TruePositive + estimate.FalsePositive
if estimate.PositiveAlerts == 0 {
return AlertEstimate{}, errors.New("posterior is undefined when the model predicts no positive alerts")
}
estimate.ProbabilityEventGivenAlert = estimate.TruePositive / estimate.PositiveAlerts
return estimate, nil
}
The same detector means something different in a rarer population.
Keep sensitivity at 90% and false-positive rate at 1%, but imagine the event occurs in
only 0.1% of comparable windows. In a new cohort of 10,000 windows, 10 are affected and
9,990 are not. The expected true-positive count is 10 × 0.90 = 9; the
expected false-positive count is 9,990 × 0.01 = 99.9. The positive-alert
denominator is 9 + 99.9 = 108.9, so the modeled PPV is 9 / 108.9 ≈ 8.3%.
The detector did not change; the population did. This is why a published “90% accurate” claim is incomplete. Accuracy itself mixes several outcomes and depends on how common the event is. For alert interpretation, preserve base rate, sensitivity, false-positive rate, cohort definition, and collection period separately.
In practice, the base rate can vary by service, region, time of day, release phase, and event definition. The detector’s sensitivity and false-positive rate can also shift when its input distribution or threshold changes. Before trusting a posterior, ask when and where each rate was measured, how labels were assigned, and whether alert rates drift across cohorts.
Windows with the defined event.
Expected detections among affected windows.
Expected alerts among unaffected windows.
Expected true alerts divided by all expected alerts.