“Errors are up” needs a population, a window, and a comparison.
To practice, use this illustrative dashboard snapshot: 68 of 1,000 checkout requests failed in the five minutes after the alert (6.8%). In the comparable five-minute window before it, 2 of 1,000 failed (0.2%). The deployment completed between those windows. These figures are authored teaching data; they are not observations from a live service.
Ask which users, routes, regions, and request types are affected. Is latency rising along with failures? Are attempts counted once, or can client retries inflate the denominator? Does the failure prevent purchase, or is a noncritical status panel slow? Impact determines whether to mitigate immediately and what comparison is meaningful.
Compare a precise first noteReveal after writing your own
- Observed
- In this teaching snapshot, 68/1,000 checkout requests failed in the alert window, versus 2/1,000 in the prior comparable window.
- Known context
- A deployment completed at 14:04. The time ordering is known; causation is not.
- Unknown
- Whether failures cluster by version, region, route, dependency, or client; and whether retries affect the count.
- Next question
- Do failures correlate with the new version after controlling for route and region, and do dependency signals show a matching change?
Put events in order before you connect them.
Align the alert, first failing request, deployment start and completion, dependency alarms, traffic shifts, and any operator changes. Normalize timestamps and time zones. Preserve request IDs, release identifiers, and the original metric window so later comparisons refer to the same events.
Timelines expose gaps: an error could precede the rollout; a dependency slowdown may begin first; or the failure may appear only as traffic shifts to new instances. A timeline supports a causal story, but adjacency alone does not establish one.
| Time | Record | What it establishes |
|---|---|---|
| 13:59–14:04 | Illustrative baseline window, 2/1,000 failed | Comparison rate for this exercise; validate equivalent traffic and counting. |
| 14:04 | Deployment marked complete | Controller reported completion; not proof every instance behaves correctly. |
| 14:05–14:10 | Illustrative alert window, 68/1,000 failed | Failure rate rose in the chosen window; cause remains open. |
| 14:10 | On-call receives alert | Detection time, which may lag the onset. |
A useful hypothesis predicts what you should observe.
Do not collect evidence without a question. Write down a few explanations that account for the same report, then state what each predicts. Choose checks that could make one explanation less likely as well as more likely.
| Possible explanation | Prediction | Evidence that would change belief |
|---|---|---|
| New application version has a regression | Failures are more common on the new version for comparable routes and regions. | Version-tagged request outcome and error type, with traffic share and denominator. |
| Payment dependency is degraded | Failures cross application versions and align with dependency timeouts or saturation. | Dependency latency/error signals and request traces spanning the call. |
| Traffic or configuration changed | Failures cluster after a routing, secret, flag, or traffic-shift event. | Change records, instance configuration identity, and results by route/region. |
Ask for the smallest safe check that separates explanations.
A version comparison by route may tell you whether the new build is implicated; dependency traces may show a shared downstream failure. Check the denominator and traffic mix before interpreting either. A chart aggregated across versions can hide a clear split, while a few handpicked traces can exaggerate one.
Good troubleshooting names the question, the expected observations, the owner, and the next decision. Prefer read-only queries and narrowly scoped comparisons first. Preserve raw events before dashboards roll up or logs expire. Change one variable at a time when the service is stable enough; during a severe incident, choose the fastest safe mitigation and document the tradeoff.
- Compare
- Failure rate and error class by application version, route, and region over the same interval.
- Control
- Check request counts, rollout traffic share, and client retry accounting before comparing percentages.
- Decision it informs
- Whether to pause or reverse rollout, investigate a shared dependency, or continue to another hypothesis.
Mitigation can be justified before root cause is certain.
Empirical thinking does not mean waiting for perfect proof while users are harmed. It means choosing an action for an explicit reason, considering its downside, and checking whether the outcome matches the prediction. Pausing rollout can stop exposure to a suspected version. Rolling back may restore service, but can also remove a needed fix or leave incompatible data changes. A feature flag, traffic shift, or dependency failover may be safer in a particular system.
Before acting, say what impact warrants the change, how to reverse it, which signal will show improvement, and who is watching. Record the pre-change state and decision time. Afterward, verify user-facing outcomes; “the dashboard turned green” is not enough if the affected requests still fail.
Choose a response under uncertainty
The illustrative error increase began just after the deployment. Checkout is still failing for users. What is the best next move?
A good handoff preserves uncertainty and makes the next step obvious.
Keep a short incident record with impact and window, verified observations, current hypotheses, actions and their results, remaining risks, evidence links, and a named next owner. Separate mitigation from the later root-cause analysis. Once service is stable, test the causal explanation against the timeline and look for evidence that could disconfirm it.
For the example: the failure rate increased in the post-deployment window; the deployment is a candidate cause; no version-sliced evidence has yet established that link. The next step is to compare outcomes by version, route, and region while checking dependency signals. If impact worsens before that check completes, act on the agreed mitigation threshold and continue gathering evidence.
References: Google SRE, Effective Troubleshooting and Managing Incidents; Google’s Incident Management Guide. Accessed 2026-10-01.