← Architecture
Failure and evidence What working means

Defining success

Count what the user got.

Somewhere you have a dashboard that says a service is up. It was right when someone built it, for the question they asked. Let’s ask a better question of one small search service, decide what “working” means to the person using it, and then replay a bad afternoon to see which definitions notice.

The skill to keep: Write down the good event, the denominator, the target, and the burn rate that pages, before you build the dashboard or change the system, and test the definition against a day you know went wrong.

TypeScriptGoOne search, four definitions, two recorded status checks.

01 / The prompt

“Add a status page for the campsite search.”

A state park lets campers search for open sites across two loops, the lake loop and the ridge loop, each with its own reservation backend. You ask an agent for a status page. What comes back works: it asks the load balancer’s health probe once a minute and shows a green light while the probe gets a 200.

One afternoon the lake loop’s backend fails. The search keeps answering 200, with the ridge sites it could find. A camper who asked only for the lake loop is told there are no sites. The status page is green all afternoon, because nothing it looks at went wrong.

Google’s SRE book puts the fix in one sentence: “Start by thinking about (or finding out!) what your users care about, not what you can measure” (Service Level Objectives). The prompt never said what a camper cares about, so the agent measured what was easy.

The question it never answered: what counts as a search that worked? Until that is written down, “is it up?” has as many answers as there are dashboards.

02 / Name the move

A good event, a denominator, a target, and a rate that pages.

Defining success means deciding, before you build a dashboard or change a system, what counts as a good event for the user, which events are counted at all, what share must be good, and how fast the rest may pile up before a person is woken. The SRE book calls the measure an SLI, “a carefully defined quantitative measure of some aspect of the level of service that is provided”, and the target an SLO, “a target value or range of values for a service level that is measured by an SLI.”

Count what the user got, out of what the user asked for. Decide the target and the paging rule before you look at the graph.

Here are four ways to count the same day. Only one of them is about a camper.

Four definitions of a good event for the campsite search
DefinitionCountsGood when
Uptime probeHealth probes, one a minuteThe probe got a 200.
HTTP statusSearchesThe status was not 5xx.
Has resultsSearchesA 200 that showed at least one site.
The camper’s outcomeSearchesA 200, not partial, within one second, whatever the number of sites.

Words to put in a prompt or a review

Good event
One request that gave the user what they came for, by a rule you wrote down.
Denominator
What you count good events out of. Searches, not probes; requests, not minutes.
SLI and SLO
The share of good events, and the share you promise: 99.5% over 30 days.
Error budget
The bad events the target allows: 0.5% of a month’s searches.
Burn rate
How fast you are spending it. A burn rate of 1 spends exactly the budget in 30 days.
Valid empty
A correct answer with nothing in it, like “sold out”. Good, not bad.
Why burn rates, not a thresholdThe SRE Workbook’s paging rules

The Workbook defines burn rate as “how fast, relative to the SLO, the service consumes the error budget” and recommends paging when 2% of a 30-day budget goes in an hour (a burn rate of 14.4, checked over the last hour and the last 5 minutes) or 5% in 6 hours (6, over 6 hours and 30 minutes) (Alerting on SLOs). The long window keeps a one-minute blip from paging; the short window lets the page stop soon after the problem does. The Workbook’s table is written for a 99.9% target; the burn rates carry over to this lesson’s 99.5% unchanged, because they are ratios.

03 / Follow one day

Replay one bad afternoon through four definitions.

The same generated day four times: sixty searches a minute, a quarter of them for sold-out weekends, and the lake backend failing from 13:00 to 14:30. Each chapter counts it differently. Watch when the pager fires, and when it does not. Open Try it to break the day a different way.

Failure and evidence

Is the campsite search working?

Uptime probe · Good: the health probe got a 200.

The day so far, per 10 minutes 00:00

On targetSome bad10% or more bad

Pager

Quiet

Burn rate, last hour

0.0

Month’s budget spent today

0%

Status page

All systems operational

01/ 04
Uptime probe

The status page asks the load balancer’s health probe once a minute. It is green all day, including while campers are told the lake loop has no sites.

00:00: 0% of the last hour was bad, a burn rate of 0.0. 0% of the month’s error budget spent today.

Reduced motion: choose a scene to see its completed state.

Read this scene

00:00: 0% of the last hour was bad, a burn rate of 0.0. 0% of the month’s error budget spent today.

Uptime probe. At 00:00 the pager is quiet; the burn rate over the last hour is 0.0; 0% of the month’s error budget has been spent today.

Watch restarts the story when you come back. Step through keeps your step. Try it replays a fresh day every time.

04 / Read it in code

A classifier, a burn rate, and a search that tells the truth.

Basic form is the definition itself, one function. In the wild replays a day through it and applies the paging rules. At the call site is the part that is easiest to forget: the search has to record whether its answer was partial, or no definition can tell “sold out” from “we could not check”.

The whole decision in one function: for each definition, which log lines count, and which of those are good. The camper’s outcome counts a complete “sold out” as good and a partial answer as bad.

TypeScriptReading
slo.ts
export const SLOW_MS = 1000;

/**
 * What counts as a good event under each definition. `null` means the event is not in
 * this definition's denominator at all.
 */
export function classify(event: Event, definition: Definition): 'good' | 'bad' | null {
	if (definition === 'uptime')
		return event.path === 'health' ? (event.status === 200 ? 'good' : 'bad') : null;
	if (event.path !== 'search') return null;
	switch (definition) {
		case 'http': // the server did not error
			return event.status < 500 ? 'good' : 'bad';
		case 'has-sites': // the camper saw at least one site
			return event.status === 200 && event.sites > 0 ? 'good' : 'bad';
		case 'outcome': // the camper got a complete, true answer in time, even if it is "sold out"
			return event.status === 200 && !event.partial && event.ms <= SLOW_MS ? 'good' : 'bad';
	}
}
GoAlongside
main.go
const SlowMs = 1000

// Verdict is "good", "bad", or "" when the event is not in this definition's denominator.
func Classify(e Event, definition string) string {
	if definition == "uptime" {
		if e.Path != "health" {
			return ""
		}
		return verdict(e.Status == 200)
	}
	if e.Path != "search" {
		return ""
	}
	switch definition {
	case "http": // the server did not error
		return verdict(e.Status < 500)
	case "has-sites": // the camper saw at least one site
		return verdict(e.Status == 200 && e.Sites > 0)
	default: // "outcome": a complete, true answer in time, even if it is "sold out"
		return verdict(e.Status == 200 && !e.Partial && e.Ms <= SlowMs)
	}
}

func verdict(good bool) string {
	if good {
		return "good"
	}
	return "bad"
}
The behavior these examples promiseChecked by 19 shared cases in TypeScript and Go
  • Uptime counts health probes only; the other three count searches only. The camper’s outcome is good at exactly 1000 ms and bad at 1001.
  • The target is 99.5% over 30 days, so 5 bad in every 1000 is on target. A day’s burn rate is its bad share divided by 0.005; its share of the month’s budget divides by 30 days’ worth.
  • At each minute, a rule holds when the bad share over its long window and over its short window both reach its burn rate: 14.4 over 60 and 5 minutes, or 6 over 360 and 30. The pager fires in any minute where either holds.
  • The search asks each loop it needs. A loop that fails is left out and the answer is marked partial; only when every loop fails is it a 503.

Every expectation in the shared cases was produced by a separate model written from these rules, kept beside the examples in examples/model/, not copied from either implementation. Besides the sixteen generated days, three hand-built days sit exactly on the paging boundaries.

Reading the TypeScriptA union verdict and integer burn rates

classify returns 'good' | 'bad' | null, so “not in this denominator” cannot be mistaken for either. The burn-rate check compares bad × 1000 × 10 with burn × 10 × 5 × total in whole numbers, so “exactly 14.4” means the same thing in every language.

Reading the GoPointers for never, goroutines per loop

FirstPage is a *int, nil when the pager never fires, which encodes to JSON null like the TypeScript. The handler asks each loop in its own goroutine and counts failures under a mutex before deciding between 200 and 503.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types slo.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print, for the day of the lake outage:

go.mod
module heyrian.dev/lessons/defining-success

go 1.23
uptime: 0 bad of 1440, burn 0.00, 0% of the month's budget, never pages
http: 0 bad of 86400, burn 0.00, 0% of the month's budget, never pages
has-sites: 22950 bad of 86400, burn 53.13, 177% of the month's budget, pages at 00:00 (fast)
outcome: 4050 bad of 86400, burn 9.38, 31% of the month's budget, pages at 13:05 (fast)

The traffic is generated; every count and time is computed from it by the code above.

05 / Review the agent’s diff

“Fixes the flaky search alert.”

The complaint is fair: an alert that pages every night on sold-out weekends gets ignored, and an ignored alert is worse than none. Read what the fix counts before you merge it.

The agent’s pull request

“Fixes the flaky search alert. It was paging every night because sold-out weekends return no sites. It has been quiet for 24 hours.”

// monitoring/slo.ts
			export function isGood(e: SearchEvent): boolean {
			(removed)  return e.status === 200 && e.sites > 0;
			(added)  return e.status < 500;
			}
			
You are reviewing this change. What do you do?

06 / How it fails

A definition fails by missing the outage or by crying wolf.

A definition of success has failure modes of its own. Here is each one for this search, and what the shared cases show.

How a definition of success goes wrong, one generated day each
What goes wrongWhat the on-call engineer seesThe fixFrom the cases
It counts the probe, not the searchGreen all afternoon while lake-loop campers are told “no sites”.Count searches.uptime, lake outage: first page never
It counts errors, and the failure is a 200Green again.Make the service say partial, and count partial as bad.http, lake outage: 0 bad
It counts a valid empty as a failureA page at midnight on a normal day, and every night after.Count “sold out” as good.has-sites, normal day: burn 50.00
It does not count slownessGreen while every search takes about three seconds.Put a time limit in the good event.http, slow day: 0 bad; outcome: 5400
It pages on one bad minuteA page that clears before anyone logs in.Require the long window too.five-minute burst: fast rule 6 minutes
It sits exactly on the lineA page at exactly 14.4 or exactly 6.Decide which side the line is on, and test it.exact 14.4: pages at 00:00; exact 6: 360 minutes

When a definition cries wolf, the tempting fix is a lower target or a looser threshold. Fix the good event instead: the target describes the service, and it should not have to absorb a mistake in what you count. And when a definition misses a failure you know happened, the fix is often in the service, not the monitoring: here, recording partial.

Knowing which request went wrong, and why, is Observability across a request. Deciding what to do about a dependency that keeps failing is Containing failure.

07 / Is it worth it?

You pay for a partial flag and a page of decisions. Here is what they buy.

The probe-based status page costs nothing to build and is right about the load balancer. Hold both up against the changes a search service gets.

The same four changes, against a probe and against a written definition
ChangeUptime probeThe camper’s outcome
A second client: the park’s phone appThe probe does not know the app exists.The app’s searches are searches; they count.
Replace a loop’s reservation backendNothing to change, and nothing to learn.The burn rate says whether the new one is better, from the first day.
Change a rule: one second becomes 800 msThere is no rule.One constant, and the replayed bad day says what it would have done.
A second team takes over the ridge loopThey get a green light that was never about them.They inherit a definition, a budget, and a test.

This lesson’s technique is itself the measurement plan, so the plan for adopting it is short. Before you switch definitions:

  • Replay the last month’s logs through the new definition and write down every minute it would have paged. Match each one to an incident you know of, or to a night nobody was woken for.
  • Keep both definitions side by side for two weeks. Accept the new one when it has paged for every real incident and for nothing else.
  • Look at the budget spent per week. A definition that spends half the month’s budget on normal days is counting something that is not a failure.

This page did not replay a real park’s logs, so it has no numbers of that kind to give you. The before picture of a change is Baseline before you change.

08 / Ask for it

Two status checks, seven moments from four days.

We asked two agents, both running Claude Sonnet, for this status check. Both got the same description of the log, including that a search can correctly return zero sites. One prompt stopped there. The other added one paragraph: decide what counts as a good search from a camper’s point of view, which lines count, the 30-day target, and how fast the budget may burn before someone is paged; write it to SLO.md and implement it. It named no latency limit, no rule for partial answers, and no number. Then a script fed both checks the same logs, cut at seven moments, and compared what they said with what this lesson’s definition says.

What each status check said, run 2026-09-23, beside what this lesson’s definition does at the same moment
The log, and the momentThis lessonPlain promptArchitecture prompt
A normal day, at 23:59QuietQuietQuiet
A normal day, at 03:00 (sold-out searches all night)QuietQuietQuiet
Lake backend failing, answers marked partial, at 13:20PagesQuiet ✗Pages
Every search answering 503, at 13:20PagesPagesPages
Every search 2.5 s slower (2.7 to 3.1 s), at 13:20PagesQuiet ✗Pages
One minute of 503s at 12:00, at 12:03QuietQuietQuiet
After the lake outage ended, at 16:00QuietQuietQuiet

Both handled “sold out” correctly, because the prompt told them to. Both paged when every search failed, and both stayed quiet at night and through a one-minute blip. The difference is in the two failures that are not errors.

The plain build looked at 75 partial answers out of 100 and decided the search was healthy: “backend needs attention but this can wait for business hours.” That is a real decision, and a defensible one for someone who thinks of the search as a server. For a camper who asked for the lake loop, it said there were no sites. It also ignored searches that took up to 3.1 seconds, because its idea of too slow started at five.

The architecture build wrote a good event that counts partial answers as bad and puts a limit on time, and it paged for both. Its time limit was three seconds, and the slow afternoon’s searches took between 2.7 and 3.1, so it caught that one by a small margin: a burn rate of 16.7 against its 14.4.

status.ts · plain prompt
// 5. Degraded but working: one reservation backend is down and results are
//    partial, but the search endpoint itself is answering fine. This is
//    the graceful-degradation path working as intended.
else if (searchTotal >= MIN_PARTIAL_SAMPLE && searchPartialRate >= PARTIAL_RATE_DEGRADED) {
  healthy = true;
  page = false;
  reason = `${searchPartial}/${searchTotal} recent searches returned partial results (a reservation backend is down) — search itself is healthy; backend needs attention but this can wait for business hours`;
}
status.ts · architecture prompt
const LATENCY_THRESHOLD_MS = 3000;
const TARGET = 0.995; // 30-day SLO target -> 0.5% error budget
// …
function isGood(l: LogLine): boolean {
  return l.status === 200 && l.partial !== true && l.ms <= LATENCY_THRESHOLD_MS;
}

Nobody told either agent what a camper counts as a failure, so each decided. Asking for the procedure made the agent decide out loud, in SLO.md, where a reviewer can disagree. The line the runs showed was missing is the good event itself: a search is good when it is a complete answer within n ms; a partial answer is bad even when it is a 200.

How the runs were made and checkedOne run each, recorded as written
  • The plain prompt was sent twice. The first agent stalled after creating its folder and wrote nothing; the same prompt, word for word, went to a fresh agent after the architecture agent had finished. The builds share no files and no ports, so the order could not affect them.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples, including the architecture build’s SLO.md. The sample logs each agent made to test with are left out. The checker restores each build into a fresh folder per moment.
  • The “This lesson” column is what this lesson’s own definition says at that minute; the lesson’s tests check all seven. A ✗ marks a status check that disagrees with it.
  • This is one sample of each prompt, not a measurement of a model. The transcript audit and every number are in the run notes beside the examples.

09 / Hold it there

Make a known bad day part of the test suite.

A definition drifts the same way code does: someone quiets a noisy alert, someone adds a new kind of answer. Three checks keep it honest.

  1. The rules someone already tested

    Take the paging rules from the SRE Workbook rather than inventing thresholds: a burn rate of 14.4 over an hour and five minutes, and 6 over six hours and thirty minutes (Alerting on SLOs). Their long windows ignore blips; their short windows let a page clear when the problem does.

  2. A test that replays a bad day

    The definition is code, so test it like code. This lesson’s suite replays the lake outage, a 503 afternoon, a slow afternoon, a normal night of sold-out searches, and a one-minute blip, and asserts which of them page at which minute. Change what counts as good and a test says what you changed.

    slo.spec.ts
    describe('the moments the recorded runs were checked at', () => {
    	// The checker's logs: 20 searches a minute, cut at `until`, with an optional one-minute blip.
    	const cases: [Outage, number, number | null, boolean][] = [
    		['none', 1439, null, false],
    		['none', 180, null, false],
    		['swallowed', 800, null, true],
    		['errors', 800, null, true],
    		['slow', 800, null, true],
    		['none', 723, 720, false],
    		['swallowed', 960, null, false]
    	];
  3. A check that the service records what the definition needs

    A definition built on partial is only as good as the code that sets it. The call-site tests fail a loop on purpose and assert the event says partial; the checker in section 08 feeds real log lines to whole status checks and compares what they say.

    check-runs.mjs
    const questions = [
    	{ name: 'a normal day, at 23:59', log: () => day('none', 1439), lessonPages: false },
    	{ name: 'a normal day, at 03:00 (sold-out searches all night)', log: () => day('none', 180), lessonPages: false },
    	{ name: 'lake backend failing, answers marked partial, at 13:20', log: () => day('swallowed', 800), lessonPages: true },
    	{ name: 'every search answering 503, at 13:20', log: () => day('errors', 800), lessonPages: true },
    	{ name: 'every search 2.5 s slower (2.7 to 3.1 s), at 13:20', log: () => day('slow', 800), lessonPages: true },
    	{ name: 'one minute of 503s at 12:00, at 12:03', log: () => day('none', 723, 720), lessonPages: false },
    	{ name: 'after the lake outage ended, at 16:00', log: () => day('swallowed', 960), lessonPages: false }
    ];
Build UIs?Core Web Vitals are a definition of success with a threshold and a percentile. Your search page can report its own.

Where it already is in your components

Core Web Vitals are exactly this move, done for you. Interaction to Next Paint defines a good event, “An INP below or at 200 milliseconds means a page has good responsiveness,” and a target over a denominator, “the 75th percentile of page loads recorded in the field, segmented across mobile and desktop devices” (web.dev, INP). When you watch that number, you are watching an SLO someone else wrote.

When you have to own it

No vital knows what a campsite search is. The results component does: it knows whether the answer was complete, whether it said “sold out”, and how long the camper waited. So it can report one good-or-bad event per search from where the camper sits, and it can show “some loops did not answer” instead of an empty list that reads as sold out.

A search button that reports one event per search: good when the answer came back complete within a second.

ReactAlready in your code
Search.tsx
import { useState } from 'react';

// Measure the outcome where the user sees it: did the search come back complete, and fast
// enough? Send one good-or-bad event per search, and let the server count them.
function report(good: boolean, ms: number) {
	navigator.sendBeacon('/rum', JSON.stringify({ name: 'search', good, ms: Math.round(ms) }));
}

export default function Search() {
	const [message, setMessage] = useState('');

	async function search(nights: string) {
		const started = performance.now();
		try {
			const response = await fetch(`/search?nights=${encodeURIComponent(nights)}`);
			const body = (await response.json()) as { sites: unknown[]; partial: boolean };
			const ms = performance.now() - started;
			report(response.ok && !body.partial && ms <= 1000, ms);
			setMessage(`${body.sites.length} sites`);
		} catch {
			report(false, performance.now() - started);
			setMessage('Search failed. Try again.');
		}
	}

	return (
		<p>
			<button onClick={() => search('2026-10-10')}>Search</button>{' '}
			<span role="status">{message}</span>
		</p>
	);
}

10 / Make the call

A probe tells you the server is there. Only a definition tells you it works.

Keep the uptime probe: it answers a real question, whether the load balancer can reach the process, and it is the right signal for a service with no users yet. Write a definition of success when people depend on the answers, when an answer can be wrong without being an error, or when you are about to change something and need to know whether it helped.

Revisit the definition when the service learns a new kind of answer, when a new client arrives, or when the pager has gone off for something nobody needed to fix.

Take it with you

Explain it without saying “SLO”: “A search worked if the camper got a complete answer within a second, even if the answer was sold out. We promise that for 99.5% of searches over a month, and we wake someone when we are spending that allowance fourteen times too fast.” Then pick one service you run and write its good event in one sentence.

Paste into your next prompt, and fill in the blanks

Before building monitoring for <feature>, decide and write down:
- The user's outcome: a <request> is good when <complete, correct answer>
  within <time>. <A valid empty answer, such as "sold out"> counts as good.
  <A partial or wrong answer> counts as bad even when the status is 200.
- The denominator: every <user request>. Not health probes, not minutes.
- The target: <99.5%> good over <30 days>.
- Paging: when the budget burns 14.4 times too fast over both the last hour
  and the last 5 minutes, or 6 times over 6 hours and 30 minutes.
Make the service record whatever the definition needs, such as a partial
flag, and add a test that replays a known bad day and a normal day.
Connections to follow nextRelated lessons

Take the search into your editor. Add a third loop that is often slow, and decide whether a search that waited for it counts as good.

Back to architecture →