← Architecture
Failure and evidence Blast radius

Containing failure

Its bad day stays its own.

Most pages you build pull in something they could live without: a badge, a widget, a recommendation, a count. On a good day it costs a few milliseconds. Let’s find out what it costs on a bad one, and how to make sure a feature you could live without never takes the rest of the site down with it.

TypeScriptGoOne site, one optional API, two recorded builds.

01 / The prompt

“Show today’s snow report in the site header.”

You run the website for a ski resort: lift status, trail status, lift-ticket prices. You ask an agent to add today’s snow report to the header, from the weather API the resort pays for. What comes back works. Every page now says “142 cm, updated 06:00” at the top, and the lift-ticket page still loads in a fifth of a second.

The header is shared, so every page now makes one more call before it can answer. Nothing in the prompt said what a page should do while that call is slow. So each page does the default: it waits.

A server can only work on so many requests at once. On a busy morning, if the weather API takes 8 seconds to answer, every page holds its place for 8 seconds. The places run out, and new requests queue behind pages that are only waiting for a snow report. Soon nobody can buy a lift ticket, and the cause is a feature nobody would miss for a morning.

That is the question the prompt never answered: how much of the site is the snow report allowed to take down with it?

02 / Name the shape

Its bad day stays its own.

Containing failure means deciding in advance how far one dependency’s failure may spread, and building the limits that stop it there. The distance it spreads is its blast radius. Here, the snow report’s blast radius is every page.

An optional dependency gets a limit on how long it may hold a page, a fallback for when that runs out, a cap on how much of the server it may use at once, and a switch that stops calling it while it is failing.

Here is what each part of a page needs, and what it can do without.

What each part of a page depends on
Part of the pageComes fromWithout it
Lift statusThe resort’s lift serviceThere is no page. Say so.
Trail statusThe resort’s lift serviceThere is no page. Say so.
Lift-ticket pricesThe resort’s price serviceThere is no page, and this is the one that makes money.
Snow report, in every headerA third-party weather APIThe last known report, marked as possibly out of date. Every page still works.

Words to put in a prompt or a review

Blast radius
Everything that stops working when one part fails.
Deadline
How long a page will wait for a dependency before going on without it.
Fallback
What the page shows instead: a stale value, a placeholder, or nothing.
Circuit breaker
Stops calling after repeated failures, then lets one probe test for recovery.
Bulkhead
A cap on how much of the server one dependency may hold at once.
Load shedding
Refusing work beyond the cap at once, instead of queuing it.
Where the names come fromCircuits, ships, and a book about production

Martin Fowler credits the circuit breaker to Michael Nygard: “In his excellent book Release It, Michael Nygard popularized the Circuit Breaker pattern to prevent this kind of catastrophic cascade,” and describes the third state as “half open … ready to make a real call as trial to see if the problem is fixed” (CircuitBreaker). A bulkhead is the wall between compartments in a ship’s hull, so a leak floods one compartment and not the ship.

03 / Follow one busy minute

Watch twenty slots fill with pages waiting on a snow report.

Twenty page requests a second, and a server that works on twenty at a time. From 10 seconds to 30 seconds, the weather API takes 8 seconds to answer. First the shared build, where every page waits for its header. Then the contained build, the same minute. Open Try it to switch each mechanism on and off yourself.

Failure and evidence

What happens to every page when the header’s API slows down?

Shared build · 7 s into the morning

Server slots

  1. L
  2. T
  3. $

Own dataWaiting on weatherFreeL lifts · T trails · $ tickets

What readers see

Waiting for a slot
0
Last lift-ticket page
150 ms
Header on the last page
142 cm, updated 06:00
01/ 04
A busy morning

Twenty page requests a second. Each page asks for its own data and, for the header, today’s snow report. Most slots finish in a fifth of a second.

3 of 20 slots busy, 3 of them waiting on the weather API; 0 requests queued. The last lift-ticket page took 150 ms.

Reduced motion: choose a scene to see its completed state.

Read this scene

3 of 20 slots busy, 3 of them waiting on the weather API; 0 requests queued. The last lift-ticket page took 150 ms.

Shared build, 7 seconds in. 3 of 20 slots busy, 3 waiting on the weather API. 0 requests waiting for a slot. Last lift-ticket page: 150 ms. Breaker: closed.

Watch restarts the story when you come back. Step through keeps your step. Try it runs a fresh minute every time.

04 / Read the shape

A breaker, a guard, and a header that can go without.

Basic form is the circuit breaker alone, a small state machine. In the wild it sits inside a guard with a bulkhead and a deadline, and every weather call passes through the guard. At the call site the page asks the guard, and takes the fallback when the answer is no.

Notice the order inside the guard: bulkhead, then breaker, then the call with its deadline. The cheapest “no” comes first, and a call the breaker lets through is always one the server has room for.

The circuit breaker on its own. It counts failures in a row while closed, opens at five, refuses calls for 5 seconds, then lets exactly one probe through. The probe’s result closes it or opens it again.

TypeScriptReading
containment.ts
/** Stops calling a dependency after repeated failures, then lets one probe test it. */
export class CircuitBreaker {
	state: 'closed' | 'open' | 'half-open' = 'closed';
	opens = 0;
	#failures = 0;
	#openedAt = 0;
	#probing = false;
	readonly threshold: number;
	readonly cooldownMs: number;

	constructor(threshold = 5, cooldownMs = 5000) {
		this.threshold = threshold;
		this.cooldownMs = cooldownMs;
	}

	/** May a call go out now? A probe is the one call allowed while half-open. */
	allow(now: number): 'call' | 'probe' | 'reject' {
		if (this.state === 'open' && now - this.#openedAt >= this.cooldownMs) {
			this.state = 'half-open';
			this.#probing = false;
		}
		if (this.state === 'closed') return 'call';
		if (this.state === 'half-open' && !this.#probing) {
			this.#probing = true;
			return 'probe';
		}
		return 'reject';
	}

	/** Report how a call ended. Only calls made while closed, and the probe, count. */
	record(ok: boolean, probe: boolean, now: number): void {
		if (probe) {
			if (ok) this.#close();
			else this.#open(now);
		} else if (this.state === 'closed') {
			this.#failures = ok ? 0 : this.#failures + 1;
			if (this.#failures >= this.threshold) this.#open(now);
		}
	}

	#open(now: number) {
		this.state = 'open';
		this.opens++;
		this.#openedAt = now;
		this.#failures = 0;
		this.#probing = false;
	}

	#close() {
		this.state = 'closed';
		this.#failures = 0;
		this.#probing = false;
	}
}
GoAlongside
main.go
// CircuitBreaker stops calling a dependency after repeated failures, then lets one probe test it.
type CircuitBreaker struct {
	State     string // "closed", "open", or "half-open"
	Opens     int
	Threshold int
	Cooldown  int64 // milliseconds
	failures  int
	openedAt  int64
	probing   bool
}

func NewCircuitBreaker(threshold int, cooldownMs int64) *CircuitBreaker {
	return &CircuitBreaker{State: "closed", Threshold: threshold, Cooldown: cooldownMs}
}

// Allow says whether a call may go out now: "call", "probe" (the one call allowed while
// half-open), or "reject".
func (b *CircuitBreaker) Allow(now int64) string {
	if b.State == "open" && now-b.openedAt >= b.Cooldown {
		b.State, b.probing = "half-open", false
	}
	if b.State == "closed" {
		return "call"
	}
	if b.State == "half-open" && !b.probing {
		b.probing = true
		return "probe"
	}
	return "reject"
}

// Record reports how a call ended. Only calls made while closed, and the probe, count.
func (b *CircuitBreaker) Record(ok, probe bool, now int64) {
	switch {
	case probe && ok:
		b.State, b.failures, b.probing = "closed", 0, false
	case probe:
		b.open(now)
	case b.State == "closed":
		if ok {
			b.failures = 0
		} else {
			b.failures++
		}
		if b.failures >= b.Threshold {
			b.open(now)
		}
	}
}

func (b *CircuitBreaker) open(now int64) {
	b.State, b.openedAt, b.failures, b.probing = "open", now, 0, false
	b.Opens++
}
The behavior these examples promiseChecked by 48 shared cases in TypeScript and Go
  • The breaker counts failures in a row while closed. A success resets the count. At five it opens for 5 seconds and refuses calls. After that, the first call is a probe and the rest are refused until the probe ends; the probe’s success closes the breaker, and its failure opens it again. Results from calls made before it opened are ignored.
  • The bulkhead allows four weather calls at once. A fifth does not wait; the page takes the fallback. A breaker refusal gives the bulkhead place back.
  • With the deadline, a call that would take longer than 300 ms ends at 300 ms as a failure. Without the fallback, any weather failure fails the page.
  • The simulation: twenty slots, a request every 50 ms for a minute cycling through lifts (20 ms), trails (40 ms), and tickets (150 ms), each also making the header’s weather call. A page holds its slot until both calls are done. The weather API answers in 120 ms, except from 10 s to 30 s, when it takes 8 seconds or fails in 5 ms.

Every expectation in the shared cases was produced by a separate model written from these rules, kept beside the examples in examples/model/, not copied from either implementation. It covers all sixteen combinations of the four mechanisms under three weather conditions.

Reading the TypeScriptPrivate fields, a union ticket, AbortSignal

The breaker’s counters are private # fields, so only allow and record can change them. Ticket is a union on go: a caller has to check it before it can read probe, and leave ignores a ticket that never went out. AbortSignal.timeout enforces the deadline on the real fetch.

Reading the GoA struct ticket and a context deadline

Ticket is a struct whose Go field plays the same role as the TypeScript union. context.WithTimeout puts the deadline on the request, and the handler uses r.Context(), so a reader who closes the tab also cancels the weather call. A deadline of 0 means none.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types containment.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/containing-failure

go 1.23
shared, weather normal: p95 150 ms, tickets p95 150 ms, 0 degraded, 0 failed, 1200 weather calls, breaker opened 0
shared, weather slow: p95 19760 ms, tickets p95 19700 ms, 0 degraded, 0 failed, 1200 weather calls, breaker opened 0
shared, weather down: p95 150 ms, tickets p95 150 ms, 0 degraded, 400 failed, 1200 weather calls, breaker opened 0
contained, weather slow: p95 150 ms, tickets p95 150 ms, 432 degraded, 0 failed, 779 weather calls, breaker opened 4
contained, weather down: p95 150 ms, tickets p95 150 ms, 412 degraded, 0 failed, 798 weather calls, breaker opened 4

These are simulated milliseconds from a model of one server, not a load test. They show how the mechanisms change the shape of a bad minute, not what your server would measure.

05 / Review the agent’s diff

“Pages no longer fail when the weather API is down.”

A bug report said every page showed an error when the weather API was down, and this change fixes that. Read what it does when the API is slow instead.

The agent’s pull request

“Pages no longer fail when the weather API is down. The snow report retries three times, then the header just leaves it out. All 18 tests pass.”

// src/routes/+layout.server.ts
			export async function load({ fetch }) {
			(removed)  const snow = await fetch(WEATHER_URL).then((r) => r.json());
			(added)  let snow = null;
			(added)  for (let attempt = 1; attempt <= 3 && !snow; attempt++) {
			(added)    try {
			(added)      const response = await fetch(WEATHER_URL);
			(added)      if (response.ok) snow = await response.json();
			(added)    } catch {}
			(added)  }
			  return { snow };
			}
			
You are reviewing this change. What do you do?

06 / How it fails

Down is the easy case. Slow is the one that spreads.

Here is each way the snow report can go wrong, what a skier sees, and what the contained build does. The numbers come from the shared cases, one busy minute each.

Failure modes of the header’s weather call, over one busy minute
What goes wrongWhat a skier seesWhat the contained build doesShared → contained
Slow: 8 secondsEvery page, on time, with “snow report unavailable” and the last known depth.Stops waiting at 300 ms, then the breaker stops calling for 5 s at a time.lift tickets p95 19,700 → 150 ms
Down: fails in 5 msThe same stale header.Takes the fallback; the breaker opens after five failures.failed pages 400 → 0
Slow, with a breaker but no deadlineEvery page, slow.Nothing. An 8-second answer is still a success, so the breaker never counts a failure.breaker opens 0
Slow, with a deadline but no fallbackAn error page, quickly.Fails fast. Better than waiting, worse than degrading.failed pages 400
Slow, and the bulkhead is fullThe stale header.Skips the call: four are already waiting.weather calls 1,200 → 732
The probe failsThe stale header, a little longer.Opens again for another 5 seconds.breaker opens 4 in 20 s
The API recoversFresh reports again.The next probe after the cooldown succeeds and the breaker closes. Up to 5 s late.see chapter 4 of the story

Two rows are worth a second look. A breaker only sees failures. Without a deadline, a slow answer is a success, however late, and the breaker stays closed while the slots fill. And a deadline without a fallback fails fast, which is better than waiting but still puts an error page in front of someone buying a lift ticket. The mechanisms work as a set.

Each has a lesson of its own: Timeouts, deadlines, and races, Backpressure and queues, and on the client side, Error boundaries in UI.

07 / Is it worth it?

You pay for a stale header and four small parts. Here is what they buy.

The shared build is simpler and shows the freshest report on every page, on a good day. Hold both builds up against the changes every site eventually gets.

The same four changes, made to each build
ChangeShared buildContained build
A second client: the resort’s phone appThe app’s own calls add to the weather API’s load on its worst day.If the app goes through the same guard, it gets the same limits.
Replace the weather vendorThe new vendor’s bad days are every page’s bad days.Same guard, new URL. The limits do not change.
Change a rule: a 500 ms deadlineThere is no deadline to change.One constant. No other difference.
A second team takes over the headerTheir dependency can still take down the ticket page.Its limits are written down, and a test holds them.

Before you add containment, write down what you will measure and what you would accept, so it is judged by more than one good afternoon:

  • Page time at the 95th percentile, per page, split by whether the header was fresh or stale. The lift-ticket page should not move when the weather API has a bad day.
  • Share of pages with a stale header, per hour. This is the price you are paying; it should be near zero on a normal day.
  • Calls per second to the weather API while it is failing. With a breaker it should drop to about one probe per cooldown.
  • Slots in use, or requests waiting, from the server. A queue that grows while one dependency is slow is the signal this lesson is about.

This page did not run the site for real skiers, so it has no numbers to give you. The simulation’s numbers are a model, and the page says so where it shows them. Choosing the threshold is Defining success; taking the before picture is Baseline before you change.

08 / Ask for it

Two prompts, two builds, one weather API that stops answering.

We sent two agents the same request for this site at the same time, both running Claude Sonnet. One prompt described the three pages, the header, and the services. The other added two sentences: the snow report is optional, so no page may be slow or fail because of it; and a struggling weather API must not get a full stream of requests, and recovery should be noticed. It named no deadline, fallback, breaker, or cap. Then a script started each build against fake services and broke the weather API.

What the checker found, run 2026-09-23
What the checker didPlain promptArchitecture prompt
An ordinary page200 after 10 ms, with the report200 after 8 ms, with the report
The weather API never answersNo reply within 40.0 s200 after 1.2 s, without the report
50 pages at once while it never answers0 of 50 answered within 2 s; slowest 40.0 s; 50 weather calls50 of 50 answered within 2 s; slowest 1.2 s; 50 weather calls
The weather API answers 500200 after 9 ms, without the report200 after 10 ms, without the report
The weather API refuses connections200 after 8 ms, without the report200 after 9 ms, without the report
Ten seconds of 500s, 100 pages, then it recovers100 of 100 pages 200; 100 weather calls; report back after 5 ms100 of 100 pages 200; 4 weather calls; report back after 5.3 s

The plain build contained errors and not waiting. It wrapped the weather call in a try and rendered “Snow report unavailable” when it failed, so a 500 or a refused connection cost the page nothing. But its fetch has no deadline, so when the weather API stopped answering, every page stopped answering with it: none of fifty pages came back in the 40 seconds the checker waited.

The architecture build held its 1.2-second deadline on every page, and under steady failure its breaker cut 100 calls down to 4. It got the burst wrong. When fifty pages arrived together, all fifty weather calls went out before the first one had failed, because a breaker learns from calls that have finished. Only a cap on calls in flight, the bulkhead, stops a burst at the door.

server.ts · plain prompt
async function fetchJson<T>(url: string): Promise<T> {
  const res = await fetch(url);
  if (!res.ok) {
  // …

async function renderSnowHeader(): Promise<string> {
  try {
    const snow = await fetchJson<SnowReport>(`${WEATHER_URL}/snow`);
    const text = `${snow.depthCm} cm, updated ${snow.updated}`;
    return `<header class="snow-report"><p>${escapeHtml(text)}</p></header>`;
  } catch {
    return `<header class="snow-report"><p>Snow report unavailable</p></header>`;
  }
}
server.ts · architecture prompt
  if (this.state === "open") {
    const cooldownElapsed = Date.now() - this.openedAt >= this.cooldownMs;
    if (cooldownElapsed && !this.probing) {
      // Cooldown is over: let one probe request through to check for
      // recovery. Everyone else in the meantime still gets the cache.
      this.state = "half-open";
      return this.probe();
    }
    return { report: this.cache, stale: true };
  }

  if (this.state === "half-open") {
    // A probe is already in flight (or just finished on another
    // request); don't pile on additional live calls.
    return { report: this.cache, stale: true };
  }

  // closed: call normally.
  return this.probe();
}

Both agents turned “optional” into a fallback, and only one turned “promptly” into a number. The line the runs showed was missing is the one that makes “slow” a failure the other mechanisms can see: wait at most n ms for it, and never have more than k calls to it in flight.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. The only differences were the Architecture block and the output folder.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. The checker restores them into a temporary folder and starts a fresh server for every question, with fake lift, price, and weather services.
  • This is one sample of each prompt, not a measurement of a model. Another run could land differently. The transcript audit and every number are in the run notes beside the examples.

09 / Hold it there

Make a hanging dependency part of the test suite.

Containment is easy to undo by accident: one await in the wrong place, one new widget in the header. Three kinds of check keep it in place.

  1. The framework’s own boundaries

    SvelteKit streams a promise that a server load returns without awaiting, so the page renders and the header fills in later, with two caveats from its docs: streaming “will only work when JavaScript is enabled”, and on platforms that buffer responses “the page will only render once all promises resolve” (Streaming with promises). A <svelte:boundary> catches errors “during rendering or while running effects”, but not “in event handlers or after a setTimeout or async work” (svelte:boundary). React’s Suspense and error boundaries play the same parts. None of them adds a deadline or a breaker; they decide what the reader sees.

  2. A test that hangs on purpose

    The rule “the header never holds a page longer than its deadline” is a test, not a review comment. This lesson’s suite starts a fake weather API that does not answer for a second, and asserts the layout load returns the last known report in under 900 ms. Delete the deadline and the test fails.

    containment.spec.ts
    it('stops waiting at the deadline and shows the last known report', async () => {
    	mode = 'hang';
    	const started = Date.now();
    	expect(await loadHeader(new WeatherGuard(ALL_ON), base, last)).toEqual({
    		snow: last,
    		stale: true
    	});
    	expect(Date.now() - started).toBeLessThan(900);
    });
  3. A check on what actually happens under load

    A unit test sees one request. The failure in this lesson needs many at once. The checker in section 08 sends fifty pages while the weather API hangs and counts how many answer, and counts the calls that reach the weather API while it fails.

    check-runs.mjs
    async '50 pages at once while the weather API never answers'() {
    	weather.mode = 'hang';
    	const replies = await Promise.all(Array.from({ length: 50 }, () => page()));
    	const answered = replies.filter((r) => typeof r.status === 'number');
    	return {
    		answeredWithin2s: answered.filter((r) => r.ms <= 2000).length,
    		answered: answered.length,
    		slowestMs: Math.max(...replies.map((r) => r.ms)),
    		weatherCalls: weather.calls
    	};
    },
Build UIs?Every error boundary you have written contains a failure. The header is where you decide what it shows.

Where it already is in your components

An error boundary is containment for rendering. React’s error boundaries and Svelte’s <svelte:boundary> put a wall around part of the tree: an error inside replaces that part with a fallback and leaves the rest of the page alone. A Suspense boundary does the same for waiting: the part that is not ready shows a placeholder, and the rest renders. You have been drawing blast radii every time you chose where to put one.

When you have to own it

A boundary decides what the reader sees; it does not decide how long the server waits. The snow report’s header needs both. The layout load returns the report as a promise it does not await, with a 300 ms deadline and the last known report inside it, and the header renders three states: loading, fresh, and stale. The stale state says so, because a skier deciding whether to drive up should know the depth might be yesterday’s.

A header with a boundary around the snow report. If the report throws while rendering, only the report is replaced.

ReactAlready in your code
Header.tsx
import { Component, Suspense, type ReactNode } from 'react';

// React draws the boundary for you: an error inside it replaces only what it wraps.
class Boundary extends Component<
	{ fallback: ReactNode; children: ReactNode },
	{ failed: boolean }
> {
	state = { failed: false };
	static getDerivedStateFromError() {
		return { failed: true };
	}
	render() {
		return this.state.failed ? this.props.fallback : this.props.children;
	}
}

export default function Header({ snow }: { snow: ReactNode }) {
	return (
		<header>
			<strong>Pine Ridge</strong>
			<Boundary fallback={<span>Snow report unavailable</span>}>
				<Suspense fallback={<span>Loading snow report…</span>}>{snow}</Suspense>
			</Boundary>
		</header>
	);
}

10 / Make the call

Contain what is optional. Fail fast on what is not.

Containment is for dependencies a page can live without: a widget, a count, a recommendation, a snow report. For the page’s own data, the lift-ticket prices, there is nothing to degrade to. A deadline still helps there, so the page fails in a second instead of hanging for a minute, but a stale price is not an answer.

Keep it simple when there is one user and no shared capacity, such as a script or an internal tool that only you run, or when the dependency is part of the same process and fails with it. Revisit the limits when traffic grows, when the vendor changes, or when a second page starts calling the same API.

Take it with you

Explain it without saying “circuit breaker” or “bulkhead”: “The snow report can make a page wait 300 milliseconds and no longer. If it keeps failing, we stop asking for a few seconds at a time, and never let it hold more than four requests. Pages show yesterday’s report instead.” Then open your own layout or shared header and find everything it awaits.

Paste into your next prompt, and fill in the blanks

<Dependency> is optional for <pages>. When it is slow, failing, or
unreachable, those pages still render within <budget> and show <fallback>.
- Wait at most <deadline> for it.
- Run at most <n> calls to it at once; beyond that, use the fallback without calling.
- After <k> failures in a row, stop calling it for <cooldown>, then let one
  call through to find out whether it has recovered.
- A slow answer counts as a failure once it passes the deadline.
Add a test where <dependency> never answers, and assert that the page still
renders within <budget>.
Connections to follow nextRelated lessons

Take the site into your editor. Add a second optional widget to the header, a lift-queue camera, and decide whether it shares the snow report’s bulkhead or gets its own.

Back to architecture →