← Concepts & practices
Concept Errors, results, and recovery

Retry, backoff & idempotency

Try again later, not all at once.

You already retry things: a fetch that failed, a socket that dropped, a save that timed out. One retry after a blip is a good idea. Let’s follow a webhook forwarder from that one good idea to twelve events retrying against the same outage, and find the rules that make retries help instead of pile on.

TypeScriptGo One forwarder, two implementations.

01 / The idea

Retrying a blip is a fair start.

You’re building a webhook forwarder: when something happens in your app, it posts an event to a customer’s URL. Networks blip, so when a post fails with a 503 you try again straight away, up to eight times. For one event and a brief failure, that loop recovers before anyone notices.

Read the first retry loopTypeScript · the version this lesson starts from
forwarder.ts
// The first version: one event, a brief blip, so try again straight away.
export async function forwardWithRetry(send: () => Promise<number>, attempts = 8): Promise<number> {
	let status = 0;
	for (let attempt = 1; attempt <= attempts; attempt++) {
		status = await send();
		// Stop on anything that isn't overload or an outage.
		if (status < 500 && status !== 429) return status;
	}
	return status;
}

Go’s version is the same loop. Both languages meet again at retryDelay in section 02.

Then the destination goes down while twelve events are in flight. Every event follows the same helpful rule against the same failing service: fail, try again, fail again. Twelve loops that each look careful add up to a steady stream of requests at a service that’s trying to come back. None of them stops because retrying is pointless, only because it ran out of tries.

Backoff means waiting longer after each failure. Jitter means picking a random moment inside that wait, so callers that failed together don’t come back together. And a retry is only safe when the receiver can tell a repeat from new work: that’s idempotency. Marc Brooker’s post on the AWS Architecture Blog is blunt about the second part: jittered backoff “should be considered a standard approach for remote clients.”

If you’ve used TanStack Query, you’ve already shipped this. A failed query retries three times in the browser, waiting Math.min(1000 * 2 ** failureCount, 30000) milliseconds: one second, two, four, and never more than thirty. Section 05 writes that default out, then builds a reconnect you have to own.

02 / See the shape

Wait longer, spread out, and know when to stop.

The basic form is the spacing rule: how long to wait before the next try. In the wild adds the decision to try at all, with a reason for every stop. At the call site runs one outage through all three policies.

Both languages run the same model on a logical clock, with no real network or sleeps, and replay the same 18 shared schedules.

The spacing rule. Immediate retries wait one tick, backoff doubles up to a cap, and jitter picks a random point inside that window.

TypeScriptReading
forwarder.ts
// When to try again: nothing between tries, a doubling wait, or a random point inside it.
export function retryDelay(
	policy: Policy,
	failedAttempts: number,
	random: number
): { delay: number; random: number } {
	const window = Math.min(8, 2 ** Math.min(failedAttempts, 3)); // 2, 4, 8, 8… ticks.
	if (policy === 'immediate') return { delay: 1, random }; // Next tick: no deliberate backoff.
	if (policy === 'backoff') return { delay: window, random };
	const next = (Math.imul(random, 1664525) + 1013904223) >>> 0;
	const unit = next / 4294967296;
	return { delay: Math.max(1, Math.ceil(unit * window)), random: next };
	// Full jitter rounded up to our one-tick clock. The first tick is the minimum.
}
GoAlongside
forwarder.go
// When to try again: nothing between tries, a doubling wait, or a random point inside it.
func retryDelay(policy string, failedAttempts int, random uint32) (int, uint32) {
	window := 1 << min(failedAttempts, 3) // 2, 4, 8, 8… ticks.
	if policy == "immediate" {
		return 1, random
	} // Next tick: no deliberate backoff.
	if policy == "backoff" {
		return window, random
	}
	next := random*1664525 + 1013904223 // uint32 wrapping is intentional.
	unit := float64(next) / 4294967296.0
	return max(1, int(math.Ceil(unit*float64(window)))), next
	// Full jitter rounded up to the model's one-tick clock.
}
Reading the TypeScriptA repeatable random stream

retryDelay returns the wait and the next random state, so every event carries its own repeatable stream. Math.imul multiplies as 32-bit integers, and >>> 0 turns the result back into an unsigned number. The seed keeps the film and the tests repeatable; production code can use Math.random().

Unless it stops the event, afterFailure changes only when it’s next due and its random state. The key and the deadline stay as they were, so a retry is the same operation again.

Reading the Gouint32 wraps on purpose

random*1664525 + 1013904223 overflows a uint32, and the Go spec defines unsigned arithmetic as wrapping, which matches TypeScript’s >>> 0.

1 << min(failedAttempts, 3) gives windows of 2, 4, 8, then 8 ticks. min and max have been built in since Go 1.21.

03 / Follow the outage

Watch twelve events meet one outage.

Five steps, every bar from running the forwarder you just read. Each column is one tick: first attempts at the bottom, retries stacked on top, and a strip underneath for whether the destination was up. Before each step, guess how many attempts it takes to deliver all twelve.

In Try it, break and restore the destination yourself, then switch policies on the same history.

Retry and backoff

One outage, three retry policies.

Retry immediately1 event

One event. Down for 1 tick, then healthy

attempts
2
delivered
1
stopped
0
most retries in a tick
1

Retry immediately, 1 event. One event. Down for 1 tick, then healthy. Attempts per tick: t0: 1, t1: 1. 2 attempts, 1 delivered, 0 stopped, peak 1 retries in a tick. One failure, one retry, delivered. For a single event and a brief blip, retrying straight away is fine.

01/ 05
Retry one event after a blip

A blip, one retry.

One event hits a 503, tries again next tick, and gets through.

Reduced motion: choose a scene to see its completed state.

Read this scene

One event hits a 503, tries again next tick, and gets through.

Retry immediately, 1 event. One event. Down for 1 tick, then healthy. Attempts per tick: t0: 1, t1: 1. 2 attempts, 1 delivered, 0 stopped, peak 1 retries in a tick. One failure, one retry, delivered. For a single event and a brief blip, retrying straight away is fine.

Watch restarts when you return. Step through keeps your selected step. Try it starts a fresh batch of twelve events each time you open it.

What backoff and jitter buy you

Now put names on what you just watched. These are the words you’ll hear in a design review, and each one points at something on this page.

Fewer wasted attempts
The same outage costs 60 attempts retrying immediately, 48 with backoff, and 39 with jitter. All twelve events are delivered every time.
Retries that don’t arrive together
Without jitter, all twelve retry at t2 and again at t6. With it, no tick sees more than 9 retries.
Bounded work
An attempt limit and a deadline end every event. In step 5 the destination never comes back, and the forwarder still stops.
Stops that explain themselves
Each stop names its reason: a request that can’t change, a repeat that isn’t safe, the attempt limit, or the deadline.
Safe to repeat, or not repeated
Every try carries the event’s same key, and the forwarder won’t retry at all unless the receiver can recognize a repeat.

The review words are exponential backoff, jitter, and the thundering herd they break up. The promise that a repeat is harmless is idempotency. Section 08 covers what they cost.

04 / Try a decision

Two careful retry loops multiply.

A worker forwards each event with its own loop of up to three attempts. Later the team adopts an HTTP client that retries 503s by itself, also up to three attempts. Both loops are in layered.ts, and the lesson’s tests count what they send.

The destination stays down. How many requests does one event send?

The worker retries up to 3 attempts. Months later it switches to an HTTP client that also retries 503s, up to 3 attempts. Neither loop changed; both are reasonable on their own.

05 / Give it a real job

The worker owns the retries. The receiver makes them safe.

In a real forwarder, events wait in a durable queue. A worker takes one, sends it, and decides from the reply whether to retry, when, and when to give up. The customer’s endpoint does the other half: it recognizes a repeat by the event’s ID and doesn’t do the work twice.

Worker

Owns the budget

One place decides the wait, the attempt limit, and the deadline.

Destination

Says when to wait

A 429 or 503 can carry Retry-After, and that wait comes first.

Receiver

Makes repeats safe

It records the event IDs it has handled and skips a repeat.

An event that stops still needs somewhere to go. A real forwarder records it as failed, with its reason, so someone can fix the endpoint and send it again with the same key. Retry-After can be a number of seconds or a date, and the forwarder treats it as the shortest wait it’s allowed.

The example leaves out the queue, real HTTP, parsing Retry-After, and storing failed events. None of those change who owns the decision to try again.

Build UIs?The data-fetching library you use already retries for you, and one day a live connection makes you write the backoff yourself.

Where it already is in your components

TanStack Query retries a failed query three times in the browser, doubling the wait from one second up to a thirty-second cap, and pauses those retries while the device is offline. You rarely notice, except as a spinner that lasts a few seconds longer before the error.

The textbook panes write those defaults out with one change: they only retry failures that might pass next time. A 404 won’t, so it shows straight away. Put a retry loop inside queryFn as well and you’re back in section 04, with the attempts multiplying.

When you have to own it

Now it’s a live notifications badge over a WebSocket. When the server restarts, every open tab loses its connection in the same second. If they all reconnect one second later, they arrive together: the crowd from step 3, made of your users’ tabs. So each tab draws a random wait inside a doubling window, capped at thirty seconds.

A connection that opens resets the count. After eight failures in a row the badge stops and offers Try again instead of retrying forever. While the browser reports it’s offline, waiting can’t help, so the badge waits for the online event instead. MDN calls navigator.onLine “inherently unreliable”, so the badge uses it only to pause, never to give up.

reconnect.ts
// One place decides how a live connection comes back. Components only call these.

export const maxFailures = 8;

// Full jitter: a random wait between zero and a doubling window, capped at 30 seconds.
// Every open tab draws its own wait, so they don't all reconnect in the same second.
export function reconnectDelay(failures: number, random: () => number = Math.random): number {
	const window = Math.min(30_000, 1_000 * 2 ** failures);
	return Math.round(random() * window);
}

export type NextStep = 'wait-for-network' | 'retry' | 'give-up';

// Waiting can't fix a missing network, and trying forever isn't a plan either.
export function nextStep(failures: number, online: boolean): NextStep {
	if (!online) return 'wait-for-network';
	return failures >= maxFailures ? 'give-up' : 'retry';
}

An order list whose query retries three times with doubling waits capped at thirty seconds, and only for failures that might pass next time. TanStack Query in React and Svelte.

ReactAlready in your code
OrderList.tsx
import { useQuery } from '@tanstack/react-query';

type Order = { id: string; total: number };

class HttpError extends Error {
	status: number;
	constructor(status: number) {
		super(`HTTP ${status}`);
		this.status = status;
	}
}

export function OrderList() {
	const orders = useQuery({
		queryKey: ['orders'],
		queryFn: async (): Promise<Order[]> => {
			const response = await fetch('/api/orders');
			if (!response.ok) throw new HttpError(response.status);
			return response.json();
		},
		// Retry what might pass next time: no response, overload, or an outage. Not a 404.
		retry: (failureCount, error) => {
			const mightChange =
				!(error instanceof HttpError) || error.status === 429 || error.status >= 500;
			return mightChange && failureCount < 3;
		},
		// TanStack Query's default wait, written out: 1s, 2s, 4s… capped at 30s.
		retryDelay: (failureCount) => Math.min(1000 * 2 ** failureCount, 30_000)
	});

	if (orders.isPending) return <p>Loading orders…</p>;
	if (orders.isError) return <p role="alert">Couldn’t load orders: {orders.error.message}</p>;
	return (
		<ul>
			{orders.data.map((order) => (
				<li key={order.id}>{order.total}</li>
			))}
		</ul>
	);
}

06 / Recognize it elsewhere

Anywhere something tries again on your behalf.

You’ve met all of these. For each one, find who decides the wait and what makes a repeat safe.

Familiar retries, what repeats, and how each one waits
Where you’ve seen itWhat tries againHow it waits, and what keeps it safe
TanStack QueryA failed queryThree retries in the browser, doubling from 1s to a 30s cap. Reads are safe to repeat.
Stripe webhooksAn event your endpoint didn’t acceptExponential backoff for up to three days in live mode. Your endpoint skips event IDs it has already handled.
Stripe’s APIA POST whose reply never cameYou retry with the same Idempotency-Key, and Stripe replays the first result.
A chat or notifications socketA dropped connectionYour code: a jittered wait that resets once connected.

One caller retrying once after a blip doesn’t need a policy. It becomes one when many callers can fail at the same moment.

07 / Already in your toolbox

Your tools already retry this way.

Three places to look. For each one, find the wait, the limit, and what makes a repeat safe.

AWS Architecture Blog · Exponential Backoff and Jitter

Marc Brooker’s 2015 simulation behind “full jitter”: sleep = random(0, min(cap, base * 2 ** attempt)). Its numbers come from a different workload, but the clusters it shows are step 3 of this lesson.

Read the post ↗

TanStack Query · retryer.ts

The retry loop behind every query: a default of three retries in the browser and none on the server, the doubling delay capped at 30 seconds, and a pause while offline.

Read the retryer source ↗

Stripe · Idempotent requests

Stripe saves the status and body of the first request for a key, “regardless of whether it succeeds or fails”, and returns it for retries. Keys can be removed after 24 hours, and reusing one with different parameters is an error.

Read the reference ↗
A useful counterexample: a Pay buttonWhen not to retry automatically

A customer taps Pay and the request times out. The charge may already have gone through. Retrying without a key could charge them twice, and a spinner that quietly retries hides the question. Show that the outcome is unknown and check the order first.

With an idempotency key created once for that payment, a retry with the same key is safe. That’s exactly the job Stripe’s keys do.

08 / The parts to watch

Every retry is more work somewhere.

These are the places it still goes wrong.

Retries stack across layers

A client library, an SDK, a proxy, and your own loop may each retry. Their limits multiply, as section 04 showed. Decide which layer owns retries and give the others a single attempt.

A timeout doesn’t mean it didn’t happen

The request may have reached the server and done its work before the reply was lost. Only retry what the receiver can recognize as a repeat. The forwarder refuses to retry when repeatSafe is false.

Some failures won’t change

A payload the destination rejects as invalid will be rejected again. The forwarder stops those after one attempt. Which replies are final is the API’s contract, not a rule that every 4xx is.

The deadline doesn’t restart

The deadline is fixed when the event is created. Resetting it after each failure would let an event retry forever, one fresh deadline at a time.

Per-event limits don’t cap the total

Eight attempts per event is still eight thousand attempts for a thousand events. A retry budget caps retries across every request to a destination, and a circuit breaker stops sending for a while when failures keep coming. This model has neither.

Random waits make tests flaky

Pass the randomness in, like the forwarder’s seeded stream, so a test can replay one exact schedule.

09 / Make the call

What would you have to change tomorrow?

Give both designs a plausible change and follow the work it creates.

How a change affects an immediate retry loop and a policy with backoff, jitter, and limits
The changeRetry immediatelyBackoff, jitter, and limits
One request, one brief blipRetries next tick and gets through first.Also recovers, after a random wait of one or two ticks.
The destination goes down with twelve events in flight60 attempts, with all twelve retrying in the same tick.39 attempts, and never more than 9 retries in a tick.
It stays downEight attempts back to back, then an error with no reason.Each event stops by its deadline and records why.
The receiver can’t recognize repeatsRetries anyway, so a lost reply can deliver an event twice.Refuses to repeat; the event stops as unsafe.
Someone adds a retrying HTTP clientAttempts multiply.Attempts multiply here too. Only choosing one owner fixes it.

Reach for backoff with jitter when many callers can fail at the same moment, and put a limit and a deadline on every retry. Twelve events meeting one outage is the moment.

Keep the simple loop for one caller retrying one cheap, safe request. A command-line tool fetching a file once doesn’t make a crowd.

The question I’d leave beside the code is: who else is retrying this right now, and is it safe to send twice?

10 / Take the idea with you

Explain the forwarder without saying “backoff.”

“When a send fails, wait before trying again, a bit longer each time and at a random moment so everyone doesn’t come back together. Stop when it can’t work or time’s up, and only repeat what the receiver can recognize as a repeat.” In a review, the words are exponential backoff, jitter, retry budget, and idempotency key.

Before moving on, jot down why twelve careful loops flooded the destination, why two retry layers sent nine requests, and one place in your own code that retries, along with whoever else might be retrying the same request.

Connections to follow nextRelated lessons

Take the forwarder into your editor. Add a retry budget shared by all twelve events, say four retries a tick, and see which policy delivers everything first.

Back to Concepts & practices →