← Architecture
Events, queues, and workflows Undo across companies

Sagas and compensation

Write down what you did, so you can undo it.

Every trip-booking site you have used holds a room, charges your card, and books a car, at three companies that know nothing of each other. When the car is gone, somebody has to give the money back and let the room go. Let’s watch who remembers to.

TypeScriptGoOne weekend, two designs, a recovery pass, two recorded builds.

01 / The prompt

“Book the weekend: hold the hotel, charge the card, book the car, and undo it all if anything fails.”

A small travel site sells one package: two nights at a partner hotel and a rental car, for $330. Mei books it. The obvious handler calls the hotel, the card processor, and the car firm in turn, pushes an undo onto a list after each, and runs the list backward in a catch block. When the car firm says “sold out”, it refunds the card and releases the room. Every test passes.

Then a deploy restarts the server right after the charge. A catch block cannot run in a process that is gone, and the list of undos was in its memory. In the lesson’s example the site ends with no record of the trip, while the hotel holds a room (T-1 held) and the card shows T-1 +330. Nobody is going to give that money back, because nothing remembers taking it.

The prompt never asked the question: when the process dies, or a company does not answer, halfway through, what does the site remember about what it already did, and who undoes it?

The idea has a name and a paper. Hector Garcia-Molina and Kenneth Salem called a long transaction a saga “if it can be written as a sequence of transactions that can be interleaved with other transactions”, where the system guarantees “that either all the transactions in a saga are successfully completed or compensating transactions are run to amend a partial execution” (Sagas, SIGMOD 1987). Here the transactions are other companies’ API calls, and the compensations are new requests to them.

02 / Name the shape

Write each step down before you take it. Undo from what you wrote, never from memory.

A saga is a business action made of steps that each commit on their own, with a compensation for every step: a new action that cancels its effect. The steps and their outcomes go in a log in your own database, so the undo list survives the process. A recovery pass reads the log on startup and on a timer, and finishes every saga that is not finished, forward or backward.

Log the step, call it, log what came back. If a step fails, undo every step that may have happened, newest first, by its key. A timeout is unknown, never failed: send the step again with the same key to learn what happened, and undo it only if it did not happen or you still cannot tell. An undo that fails is retried until it succeeds.

Consistency without a shared transaction made the same move inside one shop: record the action first and let a recovery pass finish it. A saga is what that becomes when the steps belong to other companies and each undo is a request with its own rules and its own failures. Who owns what:

Who owns each part of a booked weekend
PartOwnerUndone by
The trip and its step logOur databaseNothing: it is the record of what to undo
The room holdThe hotelA release, keyed by the trip
The $330 chargeThe card processorA refund: a second line on Mei’s statement
The car bookingThe car firmA cancel by key, safe even if no booking exists
Finishing unfinished tripsThe recovery passRuns on startup and on a timer

Words to put in a prompt or a review

Saga
A business action made of steps that each commit alone, with an undo for each.
Compensation
The undo: a new action, such as a refund, that cancels a step’s effect.
Step log
What the saga started and what came back, in its own database. The undo list.
Unknown outcome
A call with no answer. It may have happened, so it is sent again with the same key to find out, and undone if the company still cannot say it was done.
Recovery pass
The loop that finishes every saga the log says is unfinished, after a crash too.
Orchestrator
One component that runs the steps and owns the log, as in this lesson.
Orchestration or choreography?Who holds the log

This lesson runs one orchestrator: the travel site calls each company and keeps the log. In choreography there is no coordinator; each service reacts to the previous one’s event and publishes its own, and each owns its part of the undo. The log is then spread across every service’s events, and “which trips are half done?” becomes a question you answer by joining them. Choreography suits steps owned by separate teams that already publish events; an orchestrator suits one team that needs to see, and finish, every booking in one place. The example does not model choreography.

Which step goes last?Order is a decision

Put the step you cannot undo, or can undo only at a cost, last, after every step that is likely to fail. If the car firm charged a fee to cancel, you would book it after the charge had succeeded, so a declined card never costs a cancellation. The example’s three undos are free, so its order is the one travelers read on the page: room, payment, car.

03 / Half a trip

Same booking, same fault. Which site leaves something behind?

Each column runs one of the lesson’s travel sites on the same steps, and checks every trip against the hotel, the card, and the car firm: the whole weekend, or none of it. Watch four situations, then open Try it and break them yourself.

Sagas and compensation

A weekend booked at three companies

Undo in a catch block

Reply to Mei: failed

Our database

  • trip T-1 failed

Hotel

  • T-1 released

Card

  • T-1 +330
  • T-1 -330

Car firm

  • no booking

Saga with a log

Reply to Mei: canceled

Our database

  • trip T-1 canceled
  • log charge started
  • log charge done
  • log car started
  • log car failed
  • log charge compensated
  • log hotel compensated

Hotel

  • T-1 released

Card

  • T-1 +330
  • T-1 -330

Car firm

  • no booking
01/ 04
The car is sold out

Mei books T-1; the car firm is sold out.

Mei books the weekend. The hotel holds the room, the card is charged $330, and the car firm has nothing left. Both sites undo: a refund line on the card and the room released. The first version works when a step says no.

Reduced motion: choose a scene to see its completed state.

Read this scene

Mei books the weekend. The hotel holds the room, the card is charged $330, and the car firm has nothing left. Both sites undo: a refund line on the card and the room released. The first version works when a step says no.

Undo in a catch block: trip T-1 failed; hotel T-1 released; card T-1 +330, T-1 -330; car no booking.

Saga with a log: trip T-1 canceled; hotel T-1 released; card T-1 +330, T-1 -330; car no booking.

Watch restarts the story when you come back. Step through shows where each chapter ends. Try it starts a new site whenever you change the design or press Reset.

04 / Read the shape

Steps with their undos, and two passes that read only the log.

Basic form is the list of steps. In the wild is the forward and backward pass. At the call site shows who calls each. Notice that the backward pass decides what to undo from the log, not from anything held in a variable.

The steps: each names the call that does it and the call that undoes it. Every call carries the trip id as its key, so a repeat is safe, and so is undoing a step that never happened.

TypeScriptReading
booking.ts
/**
 * Each step names the call that does it and the call that undoes it. Every
 * call carries the trip id as its key, so sending one twice is safe, and so
 * is undoing one that never happened.
 */
export const steps: Step[] = [
	{
		name: 'hotel',
		run: (p, trip) => (p.hotel.hold(trip) ? 'done' : 'failed'),
		undo: (p, trip) => p.hotel.release(trip)
	},
	{
		name: 'charge',
		run: (p, trip) => (p.payments.charge(trip, PRICE) === 'charged' ? 'done' : 'failed'),
		undo: (p, trip) => p.payments.refund(trip)
	},
	{
		name: 'car',
		run: (p, trip) => {
			const answer = p.cars.book(trip);
			return answer === 'booked' ? 'done' : answer === 'timeout' ? 'unknown' : 'failed';
		},
		undo: (p, trip) => p.cars.cancel(trip)
	}
];
GoAlongside
booking.go
// Steps names, for each step, the call that does it and the call that undoes
// it. Every call carries the trip id as its key, so sending one twice is safe,
// and so is undoing one that never happened.
var Steps = []Step{
	{
		Name: "hotel",
		Run: func(p *Providers, trip string) string {
			if p.Hotel.Hold(trip) {
				return "done"
			}
			return "failed"
		},
		Undo: func(p *Providers, trip string) bool { return p.Hotel.Release(trip) },
	},
	{
		Name: "charge",
		Run: func(p *Providers, trip string) string {
			if p.Payments.Charge(trip, Price) == "charged" {
				return "done"
			}
			return "failed"
		},
		Undo: func(p *Providers, trip string) bool { return p.Payments.Refund(trip) },
	},
	{
		Name: "car",
		Run: func(p *Providers, trip string) string {
			switch p.Cars.Book(trip) {
			case "booked":
				return "done"
			case "timeout":
				return "unknown"
			}
			return "failed"
		},
		Undo: func(p *Providers, trip string) bool { return p.Cars.Cancel(trip) },
	},
}
The catch-block handlerWhat the prompt gets back

It is correct for every failure a company reports. It has nothing for a failure the process itself has, a timeout it cannot interpret, or an undo that fails.

booking.ts
/** The obvious handler: undo in a catch block, from a list kept in memory. */
export class TryCatchSite extends TravelSite {
	protected bookTrip(request: BookRequest): Reply {
		const { hotel, payments, cars } = this.providers;
		const key = request.trip;
		const undo: (() => boolean)[] = [];
		try {
			hotel.hold(key);
			undo.push(() => hotel.release(key));
			this.crashAt(request, 'after-hotel');
			if (payments.charge(key, PRICE) !== 'charged') throw new Failed('card declined');
			undo.push(() => payments.refund(key));
			this.crashAt(request, 'after-charge');
			if (cars.book(key) !== 'booked') throw new Failed('no car');
			this.crashAt(request, 'after-car');
			this.db.trips.push({ id: key, status: 'confirmed' });
			return 'confirmed';
		} catch (error) {
			if (!(error instanceof Failed)) throw error; // a dead process runs no catch block
			for (const step of undo.reverse()) step(); // a refund that fails is logged and forgotten
			this.db.trips.push({ id: key, status: 'failed' });
			return 'failed';
		}
	}
}
The recovery passOn startup and on a timer

A trip left pending resumes at its first step that is not done: a step logged as started or unknown is sent again with the same key, and a step logged as failed starts the undo. A trip left canceling continues its undo.

booking.ts
/**
 * On startup and on a timer: finish every trip the log says is not finished.
 * A step that failed is undone; one only started, or unknown, is sent again.
 */
recover(): void {
	for (const { id, status } of structuredClone(this.db.trips)) {
		if (status === 'canceling') this.compensate(id);
		if (status !== 'pending') continue;
		const next = steps.findIndex((s) => this.last(id, s.name) !== 'done');
		const event = next < 0 ? 'done' : this.last(id, steps[next].name);
		if (event === 'failed') this.compensate(id);
		else this.forward({ trip: id }, next < 0 ? steps.length : next);
	}
}
The behavior these examples promiseChecked by 18 shared scenarios
  • Both designs undo correctly when a company says no: a sold-out car or a declined card leaves nothing held and nothing charged.
  • The catch block leaves half a trip after a crash between steps, after a timeout from a car firm that did book, and after a refund that fails once.
  • The saga finishes a crashed booking from its log, asks a car firm that timed out again by key and keeps the car it had booked, cancels by key when the firm stays silent, and keeps an undo that failed until it succeeds; while it waits, the trip is canceling, not finished.
  • A refund is a second line on the statement, never a deleted charge.

Every expectation was generated by a separate model written from the contract in the examples’ README, not copied from either implementation, and it is kept beside the examples.

Reading the TypeScriptA crash that skips the catch

A crash is a thrown Crash. The catch-block handler rethrows it before its undo loop, because a killed process runs no catch; book reports it as crashed. The database is a plain object that outlives the call, which is what a real database does for a real process.

Reading the GoThe same crash, as an error

A crash is errCrash, returned by the step where the process dies; the catch-block handler returns crashed before its undo loop can run. The providers keep their records in small insertion-ordered slices, so the output is the same on every run.

Run it yourselfNo dependencies

Save the complete files at the paths in their banners. Then run node --experimental-strip-types run.ts (Node 22.18 or later), or go run . in the Go folder. Both print:

try-catch · the car is sold out: reply failed, trip [T-1 failed], card [T-1 +330 T-1 -330], hotel [T-1 released], car [] -> all or nothing
saga · the car is sold out: reply canceled, trip [T-1 canceled], card [T-1 +330 T-1 -330], hotel [T-1 released], car [] -> all or nothing
try-catch · a crash after the charge: reply crashed, trip [], card [T-1 +330], hotel [T-1 held], car [] -> hotel-held:T-1, charged:T-1
saga · a crash after the charge: reply crashed, trip [T-1 confirmed], card [T-1 +330], hotel [T-1 held], car [T-1 booked] -> all or nothing
try-catch · the car firm times out after booking: reply failed, trip [T-1 failed], card [T-1 +330 T-1 -330], hotel [T-1 released], car [T-1 booked] -> car-booked:T-1
saga · the car firm times out after booking: reply confirmed, trip [T-1 confirmed], card [T-1 +330], hotel [T-1 held], car [T-1 booked] -> all or nothing
try-catch · refunds down, then back: reply failed, trip [T-1 failed], card [T-1 +330], hotel [T-1 released], car [] -> charged:T-1
saga · refunds down, then back: reply canceling, trip [T-1 canceled], card [T-1 +330 T-1 -330], hotel [T-1 released], car [] -> all or nothing

05 / Review the agent’s diff

“No answer in five seconds counts as sold out.”

Travelers complained about a spinner, and the agent found the car firm hanging. Read what its fix writes in the log.

The agent’s pull request

“Travelers waited up to a minute when the car firm hung. I added a 5-second timeout and treat it like sold out, so the saga compensates and the traveler gets an answer fast. All saga tests pass.”

// steps.ts
			{
			  name: 'car',
			(removed)   run: async (trip) => {
			(removed)     const answer = await cars.book(trip, { key: `${trip}:car` });
			(removed)     return answer.ok ? 'done' : answer.timedOut ? 'unknown' : 'failed';
			(removed)   },
			(added)   run: async (trip) => {
			(added)     const answer = await cars.book(trip, { key: `${trip}:car`, timeoutMs: 5000 });
			(added)     return answer.ok ? 'done' : 'failed'; // no answer in 5 s counts as sold out
			(added)   },
			  undo: (trip) => cars.cancelByKey(`${trip}:car`),
			},
			
The saga tests use a fake car firm that answers at once. What do you do with this change?

06 / How it fails

Every failure is a step taken and not written down, or an undo not taken.

Each row except the last two is a shared scenario the tests run.

Failure modes of booking a weekend at three companies
What happensWhat Mei seesUndo in a catch blockSaga with a log
A company says no (sold out, declined)“Could not book”, and a refund lineUndone correctly.Undone correctly.
The process dies between stepsAn error page, then a chargeHalf a trip and no record of it.Recovery resends the step by key and finishes forward.
The car firm books, and the reply is lost“Could not book”Refunded, with a car still booked.Logged as unknown, sent again by key, and the firm answers “booked”: confirmed.
The car firm times out without booking, and again when asked“Could not book”Undone correctly, by luck.Still unknown after the resend, so undone; the cancel by key finds nothing and succeeds.
A refund fails while undoing“Could not book”, and no refundLogged once and forgotten; the money is kept.The trip stays canceling; the next pass refunds.
The process dies while undoingAn error pageThe rest of the undo list is gone.Recovery continues the undo from the log. Not modeled in the example.
A company is slow, not downA long spinnerWaits as long as the client does.Waits until its deadline, then treats the step as unknown. Not modeled in the example.

The pieces underneath have their own lessons: Retry, backoff & idempotency for sending a step again safely, and Timeouts, deadlines & races for why a call with no answer is neither yes nor no.

07 / Is it worth it?

A log, a recovery pass, and an undo for every step, against a catch block.

The shared four changes, against the catch block
ChangeUndo in a catch blockSaga with a log
A second entry point: a partner sells the package through our APIA second handler with its own undo list, or a shared one with the same gaps.It starts the same saga; recovery finishes both.
Replacing a company: a new car firmThe call and its undo change in place.One step’s call and undo change. About the same.
A new rule: travel insurance as a fourth stepAnother undo in the list, and another crash point nobody handles.Another step with its undo; the log and recovery cover it.
A second team: payments owns refundsRefund failures stay invisible to them.Trips stuck canceling on a refund are a list they can own.

The costs are real: a log table, a recovery pass to run and watch, an undo to write and test for every step, and a status the traveler can see (“canceling”) that a catch block never needed. When every step lives in one database, a transaction is simpler and better. When only one call leaves it, Consistency’s record-and-recover is enough.

Measure before you change anything, with the catch block still running:

  • Half trips per week: a daily reconciliation of our trips against each company’s records, counting holds, charges, and cars with no confirmed trip. This is the baseline, and the number the saga should take to zero.
  • Trips not finished, and the age of the oldest: the saga’s new number, alerted on when it passes a few recovery passes.
  • Time until Mei gets an answer, at the 95th percentile, before and after, so a deadline on slow companies is a measured choice.

This lesson did not measure a real travel site, and gives no numbers.

08 / Ask for it

One brief, two prompts.

Two agents running Claude Sonnet each got the brief from section 01, with the three companies’ APIs and one product sentence: “A traveler gets the whole package or none of it: never charged for a package they do not have, and never left holding a room or a car they did not get.” One prompt added a Saga block: log each step before and after the call, key every call, undo newest first, treat a timeout as unknown and undo it, retry a failed undo until it succeeds, and run a recovery pass on startup and on a timer. A script played the three companies for both builds and killed the servers along the way.

What the checker found, run 2026-09-23
QuestionPlain promptSaga prompt
Five trips, nothing wrong5 of 5 confirmed and whole5 of 5 confirmed and whole
The car is sold outfailed: 0 room, $0 charged, 0 carcanceled: 0 room, $0 charged, 0 car
The card is declinedfailed: 0 room, $0 charged, 0 carcanceled: 0 room, $0 charged, 0 car
The car firm books, and never answersanswered “confirmed” in 3.2 s; confirmed: 1 room, $330 charged, 1 caranswered “canceled” in 4.1 s; canceled: 0 room, $0 charged, 0 car
The card processor charges, and never answersanswered “confirmed” in 3.2 s; confirmed: 1 room, $330 charged, 1 caranswered “canceled” in 4.1 s; canceled: 0 room, $0 charged, 0 car
Refunds down for 30 seconds while undoingWhile down: failed: 0 room, $0 charged, 0 car. After: failed: 0 room, $0 charged, 0 car; 0 refund callsWhile down: pending: 1 room, $330 charged, 0 car (half a trip). After: canceled: 0 room, $0 charged, 0 car; 13 refund calls
Server killed while the car firm holds the booking, then restarted20 s later: pending: 1 room, $0 charged, 1 car (half a trip)20 s later: confirmed: 1 room, $330 charged, 1 car
Server killed while the card processor holds the charge, then restarted20 s later: pending: 1 room, $330 charged, 1 car (half a trip)20 s later: confirmed: 1 room, $330 charged, 1 car
Its own tests15 of 15 pass16 of 16 pass

The plain build was good while it lived. The product sentence did its work: the agent logged the trip, keyed every call, retried a timeout with the same key until the company answered, and unwound what it had secured when a company said no. It also made the move from this lesson’s “Which step goes last?” disclosure on its own: it charges the card last, so a sold-out car or a refund outage never touches money.

lib/app.ts · plain prompt
/**
 * Runs the booking saga for a trip that has already been inserted as
 * "pending". Order: hotel hold, then car booking, then the charge — the
 * traveler is only ever charged once both the room and the car are actually
 * held, and any step that fails unwinds everything secured before it.
 */
async function runBookingSaga(db: DatabaseSync, resolved: ResolvedOptions, tripId: string): Promise<void> {

What it did not have was anything that runs after a restart. Killed mid-booking, it came back and listened, and the trip stayed pending with a room held and a car booked, and in the second case with the card charged as well, for as long as the checker watched. Its whole undo lived in the request that died. The server’s startup is these three lines:

server.ts · plain prompt
app.server.listen(PORT, () => {
  console.log(`trips server listening on port ${PORT}, db at ${TRIPS_DB}`);
});

The saga build ran a recovery pass before it listened, and finished both killed bookings forward, from its log.

server.ts · saga prompt
await runRecoveryPass(db, cfg);
startRecoveryLoop(db, cfg, recoveryIntervalMs);

The saga build paid for the rule it was given. Told that a timeout is unknown and must be undone, it undid every trip whose company went quiet, including the two where the car firm or the card processor had in fact done the work. The traveler got a clean “canceled” after four seconds, where the plain build, resending with the same key, learned the truth and confirmed the trip. Both are whole; only one sold a weekend.

src/saga.ts · saga prompt
  } else if (outcome.outcome === "timeout") {
    setForwardState(db, tripId, step, "timeout");
    // Unknown is not failed, but it is also not "keep going": commit to
    // undoing everything, newest first. This decision is sticky -- even
    // if a later replay of this very step turns out to have succeeded
    // after all, we undo it rather than resuming forward progress.
    setTripDirection(db, tripId, "rollback");
  } else {

The missing lines are two, one per build: run a recovery pass on startup that finishes every trip the log says is unfinished, and before undoing a step that timed out, send it again with the same key to learn what happened, and undo only if it did not happen or cannot be finished. The prompt snippet in section 10 carries both.

How the runs were made and checkedTwo builds, recorded as written
  • Both agents were launched at the same time from empty folders; neither was told about the other, the lesson, or the checker.
  • Both builds are kept byte for byte with checksums. For every question the checker restores a build into a fresh folder with its own database, starts the server, and plays the hotel, the card processor, and the car firm, each honoring idempotency keys.
  • The checker ran twice. The first run covered the plain build alone, with six questions, while the saga agent was still working; the second added two card-processor questions and covered both builds, and is the one this table reads. The plain build’s answers to the first six were the same in both.
  • The saga agent wrote three log files to /tmp during its own check, against the prompt, and deleted them. Neither agent stopped a process by name or pattern.
  • One run of each prompt is a sample, not a measurement of the model.

09 / Hold it there

A saga breaks when someone calls a company outside it. Three checks notice.

  1. The runtime’s door: no cleanup runs on a kill

    Node’s documentation is plain about it: “'SIGKILL' cannot have a listener installed, it will unconditionally terminate Node.js on all platforms” (Node.js, signal events). No finally, exit handler, or shutdown hook can be where an undo lives. Whatever must be undone has to be written down before the call that makes it necessary.

  2. Only the saga imports the companies’ clients

    The day a handler calls the car firm directly, there is a step with no log entry and no undo. An import rule that allows the hotel, payments, and car clients in the saga’s module and nowhere else keeps them in, the way Enforcement layer enforces import rules.

  3. Tests that stop the process, and a reconciliation that counts

    The shared scenarios crash after each step, time out a car firm that did book, and take refunds down mid-undo. In production, the daily reconciliation from section 07 runs against the companies’ own records; with a saga it should find nothing but trips still pending or canceling.

There is no frontend version of this lesson: the saga runs on the server, between it and three companies a browser never talks to. What the traveler’s page owns is an honest “pending” and a retry that does not book twice, which Consistency without a shared transaction covers in its frontend row.

10 / Make the call

If you cannot roll it back, write it down before you do it.

Use a saga when one action takes several steps at systems you do not own, and each has an undo that is an action of its own: bookings, charges, shipments, accounts at partners. Keep the catch block when a lost undo costs nothing, say a cache warmed for nobody. Keep one transaction whenever the steps share a database. Reopen the design when steps are owned by separate teams that already publish events; then the choreographed form may fit better, or a workflow engine that keeps the log for you.

Take it with you

Explain it without saying “saga”: “Before we call the hotel, we write ‘calling the hotel’ in our notebook, and after, what it said. If anything goes wrong we read the notebook backward and cancel each thing it says we did. If we are not sure whether something happened, we ask again with the same reference number, and if they still cannot say, we cancel it anyway, because canceling nothing is harmless. If we get interrupted, we pick up the notebook where it left off.” Then find a handler in your own code that calls two outside services in a row, and ask what it knows about the first one if the process dies before the second.

Paste into your next prompt, and fill in the blanks

Booking [a weekend package] calls [the hotel, the card processor, and the car firm], in that order; no transaction covers them.
Write the booking and a log of its steps to our database before calling anyone: each step as started before the call, and what came back after.
Send every call with an idempotency key made from the booking id and the step, and undo by the same key.
If a step fails, undo the steps that happened, newest first.
A timeout is unknown, not failed: send the step again with the same key to learn what happened, and undo it if it did not happen or the booking cannot finish.
An undo that fails is retried until it succeeds, and the booking stays [canceling] until then.
A recovery pass on startup and every [30 seconds] finishes each unfinished booking from the log.
Answer the traveler within [10 seconds]: [confirmed], [canceled], or [pending] with an email when it settles.
Connections to follow nextRelated lessons

Take the saga into your editor. Make the car firm charge a cancellation fee after the booking, move the car step so a declined card never pays it, and add a scenario that proves it.

Back to architecture →