01 / The prompt
“Build me a parcel tracking page.”
A courier has three internal services. Shipments knows the parcel: where from, where to, which service. Scans knows every time a depot scanned it. Estimates predicts when it will arrive. You ask for the page a customer opens from the link in their email, and what comes back works: the parcel, its journey, a delivery window.
Somewhere in that build is a server that calls the three services and hands the page the result. That server is the subject of this lesson. The prompt never said what it should wait for, what it should do when one service is slow or down, or which of the services’ fields a customer should see.
The pattern has a name because teams kept meeting the same problem. Sam Newman describes it as having, instead of a general-purpose API, “one backend per user experience,” a name he credits to Phil Calçado at SoundCloud. On failure he asks: “if only the Inventory service was down, wouldn’t it be better to just degrade the functionality we pass back to the client?” (Pattern: Backends For Frontends, checked 23 September 2026). That question is the one the prompt left open.
02 / Name the shape
The page’s server answers for the page.
A backend for frontend is a server-side layer that belongs to one screen or one client. It calls the services behind it, and it owns the contract with that screen: which request the browser makes, what comes back, how long the screen waits, and what it shows when a part is missing. It does not own the facts. Shipments still decides where the parcel is going; the backend for frontend decides how the page says it.
The services decide what is true. The page’s server decides what the page waits for, what it shows, and what never leaves the server.
| What | Owner | Why |
|---|---|---|
| Where the parcel is going, its service level | Shipments | It is the record. The route copies two fields and never computes them. |
| Each scan and its code | Scans | Depots write them. The route reads them and never edits them. |
“Out for delivery” instead of OFD | Backend for frontend | Words are a screen decision. One map, with a fallback, in one place. |
| How long the page waits for each part | Backend for frontend | Only the screen knows what it can live without. |
| What the page says when a part is missing | Backend for frontend | A named gap, not a blank page and not a guess. |
| Service tokens and internal ids | Backend for frontend | They stay on the server. The reply carries only what the screen shows. |
| Rendering, and asking again | Browser | It makes one request, shows each state the reply can be in, and offers a reload. |
Words to put in a prompt or a review
- Backend for frontend (BFF)
- A server layer owned by one screen or client, shaped for it.
- Fan-out
- One incoming request that becomes several calls to other services.
- Deadline
- How long one call may take before it is cancelled and treated as missing.
- Required part
- Without it there is no page, so its failure is the page’s failure.
- Partial response
- A reply that says what it has, and names what it does not have and why.
- Pass-through
- A route that forwards what the services said, in their shapes.
- View model
- The data a screen renders, in the screen’s words and nothing more.
How many backends for frontends?One per experience, and where it lives
Newman’s rule of thumb, on the same page, is “one experience, one BFF”: if the iOS and Android apps are very similar, one BFF can serve both, and a website with a different job gets its own. In this lesson the app and the website share one route with two views, chosen by a query parameter; when the two screens start to pull the route in different directions, split it.
It does not have to be a separate service. A SvelteKit +page.server.ts load or
a Next.js server component that calls three APIs is a backend for frontend for one page. A GraphQL
layer in front of the services is another way to let a screen ask for its shape; it moves the
shaping into the query and leaves the deadline and missing-part questions where they were.
03 / Follow one page load
Watch the same page wait for two different things.
First the pass-through route: it asks all three services and waits for every answer. Then the backend for frontend: each part has a deadline, and a missing part is named instead of failing the page. The estimate service is the one that misbehaves. Open Try it and choose what each service does today.
What does the page wait for?
Pass-through route · 0 ms
Shipmentswaiting
Scanswaiting
Estimatewaiting
What the browser receives
Waiting for the route…
The pass-through route asks all three and waits for every answer.
Three requests leave at once. The page waits for the route, and the route waits for all three.
Reduced motion: choose a scene to see its completed state.
Read this scene
Three requests leave at once. The page waits for the route, and the route waits for all three.
Pass-through route at 0 ms. Shipments: waiting; Scans: waiting; Estimate: waiting. The browser is still waiting.
Watch restarts the story when you come back. Step through keeps your step. Try it runs the two routes fresh each time you load the page.
04 / Read the shape
A deadline per part, and a reply per screen.
Basic form is one call under one deadline, with an outcome that has a name. In the wild is the route: which part is required, which parts may be missing, and the reply shaped for the screen that asked, beside the pass-through route it replaces. At the call site the browser makes one request and the server’s handler answers it.
Notice what trackingPage awaits first: only Shipments, the part without which there
is no page. The other two calls started at the same moment and are still running under their own
deadlines.
One dependency under one deadline, with an outcome that has a name: a value, or timed-out, failed, or not-found. When the deadline passes the call is cancelled, not left running.
export type Outcome<T> =
{ ok: true; value: T } | { ok: false; reason: 'timed-out' | 'failed' | 'not-found' };
// One dependency, one deadline, and an outcome with a name. When the deadline
// passes, the call is cancelled, not left running.
export async function withDeadline<T>(
clock: Clock,
limit: number,
parent: AbortSignal,
call: (signal: AbortSignal) => Promise<T>
): Promise<Outcome<T>> {
const request = new AbortController();
const stop = () => request.abort(parent.reason);
parent.addEventListener('abort', stop, { once: true });
const timer = new AbortController();
try {
return await Promise.race([
call(request.signal).then((value): Outcome<T> => ({ ok: true, value })),
clock.sleep(limit, timer.signal).then((): Outcome<T> => {
request.abort(new Error('deadline'));
return { ok: false, reason: 'timed-out' };
})
]);
} catch (error) {
return {
ok: false,
reason: error instanceof ServiceError && error.status === 404 ? 'not-found' : 'failed'
};
} finally {
timer.abort();
parent.removeEventListener('abort', stop);
}
} // Outcome is a value, or the name of why there is none.
type Outcome[T any] struct {
Value T
Reason string // "", "timed-out", "failed", or "not-found"
}
// WithDeadline starts one dependency under its own deadline. When the deadline
// passes, the call's context is cancelled, not left running.
func WithDeadline[T any](ctx context.Context, limit time.Duration, call func(context.Context) (T, error)) <-chan Outcome[T] {
done := make(chan Outcome[T], 1)
go func() {
ctx, cancel := context.WithTimeout(ctx, limit)
defer cancel()
value, err := call(ctx)
var failure *ServiceError
switch {
case err == nil:
done <- Outcome[T]{Value: value}
case errors.Is(err, context.DeadlineExceeded):
done <- Outcome[T]{Reason: "timed-out"}
case errors.As(err, &failure) && failure.Status == 404:
done <- Outcome[T]{Reason: "not-found"}
default:
done <- Outcome[T]{Reason: "failed"}
}
}()
return done
} The behavior these examples promiseChecked by 12 shared scenarios, each run through both routes
- All three calls start at once. Shipments and Scans have 800 ms each, the estimate 300
ms. A call past its deadline is cancelled and counts as
timed-out. - No parcel from Shipments, for any reason, means no page: a 404 when Shipments says there
is no such parcel, otherwise 502 with
part: "shipment". The other calls are cancelled at that moment. - A missing Scans or estimate gives a 200 with
status: "partial"and amissingentry naming the part and the reason. The page answers when the last part has answered or run out of time. - The reply has the parcel’s id, route, and service; the newest scan in words; the full history only for the website; the delivery window. No account, staff id, recipient, confidence, or model name.
- A scan code the route does not know reads “Update at <depot> depot”.
- The pass-through route answers with the three payloads as sent, or 502 as soon as any call fails, leaving the others running.
Every expectation in the shared cases, down to the millisecond each route answers and the state each call was in, was produced by a separate model written from these rules and kept beside the examples, not copied from either implementation. The tests run on a clock they control, so the times are exact.
Reading the TypeScriptAbortController, Promise.race, and a clock
withDeadline races the call against a timer and aborts the call’s AbortController when the timer wins, so a well-behaved client such as fetch closes the connection. A second controller for the whole page aborts
every call at once when Shipments fails. The clock is passed in so the tests can run the
same code on virtual time; in production it is setTimeout.
Reading the Gocontext.WithTimeout, and synctest
Each part runs in a goroutine under context.WithTimeout, and returns its
outcome on a buffered channel so it never blocks after the route has stopped listening. defer cancel() on the page’s context is the whole cancellation story. The
tests run inside testing/synctest, where timers use a fake clock, so 1.5
seconds takes no time and elapsed times are exact.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types tracking.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:
module heyrian.dev/lessons/backend-for-frontend
go 1.25
bff app · 200 complete · Out for delivery, Sep 23 07:48 · arriving 10:00–12:00 · 203 bytes bff app, slow estimate · 200 partial · Out for delivery, Sep 23 07:48 · no estimate (timed-out) · 226 bytes pass-through, slow estimate · 200 pass-through · 3 raw payloads · 736 bytes pass-through, estimate down · 502 upstream failed bff web, new scan code · 200 complete · Update at Leeds depot, Sep 23 08:30 · arriving 10:00–12:00 · 7 scans listed · 483 bytes
05 / Review the agent’s diff
“The estimate always loads now.”
Customers seeing “we can’t estimate a delivery time” is a real complaint, and the fix is one number. Read what that number controls before you decide.
06 / How it fails
Three services give the page three ways to be late.
Every row is one of the shared scenarios, run through both routes. The times are the route’s own, from the tests’ virtual clock.
| What goes wrong | What the customer sees | What the backend for frontend does | Pass-through |
|---|---|---|---|
| Slow optional part: the estimate takes 1.5 s | The parcel at 300 ms, and “we can’t estimate a delivery time right now”. | Cancels the estimate call at its deadline and names the gap. | Answers at 1,500 ms, with everything. |
| Slow required part: Shipments takes 1 s | An error at 800 ms that says tracking is unavailable. | Answers 502 at the deadline instead of waiting on. | Answers at 1,000 ms, with everything. |
| Unreachable optional part: the estimate is down | The parcel, with the estimate marked as failed. | Answers at 60 ms, when Scans is in. | 502 at 50 ms. No parcel. Scans is left running. |
| Unreachable required part: Shipments is down | An error page, at once. | 502 at 30 ms, and cancels the other two calls. | 502 at 30 ms, and leaves the other two running. |
| Half there: estimate slow and Scans down | The parcel, with two named gaps. | 200 partial at 300 ms, missing scans (failed) and eta (timed-out). | 502 at 30 ms. |
| Wrong: a scan code the route has never seen | “Update at Leeds depot”. | Falls back to where the parcel is, and keeps the rest of the page. | Forwards HLD; each screen decides what to print. |
| Wrong: no parcel with that number | “No parcel with that number”. | 404 at 25 ms, and cancels the other calls. | 502, as if a service broke. |
| Duplicated: the customer reloads | The same page again. | Nothing to guard: every call is a read. Not modeled in the cases. | The same. |
Two ideas carry the table: a deadline is only half a timeout unless the call is cancelled too, and a partial answer is honest only when it names what is missing. Both have their own lessons: Timeouts, deadlines, and races, Cancellation propagation, and Aggregate and partial failure.
07 / Is it worth it?
You pay for a layer. Here is what it buys.
The pass-through route is shorter and has nothing to configure. Hold both up against the kinds of change every app gets.
| Change | Pass-through | Backend for frontend |
|---|---|---|
| A second client: the website wants the full history | Nothing to change on the server. Each client reshapes 736 bytes itself and carries its own copy of the code-to-words map. | A second view in one place: 438 bytes for the website, 203 for the app. |
| Replace a dependency: Estimates ships a new response shape | Every client reads the new shape, including app versions already on phones. | One change in the route. The screens’ contract does not move. |
| Change a rule: hide estimates the model is unsure of | Each client has to learn it, and old app versions never will. | One condition in the route, deployed once. |
| A second team takes over the app | They depend on three services’ shapes and three teams’ release notes. | They own their route. The risk moves: keep domain rules, such as whether a parcel is late, in the services and out of the route. |
Before you add the layer, write down what you will measure and the result you would accept, so it is judged by something other than how the diagram looks:
- Time until the parcel is on screen, at the 95th percentile, measured in the browser, before and after. This is what the deadlines are for.
- Share of pages answered partial, by part and reason. If the estimate misses 300 ms on a large share of loads, fix the estimate service or stream it; do not raise the deadline in the dark.
- Bytes per page on the app, and calls still running after the page answered. The second should be zero.
This page did not run the tracking page for real customers, so it has no production numbers. The checker’s timings in section 08 are one sample on one machine, not a benchmark.
08 / Ask for it
Two prompts, two builds, one browser.
We sent two agents the same request for this page at the same time, both running Claude Sonnet, with the three services running for them to call. One prompt described the page. The other added an Architecture block: one data request shaped for the page, no field the page does not show, Shipments required, a deadline per call that cancels it, and named gaps. Then a script ran each build behind services it controlled and opened the page in Chromium, once per question.
| Question | Plain prompt | Architecture prompt |
|---|---|---|
| Data requests the browser makes | None: the server sends a finished page | /api/track/PX-4471 200 669 bytes |
| Internal ids a customer can read on the page | acct_77, u-2291, u-3307 | None |
| How the three calls start | One after another | Shipments first, then the other two together |
| The estimate takes 1.5 s: parcel shown after | 1,659 ms | 382 ms |
| The estimate never answers: parcel shown after | 4,164 ms | 377 ms |
| Scans take 1.2 s: parcel shown after | 1,407 ms | 878 ms |
| The estimate is down: what the page says | Estimated arrival is unavailable right now. | We couldn't load the delivery estimate right now (the service failed). |
| Scans send a code neither prompt listed, HLD | HLD | Status update (HLD) |
| Shipments is down | Tracking is temporarily unavailable | Could not look up parcel "PX-4471": the shipments service failed. |
The plain prompt did not produce the browser calling three services. It produced a server that called them, held the tokens, and sent the browser one finished page: a backend for frontend in all but name. What it never decided was what to wait for. It called the services one after another, each with a four-second timeout, so the page waited for the estimate every time the estimate was slow. And it showed every field it was given, which put the sender’s account id and the depot staff ids on a customer’s phone.
The architecture build did what its block asked, and the checker could see each part of it: one data request, no internal field in the reply, the estimate call hung up at 300 ms, and a page that says which part is missing. It also did one thing the block never forbade: it waited for Shipments before starting the other two calls, so every load pays Shipments’ time first.
let scans: Scan[] = [];
let scansUnavailable = false;
try {
scans = await getScans(parcelId);
} catch (err) {
console.error("scans lookup failed:", err);
scansUnavailable = true;
}
let eta: Eta | null = null;
let etaUnavailable = false;
try {
eta = await getEta(parcelId);
} catch (err) {
console.error("eta lookup failed:", err);
etaUnavailable = true;
}
async function handleTrackApi(id: string): Promise<TrackApiResult> {
const shipmentsResult = await callService(
`${SHIPMENTS_URL}/shipments/${encodeURIComponent(id)}`,
SHIPMENTS_AUTH,
SHIPMENTS_TIMEOUT_MS,
);
if (!shipmentsResult.ok) {
if (shipmentsResult.status === 404) {
return { status: 404, body: { error: `No parcel found for "${id}".` } };
}
return {
status: 502,
body: {
error: `Could not look up parcel "${id}": the shipments service ${shipmentsResult.reason}.`,
},
};
}
const [scansResult, etaResult] = await Promise.all([
callService(`${SCANS_URL}/scans/${encodeURIComponent(id)}`, SCANS_AUTH, SCANS_TIMEOUT_MS),
callService(`${ETA_URL}/eta/${encodeURIComponent(id)}`, ETA_AUTH, ETA_TIMEOUT_MS),
]); Two lines were missing from both prompts. One is timing: start every call at once; nothing waits for another part unless it needs that part’s answer. The other is hidden in a product sentence. “Sees the parcel, where it has been, and when it should arrive” reads like a feature, and the plain build took it to mean every field of the parcel. The line that closes it names the fields: the page shows these fields and no others.
Neither build hid a code it did not know. The plain build’s status badge showed HLD; the architecture build’s fallback kept it in brackets. The prompt asked
for a “sensible fallback” and did not say what a customer should read.
How the runs were made and checkedOne run each, recorded as written
- Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. The only differences were the Architecture block and the output folder.
- The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. For every question the checker restores a build into a temporary folder, starts its own fake services with that question’s behavior, and opens the page at phone size in Chromium, recording every response the browser receives.
- The checker’s first run reported the plain build’s estimate message as “Estimated arrival”, which is the section’s heading. It now reads the line after the heading too. Both runs are kept.
- Both agents wrote something outside their folders against the prompt: the plain agent saved two test pages to a system temporary folder, and the architecture agent tried to write a log at the filesystem root, which failed. Neither touched the other’s folder.
- This is one sample of each prompt, not a measurement of a model. The timings are one machine’s, a few milliseconds either way between checker runs.
09 / Hold it there
Keep the tokens behind the door, and the deadlines under test.
The next change can quietly undo this: a component that fetches Scans directly because it was quicker, a deadline raised to make a complaint go away. Three kinds of check keep the shape.
The framework’s own door
Service tokens belong in server-only configuration. SvelteKit refuses to let code that reaches the browser import
$env/static/private,$lib/server, or*.server.tsfiles, and names the import chain when it does (server-only modules). In Next.js,import 'server-only'makes importing a module into a Client Component a build error, and environment variables without theNEXT_PUBLIC_prefix are empty strings in the browser (Server and Client Components).An import rule an agent cannot argue with
Only the page’s server route, and the server folder it lives beside, may import the service clients. A component that wants Scans asks the route. Enforcement layer runs rules like this against real code, and Architecture as rules writes them from one declaration. This rule was not run against this lesson’s files, which have no app around them.
.dependency-cruiser.cjs // .dependency-cruiser.cjs module.exports = { forbidden: [ { name: 'only-the-page-route-calls-services', comment: 'Service clients hold tokens. Components and browser code ask the route instead.', severity: 'error', from: { path: '^src/', pathNot: '\\+page\\.server\\.ts$|^src/lib/server/' }, to: { path: '^src/lib/server/services/' } } ] };A check on what actually happens
Rules see imports, not timing and not payloads. So check the behavior: open the page with a slow service behind it and time it, and search everything the browser received for tokens and internal ids. The checker in section 08 does both, and the lesson’s own tests assert that the estimate call is cancelled at 300 ms and that no internal field is in the reply.
check-runs.mjs /** Open a page in Chromium and record everything the browser receives. */ async function visit(browser, path, { waitFor = 'Leeds', limit = 12_000 } = {}) { const page = await browser.newPage({ viewport: { width: 390, height: 844 } }); const received = []; page.on('response', async (response) => { const request = response.request(); let body = ''; try { body = await response.text(); } catch { /* redirects and aborted bodies have none */ } received.push({ url: response.url(), type: request.resourceType(), status: response.status(), bytes: Buffer.byteLength(body), body }); }); const started = Date.now(); const document = await page.goto(`${BASE}${path}`, { waitUntil: 'commit', timeout: limit }); let shownAfter = null;
Your load function is already oneA page’s server load that calls several APIs is a backend for frontend. The deadline is the part you own.
Where it already is in your components
A SvelteKit +page.server.ts load that calls three APIs and returns what the
page renders is a backend for frontend for one page. So is a Next.js server component that
awaits three fetches before it returns markup. They run on the server, so they can hold
tokens; they decide the shape the component receives; and whatever they await is
what the page waits for.
Most of them are written like the pass-through route: one Promise.all, every
field passed down. That is fine until one of the three is slow.
When you have to own it
The day the estimate service slows down, you decide what the page does without it. Keep
the required parts under a deadline and render them, and send the optional part
separately. SvelteKit streams a promise a server load returns without awaiting, “useful if
you have slow, non-essential data, since you can start rendering the page before all the
data is available” (Streaming with promises); streaming needs JavaScript in the browser. In React, the same move is a <Suspense> boundary around the component that reads the estimate.
The tracking page as most of us write it: the server calls the three services, and the component renders the parcel, the latest scan, and the estimate, with a message for a missing part.
// app/track/[id]/page.tsx: a Next.js server component. It runs on the server,
// so it can hold service tokens and call three internal APIs. The browser gets
// only the HTML this returns.
import { notFound } from 'next/navigation';
import { shipments, scans, estimates } from '@/server/services';
export default async function TrackingPage({ params }: { params: Promise<{ id: string }> }) {
const { id } = await params;
const [parcel, history, estimate] = await Promise.allSettled([
shipments.get(id),
scans.list(id),
estimates.get(id)
]);
if (parcel.status === 'rejected') notFound();
const latest = history.status === 'fulfilled' ? history.value.at(-1) : undefined;
return (
<main>
<h1>Parcel {parcel.value.id}</h1>
<p>
{parcel.value.from} → {parcel.value.to} · {parcel.value.service}
</p>
<p>
{latest ? `${latest.text}, ${latest.time}` : 'Tracking history is unavailable right now.'}
</p>
<p>
{estimate.status === 'fulfilled'
? `Arriving ${estimate.value.window}`
: 'We can’t estimate a delivery time right now.'}
</p>
</main>
);
}
10 / Make the call
Add the layer when a screen has something to decide.
Let the browser call a service directly when there is one service, it was built for browsers, its answer is what the screen shows, and there is no token to hide: a public read API behind a CDN, say. Then a layer in between only adds a hop. That is plain client–server.
Add a backend for frontend when the screen needs several services, when one of them can be slow and the screen can live without it, when fields must not reach the browser, or when a second client wants a different shape. Reopen the decision when the route starts deciding things that are true for every client, such as whether a parcel is late: that rule belongs in a service.
Take it with you
Explain it without saying “backend for frontend”: “The page asks its own server once. That server calls the others at the same time, gives each a time limit, and tells the page what it got, in the page’s words, and what it could not get.” Then open the last page load you wrote that calls more than one API, and find what it awaits.
Paste into your next prompt, and fill in the blanks
The browser makes one request for <this screen>'s data, to one endpoint shaped for it. It never calls <the services> and never receives their tokens. The reply carries only the fields the screen shows: <list them>. <The required part> is required; without it, answer with a clear error. <The optional parts> are optional. Start every call at once. Give each a deadline: <n> ms for <part>. When a deadline passes, cancel that call and answer without it. The reply names each missing part and why: timed out or failed. The services own the facts; this route only chooses, translates, and shapes them.
Connections to follow nextRelated lessons
- Client–server architecture is the line this lesson adds a layer behind: the browser asks, the server decides.
- REST, RPC, or GraphQL is the next choice: the shape of the contract between the screen and its server.
- Timeouts, deadlines, and races explains why a timer without cancellation still leaves work running.
- The mapping layer is the move inside the route: service shapes in, screen shapes out.
- Containing failure takes the optional-part idea to a whole system.