01 / The prompt
“Build the API for our library app.”
The app has three things to do. A patron searches for “garden” and sees each book with how many copies are on the shelf. They open a book and see every copy, branch by branch, and how many holds are waiting. They place a hold. Ask an agent for the API behind those screens and you will get one of three plausible answers:
- REST.
GET /books?q=garden,GET /books/b1/copies,POST /books/b2/holds. Nouns, HTTP verbs, and status codes. - RPC.
searchCatalog({ q }),getBookPage({ id }),placeHold({ bookId }). Named operations, often through tRPC or a framework’s server functions. - GraphQL. One endpoint and a schema, and each screen sends a query naming the fields it wants.
All three work. The prompt never said what decides between them: how many round trips a phone may make per screen, what an HTTP cache may keep, who else will call this API, and what happens when a hold request is sent twice. Those are the questions this lesson answers, with the same catalog built all three ways.
02 / Name the choice
Each style decides what the contract names.
REST names resources and uses HTTP’s verbs on them; a response is a representation of the resource, and HTTP’s caching and retry rules apply as written. RPC names operations; each call says what it wants done and returns what that caller needs. GraphQL names a typed graph; a client selects the fields it wants, and the server runs a resolver for each one.
The style decides what a client can ask for and what HTTP can do for you. It does not decide round trips, what is cached, or whether a retry is safe. You write those in any style.
| Question | REST | RPC | GraphQL |
|---|---|---|---|
| What the client relies on | URLs, verbs, status codes | Operation names and argument shapes | The schema, and its own queries |
| Who shapes a screen’s data | The client, from resources | The server, one operation per need | The client, by selection |
| What an HTTP cache can do | Keep any GET marked max-age | Nothing here: every call is a POST | Nothing here: queries are POSTs. A server may accept GET for queries. |
| What a new screen costs | More requests, or a new resource | A new operation | A new query; the server changes only if the graph does |
| Where cost hides | Round trips and unused fields | Operations that multiply | Resolvers: each field reads, and queries can nest |
| What makes a retried write safe | A key, such as an Idempotency-Key header | A key, as an argument | A key, as a mutation input |
Words to put in a prompt or a review
- Resource
- A thing with a URL, such as a book or its holds.
- Operation
- A named action with arguments, such as
placeHold. - Selection
- The fields a GraphQL query asks for, and nothing else.
- Resolver
- The server function behind one GraphQL field. Selected fields cost reads.
- Round trip
- One request and its reply. On a phone, the expensive part.
- Cache-Control
- The header that says whether, and for how long, a cache may reuse a reply.
- Idempotency key
- A value the client sends so the server can recognize a retry.
They combineMost real APIs are a mix
The styles are not exclusive. A REST API can add one resource shaped for a screen, such as /availability?ids=b1,b2,b3, which is how section 05 ends. An app can use RPC
server functions for its own screens and publish REST for partners. A GraphQL server can
accept GET for queries; graphql.org notes that a client “may be encouraged to leverage
this option to facilitate HTTP caching or edge caching in a content delivery network,” and
that mutations must use POST (Serving over HTTP). Choose the style for the contract most of your callers depend on, and bend it where a
screen needs it.
03 / Follow one screen
Watch the same search screen go out three ways.
The phone searches for “garden” and shows three books with their copies on the shelf. First in each style alone, then loaded twice through an HTTP cache, then a hold whose reply was lost. Open Try it to load any screen in all three styles at once, or write your own GraphQL query and see what it costs.
What does each style put on the wire?
REST
- Nothing sent yet
0 requests reached the server
The search screen asks for the search resource, then each book’s availability.
Requests go out…
Reduced motion: choose a scene to see its completed state.
Read this scene
Requests go out…
REST: nothing sent yet.
Watch restarts the story when you come back. Step through keeps your step. Try it runs a fresh catalog for every screen you load.
04 / Read the three shapes
One core, three doors, the same screens.
Basic form is the client: the search screen in each style. In the wild is the server: three doors over one catalog, including enough of GraphQL to parse a query, check it against the schema and a depth limit, and resolve it field by field. At the call site is placing a hold, where a retry happens.
Notice that Library does not change between styles. Everything the lesson measures,
requests, bytes, reads, cache hits, and duplicate holds, is decided in the doors and the clients.
The same search screen asked three ways, from the client: a search resource and then each book’s availability; one operation named for the screen; one query that selects four fields.
// The same search screen asked three ways. Each line is a book's title,
// author, and how many copies are on the shelf.
const line = (b: { title: string; author: string; available: number; total: number }) =>
`${b.title} · ${b.author} · ${b.available} of ${b.total} in`;
// REST: the search resource, then each book's availability resource.
export function searchScreenRest(send: Send, q: string): string[] {
const books = JSON.parse(
send({ method: 'GET', path: `/books?q=${encodeURIComponent(q)}` }).body
) as Book[];
return books.map((b) => {
const a = JSON.parse(send({ method: 'GET', path: `/books/${b.id}/availability` }).body);
return line({ ...b, ...a });
});
}
// RPC: one named operation that returns what this screen shows.
export function searchScreenRpc(send: Send, q: string): string[] {
const res = send({ method: 'POST', path: '/rpc/searchCatalog', body: JSON.stringify({ q }) });
return (JSON.parse(res.body) as Parameters<typeof line>[0][]).map(line);
}
// GraphQL: one query that selects the fields this screen shows.
export function searchScreenGraphql(send: Send, q: string): string[] {
const query = `{ search(q: ${JSON.stringify(q)}) { title author available total } }`;
const res = send({ method: 'POST', path: '/graphql', body: JSON.stringify({ query }) });
return (JSON.parse(res.body).data.search as Parameters<typeof line>[0][]).map(line);
} // The same search screen asked three ways. Each line is a book's title,
// author, and how many copies are on the shelf.
type searchItem struct {
ID string `json:"id,omitempty"`
Title string `json:"title"`
Author string `json:"author"`
Available int `json:"available"`
Total int `json:"total"`
}
func line(b searchItem) string {
return fmt.Sprintf("%s · %s · %d of %d in", b.Title, b.Author, b.Available, b.Total)
}
// SearchScreenREST reads the search resource, then each book's availability resource.
func SearchScreenREST(send Send, q string) []string {
var books []Book
json.Unmarshal([]byte(send(Req{Method: "GET", Path: "/books?q=" + url.QueryEscape(q)}).Body), &books)
lines := []string{}
for _, b := range books {
var a searchItem
json.Unmarshal([]byte(send(Req{Method: "GET", Path: "/books/" + b.ID + "/availability"}).Body), &a)
lines = append(lines, line(searchItem{Title: b.Title, Author: b.Author, Available: a.Available, Total: a.Total}))
}
return lines
}
// SearchScreenRPC calls one named operation that returns what this screen shows.
func SearchScreenRPC(send Send, q string) []string {
body, _ := json.Marshal(map[string]string{"q": q})
var items []searchItem
json.Unmarshal([]byte(send(Req{Method: "POST", Path: "/rpc/searchCatalog", Body: string(body)}).Body), &items)
lines := []string{}
for _, b := range items {
lines = append(lines, line(b))
}
return lines
}
// SearchScreenGraphQL sends one query that selects the fields this screen shows.
func SearchScreenGraphQL(send Send, q string) []string {
quoted, _ := json.Marshal(q)
body, _ := json.Marshal(map[string]string{"query": "{ search(q: " + string(quoted) + ") { title author available total } }"})
var out struct {
Data struct{ Search []searchItem } `json:"data"`
}
json.Unmarshal([]byte(send(Req{Method: "POST", Path: "/graphql", Body: string(body)}).Body), &out)
lines := []string{}
for _, b := range out.Data.Search {
lines = append(lines, line(b))
}
return lines
} The behavior these examples promiseChecked by 30 shared scenarios
- Search matches title or author, ignoring case. A screen line reads “title · author · 2 of 3 in”. The book page lists every copy and “0 holds waiting”.
- REST:
/books?q=and/books/:idsend whole books andmax-age=60;/availability,/copies, and/holdssendno-store. RPC and GraphQL are POSTs. - Each search, book lookup, copies read, and queue read counts as one read. GraphQL’s
availableandtotaleach read the copies. - A shared HTTP cache keeps a GET whose reply says
max-age, by URL. - A hold adds one to the queue. The same key from the same patron returns the first hold and writes nothing; the same key from another patron is a new hold.
- GraphQL rejects a query deeper than four levels, or a field the schema lacks, before reading.
Every expectation in the shared cases, down to the bytes each request sends and receives, was produced by a separate model written from these rules and kept beside the examples, not copied from either implementation.
Reading the TypeScriptA schema as a table of resolvers
The GraphQL schema is a record from type name to field name to a resolver and the type
it returns, so checking a query and running it walk the same table. deepest measures nesting before anything runs. The parser is a regular expression
that splits the query into tokens and a small recursive function over them.
Reading the GoKeeping the selection’s order
Go maps have no order, and a GraphQL reply lists fields in the order the query selected
them, so resolved objects are an object slice of key and value pairs with
its own MarshalJSON. REST and RPC replies are structs with JSON tags, whose
field order is the declaration order, which is what makes the byte counts match across
languages.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types library.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:
module heyrian.dev/lessons/api-styles
go 1.23
search "garden" · rest: 4 requests, 0 bytes sent, 603 received, 4 reads search "garden" · rpc: 1 request, 14 bytes sent, 285 received, 4 reads search "garden" · graphql: 1 request, 70 bytes sent, 275 received, 7 reads search twice through an HTTP cache · requests reaching the server: rest 7, rpc 2, graphql 2 hold sent twice, no key · queue positions: rest 3, 4; rpc 3, 4; graphql 3, 4 similar books, four deep · graphql: query depth 6 exceeds 4
05 / Review the agent’s diff
“The search screen is faster now.”
Four round trips on a phone is a real complaint, and the fix is one header. Read what that header lets a cache do before you decide.
06 / How it fails
Each style fails in the place it put the work.
Rows backed by a shared case or a native test say so; the rest are marked as authored.
| What goes wrong | REST | RPC | GraphQL |
|---|---|---|---|
| Slow network: round trips add up (cases) | 4 requests for the search screen, 3 for the book page | 1 each | 1 each |
| Slow server: work multiplies (cases, native) | 4 reads for the search screen | 4 reads | 7 reads, and 12 for similar books two deep; a deeper query is rejected |
| Stale: live data in a cached reply (exercise) | Possible, one header away | Not by HTTP: nothing is cached | Not by HTTP: nothing is cached |
| Wrong: a client asks for what is not there (native) | 404 for a missing book | 404 for a missing book or an unknown operation | An error naming the field, before any read |
| Duplicated: a hold sent twice (cases) | Two holds in every style. With a key, one. | ||
| Half-done: one of several requests fails (authored) | The screen has books with no availability; it must say so per row | All or nothing | Partial data with an error per field, if resolvers allow it |
The duplicated row is the one people get wrong: no URL, operation name, or mutation makes a retried write safe by itself. Idempotency and at-least-once explains why a retry is the normal case, and Caching and invalidation covers what a cache may keep.
07 / Is it worth it?
The same four changes, three ways.
A choice is worth making on purpose when the changes your app will get land differently. Here is how each style takes them.
| Change | REST | RPC | GraphQL |
|---|---|---|---|
| A second client: a branch kiosk with its own screens | Reuses the resources, and pays their round trips | New operations for its screens | New queries; no server change |
| Replace a dependency: holds move to a vendor system | A change behind Library. No difference between styles. | ||
| Change a rule: “author” becomes a list of authors | Add authors, keep author; the server cannot tell who still
reads it | A second version of each operation that returns it | Add authors, deprecate author, and count the queries that
still select it |
| A second team owns holds | Their own resources under one prefix | Their own operations | One schema both teams must agree on |
Before you choose, or change, decide what you will measure and the result you would accept:
- Requests and bytes per screen, at the 95th percentile, on a phone. The shared cases pin these for this catalog; your app has its own.
- Cache hit ratio for the reads you marked cacheable. If it is near zero, caching is not a reason to prefer REST here.
- Server reads per request, by operation or query. A GraphQL query that reads twelve times for three books is a resolver-batching problem before it is a style problem.
- Duplicate writes per thousand hold requests, before and after adding keys.
This page did not measure a real app, so it gives no production numbers. The counts on this page are the catalog’s own, from the shared cases.
08 / Ask for it
Two prompts, two builds, one browser.
We sent two agents the same request at the same time, both running Claude Sonnet. One prompt described the app. The other added an Architecture block: choose the style on purpose and write down why, the search screen in one request, what a shared cache may and may not keep, a key on every hold, and additive changes only. A script then opened each build in Chromium and asked the same five questions.
| Question | Plain prompt | Architecture prompt |
|---|---|---|
| The search screen for “garden” | page GET / (200, 2599 B) | page GET / (200, 2703 B); data GET /api/books (200, 480 B) |
| The book page | page GET /book/b1 (200, 4105 B) | page GET /book/b1 (200, 4103 B); data GET /api/books/b1 (200, 405 B) |
| Responses a shared cache may keep | None of 2 | None of 4 |
| One click on Place hold, then the identical request again | 1 hold | 1 hold |
| Another patron sends the same request and key | 2 holds, one each | 1 hold between them |
The plain build did not choose REST, RPC, or GraphQL. It rendered every page on the server and sent HTML. For one app with two screens and no other client, that is a fourth answer and a good one: one request per screen, no API to version. It also made holds safe to retry, by a product rule rather than a key: a patron cannot hold the same book twice.
The architecture build chose REST, and wrote down why in terms this lesson would accept.
Then its own requirements collided: the search screen had to arrive in one request with the
copies on the shelf, those counts could never be cached, and the search could be. It chose
freshness, marked every data response no-store, and explained that too.
`GET /api/books` and `GET /api/books/:id` both include live numbers — the
on-shelf copy count and, on the book page, the holds-waiting count — inside
the same response the rest of the page data rides in. HTTP cache directives
apply to a whole response, not to individual fields, and the spec is
explicit that those numbers must never come from a cache. So both endpoints
are sent with `Cache-Control: no-store`. This is the strict, honest way to
satisfy "must never be served from a cache" without splitting a single
"get everything in one request" response into two round trips. The 60-second
shared-cache allowance the spec grants is a ceiling, not a requirement — it const already = holdIdempotency.get(idempotencyKey);
if (already) {
// Same key seen before: return the original result, don't place
// a second hold.
sendJson(res, already.status, { ...(already.body as object), replayed: true });
return;
}
const patronId = getPatronId(req); Look at the order in that second excerpt: the key is checked before the patron is read. When
another patron sent a request with the same key, the checker got back the first patron’s
hold, marked replayed: true. The line both prompts were missing is about scope: a key is remembered per patron, and the same key from someone else is a new request. The plain build got there by accident, through its per-patron rule:
const alreadyHeld = holders.has(patronId);
holders.add(patronId);
sendJson(res, 200, { holds: holders.size, alreadyHeld }); How the runs were made and checkedOne run each, recorded as written
- Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker.
- The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s
examples. For every question the checker restores a build, starts it fresh, and opens it
at phone size in Chromium, recording every page and data response with its size and
Cache-Control. - The retry questions click “Place hold” once, capture the request the page sent, and send it again unchanged, then with another patron’s header.
- The architecture agent kept its test server’s log and process id in the system’s temporary folder, against the prompt, and deleted them. Neither agent touched the other’s folder.
- This is one sample of each prompt, not a measurement of a model.
09 / Hold it there
Put the cache rule and the key where a test can see them.
A style chosen on purpose drifts one endpoint at a time: a header added to make a screen faster, a mutation added without a key. Three kinds of check keep it.
HTTP’s own rules
REST inherits them, and so does anything else sent over HTTP.
no-store“indicates that a cache MUST NOT store any part of either the immediate request or the response” (RFC 9111). “Responses to POST requests are only cacheable when they include explicit freshness information” and a matchingContent-Location, which is why RPC and GraphQL over POST get no HTTP caching here. And a method is idempotent “if the intended effect on the server of multiple identical requests with that method is the same as the effect for a single such request”; POST is not one of them (RFC 9110). A hold is a POST in every style, so it needs a key.A test that pins the contract
Headers are part of the contract, so test them like fields: live data is
no-store, stable reads carrymax-age, and nothing cacheable carries a live count. For GraphQL, the same kind of test sends a query one level too deep and expects a rejection with no reads, as this lesson’s own specs do. Run tests like this in CI, the way Enforcement layer runs its rules.catalog.spec.ts // catalog.spec.ts it('never lets a shared cache keep live availability', async () => { for (const path of ['/books/b1/availability', '/books/b1/holds', '/books/b1/copies']) { const res = await fetch(`${base}${path}`); expect(res.headers.get('cache-control')).toBe('no-store'); } const search = await fetch(`${base}/books?q=garden`); expect(search.headers.get('cache-control')).toMatch(/max-age=60/); expect(await search.text()).not.toMatch(/available|on the shelf/); });A check on what actually happens
Tests see the server. A browser sees the screen. The checker in section 08 opens each page, counts its requests and cache headers, and replays a hold request as another patron, which is how the global key store showed up.
check-runs.mjs async "another patron's request with the same key"(browser) { const { extra } = await visit(browser, '/book/b2', async (page) => { const before = holdLine(await page.evaluate(() => document.body.innerText)); const captured = page.waitForRequest((r) => r.method() === 'POST', { timeout: 5000 }); await page.getByRole('button', { name: /place hold/i }).click(); const request = await captured; await page.waitForLoadState('networkidle'); const headers = { ...request.headers(), 'x-patron': 'p-311' }; const other = await page.request.fetch(request.url(), { method: 'POST', headers, data: request.postData() ?? undefined }); await page.reload({ waitUntil: 'networkidle' }); const after = holdLine(await page.evaluate(() => document.body.innerText)); return { otherStatus: other.status(), otherBody: (await other.text()).slice(0, 200), before, after, created: holdCount(after) - holdCount(before) };
Every fetch in your components picked oneA row that fetches its own data is where the style’s cost shows up first.
Where it already is in your components
Every data call in your components is one of these. fetch('/api/books?q=') is
REST; a tRPC procedure or a SvelteKit or Next.js server function is RPC; an Apollo or urql useQuery with a query string is GraphQL. The component tree often decides the round
trips for you: a list whose rows each fetch their own availability is the REST search screen
above, four requests for three books.
When you have to own it
The day that list is slow on a phone, you choose where the fix goes. Keep the cacheable part cacheable and batch the live part into one request; or add one operation for the screen; or move the screen to a query and pay for resolvers that batch. The samples show the first, which keeps the style and removes the round trips.
A search list whose rows each fetch their own availability. It reads well, and makes one request per book.
// SearchResults.tsx. Each row fetches its own availability, which reads well
// and quietly makes one request per book: three rows, four requests.
import { useEffect, useState } from 'react';
type Book = { id: string; title: string; author: string };
function Availability({ id }: { id: string }) {
const [text, setText] = useState('…');
useEffect(() => {
fetch(`/books/${id}/availability`)
.then((r) => r.json())
.then((a: { available: number; total: number }) =>
setText(`${a.available} of ${a.total} in`)
);
}, [id]);
return <span>{text}</span>;
}
export function SearchResults({ q }: { q: string }) {
const [books, setBooks] = useState<Book[]>([]);
useEffect(() => {
fetch(`/books?q=${encodeURIComponent(q)}`)
.then((r) => r.json())
.then(setBooks);
}, [q]);
return (
<ul>
{books.map((b) => (
<li key={b.id}>
{b.title} · {b.author} · <Availability id={b.id} />
</li>
))}
</ul>
);
}
10 / Make the call
Choose for your callers, then write down what would change it.
One app, one team, screens that map to actions: start with RPC, through your framework’s server functions or tRPC, or with pages rendered on the server, as the plain build did. You accept that a second client will need its own operations.
Several clients, or anyone outside your team, and reads worth caching: REST, with a resource shaped for a screen where round trips hurt. You accept extra requests and fields the screen does not use.
Many screens needing different shapes of a real graph, and a team that can own the schema: GraphQL, with a depth or cost limit and batched resolvers from the first day. You accept no HTTP cache for POSTed queries, and resolver costs you have to watch.
Reconsider when a second client arrives, when requests per screen or reads per request show up in your measurements, or when you need to retire a field and cannot tell who reads it.
Take it with you
Explain it without saying “REST”, “RPC”, or “GraphQL”: “Either the client asks for things by address and the web’s caches help, or it calls actions the server wrote for it, or it asks for exactly the fields it wants and the server works each one out.” Then open the screen you touched last and count its requests.
Keep a note of the decision
Why: <the screens and clients that call this API> What: <REST, RPC, or GraphQL>, at <where the contract is written down> Constraint: <round trips per screen, what may be cached, who else calls it> Fallback: <what you accepted, such as extra requests or no HTTP cache> Reconsider when: <a second client, a slow screen, a field you cannot retire>
Paste into your next prompt, and fill in the blanks
Use <REST | RPC | GraphQL> for this API, because <the screens and clients>. Write the choice and the reason in <API.md>. <Screen> gets what it shows in <n> requests. A shared cache may keep <stable reads> for <n> seconds. <Live data> is never cached, so it is not in the same response as anything that is. Every write that a client may retry takes a key, scoped to the <user> who sent it; the same key from the same <user> returns the first result. <GraphQL only:> reject queries deeper than <n>, and batch reads per request.
Connections to follow nextRelated lessons
- Backend for frontend is where a screen-shaped operation usually lives.
- API contracts covers what a client may rely on, in any style.
- Versioning and compatibility is the “author becomes authors” change in depth.
- Idempotency and at-least-once is the hold sent twice.
- Polling, server-sent events, or WebSockets is the next choice: how a screen stays current once it has loaded.