← Architecture
Client and server Choose what the contract names

REST, RPC, or GraphQL

Resources, operations, or a graph to select from.

Every fetch your components make speaks one of three dialects: a resource, a named operation, or a query that picks its fields. Usually the first endpoint decided it. Let’s build one small library catalog all three ways and count what each costs a phone.

TypeScriptGoOne catalog, three API styles, two recorded builds.

01 / The prompt

“Build the API for our library app.”

The app has three things to do. A patron searches for “garden” and sees each book with how many copies are on the shelf. They open a book and see every copy, branch by branch, and how many holds are waiting. They place a hold. Ask an agent for the API behind those screens and you will get one of three plausible answers:

  • REST. GET /books?q=garden, GET /books/b1/copies, POST /books/b2/holds. Nouns, HTTP verbs, and status codes.
  • RPC. searchCatalog({ q }), getBookPage({ id }), placeHold({ bookId }). Named operations, often through tRPC or a framework’s server functions.
  • GraphQL. One endpoint and a schema, and each screen sends a query naming the fields it wants.

All three work. The prompt never said what decides between them: how many round trips a phone may make per screen, what an HTTP cache may keep, who else will call this API, and what happens when a hold request is sent twice. Those are the questions this lesson answers, with the same catalog built all three ways.

02 / Name the choice

Each style decides what the contract names.

REST names resources and uses HTTP’s verbs on them; a response is a representation of the resource, and HTTP’s caching and retry rules apply as written. RPC names operations; each call says what it wants done and returns what that caller needs. GraphQL names a typed graph; a client selects the fields it wants, and the server runs a resolver for each one.

The style decides what a client can ask for and what HTTP can do for you. It does not decide round trips, what is cached, or whether a retry is safe. You write those in any style.

What each style commits you to, on this catalog
QuestionRESTRPCGraphQL
What the client relies onURLs, verbs, status codesOperation names and argument shapesThe schema, and its own queries
Who shapes a screen’s dataThe client, from resourcesThe server, one operation per needThe client, by selection
What an HTTP cache can doKeep any GET marked max-ageNothing here: every call is a POSTNothing here: queries are POSTs. A server may accept GET for queries.
What a new screen costsMore requests, or a new resourceA new operationA new query; the server changes only if the graph does
Where cost hidesRound trips and unused fieldsOperations that multiplyResolvers: each field reads, and queries can nest
What makes a retried write safeA key, such as an Idempotency-Key headerA key, as an argumentA key, as a mutation input

Words to put in a prompt or a review

Resource
A thing with a URL, such as a book or its holds.
Operation
A named action with arguments, such as placeHold.
Selection
The fields a GraphQL query asks for, and nothing else.
Resolver
The server function behind one GraphQL field. Selected fields cost reads.
Round trip
One request and its reply. On a phone, the expensive part.
Cache-Control
The header that says whether, and for how long, a cache may reuse a reply.
Idempotency key
A value the client sends so the server can recognize a retry.
They combineMost real APIs are a mix

The styles are not exclusive. A REST API can add one resource shaped for a screen, such as /availability?ids=b1,b2,b3, which is how section 05 ends. An app can use RPC server functions for its own screens and publish REST for partners. A GraphQL server can accept GET for queries; graphql.org notes that a client “may be encouraged to leverage this option to facilitate HTTP caching or edge caching in a content delivery network,” and that mutations must use POST (Serving over HTTP). Choose the style for the contract most of your callers depend on, and bend it where a screen needs it.

03 / Follow one screen

Watch the same search screen go out three ways.

The phone searches for “garden” and shows three books with their copies on the shelf. First in each style alone, then loaded twice through an HTTP cache, then a hold whose reply was lost. Open Try it to load any screen in all three styles at once, or write your own GraphQL query and see what it costs.

REST, RPC, or GraphQL

What does each style put on the wire?

REST

  1. Nothing sent yet

0 requests reached the server

01/ 05
REST: resources

The search screen asks for the search resource, then each book’s availability.

Requests go out…

Reduced motion: choose a scene to see its completed state.

Read this scene

Requests go out…

REST: nothing sent yet.

Watch restarts the story when you come back. Step through keeps your step. Try it runs a fresh catalog for every screen you load.

04 / Read the three shapes

One core, three doors, the same screens.

Basic form is the client: the search screen in each style. In the wild is the server: three doors over one catalog, including enough of GraphQL to parse a query, check it against the schema and a depth limit, and resolve it field by field. At the call site is placing a hold, where a retry happens.

Notice that Library does not change between styles. Everything the lesson measures, requests, bytes, reads, cache hits, and duplicate holds, is decided in the doors and the clients.

The same search screen asked three ways, from the client: a search resource and then each book’s availability; one operation named for the screen; one query that selects four fields.

TypeScriptReading
library.ts
// The same search screen asked three ways. Each line is a book's title,
// author, and how many copies are on the shelf.
const line = (b: { title: string; author: string; available: number; total: number }) =>
	`${b.title} · ${b.author} · ${b.available} of ${b.total} in`;

// REST: the search resource, then each book's availability resource.
export function searchScreenRest(send: Send, q: string): string[] {
	const books = JSON.parse(
		send({ method: 'GET', path: `/books?q=${encodeURIComponent(q)}` }).body
	) as Book[];
	return books.map((b) => {
		const a = JSON.parse(send({ method: 'GET', path: `/books/${b.id}/availability` }).body);
		return line({ ...b, ...a });
	});
}

// RPC: one named operation that returns what this screen shows.
export function searchScreenRpc(send: Send, q: string): string[] {
	const res = send({ method: 'POST', path: '/rpc/searchCatalog', body: JSON.stringify({ q }) });
	return (JSON.parse(res.body) as Parameters<typeof line>[0][]).map(line);
}

// GraphQL: one query that selects the fields this screen shows.
export function searchScreenGraphql(send: Send, q: string): string[] {
	const query = `{ search(q: ${JSON.stringify(q)}) { title author available total } }`;
	const res = send({ method: 'POST', path: '/graphql', body: JSON.stringify({ query }) });
	return (JSON.parse(res.body).data.search as Parameters<typeof line>[0][]).map(line);
}
GoAlongside
main.go
// The same search screen asked three ways. Each line is a book's title,
// author, and how many copies are on the shelf.
type searchItem struct {
	ID        string `json:"id,omitempty"`
	Title     string `json:"title"`
	Author    string `json:"author"`
	Available int    `json:"available"`
	Total     int    `json:"total"`
}

func line(b searchItem) string {
	return fmt.Sprintf("%s · %s · %d of %d in", b.Title, b.Author, b.Available, b.Total)
}

// SearchScreenREST reads the search resource, then each book's availability resource.
func SearchScreenREST(send Send, q string) []string {
	var books []Book
	json.Unmarshal([]byte(send(Req{Method: "GET", Path: "/books?q=" + url.QueryEscape(q)}).Body), &books)
	lines := []string{}
	for _, b := range books {
		var a searchItem
		json.Unmarshal([]byte(send(Req{Method: "GET", Path: "/books/" + b.ID + "/availability"}).Body), &a)
		lines = append(lines, line(searchItem{Title: b.Title, Author: b.Author, Available: a.Available, Total: a.Total}))
	}
	return lines
}

// SearchScreenRPC calls one named operation that returns what this screen shows.
func SearchScreenRPC(send Send, q string) []string {
	body, _ := json.Marshal(map[string]string{"q": q})
	var items []searchItem
	json.Unmarshal([]byte(send(Req{Method: "POST", Path: "/rpc/searchCatalog", Body: string(body)}).Body), &items)
	lines := []string{}
	for _, b := range items {
		lines = append(lines, line(b))
	}
	return lines
}

// SearchScreenGraphQL sends one query that selects the fields this screen shows.
func SearchScreenGraphQL(send Send, q string) []string {
	quoted, _ := json.Marshal(q)
	body, _ := json.Marshal(map[string]string{"query": "{ search(q: " + string(quoted) + ") { title author available total } }"})
	var out struct {
		Data struct{ Search []searchItem } `json:"data"`
	}
	json.Unmarshal([]byte(send(Req{Method: "POST", Path: "/graphql", Body: string(body)}).Body), &out)
	lines := []string{}
	for _, b := range out.Data.Search {
		lines = append(lines, line(b))
	}
	return lines
}
The behavior these examples promiseChecked by 30 shared scenarios
  • Search matches title or author, ignoring case. A screen line reads “title · author · 2 of 3 in”. The book page lists every copy and “0 holds waiting”.
  • REST: /books?q= and /books/:id send whole books and max-age=60; /availability, /copies, and /holds send no-store. RPC and GraphQL are POSTs.
  • Each search, book lookup, copies read, and queue read counts as one read. GraphQL’s available and total each read the copies.
  • A shared HTTP cache keeps a GET whose reply says max-age, by URL.
  • A hold adds one to the queue. The same key from the same patron returns the first hold and writes nothing; the same key from another patron is a new hold.
  • GraphQL rejects a query deeper than four levels, or a field the schema lacks, before reading.

Every expectation in the shared cases, down to the bytes each request sends and receives, was produced by a separate model written from these rules and kept beside the examples, not copied from either implementation.

Reading the TypeScriptA schema as a table of resolvers

The GraphQL schema is a record from type name to field name to a resolver and the type it returns, so checking a query and running it walk the same table. deepest measures nesting before anything runs. The parser is a regular expression that splits the query into tokens and a small recursive function over them.

Reading the GoKeeping the selection’s order

Go maps have no order, and a GraphQL reply lists fields in the order the query selected them, so resolved objects are an object slice of key and value pairs with its own MarshalJSON. REST and RPC replies are structs with JSON tags, whose field order is the declaration order, which is what makes the byte counts match across languages.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types library.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/api-styles

go 1.23
search "garden" · rest: 4 requests, 0 bytes sent, 603 received, 4 reads
search "garden" · rpc: 1 request, 14 bytes sent, 285 received, 4 reads
search "garden" · graphql: 1 request, 70 bytes sent, 275 received, 7 reads
search twice through an HTTP cache · requests reaching the server: rest 7, rpc 2, graphql 2
hold sent twice, no key · queue positions: rest 3, 4; rpc 3, 4; graphql 3, 4
similar books, four deep · graphql: query depth 6 exceeds 4

05 / Review the agent’s diff

“The search screen is faster now.”

Four round trips on a phone is a real complaint, and the fix is one header. Read what that header lets a cache do before you decide.

The agent’s pull request

“The search screen was slow on phones: four requests every time. Availability is now cacheable too. All 33 tests pass.”

// server/rest.ts, GET /books/:id/availability
			return json(200, availability(lib.copies(id)), {
			(removed)  'cache-control': 'no-store'
			(added)  'cache-control': 'max-age=60' // the CDN can answer these
			});
			
You are reviewing this change. What do you do?

06 / How it fails

Each style fails in the place it put the work.

Rows backed by a shared case or a native test say so; the rest are marked as authored.

Failure modes of the catalog, by style
What goes wrongRESTRPCGraphQL
Slow network: round trips add up (cases)4 requests for the search screen, 3 for the book page1 each1 each
Slow server: work multiplies (cases, native)4 reads for the search screen4 reads7 reads, and 12 for similar books two deep; a deeper query is rejected
Stale: live data in a cached reply (exercise)Possible, one header awayNot by HTTP: nothing is cachedNot by HTTP: nothing is cached
Wrong: a client asks for what is not there (native)404 for a missing book404 for a missing book or an unknown operationAn error naming the field, before any read
Duplicated: a hold sent twice (cases)Two holds in every style. With a key, one.
Half-done: one of several requests fails (authored)The screen has books with no availability; it must say so per rowAll or nothingPartial data with an error per field, if resolvers allow it

The duplicated row is the one people get wrong: no URL, operation name, or mutation makes a retried write safe by itself. Idempotency and at-least-once explains why a retry is the normal case, and Caching and invalidation covers what a cache may keep.

07 / Is it worth it?

The same four changes, three ways.

A choice is worth making on purpose when the changes your app will get land differently. Here is how each style takes them.

The same four changes, made in each style
ChangeRESTRPCGraphQL
A second client: a branch kiosk with its own screensReuses the resources, and pays their round tripsNew operations for its screensNew queries; no server change
Replace a dependency: holds move to a vendor systemA change behind Library. No difference between styles.
Change a rule: “author” becomes a list of authorsAdd authors, keep author; the server cannot tell who still reads itA second version of each operation that returns itAdd authors, deprecate author, and count the queries that still select it
A second team owns holdsTheir own resources under one prefixTheir own operationsOne schema both teams must agree on

Before you choose, or change, decide what you will measure and the result you would accept:

  • Requests and bytes per screen, at the 95th percentile, on a phone. The shared cases pin these for this catalog; your app has its own.
  • Cache hit ratio for the reads you marked cacheable. If it is near zero, caching is not a reason to prefer REST here.
  • Server reads per request, by operation or query. A GraphQL query that reads twelve times for three books is a resolver-batching problem before it is a style problem.
  • Duplicate writes per thousand hold requests, before and after adding keys.

This page did not measure a real app, so it gives no production numbers. The counts on this page are the catalog’s own, from the shared cases.

08 / Ask for it

Two prompts, two builds, one browser.

We sent two agents the same request at the same time, both running Claude Sonnet. One prompt described the app. The other added an Architecture block: choose the style on purpose and write down why, the search screen in one request, what a shared cache may and may not keep, a key on every hold, and additive changes only. A script then opened each build in Chromium and asked the same five questions.

What the checker found, run 2026-09-23
QuestionPlain promptArchitecture prompt
The search screen for “garden”page GET / (200, 2599 B)page GET / (200, 2703 B); data GET /api/books (200, 480 B)
The book pagepage GET /book/b1 (200, 4105 B)page GET /book/b1 (200, 4103 B); data GET /api/books/b1 (200, 405 B)
Responses a shared cache may keepNone of 2None of 4
One click on Place hold, then the identical request again1 hold1 hold
Another patron sends the same request and key2 holds, one each1 hold between them

The plain build did not choose REST, RPC, or GraphQL. It rendered every page on the server and sent HTML. For one app with two screens and no other client, that is a fourth answer and a good one: one request per screen, no API to version. It also made holds safe to retry, by a product rule rather than a key: a patron cannot hold the same book twice.

The architecture build chose REST, and wrote down why in terms this lesson would accept. Then its own requirements collided: the search screen had to arrive in one request with the copies on the shelf, those counts could never be cached, and the search could be. It chose freshness, marked every data response no-store, and explained that too.

API.md · architecture prompt
`GET /api/books` and `GET /api/books/:id` both include live numbers — the
on-shelf copy count and, on the book page, the holds-waiting count — inside
the same response the rest of the page data rides in. HTTP cache directives
apply to a whole response, not to individual fields, and the spec is
explicit that those numbers must never come from a cache. So both endpoints
are sent with `Cache-Control: no-store`. This is the strict, honest way to
satisfy "must never be served from a cache" without splitting a single
"get everything in one request" response into two round trips. The 60-second
shared-cache allowance the spec grants is a ceiling, not a requirement — it
server.ts · architecture prompt
const already = holdIdempotency.get(idempotencyKey);
if (already) {
  // Same key seen before: return the original result, don't place
  // a second hold.
  sendJson(res, already.status, { ...(already.body as object), replayed: true });
  return;
}

const patronId = getPatronId(req);

Look at the order in that second excerpt: the key is checked before the patron is read. When another patron sent a request with the same key, the checker got back the first patron’s hold, marked replayed: true. The line both prompts were missing is about scope: a key is remembered per patron, and the same key from someone else is a new request. The plain build got there by accident, through its per-patron rule:

server.ts · plain prompt
const alreadyHeld = holders.has(patronId);
holders.add(patronId);

sendJson(res, 200, { holds: holders.size, alreadyHeld });
How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. For every question the checker restores a build, starts it fresh, and opens it at phone size in Chromium, recording every page and data response with its size and Cache-Control.
  • The retry questions click “Place hold” once, capture the request the page sent, and send it again unchanged, then with another patron’s header.
  • The architecture agent kept its test server’s log and process id in the system’s temporary folder, against the prompt, and deleted them. Neither agent touched the other’s folder.
  • This is one sample of each prompt, not a measurement of a model.

09 / Hold it there

Put the cache rule and the key where a test can see them.

A style chosen on purpose drifts one endpoint at a time: a header added to make a screen faster, a mutation added without a key. Three kinds of check keep it.

  1. HTTP’s own rules

    REST inherits them, and so does anything else sent over HTTP. no-store “indicates that a cache MUST NOT store any part of either the immediate request or the response” (RFC 9111). “Responses to POST requests are only cacheable when they include explicit freshness information” and a matching Content-Location, which is why RPC and GraphQL over POST get no HTTP caching here. And a method is idempotent “if the intended effect on the server of multiple identical requests with that method is the same as the effect for a single such request”; POST is not one of them (RFC 9110). A hold is a POST in every style, so it needs a key.

  2. A test that pins the contract

    Headers are part of the contract, so test them like fields: live data is no-store, stable reads carry max-age, and nothing cacheable carries a live count. For GraphQL, the same kind of test sends a query one level too deep and expects a rejection with no reads, as this lesson’s own specs do. Run tests like this in CI, the way Enforcement layer runs its rules.

    catalog.spec.ts
    // catalog.spec.ts
    it('never lets a shared cache keep live availability', async () => {
    	for (const path of ['/books/b1/availability', '/books/b1/holds', '/books/b1/copies']) {
    		const res = await fetch(`${base}${path}`);
    		expect(res.headers.get('cache-control')).toBe('no-store');
    	}
    	const search = await fetch(`${base}/books?q=garden`);
    	expect(search.headers.get('cache-control')).toMatch(/max-age=60/);
    	expect(await search.text()).not.toMatch(/available|on the shelf/);
    });
  3. A check on what actually happens

    Tests see the server. A browser sees the screen. The checker in section 08 opens each page, counts its requests and cache headers, and replays a hold request as another patron, which is how the global key store showed up.

    check-runs.mjs
    async "another patron's request with the same key"(browser) {
    	const { extra } = await visit(browser, '/book/b2', async (page) => {
    		const before = holdLine(await page.evaluate(() => document.body.innerText));
    		const captured = page.waitForRequest((r) => r.method() === 'POST', { timeout: 5000 });
    		await page.getByRole('button', { name: /place hold/i }).click();
    		const request = await captured;
    		await page.waitForLoadState('networkidle');
    		const headers = { ...request.headers(), 'x-patron': 'p-311' };
    		const other = await page.request.fetch(request.url(), { method: 'POST', headers, data: request.postData() ?? undefined });
    		await page.reload({ waitUntil: 'networkidle' });
    		const after = holdLine(await page.evaluate(() => document.body.innerText));
    		return { otherStatus: other.status(), otherBody: (await other.text()).slice(0, 200), before, after, created: holdCount(after) - holdCount(before) };
Every fetch in your components picked oneA row that fetches its own data is where the style’s cost shows up first.

Where it already is in your components

Every data call in your components is one of these. fetch('/api/books?q=') is REST; a tRPC procedure or a SvelteKit or Next.js server function is RPC; an Apollo or urql useQuery with a query string is GraphQL. The component tree often decides the round trips for you: a list whose rows each fetch their own availability is the REST search screen above, four requests for three books.

When you have to own it

The day that list is slow on a phone, you choose where the fix goes. Keep the cacheable part cacheable and batch the live part into one request; or add one operation for the screen; or move the screen to a query and pay for resolvers that batch. The samples show the first, which keeps the style and removes the round trips.

A search list whose rows each fetch their own availability. It reads well, and makes one request per book.

ReactAlready in your code
SearchResults.tsx
// SearchResults.tsx. Each row fetches its own availability, which reads well
// and quietly makes one request per book: three rows, four requests.
import { useEffect, useState } from 'react';

type Book = { id: string; title: string; author: string };

function Availability({ id }: { id: string }) {
	const [text, setText] = useState('…');
	useEffect(() => {
		fetch(`/books/${id}/availability`)
			.then((r) => r.json())
			.then((a: { available: number; total: number }) =>
				setText(`${a.available} of ${a.total} in`)
			);
	}, [id]);
	return <span>{text}</span>;
}

export function SearchResults({ q }: { q: string }) {
	const [books, setBooks] = useState<Book[]>([]);
	useEffect(() => {
		fetch(`/books?q=${encodeURIComponent(q)}`)
			.then((r) => r.json())
			.then(setBooks);
	}, [q]);
	return (
		<ul>
			{books.map((b) => (
				<li key={b.id}>
					{b.title} · {b.author} · <Availability id={b.id} />
				</li>
			))}
		</ul>
	);
}

10 / Make the call

Choose for your callers, then write down what would change it.

One app, one team, screens that map to actions: start with RPC, through your framework’s server functions or tRPC, or with pages rendered on the server, as the plain build did. You accept that a second client will need its own operations.

Several clients, or anyone outside your team, and reads worth caching: REST, with a resource shaped for a screen where round trips hurt. You accept extra requests and fields the screen does not use.

Many screens needing different shapes of a real graph, and a team that can own the schema: GraphQL, with a depth or cost limit and batched resolvers from the first day. You accept no HTTP cache for POSTed queries, and resolver costs you have to watch.

Reconsider when a second client arrives, when requests per screen or reads per request show up in your measurements, or when you need to retire a field and cannot tell who reads it.

Take it with you

Explain it without saying “REST”, “RPC”, or “GraphQL”: “Either the client asks for things by address and the web’s caches help, or it calls actions the server wrote for it, or it asks for exactly the fields it wants and the server works each one out.” Then open the screen you touched last and count its requests.

Keep a note of the decision

Why: <the screens and clients that call this API>
What: <REST, RPC, or GraphQL>, at <where the contract is written down>
Constraint: <round trips per screen, what may be cached, who else calls it>
Fallback: <what you accepted, such as extra requests or no HTTP cache>
Reconsider when: <a second client, a slow screen, a field you cannot retire>

Paste into your next prompt, and fill in the blanks

Use <REST | RPC | GraphQL> for this API, because <the screens and clients>.
Write the choice and the reason in <API.md>.
<Screen> gets what it shows in <n> requests.
A shared cache may keep <stable reads> for <n> seconds. <Live data> is never cached,
so it is not in the same response as anything that is.
Every write that a client may retry takes a key, scoped to the <user> who sent it;
the same key from the same <user> returns the first result.
<GraphQL only:> reject queries deeper than <n>, and batch reads per request.
Connections to follow nextRelated lessons

Take the catalog into your editor. Add /availability?ids= to the REST door and count the search screen’s requests again, in both languages.

Back to architecture →