01 / The prompt
“Move the sheet-music API to functions. It’s idle most of the day.”
A community choir keeps its sheet music in a small web service. A member asks for their part transposed into their key, and gets a PDF back in about two seconds. Three things happen to it:
- Rehearsal night. At 7 pm forty singers ask for their part at once. The rest of the week, almost nothing.
- The free plan. Twenty transpositions per member per day; the twenty-first answers 429.
- The rehearsal pack. Once a month the director exports every part for the month’s music. It takes 18 minutes.
The request is reasonable: the traffic is bursty and mostly idle, which is what functions are for. An agent given it will answer one of three ways, and all three work in a demo with one singer:
- The same code, moved to functions as asked: each route becomes a function, and everything else stays where it was.
- Keep the long-running server, because it already works and the idle hours cost little.
- Functions for the requests, with the quota counted in a shared store and the rehearsal pack handed to a queue and a worker.
The first changes two things the code quietly relied on. The quota counter lived in the process’s memory, and there is no longer one process. The export ran inside the request, and a function has a time limit shorter than 18 minutes. The second keeps both working until the day a second copy of the server runs. The third costs a store and a queue.
The prompt never asked: where does the state live, and how long may one piece of work run, once there is more than one process?
02 / Name the choice
Two runtimes, three questions.
A long-running server is a process you start that stays up and handles many requests, several at a time, sharing one memory until the next deploy restarts it. A function is code the platform runs per request, in execution environments it starts and stops: one request at a time per environment, a new environment (a cold start) when none is free, idle ones reclaimed, and a hard time limit on each invocation.
Answer three questions, in this order. Can the work take longer than a request can wait? Then it goes to a queue and a worker, on either runtime. Do requests have to agree on something they count or remember? Then it lives in a shared store, on either runtime. Only then choose the runtime, from the traffic.
| Long-running server | Functions | |
|---|---|---|
| Forty requests at once | Queue for its slots, unless you run more servers | More environments start, each with a cold start |
| The first request after a quiet spell | Warm | Often cold: idle environments are reclaimed |
| Process memory | One per server, until a restart or a second server | One per environment, gone when it is reclaimed |
| Work that runs for 18 minutes | Runs, but the front door stopped waiting long before | Stopped at the function’s time limit |
| Cost when idle | The server, all week | Close to nothing; you pay per invocation |
| What you operate | Processes, scaling, health checks, restarts | Configuration and limits; the platform owns the rest |
Words to put in a prompt or a review
- Execution environment
- One running copy of a function, with its own memory. It handles one request at a time.
- Cold start
- Starting a new environment before the handler runs. The first request waits for it.
- Function timeout
- The longest one invocation may run before the platform stops it.
- Front door timeout
- How long the load balancer or gateway waits for an answer before it replies 504.
- Queue and worker
- The request records the work and answers 202; a separate process does it later.
- Shared store
- Redis, a database row, a platform key-value store: somewhere every copy sees the same value.
What the platforms actually promiseAWS Lambda and an AWS load balancer, as examples
AWS Lambda’s function timeout is “900 seconds (15 minutes)” (Lambda quotas, fetched 23 September 2026). Its docs say a cold start’s duration “varies from under 100 ms to over 1 second”, and that “You should not assume that the execution environment will persist indefinitely” (Lambda execution environment, fetched 23 September 2026). For an Application Load Balancer’s idle timeout, “The default is 60 seconds” (Application Load Balancers, fetched 23 September 2026).
The lesson’s model uses the 15-minute limit and the 60-second front door. Its 1.2-second cold start and 10-minute idle reclaim are stated inputs, not measurements of any platform. Other platforms set different numbers; the order is what matters: a front door that gives up before the function limit, and a monthly job longer than both.
03 / Follow one evening
Watch each runtime pay for a different evening.
The same service on a server and on functions, through four evenings: rehearsal night, a quiet evening, ro spending her quota, and the monthly pack. Then the same two runtimes with the quota in a shared store and the pack on a queue. Step through, or open Try it and run any evening on any build.
Which runtime pays for which evening?
Forty singers at 7 pm
Server, quota in memory, export inline
- Slowest reply
- 20.0 s
- Cold starts
- 0
- Allowed / refused
- 40 / 0
Functions, quota in memory, export inline
- Slowest reply
- 3.2 s
- Cold starts
- 40
- Allowed / refused
- 40 / 0
Forty singers ask for their part at 7 pm.
The requests arrive.
Reduced motion: choose a scene to see its completed state.
Read this scene
The requests arrive.
Rehearsal night. Forty singers at 7 pm. Server, quota in memory, export inline: slowest 20000 ms, 0 cold starts. Functions, quota in memory, export inline: slowest 3200 ms, 40 cold starts.
Watch restarts the story when you come back. Step through keeps your step. Try it runs any workload on any build.
04 / Read the shape
The plan comes first. The runtime comes last.
Basic form is the plan, from three facts about the work. In the wild is the service on each runtime, with one line that decides where the quota lives. At the call site the front door sends the same requests to either.
The plan, in order: work longer than a request can wait goes to a queue and a worker; state that requests must agree on goes to a shared store; only then does the traffic shape pick the runtime.
export interface Job {
/** How long the work takes at worst. */
worstMs: number;
/** Whether requests must agree on something they count or remember. */
sharedState: boolean;
/** How many arrive at once, compared with the server's slots. */
burst: number;
}
/**
* The runtime question has three parts, and only one of them is “server or
* function”. Work longer than a request can wait goes to a queue and a
* worker, whichever runtime answers the request. State that requests must
* agree on goes to a shared store, whichever runtime holds the process.
* Only then does the traffic shape pick the runtime.
*/
export function plan(job: Job) {
return {
completion:
job.worstMs > FRONT_MS ? ('queue and worker' as const) : ('in the request' as const),
state: job.sharedState ? ('shared store' as const) : ('process memory is fine' as const),
runtime: job.burst > SLOTS ? ('function' as const) : ('either' as const)
};
} type Job struct {
WorstMs int
SharedState bool
Burst int
}
type Plan struct{ Completion, State, Runtime string }
// PlanFor answers the three parts of the runtime question in order: long
// work goes to a queue, shared state goes to a store, and only then does the
// traffic shape pick the runtime.
func PlanFor(j Job) Plan {
p := Plan{"in the request", "process memory is fine", "either"}
if j.WorstMs > FrontMs {
p.Completion = "queue and worker"
}
if j.SharedState {
p.State = "shared store"
}
if j.Burst > Slots {
p.Runtime = "function"
}
return p
} The behavior these examples promiseChecked by 6 shared scenarios for four builds in TypeScript and Go
- A server has four slots. A request takes the slot that frees first and waits for it; its memory is shared by every request until a deploy restarts it.
- A function platform gives each request an idle environment, or starts one with a 1,200 ms cold start. Each environment has its own memory. One idle for 10 minutes is reclaimed; a deploy replaces them all.
- A function invocation longer than 900,000 ms is stopped. The front door answers 504 after 60,000 ms, whatever is still running behind it.
- Unguarded, the quota counts in process memory and the export runs in the request. Guarded, the quota counts in a store every copy shares, and the export is enqueued in 50 ms and done by one worker.
Every expected result was produced by a separate model written from these rules, kept
beside the examples in model/cases.py, not copied from either implementation.
Reading the TypeScriptPrivate fields and one line that decides
#instances and #slots are private class fields, so nothing
outside Service can peek at a runtime’s memory. The line const counts = this.guarded ? this.#shared : (instance?.counts ?? this.#serverCounts) is the whole hazard: three places a count can live, and only one of them is shared.
Reading the GoA pointer for the environment and a map per copy
place returns a *instance, nil on a server, so
the same code path handles both runtimes. Each instance owns its own counts map, which is exactly what a function environment’s memory is.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types sheet-music.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:
module heyrian.example/servers-vs-functions
go 1.22
server: 40 at once, slowest 20000 ms · ro gets 20 of 25 · export gateway-timeout function: 40 at once, slowest 3200 ms · ro gets 25 of 25 · export gateway-timeout function (guarded): 40 at once, slowest 3200 ms · ro gets 20 of 25 · export accepted
Moving to functions fixed the burst and broke the quota. Only the guarded build gets all three right, and the runtime is not what fixed the export.
05 / Review the agent’s diff
“Drop the Redis call from the hot path.”
A store call on every transposition looks like waste, and the quota test still passes after this change. Read where the new counter lives.
06 / How it fails
Each runtime fails where it hides a cost.
Here is each way the service can go wrong, what a singer or the director sees, and what the guarded plan does.
| What goes wrong | What they see | What the guarded plan does | Backed by |
|---|---|---|---|
| Burst on a server | The last singers wait 20 s for four slots. | Runs on functions, or sizes the servers for 7 pm. | Case “forty singers at 7 pm” |
| Cold after a quiet spell | 3.2 s instead of 2 s, on functions. | Accepts it: a choir can wait a second. | Case “a request after a quiet quarter hour” |
| State per environment | ro gets 25 of her 20. | Counts in a shared store with an atomic increment. | Case “one member sends 25, five at a time” |
| State reset by a deploy | ro gets 5 more after a release, on either runtime. | The store outlives deploys. | Case “20 transpositions, a deploy, then 5 more” |
| Work outlives the request | A 504 at 60 s. The server finishes the pack unseen; the function is stopped at 15 minutes and never does. | Enqueues it; a worker finishes it; a status route says when. | Case “the rehearsal-pack export” |
| The shared store is down | No transpositions at all. | Nothing more; that is the price of one count everybody agrees on. Decide whether to fail open or closed. | Authored |
Work queues and background jobs covers the worker side of the export, and Containing failure what to do when the store the quota depends on is slow or down.
07 / Is it worth it?
A store and a queue cost two moving parts.
The guarded plan adds a network call to every transposition and a worker to run and watch. Hold the plain service on each runtime and the guarded one up against the changes this service will get.
| Change | Server, memory | Functions, memory | Guarded, either runtime |
|---|---|---|---|
| A second choir joins; rehearsal night doubles | A second server, and now two quota counters. | Works; the quota was already wrong. | Works; one counter. |
| Move from functions to a server, or back | The quota and the export change behavior with the move. | Only latency and cost change. | |
| A second long job: a yearly archive | Another 504 nobody sees finish. | Another job stopped at 15 minutes. | A second job type on the same queue. |
| Deploy three times a day | Quota resets three times a day. | Quota resets three times a day. | No change. |
Before moving a service between runtimes, decide what you will look at:
- Latency at the 95th percentile, per route, split by whether the request waited for a cold start. On functions, the cold ones are the number to watch.
- Requests allowed per member per day, read from the access log and compared with the limit. The accepted number over the limit is zero.
- 504s at the front door, and jobs enqueued against jobs finished. The accepted difference is zero; a 504 on a route that should be quick points at work that belongs on a queue.
- Cost per month at the real traffic, including the store and the worker.
This page did not measure a real service, so it has no numbers to give you. The model counts waits and outcomes from stated inputs; it does not time a platform.
08 / Ask for it
Two prompts, one platform, the same two builds.
We sent two agents the same request for this service at the same time, both running Claude Sonnet. Each folder held a small local function platform, with a README that said how it behaves: an environment per concurrent request with its own memory, cold starts, reclaims, a 15-minute limit, a 60-second front door, a shared key-value store, and a queue whose workers have no time limit. The shared prompt described the service and nothing about memory or queues. The architecture prompt added the quota in the store and the pack on the queue. Then a script ran each build on a fresh platform with its clock scaled down, and did the same to a control build we wrote with the quota in module memory and the pack in the request.
| What the checker did | Plain prompt | Architecture prompt | Control (not an agent) |
|---|---|---|---|
| Forty singers ask at once | 40 of 40 got their part; 40 cold starts | 40 of 40 got their part; 40 cold starts | 40 of 40 got their part; 40 cold starts |
| ro sends 25, five at a time | ro got 20 of 25, refused 5 | ro got 20 of 25, refused 5 | ro got 25 of 25, refused 0 |
| ro uses 20, a deploy, then 5 more | after the deploy ro got 0 more of 5 (she had used 20) | after the deploy ro got 0 more of 5 (she had used 20) | after the deploy ro got 5 more of 5 (she had used 20) |
| The director asks for the pack | answered 202 in 0.1 s; the director got the pack after 4.3 s | answered 202 in 0.1 s; the director got the pack after 4.3 s | answered 504 in 1.5 s; the pack never finished |
The two agents built the same design. The plain agent was never told about memory or queues, but the platform’s README said that “anything kept in module memory belongs to that environment only”, and gave both limits. It read that, counted the quota with the store’s atomic increment, and enqueued the pack with a comment that it would outlive both limits. The failure the story is built on happened only in the control, which is there to show the checker can see it.
That is the same lesson the rendering runs taught: an agent told how the runtime behaves designs for it. The plain prompt worked because the platform’s documentation did the architecture prompt’s job. Your platform’s documentation is not in your agent’s folder unless you put it there, and a Next.js route handler looks the same whether it will run once on a laptop or in forty environments on rehearsal night.
const jobId = await queue.enqueue('rehearsal-pack', { choir, month });
return new Response(JSON.stringify({ jobId, status: 'queued' }), {
status: 202,
headers: { 'content-type': 'application/json' }
}); const quotaKey = `transpose-quota:${member}:${day}`;
const used = await kv.incr(quotaKey, ONE_DAY_SECONDS);
if (used > DAILY_LIMIT) {
return new Response('Daily transposition limit reached (20/day on the free plan)', { status: 429 });
} So the line worth adding is the one about the runtime, not the design: each function environment has its own memory and a deploy starts fresh ones; functions are stopped at N seconds and the front door waits M. With that written down, both agents chose the store and the queue themselves.
How the runs were made and checkedOne run each, one checker run
- Both agents received the prompts word for word, in fresh contexts, at the same time. The only differences were the Architecture block and the output folder. The shared part told them to stop any process by its PID; neither used a pattern kill, and nothing was left running.
- Both looked outside their folders first, for a
platform/folder in the repository or on the disk, and found nothing they used. The plain agent listed its folder’s parent, where we had left the prompt files, and so saw the name of the architecture prompt’s file. It did not open it. That was our mistake in setting up the runs, and is recorded in the run notes. - The checker scales the platform’s clock so a run takes seconds: a 100 ms cold start, a 1.5 s front door, a 3 s function limit, and a 4 s pack, in the same order as the real ones. Forty cold starts for forty simultaneous singers is the platform, not the builds.
- One run of each prompt is a sample, not a measurement of a model.
09 / Hold it there
Make module state and long work in a function hard to ship.
Both hazards are one line: a new Map() at the top of a route file, or an import of
the long job into a handler. Three kinds of check catch them.
The platform’s own limit, written in the route
Next.js lets a route set
export const maxDuration = 5, “the maximum execution time (in seconds) for server-side logic in a route segment”, and says deployment platforms can use it “to add specific execution limits” (maxDuration, Next.js 16.3.6, fetched 23 September 2026). SvelteKit’s@sveltejs/adapter-vercel6.3.4 takesexport const config = { maxDuration }per route, typed as the “Maximum execution duration (in seconds) that will be allowed for the Serverless Function” in its own type definitions. A short limit on a quick route turns a slow export into a fast, visible failure instead of a long, invisible one.Rules a check enforces
Only a worker may import the pack renderer. Run against a small fixture, dependency-cruiser flagged the function that imported it and passed the one that enqueued it. It cannot see a top-level
Map, so a second check of a dozen lines reads the function files and flagged the one that kept a counter in module memory. Architecture as rules covers turning a sentence like these into a check..dependency-cruiser.cjs // .dependency-cruiser.cjs module.exports = { forbidden: [ { name: 'no-long-work-in-functions', comment: 'The rehearsal pack runs for 18 minutes. Only a worker may run it.', severity: 'error', from: { path: '^src/functions/' }, to: { path: '^src/lib/rehearsal-pack\\.mjs$' } } ] };npx depcruise src error no-long-work-in-functions: src/functions/rehearsal-pack.mjs → src/lib/rehearsal-pack.mjs x 1 dependency violations (1 errors, 0 warnings). 8 modules, 5 dependencies cruised.check-function-state.mjs // check-function-state.mjs // A function's module memory belongs to one environment. Nothing in // src/functions/ may keep mutable state at the top level of a module. for (const name of readdirSync('src/functions')) { const file = `src/functions/${name}`; readFileSync(file, 'utf8').split('\n').forEach((line, i) => { if (/^(let|var)\s/.test(line) || /^(export\s+)?const\s+\w+\s*=\s*new\s+(Map|Set|Array)\b/.test(line)) problems.push(`${file}:${i + 1}: module-level state: ${line.trim()}`); }); }node check-function-state.mjs error src/functions/transpose-local.mjs:2: module-level state: const used = new Map(); // cheaper than a store call 1 problem in function modules.A check on what actually happens
Source checks miss a counter hidden in a helper module the function imports. So send the requests: five at a time from one member, against a platform that starts a fresh environment per concurrent request, and fail if more than 20 get through. The lesson’s spec does it against the model; the checker in section 08 does it to the recorded builds.
check-runs.mjs async 'ro sends 25, five at a time (limit 20)'() { const results = []; for (let wave = 0; wave < 5; wave++) results.push(...(await Promise.all(Array.from({ length: 5 }, () => transpose('ro'))))); const { environments } = await stats(); return { results, environments, verdict: `ro got ${count(results, 'ok')} of 25, refused ${count(results, '429')}` }; },
Where this lives in Next.js and SvelteKitEvery route handler already runs as a function on Vercel or Netlify. The quota and the export are where it shows.
Where it already is in your components
Deploy a Next.js app or a SvelteKit app with a serverless adapter, and every route.ts and +server.ts runs as a function. A module-level const cache = new Map() in one of them works in dev, where there
is one process, and becomes one cache per environment in production. That is fine for a
cache any copy can rebuild, and wrong for a count every request must agree on.
When you have to own it
The first route that counts something across requests, or that can run for longer than a
minute. The count goes to a store with an atomic increment. The long work goes to a queue;
the route answers 202 with a job id. Each route states its own maxDuration, so a slow one fails fast.
The transposition route: short and bursty, so a function fits. The daily quota is an atomic increment in a shared store, and the route sets maxDuration to 10 seconds.
// app/api/parts/[id]/transpose/route.ts · Next.js route handler
// A short request that bursts on rehearsal night: a function is a good fit.
// The quota lives in a shared store, because each function environment has
// its own memory and a counter there would allow 20 per environment.
export const maxDuration = 10; // seconds; a transposition takes about two
interface QuotaStore {
/** Atomically add one to `key` and return the new count; the key expires after `ttlSeconds`. */
incr(key: string, ttlSeconds: number): Promise<number>;
}
declare const quota: QuotaStore; // Redis, a database row, or your platform's KV
declare function transposePart(id: string, key: string): Promise<Blob>;
export async function POST(
request: Request,
{ params }: { params: Promise<{ id: string }> }
): Promise<Response> {
const member = request.headers.get('x-member-id');
if (!member) return new Response('Sign in first', { status: 401 });
const today = new Date().toISOString().slice(0, 10);
if ((await quota.incr(`transpose:${member}:${today}`, 86_400)) > 20)
return new Response('Daily limit reached', { status: 429 });
const { key } = (await request.json()) as { key: string };
const pdf = await transposePart((await params).id, key);
return new Response(pdf, { headers: { 'content-type': 'application/pdf' } });
}
10 / Make the call
Queue the long work, share the state, then pick the runtime.
For the choir I would run the transposition route as a function, count the quota in a shared store, and send the rehearsal pack to a queue with a worker. I accept a cold start of about a second on the first request after a quiet spell, and a store call on every transposition. I would reconsider when traffic becomes steady enough that a small server is cheaper than the invocations, when the first request of the evening must be fast (keep some environments warm, or run a server), or when the service needs long-lived connections, such as live cursors in a shared score.
A server alone is fine when the traffic is steady, there is one process, and nothing runs longer than the front door waits. The day there are two processes, the store is needed anyway.
Keep a note of the decision
- Why
- Bursty, mostly idle traffic; a per-member daily limit; one job a month that runs for 18 minutes.
- What
- Functions for the requests, the quota in a shared store with an atomic increment, and the pack on a queue with a worker and a status route.
- Constraint
- Each function environment has its own memory and is replaced on deploy and when idle; a function stops at 15 minutes; the front door waits 60 seconds.
- Fallback
- A cold start on the first request after a quiet spell; a store call per transposition; if the store is down, no transpositions.
- Reconsider when
- Traffic becomes steady enough that a server is cheaper, the first request must be fast, or the service needs long-lived connections.
Take it with you
Explain it without saying “serverless”: “Our code can run in one program that stays up, or in many short-lived copies started per request. Before choosing, we decide where anything we remember between requests is kept, and where work that takes longer than a request goes.” Then find a counter, a cache, or a timer your own server keeps in memory, and ask what happens to it when two copies run.
Paste into your next prompt, and fill in the blanks
Runtime: <functions | a long-running server>, because <traffic shape>. Nothing that must hold across requests lives in process memory: each function environment has its own, and a deploy starts fresh ones. <Counts, limits, sessions> live in <shared store>, updated atomically. Work that can take longer than <front door timeout> is enqueued with <queue>; the request answers 202 with a job id, a worker does the work, and <status route> says whether it is queued, running, or done. Each function sets its own time limit (<seconds>) and does one short thing.
Connections to follow nextRelated lessons
- Work queues and background jobs is the worker behind the rehearsal pack: retries, idempotency, and knowing it finished.
- Server, static, or client rendering is the question before this one: which pages need a server at all.
- Caching and invalidation is when process memory is a fine place for something, and when it is not.
- Expand and contract is the next thing a deploy does to you: two versions running at once.