← Architecture
Client and server Keep a screen current

Polling, server-sent events, or WebSockets

Ask on a timer, listen to a stream, or hold a conversation.

You have watched a deploy log scroll by, or a notification badge tick up without a reload. Some of those pages ask every few seconds, some keep a line open, and all of them eventually lose the connection. Let’s follow one deploy three ways, and pull the plug halfway through.

TypeScriptGoOne deploy log, three transports, one outage, two recorded builds.

01 / The prompt

“Show the deploy log as it happens.”

A deploy of web@a41f2c writes eight lines in five seconds: build started, installing, compiling, tests passed, uploading, switching traffic, health checks, finished. A developer opens the page and wants to watch them arrive. Ask an agent for it and you will get one of three plausible answers:

  • Polling. The page asks GET /deploys/7/log every second.
  • Server-sent events. The page opens an EventSource, and the server pushes each line down one long response.
  • A WebSocket. The page opens a two-way connection, and either side can send.

All three show the log on a good connection. The prompt never said how late a line may be, how many requests are acceptable, whether the page ever sends anything back, or what happens when the laptop’s connection drops for a second and a half in the middle. The last question is the one this lesson is built around.

02 / Name the choice

Ask, listen, or talk.

Polling is the client asking on a timer; each answer is a sample, as fresh as the last tick. Server-sent events are one HTTP response that never ends, carrying events from the server to the page, and the browser’s EventSource reconnects on its own. A WebSocket is a two-way connection; either side can send at any time, and reconnecting is entirely your code.

The transport decides how soon a change can arrive and which way messages can travel. None of them is a record. What the page missed while it was away is recovered only by an id it can hand back: a cursor.

What each transport gives you, and what you still write
QuestionPollingServer-sent eventsWebSocket
Which way messages travelClient asks, server answersServer to clientBoth ways
How late a change can beUp to one interval, plus a round tripOne tripOne trip
Requests while nothing changesOne per intervalNone after the firstNone after the first
After a dropped connectionThe next tick tries againThe browser reconnects and sends Last-Event-IDNothing, until your code reconnects
What you still writeA cursor, so each ask is smallAn id on every event, and a server that replays after itReconnect, backoff, and a resume message

Words to put in a prompt or a review

Staleness
How long a change waits before the screen shows it.
Cursor
The id of the last thing the client has, sent back so the server knows where to resume.
Last-Event-ID
The header a browser sends when an event stream reconnects, holding the last id it saw.
Replay
The server sending what a reconnecting client missed.
Backoff
Waiting longer between reconnect attempts, so a thousand pages do not retry at once.
Live state
What the page says about its connection: live, reconnecting, or offline.
What the browser does for server-sent eventsQuoted from the HTML Standard

The HTML Standard gives every EventSource “a reconnection time, in milliseconds”, which “must initially be an implementation-defined value, probably in the region of a few seconds”; a retry: field from the server changes it. And “the Last-Event-ID HTTP request header reports an EventSource object’s last event ID string to the server when the user agent is to reestablish the connection” (HTML Standard, server-sent events). The last event ID is whatever the server put in an id: field. No id, no header, no resume. WebSockets have no equivalent in their protocol (RFC 6455).

03 / Follow one deploy

Watch the same eight lines arrive six ways.

The top track is the server’s log; the bottom track is the developer’s screen. First polling at two speeds, then a stream, then the same stream through an outage with and without ids, then a WebSocket nobody taught to reconnect. Open Try it to choose the transport and the outage yourself.

Polling, SSE, or WebSockets

When does a line reach the screen?

Polling every 1 s · 0.0 s

Deploy #7

  1. Waiting for the first line

Written but not on screen: 1

01/ 06
Poll every second

The page asks “anything after the last line I have?” once a second.

0.0 s into the deploy.

Reduced motion: choose a scene to see its completed state.

Read this scene

0.0 s into the deploy.

Polling every 1 s, 0.0 s. On screen: nothing yet.

Watch restarts the story when you come back. Step through keeps your step. Try it runs the deploy again for every choice.

04 / Read the three shapes

A timer, a stream, and a cursor.

Basic form is polling with a cursor. In the wild is a stream that can reconnect and resume, the same code for server-sent events and a WebSocket, with the two choices that decide what survives an outage as separate flags. At the call site the TypeScript is the browser following a stream, and the Go is the real server door that replays after Last-Event-ID.

Notice where the cursor lives in each. Polling keeps last in the page. The stream keeps lastId in the page and sends it on reconnect. The Go handler reads it from a header. The transport changes; the need for “what did you last see?” does not.

Polling: ask on a timer. With a cursor, the page asks only for lines after the last id it has; without one it fetches the whole log each time.

TypeScriptReading
deploy-log.ts
// Polling: ask on a timer. With a cursor, ask only for lines after the last
// one seen; without one, fetch the whole log and replace the screen.
export function poll(
	sim: Scheduler,
	net: Network,
	log: DeployLog,
	view: View,
	interval: number,
	cursor: boolean
) {
	let last = 0;
	for (let t = 0; t <= HORIZON; t += interval)
		sim.at(t, () => {
			view.requests++;
			if (!net.up(sim.now)) return; // the request fails; the next tick tries again
			const lines = log.after(cursor ? last : 0, sim.now);
			const body = JSON.stringify(lines);
			sim.at(sim.now + net.latency, () => {
				view.bytes += size(body);
				for (const line of lines) view.receive(line, sim.now, false);
				if (lines.length) last = Math.max(last, lines.at(-1)!.id);
			});
		});
}
GoAlongside
main.go
// Poll asks on a timer. With a cursor, it asks only for lines after the last
// one seen; without one, it fetches the whole log and replaces the screen.
func Poll(sim *Scheduler, net Network, log *DeployLog, view *View, interval int, cursor bool) {
	last := 0
	for t := 0; t <= Horizon; t += interval {
		sim.At(t, func() {
			view.Requests++
			if !net.Up(sim.Now) {
				return // the request fails; the next tick tries again
			}
			from := 0
			if cursor {
				from = last
			}
			lines := log.After(from, sim.Now)
			body := encode(lines)
			sim.At(sim.Now+net.Latency, func() {
				view.Bytes += len(body)
				for _, line := range lines {
					view.Receive(line, sim.Now, false)
				}
				if len(lines) > 0 {
					last = max(last, lines[len(lines)-1].ID)
				}
			})
		})
	}
}
The behavior these examples promiseChecked by 13 shared scenarios
  • Eight lines at 0, 400, 1300, 2300, 2600, 2900, 4200, and 5000 ms. The network takes 50 ms each way; the optional outage runs from 2000 to 3400 ms.
  • Polling asks at 0 and every interval after; a request during the outage fails. With a cursor it asks for lines after the last id it has.
  • A stream closes when the outage starts. With reconnect it tries every second; with resume the server sends every line after the last id the client saw.
  • A note sent at 4.5 s goes over the WebSocket itself, or as one more HTTP request otherwise.

Every expectation in the shared cases, down to the millisecond each line reached the screen and the bytes that carried it, was produced by a separate model written from these rules and kept beside the examples, not copied from either implementation.

Reading the TypeScriptA scheduler instead of a clock

Scheduler runs callbacks in time order, so “50 ms later” is sim.at(sim.now + 50, …) and an outage is a time range. The transports are callback-driven, like EventSource and WebSocket in a browser. The call site uses the real EventSource, whose lastEventId is the id the server sent.

Reading the GoA real handler and its test

EventsHandler sets text/event-stream, replays after the Last-Event-ID header, and flushes each frame with http.NewResponseController. It follows the log through a buffered channel until the request’s context ends. Its test opens a real connection with httptest.NewServer and reads frames as they arrive.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types deploy-log.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/realtime-updates

go 1.23
poll every 1 s: 8 lines, nothing missing, worst delay 850 ms, 7 requests, 424 bytes
poll every 1 s, whole log: 8 lines, nothing missing, worst delay 850 ms, 7 requests, 1365 bytes
server-sent events: 8 lines, nothing missing, worst delay 50 ms, 1 request, 301 bytes
SSE with ids, 1.4 s drop: 8 lines, nothing missing, worst delay 1800 ms, 3 requests, 301 bytes
SSE without ids, 1.4 s drop: 5 lines, missing 4, 5, 6, worst delay 50 ms, 3 requests, 157 bytes
WebSocket, no reconnect, 1.4 s drop: 3 lines, missing 4, 5, 6, 7, 8, worst delay 50 ms, 1 request, 126 bytes

05 / Review the agent’s diff

“Nothing in the page reads the id.”

Removing unused code is good hygiene, and the page’s JavaScript really never mentions the id. Find what does before you decide.

The agent’s pull request

“Cleanup: removed the id field from the event stream. Nothing in the page reads it, and every frame is a few bytes smaller. All 17 tests pass.”

// server/events.ts
			for (const line of log.after(lastEventId)) {
			(removed)  res.write(`id: ${line.id}\ndata: ${line.text}\n\n`);
			(added)  res.write(`data: ${line.text}\n\n`); // the page never reads ids
			}
			
You are reviewing this change. What do you do?

06 / How it fails

Every transport fails at the moment the line goes quiet.

Each row is a shared scenario unless it is marked as authored.

Failure modes of one deploy log
What goes wrongWhat the developer seesWhat decides it
Slow: polling every three secondsLines up to 2,650 ms late, and “finished” not shown by 6 sThe interval. Faster ticks cost more requests.
Unreachable: a 1.4 s outage, pollingTwo failed requests, then every line, the latest 2,750 ms lateThe cursor: the next tick asks for everything after the last id.
Unreachable: the outage, a stream with idsLines 4 to 6, 1.8 s late, once eachReconnect every second, and replay after Last-Event-ID.
Lost: the outage, a stream without idsLines 4, 5, and 6 never appearNo id to send back, so the server resumes from “now”.
Stuck: the outage, a WebSocket without reconnectThe log stops at “Compiling”, and the deploy looks hungA WebSocket does not reconnect by itself.
Duplicated: replaying the whole log on reconnect (authored)Lines 1 to 3 twice, unless the page replaces what it showsReplay after an id, or replace the view; never append a full replay.
Silent: any of the above (authored)A page that looks live and is notWhether the page says “Reconnecting…”.

A reconnect is a retry, and a replayed line is a message delivered twice unless the id stops it. Idempotency and at-least-once and Backpressure and queues cover the two ideas underneath: safe repeats, and a producer faster than its reader.

07 / Is it worth it?

A stream costs a connection. Here is what it buys.

Polling with a cursor is simple, stateless, and survives the outage on its own. Hold the three up against the changes this page will get.

The same four changes, made to each transport
ChangePollingServer-sent eventsWebSocket
A second client: a CLI that tails the logThe same endpoint, in a loopcurl -N reads the stream as textNeeds a WebSocket client and the resume message
Replace a dependency: logs move to a queue serviceA change behind the log. The transport does not care where lines come from.
Change a rule: a line must appear within 1 sAn interval under a second, and a request per page per secondAlready trueAlready true
A second team adds “cancel deploy”One POSTOne POST beside the streamA message on the open connection, if the team owns its protocol

Before choosing or switching, decide what you will measure and the result you would accept:

  • Staleness: the time from a line being written to it being on screen, at the 95th percentile, measured with a timestamp in the event.
  • Requests or open connections per watching page, and what that is at your busiest hour.
  • Lines missing or shown twice after a reconnect: zero is the only acceptable number, so it belongs in a test, not a dashboard.
  • Time to recover after the connection returns.

This page did not run the log for real users, so it gives no production numbers; the counts on the page are the simulation’s own.

08 / Ask for it

Two prompts, two builds, one proxy we could cut.

We sent two agents the same request at the same time, both running Claude Sonnet. One prompt described the page. The other added an Architecture block: choose the transport on purpose and write down why, a line on screen within a second, an id on every line and a resume after a drop with every line exactly once, the full log on a reload, and a live or reconnecting status. A script then opened each build in Chromium through a proxy it could cut, and started the deploy.

What the checker found, run 2026-09-23
QuestionPlain promptArchitecture prompt
How the page follows the logServer-sent events (1 connection)Server-sent events (1 connection)
A deploy, watched from the startAll 8; worst delay 9 msAll 8; worst delay 10 ms
The connection drops from 2.0 s to 3.4 sAll 8; worst delay 2,728 msAll 8; worst delay 1,727 ms
What the page said during the dropNothingReconnecting…
Reloaded at 3.0 sAll 8; worst delay 32 msAll 8; worst delay 40 ms
Opened at 3.5 sAll 8; worst delay 35 msAll 8; worst delay 58 ms

Both agents chose server-sent events, and both builds kept every line through the outage, once each. That is not the failure we expected from the plain build. It got there a different way: on every connection its server sends the whole log so far, and the page throws away what it shows and draws it again.

public/app.js · plain prompt
const es = new EventSource(`/deploys/${deployId}/stream`);

es.addEventListener("sync", (event) => {
  const data = JSON.parse(event.data);
  clearLog();
  data.lines.forEach(appendLine);
  setStatus(data.status);
server.ts · architecture prompt
// Tell the browser to retry quickly if the connection drops.
res.write('retry: 1000\n\n');

// Catch-up: replay everything the client hasn't seen yet, in order.
for (const line of deploy.lines) {
  if (line.id > lastEventId) res.write(formatLineEvent(line));
}

What separated them was recovery. The architecture build told the browser retry: 1000, so it was back about a second after the network was, and it resumed after the last id instead of resending everything. The plain build waited for the browser’s default reconnection time, a few seconds, and said nothing while it waited: for those seconds the page looked live and was not. The architecture build’s page said “Reconnecting…”.

So the line that mattered in the architecture block was not the transport; both picked the same one. It was the status requirement, and the resume rule that keeps a reconnect cheap when the log is long: after a drop, resume after the last id and say you are reconnecting; tell the browser how soon to retry.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. For every question the checker restores a build, starts it fresh, and opens it in Chromium through a TCP proxy that it can cut for exactly 1.4 seconds. A script in the page records the first moment each line is on screen.
  • The checker’s first run, against the plain build alone, measured a late-opened page’s lines from when they were written, so a page opened at 3.5 s looked 3.5 s late. It now measures from when the page could first show them. Both runs are kept.
  • The plain agent stopped its test server with a pattern that matches every process started the same way on the machine; the architecture agent kept a log in the system’s temporary folder and deleted it. Neither touched the other’s folder, and their test servers did not overlap in time.
  • This is one sample of each prompt, not a measurement of a model.

09 / Hold it there

Test the reconnect, because the happy path always passes.

Every transport works while the connection holds, so ordinary tests pass forever. The next change that drops an id, replays too much, or forgets a status will pass them too. Three checks keep it.

  1. The platform’s own behavior

    EventSource reconnects and sends Last-Event-ID by itself, as quoted in section 02, but only if the server sends ids, and only after the reconnection time, which the server sets with retry:. A WebSocket gives you neither; if you choose one, the reconnect loop and the resume message are yours to write and test.

  2. A test that reconnects

    Put the resume in a test at the door: connect with Last-Event-ID and assert the first frame is the next line, then append one and assert it arrives. This is the lesson’s own Go test, over a real connection. Run tests like it in CI, the way Enforcement layer runs its rules.

    main_test.go
    // The real handler, over a real connection: a browser that last saw line 2
    // gets 3 onward, then new lines as they are appended.
    func TestEventsHandlerResumesFromLastEventID(t *testing.T) {
    	log := NewDeployLog()
    	for _, s := range Script[:3] {
    		log.Append(s.Text, s.At)
    	}
    	server := httptest.NewServer(EventsHandler(log))
    	defer server.Close()
    	ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
    	defer cancel()
    	req, _ := http.NewRequestWithContext(ctx, "GET", server.URL, nil)
    	req.Header.Set("Last-Event-ID", "2")
    	res, err := http.DefaultClient.Do(req)
    	if err != nil {
    		t.Fatal(err)
    	}
    	defer res.Body.Close()
    	if ct := res.Header.Get("Content-Type"); ct != "text/event-stream" {
    		t.Errorf("content type %q", ct)
    	}
    	read := bufio.NewReader(res.Body)
  3. A check on what the developer actually sees

    A unit test cannot pull a laptop’s network. A proxy can. The checker in section 08 drops every connection for 1.4 seconds and counts what reaches the screen, and how many times.

    check-runs.mjs
    /** A TCP proxy in front of the build. cut(ms) drops every connection and refuses new ones for ms. */
    function proxy() {
    	const sockets = new Set();
    	let down = false;
    	const server = createServer((client) => {
    		if (down) return client.destroy();
    		const upstream = connect(PORT, '127.0.0.1');
    		sockets.add(client).add(upstream);
    		client.pipe(upstream).pipe(client);
    		const end = () => {
    			client.destroy();
    			upstream.destroy();
    			sockets.delete(client);
    			sockets.delete(upstream);
    		};
    		client.on('error', end).on('close', end);
    		upstream.on('error', end).on('close', end);
    	});
Your refetchInterval is already one of theseEvery badge or log that updates without a reload chose a transport. The reconnect is the part you own.

Where it already is in your components

A TanStack Query refetchInterval, a setInterval around fetch, or SvelteKit’s invalidate on a timer is polling. An EventSource for notifications is server-sent events. A chat or collaborative cursor is usually a WebSocket, often inside a library that wrote the reconnect for you.

When you have to own it

The day the page must survive a dropped connection without losing or repeating anything, and tell the person watching. Give every event an id, render by id so a repeat is harmless, and show whether the connection is live. The samples show polling as it is usually written, then an EventSource that skips ids it already shows and says when it is reconnecting.

A deploy log polled every second for the whole log. Simple, and it survives a drop, at one request per second per open page.

ReactAlready in your code
DeployLog.tsx
// DeployLog.tsx. Polling, as most of us first write it: TanStack Query asks
// again every second, and the component re-renders when the answer changes.
import { useQuery } from '@tanstack/react-query';

type Line = { id: number; text: string };

export function DeployLog({ deployId }: { deployId: string }) {
	const { data: lines = [], isError } = useQuery({
		queryKey: ['deploy-log', deployId],
		queryFn: () => fetch(`/deploys/${deployId}/log`).then((r) => r.json() as Promise<Line[]>),
		refetchInterval: 1000
	});
	return (
		<section>
			{isError && <p role="status">Can’t reach the server. Trying again…</p>}
			<ol>
				{lines.map((line) => (
					<li key={line.id}>{line.text}</li>
				))}
			</ol>
		</section>
	);
}

10 / Make the call

Pick the cheapest transport that meets the freshness you promised.

Changes every few minutes, or a page nobody stares at: poll with a cursor. It survives outages on its own and needs no open connection. You accept a request per interval and a line up to one interval late.

Changes every second, flowing one way: server-sent events with an id on every event and a retry:. You accept one open connection per watching page, and a server that can replay after an id.

The page talks back continuously, such as typing, cursors, or game moves: a WebSocket, with a reconnect loop, backoff, and a resume message you write and test. You accept owning the protocol that the other two borrow from HTTP.

Reconsider when the freshness promise changes, when the page starts sending more than the odd command, or when open connections per page become the thing your servers run out of.

Take it with you

Explain it without saying “polling”, “server-sent events”, or “WebSockets”: “Either the page keeps asking, or it listens on a line that stays open, or both sides talk on one. Whichever it is, the page remembers the last thing it saw, so when the line drops it can ask for what it missed.” Then find the last screen you built that updates by itself, and unplug the network while it runs.

Keep a note of the decision

Why: <what the screen shows, and how fresh it must be>
What: <polling | server-sent events | a WebSocket>, at <which endpoint>
Constraint: <direction of messages, connections per page, outages>
Fallback: <what the page shows and does while reconnecting>
Reconsider when: <the freshness promise, two-way traffic, or connection counts change>

Paste into your next prompt, and fill in the blanks

Use <polling | server-sent events | a WebSocket> for <this screen>, because
<how often it changes, and whether the client sends anything back>.
Write the choice and the reason in <REALTIME.md>.
A change appears on screen within <n> seconds.
Every <event> has an id. After a dropped connection the page resumes after
the last id it showed, and shows every <event> exactly once, in order.
Opening the page mid-way shows everything so far, then continues.
The page says whether it is live or reconnecting.
Connections to follow nextRelated lessons

Take the log into your editor. Give the WebSocket transport a backoff that doubles each attempt, and find the outage length where it first loses to server-sent events.

Back to architecture →