01 / The prompt
“The weekly report is slow. Make it faster.”
A food bank prints a report every week: each household, how many times it came, and how many pounds of food it took home, heaviest first. With 1,200 households it has started to drag. You ask an agent to speed it up, and it comes back with a clean one-pass rewrite and a message: much faster now. You run it. It is.
Nobody timed the old version before the change, so “much faster” is a feeling with a direction. More importantly, nobody saved what the old version printed. The report is read by volunteers who decide who gets a delivery, so a household that drops off the list, or a total that moves by a tenth of a pound, is not a detail.
The question the prompt never answered: faster than what, on which data, and printing the same thing? Without a baseline taken before the change, nothing after it can answer.
02 / Name the move
Four things, written down before you touch the code.
A baseline is a record of how the system behaves and performs before a change, taken so that the change can be judged against it. It has four parts, and the order matters: the output comes before the timing, because a faster report that prints something else is a different report.
Fix the workload. Snapshot the output. Measure more than once. Decide what you would accept. Then change it, and answer the same question.
| Part | For this report | Without it |
|---|---|---|
| A fixed workload | Week 38: 1,200 households, 6,000 pickups, and a few edge cases. | Before and after ran on different data, and the difference means nothing. |
| An output snapshot | Every line printed, and its fingerprint, 9a75b00e. | A behavior change looks like a speedup. |
| Repeated samples | Fifteen timed runs after three warmups, every sample kept. | One lucky or cold run decides. |
| A rule, decided first | Same fingerprint; then every run after beats every run before. | The rule bends to fit whatever the numbers say. |
Words to put in a prompt or a review
- Baseline
- The before picture: behavior and performance, on a fixed workload.
- Workload
- The exact input both measurements run on.
- Snapshot
- What the code outputs today, saved so tomorrow’s output can be compared.
- Warmup
- Runs thrown away so caches and the JIT settle before timing starts.
- Spread
- How much repeated runs differ. A change smaller than the spread is not a result.
- Acceptance rule
- What would count as better, worse, or the same, written before the change.
Why count steps as well as timeA number the machine cannot move
Wall-clock time depends on the machine, what else it is doing, and the runtime’s mood. A count of the work, here the inner-loop steps, does not: the original takes 7,200,000 on week 38 on every machine, the refactor 6,000. Counts explain why a change is faster; timings say whether it is, on a real machine. Keep both.
03 / Follow one refactor
Timing says yes. The snapshot says look again.
First the refactor judged on timing alone. Then the same refactor with the output snapshot taken first. Then the fix, judged by the same rule. Open Try it to take a real baseline in your browser.
Was the refactor worth it?
- 1Fix the workload
- 2Snapshot the output
- 3Measure, more than once
- 4Apply the rule
Week 38 · 1,200 households, 6,000 pickups
- Steps before
- 7,200,000
- Steps · the agent’s refactor
- 6,000
- Fingerprint before
- 9a75b00e
- Fingerprint · the agent’s refactor
- not taken
Timed runs
Not measured yet.
Verdict: —
“Made the report much faster: one pass instead of a loop inside a loop.”
The agent is right about the work: 7,200,000 inner-loop steps before, 6,000 after, on week 38.
Reduced motion: choose a scene to see its completed state.
Read this scene
The agent is right about the work: 7,200,000 inner-loop steps before, 6,000 after, on week 38.
The agent’s refactor. Steps: 7200000 before, 6000 after. Fingerprint before 9a75b00e, after not taken. Verdict: not decided.
Watch restarts the story when you come back. Step through keeps your step. Try it takes a real baseline in your browser.
04 / Read it in code
A rule, a measurement, and a record.
Basic form is the verdict, written before any number exists. In the wild is the measurement, with warmup and a check that every run printed the same thing. At the call site is the record you keep, with the rule stored next to the numbers so the after-measurement answers the same question.
The rule, decided before the change: a different output fingerprint means the change is judged on behavior, not speed. Then “faster” only if every run after beat every run before; overlap is inconclusive.
export type Measure = { fingerprint: string; lines: number; samples: number[] };
export type Verdict =
| { kind: 'behavior-changed'; missing: number; changed: number }
| { kind: 'faster' | 'slower' | 'inconclusive'; medianRatio: number };
export const median = (samples: readonly number[]) => {
const s = [...samples].sort((a, b) => a - b);
const mid = Math.floor(s.length / 2);
return s.length % 2 ? s[mid] : (s[mid - 1] + s[mid]) / 2;
};
/**
* The rule, decided before the change: the output must be the same, and "faster" means the
* slowest run after the change beat the fastest run before it. Overlap is not a result.
*/
export function verdict(
before: Measure,
after: Measure,
diff: { missing: number; changed: number }
): Verdict {
if (before.fingerprint !== after.fingerprint) return { kind: 'behavior-changed', ...diff };
const medianRatio = Math.round((median(after.samples) / median(before.samples)) * 100) / 100;
if (Math.max(...after.samples) < Math.min(...before.samples))
return { kind: 'faster', medianRatio };
if (Math.min(...after.samples) > Math.max(...before.samples))
return { kind: 'slower', medianRatio };
return { kind: 'inconclusive', medianRatio };
} type Measure struct {
Fingerprint string
Lines int
Samples []float64
}
// Verdict is "behavior-changed", "faster", "slower", or "inconclusive".
type Verdict struct {
Kind string `json:"kind"`
Missing int `json:"missing,omitempty"`
Changed int `json:"changed,omitempty"`
MedianRatio float64 `json:"medianRatio,omitempty"`
}
func Median(samples []float64) float64 {
s := append([]float64(nil), samples...)
sort.Float64s(s)
mid := len(s) / 2
if len(s)%2 == 1 {
return s[mid]
}
return (s[mid-1] + s[mid]) / 2
}
func minMax(s []float64) (float64, float64) {
lo, hi := s[0], s[0]
for _, v := range s {
lo, hi = math.Min(lo, v), math.Max(hi, v)
}
return lo, hi
}
// Decide applies the rule decided before the change: the output must be the same, and
// "faster" means the slowest run after the change beat the fastest run before it.
func Decide(before, after Measure, missing, changed int) Verdict {
if before.Fingerprint != after.Fingerprint {
return Verdict{Kind: "behavior-changed", Missing: missing, Changed: changed}
}
ratio := math.Round(Median(after.Samples)/Median(before.Samples)*100) / 100
beforeMin, beforeMax := minMax(before.Samples)
afterMin, afterMax := minMax(after.Samples)
switch {
case afterMax < beforeMin:
return Verdict{Kind: "faster", MedianRatio: ratio}
case afterMin > beforeMax:
return Verdict{Kind: "slower", MedianRatio: ratio}
}
return Verdict{Kind: "inconclusive", MedianRatio: ratio}
} The behavior these examples promiseChecked by 15 shared cases in TypeScript and Go
- The report lists every household with its visits and pounds, summed in hundredths and rounded once to tenths, heaviest first, ties by name. The agent’s refactor rounds each pickup first and lists only households that came; the fixed refactor keeps the original’s behavior in one pass.
- The fingerprint is FNV-1a over the report’s text; the diff counts households missing and printed differently.
- The verdict: a different fingerprint is “behavior changed”. Otherwise “faster” when the slowest run after is below the fastest run before, “slower” in the mirror case, and “inconclusive” when they overlap.
Every expectation in the shared cases was produced by a separate model written from these
rules, kept beside the examples in examples/model/, not copied from either
implementation: three workloads in three versions each, and six verdicts on given samples,
including overlap and a single run each.
Reading the TypeScriptAn injected clock and integer pounds
measure takes its clock as an option, so the tests can count calls with a
fake one and the lab can use performance.now(). Pounds are integer
hundredths, and rounding is Math.floor((x + 5) / 10), so the three
languages that check this code agree to the last digit.
Reading the GoA function type and time.Now
Report is a function type, so the three versions and the flaky one in the
tests are interchangeable. MeasureReport takes now func() time.Time for the same reason measure takes a clock. Go’s own benchmark harness, go test -bench, repeats runs for you; the lesson writes the loop out so the
rule is visible.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types baseline.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .; add measure to take baseline records too. Both print:
module heyrian.dev/lessons/baseline-before-change
go 1.23
before: 1200 lines, fingerprint 9a75b00e, 7200000 steps after: 1176 lines, fingerprint f64824a6, 6000 steps; 24 missing, 535 changed fixed: 1200 lines, fingerprint 9a75b00e, 7200 steps; 0 missing, 0 changed
05 / Review the agent’s diff
“It is much faster now.”
The diagnosis is right: a loop over every pickup inside a loop over every household is the slow part. Read what the rewrite prints before you decide how to check it.
06 / How it fails
A baseline fails by measuring the wrong thing, or not enough of it.
Each way a before-and-after comparison can mislead, what it would have told you, and what the procedure does instead. The numbers come from the shared cases and the captured run.
| What goes wrong | What it would have said | What the procedure does | From the cases |
|---|---|---|---|
| Timing without an output check | “Much faster.” | Fingerprint first; timing only for the same output. | after: 24 missing, 535 changed |
| One run each | Whatever that run said. | Fifteen runs, all kept. | single run: “faster” |
| The runs overlap | “A bit faster” from the medians. | Inconclusive, and says so. | overlap: median ratio 0.91, “inconclusive” |
| One slow outlier | “No faster.” | Inconclusive under this strict rule; look at the samples. | outlier: “inconclusive” |
| Only the data you tested with | “Same output” on week 38. | Add edge cases: nobody came, rounding, ties, an unknown id. | small workload: 1 missing, 1 changed |
| A cold first run | A slower “before” than is fair. | Warmup runs, thrown away. | tested with a fake clock |
The strict rule, no overlap at all, is a choice, and it is the one this lesson makes because the refactor it cares about changes the work by a factor of a thousand. For a change you expect to move things by five percent, decide a statistical test in advance instead; the point is that the rule exists before the numbers do. Deciding what a user would count as success is Defining success; finding where the time goes is CPU and memory profiling.
07 / Is it worth it?
You pay a few minutes and a file. Here is what they buy.
Skipping the baseline is faster, and most of the time the change is fine. Hold both habits up against the changes this report will get.
| Change | No baseline | Baseline on record |
|---|---|---|
| A second client: a CSV export for the city | A second output nobody snapshotted. | Snapshot it too; the same record covers both. |
| Replace the JSON reader with a streaming one | “Seems fine.” | Same fingerprint and the timings, or it does not merge. |
| Change a rule: round to whole pounds | Indistinguishable from a bug. | The snapshot changes on purpose, and the new one becomes the baseline. |
| A new volunteer maintains it | No idea what “normal” was. | A file that says what it printed, how fast, and on what. |
The technique is itself a measurement plan, so this section’s plan is for the habit: over the next five performance changes you merge, count how many came with a baseline, and how many of those found a behavior change or an unconvincing speedup. Keep the habit where either number is more than zero.
The captured timings in this lesson come from one machine on one afternoon (Apple M3 Max, median 33.473 ms before, 0.456 ms after the fix). They show a distribution; they are not a benchmark.
08 / Ask for it
Two refactors, both correct. One of them can prove it.
We gave two agents, both running Claude Sonnet, the same slow report and week 38’s data, and asked them to make it faster. One prompt stopped there. The other added a paragraph: take a baseline first, record what the report prints, time it with repeated runs, decide what you would accept, measure the same way after, and write it all to BASELINE.md. Then a script ran each refactor beside the original on weeks the agents never saw.
| Workload | Plain prompt | Architecture prompt |
|---|---|---|
| Week 38 (the agents’ own data) | Same output | Same output |
| Week 39 (another week) | Same output | Same output |
| Edge cases (a household that never came, rounding, a tie, an unknown household) | Same output | Same output |
| An empty week | Same output | Same output |
| A week four times larger, median of eight runs | 353 → 75 ms, no overlap | 363 → 75 ms, no overlap |
| A written baseline | None | BASELINE.md |
The result was not the one this lesson’s own example predicts. Both agents found the loop inside a loop, replaced it with one pass, and kept the report exactly as it was, on every workload including the edge cases. Both are about as fast as each other.
The difference is what each left behind. The plain agent did half a baseline without being asked: it saved the output before and after, diffed them, then deleted both files and summed up the speed as “the difference isn’t visible yet” at this size. The architecture agent left a file a reviewer can check: the workload, the output’s hash, twenty-five timed runs, a rule it wrote before the change, and the result against that rule. It even measured Node’s startup, about 37 of its 63 milliseconds, so the numbers would not overstate the report’s own time.
const totals = new Map<string, { hundredths: number; visits: number }>();
for (const pickup of week.pickups) {
const totalsForHousehold = totals.get(pickup.household);
if (totalsForHousehold) {
totalsForHousehold.hundredths += pickup.hundredths;
totalsForHousehold.visits++;
} else {
totals.set(pickup.household, { hundredths: pickup.hundredths, visits: 1 });
}
}
const rows = week.households.map((household) => {
const totalsForHousehold = totals.get(household.id);
const hundredths = totalsForHousehold ? totalsForHousehold.hundredths : 0;
const visits = totalsForHousehold ? totalsForHousehold.visits : 0;
return { name: household.name, visits, tenths: Math.floor((hundredths + 5) / 10) };
});
## The rule, fixed before touching the code
Decided from the baseline's spread alone (stdev ≈ 1.8ms, ≈2.8% of the median), before making
any change:
- **Correctness gate (must pass, non-negotiable):** the after-change command, run against the
same `week-38.json`, must produce stdout with the same sha256 as `baseline-output.txt`. Any
difference is a regression regardless of speed, full stop.
- **Faster** = median wall-clock time over 25 timed runs (3 warmup) drops by at least 15% from
the baseline median, **and** the two runs' noise bands don't overlap
(`after.median + after.stdev < baseline.median − baseline.stdev`). A change that's faster only
within the ~2-3ms noise band doesn't count.
- **Regression** = after-median is not clearly below baseline-median by that margin, or is higher,
or the output hash differs.
## After the change Two things neither agent did are worth a prompt line. Both tested only the week they were given, though the prompt said the report had slowed down as households grew; the size that motivated the change is the size to measure. And neither tried a week with a household that never came or a pickup for an unknown household, which is where the lesson’s refactor broke. The line: take the baseline on the workload that made it slow, plus these edge cases, and keep the record in the pull request.
How the runs were made and checkedOne run each, recorded as written
- Both agents received the prompts word for word, in fresh contexts, in the same message, each in a folder seeded with the same report and data. The only differences were the Architecture paragraph and the folder.
- The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples; the output and timing files they generated are left out, and the architecture agent’s BASELINE.md quotes its numbers. The checker compares every printed line with the original’s on four workloads, then times both on a week four times larger.
- This is one sample of each prompt, not a measurement of a model. One correct plain refactor does not mean plain prompts preserve behavior; the lesson’s own example is the kind that does not.
09 / Hold it there
Make the snapshot a test, and keep the record next to the code.
A baseline taken once protects one change. Three checks make it protect the next one too.
The tools that already repeat runs
Go’s
go test -benchruns a benchmark until the timing is stable and reports time per operation, and-countrepeats it so you can see the spread. In JavaScript, Vitest has abenchmode built on the same idea. Use them for the timing half; they do not check output, so keep the snapshot test beside them.A snapshot test on the workload and the edge cases
Store what the report prints for week 38 and for a handful of edge cases, and fail the build when it changes. When the change is intended, the snapshot is updated in the same pull request, where a reviewer sees it. This lesson’s shared cases are that test for all three versions. The checker in section 08 adds the edge cases below, which both agents’ refactors passed without having tried.
check-runs.mjs const edges = { households: [ { id: 'a', name: 'Alder' }, { id: 'b', name: 'Birch' }, { id: 'c', name: 'Cedar' }, { id: 'd', name: 'Dogwood' }, { id: 'e', name: 'Elm' } ], pickups: [ { household: 'a', hundredths: 1234 }, { household: 'a', hundredths: 1234 }, // rounding once gives 24.7; rounding each gives 24.6 { household: 'c', hundredths: 1000 }, { household: 'd', hundredths: 1000 }, // a tie with Cedar before Cedar's second pickup { household: 'e', hundredths: 1245 }, { household: 'c', hundredths: 5 }, { household: 'x', hundredths: 900 } // a pickup for a household not on the list ] // Birch never came };A record with the rule in it
Keep the baseline record in the repository: workload, fingerprint, every sample, runtime, and the rule. The next change is measured against it, by the same rule, instead of against someone’s memory of how fast it used to be.
bench-typescript.json · before { "version": "before", "workload": "week 38: 1200 households, 6000 pickups", "fingerprint": "9a75b00e", "lines": 1200, "work": 7200000, "medianMs": 33.473, "samplesMs": [ 32.518, 32.965, 33.781, 33.719, 33.9, 33.141, 33.473, 34.019, 33.671, 32.802, 32.952, 33.413, 34.288, 33.774, 33.309 ], "runtime": "node v22.21.1 on darwin-arm64", "rule": "same fingerprint; the slowest run after beats the fastest run before" }
Build UIs?React and the browser already measure renders. A filter box is where you own the baseline.
Where it already is in your components
React’s <Profiler> “lets you measure rendering performance of a React
tree programmatically”, and its onRender callback reports each commit’s actualDuration. It is “disabled in the production build by default” (Profiler), so it measures development builds unless you opt in. In any framework, the browser’s performance.mark and performance.measure put your own timings in the
Performance panel. The numbers are there; the habit of keeping them is the part you add.
When you have to own it
The food bank’s household list has a filter box, and someone wants a faster filter. The baseline is the same keystrokes typed before and after, the rows each keystroke shows, and each update’s duration. The rows matter as much as the time: a faster filter that matches differently is a different feature, exactly like the report that dropped the households who did not come.
Collecting render timings: React’s Profiler, and marks and measures around a Svelte update.
import { Profiler, type ProfilerOnRenderCallback, type ReactNode } from 'react';
// React already measures renders. Keep the numbers instead of eyeballing the page: collect
// every commit's duration while you replay the same interaction, before and after a change.
export const commits: { id: string; phase: string; ms: number }[] = [];
const record: ProfilerOnRenderCallback = (id, phase, actualDuration) => {
commits.push({ id, phase, ms: Math.round(actualDuration * 100) / 100 });
};
export default function Measured({ children }: { children: ReactNode }) {
return (
<Profiler id="pickup-list" onRender={record}>
{children}
</Profiler>
);
}
10 / Make the call
Take a baseline when someone will rely on the answer.
A baseline is overkill for a change you would never describe as “faster”, for code nobody depends on yet, and for a script you run once. Take one whenever a change is justified by speed, whenever an agent rewrites something that other people read the output of, and before any change you would want to undo if it turned out worse.
Retake it when the workload grows, when the runtime or machine changes, and when the output is meant to change, so the next comparison starts from the truth.
Take it with you
Explain it without saying “baseline”: “Before changing it, I saved exactly what it printed for this week’s data, timed it fifteen times, and wrote down what I would call faster. After the change it printed the same thing and every run was quicker.” Then find the last change you merged because it felt faster, and write down what you would have measured.
Paste into your next prompt, and fill in the blanks
Before changing <component>, take a baseline and write it to BASELINE.md: - The workload: <input files or fixtures>, fixed for before and after. - The behavior: exactly what it outputs on that workload (save it, and a hash). Also run <edge cases: empty input, <boundary values>, <odd records>>. - The measurement: <n> timed runs after <k> warmups, every sample kept. - The rule, decided now: the output must be identical; "faster" means <every run after beats every run before / the median drops by <x>%>. After the change, measure the same way and report against the rule. Do not claim a speedup you did not measure.
Connections to follow nextRelated lessons
- CPU and memory profiling finds where the time goes; this lesson decides whether a change to it helped.
- Defining success chooses the numbers a user would care about before you measure them.
- Fitness functions turns a check like the snapshot into one that runs on every change.
- Spec before code is the same move for new code: write down what must be true before it exists.