Make “slow” specific enough to investigate.
A lesson catalog filters as someone searches. It works on a small dataset. With more records and a longer session, searches do more work and old result lists accumulate. We have two symptoms to investigate; we don’t yet have one explanation for both.
Profiling attributes resource use to parts of a running program. A timer tells us how long a replay took. A profile helps us investigate what consumed CPU or allocated memory during that work.
- Job
- Search a fixed catalog with the same query sequence.
- Required behavior
- Return every matching ID in input order, including all records for an empty query.
- Starting implementation
- Normalize each document on every search. Keep each result list in a debug history.
- Product constraint
- The view needs current results. Full search history is incidental debugging state, not an Undo feature.
Start with 5,000 generated records and 24 searches. Use the same dataset, query order, build, and environment for both versions. Separate setup from the repeated searches; building an index is work too.
Write down the symptom you want to improve before recording. For a real product, that might be typing-to-result latency or memory still held after closing a view. This lesson isolates search computation and result ownership; it does not measure rendering, requests, or a complete user interaction.
A workload someone else can repeat, the behavior it must preserve, and the resource you are investigating.
Ask one question of the right evidence.
Where is CPU being spent?
Sample active call stacks during the repeated searches. If most elapsed time is waiting, follow that wait with a trace instead.
What keeps getting allocated?
Inspect allocation sites. Many temporary objects can create collection work without remaining alive.
What is still being kept?
Compare equivalent lifecycle points and find the owner of surviving objects. A large heap alone does not establish a leak.
Collect CPU and memory evidence separately so the instruments interfere less with one another. A profile is an observation under a workload, not a verdict on a function. Go’s diagnostics guide explains the distinction and the interference between tools.
Choose the source comparisons and capture recipes throughout this lesson.
Capture the experiment on this page.
- Open DevTools → Performance. Start recording, run the comparison in step 4, then stop.
- Select a replay’s JavaScript task. Inspect the Main track, Bottom-up, and callers; exclude page load and the gaps between runs.
- Record which activity owns the work. In a production build, source maps help you connect generated code to its source.
The experiment includes output checks and warmup before its measured rounds. Those also appear in a recording. Separate setup, replay, and the surrounding interface. See Chrome’s Performance reference.
Capture retained memory in the browserA separate investigation
- Reset the lab and take a heap snapshot in DevTools → Memory.
- Choose “Memory: keep only current results,” run the comparison, then take another snapshot.
- Inspect the two live
SearchSessioninstances and theirhistoryarrays. Both final sessions are held intentionally for inspection. Follow a result array’s retainers back to its session. - Reset, take another snapshot, and check whether the session-owned objects disappear. Repeat at a larger search count to compare equivalent runs, rather than accumulating unrelated page activity.
The baseline keeps one list per search; the candidate keeps the last. This is a controlled retention example, not proof that the page itself leaks. Chrome snapshots trigger collection; console inspection can itself keep objects alive. A snapshot’s retaining path tells you who holds an object. See heap snapshots and retainers.
TypeScript outside the browserNode · CPU and allocation sampling
From this lesson’s examples/ directory, run the driver with Node 22.18 or later.
The source uses supported erasable TypeScript syntax. Capture one resource at a time; open
the CPU artifact in a compatible profile viewer and the heap artifact in DevTools Memory.
# From examples/ · Node 22.18+; run one profiler at a time
node --experimental-strip-types --cpu-prof --cpu-prof-name=baseline.cpuprofile profile.ts baseline
node --experimental-strip-types --cpu-prof --cpu-prof-name=index.cpuprofile profile.ts index
# Allocation sampling, in a separate process
node --experimental-strip-types --heap-prof --heap-prof-name=baseline.heapprofile profile.ts baseline The two-second driver repeats fresh sessions. CPU capture includes startup and warmup: locate the search work. Allocation samples are not a heap snapshot or a retaining-path graph. Compare per-replay work as well as totals; faster versions execute more replays in the same interval. See the Node CPU flags and heap flags.
Copy the Node capture driverprofile.ts · alongside search.ts
// Node 22.18+; run separately with --cpu-prof OR --heap-prof.
import { performance } from 'node:perf_hooks';
import { makeDocuments, makeQueries, modes, replay, SearchSession, type Mode } from './search.ts';
const mode = process.argv[2] ?? 'baseline';
if (!modes.includes(mode as Mode)) throw new Error('Mode: baseline | index | release | both');
const documents = makeDocuments(10000);
const queries = makeQueries(48);
replay(new SearchSession(documents, mode as Mode), queries); // Warmup, still visible to startup profilers.
let held: SearchSession | undefined;
let iterations = 0;
const start = performance.now();
while (performance.now() - start < 2000) {
held = new SearchSession(documents, mode as Mode);
replay(held, queries);
iterations++;
}
console.log({
mode,
iterations,
elapsedMs: performance.now() - start,
finalSession: held?.inspect()
});
// Fresh session per replay: this driver measures steady repeated work, not one endlessly growing history.
// Use the browser snapshot procedure to investigate retaining paths; a heap profile is not a snapshot.
Go: capture, then choose what pprof countsCPU · alloc_space · inuse_space
From examples/go/, use the included benchmark harness. Each operation is
a fresh session plus 48 searches over 10,000 records. The final session stays
referenced so its lifetime is defined in the heap capture.
# From examples/go/ · standard library only
go test -run '^$' -bench '^BenchmarkSearch/baseline$' -benchtime=2s -cpuprofile=baseline.cpu
go tool pprof -top baseline.cpu
# Separate capture: all allocation traffic versus still-live allocations
go test -run '^$' -bench '^BenchmarkSearch/baseline$' -benchtime=2s -memprofile=baseline.heap
go tool pprof -top -alloc_space baseline.heap
go tool pprof -top -inuse_space baseline.heap
# Repeat with /index, /release, or /both and a distinct output filename. alloc_space counts sampled allocation volume, including objects already
collected. inuse_space selects sampled live bytes. Both attribute
allocations to sites; neither is a JavaScript-style retaining-path graph. CPU “flat”
is work in the function; “cum” includes callees. Keep the sample type and workload in
your note. See runtime/pprof.
Copy the Go tests and benchmark harnesssearch_test.go · alongside search.go and go.mod
package main
import (
"encoding/json"
"os"
"reflect"
"testing"
)
func TestCases(t *testing.T) {
data, err := os.ReadFile("../cases.json")
if err != nil {
t.Fatal(err)
}
var cases []struct {
Name string
Documents []Document
Queries []string
Expected [][]int
}
if err := json.Unmarshal(data, &cases); err != nil {
t.Fatal(err)
}
for _, c := range cases {
t.Run(c.Name, func(t *testing.T) {
for _, mode := range []string{"baseline", "index", "release", "both"} {
s := NewSession(c.Documents, mode)
total := 0
for i, q := range c.Queries {
got := s.Search(q)
if !reflect.DeepEqual(got, c.Expected[i]) {
t.Fatalf("%s: %v != %v", mode, got, c.Expected[i])
}
total += len(got)
}
indexed := mode == "index" || mode == "both"
keep := mode == "baseline" || mode == "index"
norm, entries := len(c.Documents)*len(c.Queries), 0
if indexed {
norm, entries = len(c.Documents), len(c.Documents)
}
retained, batches := total, len(c.Queries)
if !keep && batches > 0 {
retained, batches = len(c.Expected[len(c.Expected)-1]), 1
}
want := Stats{norm, len(c.Documents) * len(c.Queries), total, retained, batches, entries}
if s.Inspect() != want {
t.Fatalf("%s: %+v != %+v", mode, s.Inspect(), want)
}
}
})
}
}
func TestSnapshotAndResults(t *testing.T) {
for _, mode := range []string{"baseline", "index", "release", "both"} {
docs := []Document{{4, "Factory"}}
s := NewSession(docs, mode)
docs[0].Text = "Stack"
first := s.Search("factory")
s.Search("stack")
if !reflect.DeepEqual(first, []int{4}) {
t.Fatal(first)
}
}
}
var held *SearchSession // Retain only the final session so inuse_space has a defined lifetime.
func BenchmarkSearch(b *testing.B) {
for _, mode := range []string{"baseline", "index", "release", "both"} {
b.Run(mode, func(b *testing.B) {
docs, queries := MakeDocuments(10000), MakeQueries(48)
b.ReportAllocs()
b.ResetTimer()
for i := 0; i < b.N; i++ {
held = NewSession(docs, mode)
if Replay(held, queries) != 195000 {
b.Fatal("changed result")
}
}
})
}
}
Behavior tests read the shared ../cases.json. The profiling commands skip
those tests with -run '^$'.
A recording tied to one workload and a named metric: active CPU, allocation volume, or live memory.
A big frame is a place to look. Follow its callers.
A function can be prominent because each call is costly, because it is called often, or because its children do the work. Before rewriting it, find which of those explanations fits.
The parent includes its descendants. Adding 100 + 60 + 25 + 5 would count the same samples twice. In a CPU flame graph, width represents accumulated samples; a timeline flame chart also preserves when events happened. Neither width nor color means “bad code.”
Our candidate hypothesis is specific: document text is normalized repeatedly even though the catalog has not changed. A real profile may group or inline that work differently. Connect the hot path to the repeated search before acting on it.
Memory needs another question. Result arrays are created by search, but retained by the session’s debug history. Changing the allocation site is one possibility; changing the owner’s lifetime is another. A counter that says “entries created” cannot tell you how many are still reachable.
A suspected cause, the evidence connecting it to the symptom, and the result you expect from one change.
Make the smallest change your evidence supports.
Try precomputing first. Then run a separate comparison that changes only the history. Finally, combine them. Keeping those experiments apart lets us explain which decision caused which result.
Change one thing. Keep the search the same.
The TypeScript source runs here. Your language preferences change the code you read, not this runtime.
Predict: what work moves into setup, and what new state stays alive?
Choose one change, predict its effect, then run the comparison.
Expect two views of the same workload: elapsed time, and a map of what the session still holds.
Counts describe this implementation’s operations and logical entries, not allocated bytes or a heap measurement. Timings include instrumentation, exclude rendering and data generation, and may overlap or reverse on small workloads. Both final sessions are held for inspection until reset, another run, or navigation.
Follow the change in the sourceOne workload · four modes
export class SearchSession {
private documents: Document[];
private index: string[] | null;
readonly history: number[][] = [];
normalizedDocuments = 0;
scannedDocuments = 0;
createdResultSlots = 0;
private keepHistory: boolean;
constructor(documents: Document[], mode: Mode) {
// Take a snapshot: edits require a new session and index.
this.documents = documents.map((document) => ({ ...document }));
this.keepHistory = mode === 'baseline' || mode === 'index';
this.index =
mode === 'index' || mode === 'both'
? this.documents.map((document) => {
this.normalizedDocuments++;
return normalize(document.text);
})
: null;
}
search(query: string): number[] {
const needle = normalize(query);
const ids: number[] = [];
for (let i = 0; i < this.documents.length; i++) {
this.scannedDocuments++;
let text = this.index?.[i];
if (text === undefined) {
this.normalizedDocuments++;
text = normalize(this.documents[i].text);
}
if (text.includes(needle)) ids.push(this.documents[i].id);
}
this.createdResultSlots += ids.length;
if (!this.keepHistory) this.history.length = 0;
this.history.push(ids);
return ids;
}
inspect() {
return {
normalizedDocuments: this.normalizedDocuments,
scannedDocuments: this.scannedDocuments,
createdResultSlots: this.createdResultSlots,
retainedResultSlots: this.history.reduce((sum, ids) => sum + ids.length, 0),
retainedBatches: this.history.length,
indexEntries: this.index?.length ?? 0
};
}
} type SearchSession struct {
documents []Document
index []string
history [][]int
keepHistory bool
normalizedDocuments, scannedDocuments, createdResultSlots int
}
func NewSession(documents []Document, mode string) *SearchSession {
s := &SearchSession{documents: append([]Document(nil), documents...), keepHistory: mode == "baseline" || mode == "index"}
if mode == "index" || mode == "both" {
s.index = make([]string, len(documents))
for i, document := range s.documents {
s.index[i] = normalize(document.Text)
s.normalizedDocuments++
}
}
return s
}
func (s *SearchSession) Search(query string) []int {
needle := normalize(query)
ids := make([]int, 0)
for i, document := range s.documents {
s.scannedDocuments++
var text string
if s.index != nil {
text = s.index[i]
} else {
text = normalize(document.Text)
s.normalizedDocuments++
}
if strings.Contains(text, needle) {
ids = append(ids, document.ID)
}
}
s.createdResultSlots += len(ids)
if !s.keepHistory {
s.history = nil
}
s.history = append(s.history, ids)
return ids
} baseline repeats normalization and keeps every result list. index changes preparation only. release changes history only. both combines
them. All modes snapshot the input catalog and preserve the same search results.
The counter records document normalizations, including index construction, but not query normalization. The search still visits every document: this is prepared text, not a lookup index that avoids scanning.
export function replay(session: SearchSession, queries: string[]): number {
let matches = 0;
for (const query of queries) matches += session.search(query).length;
return matches;
}
export function example(mode: Mode = 'baseline') {
const session = new SearchSession(makeDocuments(1000), mode);
const matches = replay(session, makeQueries(24));
return { mode, matches, ...session.inspect() };
} func Replay(s *SearchSession, queries []string) int {
matches := 0
for _, query := range queries {
matches += len(s.Search(query))
}
return matches
}
func main() {
for _, mode := range []string{"baseline", "index", "release", "both"} {
s := NewSession(MakeDocuments(1000), mode)
fmt.Println(mode, "matches:", Replay(s, MakeQueries(24)), "stats:", s.Inspect())
}
} The memory change deliberately gives up the debug archive. That is acceptable only because the stated product contract needs current results. If this were an audit log or an Undo history, deleting it would change the feature. We would need a different retention policy.
The CPU change also has a price: setup work, stored strings, and a freshness boundary. This sample snapshots a catalog. A live editable catalog needs an explicit rebuild or update policy; silently searching a stale index is not an optimization.
Language details that affect this investigation
TypeScript: the session holds JavaScript arrays; results are returned by reference and callers must treat them as read-only. Removing a session reference makes objects eligible for collection only if no other owner retains them.
Go: results use slices. Discarding the old history removes its references here; reslicing a larger buffer elsewhere can keep its backing allocation alive. The sample’s copied document structs still refer to immutable string data.
These differences are why logical entry counts are not cross-language byte measurements.
Complete source filesCopyable · no package dependencies
export type Document = { id: number; text: string };
export type Mode = 'baseline' | 'index' | 'release' | 'both';
export const modes: Mode[] = ['baseline', 'index', 'release', 'both'];
export const queryCycle = [
'FACTORY',
' service ',
'stack',
'missing',
'',
'go',
'memory',
'factory'
];
// ASCII-only fixture: no locale, Unicode case-folding, or search ranking contract.
export function normalize(text: string): string {
return text.replace(/^[ \t\r\n]+|[ \t\r\n]+$/g, '').toLowerCase();
}
export class SearchSession {
private documents: Document[];
private index: string[] | null;
readonly history: number[][] = [];
normalizedDocuments = 0;
scannedDocuments = 0;
createdResultSlots = 0;
private keepHistory: boolean;
constructor(documents: Document[], mode: Mode) {
// Take a snapshot: edits require a new session and index.
this.documents = documents.map((document) => ({ ...document }));
this.keepHistory = mode === 'baseline' || mode === 'index';
this.index =
mode === 'index' || mode === 'both'
? this.documents.map((document) => {
this.normalizedDocuments++;
return normalize(document.text);
})
: null;
}
search(query: string): number[] {
const needle = normalize(query);
const ids: number[] = [];
for (let i = 0; i < this.documents.length; i++) {
this.scannedDocuments++;
let text = this.index?.[i];
if (text === undefined) {
this.normalizedDocuments++;
text = normalize(this.documents[i].text);
}
if (text.includes(needle)) ids.push(this.documents[i].id);
}
this.createdResultSlots += ids.length;
if (!this.keepHistory) this.history.length = 0;
this.history.push(ids);
return ids;
}
inspect() {
return {
normalizedDocuments: this.normalizedDocuments,
scannedDocuments: this.scannedDocuments,
createdResultSlots: this.createdResultSlots,
retainedResultSlots: this.history.reduce((sum, ids) => sum + ids.length, 0),
retainedBatches: this.history.length,
indexEntries: this.index?.length ?? 0
};
}
}
export function makeDocuments(count: number): Document[] {
const topics = [
'Factory service TypeScript',
'Stack memory Go',
'Graph service Rust',
'Queue memory Go'
];
return Array.from({ length: count }, (_, id) => ({
id,
text: ` ${topics[id % topics.length]} lesson ${id} `
}));
}
export function makeQueries(count: number): string[] {
return Array.from({ length: count }, (_, i) => queryCycle[i % queryCycle.length]);
}
export function replay(session: SearchSession, queries: string[]): number {
let matches = 0;
for (const query of queries) matches += session.search(query).length;
return matches;
}
export function example(mode: Mode = 'baseline') {
const session = new SearchSession(makeDocuments(1000), mode);
const matches = replay(session, makeQueries(24));
return { mode, matches, ...session.inspect() };
}
console.log(example());
package main
import (
"fmt"
"strings"
)
type Document struct {
ID int `json:"id"`
Text string `json:"text"`
}
func normalize(text string) string { return strings.ToLower(strings.Trim(text, " \t\r\n")) }
type SearchSession struct {
documents []Document
index []string
history [][]int
keepHistory bool
normalizedDocuments, scannedDocuments, createdResultSlots int
}
func NewSession(documents []Document, mode string) *SearchSession {
s := &SearchSession{documents: append([]Document(nil), documents...), keepHistory: mode == "baseline" || mode == "index"}
if mode == "index" || mode == "both" {
s.index = make([]string, len(documents))
for i, document := range s.documents {
s.index[i] = normalize(document.Text)
s.normalizedDocuments++
}
}
return s
}
func (s *SearchSession) Search(query string) []int {
needle := normalize(query)
ids := make([]int, 0)
for i, document := range s.documents {
s.scannedDocuments++
var text string
if s.index != nil {
text = s.index[i]
} else {
text = normalize(document.Text)
s.normalizedDocuments++
}
if strings.Contains(text, needle) {
ids = append(ids, document.ID)
}
}
s.createdResultSlots += len(ids)
if !s.keepHistory {
s.history = nil
}
s.history = append(s.history, ids)
return ids
}
type Stats struct{ NormalizedDocuments, ScannedDocuments, CreatedResultSlots, RetainedResultSlots, RetainedBatches, IndexEntries int }
func (s *SearchSession) Inspect() Stats {
retained := 0
for _, ids := range s.history {
retained += len(ids)
}
return Stats{s.normalizedDocuments, s.scannedDocuments, s.createdResultSlots, retained, len(s.history), len(s.index)}
}
func MakeDocuments(count int) []Document {
topics := []string{"Factory service TypeScript", "Stack memory Go", "Graph service Rust", "Queue memory Go"}
docs := make([]Document, count)
for i := range docs {
docs[i] = Document{i, fmt.Sprintf(" %s lesson %d ", topics[i%len(topics)], i)}
}
return docs
}
func MakeQueries(count int) []string {
cycle := []string{"FACTORY", " service ", "stack", "missing", "", "go", "memory", "factory"}
queries := make([]string, count)
for i := range queries {
queries[i] = cycle[i%len(cycle)]
}
return queries
}
func Replay(s *SearchSession, queries []string) int {
matches := 0
for _, query := range queries {
matches += len(s.Search(query))
}
return matches
}
func main() {
for _, mode := range []string{"baseline", "index", "release", "both"} {
s := NewSession(MakeDocuments(1000), mode)
fmt.Println(mode, "matches:", Replay(s, MakeQueries(24)), "stats:", s.Inspect())
}
}
The fixture is ASCII substring search with whitespace trimming and stable input order. It is not locale-aware search, ranking, networking, or a framework renderer. The runtime drivers use larger fixed workloads to collect samples; do not compare their elapsed times as a language ranking.
An isolated change, preserved outputs, and an account of the work or lifetime it changed.
Check the improvement. Keep the explanation.
Run the same behavior cases before trusting a timing. Then repeat the same workload with the profiler off for timing, and capture again to check whether the suspected hot path or retained objects changed.
The browser lab checks every query’s returned IDs, warms each version once, alternates order, and reports three runs including setup. This is a small local experiment. Three runs cannot establish a production percentile or a reliable speedup on every device. If the ranges overlap, say so and gather better evidence.
Before you call it fixed
- Behavior: every query returns the same IDs in the same order, including empty and missing searches.
- Resource: compare the same sample type and workload. Fewer total allocations in a shorter run is not a per-operation improvement.
- Lifetime: repeat equivalent open → use → close cycles. Compare surviving state after cleanup; don't confuse heap size with process RSS.
- Trade-off: include index construction, retained strings, and freshness. Measure the interaction or endpoint again where the symptom was reported.
Build UIs?See where this shows up in your components.
When the search is inside a frontend
If typing still stutters after filtering gets cheaper, inspect the rest of the interaction. A browser trace can reveal rendering, style, layout, painting, or other scripting work. The lab intentionally leaves these out. Optimizing a pure function does not establish that a page feels faster.
For memory, use a repeatable lifecycle: open a panel, interact, close it, then inspect what remains. A subscription, listener, timer, cache, or debug store can outlive the view. An allocation stack tells you where a result was created; a retaining path helps explain why it survives.
Allocation churn, sustained retention, and a large but stable working set are different investigations. Chrome’s memory guide covers these distinctions. Start from the user’s symptom, then choose the tool.
What would you investigate next?
Choose the evidence you need before choosing a fix.
A request takes 900 ms end to end. Its CPU samples account for little activity. You have not inspected the network or database spans.
Leave evidence the next engineer can follow.
- What
- The symptom, fixed workload, build, runtime, and machine.
- Why
- The profile observation and the cause it made you suspect.
- Change
- One intervention and the behavior it preserves or intentionally changes.
- Evidence
- Before/after captures, repeated timings, checks, and what remains uncertain.
- Limits
- Setup cost, lifetime, omitted work, and the next condition to test.
Try writing the conclusion without “it’s faster”: what stopped happening, what stopped being retained, and what did you pay for that change?
This technique connects measurement to design judgment. The ownership question will feel familiar from Composition over inheritance. The distinction between useful history and accidental retention connects to Stack’s undo history. Profiling supplies evidence for the decision; the application still decides which behavior matters.