← Math in Practice
Concept Engineering economics and measurement

Cache hit rate and break-even cost

A cache hit is useful only when the work it avoids is worth more than the work the cache adds.

A product catalog cache reports an 85% hit rate after launch. The database CPU graph has barely moved, and a teammate proposes increasing the TTL until the dashboard looks better. Before tuning it, the team needs to know what a hit avoids, what every lookup costs, and how updates keep cached values correct.

The judgment to keep

Translate hits and misses into a named cost over the same interval. A hit-rate percentage is a traffic description; break-even depends on the work saved, lookup overhead, writes, and costs this simple model does not count.

TypeScriptGo Hit rate · explicit cost model · break-even threshold · writes and invalidation
01 / Read the cache claim

The dashboard counts hits. The decision needs avoided work.

Imagine an API serving catalog reads. Across a measured one-second interval it handles 10,000 reads and 800 writes. A representative origin read uses 8 ms of database CPU. A cache lookup uses 0.2 ms of application CPU, and processing one write's invalidation event uses 0.5 ms of CPU. The measured read hit rate is 85%. These numbers are illustrative workload inputs, not benchmark claims.

We will compare modeled CPU time in CPU-ms per second. The unit lets us add work rates, but it combines work on different machines into a load estimate; it is not a dollar bill or a latency promise. A production comparison should use like-for-like resource cost or convert CPU, memory, network, storage, and operational effort into a stated economic model.

Case file / Catalog cache reviewOne read path, plus the work needed to keep it useful.
Read volume
10,000 reads per second in the same measurement window.
Write volume
800 writes per second whose updates trigger invalidation.
Origin and lookup
8 CPU-ms per uncached read; 0.2 CPU-ms for each cache lookup.
Question
Does the measured hit rate reduce modeled compute work enough to justify the cache?
02 / Name every cost

Define the denominator before trusting the percentage.

For one interval, let R be reads, W writes that require invalidation, h the fraction of eligible reads served by the cache, Co origin CPU-ms per read, Cl cache lookup CPU-ms per read, and Ci invalidation CPU-ms per write. The denominator of the hit rate is eligible read requests: h = hits / (hits + misses). It is not all API requests unless all API requests are cacheable reads.

Then the illustrative baseline is R × Co. The modeled cached cost is R × Cl + R × (1 − h) × Co + W × Ci. Substituting the case values gives 10,000 × 0.2 + 10,000 × 0.15 × 8 + 800 × 0.5 = 14,400 CPU-ms/s. The baseline is 80,000 CPU-ms/s, so estimated savings are 65,600 CPU-ms/s across only the listed work categories.

Every cost must use the same time unit and compatible accounting boundary. If the cache lookup is measured on application hosts but origin cost is a database figure, CPU-ms can be added as a rough aggregate of compute demand only if the team accepts that proxy. For a cloud bill, price each resource on its own curve instead.

Model / workload compute

Without cache: reads × origin cost

With cache: reads × lookup cost + misses × origin cost + writes × invalidation cost

Break-even requires saved origin work > lookup work + write freshness work.

03 / Find the break-even point

The threshold is a consequence of the cost model.

Set the cached workload cost no greater than the uncached cost and solve for h. After cancelling the common read volume from the origin terms, the break-even hit rate is h > (R × Cl + W × Ci) / (R × Co). In this case, the threshold is (2,000 + 400) / 80,000 = 0.03, or 3%. At 85% the model has a large positive saving. That threshold belongs to these assumed costs and traffic; it is not a universal target for caches.

If the calculated threshold is above 100%, this cache cannot break even on the modeled work even at a perfect hit rate. If lookup cost is high or writes dominate, a high hit rate can still fail to pay. A negative estimate of saved CPU is a reason to inspect the model and workload, not proof that every cache should be removed.

04 / Account for freshness work

Writes change both the cost and the meaning of a hit.

The equation includes a simple per-write invalidation charge, but the operational effect can be larger. A write may invalidate several keys, fan out to many cache nodes, trigger a refill storm, or leave stale values visible until expiration. If updates happen more frequently than the TTL, frequent invalidation may keep the hit rate low while retaining most of the coordination cost.

Freshness is a correctness contract. Ask how stale data may be, whether writes synchronously invalidate or version entries, what happens when invalidation is lost, and whether a cache miss returns to a healthy origin. Measure write amplification and stale-read behavior alongside hit ratio. Do not raise TTL solely to increase hits if product behavior requires fresher values.

05 / Change the traffic shape

See which assumption is doing the work.

Change the measured workload and watch the modeled CPU demand. Reads and writes are rates per second; costs are CPU-ms per operation; hit rate is hits divided by eligible reads. The calculator reports only the listed compute costs, with no monetary valuation or freshness score.

Interactive estimate

Cache cost model

Uncached baseline80,000 CPU-ms/s
Modeled cached work14,400 CPU-ms/s
Estimated saving65,600 CPU-ms/s
Break-even hit rate3%

The listed compute costs are below the uncached baseline. Check omitted infrastructure costs and freshness before deciding.

06 / Practice in code

Make the denominator and validity checks visible.

The TypeScript and Go functions implement the stated model, reject invalid inputs, and return baseline, cached cost, savings, and break-even rate. Run the arithmetic by hand for the default inputs: 80,000 baseline CPU-ms/s, 14,400 modeled cached CPU-ms/s, 65,600 saved, and a 3% threshold. The result is only as complete as the cost categories supplied.

Compare the same estimate in TypeScript and Go.

Both examples validate the inputs and expose the cost model terms.

TypeScriptCache break-even · modeled CPU work
cache-cost.ts
type CacheModel = {
	readsPerSecond: number;
	writesPerSecond: number;
	originCpuMsPerRead: number;
	cacheCpuMsPerRead: number;
	invalidationCpuMsPerWrite: number;
	hitRate: number;
};

type CacheResult = {
	uncachedCpuMsPerSecond: number;
	cachedCpuMsPerSecond: number;
	savedCpuMsPerSecond: number;
	breakEvenHitRate: number;
};

function estimateCache(model: CacheModel): CacheResult {
	const values = Object.values(model);
	if (!values.every(Number.isFinite) || values.some((value) => value < 0)) {
		throw new Error('All inputs must be finite and non-negative.');
	}
	if (model.hitRate > 1) throw new Error('Hit rate must be between 0 and 1.');
	if (model.originCpuMsPerRead === 0 || model.readsPerSecond === 0) {
		throw new Error('A positive read volume and origin cost are needed for break-even.');
	}

	const uncachedCpuMsPerSecond = model.readsPerSecond * model.originCpuMsPerRead;
	const cachedCpuMsPerSecond =
		model.readsPerSecond * model.cacheCpuMsPerRead +
		model.readsPerSecond * (1 - model.hitRate) * model.originCpuMsPerRead +
		model.writesPerSecond * model.invalidationCpuMsPerWrite;
	const breakEvenHitRate =
		(model.readsPerSecond * model.cacheCpuMsPerRead +
			model.writesPerSecond * model.invalidationCpuMsPerWrite) /
		uncachedCpuMsPerSecond;
	return {
		uncachedCpuMsPerSecond,
		cachedCpuMsPerSecond,
		savedCpuMsPerSecond: uncachedCpuMsPerSecond - cachedCpuMsPerSecond,
		breakEvenHitRate
	};
}

const result = estimateCache({
	readsPerSecond: 10_000,
	writesPerSecond: 800,
	originCpuMsPerRead: 8,
	cacheCpuMsPerRead: 0.2,
	invalidationCpuMsPerWrite: 0.5,
	hitRate: 0.85
});

console.log(result);
GoCache break-even · modeled CPU work
cache-cost.go
package main

import (
	"errors"
	"fmt"
	"math"
)

type CacheModel struct {
	ReadsPerSecond            float64
	WritesPerSecond           float64
	OriginCPUMsPerRead        float64
	CacheCPUMsPerRead         float64
	InvalidationCPUMsPerWrite float64
	HitRate                   float64
}

type CacheResult struct {
	UncachedCPUMsPerSecond float64
	CachedCPUMsPerSecond   float64
	SavedCPUMsPerSecond    float64
	BreakEvenHitRate       float64
}

func estimateCache(model CacheModel) (CacheResult, error) {
	values := []float64{model.ReadsPerSecond, model.WritesPerSecond, model.OriginCPUMsPerRead, model.CacheCPUMsPerRead, model.InvalidationCPUMsPerWrite, model.HitRate}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
			return CacheResult{}, errors.New("all inputs must be finite and non-negative")
		}
	}
	if model.HitRate > 1 {
		return CacheResult{}, errors.New("hit rate must be between 0 and 1")
	}
	if model.ReadsPerSecond == 0 || model.OriginCPUMsPerRead == 0 {
		return CacheResult{}, errors.New("positive read volume and origin cost are needed for break-even")
	}

	uncached := model.ReadsPerSecond * model.OriginCPUMsPerRead
	cached := model.ReadsPerSecond*model.CacheCPUMsPerRead +
		model.ReadsPerSecond*(1-model.HitRate)*model.OriginCPUMsPerRead +
		model.WritesPerSecond*model.InvalidationCPUMsPerWrite
	breakEven := (model.ReadsPerSecond*model.CacheCPUMsPerRead +
		model.WritesPerSecond*model.InvalidationCPUMsPerWrite) / uncached
	return CacheResult{uncached, cached, uncached - cached, breakEven}, nil
}

func main() {
	result, err := estimateCache(CacheModel{10_000, 800, 8, 0.2, 0.5, 0.85})
	if err != nil {
		panic(err)
	}
	fmt.Printf("uncached %.0f CPU-ms/s; cached %.0f CPU-ms/s; saved %.0f CPU-ms/s; break-even hit rate %.1f%%\n",
		result.UncachedCPUMsPerSecond, result.CachedCPUMsPerSecond, result.SavedCPUMsPerSecond, result.BreakEvenHitRate*100)
}
07 / Measure before expanding

A model can rank a question. A production trace must answer it.

Before increasing cache scope, gather eligible reads, hits, misses, bypasses, write invalidations, origin cost per miss, lookup cost, refill behavior, stale age, and cache fleet cost over aligned windows. Segment hot and cold keys: one aggregate hit ratio can hide a small set of keys driving most origin cost.

Then compare an observed interval before and after a controlled change, accounting for traffic mix and request volume. If the cache reduces database work but shifts CPU to application hosts or network cost to another tier, state that transfer explicitly. Keep rollback conditions tied to correctness and latency objectives, not the hit-rate target alone.

Try the decision with new numbers

  1. Change to 2,000 reads/s, 1,500 invalidating writes/s, a 1 CPU-ms origin read, a 0.2 CPU-ms cache lookup, and a 0.5 CPU-ms invalidation. What is the break-even hit rate?
  2. At 60% hits, does the modeled cache save compute? Show both totals and identify which workload term dominates.
  3. Now suppose 10% of hits exceed the allowed staleness age. What new measurement and product constraint do you need before choosing a TTL?

Transfer: If a request can be served from a local in-process cache or a remote shared cache, which cost terms and failure modes change? Keep the hit denominator tied to requests eligible for each layer.