The dashboard counts hits. The decision needs avoided work.
Imagine an API serving catalog reads. Across a measured one-second interval it handles 10,000 reads and 800 writes. A representative origin read uses 8 ms of database CPU. A cache lookup uses 0.2 ms of application CPU, and processing one write's invalidation event uses 0.5 ms of CPU. The measured read hit rate is 85%. These numbers are illustrative workload inputs, not benchmark claims.
We will compare modeled CPU time in CPU-ms per second. The unit lets us add
work rates, but it combines work on different machines into a load estimate; it is not a
dollar bill or a latency promise. A production comparison should use like-for-like
resource cost or convert CPU, memory, network, storage, and operational effort into a
stated economic model.
- Read volume
- 10,000 reads per second in the same measurement window.
- Write volume
- 800 writes per second whose updates trigger invalidation.
- Origin and lookup
- 8 CPU-ms per uncached read; 0.2 CPU-ms for each cache lookup.
- Question
- Does the measured hit rate reduce modeled compute work enough to justify the cache?
Define the denominator before trusting the percentage.
For one interval, let R be reads, W writes that require
invalidation, h the fraction of eligible reads served by the cache, Co origin CPU-ms per read, Cl cache lookup CPU-ms per read, and Ci invalidation CPU-ms per write. The denominator of the hit rate is eligible read requests: h = hits / (hits + misses). It is not all API requests unless all API
requests are cacheable reads.
Then the illustrative baseline is R × Co. The modeled cached cost
is R × Cl + R × (1 − h) × Co + W × Ci.
Substituting the case values gives 10,000 × 0.2 + 10,000 × 0.15 × 8 + 800 × 0.5 = 14,400 CPU-ms/s. The baseline
is 80,000 CPU-ms/s, so estimated savings are 65,600 CPU-ms/s across
only the listed work categories.
Every cost must use the same time unit and compatible accounting boundary. If the cache lookup is measured on application hosts but origin cost is a database figure, CPU-ms can be added as a rough aggregate of compute demand only if the team accepts that proxy. For a cloud bill, price each resource on its own curve instead.
Without cache: reads × origin cost
With cache: reads × lookup cost + misses × origin cost + writes × invalidation cost
Break-even requires saved origin work > lookup work + write freshness work.
The threshold is a consequence of the cost model.
Set the cached workload cost no greater than the uncached cost and solve for h. After cancelling the common read volume from the origin terms, the break-even hit rate
is h > (R × Cl + W × Ci) / (R × Co). In
this case, the threshold is (2,000 + 400) / 80,000 = 0.03, or 3%. At 85% the
model has a large positive saving. That threshold belongs to these assumed costs and
traffic; it is not a universal target for caches.
If the calculated threshold is above 100%, this cache cannot break even on the modeled work even at a perfect hit rate. If lookup cost is high or writes dominate, a high hit rate can still fail to pay. A negative estimate of saved CPU is a reason to inspect the model and workload, not proof that every cache should be removed.
Writes change both the cost and the meaning of a hit.
The equation includes a simple per-write invalidation charge, but the operational effect can be larger. A write may invalidate several keys, fan out to many cache nodes, trigger a refill storm, or leave stale values visible until expiration. If updates happen more frequently than the TTL, frequent invalidation may keep the hit rate low while retaining most of the coordination cost.
Freshness is a correctness contract. Ask how stale data may be, whether writes synchronously invalidate or version entries, what happens when invalidation is lost, and whether a cache miss returns to a healthy origin. Measure write amplification and stale-read behavior alongside hit ratio. Do not raise TTL solely to increase hits if product behavior requires fresher values.
See which assumption is doing the work.
Change the measured workload and watch the modeled CPU demand. Reads and writes are rates per second; costs are CPU-ms per operation; hit rate is hits divided by eligible reads. The calculator reports only the listed compute costs, with no monetary valuation or freshness score.
Cache cost model
Make the denominator and validity checks visible.
The TypeScript and Go functions implement the stated model, reject invalid inputs, and return baseline, cached cost, savings, and break-even rate. Run the arithmetic by hand for the default inputs: 80,000 baseline CPU-ms/s, 14,400 modeled cached CPU-ms/s, 65,600 saved, and a 3% threshold. The result is only as complete as the cost categories supplied.
Both examples validate the inputs and expose the cost model terms.
type CacheModel = {
readsPerSecond: number;
writesPerSecond: number;
originCpuMsPerRead: number;
cacheCpuMsPerRead: number;
invalidationCpuMsPerWrite: number;
hitRate: number;
};
type CacheResult = {
uncachedCpuMsPerSecond: number;
cachedCpuMsPerSecond: number;
savedCpuMsPerSecond: number;
breakEvenHitRate: number;
};
function estimateCache(model: CacheModel): CacheResult {
const values = Object.values(model);
if (!values.every(Number.isFinite) || values.some((value) => value < 0)) {
throw new Error('All inputs must be finite and non-negative.');
}
if (model.hitRate > 1) throw new Error('Hit rate must be between 0 and 1.');
if (model.originCpuMsPerRead === 0 || model.readsPerSecond === 0) {
throw new Error('A positive read volume and origin cost are needed for break-even.');
}
const uncachedCpuMsPerSecond = model.readsPerSecond * model.originCpuMsPerRead;
const cachedCpuMsPerSecond =
model.readsPerSecond * model.cacheCpuMsPerRead +
model.readsPerSecond * (1 - model.hitRate) * model.originCpuMsPerRead +
model.writesPerSecond * model.invalidationCpuMsPerWrite;
const breakEvenHitRate =
(model.readsPerSecond * model.cacheCpuMsPerRead +
model.writesPerSecond * model.invalidationCpuMsPerWrite) /
uncachedCpuMsPerSecond;
return {
uncachedCpuMsPerSecond,
cachedCpuMsPerSecond,
savedCpuMsPerSecond: uncachedCpuMsPerSecond - cachedCpuMsPerSecond,
breakEvenHitRate
};
}
const result = estimateCache({
readsPerSecond: 10_000,
writesPerSecond: 800,
originCpuMsPerRead: 8,
cacheCpuMsPerRead: 0.2,
invalidationCpuMsPerWrite: 0.5,
hitRate: 0.85
});
console.log(result);
package main
import (
"errors"
"fmt"
"math"
)
type CacheModel struct {
ReadsPerSecond float64
WritesPerSecond float64
OriginCPUMsPerRead float64
CacheCPUMsPerRead float64
InvalidationCPUMsPerWrite float64
HitRate float64
}
type CacheResult struct {
UncachedCPUMsPerSecond float64
CachedCPUMsPerSecond float64
SavedCPUMsPerSecond float64
BreakEvenHitRate float64
}
func estimateCache(model CacheModel) (CacheResult, error) {
values := []float64{model.ReadsPerSecond, model.WritesPerSecond, model.OriginCPUMsPerRead, model.CacheCPUMsPerRead, model.InvalidationCPUMsPerWrite, model.HitRate}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return CacheResult{}, errors.New("all inputs must be finite and non-negative")
}
}
if model.HitRate > 1 {
return CacheResult{}, errors.New("hit rate must be between 0 and 1")
}
if model.ReadsPerSecond == 0 || model.OriginCPUMsPerRead == 0 {
return CacheResult{}, errors.New("positive read volume and origin cost are needed for break-even")
}
uncached := model.ReadsPerSecond * model.OriginCPUMsPerRead
cached := model.ReadsPerSecond*model.CacheCPUMsPerRead +
model.ReadsPerSecond*(1-model.HitRate)*model.OriginCPUMsPerRead +
model.WritesPerSecond*model.InvalidationCPUMsPerWrite
breakEven := (model.ReadsPerSecond*model.CacheCPUMsPerRead +
model.WritesPerSecond*model.InvalidationCPUMsPerWrite) / uncached
return CacheResult{uncached, cached, uncached - cached, breakEven}, nil
}
func main() {
result, err := estimateCache(CacheModel{10_000, 800, 8, 0.2, 0.5, 0.85})
if err != nil {
panic(err)
}
fmt.Printf("uncached %.0f CPU-ms/s; cached %.0f CPU-ms/s; saved %.0f CPU-ms/s; break-even hit rate %.1f%%\n",
result.UncachedCPUMsPerSecond, result.CachedCPUMsPerSecond, result.SavedCPUMsPerSecond, result.BreakEvenHitRate*100)
}
A model can rank a question. A production trace must answer it.
Before increasing cache scope, gather eligible reads, hits, misses, bypasses, write invalidations, origin cost per miss, lookup cost, refill behavior, stale age, and cache fleet cost over aligned windows. Segment hot and cold keys: one aggregate hit ratio can hide a small set of keys driving most origin cost.
Then compare an observed interval before and after a controlled change, accounting for traffic mix and request volume. If the cache reduces database work but shifts CPU to application hosts or network cost to another tier, state that transfer explicitly. Keep rollback conditions tied to correctness and latency objectives, not the hit-rate target alone.
Try the decision with new numbers
- Change to 2,000 reads/s, 1,500 invalidating writes/s, a 1 CPU-ms origin read, a 0.2 CPU-ms cache lookup, and a 0.5 CPU-ms invalidation. What is the break-even hit rate?
- At 60% hits, does the modeled cache save compute? Show both totals and identify which workload term dominates.
- Now suppose 10% of hits exceed the allowed staleness age. What new measurement and product constraint do you need before choosing a TTL?
Transfer: If a request can be served from a local in-process cache or a remote shared cache, which cost terms and failure modes change? Keep the hit denominator tied to requests eligible for each layer.