A short spike can dominate a chart even when the baseline is steady.
Imagine a worker queue sampled once per minute. The observed depths are 120, 130, 600, 140, 150, 160 jobs. These values are illustrative, not a
production measurement. A burst pushed one poll to 600; the next sample returned near the
prior level. A raw line shows every change, but rapid fluctuation can make a broader
movement hard to see.
A rolling average reduces noise by averaging a fixed number of recent readings. It requires keeping that window (or enough state for it) and gives older readings no influence once they leave the window. EWMA instead assigns exponentially decreasing influence to older observations, so a service can summarize history with one level and a chosen weight.
First diagnose what is fluctuating. Is the queue depth actually changing, are polls arriving at unequal intervals, or is a dashboard resampling a higher-frequency metric? Smoothing can help read a noisy display, but it cannot establish which explanation is true.
- Measurement
- Queued jobs at each poll, not jobs arriving per minute.
- Illustrative window
- Six samples one minute apart.
- Observed series
- 120, 130, 600, 140, 150, 160 jobs.
- Question
- How much should one new sample move a smoothed dashboard level?
Take α of the new reading and keep the rest of the previous estimate.
For observation xₜ, previous smoothed level sₜ₋₁, and smoothing
factor α:
sₜ = αxₜ + (1 − α)sₜ₋₁
The weights sum to one. Rewriting the update as sₜ = sₜ₋₁ + α(xₜ − sₜ₋₁) shows the same operation as moving a fraction α of
the gap between the estimate and the new observation. For the first value, a common
initialization is s₁ = x₁. This avoids inventing a prior zero; another prior
can be appropriate when it is known and documented.
At α = 0.25, start at 120 jobs. The next level is 0.25 × 130 + 0.75 × 120 = 122.5 jobs. For the burst sample it is 0.25 × 600 + 0.75 × 122.5 = 241.875 jobs. The smoothed level remains in jobs
because it is a weighted combination of queue depths.
| Poll | Observed jobs xₜ | Calculation | Level sₜ |
|---|---|---|---|
| 1 | 120 | Seed with first observation | 120.000 |
| 2 | 130 | 0.25 × 130 + 0.75 × 120 | 122.500 |
| 3 | 600 | 0.25 × 600 + 0.75 × 122.5 | 241.875 |
| 4 | 140 | 0.25 × 140 + 0.75 × 241.875 | 216.406 |
| 5 | 150 | 0.25 × 150 + 0.75 × 216.406 | 199.805 |
| 6 | 160 | 0.25 × 160 + 0.75 × 199.805 | 189.854 |
Higher α follows new evidence faster and preserves less history.
If the input stops moving and stays at a new value, the gap between the EWMA and that
value shrinks by a factor of 1 − α after every sample. After k equally spaced updates, the remaining fraction of the original gap is (1 − α)ᵏ. This is why the weights on older observations decay geometrically.
The half-life is the number of sample intervals until an observation's influence is
halved: k½ = ln(0.5) ÷ ln(1 − α). With α = 0.25, that is about 2.41 one-minute
intervals; the influence is halved after roughly 2.41 minutes when polling is regular.
This makes α easier to discuss as a response horizon. It is not a hard cutoff: older
samples continue to contribute, with progressively smaller weights.
A larger α is responsive but lets noisy samples move the estimate more. A smaller α produces a steadier line but adds lag when the underlying process changes. There is no universally correct value; choose based on the decision's timescale and inspect historical transitions and spikes.
- α = 0.10
- Half-life ≈ 6.58 samples, about 6.58 minutes.
- α = 0.25
- Half-life ≈ 2.41 samples, about 2.41 minutes.
- α = 0.50
- Half-life = 1 sample interval.
- α = 1
- Use only the latest observation; there is no prior influence or finite half-life.
A smoother changes the summary, not the system or the evidence.
An EWMA is useful for trend displays, noisy operational indicators, or forecasting a next value under a simple level model. But the smoothed line is not another observed queue depth. It can lag sudden changes, mute short incidents, and retain an after-effect after a single extreme reading. Label it as a smoothed estimate and make the raw measurements inspectable.
Diagnose a sudden movement from the raw series first: verify the unit, source, query, time range, and sampling cadence; compare related signals such as arrival rate, worker throughput, and oldest-job age. Then use the smoothed trend as one view of persistence. A lower smoothed value does not mean the spike was harmless, and a high smoothed value does not identify its cause.
Follow the same burst through fast and slow estimates.
Move α from 0.05 to 1. The observations stay fixed so the effect of the parameter is visible. The first value seeds the level; every subsequent row applies the recurrence. The half-life shown assumes exactly one minute between polls.
At each poll, the new observation receives 25% weight.
For regular one-minute samples, the calculated half-life is 2.41 sample intervals, or about 2.41 minutes.
| Poll | Observed xₜ | Smoothed sₜ |
|---|---|---|
| 1 | 120 | 120.00 |
| 2 | 130 | 122.50 |
| 3 | 600 | 241.88 |
| 4 | 140 | 216.41 |
| 5 | 150 | 199.80 |
| 6 | 160 | 189.85 |
Make the update rule and its input contract explicit.
These functions seed from the first observation, validate α in (0, 1], and
reject non-finite readings rather than silently treating a missing value as zero. Both
return the level after every sample so a reader can audit the trace. They assume a regular
cadence and all values belong to the same measured quantity and unit.
Watch the seed, validation, and recurrence for each observation.
export function ewma(values: number[], alpha: number): number[] {
if (!Number.isFinite(alpha) || alpha <= 0 || alpha > 1) {
throw new RangeError('alpha must be finite and in (0, 1]');
}
if (values.length === 0) return [];
if (!values.every(Number.isFinite)) {
throw new TypeError('every observation must be a finite number');
}
const result = [values[0]];
for (const value of values.slice(1)) {
const previous = result[result.length - 1];
result.push(alpha * value + (1 - alpha) * previous);
}
return result;
}
export function halfLifeInSamples(alpha: number): number {
if (!Number.isFinite(alpha) || alpha <= 0 || alpha >= 1) {
throw new RangeError('finite half-life requires alpha in (0, 1)');
}
return Math.log(0.5) / Math.log(1 - alpha);
}
const queueDepth = [120, 130, 600, 140, 150, 160];
console.log(ewma(queueDepth, 0.25));
// [120, 122.5, 241.875, 216.40625, 199.8046875, 189.853515625]
console.log(halfLifeInSamples(0.25));
// About 2.41 equally spaced sample intervals.
package main
import (
"errors"
"fmt"
"math"
)
func EWMA(values []float64, alpha float64) ([]float64, error) {
if math.IsNaN(alpha) || math.IsInf(alpha, 0) || alpha <= 0 || alpha > 1 {
return nil, errors.New("alpha must be finite and in (0, 1]")
}
result := make([]float64, len(values))
if len(values) == 0 {
return result, nil
}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) {
return nil, errors.New("every observation must be finite")
}
}
result[0] = values[0]
for i := 1; i < len(values); i++ {
result[i] = alpha*values[i] + (1-alpha)*result[i-1]
}
return result, nil
}
func HalfLifeInSamples(alpha float64) (float64, error) {
if math.IsNaN(alpha) || math.IsInf(alpha, 0) || alpha <= 0 || alpha >= 1 {
return 0, errors.New("finite half-life requires alpha in (0, 1)")
}
return math.Log(0.5) / math.Log(1-alpha), nil
}
func main() {
queueDepth := []float64{120, 130, 600, 140, 150, 160}
levels, err := EWMA(queueDepth, 0.25)
if err != nil {
panic(err)
}
fmt.Println(levels)
// [120 122.5 241.875 216.40625 199.8046875 189.853515625]
halfLife, err := HalfLifeInSamples(0.25)
if err != nil {
panic(err)
}
fmt.Printf("half-life: %.2f sample intervals\n", halfLife)
}
Equal sample steps are an assumption, and gaps need a policy.
The fixed-α recurrence treats every update as one equal time step. If samples arrive at
irregular intervals, the same α gives a different time-based memory: a one-minute gap and
a one-hour gap would each count as one update. For a desired continuous-time decay
constant τ and elapsed time Δt, one time-aware choice is αₜ = 1 − exp(−Δt / τ). That preserves a chosen time horizon under irregular observation times, provided τ and
timestamps are meaningful for the process. It is a modeling choice, not an automatic
correction for bad telemetry.
Missing is also not the same as zero. Replacing a missing queue sample with zero claims the queue was empty and pulls the estimate down. Skipping the update lets the previous level persist through the gap, while a time-aware update can account for elapsed time. Either choice may suit a particular measure, but record the missingness and decide explicitly. Also check whether values are instantaneous gauges, interval totals, or rates: applying the same smoother to each does not make their meanings interchangeable.
For choosing a statistic before smoothing, see Mean, median, variance, and percentiles. The sample statistic, interval, and population definition matter as much as the recurrence.