A prediction is wrong; the first task is to define what “wrong” means numerically.
Imagine a team inspecting a tiny scoring model during a training review. For one example,
an input feature has value x = 2, the labeled target is y = 1,
and the model predicts ŷ = wx + b. At the current settings w = 1.5 and b = 0, the score is 3. The team uses half squared error, L = ½(ŷ − y)², so this example's loss is 2.
These are deliberately small illustrative numbers, not benchmark data and not a complete training system. They let us trace the calculation all the way from a model parameter to a loss. That path is the teaching target: before asking which parameter to change, identify the quantity being changed and the exact output being measured.
- Input, target
x = 2,y = 1(one toy training example).- Model score
ŷ = wx + b, currently3.- Loss
L = ½(ŷ − y)² = 2.- Question
- Near
w = 1.5, how much does the loss change per unit change inw?
The derivative is the slope at a point, not a promise about every possible change.
For a function f(w), the derivative at w is the limit of the change
in output divided by the change in input as that input change approaches zero:
f′(w) = limh→0 [f(w + h) − f(w)] / h
For our fixed x, y, and b, the loss is L(w) = ½(wx + b − y)². At the current point, the score is 3 and the residual ŷ − y is 2. A small increase in w increases the score by x times that change: for a small weight change Δw, the score
changes exactly by 2Δw. Near this point the loss changes approximately by 4Δw, or 4 loss units per weight unit.
Model score for the example.
Prediction minus target.
Loss changes locally by about 4 per weight unit.
The chain rule multiplies the sensitivities along a composed path.
The weight does not enter the loss in one jump. It first affects the prediction, and the
prediction affects the loss. Write that path as w → ŷ → L. The chain rule
says dL/dw = (dL/dŷ)(dŷ/dw). In our model, dL/dŷ = ŷ − y = 2 and dŷ/dw = x = 2, so dL/dw = 2 × 2 = 4.
This factorization is useful because real models are compositions of many operations. Each operation contributes a local derivative; multiplying them follows how a small upstream change propagates to a downstream quantity. For many parameters and examples, software automatic differentiation applies these local rules systematically. The arithmetic still depends on the function and the point being evaluated.
Changing this weight changes the score.
dŷ/dw = x = 2.
dL/dŷ = ŷ − y = 2.
Multiply along the dependency path.
dL/dw = (dL/dŷ)(dŷ/dw) = (ŷ − y)x = (3 − 1) × 2 = 4
A gradient is a list of local sensitivities, one for each parameter.
The score depends on two parameters, w and b. We can
differentiate the same loss with respect to each one. Since ŷ = wx + b, the
score changes by x per unit of weight and by 1 per unit of bias. Applying the
chain rule gives ∂L/∂w = (ŷ − y)x = 4 and ∂L/∂b = (ŷ − y) = 2.
Together these partial derivatives form the gradient ∇L = (∂L/∂w, ∂L/∂b) = (4, 2). It tells how the loss changes locally along
each parameter axis. For a sufficiently small change (Δw, Δb), the
first-order approximation is ΔL ≈ 4Δw + 2Δb. That is a local approximation;
for large moves, curvature can make it inaccurate.
| Parameter | Path to loss | Derivative | Meaning here |
|---|---|---|---|
| Weight w | w → ŷ → L | ∂L/∂w = 4 | One small positive weight change raises this example's loss by about 4 times the change. |
| Bias b | b → ŷ → L | ∂L/∂b = 2 | One small positive bias change raises this example's loss by about 2 times the change. |
A correct gradient answers a narrow question; it does not identify the production cause.
- Reproduce the observation. Preserve the example, model version, preprocessing, target, and reported metric.
- Check the objective. Confirm the loss formula, reduction (sum or mean), and which examples are included.
- Trace dependencies. Write the parameter-to-output path and verify each local derivative.
- Compare numerically. Perturb one parameter slightly and compare observed loss change with the local estimate.
- Evaluate transfer. Check representative held-out cases and product-facing outcomes before attributing improvement.
Compare the chain-rule result with a finite-difference estimate.
The lab keeps the model and loss fixed, then changes the input, target, parameters, or step size. It compares the analytic gradient with a central finite difference: evaluate the loss at a small positive and negative parameter offset, then divide their difference by twice the offset. With the default values, the weight derivative is exactly 4 analytically and approximately 4 numerically.
Model: ŷ = wx + b. Loss: L = ½(ŷ − y)². The central-difference estimate varies w or b alone while holding all other inputs fixed.
Keep the forward calculation and its sensitivities visible.
Both snippets calculate one example's prediction, residual, half-squared loss, and
analytic partial derivatives. Given x = 2, y = 1, w = 1.5, b = 0, they produce prediction 3, loss 2, and gradient (4, 2).
They reject non-finite inputs and overflowed results so invalid arithmetic is visible. The
code does not update parameters or implement a full training loop.
Each computes the derivative of one half-squared-error example using the chain rule.
/** One training example: prediction = weight * input + bias; loss = 1/2 * error^2. */
export function inspectExample(input: number, target: number, weight: number, bias: number) {
if (![input, target, weight, bias].every(Number.isFinite)) {
throw new RangeError('Inputs and parameters must be finite numbers.');
}
const prediction = weight * input + bias;
const error = prediction - target;
const loss = 0.5 * error ** 2;
const dLossDWeight = error * input;
const dLossDBias = error;
if (![prediction, error, loss, dLossDWeight, dLossDBias].every(Number.isFinite)) {
throw new RangeError('Calculation exceeded the finite number range.');
}
return { prediction, error, loss, dLossDWeight, dLossDBias };
}
const result = inspectExample(2, 1, 1.5, 0);
console.log(result); // prediction=3, error=2, loss=2, gradients=(4, 2)
package main
import (
"errors"
"fmt"
"math"
)
type Example struct {
Prediction float64
Error float64
Loss float64
DWeight float64
DBias float64
}
// InspectExample differentiates one half-squared-error example analytically.
func InspectExample(input, target, weight, bias float64) (Example, error) {
values := []float64{input, target, weight, bias}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) {
return Example{}, errors.New("inputs and parameters must be finite")
}
}
prediction := weight*input + bias
errValue := prediction - target
loss := 0.5 * errValue * errValue
dWeight := errValue * input
dBias := errValue
for _, value := range []float64{prediction, errValue, loss, dWeight, dBias} {
if math.IsNaN(value) || math.IsInf(value, 0) {
return Example{}, errors.New("calculation exceeded the finite number range")
}
}
return Example{prediction, errValue, loss, dWeight, dBias}, nil
}
func main() {
result, err := InspectExample(2, 1, 1.5, 0)
if err != nil {
panic(err)
}
fmt.Printf("prediction=%.1f error=%.1f loss=%.1f gradients=(%.1f, %.1f)\n",
result.Prediction, result.Error, result.Loss, result.DWeight, result.DBias)
}
The local slope depends on the model, the data, the objective, and the point.
We used one continuous linear score, one labeled example, and a smooth half-squared loss. Real models may be nonlinear, may combine many examples, and may use objectives with boundaries or nondifferentiable points. A reported derivative can also depend on preprocessing, regularization, sample weighting, and whether the loss is summed or averaged. Those details belong to the question being answered, not in footnotes after the number.
At a point where a function is not differentiable, a single ordinary derivative may not exist; software may use a defined subgradient convention for some operations. For a very large parameter move, the first-order estimate can be poor. And even a correctly computed gradient of a training loss is not evidence by itself that user-visible quality improves. That must be measured on relevant data and outcomes.
The team changes from one example to the mean loss over 100 examples.
What extra information do you need before predicting whether the mean-loss weight derivative is positive or negative? Name a finite-difference check you would run, and one validation measurement that would still be needed before claiming the user issue improved.