← Math in Practice
Concept Math behind AI

Gradient descent and learning rate

Follow every parameter update, then check what the falling loss can actually tell you.

A ranking model's training loss has stopped improving. One engineer suggests increasing the learning rate; another says the feature scale changed in yesterday's data pipeline. Both could produce a slow or unstable trace. Before adjusting a setting, the team needs to know what a gradient step computes and which observations could distinguish the causes.

The judgment to keep

Gradient descent updates parameters in the direction that locally decreases a loss: θ := θ − η∇L(θ). The gradient supplies a direction and slope; the learning rate η sets the step size. A declining training loss says the model fits those training examples better under that loss. It does not, by itself, show that predictions improve on new data.

TypeScriptGo Loss · gradient · learning rate · optimizer trace · training and validation evidence
01 / Read the tuning incident

A training curve tells you what happened to one objective, not why.

Imagine a team retraining a ranking model after adding a new numeric feature. The training loss drops for several updates, then swings above and below its previous values. A larger learning rate might explain overshooting. So might a feature with a much larger numeric scale, a different batch mix, or noisy gradients. The curve is an observation; “the rate is too high” is a hypothesis that needs a test.

We will use a deliberately tiny model with one parameter and one loss value. This removes batches, many parameters, and data preparation so the arithmetic is visible. It teaches the update rule, not how to train a production ranking system.

Case file / Ranking model retrainThe reported training loss moves up and down after a feature pipeline change.
Observation
Loss is measured on the current training batches; its units follow the chosen loss function.
Change
A numeric feature was added and transformed in the data pipeline.
Competing explanations
Step size, feature scale, batch noise, or a changed objective/data population.
Question
What does the optimizer do on each step, and what evidence would isolate the cause?
02 / Choose a small loss surface

A one-dimensional bowl lets us check every step by hand.

Let the model have one parameter θ, with illustrative loss L(θ) = (θ − 3)². The loss is zero at θ = 3 and grows as θ moves away. This is a convex parabola with one global minimum, so it is intentionally friendlier than most real training objectives. We start at θ₀ = 0, where the loss is (0 − 3)² = 9.

The derivative is dL/dθ = 2(θ − 3). At θ = 0, the gradient is −6. Its negative sign means that a small move toward larger θ lowers this loss. The slope magnitude 6 says the loss changes steeply there per unit of θ. The derivative is not itself a recommended parameter value: it is local information about how the loss changes.

01 / LossL(θ) = (θ − 3)²

Illustrative objective; minimum loss is 0.

02 / Startθ₀ = 0

Starting loss is 9.

03 / Slope∇L(0) = −6

At zero, increasing θ initially lowers loss.

04 / Targetθ* = 3

The known minimum makes the trace easy to audit.

03 / Take a gradient step

Subtract the gradient scaled by the learning rate.

For a parameter vector θ, objective L, and positive learning rate η, the basic update is:

θₜ₊₁ = θₜ − η ∇L(θₜ)

With a single parameter, θ₀ = 0, gradient −6, and η = 0.1, the update is θ₁ = 0 − 0.1 × (−6) = 0.6. The new loss is (0.6 − 3)² = 5.76, down from 9. For this quadratic, the derivative at every new position points back toward 3.

In a model with many parameters, the gradient has one component per parameter. One learning rate scales all components in basic gradient descent. The resulting step can still be very different across directions because the loss surface may be steep in one direction and flat in another.

One update, with η = 0.1 and the illustrative loss
QuantityCalculationResult
Current parameterθ₀0
Gradient2(0 − 3)−6
Update0 − 0.1 × (−6)θ₁ = 0.6
New loss(0.6 − 3)²5.76
04 / Compare learning rates

Step size changes how quickly the parameter moves and whether it settles.

On this exact quadratic, the update simplifies to θₜ₊₁ − 3 = (1 − 2η)(θₜ − 3). The error from the minimum is multiplied by 1 − 2η every step. For 0 < η < 0.5, it approaches from the same side. For 0.5 < η < 1, the sign flips each step, so it crosses the minimum while shrinking its distance. At η = 1, it bounces between 0 and 6; above 1, the distance grows on this loss.

That exact threshold is a property of this chosen curve and update rule, not a universal learning-rate recipe. In a real model, curvature, parameter scaling, minibatch noise, momentum, adaptive optimizers, clipping, and schedules all change the observed behavior. “Lower” is not automatically “better”: a tiny rate may make progress too slowly for a useful training budget.

Six steps from θ₀ = 0. Illustrative values rounded to four decimals.
ηParameter traceObserved pattern
0.050 → 0.3 → 0.57 → 0.813 → 1.0313Steady but slow approach to 3; more updates are needed.
0.80 → 4.8 → 1.92 → 3.648 → 2.6112Crosses the minimum; oscillation shrinks on this curve.
1.10 → 6.6 → −1.32 → 8.184 → −3.2208Overshoots with growing distance; this run diverges.
05 / Inspect the optimizer trace

Change one control and read the parameter and loss at every step.

This lab calculates the same exact quadratic, L(θ) = (θ − 3)², and its derivative 2(θ − 3). Choose a rate, number of updates, and start value. The table is a deterministic model with exact gradients; it contains no training data, batch noise, or generalization result.

Loss is lower at this step count. Final θ = 2.8600; loss = 0.0196. The known minimum is θ = 3.

θₙ₊₁ = θₙ − η × 2(θₙ − 3)
StepθGradientLoss
00.0000-6.00009.0000
14.80003.60003.2400
21.9200-2.16001.1664
33.64801.29600.4199
42.6112-0.77760.1512
53.23330.46660.0544
62.8600-0.27990.0196
06 / Practice in code

Make the update and its consequences visible in the data structure.

The examples calculate the same one-parameter trace and keep the starting point, gradient, and loss with each row. The guards validate finite inputs and a bounded step count; they do not claim to be a complete machine-learning optimizer. A production optimizer must also define parameter storage, data batches, gradient computation, numerical precision, and checkpoint behavior.

Compare the same update in TypeScript and Go.

Both examples use the same quadratic, starting point, and learning-rate sweep.

TypeScriptTrace gradient descent on a quadratic
gradient-descent.ts
export type Step = { iteration: number; theta: number; gradient: number; loss: number };

// A deliberately small, deterministic loss surface: L(theta) = (theta - 3)^2.
export function gradientDescent(learningRate: number, steps: number, start = 0): Step[] {
	if (!Number.isFinite(learningRate) || learningRate <= 0) {
		throw new RangeError('learningRate must be finite and greater than zero');
	}
	if (!Number.isInteger(steps) || steps < 0 || steps > 100) {
		throw new RangeError('steps must be an integer from 0 to 100');
	}
	if (!Number.isFinite(start)) throw new TypeError('start must be finite');

	const trace: Step[] = [
		{ iteration: 0, theta: start, gradient: 2 * (start - 3), loss: (start - 3) ** 2 }
	];
	for (let iteration = 1; iteration <= steps; iteration++) {
		const previous = trace[iteration - 1].theta;
		const gradientAtPrevious = 2 * (previous - 3);
		const theta = previous - learningRate * gradientAtPrevious;
		trace.push({
			iteration,
			theta,
			gradient: 2 * (theta - 3),
			loss: (theta - 3) ** 2
		});
	}
	return trace;
}

// Example: inspect the parameter and loss at every update for three rates.
export function compareLearningRates() {
	return [0.05, 0.8, 1.1].map((rate) => ({
		learningRate: rate,
		trace: gradientDescent(rate, 6)
	}));
}
GoTrace gradient descent on a quadratic
gradient-descent.go
package main

import (
	"errors"
	"fmt"
	"math"
)

type Step struct {
	Iteration int
	Theta     float64
	Gradient  float64
	Loss      float64
}

// A deliberately small, deterministic loss surface: L(theta) = (theta - 3)^2.
func GradientDescent(learningRate float64, steps int, start float64) ([]Step, error) {
	if math.IsNaN(learningRate) || math.IsInf(learningRate, 0) || learningRate <= 0 {
		return nil, errors.New("learningRate must be finite and greater than zero")
	}
	if steps < 0 || steps > 100 {
		return nil, errors.New("steps must be from 0 to 100")
	}
	if math.IsNaN(start) || math.IsInf(start, 0) {
		return nil, errors.New("start must be finite")
	}

	gradient := 2 * (start - 3)
	trace := []Step{{Iteration: 0, Theta: start, Gradient: gradient, Loss: (start - 3) * (start - 3)}}
	for iteration := 1; iteration <= steps; iteration++ {
		previous := trace[iteration-1].Theta
		gradient = 2 * (previous - 3)
		theta := previous - learningRate*gradient
		trace = append(trace, Step{
			Iteration: iteration,
			Theta:     theta,
			Gradient:  2 * (theta - 3),
			Loss:      (theta - 3) * (theta - 3),
		})
	}
	return trace, nil
}

func main() {
	for _, rate := range []float64{0.05, 0.8, 1.1} {
		trace, err := GradientDescent(rate, 6, 0)
		if err != nil {
			panic(err)
		}
		fmt.Printf("learning rate %.2f:\n", rate)
		for _, step := range trace {
			fmt.Printf("  %d: theta=%.4f loss=%.4f\n", step.Iteration, step.Theta, step.Loss)
		}
	}
}
07 / Separate fit from generalization

A falling training loss answers a narrower question than “will this work?”

Training loss is the objective evaluated on training examples, often batches sampled from that set. If it falls, the optimizer is reducing that measured objective under the current pipeline. This can confirm that updates are being applied and can show whether the chosen objective is improving. It cannot show that the examples represent future traffic, that the labels are correct, or that a product outcome improved.

A validation metric evaluates a separate held-out set and offers evidence about performance on examples not used to compute those updates. It is still not a guarantee about production: the validation data may be stale, unrepresentative, or repeatedly used to tune decisions. Keep a final test set or later time window for a less repeatedly consulted check, and evaluate the product metric that actually matters, with slices that can reveal regressions.

If training loss falls while validation loss rises, one plausible explanation is overfitting, but first check that the two losses use compatible definitions and that preprocessing is consistent. If both stall, possible causes include a low rate, poor features, a gradient bug, an objective mismatch, or a plateau in the loss surface. The curves narrow the search; they do not uniquely identify the cause.

08 / Diagnose before tuning

The bowl is simple; real loss surfaces and gradients are not.

The learning-rate behavior above follows from a smooth, one-dimensional, deterministic quadratic and exact gradients. Real objectives have many parameters and can be non-convex. A local minimum is lower than nearby points but may not be globally best. A saddle point can have a nearly zero gradient while some directions descend and others ascend. Flat regions, steep narrow directions, and poorly scaled features can make one global rate inefficient or unstable.

With minibatch training, the gradient is an estimate based on the selected examples, so different batches can point in slightly different directions. This noise may help exploration, but it also makes individual loss values fluctuate. Large gradient norms, non-finite values, clipping, regularization, and optimizer state can all affect the trace. A plotted curve may also average or smooth values, so inspect how it was aggregated.

For the incident, make a diagnostic sequence: reproduce the preprocessing ranges; compare gradient norms and update magnitudes before and after the feature change; hold initialization, batches, and objective fixed while sweeping the rate; then compare training and validation curves and representative slices. If the feature scale changed dramatically, a controlled normalization experiment can test that hypothesis. Change one factor at a time and keep the experiment record, rather than selecting the run that merely looks smooth.

Transfer question: if the model's feature values are rescaled by 100 but the learning rate and initialization stay fixed, which evidence would you compare before deciding the optimizer is “broken”? Start with feature ranges, gradient norms, per-step update sizes, and both training and validation curves; the equation itself is unchanged, but the numerical landscape seen by the optimizer may be poorly conditioned.

References