← Math in Practice
Concept Math behind AI

Quantization and numeric precision

A smaller numeric format saves memory only if the model still does its job.

A team wants to serve a larger language model on the same accelerator. Moving weights from fp32 to a smaller format could reduce memory and improve throughput, but a benchmark shows changed outputs for a subset of prompts. The team needs to separate storage arithmetic from numerical behavior and task quality.

The judgment to keep

Range tells you how large or small a representable value can be; precision tells you how finely nearby values are distinguished. Lower-bit storage can save substantial memory, but quantization introduces rounding and sometimes clipping. Validate the resulting model on representative tasks and hardware before choosing a format.

TypeScriptGo Floating-point range · significant precision · int8 calibration · memory · task quality
01 / Read the deployment problem

Smaller weights change both the footprint and the computation.

A team is testing a model conversion to fit more concurrent requests into available device memory. The converted model is smaller, but a few outputs changed and one benchmark regressed. That does not yet identify whether the cause is overflow, rounding, a calibration range that clips activations, unsupported hardware behavior, or a task-specific sensitivity.

Begin with two separate questions. First, what storage reduction should the format provide? Second, what numerical and task changes does this exact model conversion cause on the target runtime? Parameter count answers the first roughly; evaluation and hardware measurements answer the second.

Case file / Model-serving reviewDoes int8 make this model both smaller and safe to ship?
Goal
Reduce weight memory to fit a larger batch on target hardware.
Risk
Numerical error may concentrate in sensitive layers or outputs.
Knowns
Parameter count, storage format, calibration data, target runtime and device.
Decision evidence
Peak memory, latency/throughput, representative task metrics, and output/error slices.
Record: model and conversion version, calibration data, target hardware/runtime, peak memory, latency distribution, throughput, and task metric slices.
02 / Separate range from precision

fp16 has tighter range; bf16 keeps range with coarser steps.

IEEE-style floating-point values divide bits among sign, exponent, and significand. The exponent largely controls dynamic range; significand bits control how finely values are spaced near a magnitude. “16-bit” alone does not tell you the tradeoff.

fp32 uses 32 bits, about 24 bits of significand precision, and a maximum finite value near 3.4×10³⁸. fp16 uses 16 bits and about 11 bits of significand precision; its maximum finite value is 65,504, so large intermediate values can overflow even when fp32 handled them. bf16 also uses 16 bits, with about 8 significand bits, but an exponent width similar to fp32 gives it a much wider range, around 3.4×10³⁸. It represents a broad range coarsely.

int8 is a signed 8-bit integer with codes −128 through 127. It does not encode arbitrary real numbers by itself. A quantizer maps a chosen real interval to those codes using a scale and often a zero point. Its range depends on calibration and quantization scheme; values outside the interval may clip.

fp324 bytes · broad range

More significand precision for stored values.

fp162 bytes · narrower range

Fine precision relative to bf16, lower max magnitude.

bf162 bytes · broad range

fp32-like exponent range, coarser significand.

int81 byte · calibrated codes

Resolution depends on scale and represented interval.

03 / Calculate the memory change

Storage bytes are a useful first estimate, not peak serving memory.

For one million parameters, raw weight storage is about 4,000,000 bytes (4.00 MB or 3.81 MiB) in fp32, 2.00 MB in fp16 or bf16, and 1.00 MB in int8. This idealized arithmetic excludes metadata such as scales, zero points, alignment, and container overhead.

Runtime memory also includes activations, temporary workspaces, caches, allocator fragmentation, and possibly a higher-precision copy of some tensors. An int8 checkpoint does not guarantee an int8 execution path. Similarly, shorter arithmetic may increase throughput only when the target hardware and kernels support it efficiently.

Back-of-envelope / Raw weights onlyMultiply parameter count by bytes per value.
fp324.00 MB (3.81 MiB)

4 bytes per value

fp16 / bf162.00 MB (1.91 MiB)

2 bytes per value

int81.00 MB (0.95 MiB)

1 byte per value

ScopeWeights only

Runtime overhead is not included.

04 / Calibrate and measure error

Quantization maps a real interval to discrete codes.

For a simple uniform signed int8 mapping of a calibration interval [low, high], the step size is approximately (high−low)/255. Each value is rounded to a code; reconstruction maps the code back to a nearby real number. Widen the range and each step gets coarser; narrow it and more outliers may clip. Real schemes can be symmetric or asymmetric, per-tensor or per-channel, and may treat zero points and endpoints differently.

Calibration data estimates ranges or scales. It should represent expected serving inputs, but it must remain separate from the held-out evaluation set used to decide whether the converted model meets its task goal. Evaluate both common and important rare slices: average error or an aggregate score can hide a failure in a sensitive category.

RangeWhat can be represented?

Values outside it may saturate or clip.

Step sizeHow coarse are codes?

Wider interval at fixed bits means larger steps.

CalibrationWhat data set the scale?

Mismatch can cause clipping or waste precision.

Task qualityDoes behavior still work?

Check outputs on representative labeled data.

05 / Change the calibration range

See clipping and reconstruction error separately.

Interactive lab / Educational uniform int8 mappingMap a fixed toy weight vector into signed 8-bit codes.

Input values: -1.2, -0.65, -0.1, 0.1, 0.49, 1, 1.2

Step ≈ 0.0078; clipped values 2; reconstruction RMSE 0.1069.

Originalint8 codeReconstructed
-1.20-128-1.000
-0.65-83-0.647
-0.10-13-0.098
0.10120.098
0.49620.490
1.001271.000
1.201271.000

This is a transparent toy affine mapping. It does not simulate fp16/bf16 rounding, a neural network, or any particular framework's quantizer.

06 / Validate the model behavior

Pair numerical checks with the product task and serving target.

Keep a full-precision baseline, fix evaluation prompts/examples and decoding settings, then compare the converted model. Depending on the task, inspect exact-match or ranking quality, calibration, safety, structured output validity, or human-rated quality. Also inspect output deltas and error by layer or input slice where available. A small average error in weights or activations does not guarantee a small change in model behavior.

Benchmark the actual target device and runtime. Measure peak memory, throughput at a stated batch/concurrency, and latency percentiles after warmup. Include load/unload and conversion costs if they affect the service. Check for fallback operations that silently run in a wider format or on a different execution path.

Release record: calibration set, evaluation set, quantization scheme, hardware/runtime, task metrics and slices, memory, latency, throughput, and rollback threshold.
07 / Practice in code

Make the chosen interval visible in the calculation.

These snippets implement one educational uniform signed int8 mapping and a raw parameter-storage estimate. They validate inputs and report clipping and RMSE. They are not an implementation of floating-point formats or a substitute for your model framework's quantization workflow.

Go
quantize.go
package main

import (
	"errors"
	"math"
)

type Quantized struct {
	Values  []float64
	Scale   float64
	Clipped int
	RMSE    float64
}

func QuantizeInt8(values []float64, low, high float64) (Quantized, error) {
	if len(values) == 0 || math.IsNaN(low) || math.IsInf(low, 0) || math.IsNaN(high) || math.IsInf(high, 0) || low >= high {
		return Quantized{}, errors.New("use finite values and calibration range low < high")
	}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return Quantized{}, errors.New("values must be finite")
		}
	}
	scale := (high - low) / 255
	restored := make([]float64, len(values))
	clipped := 0
	squaredError := 0.0
	for i, value := range values {
		bounded := math.Max(low, math.Min(high, value))
		if bounded != value {
			clipped++
		}
		code := math.Round((bounded-low)/scale) - 128
		reconstructed := (code+128)*scale + low
		restored[i] = reconstructed
		delta := value - reconstructed
		squaredError += delta * delta
	}
	return Quantized{restored, scale, clipped, math.Sqrt(squaredError / float64(len(values)))}, nil
}

func StorageBytes(parameters, bytesPerValue int64) (int64, error) {
	if parameters < 0 || bytesPerValue < 1 {
		return 0, errors.New("use nonnegative parameter count and positive byte width")
	}
	return parameters * bytesPerValue, nil
}
Read the design in your languages.

These choices apply to the comparisons throughout this story.