← Math in Practice
Concept Math behind AI

Matrices as transformations

A matrix turns a set of input coordinates into a new set of coordinates, using the same learned weights for every example.

A product team refreshes an embedding model and the ranking of known support queries changes. The new encoder has a projection layer: a table of learned weights that maps each input representation to another coordinate space. An engineer can inspect the shape and arithmetic without treating the weights as human-readable rules.

The judgment to keep

A matrix is a reusable linear transformation. Its rows specify output coordinates; its columns correspond to input coordinates. Check dimensions, a hand-worked example, and the exact order of composed layers before blaming the model or changing a threshold.

TypeScriptGo Matrix shapes · matrix-vector multiplication · batches · composition · neural-network layers
01 / Read the model change

A new layer can change every output while keeping the same input width.

After an encoder deployment, a set of known support queries moves in the retrieval ranking. The service team knows a projection layer changed, but the dashboard shows only final similarity scores. They want to check whether the new layer's dimensions and arithmetic are doing what the release notes claim.

We'll use two input coordinates and two output coordinates. The numbers are invented for arithmetic practice, not model weights or measured product results. The same steps work with a much wider matrix; deployed models may use several layers and nonlinear operations between them.

Case file / Projection layer reviewDoes the changed layer map inputs into the expected output shape?
Input width
Two coordinates: x = [2, 1].
Weight matrix
Two rows and two columns: W = [[2, 1], [1, 3]].
Convention
Column vector notation: y = Wx; each row makes one output.
Question
What should the output be, and how does the service apply it to a batch?
Keep the evidence: model version, input/output dimensions, weight and bias artifacts, preprocessing, and a few fixed input/output pairs.
02 / Read rows and columns

Rows decide outputs; columns accept inputs.

An m × n matrix has m rows and n columns. Under the column-vector convention, x has shape n × 1, so Wx has shape m × 1. The inner dimensions must agree: the matrix's column count must equal the input vector length. Each row produces one output coordinate.

Our matrix has shape 2 × 2; it accepts two inputs and produces two outputs. A 3 × 2 matrix would accept two inputs and produce three outputs. This shape bookkeeping is a quick way to catch an accidental transpose or incompatible model artifact.

01 / WeightsW: 2 × 2

Two output rows, two input columns.

02 / Inputx: 2 × 1

One column vector with two coordinates.

03 / Multiply(2 × 2)(2 × 1)

The two inner dimensions match.

04 / Outputy: 2 × 1

Two weighted output coordinates.

03 / Multiply one input by hand

Each output is one row dotted with the input.

For W = [[2, 1], [1, 3]] and x = [2, 1], calculate each row separately:

y₁ = 2×2 + 1×1 = 5

y₂ = 1×2 + 3×1 = 5

So Wx = [5, 5]. If the layer includes bias b = [1, −1], it adds one value to each output after the weighted sums: Wx + b = [6, 4]. Bias changes the affine transformation; it does not change the matrix's input/output shape.

Row 1[2, 1] · [2, 1]

2×2 + 1×1 = 5

Row 2[1, 3] · [2, 1]

1×2 + 3×1 = 5

Output[5, 5]

One value from each matrix row.

04 / Apply the layer to a batch

Reuse the same weights for each input row.

A service rarely encodes one item at a time. Suppose a batch stores each input as a row: X = [[2, 1], [0, 1]], with shape 2 × 2. Under our earlier column-vector convention, each column of a batch would be one example; the code and common ML libraries often store examples as rows instead. For this row-oriented storage, calculate Y = XWᵀ. This produces [[5, 5], [1, 3]].

The transpose here is bookkeeping for row-wise batch storage. The layer still applies the same transformation to both examples. Implementations use optimized matrix multiplication to perform many multiply-and-add operations efficiently, often on specialized hardware. The arithmetic contract stays inspectable even when the hardware path is not.

Batch calculationTwo examples, one shared weight matrix
Example 1
[2, 1] → [5, 5].
Example 2
[0, 1] → [1, 3].
Same parameters
Both use W = [[2, 1], [1, 3]].
Implementation convention
Document whether examples are rows or columns before comparing formulas.
05 / Compose transformations

Order matters when layers are composed.

With column vectors, applying W₁ and then W₂ gives W₂(W₁x) = (W₂W₁)x. Matrix multiplication composes the transformations, and it is generally not commutative: swapping the order can change the result. For a hand-check, let W₂ = [[1, 1], [0, 1]]. First W₁x = [5, 5]; then W₂[5, 5] = [10, 5]. Reversing the order is not the same operation.

This is one reason architecture order matters. A neural network applies learned transformations in a sequence, with nonlinear activation functions and sometimes normalization, residual connections, or other operations between them. A whole network is not necessarily one plain matrix product; biases and nonlinear steps affect how we combine the pieces.

Firstx → W₁x

[2, 1] becomes [5, 5].

ThenW₂(W₁x)

[5, 5] becomes [10, 5].

Combined(W₂W₁)x

Same order, same result.

06 / Diagnose before changing weights

Use a controlled input to locate a change in the pipeline.

  1. Reproduce. Save the exact inputs and both model artifact versions.
  2. Check shape. Confirm input width, output width, matrix orientation, and bias length.
  3. Hand-check. Run a tiny fixed vector through each row and compare to runtime output.
  4. Trace placement. Verify activation, normalization, preprocessing, and batch convention.
  5. Evaluate outcomes. Compare labeled retrieval quality and serving cost before rolling forward or back.
Transfer exercise

The encoder output width changes from 2 to 3.

What must change in the next layer's input shape? Name two checks beyond shape validation that would establish whether the new embeddings remain useful for search.

07 / Change the transform

Alter one weight and follow its effect on the output.

Use the illustrative two-input, two-output layer from the case file. These controls show the direct multiplication only; they do not represent a trained or production model.

Weight matrix W
Input vector x
y = Wx[5, 5]

Row 1: 2×2 + 1×1 = 5

Row 2: 1×2 + 3×1 = 5

08 / Practice in code

Validate the shape, then expose the weighted sums.

The examples apply a matrix to one vector, calculate a small batch, and optionally add a bias. They check dimensions and finite inputs; real model code must also manage tensor layouts, numerical behavior, model versioning, and evaluation.

Compare the same layer in TypeScript and Go.

Both examples apply W = [[2, 1], [1, 3]] to [2, 1] and [0, 1], producing [5, 5] and [1, 3].

TypeScriptMatrix-vector layer and batch
layer.ts
export type Matrix = readonly (readonly number[])[];
export type Vector = readonly number[];

function validateMatrix(matrix: Matrix, inputLength: number): number {
	if (matrix.length === 0 || matrix[0]?.length !== inputLength) {
		throw new RangeError('Matrix columns must match the input vector length.');
	}
	const width = matrix[0].length;
	if (matrix.some((row) => row.length !== width || !row.every(Number.isFinite))) {
		throw new RangeError('Matrix rows must be rectangular and finite.');
	}
	if (!Number.isSafeInteger(inputLength) || inputLength < 1) {
		throw new RangeError('Input dimension must be a positive integer.');
	}
	return matrix.length;
}

export function applyLayer(matrix: Matrix, input: Vector, bias?: Vector): number[] {
	if (!input.length || !input.every(Number.isFinite)) {
		throw new RangeError('Input vector must be nonempty and finite.');
	}
	const outputLength = validateMatrix(matrix, input.length);
	if (bias && (bias.length !== outputLength || !bias.every(Number.isFinite))) {
		throw new RangeError('Bias length must match the number of matrix rows.');
	}
	return matrix.map((row, i) =>
		row.reduce((sum, weight, j) => sum + weight * input[j], bias?.[i] ?? 0)
	);
}

export function applyBatch(matrix: Matrix, inputs: readonly Vector[]): number[][] {
	return inputs.map((input) => applyLayer(matrix, input));
}

const weights = [
	[2, 1],
	[1, 3]
];
const inputs = [
	[2, 1],
	[0, 1]
];

console.log('single input:', applyLayer(weights, inputs[0])); // [5, 5]
console.log('batch, one input per row:', applyBatch(weights, inputs)); // [[5, 5], [1, 3]]
console.log('with bias:', applyLayer(weights, inputs[0], [1, -1])); // [6, 4]
GoMatrix-vector layer and batch
layer.go
package main

import (
	"fmt"
	"math"
)

type Matrix [][]float64
type Vector []float64

func ApplyLayer(weights Matrix, input Vector, bias Vector) (Vector, error) {
	if len(input) == 0 {
		return nil, fmt.Errorf("input vector must be nonempty")
	}
	if len(weights) == 0 {
		return nil, fmt.Errorf("matrix must have at least one row")
	}
	for _, value := range input {
		if !isFinite(value) {
			return nil, fmt.Errorf("input values must be finite")
		}
	}
	output := make(Vector, len(weights))
	for i, row := range weights {
		if len(row) != len(input) {
			return nil, fmt.Errorf("row %d has %d columns; want %d", i, len(row), len(input))
		}
		if bias != nil && len(bias) != len(weights) {
			return nil, fmt.Errorf("bias length must match output dimension")
		}
		output[i] = 0
		if bias != nil {
			output[i] = bias[i]
		}
		for j, weight := range row {
			if !isFinite(weight) {
				return nil, fmt.Errorf("matrix values must be finite")
			}
			output[i] += weight * input[j]
		}
	}
	return output, nil
}

func isFinite(value float64) bool {
	return !math.IsNaN(value) && !math.IsInf(value, 0)
}

func main() {
	weights := Matrix{{2, 1}, {1, 3}}
	inputs := []Vector{{2, 1}, {0, 1}}
	for _, input := range inputs {
		output, err := ApplyLayer(weights, input, nil)
		if err != nil {
			panic(err)
		}
		fmt.Println(output)
	}
	withBias, err := ApplyLayer(weights, inputs[0], Vector{1, -1})
	if err != nil {
		panic(err)
	}
	fmt.Println("with bias:", withBias)
}
09 / Keep the model in context

A matrix explains a calculation; it does not explain the learned representation.

A matrix can be singular and have no inverse; many transformations intentionally compress dimensions, so information may be lost. Even when an inverse exists mathematically, applying it may be unstable or irrelevant to a model task. A layer with bias is affine, and a network with nonlinear activations is not generally reducible to one global matrix transformation. The matrix is still useful: it states how a particular layer combines coordinates at a particular point in the pipeline.

In model diagnosis, keep the numerical claim narrow. A shape-compatible multiplication shows that the operation can be computed. A matching hand calculation checks implementation. Neither establishes that the embeddings preserve the distinctions users need. That requires representative queries, relevance judgments, and an evaluation that reflects the product's costs.

Take the idea with you: label the shape convention, work one output row by hand, compare one or more fixed examples with runtime behavior, and evaluate model changes using task evidence.