A new layer can change every output while keeping the same input width.
After an encoder deployment, a set of known support queries moves in the retrieval ranking. The service team knows a projection layer changed, but the dashboard shows only final similarity scores. They want to check whether the new layer's dimensions and arithmetic are doing what the release notes claim.
We'll use two input coordinates and two output coordinates. The numbers are invented for arithmetic practice, not model weights or measured product results. The same steps work with a much wider matrix; deployed models may use several layers and nonlinear operations between them.
- Input width
- Two coordinates:
x = [2, 1]. - Weight matrix
- Two rows and two columns:
W = [[2, 1], [1, 3]]. - Convention
- Column vector notation:
y = Wx; each row makes one output. - Question
- What should the output be, and how does the service apply it to a batch?
Rows decide outputs; columns accept inputs.
An m × n matrix has m rows and n columns. Under the
column-vector convention, x has shape n × 1, so Wx has shape m × 1. The inner dimensions must agree: the matrix's column count
must equal the input vector length. Each row produces one output coordinate.
Our matrix has shape 2 × 2; it accepts two inputs and produces two outputs. A 3 × 2 matrix would accept two inputs and produce three outputs. This shape bookkeeping
is a quick way to catch an accidental transpose or incompatible model artifact.
Two output rows, two input columns.
One column vector with two coordinates.
The two inner dimensions match.
Two weighted output coordinates.
Each output is one row dotted with the input.
For W = [[2, 1], [1, 3]] and x = [2, 1], calculate each row
separately:
y₁ = 2×2 + 1×1 = 5
y₂ = 1×2 + 3×1 = 5
So Wx = [5, 5]. If the layer includes bias b = [1, −1], it adds
one value to each output after the weighted sums: Wx + b = [6, 4]. Bias
changes the affine transformation; it does not change the matrix's input/output shape.
2×2 + 1×1 = 5
1×2 + 3×1 = 5
One value from each matrix row.
Reuse the same weights for each input row.
A service rarely encodes one item at a time. Suppose a batch stores each input as a row: X = [[2, 1], [0, 1]], with shape 2 × 2. Under our earlier column-vector convention, each column
of a batch would be one example; the code and common ML libraries often store examples as
rows instead. For this row-oriented storage, calculate Y = XWᵀ. This produces [[5, 5], [1, 3]].
The transpose here is bookkeeping for row-wise batch storage. The layer still applies the same transformation to both examples. Implementations use optimized matrix multiplication to perform many multiply-and-add operations efficiently, often on specialized hardware. The arithmetic contract stays inspectable even when the hardware path is not.
- Example 1
[2, 1] → [5, 5].- Example 2
[0, 1] → [1, 3].- Same parameters
- Both use
W = [[2, 1], [1, 3]]. - Implementation convention
- Document whether examples are rows or columns before comparing formulas.
Order matters when layers are composed.
With column vectors, applying W₁ and then W₂ gives W₂(W₁x) = (W₂W₁)x. Matrix multiplication composes the transformations, and it
is generally not commutative: swapping the order can change the result. For a hand-check,
let W₂ = [[1, 1], [0, 1]]. First W₁x = [5, 5]; then W₂[5, 5] = [10, 5]. Reversing the order is not the same operation.
This is one reason architecture order matters. A neural network applies learned transformations in a sequence, with nonlinear activation functions and sometimes normalization, residual connections, or other operations between them. A whole network is not necessarily one plain matrix product; biases and nonlinear steps affect how we combine the pieces.
[2, 1] becomes [5, 5].
[5, 5] becomes [10, 5].
Same order, same result.
Use a controlled input to locate a change in the pipeline.
- Reproduce. Save the exact inputs and both model artifact versions.
- Check shape. Confirm input width, output width, matrix orientation, and bias length.
- Hand-check. Run a tiny fixed vector through each row and compare to runtime output.
- Trace placement. Verify activation, normalization, preprocessing, and batch convention.
- Evaluate outcomes. Compare labeled retrieval quality and serving cost before rolling forward or back.
The encoder output width changes from 2 to 3.
What must change in the next layer's input shape? Name two checks beyond shape validation that would establish whether the new embeddings remain useful for search.
Alter one weight and follow its effect on the output.
Use the illustrative two-input, two-output layer from the case file. These controls show the direct multiplication only; they do not represent a trained or production model.
Row 1: 2×2 + 1×1 = 5
Row 2: 1×2 + 3×1 = 5
Validate the shape, then expose the weighted sums.
The examples apply a matrix to one vector, calculate a small batch, and optionally add a bias. They check dimensions and finite inputs; real model code must also manage tensor layouts, numerical behavior, model versioning, and evaluation.
Both examples apply W = [[2, 1], [1, 3]] to [2, 1] and [0, 1], producing [5, 5] and [1, 3].
export type Matrix = readonly (readonly number[])[];
export type Vector = readonly number[];
function validateMatrix(matrix: Matrix, inputLength: number): number {
if (matrix.length === 0 || matrix[0]?.length !== inputLength) {
throw new RangeError('Matrix columns must match the input vector length.');
}
const width = matrix[0].length;
if (matrix.some((row) => row.length !== width || !row.every(Number.isFinite))) {
throw new RangeError('Matrix rows must be rectangular and finite.');
}
if (!Number.isSafeInteger(inputLength) || inputLength < 1) {
throw new RangeError('Input dimension must be a positive integer.');
}
return matrix.length;
}
export function applyLayer(matrix: Matrix, input: Vector, bias?: Vector): number[] {
if (!input.length || !input.every(Number.isFinite)) {
throw new RangeError('Input vector must be nonempty and finite.');
}
const outputLength = validateMatrix(matrix, input.length);
if (bias && (bias.length !== outputLength || !bias.every(Number.isFinite))) {
throw new RangeError('Bias length must match the number of matrix rows.');
}
return matrix.map((row, i) =>
row.reduce((sum, weight, j) => sum + weight * input[j], bias?.[i] ?? 0)
);
}
export function applyBatch(matrix: Matrix, inputs: readonly Vector[]): number[][] {
return inputs.map((input) => applyLayer(matrix, input));
}
const weights = [
[2, 1],
[1, 3]
];
const inputs = [
[2, 1],
[0, 1]
];
console.log('single input:', applyLayer(weights, inputs[0])); // [5, 5]
console.log('batch, one input per row:', applyBatch(weights, inputs)); // [[5, 5], [1, 3]]
console.log('with bias:', applyLayer(weights, inputs[0], [1, -1])); // [6, 4]
package main
import (
"fmt"
"math"
)
type Matrix [][]float64
type Vector []float64
func ApplyLayer(weights Matrix, input Vector, bias Vector) (Vector, error) {
if len(input) == 0 {
return nil, fmt.Errorf("input vector must be nonempty")
}
if len(weights) == 0 {
return nil, fmt.Errorf("matrix must have at least one row")
}
for _, value := range input {
if !isFinite(value) {
return nil, fmt.Errorf("input values must be finite")
}
}
output := make(Vector, len(weights))
for i, row := range weights {
if len(row) != len(input) {
return nil, fmt.Errorf("row %d has %d columns; want %d", i, len(row), len(input))
}
if bias != nil && len(bias) != len(weights) {
return nil, fmt.Errorf("bias length must match output dimension")
}
output[i] = 0
if bias != nil {
output[i] = bias[i]
}
for j, weight := range row {
if !isFinite(weight) {
return nil, fmt.Errorf("matrix values must be finite")
}
output[i] += weight * input[j]
}
}
return output, nil
}
func isFinite(value float64) bool {
return !math.IsNaN(value) && !math.IsInf(value, 0)
}
func main() {
weights := Matrix{{2, 1}, {1, 3}}
inputs := []Vector{{2, 1}, {0, 1}}
for _, input := range inputs {
output, err := ApplyLayer(weights, input, nil)
if err != nil {
panic(err)
}
fmt.Println(output)
}
withBias, err := ApplyLayer(weights, inputs[0], Vector{1, -1})
if err != nil {
panic(err)
}
fmt.Println("with bias:", withBias)
}
- PyTorch: Linear layer — the documented affine operation and weight shape convention used by one widely used ML library.
- NumPy: matrix multiplication — matrix product shape rules and batched multiplication behavior.
A matrix explains a calculation; it does not explain the learned representation.
A matrix can be singular and have no inverse; many transformations intentionally compress dimensions, so information may be lost. Even when an inverse exists mathematically, applying it may be unstable or irrelevant to a model task. A layer with bias is affine, and a network with nonlinear activations is not generally reducible to one global matrix transformation. The matrix is still useful: it states how a particular layer combines coordinates at a particular point in the pipeline.
In model diagnosis, keep the numerical claim narrow. A shape-compatible multiplication shows that the operation can be computed. A matching hand calculation checks implementation. Neither establishes that the embeddings preserve the distinctions users need. That requires representative queries, relevance judgments, and an evaluation that reflects the product's costs.