The nearest result can still be the wrong answer for the user.
A team indexes short incident notes as embedding vectors. An engineer searches for “cache miss after deploy.” The top result says “cache miss after deploy,” while another says “cache misses after release.” The first sounds exact, but a recent index change has made it land below the second result. The team wants to know whether the metric, the vectors, or the data path changed.
The following two-dimensional vectors are deliberately tiny so their arithmetic can be checked by hand. Real embedding dimensions are numerous, and their coordinates usually do not have a direct human-readable meaning. The toy vectors illustrate metric behavior; they are not embeddings produced by a model.
- Query vector
q = [1, 1], one illustrative search request.- Candidate A
[1, 0], the exact-sounding title.- Candidate B
[10, 10], same direction as the query, much larger norm.- Candidate C
[1, 2], a related note with a different direction.- Question
- Does the index rank direction, coordinate distance, or a third behavior?
Cosine removes scale; Euclidean keeps it.
A vector is an ordered list of numbers. Its magnitude (or norm) is its
length: for [x, y], the Euclidean norm is √(x² + y²). Its direction describes its orientation from the origin. A dot product
combines coordinate-wise products: a · b = Σ aᵢbᵢ.
cosine(a, b) = (a · b) ÷ (‖a‖ × ‖b‖)
Cosine similarity divides out both lengths. For ordinary nonzero real vectors its range is −1 to 1: 1 means same direction, 0 means perpendicular, and −1 means opposite direction. The values you observe depend on the model and training objective, so do not assume that the full theoretical range appears in your data.
Euclidean distance(a, b) = √(Σ (aᵢ − bᵢ)²)
Euclidean distance is nonnegative, is zero for identical vectors, and grows as coordinates separate. If one candidate is a scaled copy of another, cosine stays the same while Euclidean distance usually changes. That is a key diagnostic when model output norms vary.
Does this candidate point in a similar direction?
How far apart are the actual numeric coordinates?
Cosine is unchanged for positive scaling; Euclidean is not.
Dimensions, normalization, and model version must agree.
The disagreement is visible before any model is involved.
For q = [1, 1], its magnitude is √2. Candidate A, [1, 0], has magnitude 1 and dot product 1, so its cosine is 1 ÷ √2 ≈ 0.707. Its Euclidean distance is √((1−1)² + (1−0)²) = 1.
Candidate B is [10, 10]. Its cosine is exactly 1 because it points in the
same direction as q. Its distance is √(9² + 9²) = √162 ≈ 12.728. Candidate C
is [1, 2]: dot product 3, magnitude √5, cosine 3 ÷ (√2 × √5) ≈ 0.949, and distance 1. The metric rankings therefore differ.
| Vector | Cosine with q | Euclidean distance from q | Interpretation |
|---|---|---|---|
| A · [1, 0] | ≈ 0.707 | 1 | Shorter, less aligned. |
| B · [10, 10] | 1 | ≈ 12.728 | Same direction, far in raw coordinates. |
| C · [1, 2] | ≈ 0.949 | 1 | More aligned than A, tied distance in this example. |
If we normalize every nonzero vector to unit length first, then ‖a − b‖² = 2 − 2 cos(a,b). So minimizing normalized Euclidean distance
produces the same ranking as maximizing cosine similarity. This is not a coincidence: on
the unit sphere, the only remaining difference is angle. Squared distance and distance
also have the same ranking because square root preserves order for nonnegative values.
A good-looking score does not prove a result is relevant.
The toy ranking tells us what these formulas do, not what a production embedding means. To investigate a retrieval regression, preserve a set of representative queries and human judgments, then inspect the same query and candidates through the old and new index paths. Record model and preprocessing versions, vector dimensions, norms, normalization steps, metric configuration, filters, and tie-breaking behavior. Compare top-k overlap and relevance labels rather than looking only at one query.
A ranking metric orders candidates. A threshold turns a score into a decision, such as “show this result” or “send this request to a human.” Cosine thresholds are not universal constants: they depend on the embedding model, corpus, query style, chunking, and the cost of false matches versus missed matches. Measure precision and recall on labeled data before choosing a cutoff.
See when the metric cares about vector length.
Keep the query and toy candidates fixed, then scale candidate B along its existing direction. Change the metric to compare its ranking effect. Turn on normalization to compare all vectors on the unit sphere; then the cosine and Euclidean rankings should agree. Scores are calculated from the displayed vectors, not from an embedding model.
Query: [1, 1]
- 01 B · “cache miss after deploy” (same direction, larger norm)Current vector: [10.00, 10.00] 1.000
- 02 C · “cache misses after release”Current vector: [1.00, 2.00] 0.949
- 03 A · “cache miss after deploy”Current vector: [1.00, 0.00] 0.707
Sorted by highest cosine similarity.
Reject mismatched dimensions and undefined cosine inputs.
The implementations validate finite coordinates and equal, nonempty dimensions. Euclidean distance accepts zero vectors. Cosine rejects them because its denominator would be zero. The example reports raw metrics; it does not silently normalize, so the caller must make that choice consistently for stored vectors and queries.
Trace the dot product, norms, and squared coordinate differences; then inspect input guards.
export type Vector = readonly number[];
function validatePair(a: Vector, b: Vector): void {
if (a.length === 0 || a.length !== b.length) {
throw new RangeError('Vectors must have the same nonzero dimension.');
}
if (![...a, ...b].every(Number.isFinite)) {
throw new RangeError('Vector coordinates must be finite.');
}
}
export function dotProduct(a: Vector, b: Vector): number {
validatePair(a, b);
return a.reduce((sum, value, index) => sum + value * b[index], 0);
}
export function euclideanDistance(a: Vector, b: Vector): number {
validatePair(a, b);
const squaredDistance = a.reduce((sum, value, index) => sum + (value - b[index]) ** 2, 0);
return Math.sqrt(squaredDistance);
}
export function cosineSimilarity(a: Vector, b: Vector): number {
validatePair(a, b);
const dot = dotProduct(a, b);
const normA = Math.sqrt(dotProduct(a, a));
const normB = Math.sqrt(dotProduct(b, b));
if (normA === 0 || normB === 0) {
throw new RangeError('Cosine similarity is undefined for a zero vector.');
}
return dot / (normA * normB);
}
export function normalize(vector: Vector): number[] {
if (vector.length === 0 || !vector.every(Number.isFinite)) {
throw new RangeError('Vector must be nonempty and have finite coordinates.');
}
const norm = Math.sqrt(vector.reduce((sum, value) => sum + value * value, 0));
if (norm === 0) throw new RangeError('A zero vector has no direction to normalize.');
return vector.map((value) => value / norm);
}
const query = [1, 1];
const candidates = [
{ name: 'A · exact-sounding title', vector: [1, 0] },
{ name: 'B · same direction, larger norm', vector: [10, 10] },
{ name: 'C · related wording', vector: [1, 2] }
];
for (const candidate of candidates) {
console.log({
name: candidate.name,
cosine: cosineSimilarity(query, candidate.vector).toFixed(3),
euclidean: euclideanDistance(query, candidate.vector).toFixed(3)
});
}
// For nonzero unit vectors, squared Euclidean distance = 2 × (1 − cosine).
const unitQuery = normalize(query);
const unitCandidate = normalize(candidates[2].vector);
console.log('normalized distance:', euclideanDistance(unitQuery, unitCandidate).toFixed(3));
package main
import (
"errors"
"fmt"
"math"
)
func validatePair(a, b []float64) error {
if len(a) == 0 || len(a) != len(b) {
return errors.New("vectors must have the same nonzero dimension")
}
for _, vector := range [][]float64{a, b} {
for _, value := range vector {
if math.IsNaN(value) || math.IsInf(value, 0) {
return errors.New("vector coordinates must be finite")
}
}
}
return nil
}
func dotProduct(a, b []float64) (float64, error) {
if err := validatePair(a, b); err != nil {
return 0, err
}
var dot float64
for i, value := range a {
dot += value * b[i]
}
return dot, nil
}
func euclideanDistance(a, b []float64) (float64, error) {
if err := validatePair(a, b); err != nil {
return 0, err
}
var squaredDistance float64
for i, value := range a {
difference := value - b[i]
squaredDistance += difference * difference
}
return math.Sqrt(squaredDistance), nil
}
func cosineSimilarity(a, b []float64) (float64, error) {
dot, err := dotProduct(a, b)
if err != nil {
return 0, err
}
normA, err := dotProduct(a, a)
if err != nil {
return 0, err
}
normB, err := dotProduct(b, b)
if err != nil {
return 0, err
}
if normA == 0 || normB == 0 {
return 0, errors.New("cosine similarity is undefined for a zero vector")
}
return dot / (math.Sqrt(normA) * math.Sqrt(normB)), nil
}
func normalize(vector []float64) ([]float64, error) {
if len(vector) == 0 {
return nil, errors.New("vector must be nonempty")
}
var squaredNorm float64
for _, value := range vector {
if math.IsNaN(value) || math.IsInf(value, 0) {
return nil, errors.New("vector coordinates must be finite")
}
squaredNorm += value * value
}
if squaredNorm == 0 {
return nil, errors.New("a zero vector has no direction to normalize")
}
norm := math.Sqrt(squaredNorm)
unit := make([]float64, len(vector))
for i, value := range vector {
unit[i] = value / norm
}
return unit, nil
}
func main() {
query := []float64{1, 1}
candidates := []struct {
name string
vector []float64
}{
{"A · exact-sounding title", []float64{1, 0}},
{"B · same direction, larger norm", []float64{10, 10}},
{"C · related wording", []float64{1, 2}},
}
for _, candidate := range candidates {
cosine, err := cosineSimilarity(query, candidate.vector)
if err != nil {
panic(err)
}
distance, err := euclideanDistance(query, candidate.vector)
if err != nil {
panic(err)
}
fmt.Printf("%s: cosine %.3f, Euclidean %.3f\n", candidate.name, cosine, distance)
}
// For nonzero unit vectors, squared Euclidean distance = 2 × (1 − cosine).
unitQuery, _ := normalize(query)
unitCandidate, _ := normalize(candidates[2].vector)
distance, _ := euclideanDistance(unitQuery, unitCandidate)
fmt.Printf("normalized distance: %.3f\n", distance)
}
Cosine is common when direction is the signal; it is not a universal default.
Cosine is common in embedding retrieval because it compares angular alignment while ignoring vector magnitude. That is useful when the representation or product objective treats scale as incidental. Euclidean distance is reasonable when absolute coordinate offsets and magnitude carry meaningful signal, or when a model and index were designed and evaluated around that geometry. Dot product is another common choice: it combines alignment and magnitude and can therefore rank scaled vectors differently from cosine.
Some embedding models recommend a particular similarity function or normalization procedure. Follow the model and index contract, then validate it against the task. Switching metrics, normalizing older stored vectors, or mixing model versions can change ranking even when the query text is unchanged. A benchmark with fixed, labeled queries can tell whether those changes improve the actual retrieval outcome.
- Observed
- Top results changed after the model cutover.
- Check first
- Dimensions, model version, normalization, norms, metric, and index configuration.
- Measure
- Top-k relevance on the same labeled query set before and after rebuilding.
- Decision
- Keep the metric that satisfies the measured retrieval objective and latency/cost constraints.