← Applied algorithms
Text, names, and changes From similar letters to a review decision

Jaro–Winkler similarity

Explain why two names look alike.

Data tools already ship this score. Snowflake’s JAROWINKLER_SIMILARITY answers with a whole number from 0 to 100, and a 96 reads like “96% sure.” It isn’t. By the end of this lesson you’ll know what that kind of number counts.

Here’s a pair to count it on. A contact manager holds two records created separately: one given name is Martha, the other Marhta. Exact equality misses a plausible typo, and merging automatically would assume more than either name tells us. We want to surface the pair for review and explain which letters, order, and shared beginning put it there.

TypeScriptGoOne contact-name comparison in each language

01 / The idea

Ask what kind of similarity matters.

Jaro measures similarity using nearby matching symbols and disagreements in their order. Jaro–Winkler adds an adjustment for a shared prefix. Both give a score from 0 to 1, where higher means more alike under these rules. Neither is a probability that two records are the same person.

Our fictional contact manager already has the two candidate records, so we only compare their given names; finding candidates in a whole address book is its own job. A small preparation rule uppercases ASCII letters and collapses spaces, giving us MARTHA and MARHTA.

Levenshtein distance asks how many insertions, deletions, and substitutions would transform one name into the other. Here, changing T to H and H to T costs two edits. Jaro asks a different question: how much of each name can we match nearby, and how consistent is the order?

Levenshtein

2 edits

A minimum transformation cost. Smaller is closer.

Jaro

0.944444…

Nearby matches, with an order penalty. Larger is closer.

Jaro–Winkler

0.961111…

The same base score, boosted by the shared MAR.

The three numbers have different units. Two edits isn’t “twice as bad” as a similarity of 0.5, and 0.961111 isn’t 96% confidence. What matters is whether a score picks out pairs worth a person’s attention.

Follow the example in your languages.

Every comparison follows these choices. The experiment runs the TypeScript; the other languages implement the same bounded contract and are checked natively against the same cases.

02 / Name the rule

A letter must be nearby, and unused.

Both names have six symbols. Our matching radius is half the longer length, rounded down, minus one: 6 / 2 − 1 = 2. Negative radii clamp to zero. At each left position, look at right positions at most two places away and take the first unused equal symbol.

The left T at position 3 finds the right T at position 4. The left H at position 4 finds the right H at position 3. Each occurrence can be used once; two left letters cannot claim the same right letter.

Then compare the matched symbols in each name’s original order.

03 / Follow one operation

Six matches, two out of order.

Watch the matches, then the order check. Use Step through to inspect settled states, or Try it to change the names, prefix policy, and review cutoff.

Jaro–Winkler

Why names look alike.

AFTER ONE SCAN STEP1 match so far
A bounded search for each letterSource position 0: M. Target window 0 to 2. Matched target position 0. 1 match after this step.LeftRightM0A1R2T3H4A5M0A1R2H3T4A5

M at position 0 → right position 0

Radius 2. The dashed window covers positions 0–2. A ✓ marks a position used earlier.

01/ 04
Find nearby matches

Match each occurrence once.

Scan left to right. Take the first unused equal letter within two positions; all six letters find a match.

Reduced motion: choose a scene to see its completed state.

Read this scene

Scan left to right. Take the first unused equal letter within two positions; all six letters find a match.

Source position 0: M. Target window 0 to 2. Matched target position 0. 1 match after this step.

Watch and Step through share Martha / Marhta. Try it starts a fresh pair each time you open it; its input changes replace the completed result immediately.

All six symbols match, but their matched sequences are MARTHA and MARHTA. Positions 3 and 4 disagree. Jaro calls half that mismatch count its transposition term: t = 2 / 2 = 1.

This term summarizes order disagreement; it isn’t an edit script or a count of adjacent swaps. With three order mismatches, our implementation keeps t = 1.5 rather than rounding down.

04 / Read the shape

Keep the evidence behind the number.

The complete implementation first converts the inputs to Unicode scalar values. This excerpt then searches bounded windows and records the chosen pairs. A used flag prevents reuse, and scanning each window from its lowest index makes repeated-letter choices deterministic.

TypeScriptFind nearby, unused matches
contacts.ts
const radius = Math.max(0, Math.floor(Math.max(left.length, right.length) / 2) - 1);
const used = Array<boolean>(right.length).fill(false);
const pairs: Pair[] = [],
	steps: MatchStep[] = [];
for (let i = 0; i < left.length; i++) {
	const start = Math.min(right.length, Math.max(0, i - radius));
	const end = Math.min(right.length, i + radius + 1); // Exclusive.
	let matchIndex: number | null = null;
	for (let j = start; j < end; j++) {
		if (!used[j] && left[i] === right[j]) {
			used[j] = true;
			matchIndex = j;
			pairs.push({ left: i, right: j });
			break;
		}
	}
	steps.push({ leftIndex: i, start, end, matchIndex });
}
GoFind nearby, unused matches
contacts.go
radius := max(0, max(len(left), len(right))/2-1)
used := make([]bool, len(right))
pairs, steps := []Pair{}, []MatchStep{}
for i, symbol := range left {
	start := min(len(right), max(0, i-radius))
	end := min(len(right), i+radius+1) // Exclusive.
	matchIndex := -1
	for j := start; j < end; j++ {
		if !used[j] && symbol == right[j] {
			used[j] = true
			matchIndex = j
			pairs = append(pairs, Pair{i, j})
			break
		}
	}
	steps = append(steps, MatchStep{i, start, end, matchIndex})
}

Let m be the match count, a and b the two lengths, and t half the order mismatches. Jaro averages three contributions: m / a, m / b, and (m − t) / m. For our pair, that is (6/6 + 6/6 + 5/6) / 3 = 0.944444….

The first two terms reward coverage of each input. The third reduces the score when matched symbols appear in a different order.

The Winkler adjustment is prefix length × 0.1 × (1 − Jaro). We count at most four initial equal symbols and apply it only when the computed Jaro score is greater than 0.7. For MAR, the added amount is 3 × 0.1 × (1 − 0.944444…) = 0.016666….

TypeScriptCompute Jaro and the prefix contribution
contacts.ts
const leftOrder = pairs.map((pair) => left[pair.left]);
const rightOrder = right.filter((_, index) => used[index]);
const mismatches = leftOrder.filter((symbol, i) => symbol !== rightOrder[i]).length;
const transpositions = mismatches / 2; // May be fractional; not an edit count.
const m = pairs.length;
const jaro =
	m === 0
		? left.length === 0 && right.length === 0
			? 1
			: 0
		: (m / left.length + m / right.length + (m - transpositions) / m) / 3;
let prefix = 0;
while (prefix < Math.min(4, left.length, right.length) && left[prefix] === right[prefix])
	prefix++;
const boost = jaro > 0.7 ? 0.1 * prefix * (1 - jaro) : 0;
const winkler = jaro + boost;
GoCompute Jaro and the prefix contribution
contacts.go
leftOrder, rightOrder := []rune{}, []rune{}
for _, pair := range pairs {
	leftOrder = append(leftOrder, left[pair.Left])
}
for i, symbol := range right {
	if used[i] {
		rightOrder = append(rightOrder, symbol)
	}
}
mismatches := 0
for i, symbol := range leftOrder {
	if symbol != rightOrder[i] {
		mismatches++
	}
}
transpositions := float64(mismatches) / 2 // Retain a possible half.
m, jaro := float64(len(pairs)), 0.0
if m == 0 {
	if len(left) == 0 && len(right) == 0 {
		jaro = 1
	}
} else {
	jaro = (m/float64(len(left)) + m/float64(len(right)) + (m-transpositions)/m) / 3
}
prefix := 0
for prefix < min(4, len(left), len(right)) && left[prefix] == right[prefix] {
	prefix++
}
boost := 0.0
if jaro > 0.7 {
	boost = 0.1 * float64(prefix) * (1 - jaro)
}
winkler := jaro + boost
Reading the matching codeWindow bounds and empty inputs

Window ends are exclusive in code. A radius of 2 around position 3 scans positions 1 through 5, written as the range from 1 up to, but not including, 6.

With no matches the score is zero, returned before dividing. Two empty raw strings score one by an explicit convention; the contact wrapper checks for a blank name before it ever scores one.

Reading the TypeScriptOne object with every step, and null

analyze returns an Analysis object that keeps every intermediate value: the radius, each window step, the matched pairs, both matched orders, and the three scores. The visual reads its explanation straight from that object. A step’s matchIndex is null when its window held no unused equal symbol, and transpositions is a plain number that may be fractional.

screenPair returns analysis: null and score: null for a missing name. Invalid input throws an Error.

Reading the Go-1, pointers, and returned errors

Analyze and ScreenPair return a value and an error. A step’s MatchIndex is -1 when nothing matched.

Screening holds an *Analysis and a *float64, which stay nil for a missing name, where TypeScript uses null. main panics only if the fixed demo returns an error.

What is refusedText, lengths, and cutoffs

TypeScript throws on invalid input and Go returns an error. Every version rejects excessive lengths and nonfinite or out-of-range cutoffs. TypeScript rejects unpaired UTF-16 surrogates and Go rejects invalid UTF-8. Matching uses scalar positions, not byte offsets or user-perceived grapheme clusters.

Each name may hold up to 64 Unicode scalar values, and the review cutoff must be a number from 0 to 1. A blank prepared name isn’t an error: it returns “missing name” with no score. Both languages run the same shared cases.

The prefix rule favors disagreements later in a name. Try Martha / Nartha to remove the shared beginning, then Martha / Mzzzzz to see why a shared first letter alone doesn’t trigger the boost. The weighting is a policy: check it against the names your application actually sees.

05 / Try a decision

Keep the decision explainable.

The experiment shows both base and adjusted scores, even when only one drives the review decision. That lets a reviewer see whether the shared beginning changed the outcome. Formatting a score to six decimals is for reading; cutoff decisions use its unrounded value.

A real contact workflow would carry record identifiers, original fields, and the reason a pair was suggested. Joining records needs more identifying evidence and an explicit workflow; this lesson stops at the suggestion.

Show the score as what it is. The two names, their prepared keys, and the contributions explain the suggestion; a ring reading “96% match” would look like a confidence estimate the algorithm never produced.

The score crossed the cutoff. What follows?

Two independently created contact records say Martha and Marhta. Their adjusted given-name score is 0.961111, above a review cutoff of 0.95. No shared trusted identifier has been established.

06 / Follow the cost

The scorer is cheap. The pairs are not.

Jaro–Winkler: time and extra space, for names of a and b scalars and N contacts
OperationTimeExtra spaceWhat it assumes
Find nearby, unused matchesO(a × b + a + b)O(a + b)Each left symbol scans a window of at most 2r + 1 right symbols. The used flags, steps, and pairs are the evidence the visual reads.
Compare order, count the prefixO(a + b)O(a + b)The matched symbols in each name’s original order, then at most four equal starting symbols.
Prepare and screen one pairO(a × b + a + b)O(a + b)Preparation is one pass over each name; comparing the score with the cutoff is one step.
Screen every pair of N contactsO(N² × a × b)O(a + b)N(N − 1) / 2 comparisons, one at a time. A faster scorer doesn’t change that count; choosing candidates does.

The pair scorer is cheap; choosing which pairs to score is not. Comparing every pair among N contacts means roughly N² / 2 comparisons, and a faster scorer doesn’t change that.

What one comparison costsTime and space for the window scan

For lengths a and b, the window scan takes at most O(a × b + a + b) time and O(a + b) auxiliary space, including the evidence the visual uses. The 64-symbol input limit bounds both here.

07 / Give it a real job

“Worth reviewing” is a caller’s decision.

The scoring function knows nothing about contacts. Our wrapper prepares the two names, rejects missing name information, chooses Jaro or Jaro–Winkler, and compares the unrounded result with an inclusive review cutoff. At 0.95, the default pair crosses the cutoff only with the prefix adjustment enabled.

Keep two thresholds apart: 0.7 decides whether the prefix boost applies, and 0.95 is this example’s review policy. Neither is a universal duplicate threshold.

TypeScriptPrepare names and screen a pair
contacts.ts
export function prepareName(text: string): string {
	symbols(text);
	return text
		.replace(/[a-z]/g, (letter) => letter.toUpperCase())
		.replace(/[\t\n\v\f\r ]+/g, ' ')
		.replace(/^ | $/g, '');
}
export function screenPair(
	left: string,
	right: string,
	cutoff: number,
	usePrefix: boolean
): Screening {
	if (!Number.isFinite(cutoff) || cutoff < 0 || cutoff > 1)
		throw new Error('Review cutoff must be between 0 and 1.');
	const leftKey = prepareName(left),
		rightKey = prepareName(right);
	if (!leftKey || !rightKey)
		return { leftKey, rightKey, analysis: null, score: null, status: 'missing-name' };
	const analysis = analyze(leftKey, rightKey);
	const score = usePrefix ? analysis.winkler : analysis.jaro;
	return {
		leftKey,
		rightKey,
		analysis,
		score,
		status: score >= cutoff ? 'review' : 'below-cutoff'
	};
}
GoPrepare names and screen a pair
contacts.go
func PrepareName(text string) (string, error) {
	values, err := Symbols(text)
	if err != nil {
		return "", err
	}
	var out strings.Builder
	pendingSpace := false
	for _, r := range values {
		if r == ' ' || r == '\t' || r == '\n' || r == '\v' || r == '\f' || r == '\r' {
			pendingSpace = out.Len() > 0
			continue
		}
		if pendingSpace {
			out.WriteByte(' ')
			pendingSpace = false
		}
		if r >= 'a' && r <= 'z' {
			r -= 'a' - 'A'
		}
		out.WriteRune(r)
	}
	return out.String(), nil
}
func ScreenPair(left, right string, cutoff float64, usePrefix bool) (Screening, error) {
	if math.IsNaN(cutoff) || math.IsInf(cutoff, 0) || cutoff < 0 || cutoff > 1 {
		return Screening{}, fmt.Errorf("review cutoff must be between 0 and 1")
	}
	leftKey, err := PrepareName(left)
	if err != nil {
		return Screening{}, err
	}
	rightKey, err := PrepareName(right)
	if err != nil {
		return Screening{}, err
	}
	result := Screening{LeftKey: leftKey, RightKey: rightKey, Status: "missing-name"}
	if leftKey == "" || rightKey == "" {
		return result, nil
	}
	analysis, err := Analyze(leftKey, rightKey)
	if err != nil {
		return Screening{}, err
	}
	score := analysis.Jaro
	if usePrefix {
		score = analysis.Winkler
	}
	result.Analysis = &analysis
	result.Score = &score
	result.Status = "below-cutoff"
	if score >= cutoff {
		result.Status = "review"
	}
	return result, nil
}

Preparation changes comparison keys, never the source strings. It doesn’t remove accents, reorder names, expand nicknames, or normalize Unicode; each of those needs its own requirements. A blank prepared name returns “missing name” with no score, because two empty fields aren’t evidence that two contacts match.

In a real contact manager this runs where the records live, in the service or batch job that pairs candidates for review. Nothing in a component computes it; the page shows the score and its evidence.

What preparation does, exactlyWhitespace, empty names, and a cutoff of zero

It recognizes only ASCII letters and ASCII whitespace: space, tab, line feed, vertical tab, form feed, and carriage return.

Raw empty strings compare equal, but the wrapper checks for a blank name first. Nonblank pairs with score zero do pass a cutoff of zero: that is exactly the configured policy, and the lab makes its consequence visible.

TypeScriptUse the result without joining records
contacts.ts
export function demo(): string {
	const pair = screenPair('Martha', 'Marhta', 0.95, true);
	if (pair.score === null) return 'More name information needed.';
	return `score=${pair.score.toFixed(6)}; ${pair.status}\nKeep both contact records until identity is checked.`;
}
GoUse the result without joining records
contacts.go
func Demo() (string, error) {
	pair, err := ScreenPair("Martha", "Marhta", 0.95, true)
	if err != nil {
		return "", err
	}
	if pair.Score == nil {
		return "More name information needed.", nil
	}
	return fmt.Sprintf("score=%.6f; %s\nKeep both contact records until identity is checked.", *pair.Score, pair.Status), nil
}
func main() {
	output, err := Demo()
	if err != nil {
		panic(err)
	}
	fmt.Println(output)
}
Copy and run the complete exampleSelf-contained, no packages

Each file prints score=0.961111; review, then asks you to keep both records until identity is checked. No packages, database, or service are required. Inputs are limited to 64 Unicode scalar values per string; the browser experiment uses 12 so individual positions remain inspectable.

Save the selected file and run node --experimental-strip-types contacts.ts with Node 22.18 or newer, or go run contacts.go with Go 1.23 or newer.

TypeScriptComplete contact-name comparison
contacts.ts
// A bounded teaching implementation. Scores are not identity probabilities.
export const MAX_SYMBOLS = 64;
export function symbols(text: string): string[] {
	const out: string[] = [];
	for (const symbol of text) {
		const point = symbol.codePointAt(0)!;
		if (point >= 0xd800 && point <= 0xdfff) throw new Error('Text must be well-formed Unicode.');
		out.push(symbol);
		if (out.length > MAX_SYMBOLS) throw new Error('Use at most 64 Unicode scalar values.');
	}
	return out;
}
export type MatchStep = {
	leftIndex: number;
	start: number;
	end: number;
	matchIndex: number | null;
};
export type Pair = { left: number; right: number };
export type Analysis = {
	left: string[];
	right: string[];
	radius: number;
	steps: MatchStep[];
	pairs: Pair[];
	leftOrder: string[];
	rightOrder: string[];
	mismatches: number;
	transpositions: number;
	prefix: number;
	jaro: number;
	boost: number;
	winkler: number;
};

export function analyze(leftText: string, rightText: string): Analysis {
	const left = symbols(leftText),
		right = symbols(rightText);
	const radius = Math.max(0, Math.floor(Math.max(left.length, right.length) / 2) - 1);
	const used = Array<boolean>(right.length).fill(false);
	const pairs: Pair[] = [],
		steps: MatchStep[] = [];
	for (let i = 0; i < left.length; i++) {
		const start = Math.min(right.length, Math.max(0, i - radius));
		const end = Math.min(right.length, i + radius + 1); // Exclusive.
		let matchIndex: number | null = null;
		for (let j = start; j < end; j++) {
			if (!used[j] && left[i] === right[j]) {
				used[j] = true;
				matchIndex = j;
				pairs.push({ left: i, right: j });
				break;
			}
		}
		steps.push({ leftIndex: i, start, end, matchIndex });
	}
	const leftOrder = pairs.map((pair) => left[pair.left]);
	const rightOrder = right.filter((_, index) => used[index]);
	const mismatches = leftOrder.filter((symbol, i) => symbol !== rightOrder[i]).length;
	const transpositions = mismatches / 2; // May be fractional; not an edit count.
	const m = pairs.length;
	const jaro =
		m === 0
			? left.length === 0 && right.length === 0
				? 1
				: 0
			: (m / left.length + m / right.length + (m - transpositions) / m) / 3;
	let prefix = 0;
	while (prefix < Math.min(4, left.length, right.length) && left[prefix] === right[prefix])
		prefix++;
	const boost = jaro > 0.7 ? 0.1 * prefix * (1 - jaro) : 0;
	const winkler = jaro + boost;
	return {
		left,
		right,
		radius,
		steps,
		pairs,
		leftOrder,
		rightOrder,
		mismatches,
		transpositions,
		prefix,
		jaro,
		boost,
		winkler
	};
}

export type Screening = {
	leftKey: string;
	rightKey: string;
	analysis: Analysis | null;
	score: number | null;
	status: 'missing-name' | 'below-cutoff' | 'review';
};
export function prepareName(text: string): string {
	symbols(text);
	return text
		.replace(/[a-z]/g, (letter) => letter.toUpperCase())
		.replace(/[\t\n\v\f\r ]+/g, ' ')
		.replace(/^ | $/g, '');
}
export function screenPair(
	left: string,
	right: string,
	cutoff: number,
	usePrefix: boolean
): Screening {
	if (!Number.isFinite(cutoff) || cutoff < 0 || cutoff > 1)
		throw new Error('Review cutoff must be between 0 and 1.');
	const leftKey = prepareName(left),
		rightKey = prepareName(right);
	if (!leftKey || !rightKey)
		return { leftKey, rightKey, analysis: null, score: null, status: 'missing-name' };
	const analysis = analyze(leftKey, rightKey);
	const score = usePrefix ? analysis.winkler : analysis.jaro;
	return {
		leftKey,
		rightKey,
		analysis,
		score,
		status: score >= cutoff ? 'review' : 'below-cutoff'
	};
}

export function demo(): string {
	const pair = screenPair('Martha', 'Marhta', 0.95, true);
	if (pair.score === null) return 'More name information needed.';
	return `score=${pair.score.toFixed(6)}; ${pair.status}\nKeep both contact records until identity is checked.`;
}

console.log(demo());
GoComplete contact-name comparison
contacts.go
package main

import (
	"fmt"
	"math"
	"strings"
	"unicode/utf8"
)

const MaxSymbols = 64

func Symbols(text string) ([]rune, error) {
	if !utf8.ValidString(text) {
		return nil, fmt.Errorf("text must be well-formed Unicode")
	}
	out := []rune{}
	for _, symbol := range text {
		out = append(out, symbol)
		if len(out) > MaxSymbols {
			return nil, fmt.Errorf("use at most 64 Unicode scalar values")
		}
	}
	return out, nil
}

type MatchStep struct{ LeftIndex, Start, End, MatchIndex int } // -1 means unmatched.
type Pair struct{ Left, Right int }
type Analysis struct {
	Left, Right           []rune
	Radius                int
	Steps                 []MatchStep
	Pairs                 []Pair
	LeftOrder, RightOrder []rune
	Mismatches            int
	Transpositions        float64
	Prefix                int
	Jaro, Boost, Winkler  float64
}

func Analyze(leftText, rightText string) (Analysis, error) {
	left, err := Symbols(leftText)
	if err != nil {
		return Analysis{}, err
	}
	right, err := Symbols(rightText)
	if err != nil {
		return Analysis{}, err
	}
	radius := max(0, max(len(left), len(right))/2-1)
	used := make([]bool, len(right))
	pairs, steps := []Pair{}, []MatchStep{}
	for i, symbol := range left {
		start := min(len(right), max(0, i-radius))
		end := min(len(right), i+radius+1) // Exclusive.
		matchIndex := -1
		for j := start; j < end; j++ {
			if !used[j] && symbol == right[j] {
				used[j] = true
				matchIndex = j
				pairs = append(pairs, Pair{i, j})
				break
			}
		}
		steps = append(steps, MatchStep{i, start, end, matchIndex})
	}
	leftOrder, rightOrder := []rune{}, []rune{}
	for _, pair := range pairs {
		leftOrder = append(leftOrder, left[pair.Left])
	}
	for i, symbol := range right {
		if used[i] {
			rightOrder = append(rightOrder, symbol)
		}
	}
	mismatches := 0
	for i, symbol := range leftOrder {
		if symbol != rightOrder[i] {
			mismatches++
		}
	}
	transpositions := float64(mismatches) / 2 // Retain a possible half.
	m, jaro := float64(len(pairs)), 0.0
	if m == 0 {
		if len(left) == 0 && len(right) == 0 {
			jaro = 1
		}
	} else {
		jaro = (m/float64(len(left)) + m/float64(len(right)) + (m-transpositions)/m) / 3
	}
	prefix := 0
	for prefix < min(4, len(left), len(right)) && left[prefix] == right[prefix] {
		prefix++
	}
	boost := 0.0
	if jaro > 0.7 {
		boost = 0.1 * float64(prefix) * (1 - jaro)
	}
	winkler := jaro + boost
	return Analysis{left, right, radius, steps, pairs, leftOrder, rightOrder, mismatches, transpositions, prefix, jaro, boost, winkler}, nil
}

type Screening struct {
	LeftKey, RightKey string
	Analysis          *Analysis
	Score             *float64
	Status            string
}

func PrepareName(text string) (string, error) {
	values, err := Symbols(text)
	if err != nil {
		return "", err
	}
	var out strings.Builder
	pendingSpace := false
	for _, r := range values {
		if r == ' ' || r == '\t' || r == '\n' || r == '\v' || r == '\f' || r == '\r' {
			pendingSpace = out.Len() > 0
			continue
		}
		if pendingSpace {
			out.WriteByte(' ')
			pendingSpace = false
		}
		if r >= 'a' && r <= 'z' {
			r -= 'a' - 'A'
		}
		out.WriteRune(r)
	}
	return out.String(), nil
}
func ScreenPair(left, right string, cutoff float64, usePrefix bool) (Screening, error) {
	if math.IsNaN(cutoff) || math.IsInf(cutoff, 0) || cutoff < 0 || cutoff > 1 {
		return Screening{}, fmt.Errorf("review cutoff must be between 0 and 1")
	}
	leftKey, err := PrepareName(left)
	if err != nil {
		return Screening{}, err
	}
	rightKey, err := PrepareName(right)
	if err != nil {
		return Screening{}, err
	}
	result := Screening{LeftKey: leftKey, RightKey: rightKey, Status: "missing-name"}
	if leftKey == "" || rightKey == "" {
		return result, nil
	}
	analysis, err := Analyze(leftKey, rightKey)
	if err != nil {
		return Screening{}, err
	}
	score := analysis.Jaro
	if usePrefix {
		score = analysis.Winkler
	}
	result.Analysis = &analysis
	result.Score = &score
	result.Status = "below-cutoff"
	if score >= cutoff {
		result.Status = "review"
	}
	return result, nil
}

func Demo() (string, error) {
	pair, err := ScreenPair("Martha", "Marhta", 0.95, true)
	if err != nil {
		return "", err
	}
	if pair.Score == nil {
		return "More name information needed.", nil
	}
	return fmt.Sprintf("score=%.6f; %s\nKeep both contact records until identity is checked.", *pair.Score, pair.Status), nil
}
func main() {
	output, err := Demo()
	if err != nil {
		panic(err)
	}
	fmt.Println(output)
}

08 / Make the call

A useful comparison has a boundary.

Very short strings make the window tight. AB / BA has radius zero, no matches, and score zero. The generous handling of Marhta doesn’t promise that every adjacent swap scores high.

Repeated symbols create matching choices. We scan the left name in order, take the earliest unused equal letter in its window, then compare the matches in original order. Pin those details if a library replacement must keep existing decisions.

Use trusted identifiers when you have them, and explicit alias rules for known equivalences. Prefer edit distance when the requirement is a transformation cost. Before relying on name-based review, score a labeled set of your own candidate pairs: missed duplicates and unnecessary reviews cost different things, so include the spelling systems, name lengths, and variations your service supports.

You’ll meet this score in data and search tools. Lucene’s spellchecker ships JaroWinklerDistance for “short strings such as person names,” with the same 0.7 boost threshold. Splink, a record-linkage library, compares names at Jaro–Winkler thresholds of 0.9 and 0.7 by default. Snowflake’s function returns an integer and ignores case, where this lesson’s raw scorer is case-sensitive and leaves case to the wrapper. Elasticsearch’s term suggester offers jaro_winkler as a string distance, though its default is based on Damerau–Levenshtein, the transposition-aware edit distance the Levenshtein lesson points to.

Definitions, libraries, and the variant used herePrimary sources and compatibility details

NIST’s Jaro–Winkler entry places the similarity measure in record linkage and describes the common-prefix enhancement. Our contact manager is an authored teaching example.

The pinned strsim 0.11.1 source is a useful comparison: it also uses a four-symbol prefix, weight 0.1, and a base-score gate greater than 0.7. It divides an integer order-mismatch count by two; our implementation keeps fractional halves. For ABCDEF / BCADEF, three order mismatches give our t = 1.5 and Jaro 0.916666…; truncating to one gives 0.944444…. A shared function name doesn’t guarantee the same edge cases.

RapidFuzz’s Jaro–Winkler API exposes prefix weight, an optional input processor, and a similarity cutoff. Its cutoff can return zero for a score below the requested minimum. Our wrapper keeps the score and returns a separate review status so the reader can inspect a near miss. Treat replacement as a contract comparison, not a speed claim.

Here the prefix weight stays fixed, there is no extra long-string adjustment, and raw scoring performs no normalization. Shared cases cover empty inputs, repeated letters, half transpositions, Unicode scalars, prefix toggles, and review boundaries in every language. Sources checked 11 September 2026.

09 / Take the idea with you

Explain a name match without saying “Jaro–Winkler.”

“Line the two names up and count the letters they share close to the same place, using each letter once. Check how many of those shared letters come in a different order. Average how much of each name matched and how much stayed in order, then add a little when the names start the same way. A high number means the names look alike, not that they are the same person.”

For the next pair you inspect, ask: what matched nearby, what disagreed in order, and what was added only because the beginning matched? Then ask what additional evidence would justify joining the records. The score can explain a suggestion; the application’s evidence must justify the decision.

And the next time a tool hands you a 96, you’ll know it counted nearby letters, their order, and a shared beginning, not certainty.

Connections to follow nextRelated lessons

Copy the complete example and swap in two names from a list you work with. Before running it, predict the matched letters, the order mismatches, and whether the shared beginning adds a boost. Then compare the pair with Levenshtein: it counts a minimum set of edits, while Jaro–Winkler rewards nearby matches and a shared prefix.

Back to applied algorithms →