A “suspicious” sign-in is a prediction, not a confirmed takeover.
For this lesson, the event is a sign-in attempt later confirmed to be an account takeover by the investigation team. The validation set contains 10,000 labeled attempts: 100 takeovers and 9,900 legitimate sign-ins. That 1% prevalence is illustrative. It gives us a population and a denominator; it is not a claim about any real authentication service.
The detector assigns a risk score, and the policy sends attempts above a chosen threshold to review. That creates four outcomes. A caught takeover is a true positive. A legitimate sign-in sent to review is a false positive. A missed takeover is a false negative. A legitimate sign-in left alone is a true negative. “False” describes the prediction relative to the current label; it does not tell us whether the threshold or the label process was reasonable.
- Event
- A takeover is later confirmed for this sign-in attempt.
- Population
- 100 confirmed takeovers and 9,900 legitimate attempts.
- Score action
- Send an attempt above threshold for investigation.
- Decision
- Which threshold supports the response policy and available capacity?
The confusion matrix is a count table, before it is a score.
At the middle threshold, the detector flags 278 attempts. Eighty are labeled takeovers and 198 are labeled legitimate. Among the 9,722 attempts it does not flag, 20 are takeovers and 9,702 are legitimate. Each number describes an intersection of actual condition and model prediction; all four cells must add back to the evaluated cohort.
These are constructed counts so the arithmetic is visible. In a real evaluation, labels might arrive days later, some takeovers might never be reported, and “legitimate” may include unrecognized compromise. Record the labeling rule and observation window with the matrix.
| Actual condition | Flagged for review | Not flagged | Actual total |
|---|---|---|---|
| Takeover | 80 true positives | 20 false negatives | 100 |
| Legitimate sign-in | 198 false positives | 9,702 true negatives | 9,900 |
| Predicted total | 278 | 9,722 | 10,000 |
Flagged plus missed events.
Unnecessary reviews plus correctly unflagged attempts.
Rows and columns must reconcile to the same population.
Keep the question attached to the fraction.
Precision is TP / (TP + FP): of all attempts flagged, what
fraction were labeled takeovers? Here that is 80 / (80 + 198) = 80 / 278 ≈ 28.8%. The denominator is the flagged queue. Its
complement, FP / (TP + FP), is the share of flags that were false positives
under these labels.
Recall, also called sensitivity or true-positive rate, is TP / (TP + FN): of the actual takeovers in the cohort, what fraction did the
detector flag? Here it is 80 / (80 + 20) = 80 / 100 = 80%. The denominator is
actual takeovers, not all flags. Its complement, FN / (TP + FN), is the miss
rate.
Accuracy is (TP + TN) / N, or the share of all predictions
that match the labels. Here it is (80 + 9,702) / 10,000 = 97.82%. But if a
detector flags nothing, it gets all 9,900 legitimate attempts “right” and misses all 100
takeovers: 99% accuracy, 0% recall. When positives are rare, a large true-negative count
can dominate accuracy.
TP divided by all flagged attempts.
TP divided by all actual takeovers.
Correct predictions divided by all attempts.
More catches usually put more legitimate sign-ins in the queue.
A score threshold converts a ranking into an action. Lower it and more attempts are flagged; raise it and fewer are. That often increases recall and lowers precision at the lower threshold, but an actual precision-recall curve must be measured from scored labeled data. The three settings below are an illustrative validation snapshot, not a guarantee that every model or population follows the same neat pattern.
Compare threshold consequences
A middle setting · still requires a policy choice. The cohort stays fixed at 100 labeled takeovers and 9,900 legitimate sign-ins.
| Actual condition | Flagged | Not flagged | Total |
|---|---|---|---|
| Takeover | 80 true positives | 20 false negatives | 100 |
| Legitimate | 198 false positives | 9702 true negatives | 9,900 |
| Predicted total | 278 | 9722 | 10,000 |
These are illustrative counts, not live measurements. In production, select thresholds on a validation set, then evaluate on a separate time period or cohort before changing policy.
| Setting | TP / FP | FN | Precision | Recall | Queue size |
|---|---|---|---|---|---|
| Broad / lower | 94 / 890 | 6 | 9.6% | 94% | 984 |
| Middle | 80 / 198 | 20 | 28.8% | 80% | 278 |
| Strict / higher | 62 / 35 | 38 | 63.9% | 62% | 97 |
Before changing a cutoff, find out what moved.
A week-over-week precision drop could mean the score separates events less well. It could also mean takeovers are less prevalent, the score threshold changed, the analyst team labels more cases as legitimate, or a new region contributes different traffic. The metric is an observation; each explanation makes a different prediction. Check those predictions before turning the threshold knob.
Start with the raw four counts by week, then slice by region, device class, new versus returning account, and score band where sample sizes support a comparison. Verify that every attempt had a chance to be labeled: only investigating high-score attempts creates a selective-label problem, because missed takeovers in the unreviewed group remain unknown. Also record label delay; recent attempts may not have had enough time to become confirmed.
| Possible explanation | What it predicts | Evidence to inspect |
|---|---|---|
| Takeovers became rarer | Precision may fall while conditional recall and false-positive rate remain similar. | Prevalence from matured labels; compare the same definition and time window. |
| Score separation changed | At the same threshold, both missed takeovers and legitimate flags may shift. | Score distributions and confusion counts by label, model version, and cohort. |
| Label process changed | Recent or unreviewed cases may be counted as legitimate before evidence matures. | Label source, review coverage, appeal outcomes, and time to confirmation. |
| Traffic mix shifted | Aggregate metrics move while within-region or device metrics stay steadier. | Prevalence and rates by relevant cohort, with counts and uncertainty. |
Compute from counts and make empty denominators visible.
The implementations accept observed whole-number counts, validate that none are negative,
and calculate precision, recall, and accuracy. For the middle setting they return
precision about 0.288, recall 0.8, and accuracy 0.9782. The fractions are unitless; report them with the counts and
evaluation population so they do not float free of the data they summarize.
If no attempt was flagged, precision has a zero denominator and is undefined—not zero. If
there were no positive labels, recall is undefined. The TypeScript result uses null for those cases; Go uses explicit HasPrecision, HasRecall, and HasAccuracy flags. Neither example checks whether the labels are correct or the
cohort represents deployment.
Both implementations preserve undefined metrics when a denominator is zero.
export type ConfusionMatrix = {
truePositive: number;
falsePositive: number;
falseNegative: number;
trueNegative: number;
};
export type ClassificationMetrics = ConfusionMatrix & {
precision: number | null;
recall: number | null;
accuracy: number | null;
};
/** Calculate classification metrics from observed, non-negative whole-number counts. */
export function classificationMetrics(matrix: ConfusionMatrix): ClassificationMetrics {
for (const [name, value] of Object.entries(matrix)) {
if (!Number.isSafeInteger(value) || value < 0) {
throw new Error(`${name} must be a non-negative safe integer`);
}
}
const { truePositive: tp, falsePositive: fp, falseNegative: fn, trueNegative: tn } = matrix;
const predictedPositive = tp + fp;
const actualPositive = tp + fn;
const total = tp + fp + fn + tn;
if (!Number.isSafeInteger(total)) {
throw new Error('sum of counts must be a safe integer');
}
return {
...matrix,
precision: predictedPositive === 0 ? null : tp / predictedPositive,
recall: actualPositive === 0 ? null : tp / actualPositive,
accuracy: total === 0 ? null : (tp + tn) / total
};
}
package mathpractice
import (
"errors"
)
type ConfusionMatrix struct {
TruePositive int64
FalsePositive int64
FalseNegative int64
TrueNegative int64
}
type ClassificationMetrics struct {
ConfusionMatrix
Precision float64
Recall float64
Accuracy float64
HasPrecision bool
HasRecall bool
HasAccuracy bool
}
// ClassificationMetricsFromCounts calculates ratios from observed counts.
// A Has* flag is false when the corresponding denominator is zero.
func ClassificationMetricsFromCounts(m ConfusionMatrix) (ClassificationMetrics, error) {
counts := []int64{m.TruePositive, m.FalsePositive, m.FalseNegative, m.TrueNegative}
for _, count := range counts {
if count < 0 || count > (1<<53)-1 {
return ClassificationMetrics{}, errors.New("counts must be between 0 and 2^53-1")
}
}
// Each input is at most 2^53-1, so summing four remains within int64.
predictedPositive := m.TruePositive + m.FalsePositive
actualPositive := m.TruePositive + m.FalseNegative
total := predictedPositive + m.FalseNegative + m.TrueNegative
result := ClassificationMetrics{ConfusionMatrix: m}
if predictedPositive > 0 {
result.Precision = float64(m.TruePositive) / float64(predictedPositive)
result.HasPrecision = true
}
if actualPositive > 0 {
result.Recall = float64(m.TruePositive) / float64(actualPositive)
result.HasRecall = true
}
if total > 0 {
result.Accuracy = float64(m.TruePositive+m.TrueNegative) / float64(total)
result.HasAccuracy = true
}
return result, nil
}
The same catch and false-alarm rates can produce different precision.
Keep the middle setting's recall at 80% and false-positive rate at 2%: 198 / 9,900 = 2%. Now evaluate 10,000 attempts where takeovers are only 0.1% of the population: 10
takeovers and 9,990 legitimate attempts. The model would be expected to flag 8 takeovers
and miss 2, while also flagging about 9,990 × 0.02 = 199.8 legitimate attempts.
These are expected counts, so a fraction is appropriate in the calculation; actual observed
counts must be whole.
Expected precision becomes 8 / (8 + 199.8) ≈ 3.85%, even though recall and
false-positive rate were held constant. The detector's precision is not a permanent
property that transfers everywhere. It depends on the population mix as well as the
measured conditional rates. In smaller subgroups, sampling variation can make the observed
percentage jump around too.
Before transporting a scorecard to a new region, account age, device class, or season, ask for mature labels and the four counts within that target population. If prevalence changes, model the precision consequence; if recall or false-positive rate changes, investigate score or process drift. Do not let a new aggregate conceal a harmed subgroup.
Takeovers in the transfer population.
Expected catches at 80% recall.
Expected flags among legitimate attempts.
Expected true alerts divided by all expected flags.