§ 2.10Module 2

Precision, Recall, and F-measure

On this page

2.10 — Precision, Recall, and F-measure

Recall first. “90% of alerts were truly fraud” and “90% of fraud was detected” sound similar. Which metric describes each sentence?

Two different denominators

Precision = TP/(TP+FP) = P(actual + | predicted +)
Recall    = TP/(TP+FN) = P(predicted + | actual +)

Precision (positive predictive value) asks whether positive alerts are trustworthy. Recall (sensitivity/TPR) asks whether actual positives are found. Precision is harmed by false positives; recall is harmed by false negatives. This distinction is central under class imbalance.1

The F-measure combines precision and recall through their harmonic mean:

F₁ = 2PR/(P+R)

The harmonic mean is low when either component is low, so F1 rewards balance rather than allowing one metric to hide a near-zero other metric. More generally:

Fβ = (1+β²)PR/(β²P+R)

β>1 weights recall more; β<1 weights precision more. F1 is a threshold-dependent summary and does not include TN, so it is not a complete replacement for a confusion matrix or specificity.

Worked calculation

For TP=36, FP=10, FN=4, TN=50:

P = 36/(36+10) = 36/46 ≈ 0.783
R = 36/(36+4)  = 36/40 = 0.900
F₁ = 2(0.783)(0.900)/(0.783+0.900) ≈ 0.837

Equivalent direct form: F₁=2TP/(2TP+FP+FN)=72/(72+10+4)=72/86≈0.837; using rounded precision and recall gives the same result to three decimals. Precision says roughly 78.3% of alerts are correct; recall says 90% of actual positives are caught.

Threshold and averaging

For a score model, changing the threshold changes the confusion matrix and therefore precision, recall, and F1. Choose a threshold using validation data and the costs of actions—not the test set.

For multiclass problems, scikit-learn offers macro, weighted, and micro averaging. Macro averages class metrics equally and exposes poor minority-class performance; weighted weights by class support; micro pools decisions and can be dominated by common classes. State the averaging choice.1

Exercise

A classifier has TP=8, FP=2, FN=12. Compute precision, recall, and F1. Which metric reveals that many positives were missed?

Revealed answer

Precision =8/(8+2)=0.80. Recall =8/(8+12)=0.40. F1=2(0.8)(0.4)/(0.8+0.4)=0.64/1.2≈0.533. Recall reveals the 12 false negatives and poor detection of actual positives.

Exam lens

Say “precision denominator = predicted positives; recall denominator = actual positives.” Then write F1 and compute. Mention F1 ignores TN and depends on a threshold; use specificity/ROC or a cost metric when negatives matter.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, precision/recall/F-score metrics and f1_score API. ↩ ↩2