Precision, Recall, and F-measure
On this page
2.10 — Precision, Recall, and F-measure
Recall first. “90% of alerts were truly fraud” and “90% of fraud was detected” sound similar. Which metric describes each sentence?
Two different denominators
Precision = TP/(TP+FP) = P(actual + | predicted +)
Recall = TP/(TP+FN) = P(predicted + | actual +)
Precision (positive predictive value) asks whether positive alerts are trustworthy. Recall (sensitivity/TPR) asks whether actual positives are found. Precision is harmed by false positives; recall is harmed by false negatives. This distinction is central under class imbalance.1
The F-measure combines precision and recall through their harmonic mean:
F₁ = 2PR/(P+R)
The harmonic mean is low when either component is low, so F1 rewards balance rather than allowing one metric to hide a near-zero other metric. More generally:
Fβ = (1+β²)PR/(β²P+R)
β>1 weights recall more; β<1 weights precision more. F1 is a threshold-dependent summary and does not include TN, so it is not a complete replacement for a confusion matrix or specificity.
Worked calculation
For TP=36, FP=10, FN=4, TN=50:
P = 36/(36+10) = 36/46 ≈ 0.783
R = 36/(36+4) = 36/40 = 0.900
F₁ = 2(0.783)(0.900)/(0.783+0.900) ≈ 0.837
Equivalent direct form: F₁=2TP/(2TP+FP+FN)=72/(72+10+4)=72/86≈0.837; using rounded precision and recall gives the same result to three decimals. Precision says roughly 78.3% of alerts are correct; recall says 90% of actual positives are caught.
Threshold and averaging
For a score model, changing the threshold changes the confusion matrix and therefore precision, recall, and F1. Choose a threshold using validation data and the costs of actions—not the test set.
For multiclass problems, scikit-learn offers macro, weighted, and micro averaging. Macro averages class metrics equally and exposes poor minority-class performance; weighted weights by class support; micro pools decisions and can be dominated by common classes. State the averaging choice.1
Exercise
A classifier has TP=8, FP=2, FN=12. Compute precision, recall, and F1. Which metric reveals that many positives were missed?
Revealed answer
Precision =8/(8+2)=0.80. Recall =8/(8+12)=0.40. F1=2(0.8)(0.4)/(0.8+0.4)=0.64/1.2≈0.533. Recall reveals the 12 false negatives and poor detection of actual positives.
Exam lens
Say “precision denominator = predicted positives; recall denominator = actual positives.” Then write F1 and compute. Mention F1 ignores TN and depends on a threshold; use specificity/ROC or a cost metric when negatives matter.
Rapid revision checklist
- Can I state precision and recall in words and formulas?
- Can I calculate F1 from a confusion matrix?
- Can I explain why the harmonic mean penalizes imbalance?
- Can I distinguish F1 from accuracy and specificity?
- Can I explain macro versus weighted averaging?
Key takeaways
- Precision measures alert reliability; recall measures positive-case coverage.
- F1 balances them and becomes small when either is small.
- Threshold and averaging choices must be reported.
- F1 does not use TN, so it is not a full diagnostic summary.
Sources
Footnotes
-
scikit-learn, precision/recall/F-score metrics and
f1_scoreAPI. ↩ ↩2