Sensitivity and Specificity
On this page
2.9 — Sensitivity and Specificity
Recall first. In a screening test, which denominator belongs to “find the positives” and which belongs to “reject the negatives”? Write both fractions.
The two conditional rates
From a confusion matrix:
Sensitivity = TP/(TP+FN) = P(predicted + | actual +)
Specificity = TN/(TN+FP) = P(predicted − | actual −)
Sensitivity (true-positive rate, TPR, and recall) asks: of all actual positives, how many did we detect? Specificity (true-negative rate, TNR) asks: of all actual negatives, how many did we correctly reject? The denominators are different: actual positives for sensitivity, actual negatives for specificity.12
A false-negative-sensitive application may choose a threshold for high sensitivity; a false-positive-sensitive application may prioritize specificity. There is usually a trade-off as the score threshold moves. Neither rate alone describes prevalence or the reliability of a positive prediction—that is the role of precision/PPV.
Worked calculation
For TP=36, FN=4, FP=10, TN=50:
Sensitivity = 36/(36+4) = 36/40 = 0.90 = 90%
Specificity = 50/(50+10) = 50/60 ≈ 0.833 = 83.3%
The test detects 90% of actual positives and correctly rejects about 83.3% of actual negatives. It does not mean 90% of positive predictions are correct; that would be precision 36/(36+10).
Threshold and prevalence
A probabilistic classifier ranks cases by score. Lowering the positive threshold usually increases TP and may decrease FN (higher sensitivity), while also increasing FP (lower specificity). This is a property of the operating point, not a reason to call one metric universally better.
Sensitivity and specificity are conditional on the actual class and are therefore less directly changed by prevalence than predictive values, assuming the test’s conditional behavior remains stable. In deployment, prevalence shift and spectrum effects can still alter observed performance.2
Exercise
A test has TP=90, FN=10, FP=45, TN=855. Calculate sensitivity and specificity. Which error type is more common in count terms?
Revealed answer
Sensitivity =90/(90+10)=0.90=90%. Specificity =855/(855+45)=0.95=95%. False positives total 45, false negatives total 10, so false positives are more common, even though the specificity is high because there are many more negatives.
Exam lens
Use verbal checks: sensitivity denominator = all actual positives; specificity denominator = all actual negatives. State synonyms TPR/recall and TNR. Contrast with precision, whose denominator is predicted positives.
Rapid revision checklist
- Can I write sensitivity and specificity without swapping denominators?
- Can I identify TPR and TNR?
- Can I explain what lowering a threshold tends to do?
- Can I distinguish sensitivity from precision?
Key takeaways
- Sensitivity measures detection of actual positives.
- Specificity measures rejection of actual negatives.
- They trade off with a classification threshold and express different error priorities.
- High sensitivity does not imply high precision, especially with rare positives.
Sources
Footnotes
-
scikit-learn,
recall_scoreAPI and classification metrics guide. ↩ -
Hastie, Tibshirani & Friedman, ESL, chapter 7; NCBI Bookshelf, Sensitivity and Specificity. ↩ ↩2