Kappa Statistics
On this page
2.8 — Kappa Statistics
Recall first. If two classifiers agree on 90% of cases, is that necessarily strong agreement? What baseline agreement should be removed?
Agreement beyond chance
Cohen’s kappa compares observed agreement with agreement expected from the classifiers’ marginal class frequencies:
κ = (pₒ − pₑ)/(1 − pₑ)
pₒ= observed agreement, usually accuracy = diagonal total /N.pₑ= expected agreement if the two labelers/predictors independently followed their observed class proportions.
Thus κ=1 is perfect agreement, κ=0 means observed agreement equals the chance-agreement baseline, and negative values mean less agreement than that baseline. Kappa is not a universal “accuracy corrected for all bias”; its interpretation depends on prevalence, marginals, sampling, and the definition of the categories.12
For a model against ground truth, kappa measures agreement between model labels and reference labels. It is also commonly used for inter-rater agreement. It should be reported with the confusion matrix and task-relevant metrics, not in isolation.
Worked numerical calculation
Use the matrix from the previous note:
| Predicted + | Predicted − | Row total | |
|---|---|---|---|
| Actual + | 36 | 4 | 40 |
| Actual − | 10 | 50 | 60 |
| Column total | 46 | 54 | 100 |
Observed agreement:
pₒ = (36+50)/100 = 0.86
Expected agreement from marginals:
pₑ = (40/100)(46/100) + (60/100)(54/100)
= 0.184 + 0.324 = 0.508
Therefore:
κ = (0.86−0.508)/(1−0.508) = 0.352/0.492 ≈ 0.715
The model agrees at 86%, but after removing the agreement expected from these marginals, kappa is about 0.715. Do not attach a universal adjective such as “good” without context; published qualitative bands are conventions, not laws.
Weighted kappa and cautions
For ordered categories (for example, mild/moderate/severe), a one-category disagreement may be less serious than a large disagreement. Weighted kappa assigns disagreement weights; unweighted kappa treats all off-diagonal disagreements alike. The weighting scheme must be stated.
Kappa can behave unexpectedly when prevalence is very imbalanced or when marginal distributions differ greatly. A high accuracy can coexist with low kappa, and kappa can fall even when accuracy rises if chance agreement rises faster. Inspect sensitivity/specificity and the raw table before interpreting it.1
Exercise
Suppose pₒ=0.80 and pₑ=0.50. Compute κ. What would happen if pₒ stayed 0.80 but pₑ increased to 0.70?
Revealed answer
First κ=(0.80−0.50)/(1−0.50)=0.30/0.50=0.60. With pₑ=0.70, κ=0.10/0.30≈0.333. The same observed agreement gives lower kappa when the marginal-based chance agreement is higher.
Exam lens
Show pₒ, compute expected agreement from row/column marginals, then substitute into the formula. Explain that kappa adjusts for chance agreement under an independence model and is sensitive to prevalence/marginals.
Rapid revision checklist
- Can I define observed and expected agreement?
- Can I calculate
pₑfrom row and column proportions? - Can I compute κ from a confusion matrix?
- Can I state two reasons not to use kappa alone?
Key takeaways
- Kappa measures agreement beyond a marginal-frequency chance baseline.
κ=(pₒ−pₑ)/(1−pₑ); perfect agreement is 1 and chance-level agreement is 0.- Prevalence and marginal imbalance strongly affect interpretation.
- Report kappa with the confusion matrix and class-specific metrics.
Sources
Footnotes
-
scikit-learn,
cohen_kappa_scoreAPI. ↩ ↩2 -
Cohen, J. (1960), “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement; DOI record. ↩