§ 2.8Module 2

Kappa Statistics

On this page

2.8 — Kappa Statistics

Recall first. If two classifiers agree on 90% of cases, is that necessarily strong agreement? What baseline agreement should be removed?

Agreement beyond chance

Cohen’s kappa compares observed agreement with agreement expected from the classifiers’ marginal class frequencies:

κ = (pₒ − pₑ)/(1 − pₑ)

Thus κ=1 is perfect agreement, κ=0 means observed agreement equals the chance-agreement baseline, and negative values mean less agreement than that baseline. Kappa is not a universal “accuracy corrected for all bias”; its interpretation depends on prevalence, marginals, sampling, and the definition of the categories.12

For a model against ground truth, kappa measures agreement between model labels and reference labels. It is also commonly used for inter-rater agreement. It should be reported with the confusion matrix and task-relevant metrics, not in isolation.

Worked numerical calculation

Use the matrix from the previous note:

Predicted +Predicted −Row total
Actual +36440
Actual −105060
Column total4654100

Observed agreement:

pₒ = (36+50)/100 = 0.86

Expected agreement from marginals:

pₑ = (40/100)(46/100) + (60/100)(54/100)
   = 0.184 + 0.324 = 0.508

Therefore:

κ = (0.86−0.508)/(1−0.508) = 0.352/0.492 ≈ 0.715

The model agrees at 86%, but after removing the agreement expected from these marginals, kappa is about 0.715. Do not attach a universal adjective such as “good” without context; published qualitative bands are conventions, not laws.

Weighted kappa and cautions

For ordered categories (for example, mild/moderate/severe), a one-category disagreement may be less serious than a large disagreement. Weighted kappa assigns disagreement weights; unweighted kappa treats all off-diagonal disagreements alike. The weighting scheme must be stated.

Kappa can behave unexpectedly when prevalence is very imbalanced or when marginal distributions differ greatly. A high accuracy can coexist with low kappa, and kappa can fall even when accuracy rises if chance agreement rises faster. Inspect sensitivity/specificity and the raw table before interpreting it.1

Exercise

Suppose pₒ=0.80 and pₑ=0.50. Compute κ. What would happen if pₒ stayed 0.80 but pₑ increased to 0.70?

Revealed answer

First κ=(0.80−0.50)/(1−0.50)=0.30/0.50=0.60. With pₑ=0.70, κ=0.10/0.30≈0.333. The same observed agreement gives lower kappa when the marginal-based chance agreement is higher.

Exam lens

Show pₒ, compute expected agreement from row/column marginals, then substitute into the formula. Explain that kappa adjusts for chance agreement under an independence model and is sensitive to prevalence/marginals.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, cohen_kappa_score API. ↩ ↩2

  2. Cohen, J. (1960), “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement; DOI record. ↩