§ 3.1Module 3

Understanding Ensembles

On this page

3.1 — Understanding Ensembles

Recall first. Why might three imperfect, different classifiers outperform one classifier? Name the condition about their errors that matters more than simply having three models.

The central idea

An ensemble combines predictions from multiple base learners to produce one prediction. The base learners may be trees, linear models, or other classifiers. Combining helps when learners are individually useful but make partly different errors. Merely duplicating the same model and error pattern adds little.12

For regression, averaging M predictions is f̄(x)=M⁻¹Σ f_m(x). If errors have mean zero, variance σ², and pairwise correlation ρ, the variance of the average is approximately:

Var(f̄) = σ²[ρ + (1−ρ)/M]

As M grows, the independent part shrinks; the correlated part remains. This is why accuracy and diversity must be balanced. For classification, hard voting takes the majority class; soft voting averages predicted probabilities, which requires probabilities to be comparable/calibrated enough for the task.1

Three ways an ensemble learns

A base learner deliberately kept simple is a weak learner. “Weak” means slightly better than a baseline on the relevant task, not necessarily useless.

Worked variance intuition

Suppose three base regressors each have error variance 4. If errors are independent (ρ=0), the average has variance 4/3≈1.33. If errors are perfectly correlated (ρ=1), the average still has variance 4: every model makes the same mistake. If ρ=0.25, variance is 4[0.25+0.75/3]=2. Diversity lowers correlation and improves averaging.

Exercise

Four classifiers vote [A,A,B,B]. Is there a unique majority? If their probabilities for class A are [0.9,0.8,0.4,0.3], what is the soft-vote probability and what extra issue should be considered?

Revealed answer

Hard voting ties, so a tie rule is needed. Mean probability for A is (0.9+0.8+0.4+0.3)/4=0.60, so soft voting selects A under a 0.5 threshold. The probabilities should be reasonably calibrated/comparable; otherwise averaging scores can be misleading.

Exam lens

Define ensemble learning, explain the accuracy–diversity trade-off, then contrast bagging (parallel variance reduction), boosting (sequential error correction), and stacking/voting (output combination). Include the correlated-error intuition.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Ensemble methods User Guide. ↩ ↩2

  2. Hastie, Tibshirani & Friedman, ESL, chapters 8–10; Breiman, “Bagging Predictors,” paper. ↩