Understanding Ensembles
On this page
3.1 — Understanding Ensembles
Recall first. Why might three imperfect, different classifiers outperform one classifier? Name the condition about their errors that matters more than simply having three models.
The central idea
An ensemble combines predictions from multiple base learners to produce one prediction. The base learners may be trees, linear models, or other classifiers. Combining helps when learners are individually useful but make partly different errors. Merely duplicating the same model and error pattern adds little.12
For regression, averaging M predictions is f̄(x)=M⁻¹Σ f_m(x). If errors have mean zero, variance σ², and pairwise correlation ρ, the variance of the average is approximately:
Var(f̄) = σ²[ρ + (1−ρ)/M]
As M grows, the independent part shrinks; the correlated part remains. This is why accuracy and diversity must be balanced. For classification, hard voting takes the majority class; soft voting averages predicted probabilities, which requires probabilities to be comparable/calibrated enough for the task.1
Three ways an ensemble learns
- Bagging: fit models independently on resampled data and aggregate; mainly reduces variance.
- Boosting: fit models sequentially, giving later learners emphasis on previous errors; can reduce bias but may overfit.
- Stacking/voting: combine outputs directly or learn a meta-model; flexible but risks leakage if meta-features are not out-of-fold.
A base learner deliberately kept simple is a weak learner. “Weak” means slightly better than a baseline on the relevant task, not necessarily useless.
Worked variance intuition
Suppose three base regressors each have error variance 4. If errors are independent (ρ=0), the average has variance 4/3≈1.33. If errors are perfectly correlated (ρ=1), the average still has variance 4: every model makes the same mistake. If ρ=0.25, variance is 4[0.25+0.75/3]=2. Diversity lowers correlation and improves averaging.
Exercise
Four classifiers vote [A,A,B,B]. Is there a unique majority? If their probabilities for class A are [0.9,0.8,0.4,0.3], what is the soft-vote probability and what extra issue should be considered?
Revealed answer
Hard voting ties, so a tie rule is needed. Mean probability for A is (0.9+0.8+0.4+0.3)/4=0.60, so soft voting selects A under a 0.5 threshold. The probabilities should be reasonably calibrated/comparable; otherwise averaging scores can be misleading.
Exam lens
Define ensemble learning, explain the accuracy–diversity trade-off, then contrast bagging (parallel variance reduction), boosting (sequential error correction), and stacking/voting (output combination). Include the correlated-error intuition.
Rapid revision checklist
- Can I explain why correlated errors limit averaging?
- Can I distinguish hard and soft voting?
- Can I map bagging, boosting, and stacking to their training patterns?
- Can I state one leakage risk in stacking?
Key takeaways
- Ensembles combine imperfect learners; diversity of errors is the resource.
- Averaging reduces independent variance, not shared error.
- Bagging is parallel, boosting sequential, and stacking learns or applies a combination rule.
- Soft voting requires meaningful comparable probabilities.