§ 3.6Module 3

Bagging and Subagging

On this page

3.6 — Bagging and Subagging

Recall first. If a single tree is unstable, should we train many trees on exactly the same data and average them? What must differ between the fits?

Bagging: bootstrap aggregation

Bagging fits a base learner independently on many bootstrap samples. Each sample draws n observations with replacement from the n training observations. Aggregate predictions by:

The resamples make fitted unstable learners differ. Averaging then mainly reduces variance, while the model can be trained in parallel. A bootstrap sample contains about 63.2% of unique training observations on average; the remaining out-of-bag (OOB) observations can estimate performance for that learner. This is a standard approximation, not a replacement for a final deployment-aware test.12

Subagging

Subagging (subsample aggregation) fits learners on subsamples drawn without replacement, often smaller than the full dataset, and aggregates them. It reduces dependence and variance without bootstrap duplicates. The exact subsample size and whether all learners use equal size are design choices. The distinction to memorize is replacement: bagging with replacement; subagging without replacement.

Worked sample trace

Training IDs are {A,B,C,D,E}. A bootstrap sample might be [A,C,C,E,B]: five draws, but four unique IDs; D is OOB for that learner. A subagging sample of size 3 might be [A,D,E]: no duplicate IDs. If three regression trees predict 8, 10, and 12, bagging prediction is their mean 10. For classifiers voting [yes, no, yes], the hard vote is yes; soft vote would average probabilities if supplied.

Why aggregation helps

If each base learner has error variance and errors are not perfectly correlated, averaging reduces the independent component. Bagging is most useful for high-variance learners such as unpruned trees; it often changes bias little. It may not help a stable high-bias model, and correlated training data or leakage can make the ensemble look better than it is. OOB evaluation is useful during development, but hyperparameter selection and final evaluation still need honest separation.1

Exercise

Which sampling method can contain duplicate rows: bagging or subagging? If five trees predict classes [A,B,A,A,B], what is the hard-vote result?

Revealed answer

Bagging can contain duplicate rows because it samples with replacement; subagging does not. Class A has three votes versus two for B, so the hard-vote result is A.

Exam lens

Give the algorithm: resample → fit independently → aggregate. State with/without replacement, variance reduction, parallelism, and OOB intuition. Contrast subagging in one sentence.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Bagging meta-estimator. ↩ ↩2

  2. Breiman, “Bagging Predictors,” paper; Hastie, Tibshirani & Friedman, ESL, chapter 8. ↩