§ 3.8Module 3

Random Forest Compared with Boosting

On this page

3.8 — Random Forest Compared with Boosting

Recall first. Which ensemble trains trees independently and which makes later trees depend on earlier errors? Which one is easier to parallelize?

Side-by-side comparison

DimensionRandom forest / baggingBoosting (including XGBoost)
Tree relationshipIndependent/parallelSequential/additive
Main mechanismBootstrap rows + random feature subsetsFit next learner to loss gradient or weighted errors
Typical effectMainly lowers varianceOften lowers bias and can lower variance with regularization
Base treesOften deeper, randomized treesOften shallow/controlled trees
ParallelismHigh across treesLimited across boosting rounds
Main risksCorrelated trees, large model, residual biasNoise sensitivity, overfitting, tuning sensitivity
ControlsTrees, features/split, depth, leaf sizeLearning rate, rounds, depth/leaves, subsampling, regularization, early stopping

These are tendencies, not guarantees. A forest can be biased if features are weak; boosting can be regularized and robust; validation on the actual task decides.12

Choosing by problem behavior

Choose a forest as a strong low-tuning baseline when you want robustness, parallel training, and a model that is relatively insensitive to scaling. Choose boosting when careful validation/tuning is affordable and small structured corrections can produce high accuracy. For noisy labels and outliers, inspect boosting carefully because repeatedly emphasizing hard cases can emphasize noise. For high-dimensional sparse data, compare appropriate linear and tree methods rather than assuming either ensemble wins.

Neither model automatically provides causal explanation. Both can expose feature importance, but importance is conditional on the data, model, and correlated predictors.

Worked decision trace

A dataset has 100,000 rows, noisy labels, and a need for a quick baseline. A random forest can train trees independently, use OOB diagnostics, and provide a reasonable first result. A second experiment uses a clean, structured tabular dataset where validation can support careful tuning. Gradient boosting with shallow trees, a small learning rate, and early stopping is a plausible candidate. In both cases, split by time/group if needed, compare to a baseline, and retain a final test set.

Exercise

Match the clue to the likely first choice: (A) training rounds must be highly parallel, (B) you can tune carefully for maximum tabular accuracy, (C) later learners must focus on previous residuals.

Revealed answer

A → random forest/bagging; B → boosting is a plausible candidate; C → boosting. These are starting choices, not guarantees of performance.

Exam lens

Use the comparison table: independence vs sequence, variance vs bias, parallelism, base-tree depth, and overfitting controls. Avoid absolute statements such as “boosting is always more accurate.”

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Ensemble methods. ↩

  2. Hastie, Tibshirani & Friedman, ESL, chapters 8–10; XGBoost, parameter tuning. ↩