Random Forest Compared with Boosting
On this page
3.8 — Random Forest Compared with Boosting
Recall first. Which ensemble trains trees independently and which makes later trees depend on earlier errors? Which one is easier to parallelize?
Side-by-side comparison
| Dimension | Random forest / bagging | Boosting (including XGBoost) |
|---|---|---|
| Tree relationship | Independent/parallel | Sequential/additive |
| Main mechanism | Bootstrap rows + random feature subsets | Fit next learner to loss gradient or weighted errors |
| Typical effect | Mainly lowers variance | Often lowers bias and can lower variance with regularization |
| Base trees | Often deeper, randomized trees | Often shallow/controlled trees |
| Parallelism | High across trees | Limited across boosting rounds |
| Main risks | Correlated trees, large model, residual bias | Noise sensitivity, overfitting, tuning sensitivity |
| Controls | Trees, features/split, depth, leaf size | Learning rate, rounds, depth/leaves, subsampling, regularization, early stopping |
These are tendencies, not guarantees. A forest can be biased if features are weak; boosting can be regularized and robust; validation on the actual task decides.12
Choosing by problem behavior
Choose a forest as a strong low-tuning baseline when you want robustness, parallel training, and a model that is relatively insensitive to scaling. Choose boosting when careful validation/tuning is affordable and small structured corrections can produce high accuracy. For noisy labels and outliers, inspect boosting carefully because repeatedly emphasizing hard cases can emphasize noise. For high-dimensional sparse data, compare appropriate linear and tree methods rather than assuming either ensemble wins.
Neither model automatically provides causal explanation. Both can expose feature importance, but importance is conditional on the data, model, and correlated predictors.
Worked decision trace
A dataset has 100,000 rows, noisy labels, and a need for a quick baseline. A random forest can train trees independently, use OOB diagnostics, and provide a reasonable first result. A second experiment uses a clean, structured tabular dataset where validation can support careful tuning. Gradient boosting with shallow trees, a small learning rate, and early stopping is a plausible candidate. In both cases, split by time/group if needed, compare to a baseline, and retain a final test set.
Exercise
Match the clue to the likely first choice: (A) training rounds must be highly parallel, (B) you can tune carefully for maximum tabular accuracy, (C) later learners must focus on previous residuals.
Revealed answer
A → random forest/bagging; B → boosting is a plausible candidate; C → boosting. These are starting choices, not guarantees of performance.
Exam lens
Use the comparison table: independence vs sequence, variance vs bias, parallelism, base-tree depth, and overfitting controls. Avoid absolute statements such as “boosting is always more accurate.”
Rapid revision checklist
- Can I contrast training dependence and parallelism?
- Can I state the typical bias/variance role of each?
- Can I name controls for each ensemble?
- Can I explain why validation overrides textbook preferences?
Key takeaways
- Forests randomize and average; boosting sequentially corrects.
- Forests are easier to parallelize and often robust; boosting is more tuning-sensitive but powerful.
- Both can overfit and both need honest validation.
- “Better” depends on data, objective, noise, compute, and deployment constraints.
Sources
Footnotes
-
scikit-learn, Ensemble methods. ↩
-
Hastie, Tibshirani & Friedman, ESL, chapters 8–10; XGBoost, parameter tuning. ↩