Random Forest
On this page
3.7 — Random Forest
Recall first. Bagging trees reduces variance, but what makes every tree less correlated than ordinary bagging?
Bagged trees plus random feature selection
A random forest grows many decision trees, typically each on a bootstrap sample, and at each split considers only a random subset of features. The final prediction is majority vote/class-probability average or regression mean. Row resampling and feature subsampling both increase diversity; aggregation reduces variance.12
Ordinary bagging lets every split examine all features. If one very strong feature dominates, all trees may choose similar early splits and remain correlated. Randomly limiting candidate features makes alternative features useful in different trees. The forest sacrifices some individual-tree strength to improve the ensemble’s average.
Out-of-bag estimation and controls
For each tree, observations not selected in its bootstrap sample are OOB. Aggregate predictions for each training observation only from trees for which it was OOB, producing an OOB score. This is useful for an internal estimate, but group/time dependence and tuning still require care.
Controls include number of trees, number of features considered per split (max_features), maximum depth, minimum leaf size, class weights, and bootstrap/subsample choices. More trees generally stabilize the estimate until computational returns diminish; depth and leaf size control individual-tree variance. Feature scaling is usually unnecessary for tree thresholds.1
Worked trace
Suppose three trees predict class labels [defect, normal, defect]; the forest predicts defect by majority vote. For a regression target, predictions [12,15,18] produce 15. If a forest has five features but max_features=2, each split randomly considers two of the five; another split can draw a different pair. This is feature randomness, not random deletion of features from the whole dataset.
Limitations
Random forests are strong general-purpose baselines, but they can be large, less transparent than one tree, and poor at extrapolating beyond the target range in regression because leaves predict observed-region averages. Impurity-based importance can be biased toward high-cardinality variables; permutation importance on suitable held-out data is often a better diagnostic, but correlated features still complicate interpretation. A high OOB score does not remove the need to examine subgroup and deployment performance.
Exercise
A forest of 101 trees gives 55 votes for class 1 and 46 for class 0. What is the hard-vote prediction? If all trees are identical, would the forest gain the usual variance-reduction benefit?
Revealed answer
It predicts class 1. If all trees are identical, their errors are perfectly correlated, so voting adds little variance reduction; diversity from resampling and feature subsampling is essential.
Exam lens
Draw the two randomization levels: bootstrap rows and random feature subset per split. Then state aggregation, OOB estimate, variance reduction, and the trade-off between tree strength and correlation.
Rapid revision checklist
- Can I distinguish random forest from ordinary bagging?
- Can I explain
max_featuresat a split? - Can I calculate a majority/mean prediction?
- Can I define OOB prediction?
- Can I name one interpretability limitation?
Key takeaways
- Random forest = bagged trees plus random feature subsets at splits.
- Feature randomness lowers correlation, making averaging more effective.
- OOB predictions give an internal estimate from unused bootstrap cases.
- Forests reduce variance but trade away some interpretability and may not extrapolate.
Sources
Footnotes
-
scikit-learn, Random Forests User Guide and
RandomForestClassifierAPI. ↩ ↩2 -
Breiman, “Random Forests,” Machine Learning 45 (2001), paper. ↩