§ 3.7Module 3

Random Forest

On this page

3.7 — Random Forest

Recall first. Bagging trees reduces variance, but what makes every tree less correlated than ordinary bagging?

Bagged trees plus random feature selection

A random forest grows many decision trees, typically each on a bootstrap sample, and at each split considers only a random subset of features. The final prediction is majority vote/class-probability average or regression mean. Row resampling and feature subsampling both increase diversity; aggregation reduces variance.12

Ordinary bagging lets every split examine all features. If one very strong feature dominates, all trees may choose similar early splits and remain correlated. Randomly limiting candidate features makes alternative features useful in different trees. The forest sacrifices some individual-tree strength to improve the ensemble’s average.

Out-of-bag estimation and controls

For each tree, observations not selected in its bootstrap sample are OOB. Aggregate predictions for each training observation only from trees for which it was OOB, producing an OOB score. This is useful for an internal estimate, but group/time dependence and tuning still require care.

Controls include number of trees, number of features considered per split (max_features), maximum depth, minimum leaf size, class weights, and bootstrap/subsample choices. More trees generally stabilize the estimate until computational returns diminish; depth and leaf size control individual-tree variance. Feature scaling is usually unnecessary for tree thresholds.1

Worked trace

Suppose three trees predict class labels [defect, normal, defect]; the forest predicts defect by majority vote. For a regression target, predictions [12,15,18] produce 15. If a forest has five features but max_features=2, each split randomly considers two of the five; another split can draw a different pair. This is feature randomness, not random deletion of features from the whole dataset.

Limitations

Random forests are strong general-purpose baselines, but they can be large, less transparent than one tree, and poor at extrapolating beyond the target range in regression because leaves predict observed-region averages. Impurity-based importance can be biased toward high-cardinality variables; permutation importance on suitable held-out data is often a better diagnostic, but correlated features still complicate interpretation. A high OOB score does not remove the need to examine subgroup and deployment performance.

Exercise

A forest of 101 trees gives 55 votes for class 1 and 46 for class 0. What is the hard-vote prediction? If all trees are identical, would the forest gain the usual variance-reduction benefit?

Revealed answer

It predicts class 1. If all trees are identical, their errors are perfectly correlated, so voting adds little variance reduction; diversity from resampling and feature subsampling is essential.

Exam lens

Draw the two randomization levels: bootstrap rows and random feature subset per split. Then state aggregation, OOB estimate, variance reduction, and the trade-off between tree strength and correlation.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Random Forests User Guide and RandomForestClassifier API. ↩ ↩2

  2. Breiman, “Random Forests,” Machine Learning 45 (2001), paper. ↩