§ 3.4Module 3

Stumping

On this page

3.4 — Stumping

Recall first. What is the smallest decision tree that can still split data into two groups? Why might a boosting algorithm prefer many such weak trees?

A decision stump

A decision stump is a decision tree of depth one: one feature test at the root and two leaves. It may be written:

h(x) = +1 if xⱼ < t, otherwise −1

(or the opposite orientation). It is a weak learner because it has low capacity and usually performs only modestly better than random on a difficult problem. Its simplicity makes it fast, interpretable, and less likely to fit complex noise in one step.1

Stumps inside AdaBoost

AdaBoost fits a stump using current example weights. The stump chooses the feature and threshold with the smallest weighted classification error:

err = Σᵢ wᵢ · 1[h(xᵢ) ≠ yᵢ]
α = 1/2 log((1−err)/err)

Then update example weights (one common form):

wᵢ ← wᵢ exp(−α yᵢ h(xᵢ))

Normalize the weights to sum to 1. Misclassified examples (yᵢhᵢ=-1) are multiplied by e^α; correctly classified examples by e^(−α). The next stump sees a changed distribution and tends to address the previous errors. The final classifier is sign(Σ αₘhₘ(x)).2

Worked trace

Four examples start with weights 0.25. A stump misclassifies one, so err=0.25:

α = 0.5 log(0.75/0.25) = 0.5 log 3 ≈ 0.549

Unnormalized weights: the error gets 0.25e^0.549≈0.433; each correct example gets 0.25e^(−0.549)≈0.144. Their sum is about 0.866; after normalization the misclassified example has weight 0.5 and each correct example 1/6. The next stump is strongly encouraged to classify that difficult point correctly. This is an illustrative trace; actual libraries may use equivalent updates and numerical handling.

Why not always use deep trees?

A deep base tree can solve training errors quickly, leaving boosting with high variance and less useful correction. Stumps create a controlled sequence of simple, different corrections. However, stumps may need many rounds and can struggle with interactions unless later stumps assemble them. Validation determines whether depth one is adequate.

Exercise

If a stump has weighted error 0.4, compute its AdaBoost weight approximately. If it misclassifies an example, does that example’s next-round weight rise or fall?

Revealed answer

α=0.5 log(0.6/0.4)=0.5 log(1.5)≈0.203. A misclassified example’s multiplier is e^α>1, so its relative weight rises.

Exam lens

Define “depth-one tree,” write weighted error and α, then explain the weight update and why repeated stumps can form a strong classifier. State that a stump is a weak learner, not a complete high-capacity model.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, AdaBoost User Guide and DecisionTreeClassifier API. ↩

  2. Hastie, Tibshirani & Friedman, ESL, chapter 10; Freund & Schapire, AdaBoost paper. ↩