Stumping
On this page
3.4 — Stumping
Recall first. What is the smallest decision tree that can still split data into two groups? Why might a boosting algorithm prefer many such weak trees?
A decision stump
A decision stump is a decision tree of depth one: one feature test at the root and two leaves. It may be written:
h(x) = +1 if xⱼ < t, otherwise −1
(or the opposite orientation). It is a weak learner because it has low capacity and usually performs only modestly better than random on a difficult problem. Its simplicity makes it fast, interpretable, and less likely to fit complex noise in one step.1
Stumps inside AdaBoost
AdaBoost fits a stump using current example weights. The stump chooses the feature and threshold with the smallest weighted classification error:
err = Σᵢ wᵢ · 1[h(xᵢ) ≠ yᵢ]
α = 1/2 log((1−err)/err)
Then update example weights (one common form):
wᵢ ← wᵢ exp(−α yᵢ h(xᵢ))
Normalize the weights to sum to 1. Misclassified examples (yᵢhᵢ=-1) are multiplied by e^α; correctly classified examples by e^(−α). The next stump sees a changed distribution and tends to address the previous errors. The final classifier is sign(Σ αₘhₘ(x)).2
Worked trace
Four examples start with weights 0.25. A stump misclassifies one, so err=0.25:
α = 0.5 log(0.75/0.25) = 0.5 log 3 ≈ 0.549
Unnormalized weights: the error gets 0.25e^0.549≈0.433; each correct example gets 0.25e^(−0.549)≈0.144. Their sum is about 0.866; after normalization the misclassified example has weight 0.5 and each correct example 1/6. The next stump is strongly encouraged to classify that difficult point correctly. This is an illustrative trace; actual libraries may use equivalent updates and numerical handling.
Why not always use deep trees?
A deep base tree can solve training errors quickly, leaving boosting with high variance and less useful correction. Stumps create a controlled sequence of simple, different corrections. However, stumps may need many rounds and can struggle with interactions unless later stumps assemble them. Validation determines whether depth one is adequate.
Exercise
If a stump has weighted error 0.4, compute its AdaBoost weight approximately. If it misclassifies an example, does that example’s next-round weight rise or fall?
Revealed answer
α=0.5 log(0.6/0.4)=0.5 log(1.5)≈0.203. A misclassified example’s multiplier is e^α>1, so its relative weight rises.
Exam lens
Define “depth-one tree,” write weighted error and α, then explain the weight update and why repeated stumps can form a strong classifier. State that a stump is a weak learner, not a complete high-capacity model.
Rapid revision checklist
- Can I define a stump and draw its structure?
- Can I compute weighted error and learner weight?
- Can I show why a misclassified point receives more weight?
- Can I explain why many stumps can represent more structure than one stump?
Key takeaways
- A stump is one split with two leaves.
- AdaBoost uses weighted stump error to set each stump’s vote.
- Misclassified examples receive greater relative attention in the next round.
- Many shallow corrections can form a strong, but still regularizable, ensemble.
Sources
Footnotes
-
scikit-learn, AdaBoost User Guide and
DecisionTreeClassifierAPI. ↩ -
Hastie, Tibshirani & Friedman, ESL, chapter 10; Freund & Schapire, AdaBoost paper. ↩