§ 3.5Module 3

XGBoost

On this page

3.5 — XGBoost

Recall first. Gradient boosting adds trees sequentially. What two levers can stop each new tree from making an overly large correction?

What XGBoost is

XGBoost (eXtreme Gradient Boosting) is a regularized, efficient implementation of gradient-boosted decision trees. It builds an additive model:

ŷᵢ = Σₜ fₜ(xᵢ),  fₜ ∈ tree space

At stage t, it adds a tree that improves the chosen differentiable objective. Its engineering includes parallelized split finding, sparse/missing-value handling, and practical system optimizations; these are implementation details, not a change to the core idea of sequential gradient boosting.12

Objective and regularization

A simplified training objective is:

Obj = Σᵢ l(yᵢ, ŷᵢ) + Σₜ Ω(fₜ)
Ω(f) = γT + (λ/2)Σⱼwⱼ² + αΣⱼ|wⱼ|

l is the prediction loss, T is number of leaves, and wⱼ are leaf scores. γ penalizes adding leaves, while L2 λ and L1 α shrink leaf weights. The exact objective and defaults depend on the task/API version; the important exam distinction is loss plus explicit tree-complexity regularization.1

Using a second-order Taylor approximation, XGBoost uses gradient gᵢ and Hessian hᵢ of the loss. For a leaf with sums G=Σgᵢ and H=Σhᵢ, when L1 regularization is omitted (α=0), a common optimal leaf-weight expression is:

w* = −G/(H+λ)

When α>0, the L1 term applies soft-thresholding first: w* = −sign(G) max(|G|−α, 0)/(H+λ). A split is worthwhile only when its gain exceeds the complexity penalty γ. This explains why XGBoost is more than “just add a tree”: it chooses regularized updates using first- and second-order information.

Important controls

A validation set or cross-validation is required for choosing these. The XGBoost parameter reference describes the controls and notes that learning rate shrinkage helps prevent overfitting.1

Worked one-leaf update

Suppose a candidate leaf has G=6, H=4, and λ=2. Its regularized optimal score is:

w* = −6/(4+2) = −1

The negative update says the current predictions are too high in the gradient direction for this leaf. If λ were 0, the update would be −1.5; regularization shrinks its magnitude. A full implementation also evaluates the objective gain and whether adding the leaf passes γ.

Exercise

A team sets learning_rate=0.01 but leaves max_depth=20 and trains until the training score is perfect. Name two risks and two controls.

Revealed answer

Deep trees can still overfit despite a small learning rate, and training until perfection ignores validation performance. Limit depth/leaves, use row/feature subsampling or regularization, monitor a validation set with early stopping, and choose the number of rounds by CV. Small learning rate is not a complete anti-overfitting strategy.

Exam lens

Define XGBoost as regularized gradient-boosted trees. State additive sequential fitting, objective loss + complexity penalty, learning-rate shrinkage, and tree/leaf regularization. Separate conceptual algorithm from engineering optimizations and library defaults.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. XGBoost documentation, Parameters, Introduction to Model, and Parameter tuning. ↩ ↩2 ↩3

  2. Chen & Guestrin, “XGBoost: A Scalable Tree Boosting System,” KDD 2016, paper. ↩