ML Atlas

03 · Supervised · 5 min read · Interactive · updated

How does gradient boosting work and why is it so effective?

In short

Gradient boosting builds a model step by step: each new shallow tree corrects the errors of the trees so far, moving in the direction that lowers the loss.

What it is

Gradient boosting is an ensemble method in which the model is built as a sum of many simple models — usually shallow decision trees — added one after another. Each new tree learns to predict how much, and in which direction, the current sum needs to be corrected to reduce the loss function. The general formulation was given by Jerome Friedman (2001).

For regression with squared error the idea is remarkably simple. Start with the mean. Compute the residuals: how far each example still is from the truth. Train a small tree to predict those residuals and add its (shrunken) predictions to the model. Compute the new residuals, fit the next tree — and repeat hundreds of times.

Intuition: a golfer does not sink the ball in one stroke. The first shot sends the ball near the target, and each subsequent one corrects what is left. Gradient boosting is today one of the most effective methods for tabular data; its fast implementations are XGBoost, LightGBM and CatBoost.

Mechanism — why it works this way

Why "gradient"? The residual yᵢ − F(xᵢ) is exactly the negative gradient of the squared loss ½(y − F)² with respect to the prediction F. Fitting a tree to the residuals is therefore a step of gradient descent — not in weight space, but in function space. This observation lets you use any differentiable loss: for classification the trees are fitted to the gradient of cross-entropy, for robust regression to that of the Huber or absolute loss. Training for a specific metric becomes a matter of choosing the loss function.

The key parameter is the learning rate (shrinkage, ν): each tree is added with weight ν, typically 0.01–0.1. Smaller steps mean the model needs more trees, but also that no single tree decides the outcome, and the ensemble approaches a good function more gently. Friedman observed that a small ν almost always improves generalisation; the price is time.

Unlike a random forest, boosting primarily reduces bias, and every tree adds something new. That is why adding trees indefinitely leads to overfitting: after a while the model starts fitting noise. The number of trees is chosen by early stopping on validation data, with additional regularisation from shallow trees (depth 3–8), a minimum number of examples per leaf, and subsampling in each round (so-called stochastic gradient boosting).

Tree depth determines the order of interactions the model can capture. Stumps (depth 1) give an additive model — a sum of separate functions of each feature. Depth-3 trees can model interactions of three features at once.

Caveats: many hyperparameters that interact with one another, sequential training (hard to parallelise at the level of trees), and no extrapolation beyond the range of the data. With small datasets and near-linear relationships, a simple linear model can win.

By example

The Diabetes dataset: 442 patients, 10 features, target — disease progression after one year; training on 331, testing on 111 (75/25 split, random_state=0). The model starts from the training mean of 151.9 (RMSE 79.1). The first three depth-3 trees with a learning rate of 0.1 lower the training error to 74.6, 70.8 and 67.5. With a learning rate of 1.0, after 100 trees the model reaches a training R² of 1.00 and a test R² of −0.31 — it has learned the noise. With 0.1 the best test score (0.29) comes after just 13 trees, and by 2,000 trees it drops to 0.09. With 0.01 the optimum (0.30) comes at 156 trees and the decline afterwards is much gentler. Shallower trees (depth 2) with 0.01 give the best result, 0.33, after 194 rounds; early stopping on 20% of the training data halted on its own after 370 trees with a test R² of 0.29.

An honest punchline: in 5-fold cross-validation, plain linear regression achieves a mean R² of 0.49 on this dataset, while gradient boosting gets 0.42 (default settings) and 0.45 (learning rate 0.01, depth 2, 1,000 trees). 442 patients and a near-linear relationship are terrain where boosting's flexibility is more a cost than an asset.

In practice

  • In scikit-learn: GradientBoostingRegressor / GradientBoostingClassifier (small data) and the much faster HistGradientBoostingRegressor / HistGradientBoostingClassifier (from tens of thousands of rows upward).
  • Typical starting point: learning_rate 0.05–0.1, max_depth 3–6, a few hundred to a few thousand trees with early stopping.
  • Early stopping: n_iter_no_change=50, validation_fraction=0.1 in scikit-learn, or early_stopping_rounds in XGBoost and LightGBM.
  • subsample=0.8 (row sampling) and feature sampling reduce variance and speed up training.
  • Fix a small learning_rate first, then choose the number of trees; halving the learning rate means roughly twice as many trees.

Frequently asked questions

How does gradient boosting differ from a random forest?
A forest trains deep trees independently and averages them, which reduces variance. Boosting trains shallow trees sequentially, each correcting the errors of the previous ones, which reduces bias. Boosting usually achieves better results once tuned; a forest is more robust to poor hyperparameters.
Does gradient boosting overfit?
Yes, if you keep adding trees without limit, especially with a high learning rate. The safeguards are early stopping on validation data, a small learning rate, shallow trees and subsampling.
Why does boosting win competitions on tabular data?
Because it handles features on different scales, non-linearities, interactions and missing values without tedious preprocessing, and its implementations are very fast. Benchmark studies, including Grinsztajn et al. (2022), show that on typical tabular data tree-based models still often beat neural networks.

Sources

  • Friedman J. H. "Greedy Function Approximation: A Gradient Boosting Machine", Annals of Statistics 29(5), 2001.
  • Friedman J. H. "Stochastic Gradient Boosting", Computational Statistics & Data Analysis 38(4), 2002.
  • Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., 2009, ch. 10.
  • Grinsztajn L., Oyallon E., Varoquaux G. "Why do tree-based models still outperform deep learning on typical tabular data?", NeurIPS 2022 (Datasets and Benchmarks).
  • scikit-learn documentation, "Gradient-boosted trees": https://scikit-learn.org/stable/modules/ensemble.html#gradient-boosted-trees

See also