gboost
Gradient boosting written from scratch — watch the decision boundary evolve over boosting rounds, right in your browser.
01 Boost a boundary
Two interleaved half-moons in 2D. Each round fits one shallow tree to the gradients of the log loss, takes a fraction of the step it suggests (the learning rate), and repeats. Newton leaf values, histogram-binned splits and L2 regularisation — the same arithmetic as the Python package.
04 Overfitting: train vs validation loss
The data is split 70/30 before training. Trees fit only the training rows; the validation log loss is computed each round on held-out rows the booster never sees. Deep trees plus a high learning rate push the curves apart — that gap is overfitting.
05 Feature importance (gain) — interactive
Total split gain accumulated from the trees actually grown above, normalised to sum to 1. Hover a bar for the exact share. Remember section 03: on noisy data gain can crown a noise column — this chart measures training gain only.
Hover a row to inspect it.
06 Residual plot
Each point is a training row: horizontal axis is the model's predicted probability at the current round, vertical axis is the residual (label − prediction). A well-behaved fit scatters residuals around zero across the whole probability range; structure in the cloud means the model is missing something.
02 The loop
Start from a constant. Ask the loss which way each prediction should move and how sharply. Fit one shallow tree to those gradients. Take a fraction of the step it suggests. Repeat.
gain = G_left²/(H_left+λ) + G_right²/(H_right+λ) − G²/(H+λ)
leaf = −ΣG / (ΣH + λ)
G and H are the summed gradients and hessians in a node. The leaf value is the minimiser of the second-order expansion — a Newton step — rather than the mean of the residuals. The learning rate is the whole trick: taking a tenth of each step and growing ten times as many trees reaches a better answer, because no single tree ever gets to be confident.
03 Gain importance is not importance
A column of pure noise with many distinct values gets more chances to produce a lucky split, and gain only counts how far the training loss fell when it did. On the package's churn data, gain ranks a noise column third of six — above contract. Permutation importance, measured on held-out data, gives it 0.000.
| feature | gain | permutation |
|---|---|---|
| tenure_months | 0.359 | 0.540 |
| monthly_charge | 0.197 | 0.066 |
| reference_id (pure noise) | 0.176 | 0.000 |
| contract | 0.169 | 0.307 |