← All projects
Machine learning & AI · Python
Δ

gboost

Gradient boosting written from scratch — watch the decision boundary evolve over boosting rounds, right in your browser.

01 Boost a boundary

Two interleaved half-moons in 2D. Each round fits one shallow tree to the gradients of the log loss, takes a fraction of the step it suggests (the learning rate), and repeats. Newton leaf values, histogram-binned splits and L2 regularisation — the same arithmetic as the Python package.

0 / 60
class 0 class 1 background shade = predicted probability
training log loss per round (lower is better)

04 Overfitting: train vs validation loss

The data is split 70/30 before training. Trees fit only the training rows; the validation log loss is computed each round on held-out rows the booster never sees. Deep trees plus a high learning rate push the curves apart — that gap is overfitting.

train log loss validation log loss

05 Feature importance (gain) — interactive

Total split gain accumulated from the trees actually grown above, normalised to sum to 1. Hover a bar for the exact share. Remember section 03: on noisy data gain can crown a noise column — this chart measures training gain only.

Hover a row to inspect it.

06 Residual plot

Each point is a training row: horizontal axis is the model's predicted probability at the current round, vertical axis is the residual (label − prediction). A well-behaved fit scatters residuals around zero across the whole probability range; structure in the cloud means the model is missing something.

residual ≈ 0 band positive residual (under-predicted) negative residual (over-predicted)

02 The loop

Start from a constant. Ask the loss which way each prediction should move and how sharply. Fit one shallow tree to those gradients. Take a fraction of the step it suggests. Repeat.

gain = G_left²/(H_left+λ) + G_right²/(H_right+λ) − G²/(H+λ)
leaf = −ΣG / (ΣH + λ)

G and H are the summed gradients and hessians in a node. The leaf value is the minimiser of the second-order expansion — a Newton step — rather than the mean of the residuals. The learning rate is the whole trick: taking a tenth of each step and growing ten times as many trees reaches a better answer, because no single tree ever gets to be confident.

03 Gain importance is not importance

A column of pure noise with many distinct values gets more chances to produce a lucky split, and gain only counts how far the training loss fell when it did. On the package's churn data, gain ranks a noise column third of six — above contract. Permutation importance, measured on held-out data, gives it 0.000.

featuregainpermutation
tenure_months0.3590.540
monthly_charge0.1970.066
reference_id (pure noise)0.1760.000
contract0.1690.307
38 tests · Python 3.10–3.12 · NumPy the only dependency · Built by Umer Hashmi