← All projects
Machine learning & AI · Python
%

churn

Score a customer's churn risk with a logistic model trained live in your browser — and see exactly which features push the probability up or down.

01 Score a customer

A logistic regression is trained on ~2,000 synthetic customers generated from the same relationships as src/churn/data.py when you press Score (or change a field after the first run). Contributions are coefficient × (feature − mean) on the log-odds scale.

02 Prediction

—
low
Press Score customer

Feature contributions (log-odds)

  • No prediction yet.

03 What drives churn in this data

  1. Contract type — month-to-month customers churn about 2.6× as often as two-year ones.
  2. Tenure — 49.8% in the first six months against 15.2% after four years.
  3. Spend per month of tenure, support calls, late payments and paying by electronic cheque.

The label is drawn from a logistic model plus noise (σ ≈ 0.55), so no model passes ~0.80 ROC-AUC on this data. A model claiming 0.99 is leaking. PR-AUC is the headline number because churners are the minority class.

04 Held-out results (Python package)

modelaccuracyprecisionrecallf1roc_aucpr_auc
baseline (always stay)0.6110.0000.0000.0000.5000.389
logistic regression0.6880.5840.6900.6320.7690.684
random forest0.6850.5850.6550.6180.7530.667
gradient boosting0.7180.6610.5650.6090.7700.681

Baseline accuracy is 61% by predicting "nobody churns" — which is why accuracy alone is a bad way to judge this problem.

13 tests · Python 3.10+ · scikit-learn · Built by Umer Hashmi