← All projects
Machine learning & AI · Python
%
churn
Score a customer's churn risk with a logistic model trained live in your browser — and see exactly which features push the probability up or down.
01 Score a customer
A logistic regression is trained on ~2,000 synthetic customers generated from the same relationships as src/churn/data.py when you press Score (or change a field after the first run). Contributions are coefficient × (feature − mean) on the log-odds scale.
02 Prediction
—
low
Press Score customer
Feature contributions (log-odds)
- No prediction yet.
03 What drives churn in this data
- Contract type — month-to-month customers churn about 2.6× as often as two-year ones.
- Tenure — 49.8% in the first six months against 15.2% after four years.
- Spend per month of tenure, support calls, late payments and paying by electronic cheque.
The label is drawn from a logistic model plus noise (σ ≈ 0.55), so no model passes ~0.80 ROC-AUC on this data. A model claiming 0.99 is leaking. PR-AUC is the headline number because churners are the minority class.
04 Held-out results (Python package)
| model | accuracy | precision | recall | f1 | roc_auc | pr_auc |
|---|---|---|---|---|---|---|
| baseline (always stay) | 0.611 | 0.000 | 0.000 | 0.000 | 0.500 | 0.389 |
| logistic regression | 0.688 | 0.584 | 0.690 | 0.632 | 0.769 | 0.684 |
| random forest | 0.685 | 0.585 | 0.655 | 0.618 | 0.753 | 0.667 |
| gradient boosting | 0.718 | 0.661 | 0.565 | 0.609 | 0.770 | 0.681 |
Baseline accuracy is 61% by predicting "nobody churns" — which is why accuracy alone is a bad way to judge this problem.