← All projects
Data & analytics · Python
SI

Sales Insights

Messy orders export → cleaned analysis → RFM segments & cohort retention, computed live on a seeded sample.

Sample data

A deterministic generator mirrors sales sample: ~900 customers, 10 SKUs, 19 months of orders, guest rows, returns as negative quantities, and weighted repeat buyers.

Rows in
—
Net revenue
—
Identifiable customers
—
Returns kept
—
Guest rows
—

Views

Segment Customers Revenue Share Avg orders Avg recency (d)

Click a segment to highlight it in the chart. Quartile scores: R 4 = most recent; F/M 4 = highest. Labels match rfm_segments().

Top 15 customers in the selected (or overall) view
CustomerOrdersRevenueRecencySegment

The pipeline

orders.csv ─▶ clean (report) ─▶ analysis ─▶ CLI tables ─▶ HTML report

Cleaning reports duplicates removed, blank prices dropped, guests kept, returns kept negative, country spellings normalised.

RFM rules

  • R ≥ 3 & F ≥ 3 & M ≥ 3 → champions
  • R ≥ 3 & F ≥ 3 → loyal
  • R ≥ 3 & F ≤ 2 → new or occasional
  • R ≤ 2 & F ≥ 3 & M ≥ 3 → at risk (valuable)
  • R = 1 & F ≤ 2 → lost
  • else → needs attention

Quartiles with duplicates=drop for ties; F/M ranked first so small shops with ties still split.

15 tests · Python 3.10+ · pandas 2.0+ · Built by Umer Hashmi