LLM engineering · Python

LLM Regression Detector

A model upgrade can raise the average and still break the task your product depends on. This tool compares two model versions on the same questions, task by task, and fails the build when a task got significantly worse. Below: six real upgrades, measured on HELM Lite's recorded answers from both versions.

What one upgrade did, task by task

Only questions both versions answered are compared. Broke: the old version was right and the new one wrong; fixed: the reverse. A McNemar test on those two counts, Holm-corrected across the five tasks, decides whether a change is real; a task must also move by at least two points.

Change in accuracy, points

All six upgrades

The overall number is what a single benchmark score would have shown. The gate fails a release when any task regressed.