What one upgrade did, task by task
Only questions both versions answered are compared. Broke: the old version was right and the new one wrong; fixed: the reverse. A McNemar test on those two counts, Holm-corrected across the five tasks, decides whether a change is real; a task must also move by at least two points.
Change in accuracy, points
All six upgrades
The overall number is what a single benchmark score would have shown. The gate fails a release when any task regressed.