Where each version's failures come from
Every wrong answer is traced step by step and blamed on the first step that broke. “Every check passed” means the pipeline worked and the model was simply wrong. Versions run with a "\n\n" stop sequence are marked.
The regression, traced
Broken answers, in full: the text stops where the stop sequence cut it
Compare any two versions
The blame counts for two versions side by side. They are totals, so they show which steps got better or worse overall, not which questions changed; the traced regression above follows the questions themselves.
Does it blame the right step?
Correct GPT-4o traces, each with one step broken on purpose. A trace that still ends right has nothing to blame; of those that failed, how many were blamed on the step that was broken, for the right reason.