LLM engineering · Python

Pipeline Forensics

When an LLM pipeline gets an answer wrong, the model is only one suspect. The prompt may have lost the question, a stop sequence may have cut the answer off, the token limit may have hit, or the extractor may have read the wrong number. This tool records every step as a trace and blames the first one that broke. Below: 12 model versions' recorded GSM8K runs, and a regression it explained.

Where each version's failures come from

Every wrong answer is traced step by step and blamed on the first step that broke. “Every check passed” means the pipeline worked and the model was simply wrong. Versions run with a "\n\n" stop sequence are marked.

The regression, traced

Broken answers, in full: the text stops where the stop sequence cut it

Compare any two versions

The blame counts for two versions side by side. They are totals, so they show which steps got better or worse overall, not which questions changed; the traced regression above follows the questions themselves.

Does it blame the right step?

Correct GPT-4o traces, each with one step broken on purpose. A trace that still ends right has nothing to blame; of those that failed, how many were blamed on the step that was broken, for the right reason.