Roll out an upgrade
Each request draws a question the way the bench does (only which version got it right matters, so the page carries those counts per task) and a user whose bucket decides the arm. Every 50 requests the canary checks each guardrail's evidence that the candidate is worse than control by more than the tolerance; crossing the line rolls it back. Stages last 20,000 requests.
100 replays of each upgrade
The bench's numbers (tolerance 2 points). Extra wrong answers: those the candidate gave beyond what control would have, median over replays; a direct switch puts every request on the candidate.
An identical candidate should almost never roll back
Each of 12 versions rolled out against itself, 100 times, with no tolerance: every rollback is a false alarm. Re-running a fixed-sample test at every check (the usual dashboard) rolls back about half of them.