LLM engineering · Python

AI Feature Flags

Feature flags for AI changes: a new model version goes to 1%, 5%, 25% and 50% of users, and rolls back on its own the moment it is measurably worse, overall or on any one task. The test stays valid however often it is checked. Below you can run a rollout of six real model upgrades, using how each version actually answered HELM Lite's questions.

Roll out an upgrade

Each request draws a question the way the bench does (only which version got it right matters, so the page carries those counts per task) and a user whose bucket decides the arm. Every 50 requests the canary checks each guardrail's evidence that the candidate is worse than control by more than the tolerance; crossing the line rolls it back. Stages last 20,000 requests.

100 replays of each upgrade

The bench's numbers (tolerance 2 points). Extra wrong answers: those the candidate gave beyond what control would have, median over replays; a direct switch puts every request on the candidate.

An identical candidate should almost never roll back

Each of 12 versions rolled out against itself, 100 times, with no tolerance: every rollback is a false alarm. Re-running a fixed-sample test at every check (the usual dashboard) rolls back about half of them.