How much the prompt format matters
The same model on the same questions, with only the prompt's wording and layout changed. Highlighted rows: the best format beats HELM's default by two points or more.
Run an experiment
Up to 20,000 requests. Every 50, the platform updates each prompt's posterior and checks every pair; a prompt is dropped once it is worse than another by more than a point with always-valid evidence. With one prompt left the experiment is decided; at the horizon the best survivor serves. Each request's score is drawn from that prompt's recorded scores.
Share of traffic per prompt (last 500 requests)
Across 30 experiments × 20 replays
Shortfall: score lost against the best prompt per 1,000 requests, the chosen prompt serving once the test ends.
Identical prompts (A/A)
Every variant replaced by the default's answers, no tolerance: any drop is a false alarm. Target: under 5%.