LLM engineering · Python

Prompt A/B Platform

Prompt changes are product changes, so they should be tested like them: several prompts share live traffic, traffic shifts toward the better ones while the test runs, and a prompt is dropped only when an always-valid test shows it is worse. Below you can run an experiment on 30 real ones: ten open models, three tasks, and the prompt formats HELM Classic recorded for each.

How much the prompt format matters

The same model on the same questions, with only the prompt's wording and layout changed. Highlighted rows: the best format beats HELM's default by two points or more.

Run an experiment

Up to 20,000 requests. Every 50, the platform updates each prompt's posterior and checks every pair; a prompt is dropped once it is worse than another by more than a point with always-valid evidence. With one prompt left the experiment is decided; at the horizon the best survivor serves. Each request's score is drawn from that prompt's recorded scores.

Share of traffic per prompt (last 500 requests)

Across 30 experiments × 20 replays

Shortfall: score lost against the best prompt per 1,000 requests, the chosen prompt serving once the test ends.

Identical prompts (A/A)

Every variant replaced by the default's answers, no tolerance: any drop is a false alarm. Target: under 5%.