LLM engineering · Python

LLM Cost Autopilot

Most prompts don't need the most expensive model. This router predicts, for every model, the chance it answers a prompt correctly and picks the cheapest one that is likely enough to. Everything below is measured on 2,197 held-out questions from HELM Lite, with each provider's list prices applied to HELM's real token counts.

Cost against accuracy

Each dot is a strategy run on the same held-out questions (GSM8K, LegalBench, MATH, MMLU, OpenBookQA). The line is the learned router at every confidence setting. The oracle knows in hindsight which cheap model was right: a ceiling, not a strategy. Tap a dot for its numbers.

Set the router's confidence

The router sends a prompt to the cheapest model whose predicted chance of being right clears this confidence (and to the model it trusts most when none does). Higher is safer and dearer. The marked setting is the one chosen on the training half: the lowest that matched gpt-4o's accuracy there.

Where the router sends traffic

By task

The router at its chosen confidence against sending everything to gpt-4o.