LLM infrastructure · Python

LLM Gateway

One OpenAI-compatible endpoint in front of several model providers. Each provider gets a circuit breaker fed by a rolling health window; open breakers are skipped, half-open ones are probed until they recover on their own. Latency-sensitive requests are hedged, deferrable ones are queued through an outage. Below is the real benchmark — the same 1,440 requests sent three ways while providers fail on a schedule — replayed second by second.

Three ways to survive an outage

The same 1,440 requests (5/s interactive, 1/s batch, two tenants, three features) sent through three setups while providers fail on a schedule. Pick a metric to compare; the gateway row is highlighted throughout.

Answered requests, every 5 seconds

Share of interactive requests answered in each 5-second window. The scripted outages are shaded: alpha returns 503s from 60–120 s, every provider is down from 100–115 s, and alpha turns 6 s slower from 150–210 s. Watch the direct and retry lines fall into each window while the gateway fails over and stays up.

How the breaker reacted

Every state change the gateway made on its own, with no one touching it. A breaker opens after repeated failures or a blown latency budget, waits out a cooldown, then probes with a trickle of traffic before trusting a provider again.

What it cost

Staying up is not free: failover sends some traffic to a pricier provider, and hedged requests pay for a second call that sometimes loses. The gateway spent $1.82 against $1.70 for the direct call — up 7% — and every dollar is attributed to a tenant and a feature.

How it works

Requests come in on POST /v1/chat/completions, the same shape as the OpenAI API. A router picks a provider by request class: interactive work is answered now and hedged once if the first provider is slow; batch work is deferrable, so it is queued through an outage and drained when capacity returns.

Each provider carries a circuit breaker fed by a rolling health window in Redis, shared across every gateway replica. Five failures in a row, or a p95 latency over its budget, trips the breaker open and traffic skips that provider. After a cooldown the breaker goes half-open and a few probe requests test the water; three clean probes close it again. All of this is automatic — the state changes in the table above were made by the gateway, not by a human.

The providers here are simulated, failing on a fixed schedule so the recovery logic can be measured repeatably. The numbers measure the gateway's behaviour, not any model's quality or any vendor's real uptime. Reproduce the run with python -m bench.chaos --fake-redis (about four minutes). The raw output is bench-v2.txt / bench-v2.json.