LLM API Gateway Failover Configuration Guide
LLM API gateway failover configuration is the practice of setting up automatic switching between upstream providers, regions, or model endpoints when one becomes slow, rate-limited, or unavailable. The goal is to keep your application responding even when a single provider has an outage or degraded latency, without the end user noticing anything besides maybe a slightly slower response.
In practice this means three things working together: detection (how you know an endpoint is unhealthy), routing logic (where the request goes next), and retry policy (how aggressively you re-attempt before giving up). Get any one of these wrong and you either fail over too late, too often, or in a way that duplicates work and burns through quota. Below is a concrete setup you can adapt whether you're running your own gateway in front of multiple providers or configuring failover inside a managed layer.
Why failover matters for LLM traffic specifically
LLM APIs fail differently than typical REST backends. You don't just get 500s — you get:
- 429 rate limit errors that are transient and provider-specific
- Slow streaming starts where the connection opens but tokens trickle or stall
- Partial completions where a stream dies mid-response
- Regional degradation where one data center is fine and another is backed up
A naive failover setup that only checks HTTP status codes misses most of these. You need timeout-based failover (not just error-based) and you need to treat a stalled stream as a failure condition even if the initial response was 200.
Core components of a failover configuration
1. Health checks per upstream
Each upstream (provider, region, or API key pool) needs an independent health signal. A simple rolling-window approach works well:
window: last 60 seconds
threshold: 5 errors OR p95 latency > 8000ms
action: mark upstream "degraded" for 30s cooldown
Avoid single-request health checks — one slow request shouldn't yank an entire upstream out of rotation. Use a sliding error rate instead.
2. Timeout tuning before retry
Set two timeouts, not one:
- Connect/first-byte timeout — how long you wait for the first token or header. 5–10 seconds is reasonable for most chat completions.
- Total request timeout — a hard ceiling for the whole response, scaled to expected output length.
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 8000);
try {
const res = await fetch(upstreamUrl, {
signal: controller.signal,
headers: { Authorization: `Bearer ${apiKey}` },
method: "POST",
body: JSON.stringify(payload),
});
clearTimeout(timeout);
return res;
} catch (err) {
clearTimeout(timeout);
return failoverToNext(payload);
}
3. Retry policy with backoff
Retries should be bounded and jittered so a failing upstream doesn't get hammered by every client retrying at the same instant.
async function withRetry(fn, attempts = 3) {
for (let i = 0; i < attempts; i++) {
try {
return await fn();
} catch (err) {
if (i === attempts - 1) throw err;
const delay = Math.min(2 ** i * 200 + Math.random() * 200, 2000);
await new Promise((r) => setTimeout(r, delay));
}
}
}
Cap retries at 2–3 attempts per upstream before moving to the next one in the failover chain. More than that and you're adding latency without meaningfully improving success rate.
4. Failover order and fallback chain
Define an explicit priority list rather than random selection:
1. primary region / primary key pool
2. secondary region / secondary key pool
3. alternate model tier (if provider supports it)
4. queue + return 503 with Retry-After
Step 4 matters — don't let failover silently degrade into infinite retries. If everything is down, fail fast and tell the client when to retry.
Where a managed gateway helps
Building and maintaining this logic — health windows, cooldowns, jittered retries, stream-stall detection — is real ongoing work, especially once you're running it across a team with shared budgets and multiple API keys. This is one of the reasons teams move to a managed layer like SubToAPI: it turns your Claude access into a stable HTTPS API with sub_live_... application keys, so your application code talks to one consistent endpoint while the underlying request handling, streaming, and usage metadata are managed for you.
A typical setup looks like this against the SubToAPI endpoint:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this ticket."}]
}'
Your application still needs its own retry/timeout wrapper around this call — that part of failover design doesn't go away just because you're using a managed API layer — but you no longer have to build provider-level health checks, key rotation, or regional routing yourself. See the quickstart and the messages and streaming docs for request formats, and tools if your failover paths also need to preserve tool-calling behavior consistently across retries.
Testing your failover configuration
Don't wait for a real outage to find out your config is wrong. Run these checks before shipping:
- Kill switch test: manually block one upstream and confirm traffic shifts within your configured cooldown window.
- Latency injection: add artificial delay to simulate a slow-but-not-dead upstream and verify timeout thresholds trigger correctly.
- Stream-stall test: open a stream, stop sending tokens mid-way, confirm your client detects the stall and retries rather than hanging.
- Load test the fallback path: your secondary upstream needs enough headroom to absorb full traffic, not just overflow — if it can't, failover just moves the outage instead of fixing it.
Log every failover event with the upstream that failed, the reason, and the upstream it routed to. Without this you'll have no way to tell if your thresholds are too aggressive or too lax.
questions
What's the difference between failover and load balancing for LLM APIs? Load balancing distributes healthy traffic across multiple upstreams for throughput; failover specifically reroutes traffic away from an unhealthy upstream. A good gateway does both — balance under normal conditions, failover when one path degrades.
How long should a cooldown period be before retrying a failed upstream? Start with 30 seconds for rate-limit errors and 60–120 seconds for outright failures or timeouts. Too short and you flap back into a still-struggling upstream; too long and you waste capacity once it recovers.
Should failover logic live in application code or at the gateway layer? Both, at different levels: the gateway handles upstream/provider-level routing and health checks, while your application code should still implement its own request-level timeout and retry wrapper in case the gateway connection itself is slow.