← Blog

LLM API Gateway Failover Configuration Guide

2026-10-01 · 5 min read · SubToAPI Team

LLM API gateway failover configuration is the practice of setting up automatic switching between upstream providers, regions, or model endpoints when one becomes slow, rate-limited, or unavailable. The goal is to keep your application responding even when a single provider has an outage or degraded latency, without the end user noticing anything besides maybe a slightly slower response.

In practice this means three things working together: detection (how you know an endpoint is unhealthy), routing logic (where the request goes next), and retry policy (how aggressively you re-attempt before giving up). Get any one of these wrong and you either fail over too late, too often, or in a way that duplicates work and burns through quota. Below is a concrete setup you can adapt whether you're running your own gateway in front of multiple providers or configuring failover inside a managed layer.

Why failover matters for LLM traffic specifically

LLM APIs fail differently than typical REST backends. You don't just get 500s — you get:

A naive failover setup that only checks HTTP status codes misses most of these. You need timeout-based failover (not just error-based) and you need to treat a stalled stream as a failure condition even if the initial response was 200.

Core components of a failover configuration

1. Health checks per upstream

Each upstream (provider, region, or API key pool) needs an independent health signal. A simple rolling-window approach works well:

window: last 60 seconds
threshold: 5 errors OR p95 latency > 8000ms
action: mark upstream "degraded" for 30s cooldown

Avoid single-request health checks — one slow request shouldn't yank an entire upstream out of rotation. Use a sliding error rate instead.

2. Timeout tuning before retry

Set two timeouts, not one:

const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 8000);

try {
  const res = await fetch(upstreamUrl, {
    signal: controller.signal,
    headers: { Authorization: `Bearer ${apiKey}` },
    method: "POST",
    body: JSON.stringify(payload),
  });
  clearTimeout(timeout);
  return res;
} catch (err) {
  clearTimeout(timeout);
  return failoverToNext(payload);
}

3. Retry policy with backoff

Retries should be bounded and jittered so a failing upstream doesn't get hammered by every client retrying at the same instant.

async function withRetry(fn, attempts = 3) {
  for (let i = 0; i < attempts; i++) {
    try {
      return await fn();
    } catch (err) {
      if (i === attempts - 1) throw err;
      const delay = Math.min(2 ** i * 200 + Math.random() * 200, 2000);
      await new Promise((r) => setTimeout(r, delay));
    }
  }
}

Cap retries at 2–3 attempts per upstream before moving to the next one in the failover chain. More than that and you're adding latency without meaningfully improving success rate.

4. Failover order and fallback chain

Define an explicit priority list rather than random selection:

1. primary region / primary key pool
2. secondary region / secondary key pool
3. alternate model tier (if provider supports it)
4. queue + return 503 with Retry-After

Step 4 matters — don't let failover silently degrade into infinite retries. If everything is down, fail fast and tell the client when to retry.

Where a managed gateway helps

Building and maintaining this logic — health windows, cooldowns, jittered retries, stream-stall detection — is real ongoing work, especially once you're running it across a team with shared budgets and multiple API keys. This is one of the reasons teams move to a managed layer like SubToAPI: it turns your Claude access into a stable HTTPS API with sub_live_... application keys, so your application code talks to one consistent endpoint while the underlying request handling, streaming, and usage metadata are managed for you.

A typical setup looks like this against the SubToAPI endpoint:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Summarize this ticket."}]
  }'

Your application still needs its own retry/timeout wrapper around this call — that part of failover design doesn't go away just because you're using a managed API layer — but you no longer have to build provider-level health checks, key rotation, or regional routing yourself. See the quickstart and the messages and streaming docs for request formats, and tools if your failover paths also need to preserve tool-calling behavior consistently across retries.

Testing your failover configuration

Don't wait for a real outage to find out your config is wrong. Run these checks before shipping:

Log every failover event with the upstream that failed, the reason, and the upstream it routed to. Without this you'll have no way to tell if your thresholds are too aggressive or too lax.

questions

What's the difference between failover and load balancing for LLM APIs? Load balancing distributes healthy traffic across multiple upstreams for throughput; failover specifically reroutes traffic away from an unhealthy upstream. A good gateway does both — balance under normal conditions, failover when one path degrades.

How long should a cooldown period be before retrying a failed upstream? Start with 30 seconds for rate-limit errors and 60–120 seconds for outright failures or timeouts. Too short and you flap back into a still-struggling upstream; too long and you waste capacity once it recovers.

Should failover logic live in application code or at the gateway layer? Both, at different levels: the gateway handles upstream/provider-level routing and health checks, while your application code should still implement its own request-level timeout and retry wrapper in case the gateway connection itself is slow.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →