← Blog

API Gateway Rate Limiting for Multiple LLMs

2026-10-01 · 6 min read · SubToAPI Team

If you're routing requests to more than one LLM provider through a single gateway, rate limiting stops being a single number you can hardcode. Each provider has its own limits (requests per minute, tokens per minute, concurrent requests), your own application has its own limits you want to enforce on top, and your users or teams probably need their own quotas as well. The question "how do I rate limit an API gateway across multiple LLMs" really has three sub-problems: respecting upstream limits, enforcing your own policy limits, and doing both without silently dropping requests or blowing your budget.

This article covers the practical architecture for that: where to put the limiter, which algorithm to use, how to handle multiple providers with different limit shapes, and how to avoid the two failure modes everyone hits — getting 429'd by the provider, or letting one noisy client exhaust a shared quota.

Why rate limiting gets harder with multiple LLMs

A single-provider setup is simple: you know the provider's requests-per-minute and tokens-per-minute ceiling, you track usage against it, and you throttle before you hit it. Add a second or third provider and you get:

The fix is layering: one rate limiter that tracks each upstream provider's real limits, and a second layer that allocates slices of your overall capacity to callers, teams, or API keys.

Pick the right algorithm per layer

Token bucket works well for the upstream-facing layer because LLM limits are usually expressed as a rate (tokens/minute, requests/minute) with some burst tolerance. A bucket that refills continuously and allows short bursts matches how providers actually enforce limits.

Sliding window counters work better for the user-facing layer where you want predictable, auditable quotas ("100 requests per hour per API key") without the burst allowance of a token bucket.

Fixed concurrency limits (a simple semaphore) are necessary for providers that cap concurrent requests separately from rate — streaming responses in particular tie up a connection for the duration of generation, not just the duration of the initial call.

A minimal per-provider token bucket looks like this:

class TokenBucket {
  constructor(capacity, refillPerSecond) {
    this.capacity = capacity;
    this.tokens = capacity;
    this.refillPerSecond = refillPerSecond;
    this.lastRefill = Date.now();
  }

  tryConsume(cost = 1) {
    const now = Date.now();
    const elapsed = (now - this.lastRefill) / 1000;
    this.tokens = Math.min(this.capacity, this.tokens + elapsed * this.refillPerSecond);
    this.lastRefill = now;

    if (this.tokens >= cost) {
      this.tokens -= cost;
      return true;
    }
    return false;
  }
}

const providerBuckets = {
  anthropic: new TokenBucket(4000, 66), // tokens/min expressed as tokens/sec
  openai: new TokenBucket(3000, 50),
};

Run this check before dispatching a request. If it fails, queue the request with a short delay and retry, or return a 429 to the caller with a Retry-After header — don't just forward the request and let the upstream reject it, because that wastes a round trip and can trigger stricter provider-side throttling on repeated violations.

Layering per-key and per-team limits on top

Once upstream limits are respected, add a second check keyed by whoever is calling your gateway — an application, a team, or an individual API key. This is where most of the operational value is, because it's what prevents one integration from consuming the whole shared quota.

async function checkLimits(providerName, keyId, estimatedTokens) {
  const providerOk = providerBuckets[providerName].tryConsume(estimatedTokens);
  if (!providerOk) return { allowed: false, reason: "provider_limit" };

  const keyOk = await keyLimiter.consume(keyId, estimatedTokens);
  if (!keyOk) return { allowed: false, reason: "key_limit" };

  return { allowed: true };
}

Two details matter here:

  1. Estimate token cost before the call, not after. You can't rate limit accurately on output tokens you haven't generated yet — estimate from input length and max_tokens, then reconcile against actual usage once the response completes.
  2. Make limits visible to callers. Return remaining quota in response headers (X-RateLimit-Remaining, X-RateLimit-Reset) so client code can back off intelligently instead of retrying blindly.

Handling streaming and retries

Streaming responses complicate rate limiting because a single request can hold a connection open for tens of seconds while tokens trickle in. Count the concurrent streaming connection against your concurrency limit for its full duration, not just at request start, and release it only when the stream closes or errors out.

For retries, use exponential backoff with jitter and cap the number of attempts. If a provider is rate limiting you, hammering it with retries makes recovery slower for everyone sharing that upstream key.

Where SubToAPI fits

Building and maintaining this layering — per-provider buckets, per-key quotas, streaming-aware concurrency, usage reconciliation — is a real engineering investment if you're doing it just to manage access to Claude. SubToAPI turns your existing Claude access into an HTTPS API with application keys, usage metadata, and team seats already built in, so you're not maintaining your own limiter just to hand out scoped access internally. You generate sub_live_... keys per app or team member from one dashboard, and usage tracking is already in place rather than something you bolt on yourself.

If you're routing to Claude specifically and want the key management and quota visibility without building the gateway layer from scratch, check the quickstart or the messages and streaming docs. Plans start at Solo for single developers, with Team and Scale tiers for per-seat access — see pricing.

Practical checklist

Questions

Do I need separate rate limiters for each LLM provider, or can I use one global limit? Separate limiters per provider. Each provider enforces limits differently (tokens/minute vs. requests/minute vs. concurrency), and a single global number will either under-utilize fast providers or get you throttled on stricter ones.

Should rate limiting happen before or after estimating token usage? Before. Check your limiter against an estimated token cost (input tokens plus max_tokens) before dispatching the request, then reconcile against actual usage once the response completes so your counters stay accurate over time.

How do I rate limit streaming requests differently from regular requests? Count a streaming request against your concurrency limit for its entire open-connection duration, not just at the start. A single long-running stream should hold a concurrency slot until it closes, just like a regular request holds one until it returns.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →