Claude API Kubernetes Scaling Setup Guide
If you're running Claude API calls from workloads inside Kubernetes, the scaling problem isn't about Claude itself — Anthropic's API scales fine on its own — it's about how your pods, connections, and concurrency limits behave as you scale horizontally. This article walks through a practical setup: HPA configuration for LLM-calling services, connection pooling, handling 429s under concurrent load, and where a gateway layer fits in so you don't rebuild rate-limit logic in every pod.
The short answer: treat your Claude API-calling service like any I/O-bound service (scale on concurrency/latency, not CPU), centralize your API key and rate-limit handling behind a single internal endpoint, and add retry/backoff at the pod level so autoscaling doesn't just multiply your 429 errors.
Why CPU-based autoscaling fails for LLM workloads
Most Kubernetes HPA setups default to CPU or memory thresholds. For a service that mostly waits on network I/O to Claude's API, this is the wrong signal. A pod sitting at 5% CPU utilization can still be holding 200 concurrent streaming connections and be seconds away from falling over.
Instead, scale on:
- Concurrent request count (custom metric via Prometheus adapter)
- Request latency p95 (if latency climbs, you need more replicas or you're hitting upstream rate limits — more replicas won't fix the latter)
- Queue depth if you're using a work queue in front of your Claude calls
A minimal custom-metrics HPA looks like this:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: claude-worker-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: claude-worker
minReplicas: 2
maxReplicas: 20
metrics:
- type: Pods
pods:
metric:
name: claude_active_requests
target:
type: AverageValue
averageValue: "15"
This scales based on active in-flight requests per pod, which actually correlates with how close you are to saturating outbound connections to Claude's API.
The real bottleneck: rate limits, not compute
Here's the thing most teams miss: scaling pods horizontally doesn't scale your Claude API rate limit. Anthropic's limits are per-account (requests per minute, tokens per minute), not per-pod. If you spin up 20 replicas each making independent calls with the same API key, you're not increasing throughput — you're increasing the rate of 429 responses.
Before building a complex Kubernetes autoscaling setup, answer this question: what is actually your bottleneck — pod capacity or API rate limit? If it's the rate limit, more pods make things worse, not better, because you get more concurrent retry storms hitting the same ceiling.
Two ways to handle this:
- Centralized rate limiting — a single internal service (or sidecar) that all pods call through, which enforces your actual Claude API limits and queues/backpressures requests before they hit Anthropic.
- A managed API layer that handles rate-limit-aware routing for you, so your pods just make normal HTTPS calls without needing to coordinate.
This is one of the reasons teams put something like SubToAPI in front of Claude: it exposes a stable HTTPS endpoint with its own application API keys (sub_live_...), so your Kubernetes pods authenticate against one system, and you're not managing per-pod retry coordination against Anthropic's raw limits yourself. See the quickstart for the request shape.
Connection pooling and keep-alive
Each pod making HTTPS calls to an external API should reuse connections. Without keep-alive tuning, you pay TLS handshake overhead on every request, which under autoscaling turns into a meaningful chunk of your p95 latency.
In Node.js:
import https from "https";
const agent = new https.Agent({
keepAlive: true,
maxSockets: 50,
maxFreeSockets: 10,
});
async function callClaude(payload) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
agent,
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(payload),
});
return res.json();
}
Set maxSockets based on your pod's expected concurrency, not an arbitrary number — if each pod handles 20 concurrent requests, 50 sockets gives headroom without over-provisioning.
Pod-level resource requests for I/O-bound workers
Even though these are I/O-bound, don't set requests/limits to near-zero. Streaming responses (see streaming docs) hold open connections and buffer tokens, which does consume memory. A reasonable starting point for a worker handling streaming Claude calls:
resources:
requests:
cpu: "100m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
Tune this after watching actual memory usage under load — streaming many long responses concurrently will push memory higher than a batch of short completions.
Graceful handling of retries and backoff
Autoscaling without proper backoff logic just means more pods retrying failed requests simultaneously. Implement exponential backoff with jitter at the application level inside each pod:
async function callWithRetry(fn, attempts = 4) {
for (let i = 0; i < attempts; i++) {
try {
return await fn();
} catch (err) {
if (err.status !== 429 || i === attempts - 1) throw err;
const delay = 2 ** i * 200 + Math.random() * 200;
await new Promise((r) => setTimeout(r, delay));
}
}
}
This matters more as replica count grows — without jitter, synchronized retries across pods create thundering-herd spikes right after a rate-limit window resets.
Readiness and liveness probes for streaming workers
If your pods maintain long-lived streaming connections, a standard HTTP liveness probe can misreport health. Use a lightweight separate health endpoint that doesn't depend on an active Claude connection, and set terminationGracePeriodSeconds high enough (30–60s) that in-flight streams can finish before a pod is killed during scale-down.
FAQ
Should I scale Kubernetes pods based on Claude API latency?
Yes, partially — latency is a useful signal, but check whether rising latency is caused by pod saturation or by you approaching Anthropic's rate limits. If it's the latter, adding pods won't help; you need centralized rate-limit management instead.
Does Kubernetes autoscaling increase my Claude API rate limit?
No. Rate limits are enforced per account/API key by Anthropic, not per pod. Horizontal scaling increases your request concurrency but not your allowed throughput — you need to manage that centrally, for example through a gateway layer or by requesting a higher limit from Anthropic directly.
What's the simplest way to avoid rate-limit chaos across many pods?
Route all pods through a single internal service or external API layer that enforces consistent rate limiting and retry behavior, rather than letting each pod manage its own retries independently. This is also simpler to monitor — see /docs/messages for the request/response shape if you're routing through SubToAPI.