← Blog

Claude API Kubernetes Deployment Best Practices

2026-10-01 · 5 min read · SubToAPI Team

Running Claude-powered services on Kubernetes introduces a specific set of problems that don't show up in a local script or a single container: secrets have to be rotated across pods without downtime, autoscaling needs to react to LLM-shaped load (slow, bursty, token-bound) rather than CPU load, and streaming responses have to survive ingress timeouts and load balancer idle limits. Getting these wrong means dropped connections mid-stream, leaked API keys in ConfigMaps, or pods that scale on the wrong signal entirely.

This guide covers the concrete best practices for deploying applications that call the Claude API from a Kubernetes cluster: secret management, resource and autoscaling configuration, network and timeout tuning for streaming, and observability you actually need in production.

Store Claude API keys as Kubernetes Secrets, never ConfigMaps

This sounds obvious but it's the most common misconfiguration in review. API keys should live in a Secret object, mounted as an environment variable or file, and never checked into a Helm values.yaml that gets committed to git.

apiVersion: v1
kind: Secret
metadata:
  name: claude-api-key
  namespace: production
type: Opaque
stringData:
  CLAUDE_API_KEY: sk-ant-xxxxxxxxxxxx

Reference it in your deployment:

envFrom:
  - secretRef:
      name: claude-api-key

For anything beyond a single service, use a secrets operator (External Secrets Operator, Sealed Secrets, or your cloud provider's CSI driver) so the key lives in a proper vault and Kubernetes only holds a reference. If you're running multiple teams or environments against the same underlying Claude access, consider issuing scoped, per-service keys instead of sharing one credential across every deployment — it makes revocation and usage auditing far easier when something goes wrong. SubToAPI generates separate sub_live_... application keys per service on top of your Claude access, so each Kubernetes deployment gets its own key, its own usage metadata, and can be revoked independently without touching the others — see /docs/quickstart.

Size resource requests and limits around token throughput, not CPU

Pods that call the Claude API are mostly I/O-bound while waiting on the response, so they don't need large CPU requests. What they need is enough memory headroom for concurrent request buffering, especially if you're streaming and holding partial responses or building up tool-use state.

resources:
  requests:
    cpu: "100m"
    memory: "256Mi"
  limits:
    cpu: "500m"
    memory: "512Mi"

Set limits conservatively and monitor actual usage under load before tightening them. Under-provisioned memory limits cause OOMKills that look like random pod restarts and are hard to correlate with Claude traffic unless you're logging request IDs.

Autoscale on concurrency, not CPU

The default HorizontalPodAutoscaler metric (CPU utilization) is a poor signal for Claude workloads because the bottleneck is almost always concurrent in-flight requests or queue depth, not processor usage. Use a custom metric or KEDA scaler tied to request concurrency or an external queue length:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: claude-worker-scaler
spec:
  scaleTargetRef:
    name: claude-worker
  minReplicaCount: 2
  maxReplicaCount: 20
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        metricName: claude_inflight_requests
        threshold: "10"
        query: sum(claude_inflight_requests)

Keep minReplicaCount at 2 or more for anything user-facing — a single pod means a deploy or node drain causes visible downtime for every in-flight Claude request.

Tune timeouts for streaming responses end-to-end

If your service streams Claude responses back to clients (SSE or chunked transfer), every hop in the chain needs a timeout long enough to hold the connection open: the ingress controller, the service mesh sidecar if you run one, and the load balancer. NGINX ingress, for example, defaults to a 60-second proxy read timeout, which will cut off a long streaming completion mid-response.

metadata:
  annotations:
    nginx.ingress.kubernetes.io/proxy-read-timeout: "300"
    nginx.ingress.kubernetes.io/proxy-send-timeout: "300"

If you're on a service mesh (Istio, Linkerd), check the mesh's own timeout defaults separately — they override ingress settings and are a frequent source of "streaming works locally, breaks in cluster" bugs. If your application doesn't strictly need token-by-token output, consider buffering and returning a single response instead; it removes an entire class of infrastructure timeout issues at the cost of perceived latency. SubToAPI's streaming endpoint is a drop-in SSE interface over HTTPS, so the same timeout tuning applies whether you're calling Anthropic directly or through SubToAPI — see /docs/streaming.

Use readiness probes that check downstream dependency health, not just process liveness

A pod that's running fine but can't reach Claude because of DNS or network policy issues should fail readiness, not keep receiving traffic. A simple pattern:

readinessProbe:
  httpGet:
    path: /healthz/claude
    port: 8080
  periodSeconds: 15
  failureThreshold: 2

Have /healthz/claude make a lightweight check — a cached "last successful call" timestamp rather than a real API call on every probe — so you're not burning request quota on health checks.

Apply NetworkPolicies and egress rules deliberately

If your cluster enforces default-deny egress (it should), explicitly allow HTTPS egress to api.anthropic.com or your API gateway's domain. This is easy to forget and causes confusing failures where the pod is healthy but every outbound call times out.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-claude-egress
spec:
  podSelector:
    matchLabels:
      app: claude-worker
  policyTypes:
    - Egress
  egress:
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
      ports:
        - protocol: TCP
          port: 443

Centralize key management and usage visibility across services

As the number of services calling Claude grows — a chat backend, a batch summarization job, an internal tool — tracking which pod used how many tokens becomes a real operational need, not a nice-to-have. Rather than wiring custom logging into every deployment, route Claude traffic through a layer that gives you per-key usage metadata and team-level dashboards out of the box. That's a core part of what SubToAPI provides: swap your Anthropic key for a sub_live_... key per service, keep the same request format against https://api.subtoapi.app/v1/messages, and get usage breakdowns without building your own telemetry pipeline (/docs/messages, /pricing).

Questions

Do I need a service mesh to deploy Claude API workloads on Kubernetes? No. A mesh helps with mTLS, retries, and observability at scale, but a straightforward Deployment with proper Secrets, resource limits, and an HPA/KEDA scaler is sufficient for most Claude-calling services.

Why do my streaming Claude responses cut off after exactly 60 seconds in Kubernetes? That's almost always the default ingress proxy read timeout (NGINX defaults to 60s). Raise proxy-read-timeout and proxy-send-timeout annotations, and check any service mesh timeout settings separately.

Should I run one shared API key across all pods or one per service? Prefer one key per service or team. It limits blast radius if a key leaks, makes usage auditing per-deployment possible, and lets you revoke a single service without rotating credentials cluster-wide.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →