Claude API Kubernetes Deployment Example
Claude API Kubernetes Deployment Example
If you're running a service that calls the Claude API and you want it orchestrated by Kubernetes, the setup is mostly standard: a containerized Node.js or Python app, a Secret holding your API key, a Deployment with sensible resource limits, and a Service to expose it. The part that trips people up isn't Kubernetes itself — it's handling API key storage, timeouts for streaming responses, and autoscaling around a backend that has its own rate limits.
This article walks through a complete, working example: Dockerfile, Kubernetes manifests, secret management, health checks, and horizontal pod autoscaling for a service that proxies requests to the Claude API. The same pattern applies whether you're calling Claude directly or going through a gateway like SubToAPI that normalizes the API into a plain HTTPS endpoint.
The application
Assume a minimal Express app that wraps Claude calls behind your own /chat endpoint:
import express from "express";
const app = express();
app.use(express.json());
app.get("/healthz", (req, res) => res.sendStatus(200));
app.post("/chat", async (req, res) => {
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.CLAUDE_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: req.body.messages,
}),
});
const data = await response.json();
res.json(data);
});
app.listen(3000);
Keep the /healthz route — Kubernetes liveness and readiness probes need something that doesn't depend on the Claude API being reachable, otherwise a transient upstream outage takes down your pods.
Dockerfile
FROM node:20-alpine
WORKDIR /app
COPY package*.json ./
RUN npm ci --production
COPY . .
EXPOSE 3000
CMD ["node", "server.js"]
Build and push to your registry:
docker build -t registry.example.com/claude-proxy:1.0 .
docker push registry.example.com/claude-proxy:1.0
Storing the API key as a Secret
Never bake credentials into the image or commit them to manifests. Create a Kubernetes Secret:
kubectl create secret generic claude-api-key \
--from-literal=CLAUDE_API_KEY=sk-ant-xxxxxxxx \
-n default
If you use SubToAPI instead of calling Anthropic directly, the same pattern works — just swap the secret value for a sub_live_... key and point the app at https://api.subtoapi.app/v1/messages. This is useful when multiple services or teams need isolated, revocable keys without sharing one Claude account; see the quickstart for the endpoint shape.
Deployment manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: claude-proxy
labels:
app: claude-proxy
spec:
replicas: 2
selector:
matchLabels:
app: claude-proxy
template:
metadata:
labels:
app: claude-proxy
spec:
containers:
- name: claude-proxy
image: registry.example.com/claude-proxy:1.0
ports:
- containerPort: 3000
envFrom:
- secretRef:
name: claude-api-key
resources:
requests:
cpu: "100m"
memory: "128Mi"
limits:
cpu: "500m"
memory: "256Mi"
readinessProbe:
httpGet:
path: /healthz
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet:
path: /healthz
port: 3000
initialDelaySeconds: 10
periodSeconds: 15
Keep resource requests modest — the proxy itself is lightweight, it's mostly waiting on network I/O to the Claude API. Don't oversize CPU limits expecting heavy compute; you're bottlenecked by upstream latency, not local processing.
Service and Ingress
apiVersion: v1
kind: Service
metadata:
name: claude-proxy
spec:
selector:
app: claude-proxy
ports:
- port: 80
targetPort: 3000
type: ClusterIP
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: claude-proxy
annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "120"
nginx.ingress.kubernetes.io/proxy-send-timeout: "120"
spec:
rules:
- host: claude-proxy.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: claude-proxy
port:
number: 80
The timeout annotations matter if you support streaming or long completions — the default 60-second Nginx ingress timeout will cut off slow responses mid-stream. For streamed output specifically, make sure your ingress also disables buffering (nginx.ingress.kubernetes.io/proxy-buffering: "off") so chunks reach the client as they arrive rather than batched at the end.
Horizontal Pod Autoscaling
Since the workload is I/O-bound, CPU-based autoscaling is a reasonable default, but if your traffic is bursty (e.g., chatbot traffic spikes), scale on request concurrency instead using something like KEDA with a custom metric. A simple CPU-based HPA:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: claude-proxy-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: claude-proxy
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
One thing to watch: scaling pods doesn't scale your Claude rate limits. If you're hitting a ceiling on requests per minute regardless of replica count, the fix is on the API side, not Kubernetes — either request a limit increase, shard traffic across multiple keys, or route through a service with its own key management. This is one of the reasons teams put something like SubToAPI in front of Claude instead of managing raw API keys per service: each application or team gets its own sub_live_... key with independent usage tracking, so a traffic spike in one service doesn't exhaust a shared quota for everyone else.
Rolling updates and config changes
Updating the image is standard:
kubectl set image deployment/claude-proxy claude-proxy=registry.example.com/claude-proxy:1.1
Rotating the API key requires a pod restart since env vars aren't hot-reloaded:
kubectl create secret generic claude-api-key \
--from-literal=CLAUDE_API_KEY=sk-ant-new-key \
--dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deployment/claude-proxy
Observability
Add structured logging around your Claude calls — latency, token usage if returned, and error codes — and ship logs to your existing stack (Loki, ELK, Datadog). If you're using SubToAPI, usage metadata is already surfaced per key in the dashboard, which removes the need to build your own token-tracking layer on top of the proxy; see docs/messages for the response fields.
FAQ
Do I need a dedicated proxy service, or can client apps call Claude directly from pods? Either works, but a dedicated proxy centralizes the API key, lets you add retry/backoff logic once, and makes rate-limit handling consistent across all consumers instead of duplicating it in every service.
How do I handle Claude's streaming responses behind an Ingress? Disable response buffering on your ingress controller and increase read/send timeouts well beyond the default 60 seconds, otherwise streamed chunks get buffered or cut off before the client receives them.
Can I run this setup without managing raw Anthropic keys in Kubernetes Secrets? Yes — point your app at a gateway like SubToAPI and store a sub_live_... key instead; it gives you the same HTTPS interface with streaming and tool use support, plus per-key usage limits you can set from your signup dashboard.