← Blog

Claude API Kubernetes Deployment Example

2026-10-04 · 5 min read · SubToAPI Team

Claude API Kubernetes Deployment Example

If you're running a service that calls the Claude API and you want it orchestrated by Kubernetes, the setup is mostly standard: a containerized Node.js or Python app, a Secret holding your API key, a Deployment with sensible resource limits, and a Service to expose it. The part that trips people up isn't Kubernetes itself — it's handling API key storage, timeouts for streaming responses, and autoscaling around a backend that has its own rate limits.

This article walks through a complete, working example: Dockerfile, Kubernetes manifests, secret management, health checks, and horizontal pod autoscaling for a service that proxies requests to the Claude API. The same pattern applies whether you're calling Claude directly or going through a gateway like SubToAPI that normalizes the API into a plain HTTPS endpoint.

The application

Assume a minimal Express app that wraps Claude calls behind your own /chat endpoint:

import express from "express";

const app = express();
app.use(express.json());

app.get("/healthz", (req, res) => res.sendStatus(200));

app.post("/chat", async (req, res) => {
  const response = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": process.env.CLAUDE_API_KEY,
      "anthropic-version": "2023-06-01",
      "content-type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      messages: req.body.messages,
    }),
  });
  const data = await response.json();
  res.json(data);
});

app.listen(3000);

Keep the /healthz route — Kubernetes liveness and readiness probes need something that doesn't depend on the Claude API being reachable, otherwise a transient upstream outage takes down your pods.

Dockerfile

FROM node:20-alpine
WORKDIR /app
COPY package*.json ./
RUN npm ci --production
COPY . .
EXPOSE 3000
CMD ["node", "server.js"]

Build and push to your registry:

docker build -t registry.example.com/claude-proxy:1.0 .
docker push registry.example.com/claude-proxy:1.0

Storing the API key as a Secret

Never bake credentials into the image or commit them to manifests. Create a Kubernetes Secret:

kubectl create secret generic claude-api-key \
  --from-literal=CLAUDE_API_KEY=sk-ant-xxxxxxxx \
  -n default

If you use SubToAPI instead of calling Anthropic directly, the same pattern works — just swap the secret value for a sub_live_... key and point the app at https://api.subtoapi.app/v1/messages. This is useful when multiple services or teams need isolated, revocable keys without sharing one Claude account; see the quickstart for the endpoint shape.

Deployment manifest

apiVersion: apps/v1
kind: Deployment
metadata:
  name: claude-proxy
  labels:
    app: claude-proxy
spec:
  replicas: 2
  selector:
    matchLabels:
      app: claude-proxy
  template:
    metadata:
      labels:
        app: claude-proxy
    spec:
      containers:
        - name: claude-proxy
          image: registry.example.com/claude-proxy:1.0
          ports:
            - containerPort: 3000
          envFrom:
            - secretRef:
                name: claude-api-key
          resources:
            requests:
              cpu: "100m"
              memory: "128Mi"
            limits:
              cpu: "500m"
              memory: "256Mi"
          readinessProbe:
            httpGet:
              path: /healthz
              port: 3000
            initialDelaySeconds: 5
            periodSeconds: 10
          livenessProbe:
            httpGet:
              path: /healthz
              port: 3000
            initialDelaySeconds: 10
            periodSeconds: 15

Keep resource requests modest — the proxy itself is lightweight, it's mostly waiting on network I/O to the Claude API. Don't oversize CPU limits expecting heavy compute; you're bottlenecked by upstream latency, not local processing.

Service and Ingress

apiVersion: v1
kind: Service
metadata:
  name: claude-proxy
spec:
  selector:
    app: claude-proxy
  ports:
    - port: 80
      targetPort: 3000
  type: ClusterIP
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: claude-proxy
  annotations:
    nginx.ingress.kubernetes.io/proxy-read-timeout: "120"
    nginx.ingress.kubernetes.io/proxy-send-timeout: "120"
spec:
  rules:
    - host: claude-proxy.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: claude-proxy
                port:
                  number: 80

The timeout annotations matter if you support streaming or long completions — the default 60-second Nginx ingress timeout will cut off slow responses mid-stream. For streamed output specifically, make sure your ingress also disables buffering (nginx.ingress.kubernetes.io/proxy-buffering: "off") so chunks reach the client as they arrive rather than batched at the end.

Horizontal Pod Autoscaling

Since the workload is I/O-bound, CPU-based autoscaling is a reasonable default, but if your traffic is bursty (e.g., chatbot traffic spikes), scale on request concurrency instead using something like KEDA with a custom metric. A simple CPU-based HPA:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: claude-proxy-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: claude-proxy
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 60

One thing to watch: scaling pods doesn't scale your Claude rate limits. If you're hitting a ceiling on requests per minute regardless of replica count, the fix is on the API side, not Kubernetes — either request a limit increase, shard traffic across multiple keys, or route through a service with its own key management. This is one of the reasons teams put something like SubToAPI in front of Claude instead of managing raw API keys per service: each application or team gets its own sub_live_... key with independent usage tracking, so a traffic spike in one service doesn't exhaust a shared quota for everyone else.

Rolling updates and config changes

Updating the image is standard:

kubectl set image deployment/claude-proxy claude-proxy=registry.example.com/claude-proxy:1.1

Rotating the API key requires a pod restart since env vars aren't hot-reloaded:

kubectl create secret generic claude-api-key \
  --from-literal=CLAUDE_API_KEY=sk-ant-new-key \
  --dry-run=client -o yaml | kubectl apply -f -
kubectl rollout restart deployment/claude-proxy

Observability

Add structured logging around your Claude calls — latency, token usage if returned, and error codes — and ship logs to your existing stack (Loki, ELK, Datadog). If you're using SubToAPI, usage metadata is already surfaced per key in the dashboard, which removes the need to build your own token-tracking layer on top of the proxy; see docs/messages for the response fields.

FAQ

Do I need a dedicated proxy service, or can client apps call Claude directly from pods? Either works, but a dedicated proxy centralizes the API key, lets you add retry/backoff logic once, and makes rate-limit handling consistent across all consumers instead of duplicating it in every service.

How do I handle Claude's streaming responses behind an Ingress? Disable response buffering on your ingress controller and increase read/send timeouts well beyond the default 60 seconds, otherwise streamed chunks get buffered or cut off before the client receives them.

Can I run this setup without managing raw Anthropic keys in Kubernetes Secrets? Yes — point your app at a gateway like SubToAPI and store a sub_live_... key instead; it gives you the same HTTPS interface with streaming and tool use support, plus per-key usage limits you can set from your signup dashboard.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →