Claude API Reverse Proxy Configuration Guide
A reverse proxy in front of the Claude API sits between your application (or your users' browsers) and Anthropic's endpoints. It forwards requests, injects your real x-api-key, and lets you add logging, rate limiting, caching, or multi-tenant auth without touching your app code every time you change providers or policies.
The core configuration challenge is that Claude API responses can be streamed via server-sent events (SSE), and most default proxy setups buffer responses instead of forwarding them chunk by chunk. Get that wrong and your chat UI looks frozen until the full response arrives. Below is a working configuration for the two most common setups — nginx and a custom Node.js proxy — plus what to check regardless of which stack you use.
Why put a proxy in front of Claude at all
Common reasons teams do this:
- Never ship the real Claude API key to client-side code. Mobile apps, browser extensions and frontend SPAs should never hold a provider key.
- Centralized rate limiting and quotas per user, team or feature, independent of Anthropic's own limits.
- Request/response logging for debugging, billing, or audit trails.
- Header and body normalization if you're routing requests from multiple internal services with slightly different formats.
- Swapping providers or models behind a stable internal endpoint without redeploying every consumer.
If your only goal is "don't expose the key and get usage tracking per app," building this yourself is optional — a hosted proxy like SubToAPI already does key issuance, streaming and usage metadata out of the box (see /docs/quickstart). But if you need custom routing logic, self-hosting is the right call, and the configuration below applies whether you're proxying Claude directly or another upstream.
Baseline requirements for any Claude API proxy
Regardless of stack, your configuration needs to handle:
- Streaming passthrough — disable buffering for SSE responses.
- Long-lived connections — increase read/proxy timeouts beyond default 60s for long completions.
- Correct header forwarding —
content-type: application/json,anthropic-version, and your injectedx-api-key. - Error passthrough — forward Anthropic's status codes and error bodies instead of masking them with generic 500s.
- CORS if browsers call the proxy directly.
nginx configuration example
location /claude/ {
proxy_pass https://api.anthropic.com/;
proxy_http_version 1.1;
proxy_set_header Host api.anthropic.com;
proxy_set_header x-api-key $anthropic_api_key;
proxy_set_header anthropic-version "2023-06-01";
proxy_set_header Content-Type application/json;
# Required for streaming responses
proxy_buffering off;
proxy_cache off;
chunked_transfer_encoding on;
# Long completions and slow first-byte streams
proxy_read_timeout 300s;
proxy_connect_timeout 10s;
# Strip inbound auth so clients can't override your key
proxy_set_header Authorization "";
}
Set $anthropic_api_key from an environment-backed nginx variable (via njs or an included .conf file with restricted permissions), never hardcode it in a file that ends up in version control.
Node.js reverse proxy example
If you need per-request logic — rate limiting by user ID, injecting a system prompt, logging token usage — a thin Node proxy gives you more control than nginx alone:
import express from "express";
import fetch from "node-fetch";
const app = express();
app.use(express.json());
app.post("/claude/v1/messages", async (req, res) => {
const upstream = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify(req.body),
});
// Passthrough streaming
if (req.body.stream) {
res.setHeader("content-type", "text/event-stream");
res.setHeader("cache-control", "no-cache");
res.setHeader("connection", "keep-alive");
upstream.body.pipe(res);
return;
}
const data = await upstream.json();
res.status(upstream.status).json(data);
});
app.listen(8080);
Key details:
upstream.body.pipe(res)forwards SSE chunks as they arrive instead of waiting for the full response.- Status codes from Anthropic (
400,429,529) are forwarded rather than swallowed, so client-side error handling still works. - Never trust an
Authorizationheader from the inbound request — always inject the real key server-side.
Adding rate limits and logging
Once the base proxy works, wrap the handler with per-key or per-user middleware:
app.post("/claude/v1/messages", rateLimit(userId), async (req, res) => {
const start = Date.now();
// ...proxy logic...
logUsage({ userId, latencyMs: Date.now() - start, model: req.body.model });
});
This is also the point where most teams realize they're rebuilding a small API gateway — key issuance, per-app quotas, streaming, usage dashboards. SubToAPI packages that layer as a hosted service: you get application keys in the format sub_live_..., per-key usage metadata, and streaming support without maintaining the proxy yourself. Example call against it:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"stream": true,
"messages": [{"role": "user", "content": "Summarize this proxy config."}]
}'
Full request/response shape is in /docs/messages and streaming behavior in /docs/streaming. Tool use requests proxy the same way — see /docs/tools.
Checklist before going to production
- Confirm streaming works end-to-end with a real client, not just curl with
-N. - Set read timeouts above your p99 completion time, including tool-use round trips.
- Strip or overwrite any client-supplied auth headers before forwarding.
- Log status codes and latency, not full request/response bodies, unless you have a retention policy.
- Return Anthropic's original error JSON where possible so client-side error handling stays consistent.
FAQ
Does a reverse proxy add noticeable latency to Claude API calls? A well-configured proxy in the same region as your backend adds low single-digit milliseconds. Cross-region hops or synchronous logging on the critical path are the usual causes of real slowdowns.
Can I proxy streaming responses through nginx without extra work? No — you must set proxy_buffering off and avoid proxy_cache, otherwise nginx buffers the full response before sending it, breaking the streaming UX even though the request itself succeeds.
Should I build my own proxy or use a hosted one like SubToAPI? Build your own if you need custom per-tenant routing or provider-swapping logic. Use a hosted option if you just need app-scoped API keys, streaming, and usage tracking without maintaining infrastructure — see /pricing for plan details.