← Blog

Claude API Reverse Proxy Nginx Config Guide

2026-10-05 · 5 min read · SubToAPI Team

Claude API reverse proxy nginx config: the short answer

If you're searching for this, you're probably trying to put nginx between your application and api.anthropic.com so you can hide your API key from client-side code, add rate limiting, log requests, handle CORS, or centralize auth for a team. That's a reasonable architecture, and nginx can do it — but there are a few gotchas specific to proxying a streaming, API-key-authenticated service like Claude's, and they trip people up if you don't know about them ahead of time.

Below is a working nginx config you can adapt, followed by the specific issues (SSE streaming, timeouts, header handling) that are unique to proxying an LLM API rather than a typical REST backend. At the end we also cover why some teams skip building this themselves entirely.

Why reverse proxy the Claude API at all

Common reasons to put nginx in front of the Anthropic API instead of calling it directly from your app:

A working nginx config

Here's a baseline config that proxies /v1/messages requests to the Anthropic API, injects the API key server-side, and is configured correctly for streaming:

server {
    listen 443 ssl;
    server_name proxy.yourdomain.com;

    ssl_certificate     /etc/letsencrypt/live/proxy.yourdomain.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/proxy.yourdomain.com/privkey.pem;

    location /v1/messages {
        # Strip any client-supplied auth, inject your real key
        proxy_set_header x-api-key "YOUR_ANTHROPIC_API_KEY";
        proxy_set_header anthropic-version "2023-06-01";
        proxy_set_header Content-Type "application/json";
        proxy_set_header Host api.anthropic.com;

        proxy_pass https://api.anthropic.com/v1/messages;

        # Required for SSE streaming responses
        proxy_http_version 1.1;
        proxy_buffering off;
        proxy_cache off;
        chunked_transfer_encoding on;

        # Streaming responses can run long
        proxy_read_timeout 300s;
        proxy_send_timeout 300s;

        # Basic CORS for browser clients
        add_header Access-Control-Allow-Origin "https://yourapp.com" always;
        add_header Access-Control-Allow-Headers "Content-Type, Authorization" always;
        add_header Access-Control-Allow-Methods "POST, OPTIONS" always;

        if ($request_method = OPTIONS) {
            return 204;
        }
    }
}

The parts that actually matter

proxy_buffering off is non-negotiable for streaming. If you leave buffering on, nginx waits to accumulate the full response before forwarding it, which defeats the entire purpose of Claude's streaming mode — your client sees nothing until the full response is done, then gets it all at once. Same with proxy_cache off.

proxy_http_version 1.1 plus chunked_transfer_encoding on are needed so chunked SSE data passes through correctly rather than getting mangled or buffered by default HTTP/1.0 proxy behavior.

Timeouts need to be generous. Long completions with extended thinking or large max_tokens can run for well over nginx's default 60-second proxy_read_timeout. If the connection dies mid-stream, your client gets a truncated response with no clear error. Set this explicitly based on your expected worst-case generation time.

Never let the client set x-api-key directly. If your proxy's only job is header injection and you don't strip incoming auth headers, a misconfigured location block can let a client override your key or leak it back in error responses. Explicitly set the header server-side and don't pass through whatever the client sent.

CORS headers need always so they're added even on error responses (4xx/5xx), not just 2xx — otherwise your browser console fills up with CORS errors that mask the real problem.

Rate limiting per client

If you want per-API-key or per-IP limits on top of Anthropic's own account-level limits, add a limit_req_zone:

limit_req_zone $http_x_client_id zone=per_client:10m rate=10r/m;

location /v1/messages {
    limit_req zone=per_client burst=5 nodelay;
    # ... rest of config above
}

This requires your clients to send an identifying header (X-Client-Id), which you validate and strip before forwarding upstream.

What this setup doesn't give you

A bare nginx proxy solves key isolation and basic rate limiting, but it doesn't give you per-team API keys, usage dashboards, seat management, or tool-use/streaming handling baked in — you'd be writing and maintaining all of that yourself in Lua/njs modules or a sidecar app, and keeping it in sync as Anthropic's API evolves.

This is the gap SubToAPI fills: it turns your existing Claude access into a proper HTTPS API with scoped sub_live_... keys per application, streaming and tool use handled correctly out of the box, usage metadata per key, and team seats — without you maintaining nginx config or an auth layer by hand. If you're building the proxy mainly to issue safe keys to multiple apps or teammates, it's worth comparing against rolling your own. Check the pricing and quickstart to see if it replaces what you were about to build.

If you do want to call the Claude API through SubToAPI directly instead of proxying Anthropic yourself, the request shape is a drop-in equivalent:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the messages docs and streaming docs for the full reference.

Testing your proxy

Before shipping, verify streaming actually works end to end:

curl -N https://proxy.yourdomain.com/v1/messages \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-4-5","max_tokens":256,"stream":true,"messages":[{"role":"user","content":"Count to 10"}]}'

The -N flag disables curl's own output buffering. If you see tokens arrive one chunk at a time rather than all at once at the end, your proxy_buffering off setting is working.

questions

Does nginx add latency to Claude API requests? Minimal — typically a few milliseconds for the extra hop, assuming the proxy is geographically close to your app servers. The bigger latency risk is misconfigured buffering, which can make streaming responses feel much slower than they are.

Can I use nginx to switch between multiple Anthropic API keys for failover? Yes, using upstream blocks with multiple backend definitions and proxy_next_upstream, though you're responsible for detecting rate-limit errors (429s) and triggering failover logic yourself — nginx won't understand Anthropic's specific error codes natively.

Is a reverse proxy enough to expose the Claude API safely to a mobile app? It's a solid first step for hiding your raw API key, but you still need to add your own request authentication (e.g., JWT or API key per client) in front of nginx, or any user of your app could call your proxy and spend your Anthropic credits.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →