How to Proxy Claude API Requests: A Practical Guide
Proxying Claude API requests means routing calls through your own server (or a managed service) instead of hitting Anthropic's endpoint directly from client code. You do this to hide your real API key from end users, add logging or rate limiting, normalize responses for your app, or issue separate keys to different teams and products without sharing one master credential.
The short answer: you build a thin HTTP server that accepts requests from your app, forwards them to Anthropic with your actual credentials, and streams the response back. Below is exactly how to do that, what breaks when you try it at scale, and when it's faster to use a managed proxy instead of maintaining your own.
Why proxy Claude API calls at all
Calling the Claude API directly from a browser or mobile app is a non-starter — your key would be exposed in the request headers of every client. Even in server-to-server setups, teams end up building a proxy layer for a few recurring reasons:
- Key isolation. One Anthropic account, multiple internal apps or customers, each with its own revocable key.
- Usage tracking. You need to know which product, team, or customer is burning tokens, not just a total on Anthropic's dashboard.
- Rate limiting and quotas. Per-user or per-app limits that Anthropic's own rate limits don't give you out of the box.
- Request shaping. Injecting system prompts, enforcing model choice, stripping or validating tool definitions before they reach the model.
- Streaming and tool use pass-through. Your proxy has to handle Server-Sent Events and tool-call payloads correctly, not just plain JSON.
If you only need one or two of these, a few lines of middleware might be enough. If you need all of them reliably, you're building (or buying) a real API gateway.
Building a minimal Claude API proxy
Here's a bare-bones Node.js/Express proxy that forwards requests to Anthropic, adds a simple API key check, and logs token usage:
import express from "express";
import fetch from "node-fetch";
const app = express();
app.use(express.json());
const INTERNAL_KEYS = new Set(["app-1-key", "app-2-key"]);
app.post("/v1/messages", async (req, res) => {
const incomingKey = req.headers["x-internal-key"];
if (!INTERNAL_KEYS.has(incomingKey)) {
return res.status(401).json({ error: "invalid key" });
}
const upstream = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify(req.body),
});
const data = await upstream.json();
console.log("usage:", data.usage, "caller:", incomingKey);
res.status(upstream.status).json(data);
});
app.listen(3000);
This works for basic, non-streaming calls. The moment you need more, the complexity grows fast:
- Streaming requires piping the SSE body through without buffering, handling client disconnects, and forwarding heartbeat events correctly.
- Tool use means parsing
tool_useblocks, round-trippingtool_resultmessages, and validating JSON schemas before they're sent upstream — a bug here silently breaks agent workflows. - Retry logic for 429s and 5xxs needs backoff that respects Anthropic's
retry-afterheader, or you'll make rate limiting worse, not better. - Multi-tenant keys need their own storage, rotation, and revocation — not a hardcoded
Setin your server file. - Observability means structured logs per key, per model, per token count, searchable later — not
console.log.
None of this is hard in isolation. It's the combination, maintained over time, that turns a "quick proxy" into an on-call liability.
Streaming through a proxy correctly
If you're forwarding streaming responses, don't buffer the body before sending it to the client — pipe it directly:
app.post("/v1/messages", async (req, res) => {
const upstream = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({ ...req.body, stream: true }),
});
res.setHeader("Content-Type", "text/event-stream");
upstream.body.pipe(res);
});
This is the minimum. A production proxy also needs to handle partial chunks that split across TCP packets, client aborts mid-stream, and timeouts that don't leave open connections hanging on the upstream side.
Using a managed proxy instead
If the goal is just "give my app a Claude endpoint I control, with keys, logs, and usage data," building and maintaining this yourself is often more work than the feature is worth. SubToAPI is built specifically for this: it sits between your Claude access and your applications, issuing scoped sub_live_... keys per app or team, handling streaming and tool use correctly, and giving you usage metadata without you writing a proxy server at all.
In practice that means:
- You generate an application key from the dashboard instead of hardcoding Anthropic credentials into every service.
- Streaming and tool-call pass-through are handled for you — see /docs/streaming and /docs/tools.
- Each team or project gets its own key and visible usage, instead of one shared credential and a spreadsheet.
- Requests go through the same
/v1/messages-style interface documented at /docs/messages, so migrating existing proxy code is mostly a base-URL change.
A minimal call looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 512,
"messages": [{"role": "user", "content": "Summarize this proxy setup."}]
}'
Plans start at €9/month for solo use, with per-seat pricing for teams on /pricing, and a free trial at /signup if you want to compare it against a hand-rolled proxy before committing. The quickstart at /docs/quickstart covers the full setup in a few minutes.
Which approach to pick
Build your own proxy if you need very specific request transformations, have an existing gateway infrastructure you must integrate with, or you're only forwarding a handful of simple, non-streaming calls. Reach for a managed layer like SubToAPI if you want per-app keys, streaming, tool use, and usage tracking without owning the operational burden of running and patching that server yourself.
FAQ
Is proxying the Claude API against Anthropic's terms?
No — proxying is a normal architectural pattern for hiding credentials and adding infrastructure like logging or rate limiting. What matters is that your usage still complies with Anthropic's acceptable use policies regardless of how requests are routed.
Does a proxy add latency to Claude API calls?
Yes, a small amount — typically single-digit to low double-digit milliseconds for the extra network hop, assuming your proxy is reasonably close to Anthropic's infrastructure and isn't doing heavy synchronous processing per request.
Can a proxy handle streaming and tool use, not just plain text?
Yes, but it has to be built for it specifically — piping SSE chunks without buffering and correctly round-tripping tool_use/tool_result blocks. A basic request-response proxy will break both unless it's designed with these in mind from the start.