LLM API Proxy Server Setup Guide
An LLM API proxy server sits between your application and the underlying model provider, handling authentication, request routing, logging, and sometimes caching or rate limiting. Setting one up means deciding where that proxy lives, how it issues credentials to your apps, and what it does with every request before forwarding it upstream.
This guide walks through the actual setup process: picking an architecture, configuring authentication, adding observability, and handling the operational details most tutorials skip — retries, streaming passthrough, and key rotation. If you're building this for a team that's already using Claude, you can also skip most of the DIY work with a hosted option like SubToAPI, which we'll cover at the end.
Why Put a Proxy in Front of an LLM API
Calling a provider's API directly from every service, script, and client app works fine for a single prototype. It breaks down fast once you have more than one consumer:
- Credential sprawl. The provider key ends up in multiple codebases, CI pipelines, and local
.envfiles, making rotation a nightmare. - No shared usage visibility. You can't see which internal team or app is burning through tokens without parsing logs from five different places.
- Inconsistent retry/rate-limit handling. Every service reimplements its own backoff logic, usually badly.
- Vendor lock-in at the code level. If you ever want to swap models or add a fallback, you're rewriting call sites everywhere.
A proxy server fixes this by becoming the single place that holds the real provider credential, issues scoped application keys, and normalizes request/response handling.
Core Components of an LLM API Proxy
Before writing any code, map out what the proxy actually needs to do:
- Authentication layer — validates incoming API keys from your apps and maps them to a real provider credential.
- Request router — forwards the validated request to the correct upstream endpoint (chat completions, embeddings, etc.).
- Streaming passthrough — relays Server-Sent Events or chunked responses without buffering the whole payload.
- Logging/metadata capture — records token usage, latency, and status codes per key.
- Error normalization — converts upstream errors into a consistent shape your clients can parse.
Setting Up a Minimal Proxy
If you're rolling your own, a simple reverse-proxy pattern with a lightweight Node server is enough to start. Here's a stripped-down example using Express:
import express from "express";
import fetch from "node-fetch";
const app = express();
app.use(express.json());
const KEY_MAP = {
"app_key_123": process.env.PROVIDER_API_KEY
};
app.post("/v1/messages", async (req, res) => {
const appKey = req.headers["authorization"]?.replace("Bearer ", "");
const providerKey = KEY_MAP[appKey];
if (!providerKey) {
return res.status(401).json({ error: "invalid_api_key" });
}
const upstream = await fetch("https://api.provider.com/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${providerKey}`,
"Content-Type": "application/json"
},
body: JSON.stringify(req.body)
});
res.status(upstream.status);
upstream.body.pipe(res);
});
app.listen(3000);
This covers the basics: key translation and passthrough. In production you'll need to add:
- Streaming support that doesn't buffer chunks before forwarding them
- Per-key rate limiting so one app can't exhaust your provider quota
- Structured logging with token counts pulled from response metadata
- TLS termination if you're not already behind a load balancer that handles it
- Health checks and timeouts so a slow upstream doesn't cascade into your proxy hanging
Handling Streaming Correctly
Streaming is where most homemade proxies break. If you buffer the full response before sending it to the client, you lose the entire point of streaming — your client waits just as long as a non-streamed call. Use a direct pipe (as in the example above) or, if you're using a framework with its own body parsing, make sure chunked transfer encoding is preserved end to end. Test this explicitly with a slow network simulation; a proxy that works fine on localhost can silently buffer under a reverse proxy like Nginx unless proxy_buffering off is set.
Key Rotation and Scoping
A proxy's main security value is that application keys are not the same as the provider credential. Design your key scheme so that:
- Each app or team gets its own key, independent of the upstream provider key
- Revoking one app key doesn't require rotating the provider credential
- Keys can be scoped to specific models or rate limits if your provider supports it
Store the mapping in a database, not a hardcoded object like the example above — you want to revoke and issue keys without redeploying the proxy.
Observability: What to Log Per Request
At minimum, capture:
- Timestamp, app key, and model used
- Input/output token counts (most providers return these in response metadata)
- Latency and HTTP status
- Whether the request was streamed
This data is what lets you answer "which app is costing us the most" without guessing, and it's the foundation for any later rate-limiting or billing logic.
Skipping the DIY Work
If you're already using Claude and just need the proxy layer — API keys, streaming, tool use, usage metadata, and team seats — building and maintaining this yourself is a lot of infrastructure for what is ultimately commodity plumbing. SubToAPI turns your existing Claude access into a clean HTTPS API with sub_live_... application keys, so you get the proxy pattern described above without running a server. Check the quickstart to see the request format, or the streaming docs and tool use docs if those are your main requirements. Plans start at €9/month for solo use, with team seat pricing for larger setups, and there's a free trial at signup.
Final Checklist Before Going to Production
- Provider credential is never exposed to client apps, only to the proxy
- Streaming passes through without buffering
- Every app has its own revocable key
- Usage metadata is logged per request, not just aggregated
- Retries and timeouts are configured on the upstream call, not left to clients to figure out
FAQ
Do I need a proxy if I only have one app calling the LLM API? Not strictly — direct calls work fine for a single consumer. A proxy starts paying off once you have multiple apps, team members, or need centralized usage tracking.
Can a proxy server break streaming responses? Yes, if it buffers the full response before forwarding. Make sure buffering is disabled at every layer (app server, reverse proxy, CDN) for streamed endpoints.
Is it cheaper to build a proxy or use a hosted one? Building is "free" in licensing cost but costs ongoing engineering time for maintenance, scaling, and security. For small teams, a hosted option like SubToAPI is usually cheaper once you count engineering hours.