Self-Hosted LLM Gateway: Open Source Options Compared
If you're searching for a self-hosted LLM gateway that's open source, you're probably trying to put a single API in front of one or more model providers — normalizing requests, adding auth, tracking usage, maybe routing between models — without depending on a third-party SaaS for that layer. The short answer: several solid open source projects exist (LiteLLM, Portkey's gateway, Helicone's proxy, and a few others), and self-hosting one is a reasonable option if you have the operational capacity to run, patch, and monitor another service.
The longer answer depends on what you actually need. A gateway that just proxies requests and logs them is trivial to run. A gateway that handles per-team API keys, streaming, tool-use passthrough, retries, spend limits, and multi-provider routing is a real piece of infrastructure with its own uptime and security requirements. This article walks through the main open source options, what self-hosting actually costs you in engineering time, and when it makes more sense to use a managed layer instead.
What an LLM Gateway Actually Does
At minimum, an LLM gateway sits between your application and the model provider(s) and gives you:
- A stable internal API so you're not hardcoding provider-specific request shapes everywhere in your codebase
- Centralized API key management — one place to issue, rotate, and revoke keys per app or team
- Usage tracking — tokens in/out, cost, latency, per-key breakdowns
- Optional routing — sending requests to different models/providers based on rules, load, or cost
- Rate limiting and spend caps so one misbehaving app can't blow through your budget
Some gateways also handle streaming passthrough, tool/function calling, and request retries with backoff. Whether you need all of this depends on how many apps or teams are hitting the model, and whether cost and access control are already a problem for you.
Open Source Self-Hosted Options
LiteLLM Proxy is probably the most widely used option. It supports 100+ providers behind an OpenAI-compatible API, has built-in cost tracking, budget limits per key, and a Postgres-backed admin UI. It's a good fit if you need multi-provider routing (e.g., fallback from one model to another) and are comfortable running Python services plus a database.
Portkey's open source gateway focuses on reliability features — retries, fallbacks, caching, and observability hooks — with a lighter footprint than LiteLLM's full proxy. It's a good choice if your main pain point is flaky upstream calls rather than multi-tenant key management.
Helicone started as an observability proxy and has grown gateway-like features (caching, rate limiting). If logging and cost visibility are your primary driver, it's worth a look, though it's less focused on being a full API abstraction layer.
Cloudflare AI Gateway is technically not self-hosted in the "run it on your own servers" sense, but it's worth mentioning as a middle ground — you configure it, but Cloudflare runs the infrastructure.
All of these are legitimate projects. None of them are hard to get a basic setup running; a single docker run will get LiteLLM proxying requests in under ten minutes. The real cost shows up later.
What Self-Hosting Actually Costs You
The setup step is the easy part. The ongoing cost is what determines whether self-hosting is the right call:
- Uptime ownership. If your gateway goes down, every app behind it goes down. That's now your incident to own, at 3am, with your own alerting.
- Database and state. Most of these gateways need Postgres or Redis for keys, usage logs, and rate-limit counters. That's another stateful service to back up and patch.
- Security surface. You're now storing provider API keys and issuing your own sub-keys. A misconfigured gateway is a credential leak waiting to happen.
- Upgrades. Open source projects move fast. Staying current means tracking breaking changes in request/response shapes, especially around streaming and tool calling.
- Scaling. A gateway that works fine for 50 requests/minute needs load testing and horizontal scaling once you're past a few hundred.
None of this is a reason to avoid self-hosting — plenty of teams run LiteLLM in production without issues. But it's real engineering time, not a one-off setup task, and it's worth being honest about that before committing.
When Self-Hosting Makes Sense
Self-host an open source gateway if:
- You need to route across multiple providers (not just one model family) as a core requirement
- You already run infrastructure like this (you have on-call, monitoring, and deploy pipelines in place)
- Data residency or compliance requires the proxy layer to run inside your own network
- You want full control over routing logic and are willing to maintain it
When a Managed Option Makes More Sense
If your actual need is simpler — turn your existing Claude access into a clean HTTPS API with per-app keys, streaming, and usage metadata, without running a database and a proxy service yourself — a managed layer gets you there faster. SubToAPI does exactly that: application API keys (sub_live_...), streaming, tool use, and usage metadata in one dashboard, on top of the Claude access you already have. There's no infrastructure to patch and no proxy to keep online yourself.
A basic streaming request looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"stream": true,
"messages": [{"role": "user", "content": "Summarize this PR."}]
}'
You get team seats and per-key usage out of the box, which is usually the part teams build first when they self-host a gateway anyway. Plans start with Solo at €9/month, Team at €19/seat, and Scale at €49/seat, with a free trial at /signup. See /pricing for the full breakdown or /docs/quickstart to get a key running in a few minutes.
If your requirement is genuinely multi-provider routing across many model vendors, an open source gateway you run yourself is still the right tool. If your requirement is "give my apps and teammates clean, metered API access to the Claude access we already pay for," a managed layer is the faster and lower-maintenance path.
FAQ
Is LiteLLM free to self-host? Yes, the core proxy is open source (MIT licensed) and free to run yourself. Costs come from the infrastructure you run it on — compute, database, monitoring — not licensing.
Can I use an open source gateway with Claude specifically? Yes, most of these gateways (LiteLLM, Portkey, Helicone) support Anthropic's API alongside other providers, so you can route Claude requests through them like any other model.
What's the difference between an LLM gateway and just calling the provider API directly? A gateway adds a layer for key management, usage tracking, and routing across your own apps or teams. Calling the provider directly works fine for a single app, but gets harder to manage once multiple teams or projects share the same underlying access — see /docs/messages for how a managed gateway like SubToAPI structures this.