← Blog

Self-Hosted LLM Gateway: Open Source Options Compared

2026-09-29 · 5 min read · SubToAPI Team

If you're searching for a self-hosted LLM gateway that's open source, you're probably trying to put a single API in front of one or more model providers — normalizing requests, adding auth, tracking usage, maybe routing between models — without depending on a third-party SaaS for that layer. The short answer: several solid open source projects exist (LiteLLM, Portkey's gateway, Helicone's proxy, and a few others), and self-hosting one is a reasonable option if you have the operational capacity to run, patch, and monitor another service.

The longer answer depends on what you actually need. A gateway that just proxies requests and logs them is trivial to run. A gateway that handles per-team API keys, streaming, tool-use passthrough, retries, spend limits, and multi-provider routing is a real piece of infrastructure with its own uptime and security requirements. This article walks through the main open source options, what self-hosting actually costs you in engineering time, and when it makes more sense to use a managed layer instead.

What an LLM Gateway Actually Does

At minimum, an LLM gateway sits between your application and the model provider(s) and gives you:

Some gateways also handle streaming passthrough, tool/function calling, and request retries with backoff. Whether you need all of this depends on how many apps or teams are hitting the model, and whether cost and access control are already a problem for you.

Open Source Self-Hosted Options

LiteLLM Proxy is probably the most widely used option. It supports 100+ providers behind an OpenAI-compatible API, has built-in cost tracking, budget limits per key, and a Postgres-backed admin UI. It's a good fit if you need multi-provider routing (e.g., fallback from one model to another) and are comfortable running Python services plus a database.

Portkey's open source gateway focuses on reliability features — retries, fallbacks, caching, and observability hooks — with a lighter footprint than LiteLLM's full proxy. It's a good choice if your main pain point is flaky upstream calls rather than multi-tenant key management.

Helicone started as an observability proxy and has grown gateway-like features (caching, rate limiting). If logging and cost visibility are your primary driver, it's worth a look, though it's less focused on being a full API abstraction layer.

Cloudflare AI Gateway is technically not self-hosted in the "run it on your own servers" sense, but it's worth mentioning as a middle ground — you configure it, but Cloudflare runs the infrastructure.

All of these are legitimate projects. None of them are hard to get a basic setup running; a single docker run will get LiteLLM proxying requests in under ten minutes. The real cost shows up later.

What Self-Hosting Actually Costs You

The setup step is the easy part. The ongoing cost is what determines whether self-hosting is the right call:

None of this is a reason to avoid self-hosting — plenty of teams run LiteLLM in production without issues. But it's real engineering time, not a one-off setup task, and it's worth being honest about that before committing.

When Self-Hosting Makes Sense

Self-host an open source gateway if:

When a Managed Option Makes More Sense

If your actual need is simpler — turn your existing Claude access into a clean HTTPS API with per-app keys, streaming, and usage metadata, without running a database and a proxy service yourself — a managed layer gets you there faster. SubToAPI does exactly that: application API keys (sub_live_...), streaming, tool use, and usage metadata in one dashboard, on top of the Claude access you already have. There's no infrastructure to patch and no proxy to keep online yourself.

A basic streaming request looks like this:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 1024,
    "stream": true,
    "messages": [{"role": "user", "content": "Summarize this PR."}]
  }'

You get team seats and per-key usage out of the box, which is usually the part teams build first when they self-host a gateway anyway. Plans start with Solo at €9/month, Team at €19/seat, and Scale at €49/seat, with a free trial at /signup. See /pricing for the full breakdown or /docs/quickstart to get a key running in a few minutes.

If your requirement is genuinely multi-provider routing across many model vendors, an open source gateway you run yourself is still the right tool. If your requirement is "give my apps and teammates clean, metered API access to the Claude access we already pay for," a managed layer is the faster and lower-maintenance path.

FAQ

Is LiteLLM free to self-host? Yes, the core proxy is open source (MIT licensed) and free to run yourself. Costs come from the infrastructure you run it on — compute, database, monitoring — not licensing.

Can I use an open source gateway with Claude specifically? Yes, most of these gateways (LiteLLM, Portkey, Helicone) support Anthropic's API alongside other providers, so you can route Claude requests through them like any other model.

What's the difference between an LLM gateway and just calling the provider API directly? A gateway adds a layer for key management, usage tracking, and routing across your own apps or teams. Calling the provider directly works fine for a single app, but gets harder to manage once multiple teams or projects share the same underlying access — see /docs/messages for how a managed gateway like SubToAPI structures this.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →