LLM API Gateway Open Source Options: A 2025 Review
LLM API Gateway Open Source Options: A 2025 Review
If you're searching for an open source LLM API gateway, you're likely trying to unify multiple model providers behind one interface, add retries and fallbacks, or avoid vendor lock-in without paying for a managed platform. The short answer: the mature open source options are LiteLLM, Portkey's open source gateway, Helicone, and Kong with AI plugins — each with different tradeoffs around self-hosting effort, feature completeness, and ongoing maintenance.
This article breaks down what each option actually does, where self-hosting an open source gateway makes sense, and where it creates more work than it saves — especially if your actual need is a clean per-application API key system in front of a single provider like Claude.
What an LLM API gateway actually does
Before comparing tools, it helps to be precise about the job:
- Request routing — sending calls to the right provider/model based on config or fallback rules
- Authentication abstraction — issuing your own API keys instead of exposing raw provider keys to every service or team
- Usage tracking — logging tokens, cost, and latency per key, project, or customer
- Resilience — retries, timeouts, rate limit handling, and failover between providers
- Streaming support — proxying SSE/streaming responses without buffering the whole response
Open source tools cover these to varying degrees. None of them eliminate the operational burden of running a service — you still own uptime, scaling, and security patches.
LiteLLM
LiteLLM is the most widely adopted open source option for multi-provider routing. It exposes an OpenAI-compatible endpoint and translates requests to over 100 providers, including Anthropic, OpenAI, and various open-weight model hosts.
Strengths:
- Drop-in replacement for OpenAI SDK calls in many codebases
- Built-in retry/fallback logic between providers
- A proxy server mode with virtual API keys and budget limits
Tradeoffs:
- You run and patch the proxy yourself (Docker or Python package)
- The budgeting and usage dashboard features are basic compared to commercial billing tools
- Team/seat management and invoicing are not built in — you'd build that layer yourself
LiteLLM is a strong choice if your main problem is routing logic across providers and you're comfortable owning the deployment.
Portkey (open source gateway)
Portkey ships an open source gateway core with observability and caching, and a hosted layer for teams that don't want to self-manage it. The gateway itself handles retries, load balancing across API keys, and semantic caching.
Strengths:
- Good request/response logging out of the box
- Caching reduces redundant calls for repeated prompts
- Works as a lightweight proxy in front of existing provider keys
Tradeoffs:
- Full observability and guardrail features often require the hosted plan
- Multi-tenant billing and seat-based access control are hosted-only features, not part of the open core
Helicone
Helicone focuses more narrowly on observability — logging, cost tracking, and prompt analytics — and can run as a lightweight proxy in front of provider calls. It's less of a full gateway and more of a logging layer you insert into an existing setup.
Strengths:
- Easy to add to an existing integration with a one-line base URL change
- Good for debugging prompt/response history
Tradeoffs:
- Not designed for issuing scoped application keys or billing customers per usage
- No native tool-use or streaming-specific tooling beyond passthrough
Kong / Gloo with AI plugins
If you already run an API gateway like Kong or Gloo for non-LLM traffic, there are AI-specific plugins that add token counting, prompt templating, and basic routing to LLM backends. This fits teams standardizing all API traffic — LLM and otherwise — through one gateway layer.
Tradeoffs: these plugins are newer and less battle-tested specifically for streaming LLM responses and tool-calling patterns than purpose-built LLM gateways.
When self-hosting an open source gateway makes sense
Self-hosting is the right call when:
- You need to route across many different providers and models dynamically
- You have the engineering capacity to run, monitor, and secure another service
- Cost tracking across providers (not billing your own customers) is the primary goal
- You're comfortable building your own key-issuance and seat-management layer on top
When it doesn't
Most of the open source options above assume you're normalizing multiple providers. If your actual stack is built around Claude specifically, and what you need is:
- Per-application API keys instead of sharing one raw Anthropic key
- Usage metadata per key for billing or internal chargebacks
- Streaming and tool use working correctly without extra proxy config
- Team seats so multiple people can manage keys without sharing credentials
...then a hosted, Claude-specific API gateway can get you there faster than standing up and maintaining an open source multi-provider proxy. This is exactly what SubToAPI does: it turns your existing Claude access into a proper HTTPS API with sub_live_... keys, streaming, tool use, and per-key usage data — without you running any infrastructure.
A minimal example of calling it looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Summarize this ticket in two sentences."}
]
}'
No proxy to deploy, no Docker image to patch, no self-managed virtual key store. You create keys per application or teammate from a dashboard, and usage shows up automatically. Setup takes about the time it takes to read the quickstart.
Choosing between self-hosted and hosted
A reasonable decision framework:
- Multiple providers, cost optimization across them → open source gateway (LiteLLM is the default choice)
- One primary provider, need for clean key management and billing-ready usage data → hosted option purpose-built for that provider
- Already running a general API gateway, want to bolt on LLM traffic → Kong/Gloo AI plugins
- Mainly need prompt/response logging for debugging → Helicone as a lightweight addition
Plans for the hosted route start at Solo (€9) for individual use, Team (€19/seat) for shared key management, and Scale (€49/seat) for higher-volume teams — all listed on the pricing page, with a free trial at signup.
FAQ
Is LiteLLM free to use in production? Yes, the core proxy and SDK are open source and free to self-host. Costs come from the infrastructure you run it on and the engineering time to maintain it, not licensing fees.
Can I combine an open source gateway with a hosted service like SubToAPI? Yes — some teams route multi-provider traffic through an open source gateway and use a hosted service specifically for Claude-based applications that need clean key issuance and usage tracking without extra proxy maintenance.
Do open source gateways support streaming and tool use out of the box? Most pass through streaming responses, but tool-calling support varies by tool and provider version. Check the specific gateway's documentation against the Claude API's tool use and streaming specs before relying on it in production.