Unified LLM API for Claude and OpenAI: A Practical Guide
If you're searching for a "unified LLM API for Claude and OpenAI," you're probably trying to solve one specific problem: your product calls GPT-4 and Claude from different code paths, with different SDKs, different auth schemes, and different response shapes, and it's becoming a maintenance headache. You want one interface in your codebase that can route to either model without every feature team having to learn two APIs.
The short answer: there's no single hosted endpoint that natively speaks both Anthropic's and OpenAI's wire formats as if they were the same model family — the two companies don't share a protocol. What you actually build (or buy) is a thin abstraction layer: a consistent request/response shape in your own code, with a router underneath that calls each provider's real API and normalizes the output. This article covers how that layer is typically built, what to standardize, and where a tool like SubToAPI fits on the Claude side of that equation.
Why teams want a unified interface
Most teams don't actually need multi-provider routing on day one. They need it once one of these happens:
- A model gets deprecated or rate-limited and you need a fallback
- Pricing or quality shifts and you want to A/B test providers per feature
- Different teams in the org picked different vendors and now you're maintaining two integrations
- You want to let users choose a model without rewriting your backend
In all four cases, the fix isn't "find a magic single API" — it's isolating the provider-specific code behind a consistent internal contract so swapping providers is a config change, not a rewrite.
What to normalize
A workable unified layer typically standardizes four things:
- Auth — one env var pattern per provider, injected by your router, never scattered across call sites.
- Request shape — a common
{ model, messages, maxTokens, tools, stream }object that gets mapped to each provider's actual schema before the call. - Response shape — a common
{ content, usage, stopReason, raw }object, so your application code never has to branch on which provider answered. - Streaming — a single event format your frontend consumes, even if the underlying SSE payloads differ.
Here's a minimal router that does this for Claude (via SubToAPI) and OpenAI:
async function callLLM({ provider, model, messages, maxTokens = 1024 }) {
if (provider === "claude") {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ model, messages, max_tokens: maxTokens }),
});
const data = await res.json();
return {
content: data.content?.[0]?.text ?? "",
usage: data.usage,
stopReason: data.stop_reason,
raw: data,
};
}
if (provider === "openai") {
const res = await fetch("https://api.openai.com/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.OPENAI_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ model, messages, max_tokens: maxTokens }),
});
const data = await res.json();
return {
content: data.choices?.[0]?.message?.content ?? "",
usage: data.usage,
stopReason: data.choices?.[0]?.finish_reason,
raw: data,
};
}
throw new Error(`Unknown provider: ${provider}`);
}
Everything above the callLLM function — your prompt construction, retry logic, logging — stays provider-agnostic. That's the real win of a "unified" setup: not a single endpoint, but a single call site in your code.
Where the complexity actually lives
The hard part of multi-provider routing is rarely the request/response mapping — it's the stuff that differs structurally:
- Tool calling — Claude's tool_use blocks and OpenAI's function-calling format aren't 1:1. If you support tools in a unified layer, you need a normalization step for tool definitions and tool results on both sides, not just chat messages. (See /docs/tools if you're using SubToAPI for the Claude leg.)
- Streaming granularity — SSE event types and chunk boundaries differ between providers. Your frontend should consume one normalized stream format, not two.
- Rate limits and retries — each provider has its own limits and backoff behavior; your router needs per-provider retry policy, not a shared one.
- Usage accounting — token counts and cost per token differ, so if you want unified cost reporting, you need to compute cost per provider separately before aggregating.
Why handle the Claude side through an API layer at all
If your "Claude leg" of the router is just hitting Anthropic's raw API with a personal account, you inherit account-level auth, no per-application keys, and no usage breakdown by feature or team. That's fine for a prototype, but once you're routing production traffic from multiple internal apps through Claude, it's worth putting a proper API layer in front of it.
That's what SubToAPI does for the Claude side specifically: it turns your existing Claude access into a standard HTTPS API with sub_live_... application keys, streaming, tool use, and per-key usage metadata — so each app or team hitting Claude through your router has its own key and its own visibility, instead of one shared account token. You still write your own OpenAI call path the same way you always have; SubToAPI just makes the Claude half of the router production-grade instead of a single shared secret. Check /docs/quickstart for the setup, and /docs/streaming or /docs/tools if your router needs to normalize those specifically.
Pricing is per seat (Solo €9, Team €19/seat, Scale €49/seat — see /pricing), with a free trial at /signup, so you can wire up the Claude leg of your unified router without committing upfront.
A practical starting checklist
- Define your common request/response contract before writing any provider code
- Put each provider behind its own module, never inline API calls in business logic
- Normalize tool definitions separately from chat normalization — don't bolt it on later
- Decide your fallback policy (which provider is primary, when to fall back) explicitly, not implicitly
- Centralize API keys per provider, per application, so usage and cost are traceable
questions
Is there a single API that natively supports both Claude and OpenAI models? No — Anthropic and OpenAI have separate APIs and formats. A "unified" setup means building a thin router in your own codebase that normalizes requests and responses across both, not a single hosted multi-vendor endpoint.
Do I need a unified API if I only use one provider today? Not immediately, but structuring your Claude calls behind a clean contract from the start (via a service like SubToAPI for auth and keys) makes adding a second provider later a config change instead of a rewrite.
How do I handle tool/function calling in a unified router? Normalize tool definitions and tool results separately from chat messages — Claude's tool_use format and OpenAI's function-calling format aren't directly compatible, so this needs its own mapping layer. See /docs/tools for the Claude-side schema.