Claude API Multi-Model Routing Setup Guide
What multi-model routing actually means
Multi-model routing is the practice of sending each request to a different Claude model based on what the task needs, instead of hardcoding one model for your whole application. A support bot might route simple FAQ questions to Haiku, escalate ambiguous tickets to Sonnet, and reserve Opus for cases that require deep reasoning or long-context analysis. The goal is to spend only what each request is worth — not pay Opus prices for a one-line classification task.
The setup itself is not complicated: you need a decision function that picks a model name, a consistent request shape across models, and a fallback path for when the chosen model is unavailable or rate-limited. The rest of this guide walks through building that in practice, with code you can drop into an existing backend.
Why you'd route instead of picking one model
Before building anything, it's worth being clear on what you're optimizing for:
- Cost — Haiku-class models can be 10-20x cheaper per token than Opus-class models. Routing cheap, high-volume traffic away from the expensive model is usually the single biggest cost lever you have.
- Latency — smaller models respond faster. If you're doing autocomplete, live chat suggestions, or anything user-facing and time-sensitive, routing low-stakes requests to a fast model matters more than raw capability.
- Capability match — some tasks genuinely need the largest model: long documents, multi-step reasoning, code review across many files. Sending those to a small model just produces worse output, so you still need an escalation path.
- Resilience — if one model tier is overloaded or down, routing to an alternate model keeps your app functioning instead of failing outright.
If none of these apply to you — fixed workload, fixed budget, one model is always correct — you don't need routing. Most production apps have a mix of traffic though, which is why routing pays off quickly.
Step 1: classify the request
The router needs a signal to make a decision. Common signals, from simplest to most sophisticated:
- Endpoint or route in your app (e.g.
/classifyalways uses the small model,/analyzealways uses the large one) — zero runtime cost, but coarse. - Input length — longer prompts often correlate with more complex tasks.
- A cheap pre-classifier — use the smallest model to label the task ("simple" vs "complex") before routing the real request. This adds one extra call but is usually fast and cheap enough to be worth it.
- Explicit client hints — your frontend already knows if this is a "quick answer" button or a "deep analysis" button; pass that through as metadata.
A simple, practical approach that avoids an extra API call:
function pickModel(request) {
const { task, inputLength, requiresTools } = request;
if (task === "classify" || task === "extract" || inputLength < 200) {
return "claude-haiku";
}
if (requiresTools || inputLength > 4000 || task === "analysis") {
return "claude-opus";
}
return "claude-sonnet"; // default middle tier
}
Keep the rules simple and explicit at first. Fancy heuristics are easy to add later; what matters early on is having any router instead of a single hardcoded model.
Step 2: build a single request wrapper
Once you have a model name, route the actual call through one function so the rest of your codebase doesn't care which model was picked:
async function callClaude(model, messages, opts = {}) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model,
max_tokens: opts.maxTokens ?? 1024,
messages,
stream: opts.stream ?? false,
}),
});
if (!res.ok) {
throw new Error(`Claude request failed: ${res.status}`);
}
return res.json();
}
Then your routing logic becomes a thin layer on top:
const model = pickModel(request);
const response = await callClaude(model, messages);
If you're already running Claude through SubToAPI, this is where the setup gets simpler than managing your own provider accounts: one API key per environment, the same request shape across models, and usage metadata on every response so you can see which model each request actually used and what it cost. See the messages docs for the full request/response reference and the quickstart if you're setting this up from scratch.
Step 3: add a fallback chain
Routing isn't complete until you handle the case where your chosen model fails — rate limit, timeout, or a transient error. A basic fallback chain tries progressively cheaper (and usually more available) models:
const fallbackOrder = ["claude-opus", "claude-sonnet", "claude-haiku"];
async function callWithFallback(preferredModel, messages) {
const start = fallbackOrder.indexOf(preferredModel);
for (const model of fallbackOrder.slice(start)) {
try {
return await callClaude(model, messages);
} catch (err) {
console.warn(`Model ${model} failed, trying next tier`);
}
}
throw new Error("All models failed");
}
This is also where streaming matters for user-facing apps — if you fall back mid-conversation, make sure your UI handles a stream that restarts. The streaming guide covers chunked response handling if your router needs to support stream: true across model tiers.
Step 4: route tool-using requests deliberately
If some of your requests use tool calling, route those explicitly rather than relying on length or task-type heuristics — tool definitions add complexity that smaller models handle inconsistently. A safe default is to always send tool-using requests to your mid or top tier model, and keep your tool schemas identical across models so the router doesn't need model-specific logic. The tools documentation has the schema format if you're setting this up for the first time.
Step 5: monitor what the router is actually doing
Routing rules drift from reality fast — a "simple" classification task might start getting longer inputs over time and silently need escalation. Track, per model tier: request volume, average cost, error rate, and latency. If you're on SubToAPI, this comes from the usage metadata already attached to each response, visible per key in the dashboard — useful for seeing whether your routing rules match actual traffic patterns before you tune them further. Check pricing if you're comparing per-seat costs against your current multi-provider setup, or sign up to test routing against a live key on the free trial.
Keep the router dumb, iterate from data
The most common mistake in multi-model routing setups is over-engineering the classifier before you have any traffic data. Start with three or four explicit rules based on endpoint or task type, ship it, and look at where the router's decisions diverge from what the output quality actually needed. Tightening rules based on real usage beats guessing in advance every time.
questions
Do I need a separate API key per model? No — with a single provider key (including a SubToAPI key), you select the model per request via the model field. Keys don't need to be model-specific.
Should routing happen client-side or server-side? Server-side, almost always. It keeps your routing rules, fallback logic, and cost controls out of the client and lets you change them without shipping a new app version.
What's the simplest routing rule to start with? Route by endpoint or feature flag: fixed, low-stakes features always use the smallest model, and anything involving long documents or tool use always uses the largest. Add nuance only after you have usage data.