Claude API Middleware for Node Express: Setup Guide
What "Claude API middleware" actually means
If you're searching for this, you're probably building an Express app that calls Claude and you don't want every route handler duplicating the same boilerplate: API key checks, request formatting, error handling, retries, and logging. The fix is a middleware layer — a function (or small stack of functions) that sits between the incoming HTTP request and your Claude API call, handling the repetitive plumbing so your route handlers stay thin.
This isn't a special Anthropic SDK feature — it's a standard Express pattern applied to LLM calls. You write middleware functions that validate input, attach a configured client, catch errors consistently, and optionally stream the response back to the client. Below is a practical implementation you can drop into an existing Express app, plus the parts that are easy to get wrong (timeouts, retries, streaming, rate limits).
What the middleware layer should actually do
Before writing code, define scope. A good Claude middleware layer for Express typically handles:
- Request validation — reject malformed payloads before they hit the API and burn a request.
- Auth — verify the caller (your app's own users or internal services), separate from the Claude API key itself.
- Client injection — attach a pre-configured client instance to
reqso handlers don't re-instantiate it. - Error normalization — turn Claude API errors (rate limits, overloaded, invalid request) into consistent HTTP responses.
- Retry logic — exponential backoff for transient 429/529 errors.
- Logging/metrics — token usage, latency, and status codes for observability.
- Streaming passthrough — forward server-sent events to the client without buffering the whole response.
You don't need all of these on day one, but structuring your code so each concern is its own middleware function makes it easy to add later.
A basic middleware stack
Here's a minimal setup using express and the Anthropic SDK:
import express from "express";
import Anthropic from "@anthropic-ai/sdk";
const app = express();
app.use(express.json());
// 1. Attach a configured client to every request
function attachClient(req, res, next) {
req.claude = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
next();
}
// 2. Validate the incoming payload
function validateMessageRequest(req, res, next) {
const { messages, model } = req.body;
if (!Array.isArray(messages) || messages.length === 0) {
return res.status(400).json({ error: "messages array is required" });
}
req.claudeModel = model || "claude-3-5-sonnet-20241022";
next();
}
// 3. Normalize errors from the API
async function callClaude(req, res, next) {
try {
const response = await req.claude.messages.create({
model: req.claudeModel,
max_tokens: 1024,
messages: req.body.messages,
});
res.locals.claudeResponse = response;
next();
} catch (err) {
if (err.status === 429) {
return res.status(429).json({ error: "rate_limited", retry_after: err.headers?.["retry-after"] });
}
if (err.status === 529) {
return res.status(503).json({ error: "overloaded_try_again" });
}
console.error("claude_error", err);
return res.status(502).json({ error: "upstream_error" });
}
}
app.post("/chat", attachClient, validateMessageRequest, callClaude, (req, res) => {
res.json(res.locals.claudeResponse);
});
app.listen(3000);
This gives you a clean separation: validation fails fast, the client is reusable, and errors are mapped to sensible status codes instead of leaking raw SDK exceptions to your frontend.
Adding retries without over-engineering it
Transient errors (429, 529) are common enough that a naive retry wrapper pays for itself quickly:
async function withRetry(fn, { retries = 3, baseDelay = 500 } = {}) {
for (let attempt = 0; attempt < retries; attempt++) {
try {
return await fn();
} catch (err) {
const retryable = err.status === 429 || err.status === 529;
if (!retryable || attempt === retries - 1) throw err;
const delay = baseDelay * 2 ** attempt;
await new Promise((r) => setTimeout(r, delay));
}
}
}
Wrap the messages.create call in withRetry inside your middleware. Keep retry counts low (2–3) — Claude calls aren't free, and retrying aggressively just multiplies cost during an outage.
Streaming through middleware
If your Express app needs to stream tokens to the browser, the middleware pattern still works, but you skip buffering the full response:
app.post("/chat/stream", attachClient, validateMessageRequest, async (req, res) => {
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
try {
const stream = await req.claude.messages.create({
model: req.claudeModel,
max_tokens: 1024,
messages: req.body.messages,
stream: true,
});
for await (const event of stream) {
res.write(`data: ${JSON.stringify(event)}\n\n`);
}
res.end();
} catch (err) {
res.write(`data: ${JSON.stringify({ error: "stream_failed" })}\n\n`);
res.end();
}
});
The key detail: don't run this through a middleware that buffers res (like some compression middlewares) or the stream will arrive in one chunk instead of progressively.
Where SubToAPI fits if you'd rather not maintain this
Everything above is reasonable for a single app, but once you're running multiple services, need per-key usage tracking, or want team members issuing their own keys without sharing the root Anthropic credential, maintaining this middleware yourself gets tedious — especially the auth separation, rate-limit handling, and usage logging pieces.
SubToAPI wraps that layer for you: it turns your Claude access into an HTTPS API with scoped application keys (sub_live_...), so your Express app just calls https://api.subtoapi.app/v1/messages with Authorization: Bearer $SUBTOAPI_KEY instead of managing retries, key rotation, and usage metadata yourself. Streaming and tool use work the same way you'd expect from the native API — see the streaming docs and tool use docs. If you're just getting your Express integration running, the quickstart covers the first request end to end, and the Messages API reference documents the request/response shape your middleware needs to validate against.
Questions
Do I need middleware if I'm only calling Claude from one route? Not strictly — a single try/catch in the handler is fine. Middleware pays off once you have multiple routes or services sharing auth, validation, or retry logic.
Should retry logic live in middleware or in the SDK call itself? Either works; middleware is cleaner if multiple routes need the same retry policy. Keep retry counts low and only retry on 429/529 status codes.
Can this middleware pattern handle streaming and tool use together? Yes — treat them as separate middleware chains (one for stream: true requests, one for standard JSON responses) since streaming requires different response headers and write behavior.