Claude API Middleware for Express Apps: A Setup Guide
If you're calling Claude from an Express app, you're probably duplicating the same logic across every route that touches the model: checking API keys, validating request bodies, handling streaming, catching timeouts, and logging usage. Middleware fixes this by centralizing that logic into a single reusable layer that every Claude-related route passes through.
This article walks through building Claude API middleware for Express — from a minimal version that just injects auth headers, to a more complete setup that handles validation, error normalization, rate limiting, and logging. It also covers when it makes more sense to skip the middleware entirely and route through a managed API layer instead.
Why middleware instead of inline calls
Calling Claude directly inside each route handler works fine for a prototype. It stops working once you have more than two or three endpoints that need the model, because:
- Auth header construction gets copy-pasted everywhere
- Error handling (rate limits, timeouts, malformed responses) diverges between routes
- You lose a single place to log token usage or latency
- Streaming setup (SSE headers, chunk parsing) gets reimplemented per route
Middleware solves this by intercepting the request before it reaches your Claude-calling logic, attaching whatever context is needed (client, headers, validated payload), and giving you one place to handle failures consistently.
A minimal Claude middleware
Here's a basic middleware that validates the incoming request shape and attaches a preconfigured client to req:
const Anthropic = require("@anthropic-ai/sdk");
const anthropic = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });
function claudeMiddleware(req, res, next) {
const { messages, model } = req.body;
if (!Array.isArray(messages) || messages.length === 0) {
return res.status(400).json({ error: "messages array is required" });
}
req.claude = anthropic;
req.claudeModel = model || "claude-3-5-sonnet-20241022";
next();
}
module.exports = claudeMiddleware;
Then in your route:
app.post("/chat", claudeMiddleware, async (req, res) => {
try {
const response = await req.claude.messages.create({
model: req.claudeModel,
max_tokens: 1024,
messages: req.body.messages,
});
res.json(response);
} catch (err) {
res.status(502).json({ error: "claude_request_failed", detail: err.message });
}
});
This is enough to stop repeating validation logic, but error handling is still duplicated in every route.
Centralizing error handling
Move the try/catch into the middleware chain using Express's error-handling middleware pattern:
function errorHandler(err, req, res, next) {
if (err.status === 429) {
return res.status(429).json({ error: "rate_limited", retryAfter: err.headers?.["retry-after"] });
}
if (err.status === 401) {
return res.status(401).json({ error: "invalid_api_key" });
}
console.error("Claude API error:", err);
res.status(502).json({ error: "upstream_error" });
}
app.use(errorHandler);
Routes become simpler — just next(err) instead of inline response formatting:
app.post("/chat", claudeMiddleware, async (req, res, next) => {
try {
const response = await req.claude.messages.create({
model: req.claudeModel,
max_tokens: 1024,
messages: req.body.messages,
});
res.json(response);
} catch (err) {
next(err);
}
});
Adding request logging and usage tracking
Since the middleware already sits between every request and the model, it's the natural place to log token usage for billing or debugging:
function usageLogger(req, res, next) {
const start = Date.now();
const originalJson = res.json.bind(res);
res.json = (body) => {
const latency = Date.now() - start;
if (body?.usage) {
console.log({
route: req.path,
input_tokens: body.usage.input_tokens,
output_tokens: body.usage.output_tokens,
latency_ms: latency,
});
}
return originalJson(body);
};
next();
}
Place this before your route handlers so it wraps the response before it's sent.
Handling streaming in middleware
Streaming needs different headers and a different response lifecycle than a normal JSON response, so it's worth a dedicated middleware branch:
function streamMiddleware(req, res, next) {
if (req.body.stream) {
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
}
next();
}
The actual stream consumption still happens in the route, but keeping header setup in middleware avoids forgetting it on a new endpoint.
When middleware isn't enough
Writing your own middleware stack gets you consistent auth and error handling, but it doesn't solve a few things that come up quickly in production:
- Per-application keys. If multiple internal apps or external customers call your Claude integration, you need scoped API keys, not one shared
ANTHROPIC_API_KEYpassed around. - Team visibility. Someone needs to see usage and spend across routes and apps without grepping logs.
- Key rotation and seats. Revoking access for one app shouldn't require redeploying all of them.
This is the gap SubToAPI fills. Instead of writing and maintaining your own middleware for auth, logging, and multi-key management, you point your Express app at a single HTTPS endpoint with a sub_live_... key:
app.post("/chat", async (req, res, next) => {
try {
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: req.body.messages,
}),
});
const data = await response.json();
res.json(data);
} catch (err) {
next(err);
}
});
Each application gets its own key, usage metadata comes back with every response, and team members get dashboard visibility without you building a logging layer. Streaming and tool use work the same way as calling Claude directly — see the streaming docs and tool use docs for the request shapes.
Picking the right approach
If you're running a single Express app with one Claude integration, the custom middleware pattern above is enough — it's lightweight and you own every piece of it. If you're running several apps, giving different teams access, or need per-app usage data without building it yourself, routing through SubToAPI removes that maintenance burden. You can start with the quickstart and compare plans before deciding which fits your setup.
Questions
Does middleware replace the Anthropic SDK? No. Middleware wraps the SDK call (or an HTTP request) with validation, auth, and error handling — it doesn't replace the actual model call.
Can this middleware pattern work with streaming responses? Yes, but streaming needs separate header setup and a different response lifecycle than JSON responses, so it's usually split into its own middleware function.
Why use SubToAPI instead of writing my own key management middleware? Writing per-app keys, usage logging, and team access controls from scratch takes real engineering time. SubToAPI provides this out of the box via a single HTTPS endpoint — see the docs for setup details.