Claude API Integration with Node.js Backend: Setup Guide
Integrating Claude into a Node.js backend means adding a service layer that authenticates requests, calls the Messages API, handles streaming and errors, and exposes a clean endpoint to your frontend or mobile app. This is different from calling Claude directly from client-side code, which exposes your API key and gives you no control over rate limits, logging, or cost tracking.
This guide walks through a working setup: project structure, the actual API call, streaming responses to the client, error handling, and where to route requests if you want built-in usage tracking and per-application keys instead of managing raw provider credentials yourself.
Project Setup
You need three things: a Node.js backend (Express, Fastify, or plain http), an environment variable holding your API key, and an HTTP client. You can use fetch (built into Node 18+), axios, or an official SDK.
mkdir claude-backend && cd claude-backend
npm init -y
npm install express dotenv
claude-backend/
.env
server.js
routes/chat.js
Your .env file should never be committed to source control:
CLAUDE_API_KEY=your-key-here
PORT=3000
A Minimal Express Endpoint
Here's a basic non-streaming endpoint that forwards a message to Claude and returns the response:
// routes/chat.js
import express from "express";
const router = express.Router();
router.post("/chat", async (req, res) => {
const { message } = req.body;
if (!message) {
return res.status(400).json({ error: "message is required" });
}
try {
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.CLAUDE_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{ role: "user", content: message }],
}),
});
if (!response.ok) {
const errorBody = await response.text();
return res.status(response.status).json({ error: errorBody });
}
const data = await response.json();
res.json({ reply: data.content[0].text });
} catch (err) {
console.error("Claude API call failed:", err);
res.status(500).json({ error: "internal server error" });
}
});
export default router;
Wire it into your server:
// server.js
import express from "express";
import dotenv from "dotenv";
import chatRouter from "./routes/chat.js";
dotenv.config();
const app = express();
app.use(express.json());
app.use("/api", chatRouter);
app.listen(process.env.PORT, () => {
console.log(`Server running on port ${process.env.PORT}`);
});
This pattern keeps the API key server-side and gives you a controlled entry point for logging, rate limiting, and validation before requests reach Claude.
Streaming Responses to the Frontend
Chat UIs feel broken without streaming — users expect tokens to appear as they're generated. Node.js backends can proxy Server-Sent Events from Claude straight through to the browser:
router.post("/chat/stream", async (req, res) => {
const { message } = req.body;
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
const upstream = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.CLAUDE_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
stream: true,
messages: [{ role: "user", content: message }],
}),
});
const reader = upstream.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
res.write(decoder.decode(value));
}
res.end();
});
On the frontend, consume this with EventSource or a fetch-based reader. The key detail: don't buffer the entire response before sending — write chunks as they arrive to keep latency low.
Error Handling and Retries
Production backends need to handle rate limits (HTTP 429), overloaded errors (529), and transient network failures without crashing the request. A simple retry wrapper with exponential backoff covers most cases:
async function callClaudeWithRetry(payload, retries = 3) {
for (let attempt = 0; attempt < retries; attempt++) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.CLAUDE_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify(payload),
});
if (res.ok) return res.json();
if (res.status !== 429 && res.status !== 529) {
throw new Error(`Claude API error: ${res.status}`);
}
await new Promise((r) => setTimeout(r, 500 * 2 ** attempt));
}
throw new Error("Claude API retries exhausted");
}
Log every failed call with enough context (request ID, status code, payload size) to debug it later — you'll need this the first time a downstream consumer reports "the AI stopped responding."
Where SubToAPI Fits
The pattern above works, but it puts the raw provider API key in your backend's environment, mixes billing across every app that uses it, and gives you no per-application usage breakdown unless you build that tracking yourself.
SubToAPI sits between your Node.js backend and Claude: you generate scoped sub_live_... keys per application or environment, point your existing code at https://api.subtoapi.app/v1/messages instead of Anthropic's endpoint, and get streaming, tool use, and usage metadata without changing your integration logic. The request/response shape mirrors the Messages API, so the Express examples above work with only the base URL and key swapped — see the quickstart and messages docs for the exact payloads, and streaming docs if you're building the SSE proxy above. Plans start at €9/month with a free trial at signup; full details on pricing.
If you're running multiple services or a small team hitting Claude from different backends, per-key usage tracking and seat-based billing (Team at €19/seat, Scale at €49/seat) save you from building that dashboard yourself.
Production Checklist
Before shipping a Claude-backed endpoint:
- Validate and cap input length server-side, not just client-side
- Set a request timeout so hung connections don't pile up
- Log token usage per request if you're tracking cost per customer
- Return structured error responses, not raw provider error text, to your frontend
- Rate-limit your own
/chatendpoint independently of upstream limits
FAQs
Do I need the official Anthropic SDK, or can I use fetch directly? Either works. The SDK adds typed responses and built-in retry logic; raw fetch gives you full control over headers and streaming behavior with no extra dependency. Both call the same REST endpoint.
How do I keep my Claude API key safe in a Node.js app? Store it in environment variables, never in client-side code, and route all calls through your backend. Tools like SubToAPI let you issue separate scoped keys per application so a leaked key doesn't expose your entire account.
Can I stream Claude responses through Express to a React frontend? Yes — proxy the SSE stream from Claude's response body directly to your Express response using res.write() for each chunk, then consume it in React with EventSource or a ReadableStream reader on the fetch response.