Claude API Mock Responses for Testing: A Practical Guide
Testing code that calls Claude without burning real API credits, hitting rate limits, or waiting on network latency means mocking the API responses. The core idea is simple: intercept the HTTP call your code makes to the Claude (or Claude-compatible) endpoint and return a canned JSON payload that matches the real response shape, instead of letting the request go out over the network.
This guide covers the practical ways to do that — static fixtures with HTTP interceptors, a local mock server, record-and-replay, and the trickier case of mocking streaming and tool-use responses — so your test suite stays fast, deterministic, and free of API costs.
Why mock Claude responses at all
Calling a real LLM endpoint in your test suite has three problems:
- Cost and quota — every test run spends tokens, and CI runs multiply that fast.
- Non-determinism — model output varies between calls, which breaks assertions that expect exact text.
- Flakiness — network errors, rate limits, and latency spikes make tests unreliable, especially in parallel CI jobs.
Mocking solves all three: no network call, no tokens spent, and a response you fully control. The tradeoff is that mocks can drift from the real API shape over time, so the fixtures need to stay accurate.
Static JSON fixtures with HTTP interceptors
The most common pattern in Node.js is to intercept outgoing HTTP requests with a library like nock or msw (Mock Service Worker) and return a fixture file that mirrors the real response structure.
// fixtures/claude-response.json
{
"id": "msg_01abc123",
"role": "assistant",
"content": [{ "type": "text", "text": "Hello, how can I help?" }],
"model": "claude-3-5-sonnet",
"stop_reason": "end_turn",
"usage": { "input_tokens": 12, "output_tokens": 8 }
}
import nock from "nock";
import fixture from "./fixtures/claude-response.json";
test("handles a successful completion", async () => {
nock("https://api.subtoapi.app")
.post("/v1/messages")
.reply(200, fixture);
const result = await callClaude("Hi there");
expect(result.content[0].text).toBe("Hello, how can I help?");
});
The key is keeping fixtures close to the actual response schema. If you're integrating against SubToAPI, the Messages reference documents the exact field names and types, so your fixtures won't silently drift from what the real API returns.
Running a local mock server
For integration tests or testing multiple services against the same API, a standalone mock server is often cleaner than per-test interceptors. A tiny Express server works well:
import express from "express";
const app = express();
app.use(express.json());
app.post("/v1/messages", (req, res) => {
res.json({
id: "msg_test_001",
role: "assistant",
content: [{ type: "text", text: "Mock response for: " + req.body.messages?.[0]?.content }],
model: req.body.model,
stop_reason: "end_turn",
usage: { input_tokens: 10, output_tokens: 15 }
});
});
app.listen(4010, () => console.log("Mock Claude API on :4010"));
Point your app's base URL at http://localhost:4010 in your test environment config instead of the real endpoint. This approach is useful when several services or languages need to hit the same mock, since it's not tied to a single test runner's interceptor library.
Record and replay real responses
A hybrid approach: make a handful of real calls once, save the exact responses as fixtures, then replay them in subsequent test runs. Tools like nock's record mode or Python's vcrpy automate this.
nock.recorder.rec({
output_objects: true,
dont_print: true,
});
// make one real call here, then save nock.recorder.play() output to a file
This keeps fixtures honest because they came from an actual API response rather than a hand-written guess. The downside is you need to re-record periodically if the response schema changes — worth doing whenever you bump API versions or add new fields you depend on.
Mocking streaming responses
Streaming is the part most people get wrong when mocking, because it's not a single JSON blob — it's a sequence of server-sent events. A mock streaming endpoint needs to emit the same event sequence a real client would parse.
app.post("/v1/messages", (req, res) => {
res.setHeader("Content-Type", "text/event-stream");
const chunks = ["Hello", ", ", "world", "!"];
chunks.forEach((text, i) => {
res.write(`event: content_block_delta\n`);
res.write(`data: ${JSON.stringify({ type: "text_delta", text })}\n\n`);
});
res.write(`event: message_stop\ndata: {}\n\n`);
res.end();
});
If your application streams responses in production, test the streaming code path explicitly — a mock that only returns non-streaming JSON won't catch bugs in your chunk-parsing logic. SubToAPI's streaming docs show the exact event types and ordering to replicate if you're mocking against that API shape.
Mocking tool use
If your app relies on tool calls, your mocks need to return tool_use content blocks with realistic input payloads so your tool-execution code gets exercised in tests, not skipped.
{
"id": "msg_tool_001",
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_01xyz",
"name": "get_weather",
"input": { "location": "Berlin" }
}
],
"stop_reason": "tool_use"
}
Write a separate test that sends this fixture, confirms your code calls the right handler with the right arguments, then sends a follow-up fixture simulating the tool_result round trip. The tool use docs are a good reference for the exact block structure to mock.
When you actually want real responses
Mocks are for unit and CI tests, but you still need occasional runs against a real model to catch schema or behavior drift that a stale fixture would miss. For that, keep a small staging key with a low-cost plan — SubToAPI's Solo plan at €9/month is enough for a staging environment that exercises the real endpoint without needing a full team seat setup. Start with the quickstart or create an account at /signup if you need a second key dedicated to integration tests versus mocked unit tests.
questions
Do mocked Claude responses need to match the exact API schema? Yes. If your fixtures use field names or types that differ from the real response, your tests will pass while your production code breaks. Cross-check fixtures against the API reference periodically, especially after upgrading SDK or API versions.
Should I mock streaming responses the same way as regular ones? No — streaming needs a sequence of server-sent events in the correct order, not a single JSON object. Test your SSE parsing logic against a mock that emits real event chunks, not just a static final response.
Is record-and-replay better than hand-written fixtures? It's more accurate since the fixture comes from a real API call, but it requires periodic re-recording to stay in sync with API changes. Hand-written fixtures are easier to edit for specific edge cases like errors or empty responses.