Claude API Mock Server for Development: Setup Guide
If you're building against the Claude API, you don't want every test run, CI pipeline, or local dev session hitting the real endpoint. A Claude API mock server lets you simulate requests and responses locally — no API key charges, no rate limits, no network flakiness — while your code exercises the exact same request/response shapes it will use in production.
This matters for three practical reasons: cost (every mocked call is a call you don't pay for), speed (mock responses return in milliseconds, not seconds), and determinism (tests that depend on a live LLM's output are inherently flaky — a mock gives you the same answer every time). Below is how to set one up, what to mock, and when a mock isn't the right tool.
What you actually need to mock
The Claude API has a small number of request/response shapes worth simulating:
- Non-streaming messages — a single JSON response with
content,usage,stop_reason. - Streaming (SSE) messages — a sequence of
message_start,content_block_delta,message_delta,message_stopevents. - Tool use — responses containing
tool_useblocks that your code needs to parse and act on. - Error responses — 429 rate limits, 400 validation errors, 529 overloaded — so your retry logic is actually tested.
A good mock server covers all four. Most teams only mock the happy path and then discover their error handling is broken in production.
Option 1: a local Express/Fastify mock
The fastest way to get a mock running is a tiny HTTP server that returns canned JSON for POST /v1/messages.
// mock-server.js
import express from "express";
const app = express();
app.use(express.json());
app.post("/v1/messages", (req, res) => {
const { stream, messages } = req.body;
if (stream) {
res.setHeader("Content-Type", "text/event-stream");
res.write(`event: message_start\ndata: {"type":"message_start"}\n\n`);
res.write(`event: content_block_delta\ndata: {"delta":{"text":"Hello"}}\n\n`);
res.write(`event: message_stop\ndata: {"type":"message_stop"}\n\n`);
return res.end();
}
res.json({
id: "msg_mock_001",
type: "message",
role: "assistant",
content: [{ type: "text", text: `Echo: ${messages[0]?.content}` }],
stop_reason: "end_turn",
usage: { input_tokens: 12, output_tokens: 8 },
});
});
app.listen(4010, () => console.log("Mock Claude API on :4010"));
Point your app's baseURL at http://localhost:4010 during tests, and at the real endpoint in staging/production. Keep the base URL behind an environment variable — this is the single most important architectural decision for making mocking painless.
export CLAUDE_API_BASE_URL=http://localhost:4010
Option 2: record-and-replay fixtures
Instead of hand-writing every response, record real traffic once and replay it in tests. Tools like nock (Node) or VCR-style libraries (Python, Ruby) intercept outgoing HTTP calls and serve recorded fixtures instead.
import nock from "nock";
nock("https://api.subtoapi.app")
.post("/v1/messages")
.reply(200, {
id: "msg_fixture_1",
content: [{ type: "text", text: "Mocked reply" }],
usage: { input_tokens: 10, output_tokens: 5 },
});
This approach is good for integration tests where you want realistic payload shapes without maintaining them by hand. The downside: fixtures go stale if the API response shape changes and nobody updates them.
Option 3: contract-based mocking with OpenAPI/JSON Schema
If you want stricter guarantees, generate mock responses from a schema instead of static fixtures. Tools like Prism (Stoplight) can serve mock responses validated against an OpenAPI spec, which catches cases where your mock drifts from the real contract.
This is overkill for a solo project but worth it on a team where multiple services depend on the same mocked contract.
Testing streaming and tool use specifically
Streaming responses are the part teams most often skip mocking, and it's exactly where bugs hide — partial JSON, chunk boundaries that split a token mid-delta, connection drops mid-stream. Write at least one test that feeds your SSE parser a response split across multiple write() calls, not just one clean chunk.
For tool use, mock a response where stop_reason is tool_use and verify your code correctly extracts the tool name and input, executes it, and sends a follow-up request with the tool result appended to messages. This round-trip is where most tool-use bugs live, and it's fully testable without ever calling a real model. See /docs/tools for the request/response shapes if you're integrating this against SubToAPI.
Where mocking stops being useful
A mock server is for testing your integration code, not for evaluating model quality. If you need to know whether a prompt actually produces good output, you have to call the real API. Common mistakes:
- Relying on mocks to validate prompt engineering — mocks can't tell you if your prompt is good.
- Never running against the real API before shipping — contract drift between your mock and the live API will bite you eventually.
- Mocking at too high a level (mocking your own wrapper function instead of the HTTP boundary) — this hides bugs in your actual request serialization.
A practical middle ground: use mocks for unit tests and CI, and run a small smoke-test suite against the real API (behind a feature flag or nightly job) to catch contract drift. If you're using SubToAPI as your Claude API layer, you can run that smoke suite against a Solo-tier key (sub_live_...) cheaply — see /pricing — and keep the bulk of your test suite mocked. The /docs/quickstart guide shows the exact request format your mocks should match.
A minimal checklist for your mock server
- [ ] Non-streaming success response with realistic
usagenumbers - [ ] Streaming response split across multiple chunks
- [ ]
tool_usestop reason with a follow-up request - [ ] 429 and 529 error responses to test retry/backoff logic
- [ ] Environment variable to swap base URL between mock and real
Get this working once and your entire test suite — unit tests, CI pipelines, local dev — runs without touching a real API key or burning tokens.
Questions
Does mocking the Claude API save money on API costs? Yes — every request served by a mock is a request that never reaches the real API, so you pay nothing for it. This is most valuable in CI, where the same test suite might run dozens of times a day.
Can I use the same mock for streaming and non-streaming requests? Yes, as long as your mock checks the stream parameter in the request body and branches to return either a single JSON response or a sequence of SSE events accordingly.
Should I mock the API in production code paths? No. Mocks belong in tests and local development only. Production traffic should always hit the real endpoint — see /docs/messages for the live request format once you're ready to go beyond mocks.