Claude API Mock Server for Local Dev: Setup Guide
If you're building against the Claude API, you don't want every local test run or CI job hitting the real API. A mock server that mimics Claude's request/response shape lets you develop offline, avoid burning through credits, and reliably reproduce edge cases like rate limits, timeouts, or malformed responses. This article shows you how to build one quickly, and when to switch to a real API instead.
A Claude API mock server is just an HTTP server that implements the same endpoint shape as /v1/messages (or whatever client you're using) and returns canned or programmatically generated responses. You run it locally, point your app's base_url at it instead of the real Claude endpoint, and your integration code never knows the difference. The two main things you need to get right are matching the response schema exactly and supporting streaming if your app uses it.
Why mock instead of using a real sandbox
Testing against the live Claude API during development has real costs:
- Money — every test run, every CI pipeline execution, every refresh in dev consumes tokens.
- Speed — network round trips to a hosted LLM are slow compared to an in-memory mock, which slows down your test suite.
- Determinism — LLM outputs vary. If your tests assert on exact response content, a live model will eventually break them.
- Offline work — no internet, no API key configured yet, CI runners with no outbound access — none of these block you if you're mocking.
Mocking isn't a replacement for integration testing against the real thing before you ship. It's a layer that sits underneath that, catching logic bugs in your request building, response parsing, and error handling without touching the network.
Building a minimal mock server
Here's a small Express-based mock that implements the shape of a Messages-style endpoint:
const express = require('express');
const app = express();
app.use(express.json());
app.post('/v1/messages', (req, res) => {
const { model, messages, stream } = req.body;
if (stream) {
res.setHeader('Content-Type', 'text/event-stream');
const chunks = ['Hello', ' from', ' the', ' mock', ' server.'];
chunks.forEach((text, i) => {
res.write(`data: ${JSON.stringify({ type: 'content_block_delta', delta: { text } })}\n\n`);
});
res.write('data: [DONE]\n\n');
return res.end();
}
res.json({
id: 'msg_mock_001',
model,
role: 'assistant',
content: [{ type: 'text', text: 'This is a mocked response.' }],
usage: { input_tokens: 12, output_tokens: 8 },
stop_reason: 'end_turn'
});
});
app.listen(4000, () => console.log('Mock Claude API running on :4000'));
Point your app's client config at http://localhost:4000/v1 instead of the real API host, and everything downstream works unchanged — assuming your client only cares about the HTTP contract and not a specific SDK's internal transport.
Simulating failure modes
The real value of a mock server isn't replaying happy-path responses — it's forcing your error handling to run. Add routes or request-matching logic for:
- Rate limits — return HTTP 429 with a
retry-afterheader on every third request to test your backoff logic. - Overload errors — return a 5xx status to confirm your retry/circuit-breaker code actually fires.
- Truncated streams — close the connection mid-stream to check your client doesn't hang waiting for a
[DONE]that never arrives. - Malformed JSON — send back invalid payloads occasionally to verify your parser fails gracefully instead of crashing the process.
A simple way to do this deterministically is to key behavior off a header your tests control:
app.post('/v1/messages', (req, res) => {
const scenario = req.headers['x-mock-scenario'];
if (scenario === 'rate_limit') {
return res.status(429).set('retry-after', '2').json({ error: { type: 'rate_limit_error' } });
}
if (scenario === 'overloaded') {
return res.status(529).json({ error: { type: 'overloaded_error' } });
}
// ...normal response
});
This lets each test case request a specific failure mode without randomness, so your test suite stays stable and repeatable.
Keeping the mock in sync with reality
The biggest risk with a hand-rolled mock is schema drift — your mock stops matching what the real API actually returns, and your tests pass locally while production breaks. A few practices help:
- Record real responses once. Capture a handful of actual API responses (including streaming chunks) and use them as your mock's fixture data instead of writing JSON by hand.
- Validate against a shared contract. If you have a JSON schema or TypeScript type for the response shape, run it against both the mock and real responses in CI.
- Re-sync periodically. Treat the mock fixtures like any other dependency — review them whenever you upgrade SDK versions or change models.
When to switch to a real API
Mocking is for local dev and unit/integration tests where you control the input and expected output. It's the wrong tool for:
- Testing actual model quality or prompt behavior — a mock can't tell you if your prompt produces good answers.
- Load testing — synthetic mock latency doesn't reflect real network and inference time.
- Final pre-release verification — always run your test suite against the real API at least once before shipping.
If you want a lower-friction path to a working Claude integration without juggling separate dev and prod credentials, SubToAPI turns your existing Claude access into a standard HTTPS API with its own key (sub_live_...), streaming support, and usage metadata — so your mock server and your real backend speak the exact same request/response shape. You can build against the mock locally, then flip to the real endpoint by changing a base URL and key, following the same quickstart. Plans start at €9/month on the pricing page, with a free trial at signup.
Questions
Do I need a mock server if I already have a Claude sandbox or test account? A sandbox still makes real network calls and can still rate-limit or cost money depending on your plan. A local mock is faster, free, and lets you simulate failures a sandbox won't reliably produce on demand.
Can a mock server test streaming responses accurately? Yes, as long as it emits Server-Sent Events in the same chunk format your client expects. See streaming for the real response shape to match if you're testing against an API like SubToAPI's.
Should unit tests use the mock or should I stub the client library instead? Both work. Stubbing the client is faster and simpler for pure unit tests; a mock HTTP server is better for integration tests that need to exercise your actual networking, retry, and parsing code end to end.