Claude API Unit Testing with Mock Responses
Unit testing code that calls the Claude API means never actually calling the Claude API during your test run. Instead, you intercept the HTTP request (or the SDK call) and return a fixed, predictable response that mimics what Claude would send back. This makes your tests fast, free, deterministic, and runnable without network access or valid credentials.
The core problem with testing against a live model is that responses vary between runs, calls cost money, latency slows down CI, and rate limits or outages can make your test suite flaky for reasons that have nothing to do with your code. Mocking solves all of that by replacing the model call with a controlled stand-in while you test everything around it: your prompt construction, your error handling, your retry logic, and how your application parses the response.
What you're actually testing
When you mock the Claude API, you're not testing whether Claude gives good answers — that's not a unit test's job. You're testing:
- That your code sends the right request shape (model, messages, system prompt, tools, max_tokens)
- That your code correctly parses a successful response into whatever internal format your app uses
- That your code handles errors (429, 500, malformed JSON, timeouts) without crashing
- That streaming consumers correctly reassemble chunks into full text
- That tool-use responses trigger the right downstream function calls
Each of these needs a different mock, not just one "happy path" fixture.
Mocking at the HTTP layer
The most robust approach is mocking at the HTTP boundary rather than mocking your own wrapper functions. This catches real bugs in how you build requests and parse responses, because the mock only controls the network call, not your application logic.
In JavaScript with nock:
const nock = require("nock");
test("parses a successful Claude response", async () => {
nock("https://api.anthropic.com")
.post("/v1/messages")
.reply(200, {
id: "msg_01abc",
type: "message",
role: "assistant",
content: [{ type: "text", text: "Mocked reply" }],
stop_reason: "end_turn",
usage: { input_tokens: 12, output_tokens: 5 },
});
const result = await callClaude("What is 2+2?");
expect(result).toBe("Mocked reply");
});
In Python with respx or responses:
import responses
@responses.activate
def test_parses_successful_response():
responses.add(
responses.POST,
"https://api.anthropic.com/v1/messages",
json={
"id": "msg_01abc",
"content": [{"type": "text", "text": "Mocked reply"}],
"stop_reason": "end_turn",
"usage": {"input_tokens": 12, "output_tokens": 5},
},
status=200,
)
result = call_claude("What is 2+2?")
assert result == "Mocked reply"
If you're using SubToAPI instead of calling the model provider directly, the same pattern applies — just point your interceptor at https://api.subtoapi.app/v1/messages. The request and response shapes documented at /docs/messages are stable, so your fixtures won't need to change as your app evolves.
Building a fixture library
Rather than inlining JSON in every test, keep a small set of reusable fixtures:
// fixtures/claude.js
module.exports = {
simpleText: {
content: [{ type: "text", text: "Hello from Claude" }],
stop_reason: "end_turn",
},
toolUse: {
content: [
{
type: "tool_use",
id: "toolu_01xyz",
name: "get_weather",
input: { city: "Berlin" },
},
],
stop_reason: "tool_use",
},
rateLimited: { status: 429, body: { error: { message: "Rate limited" } } },
serverError: { status: 500, body: { error: { message: "Internal error" } } },
};
This lets you test tool-calling logic without ever invoking a real tool, and test your retry/backoff code against a 429 without waiting for an actual rate limit. If you use /docs/tools for tool definitions, build fixtures that mirror the exact tool_use block shape so your parsing code is exercised honestly.
Mocking streaming responses
Streaming is the trickiest case because you're mocking a sequence of server-sent events, not a single JSON body. You need chunks that match the real event stream shape (message_start, content_block_delta, message_stop), and your test should assert that your consumer code concatenates deltas correctly and stops on the terminal event.
const sseBody = [
'event: content_block_delta\ndata: {"delta":{"text":"Hel"}}\n\n',
'event: content_block_delta\ndata: {"delta":{"text":"lo"}}\n\n',
"event: message_stop\ndata: {}\n\n",
].join("");
nock("https://api.subtoapi.app")
.post("/v1/messages")
.reply(200, sseBody, { "Content-Type": "text/event-stream" });
See /docs/streaming for the exact event types your consumer needs to handle, so fixtures stay accurate rather than guessed.
Avoiding brittle mocks
A few practices keep mock-based tests useful instead of becoming maintenance burden:
- Assert on request bodies too. Don't just mock the response — verify the request your code sent included the right model, system prompt, and parameters.
- Keep fixtures versioned alongside the API shape. If a response field changes, update one fixture file, not forty inline mocks.
- Mix in integration tests sparingly. A small number of real calls against a trial key (see /signup) running outside your main CI loop catches drift that mocks can't.
- Test error paths explicitly. Overflowing context, invalid tool schemas, and malformed streaming chunks are common production bugs that only show up if you deliberately mock them.
If you're switching between calling Anthropic directly and routing through SubToAPI, keep your HTTP client abstraction thin enough that swapping the base URL and auth header is the only change — that's the whole point of SubToAPI's drop-in sub_live_... key model described on /pricing and the /docs/quickstart guide.
FAQ
Do I need real API credentials to run mocked tests? No. Since the mock intercepts the HTTP call before it leaves your process, you can use a placeholder string as your API key in test environments — no real key or network access required.
Should I mock the SDK or the HTTP layer? Prefer the HTTP layer when possible. Mocking the SDK method directly can hide bugs in how you construct requests, since you skip the part of your code that builds the actual payload.
How do I test rate limit and retry logic without hitting real limits? Create a fixture that returns a 429 status with the same error shape Claude's API sends, then assert your retry/backoff code behaves correctly — waits, retries, and eventually succeeds or fails gracefully.