← Blog

Claude API Unit Testing with Mock Responses

2026-10-05 · 4 min read · SubToAPI Team

Unit testing code that calls the Claude API means never actually calling the Claude API during your test run. Instead, you intercept the HTTP request (or the SDK call) and return a fixed, predictable response that mimics what Claude would send back. This makes your tests fast, free, deterministic, and runnable without network access or valid credentials.

The core problem with testing against a live model is that responses vary between runs, calls cost money, latency slows down CI, and rate limits or outages can make your test suite flaky for reasons that have nothing to do with your code. Mocking solves all of that by replacing the model call with a controlled stand-in while you test everything around it: your prompt construction, your error handling, your retry logic, and how your application parses the response.

What you're actually testing

When you mock the Claude API, you're not testing whether Claude gives good answers — that's not a unit test's job. You're testing:

Each of these needs a different mock, not just one "happy path" fixture.

Mocking at the HTTP layer

The most robust approach is mocking at the HTTP boundary rather than mocking your own wrapper functions. This catches real bugs in how you build requests and parse responses, because the mock only controls the network call, not your application logic.

In JavaScript with nock:

const nock = require("nock");

test("parses a successful Claude response", async () => {
  nock("https://api.anthropic.com")
    .post("/v1/messages")
    .reply(200, {
      id: "msg_01abc",
      type: "message",
      role: "assistant",
      content: [{ type: "text", text: "Mocked reply" }],
      stop_reason: "end_turn",
      usage: { input_tokens: 12, output_tokens: 5 },
    });

  const result = await callClaude("What is 2+2?");
  expect(result).toBe("Mocked reply");
});

In Python with respx or responses:

import responses

@responses.activate
def test_parses_successful_response():
    responses.add(
        responses.POST,
        "https://api.anthropic.com/v1/messages",
        json={
            "id": "msg_01abc",
            "content": [{"type": "text", "text": "Mocked reply"}],
            "stop_reason": "end_turn",
            "usage": {"input_tokens": 12, "output_tokens": 5},
        },
        status=200,
    )
    result = call_claude("What is 2+2?")
    assert result == "Mocked reply"

If you're using SubToAPI instead of calling the model provider directly, the same pattern applies — just point your interceptor at https://api.subtoapi.app/v1/messages. The request and response shapes documented at /docs/messages are stable, so your fixtures won't need to change as your app evolves.

Building a fixture library

Rather than inlining JSON in every test, keep a small set of reusable fixtures:

// fixtures/claude.js
module.exports = {
  simpleText: {
    content: [{ type: "text", text: "Hello from Claude" }],
    stop_reason: "end_turn",
  },
  toolUse: {
    content: [
      {
        type: "tool_use",
        id: "toolu_01xyz",
        name: "get_weather",
        input: { city: "Berlin" },
      },
    ],
    stop_reason: "tool_use",
  },
  rateLimited: { status: 429, body: { error: { message: "Rate limited" } } },
  serverError: { status: 500, body: { error: { message: "Internal error" } } },
};

This lets you test tool-calling logic without ever invoking a real tool, and test your retry/backoff code against a 429 without waiting for an actual rate limit. If you use /docs/tools for tool definitions, build fixtures that mirror the exact tool_use block shape so your parsing code is exercised honestly.

Mocking streaming responses

Streaming is the trickiest case because you're mocking a sequence of server-sent events, not a single JSON body. You need chunks that match the real event stream shape (message_start, content_block_delta, message_stop), and your test should assert that your consumer code concatenates deltas correctly and stops on the terminal event.

const sseBody = [
  'event: content_block_delta\ndata: {"delta":{"text":"Hel"}}\n\n',
  'event: content_block_delta\ndata: {"delta":{"text":"lo"}}\n\n',
  "event: message_stop\ndata: {}\n\n",
].join("");

nock("https://api.subtoapi.app")
  .post("/v1/messages")
  .reply(200, sseBody, { "Content-Type": "text/event-stream" });

See /docs/streaming for the exact event types your consumer needs to handle, so fixtures stay accurate rather than guessed.

Avoiding brittle mocks

A few practices keep mock-based tests useful instead of becoming maintenance burden:

If you're switching between calling Anthropic directly and routing through SubToAPI, keep your HTTP client abstraction thin enough that swapping the base URL and auth header is the only change — that's the whole point of SubToAPI's drop-in sub_live_... key model described on /pricing and the /docs/quickstart guide.

FAQ

Do I need real API credentials to run mocked tests? No. Since the mock intercepts the HTTP call before it leaves your process, you can use a placeholder string as your API key in test environments — no real key or network access required.

Should I mock the SDK or the HTTP layer? Prefer the HTTP layer when possible. Mocking the SDK method directly can hide bugs in how you construct requests, since you skip the part of your code that builds the actual payload.

How do I test rate limit and retry logic without hitting real limits? Create a fixture that returns a 429 status with the same error shape Claude's API sends, then assert your retry/backoff code behaves correctly — waits, retries, and eventually succeeds or fails gracefully.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →