← Blog

Claude API Mock Server for Local Dev: Setup Guide

2026-10-05 · 5 min read · SubToAPI Team

If you're building against the Claude API, you don't want every local test run or CI job hitting the real API. A mock server that mimics Claude's request/response shape lets you develop offline, avoid burning through credits, and reliably reproduce edge cases like rate limits, timeouts, or malformed responses. This article shows you how to build one quickly, and when to switch to a real API instead.

A Claude API mock server is just an HTTP server that implements the same endpoint shape as /v1/messages (or whatever client you're using) and returns canned or programmatically generated responses. You run it locally, point your app's base_url at it instead of the real Claude endpoint, and your integration code never knows the difference. The two main things you need to get right are matching the response schema exactly and supporting streaming if your app uses it.

Why mock instead of using a real sandbox

Testing against the live Claude API during development has real costs:

Mocking isn't a replacement for integration testing against the real thing before you ship. It's a layer that sits underneath that, catching logic bugs in your request building, response parsing, and error handling without touching the network.

Building a minimal mock server

Here's a small Express-based mock that implements the shape of a Messages-style endpoint:

const express = require('express');
const app = express();
app.use(express.json());

app.post('/v1/messages', (req, res) => {
  const { model, messages, stream } = req.body;

  if (stream) {
    res.setHeader('Content-Type', 'text/event-stream');
    const chunks = ['Hello', ' from', ' the', ' mock', ' server.'];
    chunks.forEach((text, i) => {
      res.write(`data: ${JSON.stringify({ type: 'content_block_delta', delta: { text } })}\n\n`);
    });
    res.write('data: [DONE]\n\n');
    return res.end();
  }

  res.json({
    id: 'msg_mock_001',
    model,
    role: 'assistant',
    content: [{ type: 'text', text: 'This is a mocked response.' }],
    usage: { input_tokens: 12, output_tokens: 8 },
    stop_reason: 'end_turn'
  });
});

app.listen(4000, () => console.log('Mock Claude API running on :4000'));

Point your app's client config at http://localhost:4000/v1 instead of the real API host, and everything downstream works unchanged — assuming your client only cares about the HTTP contract and not a specific SDK's internal transport.

Simulating failure modes

The real value of a mock server isn't replaying happy-path responses — it's forcing your error handling to run. Add routes or request-matching logic for:

A simple way to do this deterministically is to key behavior off a header your tests control:

app.post('/v1/messages', (req, res) => {
  const scenario = req.headers['x-mock-scenario'];

  if (scenario === 'rate_limit') {
    return res.status(429).set('retry-after', '2').json({ error: { type: 'rate_limit_error' } });
  }
  if (scenario === 'overloaded') {
    return res.status(529).json({ error: { type: 'overloaded_error' } });
  }
  // ...normal response
});

This lets each test case request a specific failure mode without randomness, so your test suite stays stable and repeatable.

Keeping the mock in sync with reality

The biggest risk with a hand-rolled mock is schema drift — your mock stops matching what the real API actually returns, and your tests pass locally while production breaks. A few practices help:

  1. Record real responses once. Capture a handful of actual API responses (including streaming chunks) and use them as your mock's fixture data instead of writing JSON by hand.
  2. Validate against a shared contract. If you have a JSON schema or TypeScript type for the response shape, run it against both the mock and real responses in CI.
  3. Re-sync periodically. Treat the mock fixtures like any other dependency — review them whenever you upgrade SDK versions or change models.

When to switch to a real API

Mocking is for local dev and unit/integration tests where you control the input and expected output. It's the wrong tool for:

If you want a lower-friction path to a working Claude integration without juggling separate dev and prod credentials, SubToAPI turns your existing Claude access into a standard HTTPS API with its own key (sub_live_...), streaming support, and usage metadata — so your mock server and your real backend speak the exact same request/response shape. You can build against the mock locally, then flip to the real endpoint by changing a base URL and key, following the same quickstart. Plans start at €9/month on the pricing page, with a free trial at signup.

Questions

Do I need a mock server if I already have a Claude sandbox or test account? A sandbox still makes real network calls and can still rate-limit or cost money depending on your plan. A local mock is faster, free, and lets you simulate failures a sandbox won't reliably produce on demand.

Can a mock server test streaming responses accurately? Yes, as long as it emits Server-Sent Events in the same chunk format your client expects. See streaming for the real response shape to match if you're testing against an API like SubToAPI's.

Should unit tests use the mock or should I stub the client library instead? Both work. Stubbing the client is faster and simpler for pure unit tests; a mock HTTP server is better for integration tests that need to exercise your actual networking, retry, and parsing code end to end.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →