← Blog

Claude API Test Case Generation: A Developer Guide

2026-10-09 · 5 min read · SubToAPI Team

Can Claude actually write good test cases?

Yes — Claude is well suited to test case generation because it can read a function or spec, reason about edge cases, and produce structured output (test code, JSON test data, or Gherkin scenarios) in one pass. The practical question isn't whether the Claude API can generate tests, it's how to prompt it so the output is directly usable in your test suite instead of needing heavy rewriting.

This guide covers prompt patterns for generating unit tests, edge-case tables, and mock data with the Claude API, how to keep the output in a format your test runner can consume, and how to wire this into a CI pipeline so test generation isn't a one-off manual step.

Why use an LLM for test generation

Writing exhaustive test cases by hand is repetitive and easy to under-cover. Developers typically test the happy path and one or two obvious failures, then move on. Claude is good at the part humans skip: enumerating boundary conditions, null/empty inputs, type mismatches, concurrency edge cases, and negative-path assertions — especially when you give it the function signature and a short description of intent.

The best use cases are:

It's not a replacement for integration or end-to-end testing strategy — it's a generator you review, not an oracle you trust blindly.

Prompting Claude for test case generation

The quality of generated tests depends almost entirely on how much context you give. A bare "write tests for this function" prompt produces generic, low-value output. A good prompt includes the code, the testing framework, and explicit instructions on coverage expectations.

You are generating unit tests for the function below using Jest.

Requirements:
- Cover the happy path
- Cover empty/null/undefined inputs
- Cover boundary values (0, negative numbers, max int)
- Cover type errors
- Use describe/it blocks
- Do not add explanatory comments outside the code

Function:
function calculateDiscount(price, percent) {
  if (percent < 0 || percent > 100) throw new Error("Invalid percent");
  return price - (price * percent) / 100;
}

This kind of prompt reliably produces a complete, runnable test file rather than a few scattered examples.

Calling the Claude API directly

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Generate Jest unit tests for this function, covering edge cases and boundary values: function calculateDiscount(price, percent) { if (percent < 0 || percent > 100) throw new Error(\"Invalid percent\"); return price - (price * percent) / 100; }"}
    ]
  }'

Asking for structured edge-case output

For planning before you write tests, it's often more useful to get a table of cases rather than code. Ask for JSON directly so you can feed it into a test generator script or a spreadsheet:

List edge cases for this function as a JSON array of objects with fields:
"input", "expected_behavior", "category" (boundary, type_error, null, concurrency).
Do not include any text outside the JSON array.

Claude will reliably return parseable JSON if you're explicit about the schema and tell it not to wrap the output in prose or markdown fences.

Generating tests through SubToAPI

If your team already routes Claude access through SubToAPI, the same prompts work unchanged — you just call your own HTTPS endpoint with a sub_live_... key instead of managing Anthropic credentials per developer. This matters for test generation workflows because they're often run from CI jobs, internal scripts, or shared tooling where you don't want raw provider keys floating around.

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 1024,
    messages: [
      { role: "user", content: `Generate pytest unit tests for this function, including edge cases:\n\n${functionSource}` }
    ]
  })
});

const data = await res.json();
console.log(data.content[0].text);

Because usage is tracked per application key, you can see exactly how much a test-generation job costs across a CI run or a sprint, which is useful if test generation becomes a recurring step rather than a one-off. See the quickstart and messages docs for request formats.

Wiring it into CI

A common pattern:

  1. On pull request, diff the changed functions.
  2. Send each changed function (plus surrounding context) to Claude with a test-generation prompt.
  3. Write the output to a __generated_tests__/ directory, never overwriting hand-written tests.
  4. Run the full suite; fail the build only if generated tests don't compile or run (not if they fail — a failing generated test might be flagging a real bug).
  5. Require human review before merging generated tests into the permanent suite.

Keep generated and hand-written tests in separate files. This makes it obvious what's AI-suggested versus reviewed, and avoids silently overwriting tests a developer wrote intentionally.

Streaming for large test suites

If you're generating tests for an entire module or several files at once, the response can be long enough that streaming improves perceived latency and lets you start writing output to disk incrementally. See streaming for implementation details, and tools if you want Claude to call a function that validates generated code against a linter before returning it.

Review checklist before trusting generated tests

Generated tests should pass a code review like any other PR — they're a draft, not a guarantee.

questions

Does Claude write tests in any framework? Yes, as long as you specify it. Claude handles Jest, pytest, JUnit, RSpec, and most mainstream frameworks well when the prompt names the framework and shows the function signature.

Can Claude generate test data, not just test code? Yes. Ask for structured JSON or CSV test data with explicit fields and valid/invalid examples; this works well for API payload testing and fuzzing inputs.

Is it safe to merge Claude-generated tests without review? No. Treat them as a draft — review assertions, check for invented APIs or mocks, and confirm edge cases match actual business logic before merging into your permanent suite.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →