Anthropic API Batch Processing Requests: A Guide
If you need to send the Anthropic API hundreds or thousands of prompts — classifying support tickets, summarizing documents, scoring content — sending them one request at a time is slow and expensive. Anthropic's answer to this is the Message Batches API, which lets you submit a large set of requests as a single job, get a 50% discount on token costs, and retrieve results asynchronously once processing completes.
This article covers how batch processing works on the Anthropic API, when it's the right tool versus synchronous calls or your own concurrency logic, and the practical details — limits, polling, error handling — that trip people up the first time they use it.
What "batch processing" means here
There are two different things people mean when they search for this:
- The Message Batches API — Anthropic's native batch endpoint. You upload a JSON array of requests, Anthropic processes them within 24 hours (usually much faster), and you poll for a results file.
- Client-side batching — writing your own loop or worker pool that fires many individual requests concurrently against the standard Messages endpoint, respecting rate limits.
Both are valid "batch processing" strategies, and which one you want depends on your latency requirements.
Using the Message Batches API
The flow looks like this:
- Build an array of individual message requests, each with a unique
custom_idso you can match results back to inputs. - POST the batch to the batches endpoint.
- Poll the batch status until it's
ended. - Download the results file and parse each line (it's JSONL — one JSON object per request).
A minimal example of the request shape:
{
"requests": [
{
"custom_id": "ticket-1001",
"params": {
"model": "claude-sonnet-4-5",
"max_tokens": 300,
"messages": [
{ "role": "user", "content": "Classify this support ticket: ..." }
]
}
},
{
"custom_id": "ticket-1002",
"params": {
"model": "claude-sonnet-4-5",
"max_tokens": 300,
"messages": [
{ "role": "user", "content": "Classify this support ticket: ..." }
]
}
}
]
}
The custom_id field matters more than it looks. Results come back in an unspecified order, so without it you have no reliable way to know which output belongs to which input. Always generate IDs that map directly to your own database keys.
What you get back
Each line in the results file contains the custom_id and either a successful message response or an error object. You need to handle both cases per line — a batch job finishing successfully doesn't mean every individual request inside it succeeded. Malformed prompts, content policy violations, or per-request token limit issues can still fail individually.
Limits and timing
Batches can contain tens of thousands of requests, and Anthropic processes them within a 24-hour window, though in practice most batches complete much sooner. This makes the Batches API a poor fit for anything user-facing — don't use it for a chat feature where someone is waiting on a response. It's built for backend jobs: nightly data enrichment, bulk classification, large-scale evaluation runs, dataset labeling.
When client-side concurrency is the better choice
If you need results in seconds rather than minutes or hours, batching to Anthropic's async endpoint isn't the right tool. Instead, you parallelize synchronous calls yourself, respecting your rate limits:
async function processBatch(items, concurrency = 5) {
const results = [];
const queue = [...items];
async function worker() {
while (queue.length) {
const item = queue.shift();
const response = await callModel(item);
results.push(response);
}
}
await Promise.all(Array.from({ length: concurrency }, worker));
return results;
}
This pattern gives you results as they complete, lets you retry individual failures immediately, and avoids waiting on a batch job that might take hours. The tradeoff is that you don't get the 50% batch discount, and you need to manage rate limits, backoff, and concurrency caps yourself.
Batching through SubToAPI
If you're already consuming Claude through SubToAPI, the same concurrency pattern above works directly against your sub_live_... key — you get usage metadata per call, so tracking cost and token volume across a large batch job is straightforward without building your own logging layer. For bulk jobs that don't need sub-second latency, running concurrent calls against the Messages endpoint with a worker pool like the one above is usually simpler to operate than managing a separate async batch lifecycle, especially if your job already lives inside a backend you control.
Streaming isn't useful for batch jobs (you want complete outputs to store, not partial tokens), so stick to standard request/response calls — see the quickstart for setup and the streaming docs if you also have an interactive part of your product that needs it.
Practical tips for any batching approach
- Cap your prompt size per item. Large batches amplify the cost of bloated prompts. Trim boilerplate system instructions down before scaling to thousands of calls.
- Log failures separately from successes. Partial failures are normal at scale — build your pipeline to retry just the failed
custom_ids rather than rerunning the whole batch. - Set conservative
max_tokens. Runaway output length across thousands of requests adds up fast in both cost and processing time. - Deduplicate inputs before submitting. It's common to accidentally batch the same record twice from a join or export step — check row counts before and after any transform.
- Store raw responses. Even if you only need a parsed field today, keeping the full response lets you re-extract data later without re-running the batch.
FAQ
Does the Anthropic Batches API support streaming responses?
No. Batch jobs are inherently asynchronous — you submit requests and retrieve completed results later, so there's no live token stream. If you need streaming, use the standard synchronous Messages endpoint instead.
How much cheaper is batch processing versus regular requests?
Anthropic's Message Batches API offers roughly 50% off standard token pricing in exchange for asynchronous, non-guaranteed-immediate processing (typically completing well within 24 hours).
Can I mix different models or system prompts in one batch?
Yes. Each request object in a batch carries its own full parameter set, including model, system prompt, and messages, so a single batch can contain requests destined for different models or configurations.