Claude API Batch Processing Requests: Example Guide
What "batch processing" actually means for the Claude API
When developers search for Claude API batch processing, they usually mean one of two different things. First, Anthropic's native Message Batches API, which accepts up to 100,000 requests in a single job, processes them asynchronously, and returns results within a 24-hour window at a reduced cost. This is built for large, non-urgent workloads like classifying a million support tickets overnight.
Second — and far more common in practice — is the pattern of sending a large number of individual requests to the standard messages endpoint, but doing it programmatically with controlled concurrency, retries, and rate-limit handling instead of firing them all at once or looping one-by-one. This is what most people actually need: summarizing 500 documents, generating descriptions for a product catalog, or re-scoring a dataset, and getting results back within minutes rather than waiting on an async job. The rest of this article focuses on that second pattern, since it's the one you'll build yourself most often.
Native batch API vs. concurrent request loops
Use Anthropic's Message Batches API when:
- The job is large (thousands to tens of thousands of prompts)
- You can tolerate results arriving up to 24 hours later
- Cost matters more than latency
Use a concurrent request loop when:
- You need results within seconds or minutes
- The job is a few dozen to a few thousand requests
- You want to stream or react to partial results
- You're calling Claude from inside an existing app or pipeline, not a standalone offline job
Most "batch processing" scripts developers write day-to-day fall into the second category, so that's what the examples below cover.
Example: batch processing with a Node.js concurrency pool
The naive approach — a for loop with await inside it — sends requests one at a time and wastes most of your rate limit headroom. The naive alternative — Promise.all on the entire array — fires everything at once and gets you throttled immediately. The fix is a bounded concurrency pool.
async function processBatch(items, handler, concurrency = 5) {
const results = new Array(items.length);
let index = 0;
async function worker() {
while (index < items.length) {
const current = index++;
try {
results[current] = await handler(items[current]);
} catch (err) {
results[current] = { error: err.message };
}
}
}
const workers = Array.from({ length: concurrency }, worker);
await Promise.all(workers);
return results;
}
async function summarize(text) {
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 300,
messages: [{ role: "user", content: `Summarize in two sentences: ${text}` }],
}),
});
if (response.status === 429) {
await new Promise((r) => setTimeout(r, 2000));
return summarize(text);
}
const data = await response.json();
return data.content[0].text;
}
const documents = [/* array of strings */];
const summaries = await processBatch(documents, summarize, 5);
This pattern gives you a fixed number of "workers" pulling from a shared queue, so you never have more than concurrency requests in flight. Start with a low concurrency (3–5) and increase it only after confirming you're not hitting rate limits.
Example: batch processing with curl and a shell loop
For quick, one-off jobs — reprocessing a folder of files, testing a prompt across a set of inputs — a shell loop is often faster to write than a script:
#!/bin/bash
for file in ./inputs/*.txt; do
content=$(cat "$file")
name=$(basename "$file" .txt)
curl -s https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d "$(jq -n --arg text "$content" '{
model: "claude-sonnet-4-5",
max_tokens: 500,
messages: [{role: "user", content: $text}]
}')" > "./outputs/${name}.json"
sleep 0.5
done
The sleep between calls is a crude but effective way to avoid bursting past rate limits when you don't want to build proper concurrency control for a throwaway script. For anything you'll run regularly, move to the Node.js pattern above so you get retries and error handling for free.
Handling errors and rate limits inside a batch job
A batch job that dies on the first failed request is not a batch job, it's a liability. Build in:
- Retry with backoff for
429and5xxresponses — don't retry4xxclient errors like malformed requests, they'll just fail again - Per-item error capture so one bad input doesn't kill the whole run — store the error alongside the item and keep going
- A results log written incrementally to disk (or a database), so a crash halfway through doesn't lose completed work
- Idempotency where possible — if you re-run the script, skip items that already have a saved result
async function handlerWithRetry(item, attempt = 1) {
try {
return await summarize(item);
} catch (err) {
if (attempt < 3) {
await new Promise((r) => setTimeout(r, attempt * 1000));
return handlerWithRetry(item, attempt + 1);
}
throw err;
}
}
Tracking usage across a large batch
Once you're sending hundreds or thousands of requests, token usage adds up fast, and it's easy to lose track of which job or feature is driving cost. If you're running batches through SubToAPI, every request against /v1/messages shows up in the dashboard with per-key usage, so you can see exactly what a batch job cost without cross-referencing logs manually. This is particularly useful when different team members or app features each have their own key — see the docs for how request and token metadata is reported, and the messages endpoint reference for the full request shape.
FAQ
Does the Claude API have a dedicated batch endpoint?
Anthropic offers a Message Batches API for large async jobs (up to 24-hour turnaround, cost discount). For most day-to-day bulk tasks, developers instead send concurrent requests to the standard messages endpoint, which is faster to get results from and easier to integrate into existing code.
How many requests can I send concurrently?
There's no universal number — it depends on your account's rate limits. Start with 3–5 concurrent requests, watch for 429 responses, and increase gradually. A bounded worker pool (shown above) is safer than firing all requests at once.
What's the best way to handle failures mid-batch?
Catch errors per item instead of letting one failure stop the whole job, retry only on rate-limit or server errors with backoff, and write results incrementally so a crash doesn't lose completed work. Check /docs/quickstart for request formatting details that help avoid avoidable 4xx errors in the first place.