Claude API Bulk Request Batching: A Practical Guide
If you need to run Claude over thousands of rows — summarizing documents, classifying support tickets, translating content, generating embeddings-adjacent metadata — sending requests one at a time in a loop will either take forever or trip your rate limits. Bulk request batching is how you process large volumes of Claude API calls reliably: grouping work into controlled chunks, running a bounded number of requests in parallel, and handling retries without losing progress on a multi-hour job.
This guide covers two approaches: using Anthropic's native Message Batches API for large async jobs, and building your own concurrency-controlled batching pipeline when you need lower latency or finer control over ordering and error handling.
When you actually need batching
Not every bulk job needs a batching strategy. If you're processing 20 items, a simple for loop with await is fine. Batching matters once you're dealing with:
- Thousands of items where sequential requests would take hours
- Rate limit constraints that reject bursts of concurrent calls
- Cost-sensitive workloads where async processing at a discount makes sense
- Long-running jobs that need to survive a crash or restart without redoing completed work
If your workload matches any of these, you need either the native batch API or a self-managed queue with concurrency limits.
Option 1: Anthropic's Message Batches API
Anthropic offers a dedicated batch endpoint for asynchronous, high-volume processing. You submit a set of requests as a single batch job, Anthropic processes them within a 24-hour window, and you poll for results. The main advantages are a significant cost discount compared to synchronous calls and no need to manage your own concurrency limits — Anthropic handles the throughput internally.
The batch API is the right choice when:
- Latency doesn't matter (results can arrive minutes to hours later)
- You're processing a large, fixed dataset (thousands to millions of requests)
- Cost per request matters more than turnaround time
It's the wrong choice when you need results in seconds, when your workload is interactive, or when request volume is small enough that the overhead of managing a batch job isn't worth it.
Option 2: Self-managed concurrency batching
For workloads that need faster turnaround than a 24-hour async window, or where you're calling Claude through a gateway like SubToAPI for streaming, tool use, and per-request usage metadata, you'll want to build your own batching loop with bounded concurrency.
The core pattern: chunk your dataset, run a fixed number of requests in parallel, retry failures with backoff, and track progress so a crash doesn't force a full restart.
const CONCURRENCY = 5;
async function runBatch(items, handler) {
const results = new Array(items.length);
let index = 0;
async function worker() {
while (index < items.length) {
const current = index++;
results[current] = await withRetry(() => handler(items[current]));
}
}
await Promise.all(
Array.from({ length: CONCURRENCY }, () => worker())
);
return results;
}
async function withRetry(fn, attempts = 3) {
for (let i = 0; i < attempts; i++) {
try {
return await fn();
} catch (err) {
if (i === attempts - 1) throw err;
const delay = 500 * 2 ** i;
await new Promise((r) => setTimeout(r, delay));
}
}
}
This gives you a fixed number of concurrent in-flight requests (5 in this example), automatic retries with exponential backoff, and a results array that preserves input order regardless of which request finishes first.
Calling the API inside the batch worker
async function classifyTicket(ticket) {
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 200,
messages: [
{ role: "user", content: `Classify this ticket:\n\n${ticket.text}` },
],
}),
});
if (!response.ok) throw new Error(`Request failed: ${response.status}`);
return response.json();
}
const results = await runBatch(tickets, classifyTicket);
If you're already routing traffic through SubToAPI, the same sub_live_... key works across every request in the batch, and the dashboard gives you per-key usage so you can watch token consumption and request counts climb as the job runs — useful for catching a runaway job before it burns through your monthly quota. See /docs/messages for the full request schema and /docs/quickstart if you're setting up a key for the first time.
Chunking strategy: picking concurrency and batch size
There's no universal "correct" concurrency number — it depends on your rate limits, average response time, and how latency-sensitive the job is. A few practical guidelines:
- Start conservative. Begin with concurrency of 3–5 and increase only if you're not seeing 429s or timeouts.
- Chunk large datasets. Don't hold 100,000 items in memory. Process in chunks of 500–1000, checkpoint progress after each chunk, and resume from the last checkpoint if the process dies.
- Separate read and write concerns. Fetch input data, run the batch, and write results in distinct steps so a failure in one doesn't corrupt the others.
- Log failures separately. Don't let one bad row kill the whole job — catch errors per item, log the failing IDs, and retry them in a second pass.
const CHUNK_SIZE = 500;
for (let i = 0; i < allItems.length; i += CHUNK_SIZE) {
const chunk = allItems.slice(i, i + CHUNK_SIZE);
const chunkResults = await runBatch(chunk, classifyTicket);
await saveCheckpoint(i, chunkResults);
}
Idempotency and safe retries
Retrying failed requests is only safe if repeating a call doesn't produce duplicate side effects. If your batch job writes results to a database, use the input item's ID as the write key so a retried request overwrites the same row instead of creating a duplicate. This matters more for batching than for single requests, because a job processing thousands of items will always have some fraction fail and need a retry.
Streaming vs non-streaming in bulk jobs
For bulk classification or extraction tasks, use non-streaming requests — you want the full response before moving to the next item, and streaming adds complexity with no benefit when there's no user watching output arrive in real time. Streaming makes sense for interactive use cases, not batch jobs. If part of your pipeline does need live output (e.g., a preview step before the bulk run), see /docs/streaming for the difference in request handling.
FAQ
What's the difference between the Message Batches API and building my own batching loop?
The Batches API is fully async — you submit a job and poll for results up to 24 hours later, at a lower cost per request. A self-managed loop with concurrency control gives you results in seconds to minutes and more control over retries and ordering, at standard per-request pricing.
How many concurrent requests can I safely run?
It depends on your rate limits and account tier. Start with concurrency of 3–5, monitor for 429 responses, and increase gradually. If you're routing through SubToAPI, check your plan's limits on /pricing before scaling up a batch job.
Should I use tool calling inside a bulk batch job?
Only if each item genuinely needs it — tool use adds latency and complexity per request. For straightforward classification, summarization, or extraction, a plain message request without tools is faster and simpler to retry. See /docs/tools if your batch job needs structured tool-based outputs.