Claude API Batch Processing: A Practical Guide
If you need to run Claude over thousands of documents, support tickets, or rows in a spreadsheet, sending requests one at a time is slow and sending them all at once will get you rate-limited. "Batch processing" for the Claude API means structuring many requests so they run efficiently and reliably, without overwhelming your account's concurrency and token-per-minute limits.
There are two practical approaches: Anthropic's asynchronous Batches API for large, non-urgent jobs that can wait hours for results at a reduced cost, and client-side concurrent batching for workloads where you still want results within seconds or minutes. This guide covers both, plus the retry and tracking logic you need regardless of which one you use.
When to use async batches vs. live concurrent requests
Async batch processing fits jobs like:
- Classifying or summarizing a large archive of historical documents
- Generating embeddings-adjacent text transformations for a dataset
- Any workload where a multi-hour turnaround is acceptable
You submit a collection of requests together, the provider processes them in the background, and you poll or get notified when results are ready. This is the right call when you don't need the output immediately and want to avoid babysitting rate limits yourself.
Live concurrent batching fits jobs like:
- Processing an inbox of support tickets as they arrive, in groups
- Backfilling a feature where users are waiting on results within the same session
- Any pipeline where you control concurrency yourself and need answers now
Most production systems end up using a mix: async batches for bulk historical work, concurrent requests for anything user-facing.
Building a concurrent batch processor
The core problem with "just send 500 requests with Promise.all" is that you'll blow through your concurrency limit and get a wall of 429s. Instead, cap how many requests are in flight at once and queue the rest.
async function runBatch(items, worker, concurrency = 5) {
const results = new Array(items.length);
let index = 0;
async function next() {
while (index < items.length) {
const current = index++;
try {
results[current] = await worker(items[current]);
} catch (err) {
results[current] = { error: err.message };
}
}
}
const workers = Array.from({ length: concurrency }, next);
await Promise.all(workers);
return results;
}
Then wrap each Claude call as the worker function:
async function classifyTicket(ticket) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 100,
messages: [
{ role: "user", content: `Classify this support ticket in one word: ${ticket.text}` }
]
})
});
if (!res.ok) throw new Error(`Request failed: ${res.status}`);
const data = await res.json();
return data.content[0].text;
}
const results = await runBatch(tickets, classifyTicket, 5);
Start with a concurrency of 3–5 and increase it gradually while watching for 429 responses. Your effective limit depends on your plan tier and the size of each prompt — large prompts consume more tokens per minute, so you'll hit limits sooner even at low request counts.
Handling retries and backoff
Batch jobs fail individual requests more often than single calls, simply because you're making more of them. Build retry logic with exponential backoff directly into your worker function rather than retrying the whole batch:
async function withRetry(fn, maxAttempts = 4) {
let attempt = 0;
while (attempt < maxAttempts) {
try {
return await fn();
} catch (err) {
attempt++;
if (attempt >= maxAttempts) throw err;
const delay = Math.min(1000 * 2 ** attempt, 15000);
await new Promise(r => setTimeout(r, delay));
}
}
}
Only retry on transient failures — 429 (rate limited), 500, 502, 503, 504. Don't retry on 400 or 401; those won't fix themselves and will just waste time and tokens. See /docs for the full error reference your batch logic should branch on.
Chunking large jobs
For jobs with thousands of items, don't queue everything in memory at once. Split into chunks of a few hundred, process each chunk fully, log progress, then move to the next:
function chunk(arr, size) {
const out = [];
for (let i = 0; i < arr.length; i += size) out.push(arr.slice(i, i + size));
return out;
}
for (const group of chunk(allTickets, 200)) {
const groupResults = await runBatch(group, classifyTicket, 5);
await saveResults(groupResults);
console.log(`Processed ${group.length} items`);
}
This makes jobs resumable — if something crashes halfway through, you've already persisted the earlier chunks and only need to re-run what's left.
Tracking usage across a batch run
Batch jobs are where token costs add up fastest, and it's easy to lose track of spend when you're firing hundreds of requests from a script. Every Claude API response includes usage data (input and output token counts) — log it per request so you can total cost per job, not just per account.
This is one of the reasons teams running regular batch jobs route requests through SubToAPI: every call made with an application key (sub_live_...) surfaces usage metadata in a single dashboard, so you can see exactly how much a batch run cost across all your team's keys without reconciling logs from multiple scripts or environments. Combined with per-key rate limit visibility, it makes it much easier to tune concurrency instead of guessing. Check /docs/messages for the request and response format, and /docs/quickstart if you're setting this up for the first time.
Practical tips
- Keep max_tokens tight. Batch jobs often don't need long outputs — capping
max_tokensreduces both latency and cost per item. - Validate inputs before sending. Filter out empty or malformed items client-side; it's cheaper than paying for a failed API call.
- Log request IDs. If something goes wrong at scale, you want to trace individual failures back to specific items.
- Separate read and write phases. Fetch all source data first, then run the batch, then write results — don't interleave database calls with API calls in the same loop.
questions
Does batch processing reduce the cost of each Claude API call? Client-side concurrent batching doesn't change per-token pricing — it just changes how fast requests complete. Anthropic's dedicated async Batches API does offer discounted rates in exchange for delayed, non-realtime processing.
How many requests can I run concurrently? It depends on your plan's rate limits, not a fixed number. Start with 3–5 concurrent requests, monitor for 429 responses, and increase gradually while watching your token-per-minute usage.
Should I use async batches or concurrent requests for a one-time data migration? If the job can wait hours and covers thousands of items, async batch submission is usually cheaper and simpler. For smaller jobs or anything time-sensitive, concurrent requests with retry logic get you results faster.