← Blog

Claude API Batch Processing Requests: Example Guide

2026-09-28 · 5 min read · SubToAPI Team

What "batch processing" actually means for the Claude API

When developers search for Claude API batch processing, they usually mean one of two different things. First, Anthropic's native Message Batches API, which accepts up to 100,000 requests in a single job, processes them asynchronously, and returns results within a 24-hour window at a reduced cost. This is built for large, non-urgent workloads like classifying a million support tickets overnight.

Second — and far more common in practice — is the pattern of sending a large number of individual requests to the standard messages endpoint, but doing it programmatically with controlled concurrency, retries, and rate-limit handling instead of firing them all at once or looping one-by-one. This is what most people actually need: summarizing 500 documents, generating descriptions for a product catalog, or re-scoring a dataset, and getting results back within minutes rather than waiting on an async job. The rest of this article focuses on that second pattern, since it's the one you'll build yourself most often.

Native batch API vs. concurrent request loops

Use Anthropic's Message Batches API when:

Use a concurrent request loop when:

Most "batch processing" scripts developers write day-to-day fall into the second category, so that's what the examples below cover.

Example: batch processing with a Node.js concurrency pool

The naive approach — a for loop with await inside it — sends requests one at a time and wastes most of your rate limit headroom. The naive alternative — Promise.all on the entire array — fires everything at once and gets you throttled immediately. The fix is a bounded concurrency pool.

async function processBatch(items, handler, concurrency = 5) {
  const results = new Array(items.length);
  let index = 0;

  async function worker() {
    while (index < items.length) {
      const current = index++;
      try {
        results[current] = await handler(items[current]);
      } catch (err) {
        results[current] = { error: err.message };
      }
    }
  }

  const workers = Array.from({ length: concurrency }, worker);
  await Promise.all(workers);
  return results;
}

async function summarize(text) {
  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 300,
      messages: [{ role: "user", content: `Summarize in two sentences: ${text}` }],
    }),
  });

  if (response.status === 429) {
    await new Promise((r) => setTimeout(r, 2000));
    return summarize(text);
  }

  const data = await response.json();
  return data.content[0].text;
}

const documents = [/* array of strings */];
const summaries = await processBatch(documents, summarize, 5);

This pattern gives you a fixed number of "workers" pulling from a shared queue, so you never have more than concurrency requests in flight. Start with a low concurrency (3–5) and increase it only after confirming you're not hitting rate limits.

Example: batch processing with curl and a shell loop

For quick, one-off jobs — reprocessing a folder of files, testing a prompt across a set of inputs — a shell loop is often faster to write than a script:

#!/bin/bash

for file in ./inputs/*.txt; do
  content=$(cat "$file")
  name=$(basename "$file" .txt)

  curl -s https://api.subtoapi.app/v1/messages \
    -H "Authorization: Bearer $SUBTOAPI_KEY" \
    -H "Content-Type: application/json" \
    -d "$(jq -n --arg text "$content" '{
      model: "claude-sonnet-4-5",
      max_tokens: 500,
      messages: [{role: "user", content: $text}]
    }')" > "./outputs/${name}.json"

  sleep 0.5
done

The sleep between calls is a crude but effective way to avoid bursting past rate limits when you don't want to build proper concurrency control for a throwaway script. For anything you'll run regularly, move to the Node.js pattern above so you get retries and error handling for free.

Handling errors and rate limits inside a batch job

A batch job that dies on the first failed request is not a batch job, it's a liability. Build in:

async function handlerWithRetry(item, attempt = 1) {
  try {
    return await summarize(item);
  } catch (err) {
    if (attempt < 3) {
      await new Promise((r) => setTimeout(r, attempt * 1000));
      return handlerWithRetry(item, attempt + 1);
    }
    throw err;
  }
}

Tracking usage across a large batch

Once you're sending hundreds or thousands of requests, token usage adds up fast, and it's easy to lose track of which job or feature is driving cost. If you're running batches through SubToAPI, every request against /v1/messages shows up in the dashboard with per-key usage, so you can see exactly what a batch job cost without cross-referencing logs manually. This is particularly useful when different team members or app features each have their own key — see the docs for how request and token metadata is reported, and the messages endpoint reference for the full request shape.

FAQ

Does the Claude API have a dedicated batch endpoint?

Anthropic offers a Message Batches API for large async jobs (up to 24-hour turnaround, cost discount). For most day-to-day bulk tasks, developers instead send concurrent requests to the standard messages endpoint, which is faster to get results from and easier to integrate into existing code.

How many requests can I send concurrently?

There's no universal number — it depends on your account's rate limits. Start with 3–5 concurrent requests, watch for 429 responses, and increase gradually. A bounded worker pool (shown above) is safer than firing all requests at once.

What's the best way to handle failures mid-batch?

Catch errors per item instead of letting one failure stop the whole job, retry only on rate-limit or server errors with backoff, and write results incrementally so a crash doesn't lose completed work. Check /docs/quickstart for request formatting details that help avoid avoidable 4xx errors in the first place.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →