← Blog

Claude API Async Request Queue Setup Guide

2026-10-05 · 5 min read · SubToAPI Team

If you're sending more than a handful of requests to the Claude API, you need a queue. Firing requests directly from your application code works fine in a demo, but in production you'll hit rate limits, 429s, and timeout cascades the moment traffic spikes. An async request queue decouples request submission from execution, so your app stays responsive while a worker pool processes Claude calls at a controlled rate.

This article covers how to set up that queue: the core components, a working Node.js implementation, concurrency and retry strategy, and when to reach for a managed layer instead of building your own.

Why you need a queue, not just async/await

async/await handles concurrency within a single request, but it doesn't manage how many requests run at once across your whole system. Without a queue:

A queue solves this by giving you three things: a bounded concurrency limit, a place to retry failed jobs without retrying everything, and backpressure so your producer (the part of your app generating requests) doesn't overwhelm the consumer (the part calling Claude).

Core components of an async queue

A minimal setup needs:

  1. A job store — in-memory array, Redis list, or a database table holding pending requests.
  2. A worker pool — a fixed number of concurrent workers pulling jobs and calling the API.
  3. Retry logic with backoff — for 429s and 5xx errors specifically, not for 4xx client errors.
  4. A result handler — something that receives the completion and routes it back to the caller (webhook, database write, or resolved promise).

For low-to-moderate volume, in-memory is fine. Once you need durability across restarts or multiple processes, move the job store to Redis or Postgres.

A working in-memory implementation

Here's a concurrency-limited queue using p-queue-style logic, written without the dependency so you can see what's happening:

class AsyncQueue {
  constructor({ concurrency = 5, maxRetries = 3 }) {
    this.concurrency = concurrency;
    this.maxRetries = maxRetries;
    this.queue = [];
    this.active = 0;
  }

  add(task) {
    return new Promise((resolve, reject) => {
      this.queue.push({ task, resolve, reject, attempts: 0 });
      this._next();
    });
  }

  async _next() {
    if (this.active >= this.concurrency || this.queue.length === 0) return;

    const job = this.queue.shift();
    this.active++;

    try {
      const result = await job.task();
      job.resolve(result);
    } catch (err) {
      if (job.attempts < this.maxRetries && this._isRetryable(err)) {
        job.attempts++;
        const delay = 500 * 2 ** job.attempts;
        setTimeout(() => {
          this.queue.unshift(job);
          this.active--;
          this._next();
        }, delay);
        return;
      }
      job.reject(err);
    }

    this.active--;
    this._next();
  }

  _isRetryable(err) {
    return err.status === 429 || (err.status >= 500 && err.status < 600);
  }
}

Usage against any Claude-compatible endpoint:

const queue = new AsyncQueue({ concurrency: 5, maxRetries: 3 });

async function callClaude(prompt) {
  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
      "content-type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-20250514",
      max_tokens: 1024,
      messages: [{ role: "user", content: prompt }],
    }),
  });

  if (!res.ok) {
    const err = new Error("Request failed");
    err.status = res.status;
    throw err;
  }
  return res.json();
}

const results = await Promise.all(
  prompts.map((p) => queue.add(() => callClaude(p)))
);

This gives you bounded concurrency (only 5 requests in flight), exponential backoff on retryable errors, and immediate rejection on client errors like bad request bodies.

Tuning concurrency and timeouts

A few practical rules:

Scaling beyond a single process

The in-memory queue above works well for a single Node process. Once you need multiple instances or workers that survive restarts, move the job list to Redis (using something like BullMQ) or a Postgres table with a status column (pending, processing, done, failed). The worker logic stays almost identical — you're just replacing this.queue.shift() with a database poll or a Redis BRPOP.

Where SubToAPI fits

If you're building this queue specifically to manage rate limits and concurrency on top of your own Claude API key, it's worth checking whether you actually need to own that infrastructure. SubToAPI turns your existing Claude access into an HTTPS API with application keys, built-in streaming, and usage metadata per key — so you can issue separate sub_live_... keys per worker or per customer instead of funneling everything through one rate limit.

That doesn't eliminate the need for a queue in your application — you'll still want concurrency control on your side — but it removes the guesswork around which key is consuming quota and gives you usage visibility per team or per app without building that tracking yourself. See the quickstart and messages docs for the request format.

questions

Do I need a queue if I'm only making a few requests per minute? No. If your traffic is low and predictable, direct calls with basic try/catch retry logic are simpler and sufficient. Queues earn their complexity at higher concurrency or burst traffic.

Should retries happen inside the queue or at the API client level? Inside the queue, scoped per job. This keeps backoff logic centralized and lets you track retry counts without duplicating logic across every call site in your app.

What's the difference between a request queue and a message queue like SQS? A request queue (like the one above) manages concurrency for synchronous-style API calls. A message queue like SQS decouples producers and consumers entirely and is better suited for fully asynchronous, fire-and-forget workflows with separate consumer processes.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →