← Blog

Claude API Webhook Retry Logic: A Working Example

2026-10-11 · 5 min read · SubToAPI Team

If you're calling the Claude API from a webhook handler — Stripe payment confirmations, GitHub push events, a support ticket created in Zendesk — you need retry logic that survives network blips, rate limits, and partial failures without double-processing anything or losing a request forever. This article walks through a concrete retry implementation: how to classify errors, when to back off, how to make retries idempotent, and how to stop retrying gracefully with a dead-letter queue.

The core problem with webhooks and LLM calls together is that webhook senders (Stripe, GitHub, Twilio, your own queue) have their own retry policies and timeouts, and the Claude API has its own failure modes (429 rate limits, 5xx errors, timeouts on long streaming responses). If you don't handle both layers correctly, you end up with either dropped requests or duplicate side effects — like sending the same AI-generated reply twice.

Classify errors before you retry anything

Not every failure should be retried the same way. Build a small classifier first:

function classifyError(error) {
  const status = error.status || error.response?.status;

  if (status === 429) return "rate_limit";       // retry with backoff, respect Retry-After
  if (status >= 500 && status < 600) return "server_error"; // retry with backoff
  if (status === 400 || status === 422) return "bad_request"; // do NOT retry
  if (status === 401 || status === 403) return "auth_error";  // do NOT retry, alert instead
  if (error.code === "ETIMEDOUT" || error.code === "ECONNRESET") return "network";
  return "unknown";
}

Retrying a 400 just wastes your retry budget and delays a failure you already know how to detect. Only rate_limit, server_error, and network are worth retrying automatically.

Exponential backoff with jitter

Fixed-interval retries create thundering-herd problems when multiple webhook events fail at once. Use exponential backoff with jitter, and respect Retry-After when the API provides it:

async function callClaudeWithRetry(payload, maxRetries = 5) {
  let attempt = 0;

  while (attempt <= maxRetries) {
    try {
      return await callClaude(payload);
    } catch (error) {
      const kind = classifyError(error);
      if (kind === "bad_request" || kind === "auth_error") throw error;
      if (attempt === maxRetries) throw error;

      const retryAfter = error.response?.headers?.["retry-after"];
      const base = retryAfter ? Number(retryAfter) * 1000 : 2 ** attempt * 500;
      const jitter = Math.random() * 300;
      await new Promise((r) => setTimeout(r, base + jitter));

      attempt++;
    }
  }
}

This gives you delays roughly like 500ms, 1s, 2s, 4s, 8s, with a small random offset so retries from different webhook events don't all land on the same second.

Make the retry idempotent

The part most implementations skip: if your webhook handler retries the Claude call but the first attempt actually succeeded (slow response, timeout on the client side, but the server finished), you can end up sending two replies to a customer or writing two rows to a database.

Fix this with an idempotency key tied to the original webhook event, not the retry attempt:

function idempotencyKeyFor(webhookEvent) {
  return `wh_${webhookEvent.id}`;
}

async function handleWebhook(event) {
  const key = idempotencyKeyFor(event);
  const existing = await db.lookup(key);
  if (existing) return existing.result; // already processed, skip the call entirely

  const result = await callClaudeWithRetry(buildPayload(event));
  await db.save(key, result);
  return result;
}

Store the idempotency key before you call the API, and check it at the very start of the handler — including on retries triggered by the webhook sender itself (Stripe, for example, will redeliver the same event ID if your endpoint doesn't return 200 fast enough).

Separate webhook timeouts from API call duration

Most webhook senders expect a response within a few seconds (Stripe times out around 10–20 seconds depending on the API version). A Claude API call — especially a longer completion — can take longer than that, particularly without streaming. Don't block the webhook response on the full completion:

app.post("/webhooks/support-ticket", async (req, res) => {
  res.status(200).send("accepted"); // acknowledge immediately
  await queue.enqueue("process-ticket", req.body); // process asynchronously
});

Do the actual Claude call and retry logic inside a background worker consuming that queue, not inside the HTTP handler. This decouples the webhook sender's retry policy from your own retry logic entirely — you only need to worry about one retry loop, not two overlapping ones.

Dead-letter queue for exhausted retries

When maxRetries is hit, don't drop the event. Push it to a dead-letter queue with the error context so a human (or a scheduled job) can inspect and replay it:

async function processTicket(event) {
  try {
    const result = await callClaudeWithRetry(buildPayload(event));
    await saveResult(event.id, result);
  } catch (error) {
    await deadLetterQueue.push({
      eventId: event.id,
      payload: event,
      error: error.message,
      failedAt: new Date().toISOString(),
    });
  }
}

Alert on dead-letter volume, not on individual failures — a single retry exhaustion is normal under load; a spike means something upstream (auth, quota, outage) actually broke.

Where SubToAPI fits

If you're running this kind of retry logic against an application API key through SubToAPI, the same pattern applies without changes — SubToAPI exposes standard HTTPS responses with status codes and headers you can classify the same way, so your classifyError and backoff functions don't need special-casing. Usage metadata returned per request also makes it easier to tell a rate-limit backoff apart from a genuine outage when you're debugging the dead-letter queue. See the quickstart and messages docs for request/response shapes, and streaming docs if you're moving long-running webhook calls to streamed responses instead of blocking waits.

Summary checklist

Questions

Should I retry a Claude API timeout the same way as a 500 error? Yes — treat client-side timeouts and network errors the same as server errors: retry with exponential backoff, since the request may have actually succeeded server-side and your idempotency key will prevent duplicate side effects.

How many retries is reasonable for a webhook-triggered Claude call? Three to five is typical. Beyond that you're usually masking a real outage or auth problem, and the event is better handled by a dead-letter queue than an endless retry loop.

Can I use the webhook sender's own retries instead of building my own? No — webhook senders retry the delivery of the event, not your internal Claude API call. If your handler fails mid-call, you need your own retry logic around that call, with idempotency keys so sender-level redeliveries don't cause duplicate processing.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →