← Blog

Claude API Error Handling: Status Codes Explained

2026-10-05 · 5 min read · SubToAPI Team

When a request to the Claude API fails, the HTTP status code and the error body tell you exactly what went wrong and whether retrying makes sense. Getting this right matters: treat every failure the same way and you'll either hammer the API with retries that can't succeed, or silently drop requests that would have worked on a second try.

This guide covers every status code you'll realistically encounter, what each one means in practice, and a retry strategy you can drop into production code today.

The Claude API Error Response Shape

Errors come back as JSON with a consistent structure:

{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "max_tokens: Input should be less than or equal to 4096"
  }
}

The error.type field is more useful than the status code alone for deciding why something failed, but the HTTP status code is what you should branch on for retry logic, since it's stable across providers and easy to check without parsing the body.

Status Codes and What They Mean

400 — Invalid Request

The request body is malformed: a missing required field, a parameter out of range, or an invalid combination (e.g., temperature and top_p both set when the API doesn't allow it). This is a client-side bug. Do not retry without fixing the request — retrying an identical malformed request just fails again.

401 — Authentication Error

Your API key is missing, invalid, or revoked. Check the Authorization header format and confirm the key hasn't been rotated. No amount of retrying fixes this; you need a valid key.

403 — Permission Error

The key is valid but doesn't have access to the requested resource or model. Common when a key is scoped to specific models or features. Fix the permission/scope, don't retry blindly.

404 — Not Found

The endpoint or resource (e.g., a model ID that's been deprecated) doesn't exist. Double-check the URL path and model name.

413 — Request Too Large

The request payload exceeds size limits, often from oversized attachments, images, or very long conversation history. Trim the payload; retrying as-is won't help.

429 — Rate Limited

You've exceeded requests-per-minute, tokens-per-minute, or concurrent request limits. This is the most common retryable error. The response usually includes a retry-after header — respect it. Without that header, use exponential backoff.

500 — Internal Server Error

Something failed on the provider's side. Rare, but worth a retry with backoff since it's transient by nature.

529 — Overloaded

The API is temporarily over capacity. This is also transient and retryable, but you should back off more aggressively than with a 429 since it signals systemic load, not just your own usage.

Retry Logic That Actually Works

A simple rule: retry on 429, 500, and 529. Everything else (400, 401, 403, 404, 413) is a client-side problem that won't resolve itself.

async function callWithRetry(fn, maxRetries = 4) {
  const retryable = new Set([429, 500, 529]);
  let attempt = 0;

  while (true) {
    try {
      return await fn();
    } catch (err) {
      const status = err.status;
      attempt++;

      if (!retryable.has(status) || attempt > maxRetries) {
        throw err;
      }

      const retryAfter = err.headers?.get?.('retry-after');
      const delayMs = retryAfter
        ? Number(retryAfter) * 1000
        : Math.min(1000 * 2 ** attempt, 15000) + Math.random() * 500;

      await new Promise(r => setTimeout(r, delayMs));
    }
  }
}

Key points in that logic:

Logging Errors Usefully

Don't just log the status code — log error.type and error.message alongside it, plus the request ID if the response includes one. When you're debugging a spike in 429s weeks later, "rate_limit_error" with a timestamp tells you far more than a bare "429" in your logs.

catch (err) {
  console.error({
    status: err.status,
    type: err.body?.error?.type,
    message: err.body?.error?.message,
    requestId: err.headers?.get('request-id'),
  });
}

Handling Errors Across Multiple API Keys

If you're building a product on top of Claude — issuing per-customer or per-feature API keys — error handling gets more complex. A 429 might mean one customer's key is rate limited while others are fine, and you need visibility into which key failed, how often, and whether it's a pattern worth alerting on.

This is one of the problems SubToAPI solves: it turns your Claude access into application API keys (sub_live_...) with usage metadata per key, so when something returns a 429 or 500 you can see exactly which key, team, or feature triggered it instead of digging through a single shared log stream. Error responses follow the same structure described above, so the retry logic you write today works unchanged. See the quickstart and messages docs for request/response details, or check pricing if you're evaluating plans.

Checklist for Production-Grade Error Handling

Questions

Should I retry every failed Claude API request? No. Only retry on 429, 500, and 529 — these are transient. Status codes like 400, 401, 403, and 404 indicate a problem with the request itself that retrying won't fix.

What's the difference between a 429 and a 529 error? 429 means you've hit your own rate limit (requests or tokens per minute). 529 means the API is overloaded across all users. Both are retryable, but back off more aggressively on 529.

How do I know if an error is my fault or the API's fault? Status codes in the 4xx range (400, 401, 403, 404, 413) are client-side issues — fix the request, key, or payload. 5xx and 529 are server-side and transient, so retrying with backoff is the right move.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →