← Blog

Claude API in AWS Lambda: Serverless Functions Guide

2026-09-29 · 5 min read · SubToAPI Team

Can you call Claude from AWS Lambda?

Yes. Calling the Claude API from an AWS Lambda function works the same way as calling any other HTTPS API: you send a POST request with an API key, get a response, and return it to whatever triggered the function — API Gateway, EventBridge, an SQS queue, or another Lambda. There's no special SDK requirement for Lambda specifically. The parts that actually need attention are Lambda's execution model: cold starts, timeout limits, response streaming, and how you manage secrets and concurrency when your function fans out to hundreds of concurrent invocations.

This article covers the practical setup: writing a Lambda handler that calls Claude, handling timeouts and retries correctly, dealing with streaming responses (which Lambda doesn't support well out of the box), and managing API keys and cost when your Lambda scales up. We'll use plain Node.js examples that work with any HTTPS-compatible Claude API, including a Claude subscription wrapped as an API through SubToAPI.

Basic Lambda handler calling Claude

A minimal handler that calls a Messages-style endpoint looks like this:

// index.mjs
export const handler = async (event) => {
  const body = JSON.parse(event.body || "{}");
  const userMessage = body.message || "Hello";

  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "claude-sonnet-4",
      max_tokens: 1024,
      messages: [{ role: "user", content: userMessage }]
    })
  });

  if (!response.ok) {
    const errText = await response.text();
    return { statusCode: response.status, body: errText };
  }

  const data = await response.json();
  return {
    statusCode: 200,
    body: JSON.stringify(data)
  };
};

A few Lambda-specific details matter here:

Setting the timeout and memory correctly

Lambda's default configuration will kill your function before Claude finishes responding. In your template.yaml (SAM) or CDK stack:

Resources:
  ClaudeFunction:
    Type: AWS::Serverless::Function
    Properties:
      Handler: index.handler
      Runtime: nodejs20.x
      Timeout: 60
      MemorySize: 512
      Environment:
        Variables:
          SUBTOAPI_KEY: !Ref SubToApiKeyParam

Memory doesn't affect Claude call speed much since the function is mostly waiting on network I/O, but 512 MB is a reasonable default that avoids CPU throttling during JSON parsing of large responses.

If you're calling Claude synchronously behind API Gateway, remember API Gateway has its own 29-second timeout for REST APIs (HTTP APIs allow longer). For anything that might take longer than that — long prompts, large max_tokens, or tool-use loops — put the Lambda behind an async pattern instead: API Gateway triggers Lambda, Lambda writes to SQS or invokes itself asynchronously, and the client polls or gets a webhook callback.

Handling streaming in a serverless function

Streaming is the trickiest part of running Claude in Lambda. Traditional Lambda invoked through API Gateway buffers the entire response — it can't stream tokens back to the client incrementally. If your use case needs token-by-token output, you have two options:

  1. Lambda Function URLs with response streaming (InvokeMode: RESPONSE_STREAM), available for Node.js and supported by AWS since 2023. This lets you pipe a streaming HTTP response from Claude directly through Lambda to the client.
  2. Skip streaming in Lambda entirely and return the full response once complete, handling any real-time UI needs on a different compute layer (a small always-on server, or client-side polling against a job status endpoint).

If you go with option 1, the streaming setup mirrors SubToAPI's streaming endpoint, which returns server-sent events you can forward chunk by chunk. For most backend automation use cases — summarization jobs, data enrichment pipelines, scheduled reports — you don't need streaming at all, and non-streaming Lambda is simpler and more reliable.

Managing concurrency and rate limits

Lambda scales by spinning up concurrent instances automatically, which is great for throughput but can slam your Claude API key with more concurrent requests than expected. If 500 Lambda invocations fire at once, you'll get 500 simultaneous API calls unless you throttle.

Practical mitigations:

async function callClaudeWithRetry(payload, attempt = 1) {
  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify(payload)
  });

  if (res.status === 429 && attempt <= 3) {
    await new Promise(r => setTimeout(r, attempt * 1000));
    return callClaudeWithRetry(payload, attempt + 1);
  }
  return res.json();
}

Why route through an API layer instead of raw model access

If your team already pays for Claude through a subscription rather than direct API credits, calling it from Lambda usually means someone has to build the auth translation layer, track usage per function, and rotate keys safely across environments (dev, staging, prod). SubToAPI handles that layer: it issues sub_live_... application keys per environment or per team member, so your Lambda functions in staging and production use separate, revocable keys without anyone sharing a personal login. You get usage metadata per request, which is useful for attributing Claude spend to a specific Lambda function or customer when you're running a fan-out job. Setup takes a few minutes — see the quickstart — and pricing starts at €9/month on the Solo plan, scaling to team seats via pricing if multiple developers or services need their own keys.

FAQs

Does Lambda support streaming responses from Claude? Only through Lambda Function URLs configured with RESPONSE_STREAM invoke mode. Standard API Gateway-triggered Lambdas buffer the full response before returning it, so token-by-token streaming isn't possible in that setup.

What Lambda timeout should I set for Claude API calls? At least 30–60 seconds for typical completions, adjusted for your max_tokens and prompt size. Also check any upstream timeout, like API Gateway's 29-second limit on REST APIs, which can cut off a slower Lambda invocation.

How do I avoid hitting rate limits when Lambda scales up? Set reserved concurrency on the function, buffer bursts through SQS, and implement retry-with-backoff for 429 responses. This keeps concurrent Claude calls within your API key's actual rate limit instead of scaling unchecked with Lambda's default behavior.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →