← Blog

Claude API Serverless Deployment on AWS Lambda

2026-10-03 · 5 min read · SubToAPI Team

Can you run the Claude API on AWS Lambda?

Yes, and for most request/response workloads — chat endpoints, document summarization, background enrichment jobs — Lambda is a solid fit for calling the Claude API. You get no servers to patch, automatic scaling, and pay-per-invocation billing. The practical issues you need to solve are cold start latency, Lambda's execution time limit, how you handle streaming responses, and where you store your API key securely.

This article walks through a working deployment pattern: function structure, packaging, timeout configuration, and the specific tradeoffs of running an LLM call inside a serverless function. We'll also look at when serverless isn't the right shape for Claude workloads, particularly around streaming.

Basic Lambda function calling Claude

A minimal handler that calls Claude's API and returns the response looks like this:

// index.mjs
export const handler = async (event) => {
  const body = JSON.parse(event.body || "{}");

  const response = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      messages: [{ role: "user", content: body.prompt }],
    }),
  });

  const data = await response.json();

  return {
    statusCode: 200,
    headers: { "content-type": "application/json" },
    body: JSON.stringify(data),
  };
};

Node.js 18+ runtimes include fetch natively, so you don't need to bundle an HTTP client. This keeps your deployment package small, which matters for cold start time.

Packaging and cold starts

Cold starts are the first thing people run into. A Lambda invoked after idle time has to initialize the runtime before your code runs. For a function that only does an HTTPS call, this overhead is usually under a second on Node.js — small compared to the Claude generation time itself, which is typically 2-15 seconds depending on output length.

Things that make cold starts worse:

If cold starts are a real problem for your traffic pattern (bursty, low-frequency invocations), configure provisioned concurrency on the function, or keep a scheduled warmer invocation every few minutes for latency-sensitive paths.

Timeout configuration

Lambda's default timeout is 3 seconds, which will truncate almost any real Claude call. Set the timeout explicitly based on your expected response length:

aws lambda update-function-configuration \
  --function-name claude-proxy \
  --timeout 60 \
  --memory-size 512

For long-form generation (multi-page documents, large code outputs), you may need 60-120 seconds. Lambda's hard ceiling is 900 seconds (15 minutes) — plenty for non-streaming Claude calls, but you should also set a matching timeout on API Gateway if you're fronting the function with one, since API Gateway's own timeout is capped at 29 seconds for REST APIs. This is the single most common cause of "Lambda works but my API returns a gateway timeout" bugs.

Streaming doesn't work the way you'd want

This is the real architectural limitation. Claude's streaming API sends tokens incrementally over a long-lived HTTP connection, and classic Lambda-behind-API-Gateway doesn't support that well: API Gateway buffers the response and returns it all at once, defeating the purpose of streaming.

Options if you need token-by-token delivery:

  1. Lambda response streaming via Function URLs. AWS supports awslambda.streamifyResponse() for functions invoked through a Lambda Function URL (not API Gateway), which does support chunked responses. This is the closest native option, but it adds complexity and isn't available through API Gateway integrations.
  2. Skip streaming at the Lambda layer. Many apps don't actually need token-by-token UI updates — they need the full response. If that's you, non-streaming Lambda is simpler and more reliable.
  3. Use a managed API layer instead of raw Lambda for streaming. If your product needs real streaming to the browser reliably, a persistent service (ECS/Fargate, or a managed API) handles long-lived connections more naturally than Lambda's execution model.

If your use case is mostly serverless request/response and you don't want to manage key storage, retries, and streaming plumbing yourself, SubToAPI gives you a stable HTTPS endpoint with built-in streaming support (see /docs/streaming) that you can call from a Lambda function exactly like you'd call any REST API — no SDK, no connection management on your end.

Secrets management

Never hardcode your API key in the function code or environment variables checked into source control. Use AWS Secrets Manager or SSM Parameter Store, and fetch the key at cold start (not on every invocation, to avoid extra latency and API calls):

import { SecretsManagerClient, GetSecretValueCommand } from "@aws-sdk/client-secrets-manager";

let cachedKey;

async function getApiKey() {
  if (cachedKey) return cachedKey;
  const client = new SecretsManagerClient({});
  const result = await client.send(
    new GetSecretValueCommand({ SecretId: "claude/api-key" })
  );
  cachedKey = JSON.parse(result.SecretString).apiKey;
  return cachedKey;
}

Caching the key in a module-level variable means it's only fetched once per container lifecycle, not per request.

If you're routing through SubToAPI instead of calling Anthropic directly, the same pattern applies to your sub_live_... key: store it in Secrets Manager, fetch once at cold start, and send it as a Bearer token exactly as shown in /docs/quickstart.

Concurrency and cost control

Set a reserved concurrency limit on the function so a traffic spike can't run up an unexpectedly large bill or hit your Claude rate limits all at once:

aws lambda put-function-concurrency \
  --function-name claude-proxy \
  --reserved-concurrent-executions 20

This caps how many invocations run in parallel, which indirectly throttles your outbound request rate to Claude — useful if you're on a tier with a fixed requests-per-minute limit.

When to skip Lambda entirely

Serverless Lambda makes sense for background jobs, webhook handlers, and request/response APIs with moderate traffic. It's a worse fit if you need persistent streaming connections at scale, very high request volume (where container reuse and connection pooling matter more), or sub-100ms latency requirements where cold starts are unacceptable even with provisioned concurrency. In those cases, a long-running container service — or an API layer like SubToAPI that handles the infrastructure for you, with per-seat pricing starting at /pricing — removes the operational overhead without giving up the HTTPS API model.

FAQ

Does Claude's API support streaming on Lambda? Only through Lambda Function URLs with response streaming enabled, not through API Gateway. If you need reliable token streaming, consider a persistent service or a managed API that handles streaming for you.

What timeout should I set for a Claude Lambda function? At least 60 seconds for typical responses, up to Lambda's 900-second max for long-form generation. Make sure any API Gateway in front of it has a matching or higher timeout, up to its 29-second REST API cap.

How do I avoid cold start delays for Claude calls on Lambda? Keep the deployment package small, avoid unnecessary VPC attachment, and use provisioned concurrency if your traffic is latency-sensitive and bursty.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →