← Blog

Claude API AWS Lambda Integration Guide

2026-10-06 · 5 min read · SubToAPI Team

Running Claude inside an AWS Lambda function is a common pattern for serverless chatbots, document processors, webhooks, and scheduled jobs that need an LLM call without maintaining a server. This guide walks through the architecture, the Lambda-specific constraints you'll hit (timeouts, cold starts, streaming), and working code for a function that calls the Claude API and returns a response through API Gateway.

The short version: a Lambda function makes an HTTPS request to an LLM endpoint, waits for the response, and returns it to the caller. The complexity isn't in the API call itself — it's in managing execution time limits, secrets, retries, and (if you need it) streaming through a request/response model that wasn't built for long-lived connections.

Why Lambda for Claude API calls

Lambda is a good fit when:

It's a worse fit for anything requiring persistent WebSocket connections, long multi-turn conversations held in memory, or sustained streaming to a browser — those need a container, a WebSocket API, or a function with response streaming enabled (more on that below).

Architecture options

API Gateway → Lambda → Claude API The most common setup. A client hits an API Gateway REST or HTTP API endpoint, which triggers Lambda synchronously, waits for the Claude response, and returns it as JSON.

Lambda Function URL Simpler than API Gateway if you just need a public HTTPS endpoint without the extra routing layer. Supports response streaming natively since 2023.

EventBridge/SQS → Lambda (async) For background jobs — summarizing a document after upload, classifying a support ticket — where the caller doesn't need a synchronous response.

Setting up the Lambda function

A minimal handler using Node.js 20.x:

// handler.js
export const handler = async (event) => {
  const body = JSON.parse(event.body || "{}");
  const prompt = body.prompt;

  if (!prompt) {
    return { statusCode: 400, body: JSON.stringify({ error: "prompt is required" }) };
  }

  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-3-5-sonnet",
      max_tokens: 1024,
      messages: [{ role: "user", content: prompt }],
    }),
  });

  if (!response.ok) {
    return { statusCode: response.status, body: await response.text() };
  }

  const data = await response.json();
  return {
    statusCode: 200,
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify(data),
  };
};

This example uses SubToAPI's endpoint, which wraps your existing Claude access into a plain HTTPS API with sub_live_... application keys — convenient in Lambda because you avoid managing OAuth flows or multiple provider SDKs inside the function. Check the quickstart and messages endpoint docs for the full request shape.

Managing API keys securely

Never hardcode keys in the handler. Two realistic options in Lambda:

  1. Environment variables, set via your deployment tool (SAM, CDK, Serverless Framework, Terraform). Fine for most teams, encrypted at rest by default.
  2. AWS Secrets Manager, fetched at cold start and cached in memory for the life of the execution environment. Adds latency on cold start but rotates without a redeploy.
import { SecretsManagerClient, GetSecretValueCommand } from "@aws-sdk/client-secrets-manager";

let cachedKey;

async function getApiKey() {
  if (cachedKey) return cachedKey;
  const client = new SecretsManagerClient({});
  const result = await client.send(
    new GetSecretValueCommand({ SecretId: "subtoapi/prod-key" })
  );
  cachedKey = JSON.parse(result.SecretString).key;
  return cachedKey;
}

If you're running this across a team, use separate keys per function or environment rather than one shared key — SubToAPI's pricing plans include per-seat application keys, which makes it easy to track usage and revoke access without touching other functions.

Timeouts and Lambda limits

Lambda's hard ceiling is 15 minutes, but you should set your function timeout much lower — 30 to 60 seconds for a typical Claude call, longer only if you're generating very long outputs with high max_tokens. Key points:

Streaming responses from Lambda

Standard API Gateway + Lambda integrations buffer the entire response before returning it — no token-by-token streaming. If you need streaming output in a Lambda-based setup:

SubToAPI supports streaming responses over SSE for cases where you control the client directly — see the streaming docs — but inside a traditional REST API Gateway + Lambda pattern, plan for buffered responses unless you specifically set up Function URL streaming.

Deployment snippet (SAM)

Resources:
  ClaudeFunction:
    Type: AWS::Serverless::Function
    Properties:
      Handler: handler.handler
      Runtime: nodejs20.x
      Timeout: 30
      Environment:
        Variables:
          SUBTOAPI_KEY: !Ref SubToApiKey
      Events:
        Api:
          Type: HttpApi
          Properties:
            Path: /chat
            Method: post

Deploy with sam build && sam deploy --guided, then test with curl -X POST <endpoint-url> -d '{"prompt":"Summarize this ticket"}'.

Monitoring usage

Once this is live, you'll want visibility into cost and volume per function — especially if multiple Lambdas across teams are calling the same model. SubToAPI's dashboard reports token usage and request metadata per API key, which is useful when a dozen Lambda functions all share one Claude subscription but need separate accountability.

FAQ

Does Lambda support true token-by-token streaming from Claude? Only through a Lambda Function URL with RESPONSE_STREAM invoke mode, or a WebSocket relay. Standard API Gateway integrations buffer the full response before returning it.

What timeout should I set for a Claude call in Lambda? Start with 30–60 seconds for typical completions, and keep your Lambda timeout slightly higher than your client-side fetch timeout. If generation can run longer, move to an async pattern with SQS or EventBridge instead of raising the Lambda limit.

Where should I store my Claude or SubToAPI key in Lambda? Environment variables are fine for most setups; use AWS Secrets Manager with in-memory caching if you need rotation without redeploying. Avoid hardcoding keys in the handler file under any circumstances.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →