Claude API AWS Lambda Integration Guide
Running Claude inside an AWS Lambda function is a common pattern for serverless chatbots, document processors, webhooks, and scheduled jobs that need an LLM call without maintaining a server. This guide walks through the architecture, the Lambda-specific constraints you'll hit (timeouts, cold starts, streaming), and working code for a function that calls the Claude API and returns a response through API Gateway.
The short version: a Lambda function makes an HTTPS request to an LLM endpoint, waits for the response, and returns it to the caller. The complexity isn't in the API call itself — it's in managing execution time limits, secrets, retries, and (if you need it) streaming through a request/response model that wasn't built for long-lived connections.
Why Lambda for Claude API calls
Lambda is a good fit when:
- Traffic is bursty or unpredictable and you don't want idle compute.
- The call is triggered by an event — an S3 upload, an SQS message, an API Gateway request, a cron schedule.
- You want per-request billing instead of a running server.
It's a worse fit for anything requiring persistent WebSocket connections, long multi-turn conversations held in memory, or sustained streaming to a browser — those need a container, a WebSocket API, or a function with response streaming enabled (more on that below).
Architecture options
API Gateway → Lambda → Claude API The most common setup. A client hits an API Gateway REST or HTTP API endpoint, which triggers Lambda synchronously, waits for the Claude response, and returns it as JSON.
Lambda Function URL Simpler than API Gateway if you just need a public HTTPS endpoint without the extra routing layer. Supports response streaming natively since 2023.
EventBridge/SQS → Lambda (async) For background jobs — summarizing a document after upload, classifying a support ticket — where the caller doesn't need a synchronous response.
Setting up the Lambda function
A minimal handler using Node.js 20.x:
// handler.js
export const handler = async (event) => {
const body = JSON.parse(event.body || "{}");
const prompt = body.prompt;
if (!prompt) {
return { statusCode: 400, body: JSON.stringify({ error: "prompt is required" }) };
}
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-3-5-sonnet",
max_tokens: 1024,
messages: [{ role: "user", content: prompt }],
}),
});
if (!response.ok) {
return { statusCode: response.status, body: await response.text() };
}
const data = await response.json();
return {
statusCode: 200,
headers: { "Content-Type": "application/json" },
body: JSON.stringify(data),
};
};
This example uses SubToAPI's endpoint, which wraps your existing Claude access into a plain HTTPS API with sub_live_... application keys — convenient in Lambda because you avoid managing OAuth flows or multiple provider SDKs inside the function. Check the quickstart and messages endpoint docs for the full request shape.
Managing API keys securely
Never hardcode keys in the handler. Two realistic options in Lambda:
- Environment variables, set via your deployment tool (SAM, CDK, Serverless Framework, Terraform). Fine for most teams, encrypted at rest by default.
- AWS Secrets Manager, fetched at cold start and cached in memory for the life of the execution environment. Adds latency on cold start but rotates without a redeploy.
import { SecretsManagerClient, GetSecretValueCommand } from "@aws-sdk/client-secrets-manager";
let cachedKey;
async function getApiKey() {
if (cachedKey) return cachedKey;
const client = new SecretsManagerClient({});
const result = await client.send(
new GetSecretValueCommand({ SecretId: "subtoapi/prod-key" })
);
cachedKey = JSON.parse(result.SecretString).key;
return cachedKey;
}
If you're running this across a team, use separate keys per function or environment rather than one shared key — SubToAPI's pricing plans include per-seat application keys, which makes it easy to track usage and revoke access without touching other functions.
Timeouts and Lambda limits
Lambda's hard ceiling is 15 minutes, but you should set your function timeout much lower — 30 to 60 seconds for a typical Claude call, longer only if you're generating very long outputs with high max_tokens. Key points:
- Set the Lambda timeout and a client-side fetch timeout slightly shorter than the Lambda timeout, so you get a clean error instead of an abrupt kill.
- API Gateway has its own 29-second timeout for REST/HTTP APIs. If your Claude call can exceed that, go async (SQS/EventBridge) or use a Lambda Function URL instead, which doesn't impose that limit.
- Retry transient errors (429, 503) with exponential backoff, but cap retries — a Lambda retrying indefinitely inside a 15-minute window costs real money.
Streaming responses from Lambda
Standard API Gateway + Lambda integrations buffer the entire response before returning it — no token-by-token streaming. If you need streaming output in a Lambda-based setup:
- Use a Lambda Function URL with response streaming (
InvokeMode: RESPONSE_STREAM), available for Node.js runtimes. - Or terminate streaming at the Lambda boundary and relay via WebSocket API Gateway, forwarding chunks as they arrive.
SubToAPI supports streaming responses over SSE for cases where you control the client directly — see the streaming docs — but inside a traditional REST API Gateway + Lambda pattern, plan for buffered responses unless you specifically set up Function URL streaming.
Deployment snippet (SAM)
Resources:
ClaudeFunction:
Type: AWS::Serverless::Function
Properties:
Handler: handler.handler
Runtime: nodejs20.x
Timeout: 30
Environment:
Variables:
SUBTOAPI_KEY: !Ref SubToApiKey
Events:
Api:
Type: HttpApi
Properties:
Path: /chat
Method: post
Deploy with sam build && sam deploy --guided, then test with curl -X POST <endpoint-url> -d '{"prompt":"Summarize this ticket"}'.
Monitoring usage
Once this is live, you'll want visibility into cost and volume per function — especially if multiple Lambdas across teams are calling the same model. SubToAPI's dashboard reports token usage and request metadata per API key, which is useful when a dozen Lambda functions all share one Claude subscription but need separate accountability.
FAQ
Does Lambda support true token-by-token streaming from Claude? Only through a Lambda Function URL with RESPONSE_STREAM invoke mode, or a WebSocket relay. Standard API Gateway integrations buffer the full response before returning it.
What timeout should I set for a Claude call in Lambda? Start with 30–60 seconds for typical completions, and keep your Lambda timeout slightly higher than your client-side fetch timeout. If generation can run longer, move to an async pattern with SQS or EventBridge instead of raising the Lambda limit.
Where should I store my Claude or SubToAPI key in Lambda? Environment variables are fine for most setups; use AWS Secrets Manager with in-memory caching if you need rotation without redeploying. Avoid hardcoding keys in the handler file under any circumstances.