Claude API in AWS Lambda: Serverless Functions Guide
Can you call Claude from AWS Lambda?
Yes. Calling the Claude API from an AWS Lambda function works the same way as calling any other HTTPS API: you send a POST request with an API key, get a response, and return it to whatever triggered the function — API Gateway, EventBridge, an SQS queue, or another Lambda. There's no special SDK requirement for Lambda specifically. The parts that actually need attention are Lambda's execution model: cold starts, timeout limits, response streaming, and how you manage secrets and concurrency when your function fans out to hundreds of concurrent invocations.
This article covers the practical setup: writing a Lambda handler that calls Claude, handling timeouts and retries correctly, dealing with streaming responses (which Lambda doesn't support well out of the box), and managing API keys and cost when your Lambda scales up. We'll use plain Node.js examples that work with any HTTPS-compatible Claude API, including a Claude subscription wrapped as an API through SubToAPI.
Basic Lambda handler calling Claude
A minimal handler that calls a Messages-style endpoint looks like this:
// index.mjs
export const handler = async (event) => {
const body = JSON.parse(event.body || "{}");
const userMessage = body.message || "Hello";
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [{ role: "user", content: userMessage }]
})
});
if (!response.ok) {
const errText = await response.text();
return { statusCode: response.status, body: errText };
}
const data = await response.json();
return {
statusCode: 200,
body: JSON.stringify(data)
};
};
A few Lambda-specific details matter here:
- Node 18+ runtime gives you native
fetch, so you don't need to bundle a HTTP client. - Environment variables hold the API key — never hardcode it in the handler.
- Timeout on the Lambda function needs to exceed the expected Claude response time. Default Lambda timeout is 3 seconds, which is too short for most completions. Set it to at least 30–60 seconds depending on your prompts and
max_tokens.
Setting the timeout and memory correctly
Lambda's default configuration will kill your function before Claude finishes responding. In your template.yaml (SAM) or CDK stack:
Resources:
ClaudeFunction:
Type: AWS::Serverless::Function
Properties:
Handler: index.handler
Runtime: nodejs20.x
Timeout: 60
MemorySize: 512
Environment:
Variables:
SUBTOAPI_KEY: !Ref SubToApiKeyParam
Memory doesn't affect Claude call speed much since the function is mostly waiting on network I/O, but 512 MB is a reasonable default that avoids CPU throttling during JSON parsing of large responses.
If you're calling Claude synchronously behind API Gateway, remember API Gateway has its own 29-second timeout for REST APIs (HTTP APIs allow longer). For anything that might take longer than that — long prompts, large max_tokens, or tool-use loops — put the Lambda behind an async pattern instead: API Gateway triggers Lambda, Lambda writes to SQS or invokes itself asynchronously, and the client polls or gets a webhook callback.
Handling streaming in a serverless function
Streaming is the trickiest part of running Claude in Lambda. Traditional Lambda invoked through API Gateway buffers the entire response — it can't stream tokens back to the client incrementally. If your use case needs token-by-token output, you have two options:
- Lambda Function URLs with response streaming (
InvokeMode: RESPONSE_STREAM), available for Node.js and supported by AWS since 2023. This lets you pipe a streaming HTTP response from Claude directly through Lambda to the client. - Skip streaming in Lambda entirely and return the full response once complete, handling any real-time UI needs on a different compute layer (a small always-on server, or client-side polling against a job status endpoint).
If you go with option 1, the streaming setup mirrors SubToAPI's streaming endpoint, which returns server-sent events you can forward chunk by chunk. For most backend automation use cases — summarization jobs, data enrichment pipelines, scheduled reports — you don't need streaming at all, and non-streaming Lambda is simpler and more reliable.
Managing concurrency and rate limits
Lambda scales by spinning up concurrent instances automatically, which is great for throughput but can slam your Claude API key with more concurrent requests than expected. If 500 Lambda invocations fire at once, you'll get 500 simultaneous API calls unless you throttle.
Practical mitigations:
- Set a reserved concurrency limit on the Lambda function so it can't scale past what your API key's rate limit supports.
- Use SQS as a buffer in front of Lambda with a controlled batch size, so requests to Claude are paced rather than bursty.
- Implement exponential backoff on 429 responses in your handler code, since transient rate limiting is expected under load.
async function callClaudeWithRetry(payload, attempt = 1) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify(payload)
});
if (res.status === 429 && attempt <= 3) {
await new Promise(r => setTimeout(r, attempt * 1000));
return callClaudeWithRetry(payload, attempt + 1);
}
return res.json();
}
Why route through an API layer instead of raw model access
If your team already pays for Claude through a subscription rather than direct API credits, calling it from Lambda usually means someone has to build the auth translation layer, track usage per function, and rotate keys safely across environments (dev, staging, prod). SubToAPI handles that layer: it issues sub_live_... application keys per environment or per team member, so your Lambda functions in staging and production use separate, revocable keys without anyone sharing a personal login. You get usage metadata per request, which is useful for attributing Claude spend to a specific Lambda function or customer when you're running a fan-out job. Setup takes a few minutes — see the quickstart — and pricing starts at €9/month on the Solo plan, scaling to team seats via pricing if multiple developers or services need their own keys.
FAQs
Does Lambda support streaming responses from Claude? Only through Lambda Function URLs configured with RESPONSE_STREAM invoke mode. Standard API Gateway-triggered Lambdas buffer the full response before returning it, so token-by-token streaming isn't possible in that setup.
What Lambda timeout should I set for Claude API calls? At least 30–60 seconds for typical completions, adjusted for your max_tokens and prompt size. Also check any upstream timeout, like API Gateway's 29-second limit on REST APIs, which can cut off a slower Lambda invocation.
How do I avoid hitting rate limits when Lambda scales up? Set reserved concurrency on the function, buffer bursts through SQS, and implement retry-with-backoff for 429 responses. This keeps concurrent Claude calls within your API key's actual rate limit instead of scaling unchecked with Lambda's default behavior.