Claude API Serverless Function Deployment on AWS
Deploying Claude API calls as AWS Lambda functions is straightforward once you account for three constraints that trip up most first attempts: execution timeouts, the fact that Lambda's native HTTP response model doesn't stream well, and how you manage API keys without hardcoding secrets into your deployment package.
The short answer: write a thin Lambda handler that calls the Anthropic Messages API (or a managed layer like SubToAPI), pull your key from environment variables backed by AWS Secrets Manager or SSM Parameter Store, set function timeout to at least 30-60 seconds for non-streaming calls, and put API Gateway or a Function URL in front of it. Streaming responses need a different approach since Lambda buffers output by default — more on that below.
Why Serverless Makes Sense for Claude API Workloads
Lambda is a good fit for Claude API calls when traffic is bursty or unpredictable — a chatbot backend, a document summarizer triggered by S3 uploads, or an internal tool used sporadically by a small team. You pay per invocation instead of running an idle server, and AWS handles scaling automatically when traffic spikes.
It's a worse fit for high-throughput, latency-sensitive streaming applications where cold starts and the 15-minute hard timeout become real constraints. Know which category your workload falls into before you commit to the architecture.
Basic Lambda Handler Structure
A minimal Node.js handler that calls Claude looks like this:
const https = require("https");
exports.handler = async (event) => {
const body = JSON.parse(event.body || "{}");
const payload = JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{ role: "user", content: body.prompt }],
});
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: payload,
});
const data = await response.json();
return {
statusCode: 200,
headers: { "content-type": "application/json" },
body: JSON.stringify(data),
};
};
This works, but production functions need timeout handling, retry logic for transient errors, and a secrets strategy. Lambda's default Node.js runtime includes fetch, so you don't need to bundle an HTTP client for simple calls.
Managing API Keys in a Serverless Deployment
Never commit API keys to your deployment package or set them as plaintext Lambda environment variables in source control. Two reasonable approaches:
- AWS Secrets Manager — fetch the key at cold start and cache it in a module-level variable so you're not calling Secrets Manager on every invocation.
- SSM Parameter Store with SecureString — cheaper than Secrets Manager for low-volume use, same caching pattern applies.
const { SecretsManagerClient, GetSecretValueCommand } = require("@aws-sdk/client-secrets-manager");
let cachedKey;
async function getApiKey() {
if (cachedKey) return cachedKey;
const client = new SecretsManagerClient({});
const result = await client.send(
new GetSecretValueCommand({ SecretId: "claude-api-key" })
);
cachedKey = result.SecretString;
return cachedKey;
}
If your team manages multiple applications each needing their own key, scoping, and usage tracking, generating a separate Anthropic key per app and rotating them manually gets tedious fast. SubToAPI issues sub_live_... application keys from your existing Claude access, so each Lambda function or microservice gets its own scoped key with usage visible in one dashboard — useful when you're running several serverless functions against the same underlying Claude account. See the quickstart for the setup flow.
Handling Streaming in Lambda
Standard Lambda invocations buffer the entire response before returning it to the caller, which defeats the purpose of streaming. If your use case needs token-by-token output in the browser, you have two real options on AWS:
- Lambda response streaming via Function URLs — AWS supports
awslambda.streamifyResponse()for functions invoked through a Function URL, which lets you forward Server-Sent Events as they arrive from the Claude API. - Skip streaming at the Lambda layer — return the complete response and let the client poll or use WebSockets through API Gateway if partial output matters.
exports.handler = awslambda.streamifyResponse(
async (event, responseStream) => {
const upstream = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
stream: true,
messages: [{ role: "user", content: "Summarize this article." }],
}),
});
for await (const chunk of upstream.body) {
responseStream.write(chunk);
}
responseStream.end();
}
);
This requires Node.js 18+ runtime and a Function URL configured with RESPONSE_STREAM invoke mode. API Gateway REST and HTTP APIs don't support true streaming to the client, so Function URLs are currently the practical path for this pattern on AWS.
Timeouts, Cold Starts, and Cost Control
Set your Lambda timeout based on your actual use case, not the default 3 seconds. Non-streaming Claude calls with longer outputs can take 20-40+ seconds depending on max_tokens and model. Set the timeout to a comfortable margin above your expected p99 latency, and configure your downstream client (API Gateway, ALB) timeout to match or exceed it — API Gateway's hard cap is 29 seconds for REST APIs, which is a common gotcha for longer generations.
Cold starts matter less for Claude workloads than you'd think, since the network call to the Claude API dominates total latency anyway. If cold starts are a problem for a latency-sensitive endpoint, provisioned concurrency is the fix, at the cost of paying for idle capacity.
For deploying the function itself, the Serverless Framework, AWS SAM, or AWS CDK all work fine — none of them change anything specific to the Claude API integration, so pick whichever your team already uses.
Testing Before You Deploy
Test your Claude integration logic locally before wiring it into Lambda — debugging API errors through CloudWatch logs is slower than iterating against a local script. Confirm your request format, error handling, and retry logic work against the live API first, then wrap it in the handler. The Messages API docs are a good reference for request shape regardless of which layer you're calling through.
FAQ
Can I stream Claude API responses from a Lambda function to a browser? Yes, using Lambda response streaming (awslambda.streamifyResponse) with a Function URL set to RESPONSE_STREAM invoke mode. Standard API Gateway integrations buffer the full response and won't stream token-by-token output.
What Lambda timeout should I set for Claude API calls? Set it comfortably above your expected response time — often 30-60 seconds for non-trivial outputs. Remember API Gateway REST APIs cap at 29 seconds regardless of your Lambda timeout, so use a Function URL or ALB if you need longer.
Where should I store my Claude API key in an AWS serverless deployment? Use AWS Secrets Manager or SSM Parameter Store with SecureString, fetched and cached at cold start — never hardcode keys or commit them as plaintext environment variables in your deployment config.