← Blog

Claude API Google Cloud Functions Deployment

2026-10-03 · 5 min read · SubToAPI Team

If you're looking to run Claude API calls behind a Google Cloud Functions endpoint, the core setup is straightforward: write an HTTP-triggered function, call the Anthropic API (or a gateway) from inside it, manage your API key as a secret, and handle Cloud Functions' execution limits around timeouts and streaming. The tricky parts aren't the code — they're the platform constraints that specifically affect long-running or streamed LLM responses.

This guide walks through a working deployment, the gotchas that trip people up (timeouts, cold starts, concurrent connections, and streaming support), and how to structure the function so it's maintainable once you have more than one endpoint calling Claude.

Why Cloud Functions for Claude API calls

Google Cloud Functions (2nd gen, built on Cloud Run) is a reasonable choice when you want a lightweight, stateless proxy between your frontend and Claude — for example, a form submission that triggers a Claude completion, a webhook handler that summarizes incoming data, or a scheduled job that generates reports. It scales to zero, you pay per invocation, and you don't manage servers.

It's a worse fit if you need persistent WebSocket-style streaming to the browser, very long-running agentic workflows, or sub-100ms latency — Cloud Functions has cold starts and a hard execution timeout (60 minutes max on 2nd gen with Cloud Run, but default and practical limits are much lower for HTTP functions).

Basic deployment

Here's a minimal Node.js Cloud Function that calls Claude's Messages API:

const functions = require('@google-cloud/functions-framework');

functions.http('claudeHandler', async (req, res) => {
  const { prompt } = req.body;

  const response = await fetch('https://api.anthropic.com/v1/messages', {
    method: 'POST',
    headers: {
      'x-api-key': process.env.ANTHROPIC_API_KEY,
      'anthropic-version': '2023-06-01',
      'content-type': 'application/json'
    },
    body: JSON.stringify({
      model: 'claude-sonnet-4-5',
      max_tokens: 1024,
      messages: [{ role: 'user', content: prompt }]
    })
  });

  const data = await response.json();
  res.status(200).json(data);
});

Deploy it with:

gcloud functions deploy claudeHandler \
  --gen2 \
  --runtime=nodejs20 \
  --trigger-http \
  --allow-unauthenticated \
  --entry-point=claudeHandler \
  --set-secrets=ANTHROPIC_API_KEY=anthropic-key:latest

Note the --set-secrets flag — never hardcode the API key or pass it as a plain environment variable in your gcloud deploy command, since that ends up in deployment logs. Store it in Secret Manager first:

echo -n "sk-ant-..." | gcloud secrets create anthropic-key --data-file=-

Handling timeouts and streaming

This is where most Cloud Functions + Claude integrations break down. Claude responses — especially with longer max_tokens or tool use — can take well beyond the default function timeout (60 seconds for HTTP functions unless you raise it).

Two practical fixes:

  1. Raise the timeout explicitly for Cloud Run-backed 2nd gen functions:
gcloud functions deploy claudeHandler --timeout=540s ...
  1. Avoid streaming through a Cloud Function if you can. HTTP functions buffer responses by default, which means stream: true from Claude doesn't translate cleanly into chunked responses back to the browser. You can stream function-to-function (Cloud Function to Claude) and write chunks to the response with res.write(), but you need --gen2 and careful header management (Transfer-Encoding: chunked), and behavior varies depending on whether a load balancer or Cloud CDN sits in front.

If streaming to the end user is a hard requirement, Cloud Run (not Cloud Functions) gives you more control over long-lived HTTP connections. Cloud Functions is better suited to request/response patterns where you wait for the full Claude completion and return it as JSON.

Cold starts and concurrency

Cloud Functions cold starts add 1-3 seconds on top of Claude's own latency, which is noticeable for user-facing endpoints. Minimize the impact by:

Where an API gateway simplifies things

A lot of the complexity in these deployments — key rotation, per-environment credentials, usage tracking across functions, and consistent error handling — disappears if you put a managed layer between your Cloud Function and Claude instead of calling the Anthropic API directly.

This is the gap SubToAPI fills: it turns your existing Claude access into a standard HTTPS API with application-level keys (sub_live_...), so your Cloud Function code doesn't hold a raw Anthropic key at all. You generate a scoped key per project or environment, call a single stable endpoint, and get usage metadata back with every response — useful when you have five Cloud Functions each calling Claude and want one dashboard to see what's actually being spent.

Swapping the fetch target is a one-line change:

const response = await fetch('https://api.subtoapi.app/v1/messages', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
    'content-type': 'application/json'
  },
  body: JSON.stringify({
    model: 'claude-sonnet-4-5',
    max_tokens: 1024,
    messages: [{ role: 'user', content: prompt }]
  })
});

Setup takes a few minutes — see the quickstart and Messages API reference. If you need to stream responses from a different compute layer (like Cloud Run instead of Cloud Functions), the streaming docs cover the exact headers and chunk format. There's a free trial at signup, and plans start at €9/month on the pricing page.

Production checklist

Before shipping a Cloud Function that calls Claude:

questions

Can Cloud Functions handle Claude's streaming responses? Partially. 2nd gen HTTP functions can write chunked responses with res.write(), but behavior is inconsistent behind load balancers. For true token-by-token streaming to a browser, Cloud Run gives more reliable control over long-lived connections.

What timeout should I set for a Claude API call in Cloud Functions? Start with 60-120 seconds for typical completions under 1024 tokens, and go up to 300-540 seconds if you're using large max_tokens values or multi-step tool use. Always set it explicitly rather than relying on defaults.

How do I avoid storing my Anthropic API key in plaintext during deployment? Use Google Secret Manager with --set-secrets during gcloud functions deploy, or route calls through a gateway like SubToAPI so your function only holds a scoped application key instead of your raw provider credential.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →