← Blog

Claude API WhatsApp Business Integration Guide

2026-10-03 · 5 min read · SubToAPI Team

Connecting Claude to WhatsApp Business lets you build an AI assistant that answers customer messages, handles support requests, or qualifies leads directly inside the chat app your customers already use. There is no native "Claude for WhatsApp" product — you build the integration yourself by wiring the WhatsApp Business Cloud API (or a BSP like Twilio/360dialog) to a backend that calls Claude, then sends the reply back through WhatsApp.

This guide covers the actual architecture: receiving WhatsApp messages via webhook, calling Claude with the right context, formatting replies so they render correctly in WhatsApp, and managing conversation state per contact. It also covers the practical gotchas — message length limits, 24-hour session windows, and rate limiting — that trip people up when they move from a prototype to production.

How the integration works end to end

The flow has four parts:

  1. WhatsApp sends you a webhook when a user messages your business number
  2. Your server parses the payload, extracts the message text/media and sender ID
  3. Your server calls Claude with the message plus any relevant conversation history
  4. Your server sends Claude's reply back to WhatsApp via the Send Message API

Claude never talks to WhatsApp directly. You're always the middleman, which means you're also responsible for rate limiting, retries, logging, and deciding how much conversation history to send with each request.

Setting up the WhatsApp side

If you're using Meta's Cloud API directly, you need:

If you'd rather skip Meta's verification process, Twilio and 360dialog offer the same functionality as a managed layer with simpler onboarding — the webhook/send-message pattern is the same either way, just with a different SDK.

Handling the incoming webhook

A typical webhook payload looks like this (simplified Meta Cloud API format):

{
  "entry": [{
    "changes": [{
      "value": {
        "messages": [{
          "from": "15551234567",
          "id": "wamid.xxx",
          "text": { "body": "What's the status of my order #4821?" },
          "type": "text"
        }]
      }
    }]
  }]
}

Your server extracts the sender's phone number and message body, then calls Claude:

app.post("/webhook/whatsapp", async (req, res) => {
  const msg = req.body.entry[0].changes[0].value.messages?.[0];
  if (!msg) return res.sendStatus(200);

  const userId = msg.from;
  const userText = msg.text.body;

  const reply = await callClaude(userId, userText);
  await sendWhatsAppMessage(userId, reply);

  res.sendStatus(200);
});

Respond to WhatsApp's webhook with a 200 immediately — don't wait for Claude's response to finish before acknowledging, or WhatsApp will retry the webhook and you'll process the same message twice.

Calling Claude with conversation context

WhatsApp doesn't give you built-in conversation memory — you have to store and replay it yourself, keyed by the user's phone number. A simple approach is a per-user message array in Redis or Postgres, capped at the last N turns:

async function callClaude(userId, userText) {
  const history = await getHistory(userId); // array of {role, content}

  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 600,
      system: "You are a support assistant for Acme Corp. Keep replies under 300 characters when possible — they are sent over WhatsApp.",
      messages: [...history, { role: "user", content: userText }]
    })
  });

  const data = await response.json();
  const reply = data.content[0].text;

  await appendHistory(userId, userText, reply);
  return reply;
}

Using SubToAPI here means you get a single sub_live_ API key, usage metadata per request, and a dashboard to watch spend across all your WhatsApp traffic — useful once you have more than one bot or more than one person on the team managing it. Setup is the same /v1/messages call shown in /docs/messages, so if you already have Claude API code elsewhere, migrating is a header change, not a rewrite.

Formatting replies for WhatsApp

WhatsApp doesn't render Markdown the way Claude outputs it by default. It supports a limited subset: bold, _italic_, ~strikethrough~, and monospace with triple backticks — no headers, no numbered lists with proper indentation, no tables. Tell Claude about this constraint directly in the system prompt instead of post-processing its output:

Format replies using only WhatsApp-supported markup: *bold*, _italic_, and plain line breaks for lists. Never use markdown headers, tables, or nested bullet points.

This is more reliable than trying to strip unsupported syntax after the fact, and it avoids broken-looking messages reaching customers.

Rate limits and the 24-hour window

Two WhatsApp-specific constraints matter for the Claude side:

Build retry logic that distinguishes between "Claude is rate-limited" and "WhatsApp rejected the send" — they need different backoff strategies. If you're calling Claude at meaningful volume, check /docs/streaming and /docs/quickstart for request patterns, though note that WhatsApp's Send Message API is not streaming-compatible — you always need the full response before sending, so streaming only helps if you're also showing a typing indicator in a web dashboard alongside the WhatsApp thread.

Adding business logic with tool use

Most useful WhatsApp bots need to do more than chat — look up an order, check inventory, escalate to a human. Claude's tool use lets you define functions (e.g., get_order_status, create_support_ticket) that Claude calls when the conversation needs them, rather than hardcoding intent detection. See /docs/tools for the request format — the pattern is identical whether the end channel is WhatsApp, web chat, or email.

Getting started

If you're prototyping this integration, start with the webhook and a hardcoded test number before connecting real traffic. Get the message flow working end to end, then layer in conversation history, then formatting rules, then tool use. You can sign up at /signup for a SubToAPI key and test the /v1/messages endpoint directly with curl before writing any webhook code — confirm the Claude side works in isolation first. Pricing for ongoing use is on /pricing, starting at the Solo plan for single-number bots and scaling to per-seat Team/Scale plans once multiple people manage the integration.

FAQs

Can Claude send messages to WhatsApp directly, without a backend? No. Claude only responds to API requests — it has no outbound connection to WhatsApp. You need a server that receives WhatsApp webhooks, calls Claude, and sends the reply back via WhatsApp's Send Message API.

Does WhatsApp support streaming Claude responses? No. The WhatsApp Send Message API requires a complete message body, so you must wait for Claude's full response before sending. Streaming is only useful if you're mirroring the conversation in a separate interface, like an internal dashboard.

How do I keep conversation context across multiple WhatsApp messages? Store a per-user message history (keyed by phone number) in your own database and include the recent turns in each Claude request. WhatsApp itself doesn't track conversation state for you.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →