← Blog

Build a Customer Email Classifier with the Claude API

2026-09-30 · 5 min read · SubToAPI Team

Building a customer email classifier with the Claude API means sending inbound email text to Claude and getting back a structured category — support, billing, sales, spam, urgent — that your helpdesk or workflow tool can act on automatically. The core pattern is simple: define a fixed set of labels, force Claude to return them in a predictable JSON shape using tool use, and wire that output into your ticketing system or inbox rules.

This article walks through the whole build: schema design, prompt structure, tool definitions for structured output, routing logic, and the operational details (rate limits, retries, cost) that matter once you move past a proof of concept.

Why classify emails with an LLM instead of keyword rules

Keyword-based filters break on phrasing they haven't seen before. "My card was charged twice" and "duplicate billing charge" mean the same thing but share almost no words. Claude reads for intent, not keywords, so it handles the long tail of real customer language — typos, sarcasm, mixed topics in one email — far better than regex or a bag-of-words classifier, without the training overhead of a custom ML model.

Step 1: Define your label set

Keep categories mutually exclusive and small enough to act on. A typical support inbox works well with:

Add a confidence field and a short reason field so humans can audit misclassifications later — this is the single biggest driver of trust in an automated pipeline.

Step 2: Force structured output with tool use

Free-text responses drift over time. Instead, define a tool schema and require Claude to call it. This guarantees valid JSON every time and removes brittle string parsing from your code.

{
  "name": "classify_email",
  "description": "Classify a customer support email",
  "input_schema": {
    "type": "object",
    "properties": {
      "category": {
        "type": "string",
        "enum": ["billing", "technical", "sales", "spam", "urgent"]
      },
      "confidence": { "type": "number", "minimum": 0, "maximum": 1 },
      "reason": { "type": "string" }
    },
    "required": ["category", "confidence", "reason"]
  }
}

Pass tool_choice set to that tool so Claude can't respond with plain text. See the tools docs for the exact request shape if you're calling through SubToAPI, which mirrors the Anthropic tool-use format.

Step 3: Write a tight system prompt

Resist the urge to over-explain. Give Claude the label definitions, one edge-case rule per category, and nothing else:

You classify inbound customer emails into exactly one category.
Rules:
- If the email mentions a charge, refund, or invoice, it's "billing" even if the tone is angry.
- "urgent" only applies if the sender explicitly requests a response within 24 hours
  or describes a service outage.
- Default to "technical" for bug reports, even vague ones.
- If unsure between two categories, pick the one with higher business impact.
Always call classify_email with your answer.

Long system prompts with dozens of examples tend to reduce accuracy, not improve it — Claude does better with a few sharp rules than a wall of instructions.

Step 4: Send the request

Here's a minimal call against SubToAPI, which exposes your Claude access as a standard HTTPS endpoint with an application key (sub_live_...):

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 256,
    "system": "You classify inbound customer emails...",
    "tools": [ { "name": "classify_email", ... } ],
    "tool_choice": { "type": "tool", "name": "classify_email" },
    "messages": [
      { "role": "user", "content": "Subject: Double charge on my card\n\nHi, I was billed twice this month for the same plan." }
    ]
  }'

The response includes a tool_use content block with the parsed category, confidence, and reason fields, ready to route without any regex.

Step 5: Route based on the classification

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4",
    max_tokens: 256,
    system: SYSTEM_PROMPT,
    tools: [classifyEmailTool],
    tool_choice: { type: "tool", name: "classify_email" },
    messages: [{ role: "user", content: emailBody }],
  }),
});

const data = await res.json();
const result = data.content.find((c) => c.type === "tool_use").input;

if (result.category === "urgent" && result.confidence > 0.8) {
  await pagerDuty.trigger(result);
} else {
  await helpdesk.createTicket({ tag: result.category, note: result.reason });
}

Set a confidence threshold below which tickets go to a human review queue instead of auto-routing. This catches the ambiguous 5-10% of emails without slowing down the other 90%.

Handling volume and cost

Classification is a high-volume, low-token task — most emails classify in under 300 output tokens. Batch requests where possible, and cache the system prompt if your provider supports prompt caching to cut repeated-token cost. Check pricing if you're routing this through SubToAPI; the Solo plan at €9/month is typically enough for a single inbox, while Team plans add seats for shared dashboards and usage visibility across a support team.

For getting started quickly, the quickstart guide covers key setup, and the messages docs detail every request parameter including streaming, which is useful if you later add a live "why was this classified this way" explanation in your support UI — see streaming for that pattern.

Testing and monitoring

Before going live, run the classifier against 100-200 historical emails with known correct labels and measure accuracy per category — spam and urgent tend to have the highest false-positive risk. Log every classification with its confidence score so you can spot drift if email patterns change (a new product launch, a pricing change, a bug spike) and adjust the system prompt accordingly.

Questions

Does Claude need fine-tuning to classify emails accurately? No. A well-structured prompt with tool use and 4-6 clear categories gets strong accuracy out of the box. Fine-tuning only helps once you have thousands of labeled examples in a narrow domain.

How do I stop Claude from inventing new categories? Use a tool schema with an enum field for the category and set tool_choice to force that tool call. This makes invalid categories a schema error, not a possible output.

Can I classify emails with attachments or images? Yes — Claude supports multimodal input, so you can pass image content blocks alongside the email text if attachments matter for classification, though most email triage only needs the subject and body.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →