← Blog

Building a Claude API Content Moderation Pipeline

2026-10-07 · 5 min read · SubToAPI Team

What a Claude API content moderation pipeline actually does

A content moderation pipeline built on the Claude API classifies user-generated text (or images) against policy categories — harassment, spam, self-harm, hate speech, explicit content, fraud — and returns a structured decision your application can act on: allow, flag for human review, or block. Unlike a single moderation endpoint that gives you a yes/no answer, Claude lets you encode nuanced, custom policies in a prompt, return structured JSON with confidence scores and reasoning, and adapt the categories as your platform's rules change.

This matters because generic moderation APIs are trained on broad, fixed category sets. If you run a gaming community, a marketplace, or a B2B SaaS with its own terms of service, you often need categories that don't map cleanly onto "toxicity" or "hate speech." Claude's instruction-following lets you define exactly what "allowed" and "not allowed" mean for your product, in your own words, and get consistent structured output back.

Core architecture

A production moderation pipeline has four stages:

  1. Pre-filter — cheap, fast checks (regex, block lists, length limits) that catch obvious violations without calling the model.
  2. Classification call — send the content to Claude with a system prompt defining categories and the required output schema.
  3. Decision logic — map the model's output to an action (allow / flag / block) based on confidence thresholds.
  4. Escalation and logging — route flagged content to human reviewers and store every decision for audits and appeals.

Only stage 2 needs the API. Keeping stages 1, 3, and 4 outside the model call keeps your pipeline fast and auditable.

Designing the classification prompt

The key to reliable moderation is forcing structured, parseable output. Use tool use (function calling) rather than asking the model to "respond with JSON" in free text — it's far more reliable at the schema boundary. See /docs/tools for the full tool-use reference.

{
  "name": "classify_content",
  "description": "Classify user content against moderation policy",
  "input_schema": {
    "type": "object",
    "properties": {
      "categories": {
        "type": "array",
        "items": {
          "type": "string",
          "enum": ["harassment", "hate_speech", "self_harm", "sexual_content", "spam", "fraud", "none"]
        }
      },
      "confidence": { "type": "number" },
      "action": {
        "type": "string",
        "enum": ["allow", "flag", "block"]
      },
      "reasoning": { "type": "string" }
    },
    "required": ["categories", "confidence", "action"]
  }
}

A system prompt should define each category concretely with examples of borderline cases, because models — like human reviewers — need calibration on edge cases, not just category names.

You are a content moderation classifier for a developer community forum.

Categories:
- harassment: targeted insults, threats, or repeated unwanted contact directed at a specific person
- spam: promotional links, repeated identical posts, engagement bait unrelated to the thread
- fraud: phishing links, fake giveaways, requests for payment info

Borderline cases to treat as "allow":
- Strong technical criticism of a library or company
- Sarcasm about a product, without targeting an individual
- Off-topic jokes that aren't promotional

For each message, call classify_content with your category list, a confidence score (0-1), and an action. Use "flag" for confidence between 0.4 and 0.75, "block" above 0.75, and "allow" below 0.4.

Example request

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 300,
    "system": "You are a content moderation classifier...",
    "tools": [{"name": "classify_content", "input_schema": {...}}],
    "tool_choice": {"type": "tool", "name": "classify_content"},
    "messages": [
      {"role": "user", "content": "Check out this amazing crypto giveaway, send 0.1 ETH to get 1 ETH back!"}
    ]
  }'

Forcing tool_choice to the classification tool guarantees the model always responds in the expected shape, which removes a whole class of parsing failures from your pipeline.

Threshold tuning and action mapping

Don't hardcode a single global threshold. Different categories carry different risk:

Store the raw confidence score and category array alongside the action taken, not just the final verdict. You'll need this later to retune thresholds without re-running every past message through the model.

Handling scale and latency

Moderation is usually in the critical path — a user hits "post" and expects a near-instant result. A few practical patterns:

For teams running this at volume across multiple apps or environments, SubToAPI turns your Claude access into a standard HTTPS API with per-application keys (sub_live_...), so you can isolate your moderation service's usage and rate limits from your main product traffic, and see exactly how many classification calls each app is making from one dashboard. See /docs/quickstart to get a key running in a few minutes, and /docs/messages for the full request reference.

Logging and appeals

Every moderation decision should be reproducible. Log:

This gives you an audit trail when a user appeals a block, and lets you measure precision/recall over time by sampling flagged and allowed content for human review.

Streaming isn't usually needed here

Unlike chat or completion features, moderation classification benefits from non-streaming requests — you want the full structured tool call before acting, not partial tokens. If your pipeline also has a user-facing explanation feature ("why was my post removed"), that's a separate streaming call; see /docs/streaming for that pattern.

Questions

Should I build moderation on a general-purpose model or a dedicated moderation endpoint? Use a dedicated moderation endpoint for broad, well-known categories like generic toxicity, since it's cheaper and faster. Use Claude when you need custom policy categories, nuanced context, or structured reasoning your product's rules require.

How do I avoid false positives blocking legitimate users? Set category-specific thresholds, route mid-confidence cases to "flag" instead of "block," and keep a human review queue with override logging so you can retune the prompt based on real mistakes.

Can this pipeline handle images as well as text? Yes — Claude supports image inputs, so you can send user-uploaded images through the same classification tool-use pattern described above, with a system prompt tailored to visual policy categories.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →