Building a Claude API Content Moderation Pipeline
What a Claude API content moderation pipeline actually does
A content moderation pipeline built on the Claude API classifies user-generated text (or images) against policy categories — harassment, spam, self-harm, hate speech, explicit content, fraud — and returns a structured decision your application can act on: allow, flag for human review, or block. Unlike a single moderation endpoint that gives you a yes/no answer, Claude lets you encode nuanced, custom policies in a prompt, return structured JSON with confidence scores and reasoning, and adapt the categories as your platform's rules change.
This matters because generic moderation APIs are trained on broad, fixed category sets. If you run a gaming community, a marketplace, or a B2B SaaS with its own terms of service, you often need categories that don't map cleanly onto "toxicity" or "hate speech." Claude's instruction-following lets you define exactly what "allowed" and "not allowed" mean for your product, in your own words, and get consistent structured output back.
Core architecture
A production moderation pipeline has four stages:
- Pre-filter — cheap, fast checks (regex, block lists, length limits) that catch obvious violations without calling the model.
- Classification call — send the content to Claude with a system prompt defining categories and the required output schema.
- Decision logic — map the model's output to an action (allow / flag / block) based on confidence thresholds.
- Escalation and logging — route flagged content to human reviewers and store every decision for audits and appeals.
Only stage 2 needs the API. Keeping stages 1, 3, and 4 outside the model call keeps your pipeline fast and auditable.
Designing the classification prompt
The key to reliable moderation is forcing structured, parseable output. Use tool use (function calling) rather than asking the model to "respond with JSON" in free text — it's far more reliable at the schema boundary. See /docs/tools for the full tool-use reference.
{
"name": "classify_content",
"description": "Classify user content against moderation policy",
"input_schema": {
"type": "object",
"properties": {
"categories": {
"type": "array",
"items": {
"type": "string",
"enum": ["harassment", "hate_speech", "self_harm", "sexual_content", "spam", "fraud", "none"]
}
},
"confidence": { "type": "number" },
"action": {
"type": "string",
"enum": ["allow", "flag", "block"]
},
"reasoning": { "type": "string" }
},
"required": ["categories", "confidence", "action"]
}
}
A system prompt should define each category concretely with examples of borderline cases, because models — like human reviewers — need calibration on edge cases, not just category names.
You are a content moderation classifier for a developer community forum.
Categories:
- harassment: targeted insults, threats, or repeated unwanted contact directed at a specific person
- spam: promotional links, repeated identical posts, engagement bait unrelated to the thread
- fraud: phishing links, fake giveaways, requests for payment info
Borderline cases to treat as "allow":
- Strong technical criticism of a library or company
- Sarcasm about a product, without targeting an individual
- Off-topic jokes that aren't promotional
For each message, call classify_content with your category list, a confidence score (0-1), and an action. Use "flag" for confidence between 0.4 and 0.75, "block" above 0.75, and "allow" below 0.4.
Example request
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 300,
"system": "You are a content moderation classifier...",
"tools": [{"name": "classify_content", "input_schema": {...}}],
"tool_choice": {"type": "tool", "name": "classify_content"},
"messages": [
{"role": "user", "content": "Check out this amazing crypto giveaway, send 0.1 ETH to get 1 ETH back!"}
]
}'
Forcing tool_choice to the classification tool guarantees the model always responds in the expected shape, which removes a whole class of parsing failures from your pipeline.
Threshold tuning and action mapping
Don't hardcode a single global threshold. Different categories carry different risk:
- Fraud and self-harm: bias toward flagging even at lower confidence — false positives are cheaper than false negatives.
- Spam: can tolerate a higher confidence bar before blocking, since the cost of a missed spam post is low.
- Harassment: often needs human review regardless of confidence, because context (relationship between users, prior history) matters more than the text alone.
Store the raw confidence score and category array alongside the action taken, not just the final verdict. You'll need this later to retune thresholds without re-running every past message through the model.
Handling scale and latency
Moderation is usually in the critical path — a user hits "post" and expects a near-instant result. A few practical patterns:
- Pre-filter aggressively. Block-list matches and regex for known spam patterns should never touch the API.
- Batch where you can. If moderation isn't blocking the UI (e.g., background review of uploaded images), batch multiple items into fewer calls.
- Cache repeated content. Identical or near-identical spam gets reposted constantly — hash content and cache recent verdicts.
- Use a smaller, faster model for the first pass, and escalate only ambiguous cases to a larger model for a second opinion.
For teams running this at volume across multiple apps or environments, SubToAPI turns your Claude access into a standard HTTPS API with per-application keys (sub_live_...), so you can isolate your moderation service's usage and rate limits from your main product traffic, and see exactly how many classification calls each app is making from one dashboard. See /docs/quickstart to get a key running in a few minutes, and /docs/messages for the full request reference.
Logging and appeals
Every moderation decision should be reproducible. Log:
- The raw input content (or a hash, if you can't retain the content itself for privacy reasons)
- The full model response, including reasoning
- The model version and prompt version used
- The final action and who/what took it (automated vs. human override)
This gives you an audit trail when a user appeals a block, and lets you measure precision/recall over time by sampling flagged and allowed content for human review.
Streaming isn't usually needed here
Unlike chat or completion features, moderation classification benefits from non-streaming requests — you want the full structured tool call before acting, not partial tokens. If your pipeline also has a user-facing explanation feature ("why was my post removed"), that's a separate streaming call; see /docs/streaming for that pattern.
Questions
Should I build moderation on a general-purpose model or a dedicated moderation endpoint? Use a dedicated moderation endpoint for broad, well-known categories like generic toxicity, since it's cheaper and faster. Use Claude when you need custom policy categories, nuanced context, or structured reasoning your product's rules require.
How do I avoid false positives blocking legitimate users? Set category-specific thresholds, route mid-confidence cases to "flag" instead of "block," and keep a human review queue with override logging so you can retune the prompt based on real mistakes.
Can this pipeline handle images as well as text? Yes — Claude supports image inputs, so you can send user-uploaded images through the same classification tool-use pattern described above, with a system prompt tailored to visual policy categories.