Claude API Content Moderation Filter Setup Guide
Setting up content moderation with the Claude API means building a layer that classifies user input and/or model output against your own policy categories — before or after the main generation call. Claude has baked-in safety behavior (it will refuse clearly harmful requests on its own), but that's not the same as a moderation filter you control: you need something that returns a structured verdict your application can act on, logs decisions, and lets you tune thresholds per category.
This guide covers the practical setup: how to use Claude as a classifier, how to structure the prompt so you get reliable JSON back, how to handle streaming responses, and where a second layer (like a managed gateway) helps with tracking flagged requests across a team.
Why built-in safety isn't a moderation filter
Claude's training makes it refuse or soften responses to clearly dangerous requests — violence, illegal activity, self-harm instructions, and similar categories. That protects the model itself, but it doesn't give you:
- A machine-readable verdict (allow/flag/block) you can branch on in code
- Custom categories specific to your product (spam, harassment between users, off-topic abuse, brand safety)
- Confidence scores or severity levels for a review queue
- A pre-check that runs before you spend tokens on a full generation
If your app lets users submit free text — support tickets, chat messages, community posts — you almost always want an explicit moderation step, separate from whatever the main model call does.
Step 1: Define your categories first
Before writing any code, write down the categories you actually care about. Vague categories produce vague model behavior. A workable starting set:
harassment— targeted abuse, threats, hate speechself_harm— content indicating risk to the user or otherssexual_content— explicit material not allowed in your contextspam— promotional or repetitive low-value contentpii_leak— content containing personal data that shouldn't be stored or echoed
Keep the list short. Five to eight categories is easier to get consistent classification on than twenty.
Step 2: Build a classification prompt
Use a separate, cheap call to classify input before it reaches your main prompt. Ask Claude to return strict JSON so you can parse it without guesswork:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-4-5",
"max_tokens": 200,
"system": "You are a content moderation classifier. Given user text, return ONLY valid JSON with this shape: {\"flagged\": boolean, \"categories\": {\"harassment\": number, \"self_harm\": number, \"sexual_content\": number, \"spam\": number}, \"reason\": string}. Scores are 0.0-1.0 confidence. No prose outside the JSON.",
"messages": [
{"role": "user", "content": "<insert user-submitted text here>"}
]
}'
Use a fast, cheap model for this step — classification doesn't need your best model, and running it on every submission adds up in cost and latency. Set max_tokens low since you only need a short JSON object back.
Parse the response defensively: strip markdown fences if present, validate with a schema, and fall back to "flag for human review" if parsing fails rather than silently allowing content through.
function parseVerdict(text) {
const cleaned = text.replace(/```json|```/g, "").trim();
try {
const verdict = JSON.parse(cleaned);
if (typeof verdict.flagged !== "boolean") throw new Error("bad shape");
return verdict;
} catch {
return { flagged: true, categories: {}, reason: "parse_error" };
}
}
Step 3: Set thresholds, not just booleans
Don't rely purely on the flagged boolean — use the per-category scores to decide the action:
- Score < 0.3: allow, no logging needed
- 0.3–0.7: allow but log for periodic review
- > 0.7: block or route to a human moderator
This lets you tune sensitivity per category without re-prompting the model. Harassment might need a lower block threshold than spam, for example.
Step 4: Moderate output, not just input
Input moderation catches bad prompts; it doesn't catch cases where a legitimate prompt produces a problematic response (rare with Claude, but possible with creative or long-form generation). Run a lightweight output check the same way, using the model's actual response as the input to your classifier prompt. For latency-sensitive apps, do this only for categories where false negatives are costly — you don't need to re-check every response for spam.
Handling moderation with streaming
If your main generation uses streaming (see /docs/streaming), you can't wait for the full response before showing anything to the user, but you also don't want to stream content that violates policy. Two practical patterns:
- Buffer-and-check in chunks: accumulate text in ~200-character windows, run a fast heuristic (keyword list, regex) between classifier calls, and only invoke the full classifier if the heuristic trips.
- Pre-check only: classify the prompt before generation starts, and accept that output-level issues are rare enough to catch via periodic sampling rather than real-time interruption.
Most production setups use pattern 2 for cost and latency reasons, reserving full re-classification for flagged conversations.
Where a gateway layer helps
Once moderation is running in production, you need visibility: which app or team generated the most flagged requests, how moderation calls are eating into your token budget, and who has access to adjust thresholds. This is where a thin API layer on top of your Claude access pays off. SubToAPI turns your existing Claude access into application-specific keys (sub_live_...) with usage metadata per key, so you can run your moderation classifier under its own key, track its cost separately from your main generation traffic, and give team members scoped access without sharing raw credentials. Check /docs/messages for the request format and /pricing for plan details if you're running moderation across multiple apps or team seats.
Keep a human in the loop
Automated moderation reduces volume, it doesn't eliminate judgment calls. Build a simple review queue for anything scoring in the 0.3–0.7 band, store the original text, the verdict, and the category scores, and revisit your thresholds monthly based on what reviewers actually overturn.
questions
Does Claude's API block harmful content automatically? Claude will refuse or soften responses to clearly harmful requests based on its training, but this isn't a configurable filter — you don't get scores, categories, or a verdict you can branch on in code. For that, build a separate classification call.
Should I use the same model for moderation and main generation? No. Use a fast, inexpensive model for classification since it only needs to return a short structured verdict, and reserve your main model for the actual task. This keeps moderation cost and latency low even at high volume.
How do I moderate streamed responses without delaying output? Classify the input prompt before streaming starts, then use lightweight heuristics (keyword or regex checks) on output chunks, only invoking the full classifier if a heuristic trips. Full real-time classification per chunk is usually too slow and costly for production use.