← Blog

Claude API Content Moderation Filter Setup Guide

2026-10-02 · 5 min read · SubToAPI Team

Setting up content moderation with the Claude API means building a layer that classifies user input and/or model output against your own policy categories — before or after the main generation call. Claude has baked-in safety behavior (it will refuse clearly harmful requests on its own), but that's not the same as a moderation filter you control: you need something that returns a structured verdict your application can act on, logs decisions, and lets you tune thresholds per category.

This guide covers the practical setup: how to use Claude as a classifier, how to structure the prompt so you get reliable JSON back, how to handle streaming responses, and where a second layer (like a managed gateway) helps with tracking flagged requests across a team.

Why built-in safety isn't a moderation filter

Claude's training makes it refuse or soften responses to clearly dangerous requests — violence, illegal activity, self-harm instructions, and similar categories. That protects the model itself, but it doesn't give you:

If your app lets users submit free text — support tickets, chat messages, community posts — you almost always want an explicit moderation step, separate from whatever the main model call does.

Step 1: Define your categories first

Before writing any code, write down the categories you actually care about. Vague categories produce vague model behavior. A workable starting set:

Keep the list short. Five to eight categories is easier to get consistent classification on than twenty.

Step 2: Build a classification prompt

Use a separate, cheap call to classify input before it reaches your main prompt. Ask Claude to return strict JSON so you can parse it without guesswork:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-4-5",
    "max_tokens": 200,
    "system": "You are a content moderation classifier. Given user text, return ONLY valid JSON with this shape: {\"flagged\": boolean, \"categories\": {\"harassment\": number, \"self_harm\": number, \"sexual_content\": number, \"spam\": number}, \"reason\": string}. Scores are 0.0-1.0 confidence. No prose outside the JSON.",
    "messages": [
      {"role": "user", "content": "<insert user-submitted text here>"}
    ]
  }'

Use a fast, cheap model for this step — classification doesn't need your best model, and running it on every submission adds up in cost and latency. Set max_tokens low since you only need a short JSON object back.

Parse the response defensively: strip markdown fences if present, validate with a schema, and fall back to "flag for human review" if parsing fails rather than silently allowing content through.

function parseVerdict(text) {
  const cleaned = text.replace(/```json|```/g, "").trim();
  try {
    const verdict = JSON.parse(cleaned);
    if (typeof verdict.flagged !== "boolean") throw new Error("bad shape");
    return verdict;
  } catch {
    return { flagged: true, categories: {}, reason: "parse_error" };
  }
}

Step 3: Set thresholds, not just booleans

Don't rely purely on the flagged boolean — use the per-category scores to decide the action:

This lets you tune sensitivity per category without re-prompting the model. Harassment might need a lower block threshold than spam, for example.

Step 4: Moderate output, not just input

Input moderation catches bad prompts; it doesn't catch cases where a legitimate prompt produces a problematic response (rare with Claude, but possible with creative or long-form generation). Run a lightweight output check the same way, using the model's actual response as the input to your classifier prompt. For latency-sensitive apps, do this only for categories where false negatives are costly — you don't need to re-check every response for spam.

Handling moderation with streaming

If your main generation uses streaming (see /docs/streaming), you can't wait for the full response before showing anything to the user, but you also don't want to stream content that violates policy. Two practical patterns:

  1. Buffer-and-check in chunks: accumulate text in ~200-character windows, run a fast heuristic (keyword list, regex) between classifier calls, and only invoke the full classifier if the heuristic trips.
  2. Pre-check only: classify the prompt before generation starts, and accept that output-level issues are rare enough to catch via periodic sampling rather than real-time interruption.

Most production setups use pattern 2 for cost and latency reasons, reserving full re-classification for flagged conversations.

Where a gateway layer helps

Once moderation is running in production, you need visibility: which app or team generated the most flagged requests, how moderation calls are eating into your token budget, and who has access to adjust thresholds. This is where a thin API layer on top of your Claude access pays off. SubToAPI turns your existing Claude access into application-specific keys (sub_live_...) with usage metadata per key, so you can run your moderation classifier under its own key, track its cost separately from your main generation traffic, and give team members scoped access without sharing raw credentials. Check /docs/messages for the request format and /pricing for plan details if you're running moderation across multiple apps or team seats.

Keep a human in the loop

Automated moderation reduces volume, it doesn't eliminate judgment calls. Build a simple review queue for anything scoring in the 0.3–0.7 band, store the original text, the verdict, and the category scores, and revisit your thresholds monthly based on what reviewers actually overturn.

questions

Does Claude's API block harmful content automatically? Claude will refuse or soften responses to clearly harmful requests based on its training, but this isn't a configurable filter — you don't get scores, categories, or a verdict you can branch on in code. For that, build a separate classification call.

Should I use the same model for moderation and main generation? No. Use a fast, inexpensive model for classification since it only needs to return a short structured verdict, and reserve your main model for the actual task. This keeps moderation cost and latency low even at high volume.

How do I moderate streamed responses without delaying output? Classify the input prompt before streaming starts, then use lightweight heuristics (keyword or regex checks) on output chunks, only invoking the full classifier if a heuristic trips. Full real-time classification per chunk is usually too slow and costly for production use.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →