Claude API System Prompt Design Tips
The system prompt is the single highest-leverage input in any Claude API call. It sets the model's role, tone, constraints, and output format before the conversation even starts, and it persists across every turn without being repeated by the user. Get it right and you cut down on retries, malformed outputs, and off-topic responses. Get it wrong and you'll spend more time patching individual replies than you would have spent writing a better prompt in the first place.
This article covers concrete, testable techniques for writing Claude system prompts that hold up in production: how to structure them, what to include, what to leave out, and how to iterate on them like you would any other piece of code.
System Prompt vs. User Message: Know the Difference
In the Messages API, system is a separate top-level field, not a message with role system inside the messages array:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"system": "You are a support triage assistant for a SaaS billing product. Classify each ticket into one of: billing, bug, feature_request, other. Respond with JSON only.",
"messages": [
{"role": "user", "content": "I was charged twice this month, can you check?"}
]
}'
Anything that should apply to every turn belongs in system. Anything specific to a single exchange — the actual question, the document being analyzed, the code being reviewed — belongs in messages. Mixing these up is the most common mistake: developers stuff task-specific data into the system prompt and then wonder why it doesn't update between requests, or they repeat role instructions in every user message and burn tokens for no benefit.
Be Specific About Role, Scope, and Boundaries
Vague system prompts produce vague behavior. "You are a helpful assistant" tells the model almost nothing about what "helpful" means for your use case. Instead, define:
- Role: what the model is (a code reviewer, a support agent, a data extraction tool)
- Scope: what it should and shouldn't do (answer only questions about X, refuse requests about Y)
- Audience: who's on the other end (developers, end customers, internal staff)
- Tone: formal, terse, conversational — pick one and be explicit
A well-scoped example:
You are a code review assistant for a Python backend team.
Review only for: security issues, unhandled exceptions, and SQL injection risk.
Do not comment on formatting or style — a linter already handles that.
Keep each comment under two sentences. If no issues are found, say "No issues found."
This is far more useful than "You are an expert Python developer" because it tells Claude exactly what counts as in-scope feedback and what to skip.
Specify Output Format Explicitly
If you need structured output — JSON, markdown tables, a fixed list of fields — say so directly in the system prompt, and show the exact shape you want:
Respond only with valid JSON matching this schema:
{
"category": "billing" | "bug" | "feature_request" | "other",
"priority": "low" | "medium" | "high",
"summary": "one sentence"
}
Do not include any text outside the JSON object.
Claude follows format constraints reliably when they're stated as rules rather than implied by example alone. If you're also using tool use, keep the system prompt focused on behavior and judgment; let the tool schema itself define the structured output. See /docs/tools for how tool definitions interact with the system prompt.
Use Few-Shot Examples Sparingly, But Use Them
For tasks with subtle judgment calls — tone matching, edge-case classification, style replication — one or two examples in the system prompt outperform paragraphs of description. Keep them short and put the pattern, not the full range of edge cases, since Claude generalizes well from a small number of clean examples:
Example input: "App keeps crashing on startup after the update"
Example output: {"category": "bug", "priority": "high", "summary": "App crashes on startup post-update"}
Avoid stacking more than two or three examples — beyond that you're adding tokens without adding much signal, and long system prompts increase latency and cost on every single request.
Don't Bury Instructions in the Middle
Claude, like other long-context models, pays closer attention to content at the start and end of a prompt. Put the role definition and hard constraints first. If you have reference material — a style guide, a product FAQ, a set of rules — put it after the core instructions but before any closing reminders. If there's one rule you absolutely need followed, restate it at the very end of the system prompt as a final line.
Version and Test Your System Prompts Like Code
Treat system prompts as versioned artifacts, not throwaway strings:
- Store them in your repo, not hardcoded inline in application logic
- Keep a small test set of representative inputs and expected output shapes
- Re-run that test set whenever you change the prompt or switch models
- Log which system prompt version produced which output, so regressions are traceable
If you're calling Claude through SubToAPI, this is easier to manage because every request already carries usage metadata back in the response — you can tag prompt versions in your own logs and correlate them with token usage and latency per version without building separate instrumentation. Check /docs/messages for the response fields available.
Watch Prompt Length and Cost Together
Every token in the system prompt is billed on every single request, including ones where most of it isn't relevant. If your system prompt has grown past a few hundred tokens, look for:
- Instructions that duplicate what's already implicit in a good example
- Boilerplate disclaimers that could move to a single check outside the model call
- Reference data that should be retrieved dynamically instead of embedded permanently
A tight, well-tested 150-word system prompt usually outperforms a sprawling 800-word one that tries to cover every possible scenario up front.
Getting Started Quickly
If you're prototyping system prompts and want to skip separate key management for testing, SubToAPI lets you issue an application-scoped sub_live_ key and call the same Messages-style endpoint with streaming and tool use support. Sign up at /signup, check /pricing for plan details, and see /docs/quickstart to send your first request in a few minutes.
questions
Should the system prompt or the first user message contain task instructions? Persistent instructions — role, tone, format rules — go in system. Task-specific input for that particular request goes in the first user message. This keeps the system prompt reusable across many different queries.
How long should a Claude system prompt be? There's no hard limit, but shorter, well-structured prompts with clear rules and one or two examples typically outperform long ones. Anything beyond a few hundred words should be reviewed for redundancy, since every token is billed on each request.
Can I change the system prompt between messages in the same conversation? Yes — each API call is stateless, so you can send a different system value on every request, including mid-conversation, though changing it abruptly can confuse context continuity for the model if the new instructions conflict with earlier turns.