Claude API Prompt Templates for Apps That Ship
If you're building a product on top of Claude, you don't want to reinvent prompt structure every time you add a feature. You want a small library of tested templates — for extraction, classification, summarization, rewriting, and chat — that you can parameterize and drop into any endpoint. This article gives you exactly that: copy-paste prompt templates for the most common app use cases, plus the structural rules that make them reliable in production.
The short answer: good Claude API prompt templates for apps separate system instructions (role, constraints, output format) from user input (the variable data), use explicit output schemas (usually JSON), and include a small number of examples when the task is ambiguous. Below are templates for the patterns you'll hit most often, written so you can lift them straight into your codebase.
Why templates matter more than one-off prompts
A prompt that works once in a playground often breaks when real users type unpredictable things into it. Templates fix this by:
- Keeping instructions stable in the
systemfield while only theusercontent changes per request - Forcing structured output so your backend doesn't have to parse free text
- Encoding edge-case handling ("if the input has no date, return null") directly into the instructions
- Making prompts testable — you can snapshot a template and run it against a fixed set of inputs in CI
If you're proxying Claude through SubToAPI, these templates work unchanged against the /v1/messages endpoint — the request shape mirrors Anthropic's Messages API, so you're not maintaining two prompt libraries.
Template 1: Structured data extraction
Use this for turning free text (emails, support tickets, form submissions) into JSON your app can store.
{
"system": "Extract structured data from the user's text. Return ONLY valid JSON matching this schema: {\"name\": string|null, \"email\": string|null, \"issue_category\": \"billing\"|\"bug\"|\"feature_request\"|\"other\", \"urgency\": \"low\"|\"medium\"|\"high\"}. If a field cannot be determined, use null. Do not include explanations.",
"model": "claude-sonnet-4-5",
"max_tokens": 300,
"messages": [
{ "role": "user", "content": "Hi, I'm Dana Kim, my card got charged twice this month, pretty urgent, dana@example.com" }
]
}
The key design choice is listing enum values explicitly ("billing"|"bug"|...) so Claude doesn't invent new categories. Always validate the JSON on receipt and retry once with a correction message if parsing fails.
Template 2: Classification with confidence
For routing, tagging, or moderation features:
{
"system": "Classify the input into exactly one category: spam, support, sales, abuse, other. Respond with JSON only: {\"category\": string, \"confidence\": number between 0 and 1}.",
"model": "claude-sonnet-4-5",
"max_tokens": 100,
"messages": [
{ "role": "user", "content": "{{user_text}}" }
]
}
Keep max_tokens tight for classification tasks — there's no reason to let the model ramble, and it reduces latency and cost.
Template 3: Summarization with length control
System: Summarize the following text in {{max_sentences}} sentences or fewer. Preserve key numbers and names. Write in plain, neutral language. Do not add opinions or commentary.
User: {{document_text}}
For apps that summarize at different lengths (short preview vs. full digest), parameterize max_sentences rather than writing separate prompts — it keeps your template library small and your behavior consistent.
Template 4: Rewriting / tone transformation
Useful for "make this more formal," "shorten this," or "translate to friendly support tone" features:
{
"system": "Rewrite the user's text to match this tone: {{tone}}. Keep the same meaning and factual content. Do not add new information. Return only the rewritten text, no preamble.",
"model": "claude-sonnet-4-5",
"max_tokens": 500,
"messages": [
{ "role": "user", "content": "{{original_text}}" }
]
}
The "no preamble" instruction matters — without it, Claude often prefaces output with "Here's a more formal version:", which you then have to strip in post-processing.
Template 5: Multi-turn chat with persona
For in-app assistants, define the persona once in system and let conversation history carry context:
{
"system": "You are the support assistant for Acme Inc. Be concise, friendly, and only answer questions about Acme's product. If asked about anything else, politely redirect to the topic. Never make up pricing — if unsure, say you'll check and follow up.",
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "Do you support SSO?" },
{ "role": "assistant", "content": "Yes, SSO is available on our Team and Scale plans." },
{ "role": "user", "content": "What about SAML specifically?" }
]
}
Keep the system prompt stable across the whole session; only the messages array grows. This is cheaper to reason about and avoids subtle persona drift mid-conversation.
Template 6: Tool-calling template
When your app needs Claude to decide whether to call a function (lookup inventory, fetch a record, run a calculation), define the tool schema once and reuse the instruction pattern:
{
"system": "Use the provided tools when you need external data. Do not guess values that a tool could retrieve.",
"model": "claude-sonnet-4-5",
"max_tokens": 500,
"tools": [
{
"name": "get_order_status",
"description": "Look up the status of an order by ID",
"input_schema": {
"type": "object",
"properties": { "order_id": { "type": "string" } },
"required": ["order_id"]
}
}
],
"messages": [
{ "role": "user", "content": "Where's my order #48213?" }
]
}
See /docs/tools for the full request/response cycle, including how to return tool results back to Claude for a final answer.
Making templates production-ready
- Version them. Store each template with a version number in your repo; log which version generated each response so you can trace regressions.
- Validate output. JSON-schema-validate every structured response before writing it to your database.
- Set a token ceiling. Always cap
max_tokensto the smallest value the task needs — it limits cost and runaway generations. - Stream for chat, don't stream for extraction. Structured outputs should be awaited in full so you can validate before using them; see /docs/streaming for when streaming actually helps UX.
If you're routing these templates through SubToAPI, each request uses your sub_live_... application key, and usage per template/endpoint shows up in your dashboard — handy for seeing which feature is driving token spend. Get started at /signup, check the request format in /docs/quickstart, and see full parameter options in /docs/messages.
questions
Do I need few-shot examples in every template? No. Use examples only when the task is genuinely ambiguous (unusual output formats, edge-case category boundaries). For straightforward extraction or classification, clear instructions and an explicit schema are usually enough and keep token usage lower.
Should system prompts or user messages hold the variable data? Keep instructions and constraints in system; put the actual data (document text, user input) in user content. This separation makes templates reusable across requests without rewriting instructions each time.
How do I stop Claude from adding extra commentary around JSON output? State explicitly "return only valid JSON, no explanation" in the system prompt, and set a tight max_tokens. If commentary still appears occasionally, add a post-processing step that strips anything outside the outermost {}.