Claude API Few-Shot Examples Formatting Guide
How to Format Few-Shot Examples for the Claude API
Few-shot prompting means showing Claude a handful of input/output pairs before asking it to handle a new case. The formatting question most developers run into is: should examples go in the system prompt, inside the user message, or as separate alternating messages in the messages array? The short answer is that Claude responds best to examples formatted as actual conversation turns (alternating user and assistant messages) or, for single-turn tasks, as clearly delimited blocks inside one message using XML-style tags.
This matters because Claude is trained to pay close attention to conversational structure. Burying five examples as a wall of text in the system prompt works, but it's harder for the model to parse boundaries between examples compared to using the native message format or explicit tags. Below are the three patterns that actually produce reliable results, along with the mistakes that cause inconsistent output.
Pattern 1: Alternating Messages (Best for Chat-Style Tasks)
If your task resembles a conversation — classification, rewriting, Q&A — put each example as a real user/assistant pair in the messages array, then end with the actual query as the final user message.
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 200,
"messages": [
{"role": "user", "content": "Classify sentiment: 'This laptop is amazing.'"},
{"role": "assistant", "content": "positive"},
{"role": "user", "content": "Classify sentiment: 'The battery died in two hours.'"},
{"role": "assistant", "content": "negative"},
{"role": "user", "content": "Classify sentiment: 'It arrived on time, nothing special.'"},
{"role": "assistant", "content": "neutral"},
{"role": "user", "content": "Classify sentiment: 'Customer support ignored my emails for a week.'"}
]
}
This is the most token-efficient and reliable approach because Claude's training data includes huge volumes of real multi-turn conversations. It infers the pattern (short, single-word labels) without needing extra instructions.
Pattern 2: XML Tags in a Single Message
For tasks where you want all examples bundled into one prompt (useful when you're constructing the request programmatically and don't want to manage message arrays), wrap each example in tags like <example> with sub-tags for input and output.
Convert each product description into a JSON object with "name", "price", and "in_stock" fields.
<example>
<input>Wireless Mouse - $24.99 - In Stock</input>
<output>{"name": "Wireless Mouse", "price": 24.99, "in_stock": true}</output>
</example>
<example>
<input>USB-C Cable - $9.50 - Out of Stock</input>
<output>{"name": "USB-C Cable", "price": 9.50, "in_stock": false}</output>
</example>
Now convert this:
<input>Mechanical Keyboard - $89.00 - In Stock</input>
Claude was specifically trained to recognize and respect XML-style structure, so tags like <example>, <input>, and <output> are not just cosmetic — they help the model separate the pattern from the actual task at the end. This pattern works well whether the examples live in the system prompt or the first user message.
Pattern 3: System Prompt for Format, Examples in User Message
A common hybrid: put general instructions and the output schema in the system prompt, and keep the few-shot examples in the user message right next to the real input. This keeps the system prompt stable across requests (good for caching) while letting examples vary per call.
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 300,
"system": "You are a data extraction assistant. Always respond with valid JSON matching the schema shown in the examples. Do not add commentary.",
"messages": [
{
"role": "user",
"content": "<example>\n<input>Call John at 555-0192 tomorrow</input>\n<output>{\"contact\": \"John\", \"phone\": \"555-0192\", \"when\": \"tomorrow\"}</output>\n</example>\n\n<input>Email Sarah before Friday about the invoice</input>"
}
]
}
How Many Examples and How to Order Them
A few practical rules that apply regardless of which pattern you use:
- 3–5 examples is usually enough. More than that rarely improves accuracy and just burns tokens. If Claude still isn't matching the pattern after 5 well-chosen examples, the task likely needs clearer instructions rather than more examples.
- Cover edge cases, not just the happy path. If your real inputs include empty strings, unusual formatting, or ambiguous cases, include at least one example of that in the few-shot set.
- Keep output format perfectly consistent across examples. If one example has a trailing period and another doesn't, Claude will pick up on the inconsistency and may alternate between formats in its own output.
- Put the hardest or most unusual example last, closest to the actual query. Position matters slightly, and examples nearer the end of the prompt tend to have marginally more influence on the immediate next output.
- Match punctuation and whitespace exactly to what you want back. Claude mirrors formatting precisely, including markdown bullets, indentation, and capitalization.
Avoiding Format Drift Mid-Conversation
If you're running few-shot prompts repeatedly through an API — for example, classifying a stream of support tickets — keep the example block identical across calls and only swap out the final input. Changing the examples between requests (even slightly) makes outputs less predictable, since the model has less fixed context to anchor on. This is also where prompt caching helps if your provider supports it, since the example block is the expensive, repeated part of the prompt.
If you're routing these requests through SubToAPI, the request format is the same /v1/messages structure described above — you just send it with your sub_live_... key and get usage metadata back per call, which is useful for tracking how many tokens your few-shot examples are actually costing you across thousands of requests. See the messages documentation for the full request schema and the quickstart if you're setting this up for the first time.
When to Stop Using Few-Shot and Write Clearer Instructions
If you find yourself needing more than 5–6 examples to get consistent behavior, that's usually a sign the instructions are ambiguous rather than that the model needs more demonstrations. Try rewriting the system prompt to be explicit about edge cases first — it's cheaper in tokens and more maintainable than a long example list that needs updating every time you find a new edge case.
Questions
Should few-shot examples go in the system prompt or the messages array? Use alternating user/assistant messages for conversational tasks; use XML-tagged examples inside a single message (system or user) when you want one static block of examples reused across many different inputs.
How many few-shot examples does Claude need? 3–5 well-chosen examples covering typical cases and at least one edge case is usually sufficient. Adding more rarely improves results and increases token cost on every request.
Does example order affect Claude's output? Yes, slightly — examples closer to the final query tend to have more influence, so put your most representative or trickiest example last, right before the actual input.