← Blog

Prompt Engineering for Claude API Apps That Ship

2026-10-08 · 5 min read · SubToAPI Team

If you're building a product on top of Claude rather than just chatting with it, prompt engineering stops being a creative exercise and becomes an engineering discipline. The question isn't "what's a clever prompt?" — it's "what prompt structure produces consistent, parseable, correct output across thousands of real user inputs, including the weird ones?"

This article covers the concrete techniques that matter for production Claude API apps: how to structure prompts for reliability, how to use system prompts correctly, how to control output format, and how to test prompts before they break in front of users.

Separate instructions from data

The single most common mistake in Claude API apps is mixing instructions and user content in one unstructured blob. Claude (like every LLM) performs better when it can clearly distinguish "what you're telling me to do" from "the thing you want me to do it to."

Use the system parameter for the stable, reusable instructions — role, tone, output format, constraints. Put the variable, per-request content in the user message, and wrap it in clear delimiters:

Summarize the following support ticket in 2 sentences.
Focus on the customer's actual problem, not their tone.

<ticket>
{{ticket_text}}
</ticket>

This separation does two things: it makes your prompt template reusable across requests, and it reduces the chance that user input gets interpreted as instructions (a common source of both bugs and prompt injection).

Be explicit about output format

Claude is good at following format instructions, but only if you give it one. If your app parses the response programmatically — and most API-backed apps do — specify the exact shape you want:

{
  "sentiment": "positive | neutral | negative",
  "summary": "one sentence, max 20 words",
  "flagged": true
}

Tell Claude to return only the JSON object, with no preamble like "Here's the summary:". If you need stricter guarantees than natural-language formatting instructions can give, use Claude's tool use / function calling — defining a schema and letting Claude call a "tool" to return structured data is far more reliable than hoping it follows a format instruction every time. Our tool use docs cover how to wire this up through a Claude-compatible endpoint.

Use few-shot examples for judgment calls

For tasks with subjective or ambiguous criteria — classifying tone, deciding severity, extracting the "right" fields from messy text — instructions alone underperform examples. Show two or three input/output pairs directly in the prompt:

Example 1
Input: "App keeps crashing on launch, lost all my data"
Output: {"severity": "high", "category": "bug"}

Example 2
Input: "Love the new icon, nice touch"
Output: {"severity": "low", "category": "feedback"}

Now classify:
Input: "{{user_input}}"

Few-shot examples anchor Claude's output distribution to your specific taxonomy instead of a generic one. They're especially valuable when you've already seen edge cases in production and want to encode the correct behavior directly.

Give Claude room to reason, then extract the answer

For non-trivial tasks (multi-step reasoning, calculations, policy decisions), letting Claude think through the problem before answering improves accuracy. But if you need a clean machine-readable result, don't make your app parse free-form reasoning. Ask for a visible reasoning section followed by a clearly delimited final answer:

Think through this step by step inside <reasoning></reasoning> tags.
Then give your final answer inside <answer></answer> tags, and nothing after it.

Your application code then just extracts the content between <answer> tags with a regex or XML parser, ignoring the reasoning. This pattern gives you the accuracy benefits of chain-of-thought without sacrificing parseability.

Constrain scope explicitly

Claude will try to be helpful by default, which sometimes means answering questions you didn't want answered, or padding output with caveats and disclaimers. If your app has a narrow job, say so directly:

Explicit negative instructions (what not to do) are underused but very effective for keeping output within your app's expected boundaries.

Version and test your prompts like code

Prompts that work on your ten manual test cases often fail on the thousandth real user input. Treat prompts as versioned artifacts:

This is the same discipline you'd apply to any code that affects user-facing behavior — because that's what a prompt is.

Where the API layer fits in

Prompt engineering determines what you ask Claude. The API layer determines how reliably you can ask it — streaming, retries, timeouts, usage tracking, and giving different parts of your team or product their own scoped keys. SubToAPI turns your existing Claude access into an HTTPS API with sub_live_... application keys, so the prompts you've engineered here can be sent from any service, with streaming and usage metadata included, without each environment needing its own Anthropic account. Check the quickstart to see the request format, or the messages docs for the full parameter reference.

If you're prototyping prompts interactively before wiring them into your app, streaming support lets you see output token-by-token exactly as your end users will, which is often where format or verbosity problems first become obvious.

Questions

Does prompt engineering differ between Claude models? The core techniques — delimiters, explicit format instructions, few-shot examples — transfer across models, but verbosity and instruction-following strictness vary, so test your prompts against whichever model version you deploy.

Should I put formatting rules in the system prompt or the user message? Put stable, reusable rules (tone, output schema, constraints) in the system prompt, and keep the user message focused on the specific task or data for that request.

Is tool use better than prompting for structured output? For anything you need to parse reliably in code, yes — defining a tool schema is more robust than relying on natural-language format instructions, especially as inputs get more varied.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →