Build a Claude API Email Drafting Assistant
Build a Claude API Email Drafting Assistant
Building an email drafting assistant with the Claude API means sending a prompt that contains context (recipient, purpose, tone, key points) and letting the model return a complete, ready-to-send draft. The core work isn't the API call itself — it's structuring the prompt so Claude consistently produces emails that match your brand voice, length, and format without you babysitting every output.
This guide walks through the actual build: how to structure prompts for reliable drafts, how to handle tone and length variations, how to stream output into a UI as it's generated, and how to wire the whole thing into a production backend using the Claude API (or a proxy like SubToAPI if you want a single HTTPS key instead of managing provider credentials directly).
What the assistant actually needs to do
Before writing code, define the job clearly. A useful email drafting assistant typically needs to:
- Take unstructured input (bullet points, a rough idea, a previous thread) and turn it into a coherent email
- Match a requested tone: formal, casual, apologetic, firm, persuasive
- Respect length constraints (short follow-up vs. detailed proposal)
- Optionally reply within a thread, referencing prior messages
- Return clean text with no meta-commentary ("Here's your email:") unless explicitly asked
That last point matters more than people expect. Without constraining the output format, Claude will often wrap the draft in explanations. Fixing this is a prompt design problem, not a model limitation.
Designing the prompt
A good system prompt for this use case does three things: sets the role, defines the output contract, and gives tone/style rules.
You are an email drafting assistant. Given input points and a desired
tone, write a complete email ready to send.
Rules:
- Output only the email body (greeting through sign-off). No preamble,
no explanation, no markdown formatting.
- Keep paragraphs short. Avoid filler phrases like "I hope this email
finds you well" unless tone is "formal".
- Match the requested tone exactly: formal, casual, firm, apologetic,
persuasive, or neutral.
- If a recipient name is given, use it in the greeting. If not, use
"Hi there,".
- Default length is 120-180 words unless the user specifies otherwise.
Pass the variable content — recipient, bullet points, tone, length — in the user message, not the system prompt, so you can swap them per request without rewriting the whole prompt.
{
"model": "claude-sonnet-4-5",
"max_tokens": 500,
"system": "You are an email drafting assistant...",
"messages": [
{
"role": "user",
"content": "Recipient: Maria (client)\nTone: apologetic, firm\nPoints:\n- Shipment delayed 2 weeks\n- New delivery date: March 14\n- Offering 10% discount on next order"
}
]
}
This structured-input pattern is the single biggest lever for consistency. Free-text prompts ("write an email about a late shipment") work, but labeled fields reduce ambiguity and make it trivial to build a form-based UI on top.
Handling tone and length reliably
Claude follows explicit tone instructions well, but vague ones ("sound professional") produce inconsistent results across runs. Use a fixed vocabulary of tones in your UI (a dropdown, not free text) and map each one to a short style note in the prompt:
Tone definitions:
- formal: full sentences, no contractions, no slang
- casual: contractions okay, conversational, can use first names freely
- firm: direct, no hedging, no excessive apology
- apologetic: acknowledge the issue clearly before offering a fix
- persuasive: lead with the benefit to the recipient, end with a clear ask
For length, give a numeric target rather than "short" or "long" — models are more consistent hitting "under 100 words" than interpreting "brief."
Streaming the draft into a UI
Email drafts can take a few seconds to generate at longer lengths. Streaming the response token-by-token makes the assistant feel instant, which matters a lot for a tool people will use dozens of times a day. If you're building on SubToAPI, streaming works the same way as the standard Claude Messages API — see /docs/streaming for the event format and a working example.
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 500,
stream: true,
system: emailSystemPrompt,
messages: [{ role: "user", content: userInput }]
})
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let draft = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
draft += decoder.decode(value);
renderDraft(draft); // update the textarea live
}
Replying inside a thread
For a "reply assistant" that drafts responses to an existing email, pass the original message as context in the prompt rather than trying to fine-tune anything:
Original email from Maria:
"""
<paste original email text>
"""
Write a reply with tone: apologetic, firm.
Points to cover:
- Shipment delayed 2 weeks
- New delivery date: March 14
Claude will naturally reference specifics from the original email (dates, names, questions asked) without extra instruction, which is usually what makes an AI-drafted reply feel appropriate rather than generic.
Production considerations
A few things matter once this moves from prototype to something real users rely on:
- Rate limiting and cost control. Email drafting is a high-frequency, low-token-per-call workload. Track usage per user so one heavy user doesn't blow through your budget — SubToAPI's dashboard reports usage per API key, which makes this easier to monitor across a team. See /pricing for plan limits.
- Separate keys per environment. Use distinct application keys for staging and production so a bad prompt change in dev doesn't affect live traffic.
- Fallback on empty or malformed output. Occasionally the model may violate your output contract. Validate that the response is non-empty text and doesn't contain leftover instruction artifacts before showing it to the user.
- Team access. If multiple people on your team need their own keys and usage visibility, a seat-based setup (like SubToAPI's Team plan) is simpler than sharing one credential.
To get a working endpoint quickly, follow /docs/quickstart, then adapt the request format in /docs/messages for your email prompt structure. If the assistant needs to pull data (CRM contact info, past thread history) before drafting, look at /docs/tools for tool use patterns rather than stuffing everything into the prompt manually. You can start building with a free trial at /signup.
questions
Does Claude need fine-tuning to draft emails well? No. Prompt structure (labeled inputs for tone, recipient, and points) gets consistent results without any fine-tuning or training data.
How do I stop Claude from adding explanations before the email? Add an explicit output-format rule in the system prompt stating that only the email body should be returned, with no preamble or commentary.
Can this assistant handle replying inside an existing thread? Yes — paste the original email text into the user message as context and ask for a reply with a specified tone; Claude will reference it naturally.