Claude API Long Context Summarization Example
Summarizing a 50-page PDF, a full codebase README set, or a stack of customer call transcripts is one of the most common reasons developers reach for the Claude API. Claude's large context windows (up to 200K tokens depending on model) mean you can often paste an entire document in directly, rather than chunking it manually. This article walks through a real, working example of long context summarization, including how to structure the request, when you still need chunking, and how to keep costs predictable.
The short answer for anyone in a hurry: send the full document as a single user message with a clear instruction, use a low temperature, and ask for a structured output format (bullets, sections, or JSON) so the summary is easy to parse downstream. Below is the full pattern, plus what to do when your document actually exceeds the context window.
A basic long-document summarization request
Here's a minimal example using the Messages API directly:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 1024,
"temperature": 0,
"messages": [
{
"role": "user",
"content": "Summarize the following report in 5 bullet points, then list any risks mentioned. Be concise and factual.\n\n<document>\n'"$(cat report.txt)"'\n</document>"
}
]
}'
A few things matter more than they look:
- Wrap the document in tags like
<document>...</document>. Claude handles delimited long-form text more reliably than raw pasted content mixed with instructions. - Put instructions before the document, or repeat them after it. For very long inputs, models sometimes weight the end of the prompt more heavily — repeating a short instruction after the document ("Now summarize the above in 5 bullets") improves consistency.
- Set
temperature: 0for summarization tasks. You want faithful compression, not creative variation. - Ask for structure explicitly. "5 bullet points" or "JSON with keys
summary,key_points,risks" produces far more usable output than an open-ended "summarize this."
When the document doesn't fit in one call
Even with a 200K token context window, some inputs — full log archives, multi-document research corpora, entire codebases — won't fit, or you don't want to pay for processing the whole thing every time. The standard pattern is map-reduce summarization:
- Split the source into chunks (by section, by page, or by a fixed token count with slight overlap).
- Summarize each chunk independently ("map").
- Concatenate the chunk summaries and run a final summarization pass over those ("reduce").
async function summarizeChunk(chunk) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 400,
temperature: 0,
messages: [
{ role: "user", content: `Summarize this section in 3-4 sentences:\n\n${chunk}` }
]
})
});
const data = await res.json();
return data.content[0].text;
}
async function mapReduceSummarize(chunks) {
const partials = await Promise.all(chunks.map(summarizeChunk));
const combined = partials.join("\n\n");
return summarizeChunk(combined); // final reduce pass, reused function
}
This scales to arbitrarily large inputs and keeps each individual call well within context and cost limits. The tradeoff is that map-reduce can lose cross-chunk relationships — if section 2 references something defined in section 8, that connection might not survive independent summarization. For documents where cross-references matter, prefer a single large call over chunking whenever the input fits.
Choosing chunk size
A reasonable default is 3,000–5,000 tokens per chunk with roughly 200 tokens of overlap between chunks, so context isn't lost at boundaries. Chunking by natural document structure (headings, paragraphs) usually beats fixed-size splitting because you avoid cutting sentences or tables in half.
Getting consistent summary formats
If downstream code parses the summary, ask for JSON and validate it:
Summarize the document below. Respond with only valid JSON matching:
{"summary": string, "key_points": string[], "risks": string[]}
<document>
...
</document>
Claude follows structured output instructions well, but always parse defensively — wrap JSON.parse in a try/catch and have a fallback path (re-prompt or return the raw text) for the rare malformed response.
Cost and latency considerations
Long context calls cost more per request simply because input tokens dominate the bill for summarization tasks (you're sending a lot in, and getting a little back out). A few practical levers:
- Cache repeated documents. If you're summarizing the same source multiple times (different questions, different formats), avoid re-sending the full document each time — see prompt caching in Anthropic's docs, or reuse cached responses at your application layer.
- Trim before you send. Strip boilerplate, navigation text, and repeated headers/footers from PDFs and HTML before summarization — it's dead weight that costs tokens and adds noise.
- Use a smaller/faster model for the map step and a stronger model only for the final reduce pass, if quality allows.
If you're calling the Claude API from a product and want usage visibility per customer or per feature without building metering yourself, SubToAPI sits in front of your existing Claude access and gives you application API keys, streaming support, and per-key usage metadata — useful when summarization is a billed feature and you need to know exactly which key is consuming how many tokens. See the quickstart or the Messages endpoint docs for the request shape, which mirrors the examples above.
questions
Does Claude's context window mean I never need to chunk documents? Not always. If your document fits comfortably within the model's context window (with room left for the response), a single call is simpler and preserves cross-references better than chunking. Chunking is only necessary when the input exceeds the limit or when you want to control per-request cost.
What's the best way to summarize very long transcripts with speaker turns? Keep speaker labels in the text (e.g., "Agent:", "Customer:") and instruct Claude to preserve who said what in the summary. For transcripts over ~5,000 tokens, map-reduce by conversation segment works well.
Can I stream a long summarization response? Yes — streaming is supported for long outputs and improves perceived latency for users waiting on a multi-paragraph summary. See streaming docs for the event format if you're building this through SubToAPI's API layer.