Claude API Summarization: Long Document Example
Summarizing long documents with the Claude API
If you're trying to summarize a long document — a 50-page PDF, a legal contract, a research paper, a transcript — with the Claude API, the core challenge isn't the summarization itself. Claude is good at that. The challenge is getting the document into the request correctly: handling context window limits, choosing between a single-pass summary and a chunked map-reduce approach, and structuring the prompt so the output is useful (not a vague "this document discusses various topics" paragraph).
This article walks through a working example: a single-call summarization for documents that fit in context, and a chunked approach for documents that don't, with code you can copy directly.
Step 1: Decide if you need chunking at all
Claude models (3.5 Sonnet, Opus, Haiku) support context windows up to 200K tokens — roughly 150,000 words of English text. That's enough for most reports, contracts, and transcripts in a single request. Don't chunk unless you have to; chunking adds complexity and loses cross-section context.
Rough token estimate: divide word count by 0.75, or just count characters and divide by ~4. If your document is under ~140K tokens after that math, send it in one call and leave headroom for the output.
Single-pass summarization example
For a document that fits in context, the pattern is simple: put the full text in the user message, and use the system prompt to define the summary format.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 1024,
"system": "You summarize long documents for busy executives. Produce: a 3-sentence overview, a bulleted list of key points (max 8), and a short list of action items if any exist. Do not include filler phrases like \"this document discusses\".",
"messages": [
{"role": "user", "content": "<full document text here>"}
]
}'
The system prompt is doing most of the work here — it forces a concrete output structure instead of a generic recap. If you're routing this through SubToAPI, the request shape is identical; you just point it at https://api.subtoapi.app/v1/messages with your sub_live_... key:
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
system: "Summarize the document into: overview (3 sentences), key points (bullets), open questions (bullets).",
messages: [{ role: "user", content: documentText }]
})
});
const data = await res.json();
console.log(data.content[0].text);
See the messages endpoint docs for the full request/response schema.
Chunked (map-reduce) summarization for very long documents
When a document exceeds the context window — books, multi-document bundles, long legal filings — use a map-reduce pattern:
- Split the document into overlapping chunks (e.g., 8,000 tokens each, with ~200 tokens of overlap so you don't cut sentences mid-thought).
- Map: summarize each chunk independently.
- Reduce: feed all chunk summaries into a final call that produces the combined summary.
function chunkText(text, maxChars = 24000, overlap = 800) {
const chunks = [];
let start = 0;
while (start < text.length) {
const end = Math.min(start + maxChars, text.length);
chunks.push(text.slice(start, end));
start = end - overlap;
}
return chunks;
}
async function summarizeChunk(chunk) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 500,
system: "Summarize this excerpt in 5 bullet points, preserving names, dates, and figures exactly.",
messages: [{ role: "user", content: chunk }]
})
});
const data = await res.json();
return data.content[0].text;
}
async function summarizeLongDocument(fullText) {
const chunks = chunkText(fullText);
const chunkSummaries = await Promise.all(chunks.map(summarizeChunk));
const finalRes = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1200,
system: "Combine these section summaries into one coherent document summary with an overview, key points, and action items. Remove redundancy across sections.",
messages: [{ role: "user", content: chunkSummaries.join("\n\n---\n\n") }]
})
});
const finalData = await finalRes.json();
return finalData.content[0].text;
}
A few tuning notes:
- Keep the per-chunk prompt tight and consistent — ask for bullets, not prose, so the reduce step has structured input to work with.
- Preserve exact figures (dates, amounts, names) explicitly in the map-step instructions. Models can paraphrase numbers under compression unless told not to.
- The map step can run in parallel (as shown with
Promise.all), which keeps wall-clock time reasonable even for a 10-chunk document. - Watch your request concurrency and rate limits when fanning out chunk calls — this is where a proxy with usage visibility across the whole pipeline helps, since you can see token and request totals for the job in one place rather than reconciling multiple API keys.
Choosing output length and format
Summarization quality drops when max_tokens is set too low and Claude has to truncate mid-thought, and it drops in a different way when it's set too high and the model pads with restated content. For a one-page summary, 500–800 output tokens is usually enough. For an executive brief with sections, 1000–1500. Always specify the structure explicitly (headings, bullet count, word limit) — Claude follows explicit formatting constraints reliably.
Streaming for long summarization jobs
If you're summarizing inside a UI (not a batch job), stream the response so users see output as it's generated rather than waiting for the full summary. See the streaming guide for the SSE event format, and the quickstart if you're setting up your first API key.
Questions
Can Claude summarize a document that exceeds its context window in one call? No. You need to chunk the document, summarize each chunk (map), then combine the chunk summaries into a final summary (reduce), as shown above.
What's the best max_tokens value for document summaries? 500–800 for a short summary, 1000–1500 for a structured multi-section brief. Setting it too low truncates the response; too high encourages padding.
Does SubToAPI change how summarization prompts work? No — SubToAPI proxies the same Messages API, so your prompts, system instructions, and chunking logic are identical. It adds API keys, usage metadata, and team billing on top. See /pricing for plan details.