← Blog

Claude API Summarization for Long Documents

2026-10-02 · 5 min read · SubToAPI Team

Summarizing long documents with the Claude API comes down to one decision: does the document fit in a single context window, or do you need to split it up first. Claude's models support context windows large enough to handle most reports, contracts, and transcripts in one request — but once you're dealing with multi-hundred-page PDFs, full books, or entire codebases, you need a chunking strategy to get a coherent summary without losing information.

This article covers both cases: how to summarize a document in one call when it fits, and how to structure a map-reduce pipeline when it doesn't, along with the prompt patterns that actually produce useful summaries instead of generic paraphrasing.

Single-pass summarization

If your document fits comfortably within the model's context window, the simplest and most reliable approach is a single request. The key is being specific about what kind of summary you want — "summarize this" produces vague output, while a structured instruction produces something usable.

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "system": "You are a technical summarizer. Produce a summary with: (1) a 2-sentence overview, (2) 5-8 key bullet points, (3) any action items or decisions mentioned. Be concrete, no filler language.",
    "messages": [
      {"role": "user", "content": "<full document text here>"}
    ]
  }'

This works well for documents up to tens of thousands of words. The system prompt is doing the real work here — it forces a consistent output structure, which matters a lot if you're summarizing many documents and need the results to be comparable or parseable downstream.

Chunking for documents that exceed the context window

For books, long legal filings, or full meeting transcript archives, you need to split the document into chunks, summarize each chunk, then summarize the summaries. This is the standard map-reduce pattern for LLM summarization:

  1. Split the document into overlapping chunks (e.g., 8,000–12,000 tokens each, with a few hundred tokens of overlap so you don't cut sentences or context mid-thought).
  2. Map: send each chunk to Claude with a prompt asking for a dense, factual summary — not a polished narrative, just the key facts and claims.
  3. Reduce: concatenate the chunk summaries and send them in one final request asking Claude to merge them into a single coherent summary, removing redundancy.
async function summarizeChunk(chunk) {
  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 500,
      system: "Extract the key facts, claims, and figures from this text section. Be dense and factual, no narrative framing.",
      messages: [{ role: "user", content: chunk }]
    })
  });
  const data = await res.json();
  return data.content[0].text;
}

async function summarizeDocument(chunks) {
  const partials = await Promise.all(chunks.map(summarizeChunk));
  const merged = partials.join("\n\n---\n\n");

  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1500,
      system: "Merge these section summaries into a single coherent summary of the full document. Remove redundancy, preserve all distinct facts and figures.",
      messages: [{ role: "user", content: merged }]
    })
  });
  const data = await res.json();
  return data.content[0].text;
}

Running chunk summaries in parallel (as shown with Promise.all) is important for keeping pipeline latency reasonable — sequential chunk-by-chunk summarization on a 300-page document can take minutes.

Streaming for long outputs

If you're generating a long-form summary (not a short bullet list but a detailed multi-page synthesis), streaming matters for perceived latency. Rather than waiting for the full response, you can render the summary as it's generated:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 2000,
    "stream": true,
    "messages": [{"role": "user", "content": "Write a detailed section-by-section summary of this report: ..."}]
  }'

See /docs/streaming for the full event format if you're building a UI that displays the summary incrementally.

Practical tips that improve summary quality

Where SubToAPI fits

If you're building a summarization pipeline that multiple people on your team will use — support ticket digests, meeting notes, contract reviews — SubToAPI gives you application-level API keys (sub_live_...) instead of sharing one raw credential, plus per-key usage metadata so you can see exactly how many tokens your summarization jobs are consuming and by whom. Team and Scale plans add seats so each developer or service gets its own key without touching billing. Start with the free trial at /signup, or check /pricing for plan details, and /docs/quickstart to get a key working in minutes.

Questions

Does Claude's context window mean I never need to chunk documents? No. The context window determines the upper bound, but very long documents (books, large codebases, long transcript archives) can still exceed it, and even within the limit, extremely long single-pass summarization can lose fine-grained detail. Chunking also lets you parallelize, which is faster.

Should I summarize PDFs directly or extract text first? Extract text first in almost all cases. Clean, well-formatted plain text gives the model a clearer signal than raw PDF text that includes headers, footers, and layout artifacts. If the PDF is scanned, run OCR before summarization.

How do I keep summaries consistent across many documents? Use a fixed system prompt that defines the exact output structure (overview, bullets, action items) and apply it identically to every document. Consistency comes from prompt discipline, not from the model's "memory" — each request is independent unless you explicitly pass prior context.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →