Claude API Long Document Chunking Strategy Guide
Why chunking still matters with large context windows
Claude models support very large context windows, but that doesn't mean you should always stuff an entire document into a single request. A good Claude API long document chunking strategy balances three things: staying under token limits, keeping retrieval or summarization accurate, and controlling cost — because tokens are tokens whether they're in a 20k-word contract or a 500k-word corpus.
The short answer for most teams: split documents into semantically coherent chunks of 500–2,000 tokens, add 10–15% overlap between chunks, preserve structural metadata (headings, page numbers, section IDs), and use a map-reduce pattern for summarization or a retrieval step for Q&A. The rest of this article breaks down how to implement that in practice.
When you actually need to chunk
Before building a chunking pipeline, check if you need one at all:
- Document fits comfortably in context (with room for the response and system prompt) → send it whole. Simpler, more accurate, fewer moving parts.
- Document exceeds the context window, or you're running many documents through the same prompt and want consistent, comparable outputs → chunk it.
- You need citations or source attribution → chunk it, because tracking which chunk produced which claim is far easier than parsing a single giant response.
- You're doing retrieval (RAG) rather than full-document summarization → chunking is mandatory, since you only want to pass the relevant pieces.
Core chunking strategies
Fixed-size chunking
Split text every N tokens regardless of content boundaries. It's the simplest to implement and works fine for free-form text where structure doesn't matter much.
function fixedSizeChunks(text, chunkSize = 1000, overlap = 150) {
const words = text.split(/\s+/);
const chunks = [];
let i = 0;
while (i < words.length) {
chunks.push(words.slice(i, i + chunkSize).join(" "));
i += chunkSize - overlap;
}
return chunks;
}
The downside: it can cut sentences or ideas in half, which hurts accuracy when Claude has to reason about a chunk in isolation.
Structure-aware chunking
Split on natural boundaries first — headings, paragraphs, markdown sections, HTML tags, or page breaks — then merge small pieces up to a target token count. This preserves meaning and is the right default for contracts, manuals, reports, and anything with a table of contents.
function structureAwareChunks(sections, maxTokens = 1500) {
const chunks = [];
let current = "";
for (const section of sections) {
if ((current + section).length / 4 > maxTokens && current) {
chunks.push(current);
current = "";
}
current += section + "\n\n";
}
if (current) chunks.push(current);
return chunks;
}
A rough rule of thumb: 1 token ≈ 4 characters of English text, so a 1,500-token target is roughly 6,000 characters. Always verify with an actual tokenizer before relying on this for billing-sensitive workloads.
Sliding window with overlap
Overlap between chunks prevents context loss at boundaries — a sentence that references "the clause above" shouldn't lose that reference just because it fell on a chunk edge. 10–15% overlap is a reasonable starting point; go higher (20–25%) for dense legal or technical text where cross-references are common.
Semantic chunking
For the highest accuracy, split based on embedding similarity: group sentences that are semantically close and break where topic shifts occur. This is more expensive to compute (you need an embedding pass first) but produces chunks that map cleanly to a single idea, which improves both retrieval precision and summarization quality.
Handling documents that exceed the context window
When a single document is too large even after chunking into reasonable units, use a map-reduce pattern:
- Map: send each chunk to Claude with a focused prompt ("summarize this section, note any numbers, dates, or defined terms").
- Reduce: concatenate the chunk-level outputs and send that combined, much shorter text through a final synthesis prompt.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 500,
"messages": [
{"role": "user", "content": "Summarize this section, preserving any dates, figures, and defined terms:\n\n'"$CHUNK_TEXT"'"}
]
}'
Run this per chunk, then feed the collected summaries into a second request that produces the final document-level summary. This keeps each individual call well within token limits and lets you parallelize the map step for speed. Full request and response shapes are documented at /docs/messages.
For very long outputs — a full report regenerated from 50 chunk summaries, for example — stream the final synthesis call instead of waiting for the whole response; see /docs/streaming for the implementation details.
Keeping track of chunk metadata
Attach metadata to every chunk before sending it: source document ID, page or section number, and chunk index. This is what lets you build citations later ("this answer comes from page 14, section 3.2") instead of a black-box summary. A simple approach is to prepend a short header to each chunk's text:
[doc: contract_v2.pdf | section: 3.2 | page: 14]
<chunk text here>
Claude will pick up on this context and can reference it directly in its response if you ask it to cite sources.
Choosing chunk size in practice
- Summarization: larger chunks (1,500–3,000 tokens) reduce the number of map calls and preserve more context per summary.
- Retrieval/Q&A: smaller chunks (300–800 tokens) improve retrieval precision because each chunk represents a narrower, more specific idea.
- Structured extraction (pulling fields, tables, clauses): chunk by logical unit — one chunk per clause or table — rather than by token count, so extraction prompts stay focused.
If you're running this pipeline against the Claude API through SubToAPI, each chunk call is a normal authenticated request with your sub_live_... key, and usage metadata on every response lets you see exactly how many tokens the chunking strategy is consuming per document — useful for tuning chunk size against cost. Get a key and test it against your own documents during the free trial at /signup, and check /pricing for plan details once you're running it at volume.
Questions
What chunk size should I use for the Claude API? Start around 1,000–1,500 tokens for summarization and 300–800 tokens for retrieval/Q&A, with structural boundaries (headings, paragraphs) respected rather than cutting at a hard character count.
Do I need overlap between chunks? Yes, for most documents. 10–15% overlap prevents losing context at chunk boundaries, especially in text with cross-references like contracts, manuals, or academic papers.
How do I summarize a document longer than Claude's context window? Use map-reduce: summarize each chunk independently, then run a second pass that synthesizes all chunk summaries into one final result. See /docs/messages for request structure and /docs/streaming for streaming the final synthesis output.