Build a Document Summarization Tool with Claude
What You're Actually Building
A document summarization tool built on Claude is, at its core, three moving parts: a way to get text out of a document (PDF, DOCX, HTML, plain text), a prompt strategy that handles documents larger than a single context window, and an API layer that returns clean summaries to whatever frontend or workflow consumes them. This article walks through all three, with working code, so you can ship something that handles real-world documents rather than a demo that only works on a two-paragraph blog post.
The short answer for most teams: extract text, chunk it if it's long, send it to Claude with a summarization prompt that specifies format and length, and if you're calling the model from a product rather than a script, put it behind an HTTP API so your frontend and backend don't need SDK-specific logic scattered everywhere.
Step 1: Extract Text From the Source
Claude works with plain text (and, depending on your integration, with documents sent as file content). If you're building a general-purpose tool, you'll typically need to normalize inputs first:
- PDF — use
pdf-parse(Node) orpypdf(Python) to pull raw text. - DOCX —
mammoth(Node) converts to clean HTML/text. - HTML/web pages — strip tags, keep headings if you want structure-aware summaries.
- Plain text/Markdown — use as-is.
import pdf from "pdf-parse";
import fs from "fs";
async function extractText(filePath) {
const buffer = fs.readFileSync(filePath);
const data = await pdf(buffer);
return data.text;
}
Don't skip cleanup. Headers, footers, page numbers, and repeated boilerplate (common in exported reports) waste tokens and can confuse the summary. A simple regex pass to strip repeated lines across pages goes a long way.
Step 2: Decide Your Chunking Strategy
Claude's context window is large, but "large" isn't "unlimited," and very long inputs cost more per call and can dilute the summary quality if the document covers many unrelated topics. Two strategies work well:
Single-pass summarization — for documents that fit comfortably in one request (most reports, articles, contracts under a few hundred pages of text). Send the whole thing, ask for a structured summary.
Map-reduce summarization — for books, long transcripts, or multi-document sets:
- Split the document into sections (by heading, by page count, or by token count).
- Summarize each chunk independently.
- Combine the chunk summaries into a final prompt and ask Claude to produce one coherent summary from them.
function chunkText(text, maxChars = 12000) {
const chunks = [];
let start = 0;
while (start < text.length) {
chunks.push(text.slice(start, start + maxChars));
start += maxChars;
}
return chunks;
}
Chunk boundaries matter — cutting mid-sentence is fine, cutting mid-table or mid-list item isn't. If your source has clear section headings, split on those instead of a fixed character count.
Step 3: Write a Prompt That Controls Format
The single biggest quality lever in a summarization tool isn't the model — it's the prompt. Vague prompts ("summarize this") produce vague summaries. Be explicit about length, structure, and what to do with ambiguity:
You are summarizing a business document for someone who has not read it.
Rules:
- Output 3-5 bullet points capturing the key facts, decisions, or findings.
- Follow the bullets with a 2-sentence "so what" paragraph explaining why this matters.
- If the document contains numbers, dates, or names, preserve them exactly.
- If a section is unclear or contradictory, note it briefly rather than guessing.
- Do not add information that isn't in the source.
Document:
{{document_text}}
For map-reduce, the "reduce" step prompt should explicitly say the input is a set of partial summaries, not the original document — otherwise Claude may try to re-summarize the summaries too aggressively and lose detail:
Below are summaries of consecutive sections of a longer document.
Combine them into one coherent summary of the whole document,
removing redundancy but keeping all distinct facts.
Step 4: Call the Model
If you're calling Claude directly, this is straightforward. If you want a single HTTP endpoint your frontend, backend, or automation scripts can all call without dealing with SDK setup, rate limits, or key rotation, SubToAPI wraps your existing Claude access in a REST API. You get an application key (sub_live_...), send standard JSON, and get streaming or non-streaming responses back.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 500,
"messages": [
{"role": "user", "content": "Summarize this document in 4 bullet points:\n\n" }
]
}'
For long documents where you want the summary to appear progressively in a UI rather than after a long wait, use streaming — see /docs/streaming for the event format. This matters more than it sounds: a 20-page report can take several seconds to summarize, and a streaming response makes the tool feel responsive instead of stuck.
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 500,
stream: true,
messages: [{ role: "user", content: summarizationPrompt }],
}),
});
Full request/response shapes are in /docs/messages, and a minimal end-to-end example is in /docs/quickstart.
Step 5: Handle Batches and Multiple Documents
If your tool summarizes many documents (a folder of PDFs, an inbox of reports), run requests concurrently but cap concurrency — 5-10 in-flight requests is a reasonable default before you hit rate limits or start seeing degraded latency. Log token usage per document so you can predict cost at scale; usage metadata comes back with every response, which is useful if you're billing internal teams per document processed or just want to track spend.
Step 6: Add Guardrails
A summarization tool used in production needs a few things a demo doesn't:
- Length limits — cap input size and reject or auto-chunk documents above it.
- Fact-checking prompts for sensitive use cases — for legal or financial documents, ask Claude to flag anything it's summarizing with low confidence rather than smoothing it over.
- Consistent output format — if downstream systems parse the summary, ask for JSON output with defined fields rather than free text.
- Retry logic — network or rate-limit errors happen; retry with backoff rather than failing the whole batch.
If you're building this as part of a team tool where multiple people or services need their own scoped access, generating separate application keys per integration (one for the internal dashboard, one for a batch job, one for a Slack bot) makes usage easier to audit — see /pricing for how seats and keys map to plans.
questions
Do I need to chunk every document, or only long ones? Only chunk when the document exceeds what you can comfortably send in one request — most single-file summarization (reports, articles, contracts) works fine as one call. Reserve map-reduce chunking for books, long transcripts, or multi-file inputs.
How do I stop Claude from adding information that isn't in the source? Say so explicitly in the prompt ("do not add information that isn't in the source") and, for high-stakes use cases, ask it to quote or reference the specific part of the text a claim comes from. Explicit instructions reduce this significantly.
What's the fastest way to get a working prototype without managing SDKs or infrastructure? Send documents as plain text to a hosted Claude API endpoint like SubToAPI's /v1/messages, using an application key from /signup. You get streaming, usage tracking, and team keys without setting up your own proxy or auth layer.