← Blog

Build a Document Summarization Tool with Claude

2026-09-28 · 6 min read · SubToAPI Team

What You're Actually Building

A document summarization tool built on Claude is, at its core, three moving parts: a way to get text out of a document (PDF, DOCX, HTML, plain text), a prompt strategy that handles documents larger than a single context window, and an API layer that returns clean summaries to whatever frontend or workflow consumes them. This article walks through all three, with working code, so you can ship something that handles real-world documents rather than a demo that only works on a two-paragraph blog post.

The short answer for most teams: extract text, chunk it if it's long, send it to Claude with a summarization prompt that specifies format and length, and if you're calling the model from a product rather than a script, put it behind an HTTP API so your frontend and backend don't need SDK-specific logic scattered everywhere.

Step 1: Extract Text From the Source

Claude works with plain text (and, depending on your integration, with documents sent as file content). If you're building a general-purpose tool, you'll typically need to normalize inputs first:

import pdf from "pdf-parse";
import fs from "fs";

async function extractText(filePath) {
  const buffer = fs.readFileSync(filePath);
  const data = await pdf(buffer);
  return data.text;
}

Don't skip cleanup. Headers, footers, page numbers, and repeated boilerplate (common in exported reports) waste tokens and can confuse the summary. A simple regex pass to strip repeated lines across pages goes a long way.

Step 2: Decide Your Chunking Strategy

Claude's context window is large, but "large" isn't "unlimited," and very long inputs cost more per call and can dilute the summary quality if the document covers many unrelated topics. Two strategies work well:

Single-pass summarization — for documents that fit comfortably in one request (most reports, articles, contracts under a few hundred pages of text). Send the whole thing, ask for a structured summary.

Map-reduce summarization — for books, long transcripts, or multi-document sets:

  1. Split the document into sections (by heading, by page count, or by token count).
  2. Summarize each chunk independently.
  3. Combine the chunk summaries into a final prompt and ask Claude to produce one coherent summary from them.
function chunkText(text, maxChars = 12000) {
  const chunks = [];
  let start = 0;
  while (start < text.length) {
    chunks.push(text.slice(start, start + maxChars));
    start += maxChars;
  }
  return chunks;
}

Chunk boundaries matter — cutting mid-sentence is fine, cutting mid-table or mid-list item isn't. If your source has clear section headings, split on those instead of a fixed character count.

Step 3: Write a Prompt That Controls Format

The single biggest quality lever in a summarization tool isn't the model — it's the prompt. Vague prompts ("summarize this") produce vague summaries. Be explicit about length, structure, and what to do with ambiguity:

You are summarizing a business document for someone who has not read it.

Rules:
- Output 3-5 bullet points capturing the key facts, decisions, or findings.
- Follow the bullets with a 2-sentence "so what" paragraph explaining why this matters.
- If the document contains numbers, dates, or names, preserve them exactly.
- If a section is unclear or contradictory, note it briefly rather than guessing.
- Do not add information that isn't in the source.

Document:
{{document_text}}

For map-reduce, the "reduce" step prompt should explicitly say the input is a set of partial summaries, not the original document — otherwise Claude may try to re-summarize the summaries too aggressively and lose detail:

Below are summaries of consecutive sections of a longer document.
Combine them into one coherent summary of the whole document,
removing redundancy but keeping all distinct facts.

Step 4: Call the Model

If you're calling Claude directly, this is straightforward. If you want a single HTTP endpoint your frontend, backend, or automation scripts can all call without dealing with SDK setup, rate limits, or key rotation, SubToAPI wraps your existing Claude access in a REST API. You get an application key (sub_live_...), send standard JSON, and get streaming or non-streaming responses back.

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 500,
    "messages": [
      {"role": "user", "content": "Summarize this document in 4 bullet points:\n\n" }
    ]
  }'

For long documents where you want the summary to appear progressively in a UI rather than after a long wait, use streaming — see /docs/streaming for the event format. This matters more than it sounds: a 20-page report can take several seconds to summarize, and a streaming response makes the tool feel responsive instead of stuck.

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 500,
    stream: true,
    messages: [{ role: "user", content: summarizationPrompt }],
  }),
});

Full request/response shapes are in /docs/messages, and a minimal end-to-end example is in /docs/quickstart.

Step 5: Handle Batches and Multiple Documents

If your tool summarizes many documents (a folder of PDFs, an inbox of reports), run requests concurrently but cap concurrency — 5-10 in-flight requests is a reasonable default before you hit rate limits or start seeing degraded latency. Log token usage per document so you can predict cost at scale; usage metadata comes back with every response, which is useful if you're billing internal teams per document processed or just want to track spend.

Step 6: Add Guardrails

A summarization tool used in production needs a few things a demo doesn't:

If you're building this as part of a team tool where multiple people or services need their own scoped access, generating separate application keys per integration (one for the internal dashboard, one for a batch job, one for a Slack bot) makes usage easier to audit — see /pricing for how seats and keys map to plans.

questions

Do I need to chunk every document, or only long ones? Only chunk when the document exceeds what you can comfortably send in one request — most single-file summarization (reports, articles, contracts) works fine as one call. Reserve map-reduce chunking for books, long transcripts, or multi-file inputs.

How do I stop Claude from adding information that isn't in the source? Say so explicitly in the prompt ("do not add information that isn't in the source") and, for high-stakes use cases, ask it to quote or reference the specific part of the text a claim comes from. Explicit instructions reduce this significantly.

What's the fastest way to get a working prototype without managing SDKs or infrastructure? Send documents as plain text to a hosted Claude API endpoint like SubToAPI's /v1/messages, using an application key from /signup. You get streaming, usage tracking, and team keys without setting up your own proxy or auth layer.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →