← Blog

Claude API PDF Summarizer: Build Tutorial

2026-10-05 · 5 min read · SubToAPI Team

Building a PDF summarizer with the Claude API comes down to three steps: extract the text from the PDF, send that text to Claude with a clear summarization prompt, and handle documents too long to fit in a single request. This tutorial walks through a working implementation in Node.js, including chunking for long PDFs and a few prompt choices that noticeably improve summary quality.

Claude doesn't read PDF files directly as binary uploads in a standard chat completion — you extract the text first, then pass it as plain text in the message content. The rest of this guide shows exactly how to do that end to end.

Step 1: Extract text from the PDF

Use a library like pdf-parse to pull raw text out of the file before touching the API:

npm install pdf-parse
const fs = require("fs");
const pdfParse = require("pdf-parse");

async function extractText(filePath) {
  const buffer = fs.readFileSync(filePath);
  const data = await pdfParse(buffer);
  return data.text;
}

This gives you a single string of the document's text. Expect some noise — headers, footers, page numbers repeated throughout — but Claude handles that fine as long as the core content is intact.

Step 2: Decide if you need to chunk

Claude's context window is large, so a 20-page report often fits in one request. But scanned PDFs, OCR output, or genuinely long documents (200+ pages, legal contracts, research compilations) can exceed what you want to send in a single call, especially if you also want to control cost per request.

A simple rule of thumb: if the extracted text is under ~60,000 characters, send it in one request. Above that, split it.

function chunkText(text, maxChars = 60000) {
  const chunks = [];
  for (let i = 0; i < text.length; i += maxChars) {
    chunks.push(text.slice(i, i + maxChars));
  }
  return chunks;
}

Splitting on raw character count is crude but reliable. If you want cleaner breaks, split on paragraph boundaries (\n\n) and accumulate chunks until you hit the limit, rather than cutting mid-sentence.

Step 3: Summarize each chunk, then combine

For multi-chunk documents, use a two-pass approach: summarize each chunk individually, then run a final pass that summarizes the summaries. This keeps each request focused and avoids Claude losing track of earlier sections in a very long single prompt.

async function summarizeChunk(apiUrl, apiKey, chunk, index, total) {
  const response = await fetch(`${apiUrl}/v1/messages`, {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4",
      max_tokens: 500,
      system: "You summarize document excerpts concisely and factually. Do not add opinions or outside information.",
      messages: [
        {
          role: "user",
          content: `This is section ${index + 1} of ${total} from a longer document. Summarize the key points in 3-5 bullet points.\n\n${chunk}`,
        },
      ],
    }),
  });
  const data = await response.json();
  return data.content[0].text;
}

Then combine:

async function summarizePdf(filePath, apiUrl, apiKey) {
  const text = await extractText(filePath);
  const chunks = chunkText(text);

  const partialSummaries = [];
  for (let i = 0; i < chunks.length; i++) {
    const summary = await summarizeChunk(apiUrl, apiKey, chunks[i], i, chunks.length);
    partialSummaries.push(summary);
  }

  if (chunks.length === 1) {
    return partialSummaries[0];
  }

  const finalResponse = await fetch(`${apiUrl}/v1/messages`, {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4",
      max_tokens: 800,
      system: "Combine these section summaries into one coherent executive summary.",
      messages: [
        { role: "user", content: partialSummaries.join("\n\n") },
      ],
    }),
  });
  const finalData = await finalResponse.json();
  return finalData.content[0].text;
}

This pattern — chunk, summarize each part, then roll up — scales to documents of any length and keeps each individual request cheap and fast.

Prompt design tips that matter

A few things make a real difference in summary quality:

Running this in production

If you're calling the API from a backend service rather than a script, you'll want centralized key management, usage visibility across whichever team members are hitting the summarizer, and a stable endpoint that doesn't change based on which underlying provider setup you're using. That's the gap SubToAPI fills: it turns your existing Claude access into a standard HTTPS API with its own sub_live_... application keys, so the summarizer code above works unmodified — just point apiUrl at https://api.subtoapi.app and use a key generated from your dashboard.

For streaming partial summaries back to a frontend as they're generated instead of waiting for the full response, see the streaming docs. For the full request and response shape used above, check the messages reference, and quickstart if you're setting up a key for the first time.

Error handling you shouldn't skip

PDF text extraction can produce garbage output on scanned or image-based PDFs with no embedded text layer — you'll get an empty or near-empty string back from pdf-parse. Check for that before calling the API:

if (text.trim().length < 50) {
  throw new Error("No extractable text found — this PDF may be scanned/image-based and needs OCR first.");
}

Also handle API-level failures (rate limits, oversized requests) with retries using exponential backoff, especially when summarizing many chunks in a loop — a single failed chunk shouldn't abort the whole batch.

questions

Can Claude read a PDF file directly without extracting text first? You need to extract the text yourself before sending it as part of the message content — this tutorial uses pdf-parse, but any text extraction library works the same way.

What's the best chunk size for summarizing long PDFs? Around 60,000 characters per chunk is a safe starting point for most models; adjust based on your max_tokens budget and how granular you want each section summary to be.

How do I keep summaries consistent across many documents? Fix the system prompt (tone, format, length) and reuse it across every call — varying the prompt per document is the most common cause of inconsistent summary quality.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →