Claude API PDF Document Parsing Integration Guide
Parsing PDFs with Claude's API means sending the PDF as part of a message payload and letting the model read both the text and visual layout of the document — tables, scanned pages, charts, handwriting — without you running a separate OCR pipeline first. This guide covers exactly how to structure those requests, what limits to plan for, and how to avoid the most common integration mistakes.
If you've been searching for "claude api pdf document parsing integration," you're probably trying to decide between three approaches: sending raw extracted text, using a dedicated OCR service, or passing the PDF directly to Claude as a document block. The short answer: for most use cases — invoices, contracts, forms, research papers, scanned reports — passing the PDF directly to Claude gives better results with less code, because the model can reason about layout and visual structure, not just a flattened text dump.
How PDF parsing works with Claude's API
Claude supports PDFs as a content block type alongside text and images in the Messages API. Instead of pre-processing the file into plain text, you base64-encode the PDF (or reference it via URL, depending on your provider setup) and include it directly in the content array of your message, with a text instruction describing what you want extracted.
A typical request structure looks like this:
{
"model": "claude-sonnet-4-5",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "<base64-encoded-pdf>"
}
},
{
"type": "text",
"text": "Extract all line items from this invoice as a JSON array with fields: description, quantity, unit_price, total."
}
]
}
]
}
Claude processes each page as a combination of extracted text and a rendered image, which is why it handles scanned documents, rotated pages, and tables with merged cells better than plain-text extraction pipelines. You don't need to run Tesseract or any OCR library beforehand in most cases.
Practical integration steps
1. Convert your file source to base64. If your PDFs come from user uploads, S3, or a document management system, read them as a buffer and base64-encode before building the request body.
import fs from "fs";
const pdfBuffer = fs.readFileSync("./invoice.pdf");
const base64Pdf = pdfBuffer.toString("base64");
2. Write a specific extraction prompt. Vague prompts like "summarize this PDF" produce vague output. For structured parsing, specify the exact schema you want back, and ask for JSON only, no prose:
Extract the following fields from this contract and return valid JSON only:
- party_a
- party_b
- effective_date
- termination_clause (verbatim text)
- governing_law
3. Handle multi-page documents. Claude can process documents up to a defined page and size limit per request (check current limits in your provider's docs, as they change). For documents beyond that, split by logical sections — chapters, invoice batches, report sections — rather than arbitrary page counts, so each chunk retains context.
4. Parse and validate the response. Even with a JSON-only instruction, add a parsing step that catches malformed output and retries with a stricter prompt or lower temperature. Treat the model output like any other external API response — validate before writing to your database.
5. Cache or store extracted results. PDF parsing is usually a write-once, read-many operation. Store the structured output alongside a reference to the source file so you don't re-parse on every request.
Common pitfalls
- Sending the whole PDF when you only need one section wastes tokens and increases latency. If you know which pages matter, pre-split the PDF client-side.
- Not handling scanned image-only PDFs differently. These work fine with Claude's vision capability, but expect slightly higher variance in output accuracy on poor-quality scans — add a confidence check or human review step for critical fields like dollar amounts.
- Ignoring file size limits. Large PDFs (100+ pages, embedded images) can hit request size caps. Compress images within the PDF or split the document before sending.
- No retry/backoff strategy. Document parsing requests are often larger and slower than typical chat completions — build in reasonable timeouts and retry logic for production pipelines.
Routing PDF parsing through a managed API layer
If you're building this inside a team or product, you often need more than raw model access: per-app API keys, usage tracking per document type, and the ability to swap models without re-architecting your integration. This is where a layer like SubToAPI fits — it turns your existing Claude access into a standard HTTPS API with sub_live_... keys, so your PDF-parsing service can call /v1/messages the same way it would call any other REST API, with usage metadata and team seats managed centrally instead of scattered across individual accounts.
Example request through SubToAPI:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "..."}},
{"type": "text", "text": "Extract the total amount due and due date as JSON."}
]
}]
}'
This is especially useful if multiple services or team members need to call the same PDF-parsing pipeline — you issue separate keys per integration and track usage per key instead of sharing one account credential. See the quickstart and messages docs for the full request reference, and pricing for plan details if you're evaluating this for a team.
Checklist before going to production
- Validate base64 encoding and content-type headers match your actual file
- Set explicit
max_tokenshigh enough for your expected output size - Add JSON schema validation on the response, not just a try/catch on
JSON.parse - Log extraction confidence or add spot-check sampling for financial/legal documents
- Monitor token usage per document type to catch oversized files before they hit limits
FAQ
Can Claude parse scanned (image-only) PDFs, not just text-based ones? Yes. Claude renders PDF pages as images internally, so it can read scanned documents and handwriting to a reasonable degree, though accuracy on poor-quality scans is lower than on native text PDFs.
Do I need a separate OCR tool before sending a PDF to Claude? No, in most cases. Claude handles both the embedded text layer and visual layout natively. OCR pre-processing is only worth adding if you need extremely high-volume, low-cost extraction of simple text with no layout reasoning.
What's the best way to manage API keys for a PDF-parsing service used by multiple team members? Use a layer that issues separate application keys per integration rather than sharing one raw provider credential. SubToAPI does this through sub_live_... keys with per-key usage tracking — see /docs for setup details.