Claude API PDF Parsing Integration Guide
How PDF Parsing Works With the Claude API
Claude's API supports PDFs as native input through the Messages API — you send the file as a base64-encoded document content block alongside your text prompt, and Claude reads both the text and visual layout of each page (tables, charts, scanned text, headers) without any separate OCR step. This is different from older "extract text then feed to LLM" pipelines: Claude processes the PDF directly, which means it understands layout-dependent information like table columns, form fields, and figure captions that plain text extraction loses.
If you're integrating PDF parsing into a product, the practical questions are usually: how do you format the request, how many pages can you send, how much does it cost in tokens, and how do you get structured output back instead of free-form prose. This guide walks through all four with working code.
Sending a PDF to the API
A PDF is passed as a document content block with a base64 source, placed in the same content array as your instructions. Here's a minimal example against the Claude API directly:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "'"$(base64 -w0 invoice.pdf)"'"
}
},
{
"type": "text",
"text": "Extract the invoice number, total amount, and due date as JSON."
}
]
}]
}'
If you're routing requests through SubToAPI instead of managing raw Anthropic credentials, the request shape is the same — you just swap the endpoint and auth header:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "'"$(base64 -w0 invoice.pdf)"'"}},
{"type": "text", "text": "Extract the invoice number, total amount, and due date as JSON."}
]
}]
}'
This is useful if you want application-scoped sub_live_ keys, streaming, and usage dashboards rather than a single shared account key — see /docs/messages for the full request reference.
Structuring the Output
Free-form answers are fine for a demo but not for a pipeline. Push Claude toward predictable JSON by being explicit about the schema and telling it to skip commentary:
{
"role": "user",
"content": [
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "..." } },
{ "type": "text", "text": "Return ONLY valid JSON matching this schema, no explanation: {\"invoice_number\": string, \"total\": number, \"due_date\": \"YYYY-MM-DD\", \"line_items\": [{\"description\": string, \"amount\": number}]}" }
]
}
For higher reliability, pair this with Claude's tool use feature and define an extraction tool with a JSON schema as its input — Claude will call the tool with structured arguments instead of writing text you have to parse yourself. Details on defining tools are in /docs/tools.
Handling Multi-Page and Multi-File Documents
A few practical constraints to design around:
- Page limits. The API caps the number of pages per PDF (commonly up to 100 pages depending on model and size). For longer documents, split into chunks and process sequentially, carrying forward any needed context in your prompt.
- File size. Large PDFs (especially scanned, image-heavy ones) can hit request size limits. Compress or downsample before upload if you're hitting errors.
- Multiple documents in one request. You can include several
documentblocks in a single message if you need Claude to compare or cross-reference files — e.g., matching a purchase order against an invoice. - Mixed content. You can combine
documentblocks withimageblocks in the same request, useful when a contract PDF references an attached photo of a signed page.
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 2048,
messages: [{
role: "user",
content: [
{ type: "document", source: { type: "base64", media_type: "application/pdf", data: purchaseOrderBase64 } },
{ type: "document", source: { type: "base64", media_type: "application/pdf", data: invoiceBase64 } },
{ type: "text", text: "Compare these two documents. Flag any mismatched quantities or prices." }
]
}]
})
});
Managing Token Cost on PDF-Heavy Workloads
PDFs are tokenized per page, and page cost scales with visual complexity — a dense scanned table costs more than a single paragraph of plain text. A few ways to keep costs predictable:
- Pre-filter which pages you send. If you only need the summary page of a 40-page report, extract and send that page instead of the whole file.
- Cache repeated reference documents (policy manuals, templates you compare against repeatedly) using prompt caching where your workflow allows it.
- Set a tight
max_tokenson the response since extraction tasks rarely need long outputs. - If you're processing documents in bulk across a team, a usage dashboard that breaks down spend per key or per endpoint makes it much easier to catch a runaway pipeline before it burns through budget — this is one of the things SubToAPI's dashboard tracks per application key. See /pricing for plan details and /docs/quickstart to get a key provisioned.
Streaming for Long Documents
For large extraction jobs where you want partial results as they're generated (e.g., streaming a running summary while Claude works through a long contract), use the streaming endpoint instead of waiting for the full response. See /docs/streaming for the event format and how to parse server-sent events in your client.
A Simple Production Pattern
- Validate the uploaded file is a PDF and under your size limit before encoding.
- Base64-encode and send as a
documentblock with a schema-constrained prompt or a defined extraction tool. - Parse the JSON response, validate against your schema, and retry with a stricter prompt on failure.
- Log token usage per request so you can spot documents that are unusually expensive to process (often a sign of a bad scan or OCR-needed image rather than real text).
Questions
Can Claude read scanned (image-only) PDFs, not just text PDFs? Yes. Claude processes the PDF visually, similar to how it reads images, so it can read scanned pages and handwritten or low-quality text to a reasonable degree — accuracy depends on scan quality.
Does Claude preserve table structure when parsing PDFs? Generally yes — Claude reads the visual layout, so it can map table rows and columns correctly far better than naive text-extraction tools that lose column alignment.
Is there a file size or page limit for PDF uploads? Yes, both page count and file size are capped (page limits are commonly in the dozens-to-hundreds range depending on model). For longer documents, split into chunks or extract only the relevant pages before sending.