Claude API PDF Document Analysis Tool: A Build Guide
If you're searching for a "Claude API PDF document analysis tool," you likely want one of two things: a ready-made way to feed PDFs into Claude and get structured answers back, or a clear explanation of how to build that capability yourself using the Messages API. Both are covered here.
The short answer: Claude's Messages API accepts PDF documents as a content type alongside text and images. You base64-encode the file, attach it as a document content block, and ask Claude a question about it in the same request. Claude reads the text and layout of the PDF — including tables, charts rendered as images, and scanned pages — and responds based on what it sees. No separate OCR pipeline or PDF-to-text library is required for most use cases.
How Claude reads PDF documents
Claude treats a PDF as a multimodal input. Internally it processes each page similarly to an image while also extracting embedded text, so it can answer questions that depend on layout (tables, headers, multi-column text) as well as plain content. This matters for document analysis tools because real-world PDFs — invoices, contracts, research papers, financial statements — rarely come as clean plain text.
A minimal request looks like this against a Claude-compatible Messages endpoint:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "BASE64_ENCODED_PDF"
}
},
{
"type": "text",
"text": "Extract the invoice number, total amount, and due date as JSON."
}
]
}
]
}'
That one call does extraction, parsing, and formatting in a single step — no regex, no table-parsing library.
Building the actual tool
A practical PDF analysis tool usually needs more than a single prompt. Here's the structure most teams end up with:
- Upload and encode. Accept the file upload, convert to base64 (or stream it if your SDK supports file references).
- Chunk if needed. Claude handles multi-page PDFs well, but very large documents (100+ pages) benefit from being split by section to stay within practical token and processing limits, and to keep latency predictable.
- Prompt for structure, not prose. Ask for JSON output with a defined schema instead of free text — this makes downstream parsing trivial and reduces hallucinated formatting.
- Validate the output. Even with structured prompting, validate the returned JSON against a schema before trusting it in a pipeline.
- Cache repeat queries. If multiple users ask different questions about the same document, cache the document content block server-side so you're not re-uploading the same bytes on every request.
A structured extraction prompt example:
const response = await client.messages.create({
model: "claude-sonnet-4",
max_tokens: 2048,
messages: [{
role: "user",
content: [
{ type: "document", source: { type: "base64", media_type: "application/pdf", data: pdfBase64 } },
{ type: "text", text: `Return only valid JSON matching this shape:
{ "parties": string[], "effective_date": string, "termination_clause": string, "renewal_terms": string }` }
]
}]
});
This pattern works well for contract review, invoice processing, resume parsing, and research paper summarization — any task where the input is a document and the output needs to be machine-readable.
Handling scanned and image-heavy PDFs
Scanned contracts or forms with no embedded text layer are still readable because Claude processes pages visually. Quality depends on scan resolution: blurry or skewed scans reduce accuracy. If you're building a production tool, it's worth flagging low-confidence extractions (e.g., missing required fields in the JSON output) for human review rather than assuming every page was read perfectly.
Costs and rate limits to plan around
Document analysis is more expensive per request than plain text chat because each page consumes a meaningful number of tokens. For a PDF analysis tool used by a team, this translates into real variability in monthly spend — a 50-page contract costs noticeably more to process than a one-page invoice. If multiple people on your team are hitting the API independently, tracking usage per person (not just per API key) becomes important for budgeting.
This is where a layer like SubToAPI is useful if you're turning Claude PDF analysis into an internal tool rather than a one-off script. It converts your existing Claude access into a standard HTTPS API with sub_live_... keys, so each teammate or environment gets its own key, usage shows up per key in one dashboard, and you're not sharing a single raw credential across a codebase. The request shape is the same Messages format:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "BASE64_ENCODED_PDF" } },
{ "type": "text", "text": "Summarize the key risks in this contract in three bullet points." }
]
}
]
}'
See the messages docs for the full request reference, and the quickstart if you're setting this up for the first time. For tools that chain document analysis with other steps — like calling a database to check extracted invoice numbers against existing records — the tool use guide covers function calling patterns that pair well with document extraction prompts.
When to use streaming
For long documents where you want the summary or analysis to appear incrementally rather than waiting for the full response, streaming is worth enabling, especially in a UI context. It doesn't reduce total processing time but improves perceived responsiveness for users reviewing long contracts or reports interactively.
A note on scope
A PDF document analysis tool built this way is good at extraction, summarization, comparison, and Q&A over document content. It is not a replacement for deterministic validation — always verify extracted financial figures, dates, and legal terms against the source before they drive automated decisions.
FAQs
Does Claude need a separate OCR step to read PDFs? No. Claude processes PDF pages directly, including scanned pages without a text layer, by treating each page similarly to an image while also reading any embedded text.
What's the best way to get consistent, parseable output from PDF analysis? Ask explicitly for JSON with a fixed schema in your prompt, and validate the response against that schema before using it downstream — this is more reliable than parsing free-form text.
How should I track API costs across a team using a PDF analysis tool? Give each person or environment its own API key rather than sharing one credential. SubToAPI's dashboard shows usage per key, which makes it easier to see which workflows are driving document-processing costs.