Claude API Data Extraction from Documents Guide
Extracting structured data from documents with Claude
If you're searching for "claude api data extraction from documents," you're probably trying to pull structured fields — invoice totals, contract dates, form values, table rows — out of unstructured PDFs, scanned images, or text files and turn them into clean JSON. Claude can do this directly: you send the document (as text, base64-encoded PDF, or image) along with a prompt describing the schema you want, and Claude returns structured output you can parse programmatically.
This guide covers the practical mechanics: how to format the request, how to force consistent JSON output, how to handle multi-page documents and tables, and what to do when the model gets a field wrong. It's written for developers building extraction pipelines — invoice processing, resume parsing, receipt digitization, contract review — not for one-off manual lookups.
How document data extraction works with the Claude API
There are three common ways to get document content in front of Claude:
- Plain text — paste extracted text (from OCR, a PDF parser, or a
.txtfile) directly into the prompt. - PDF upload — send the PDF as a base64-encoded document block; Claude reads both text and layout.
- Image upload — send scanned pages as images (PNG/JPEG) when you don't have a text layer, e.g. scanned receipts.
For most structured extraction tasks, PDF or image input outperforms plain text because Claude can use visual layout — table boundaries, label positioning, checkboxes — to disambiguate fields that plain-text extraction would scramble.
Defining the schema
The single biggest factor in extraction accuracy is how precisely you describe the output schema. Vague instructions ("extract the important info") produce inconsistent results. Explicit instructions with field names, types, and fallback rules produce reliable JSON.
{
"invoice_number": "string",
"invoice_date": "YYYY-MM-DD",
"vendor_name": "string",
"line_items": [
{ "description": "string", "quantity": "number", "unit_price": "number", "total": "number" }
],
"subtotal": "number",
"tax": "number",
"total_due": "number",
"currency": "ISO 4217 code"
}
Give Claude this schema in the prompt and tell it explicitly to return null for missing fields rather than guessing — that one instruction alone eliminates a large share of hallucinated values.
Example: extracting invoice data
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"source": { "type": "base64", "media_type": "application/pdf", "data": "<BASE64_PDF>" }
},
{
"type": "text",
"text": "Extract the following fields as JSON only, no markdown fences: invoice_number, invoice_date (YYYY-MM-DD), vendor_name, line_items (description, quantity, unit_price, total), subtotal, tax, total_due, currency. If a field is missing, use null. Do not guess values."
}
]
}
]
}'
This example calls /v1/messages through SubToAPI, which proxies your existing Claude access with an sub_live_... API key and the same request shape documented in /docs/messages. If you're already on a Claude subscription and want this behind a standard HTTPS endpoint — with per-key usage tracking so you can see exactly how much document processing each client or project consumes — that's the core use case SubToAPI is built for. See /docs/quickstart to get a key in a few minutes.
Parsing the response reliably
Even with a clear schema, models occasionally wrap JSON in prose or markdown fences. Two fixes:
- Add "Respond with raw JSON only, no explanation, no code fences" to the prompt.
- On the client side, strip anything before the first
{and after the last}before callingJSON.parse().
function extractJson(text) {
const start = text.indexOf("{");
const end = text.lastIndexOf("}");
return JSON.parse(text.slice(start, end + 1));
}
For production pipelines, validate the parsed object against your schema (Zod, Joi, or a JSON Schema validator) and route anything that fails validation to a manual review queue instead of silently accepting bad data.
Handling multi-page and multi-document extraction
For long contracts or multi-page invoices, two approaches work well:
- Single call with full document: Claude's context window handles most multi-page PDFs in one request. This is simpler and preserves cross-page context (e.g., a total that references a page-3 subtotal).
- Chunked calls per page: split by page when documents are very long or when you need per-page provenance (which page a field came from). Merge results client-side.
If you need to process a batch of documents — hundreds of invoices overnight, for instance — streaming responses aren't necessary; a standard request/response loop with retries and rate limiting is simpler to operate. See /docs/streaming if you do want to stream longer extraction outputs for progress feedback in a UI.
Improving accuracy on messy documents
- Give few-shot examples. Show one example document-to-JSON pair in the prompt before the real document. This is especially effective for domain-specific formats like medical forms or legal filings.
- Ask for a confidence flag. Add a
"needs_review": booleanfield to the schema and instruct Claude to set ittruewhen any value was ambiguous or partially illegible. This gives you a cheap triage signal. - Normalize dates and numbers explicitly. Specify the exact date format and decimal convention (
.vs,) you want — don't leave it implicit. - Use tool calling for strict schemas. If you need guaranteed field names and types rather than free-text JSON, define the schema as a tool and let Claude call it; this constrains output structure more tightly than prompt instructions alone. Details are in
/docs/tools.
Building this into a production pipeline
A typical extraction pipeline looks like: ingest document → convert to base64/text → call Claude with schema prompt → validate JSON → write to database → flag needs_review rows for a human. The part most teams underestimate is observability: knowing which documents failed, which fields are low-confidence, and how extraction costs scale with document volume.
If you're routing this traffic through SubToAPI, usage metadata is returned with every response and visible per API key in the dashboard, so you can track extraction volume and cost by client, document type, or environment without building that instrumentation yourself. Plans start at €9/month for solo use, with team pricing at €19/seat and higher-volume Scale plans at €49/seat — see /pricing for details, or start a free trial at /signup.
Questions
Does Claude read scanned documents without a text layer? Yes — send scanned pages as images (PNG/JPEG) rather than relying on a PDF text layer. Claude processes the visual content directly, which works well for receipts and scanned forms, though very low-resolution scans reduce accuracy.
How do I guarantee valid JSON output every time? No LLM guarantees perfectly formed JSON on every call. Combine explicit "raw JSON only" instructions, client-side parsing that strips stray text, and schema validation with a fallback to manual review for failed cases.
Should I use tool calling or prompt-based JSON for extraction? Tool calling enforces stricter structure and is better for pipelines that need guaranteed field names and types. Prompt-based JSON is simpler to set up and sufficient for lower-stakes or exploratory extraction tasks. See /docs/tools for the tool-calling approach.