← Blog

Claude API Data Extraction from Documents Guide

2026-10-03 · 5 min read · SubToAPI Team

Extracting structured data from documents with Claude

If you're searching for "claude api data extraction from documents," you're probably trying to pull structured fields — invoice totals, contract dates, form values, table rows — out of unstructured PDFs, scanned images, or text files and turn them into clean JSON. Claude can do this directly: you send the document (as text, base64-encoded PDF, or image) along with a prompt describing the schema you want, and Claude returns structured output you can parse programmatically.

This guide covers the practical mechanics: how to format the request, how to force consistent JSON output, how to handle multi-page documents and tables, and what to do when the model gets a field wrong. It's written for developers building extraction pipelines — invoice processing, resume parsing, receipt digitization, contract review — not for one-off manual lookups.

How document data extraction works with the Claude API

There are three common ways to get document content in front of Claude:

  1. Plain text — paste extracted text (from OCR, a PDF parser, or a .txt file) directly into the prompt.
  2. PDF upload — send the PDF as a base64-encoded document block; Claude reads both text and layout.
  3. Image upload — send scanned pages as images (PNG/JPEG) when you don't have a text layer, e.g. scanned receipts.

For most structured extraction tasks, PDF or image input outperforms plain text because Claude can use visual layout — table boundaries, label positioning, checkboxes — to disambiguate fields that plain-text extraction would scramble.

Defining the schema

The single biggest factor in extraction accuracy is how precisely you describe the output schema. Vague instructions ("extract the important info") produce inconsistent results. Explicit instructions with field names, types, and fallback rules produce reliable JSON.

{
  "invoice_number": "string",
  "invoice_date": "YYYY-MM-DD",
  "vendor_name": "string",
  "line_items": [
    { "description": "string", "quantity": "number", "unit_price": "number", "total": "number" }
  ],
  "subtotal": "number",
  "tax": "number",
  "total_due": "number",
  "currency": "ISO 4217 code"
}

Give Claude this schema in the prompt and tell it explicitly to return null for missing fields rather than guessing — that one instruction alone eliminates a large share of hallucinated values.

Example: extracting invoice data

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 1024,
    "messages": [
      {
        "role": "user",
        "content": [
          {
            "type": "document",
            "source": { "type": "base64", "media_type": "application/pdf", "data": "<BASE64_PDF>" }
          },
          {
            "type": "text",
            "text": "Extract the following fields as JSON only, no markdown fences: invoice_number, invoice_date (YYYY-MM-DD), vendor_name, line_items (description, quantity, unit_price, total), subtotal, tax, total_due, currency. If a field is missing, use null. Do not guess values."
          }
        ]
      }
    ]
  }'

This example calls /v1/messages through SubToAPI, which proxies your existing Claude access with an sub_live_... API key and the same request shape documented in /docs/messages. If you're already on a Claude subscription and want this behind a standard HTTPS endpoint — with per-key usage tracking so you can see exactly how much document processing each client or project consumes — that's the core use case SubToAPI is built for. See /docs/quickstart to get a key in a few minutes.

Parsing the response reliably

Even with a clear schema, models occasionally wrap JSON in prose or markdown fences. Two fixes:

function extractJson(text) {
  const start = text.indexOf("{");
  const end = text.lastIndexOf("}");
  return JSON.parse(text.slice(start, end + 1));
}

For production pipelines, validate the parsed object against your schema (Zod, Joi, or a JSON Schema validator) and route anything that fails validation to a manual review queue instead of silently accepting bad data.

Handling multi-page and multi-document extraction

For long contracts or multi-page invoices, two approaches work well:

If you need to process a batch of documents — hundreds of invoices overnight, for instance — streaming responses aren't necessary; a standard request/response loop with retries and rate limiting is simpler to operate. See /docs/streaming if you do want to stream longer extraction outputs for progress feedback in a UI.

Improving accuracy on messy documents

Building this into a production pipeline

A typical extraction pipeline looks like: ingest document → convert to base64/text → call Claude with schema prompt → validate JSON → write to database → flag needs_review rows for a human. The part most teams underestimate is observability: knowing which documents failed, which fields are low-confidence, and how extraction costs scale with document volume.

If you're routing this traffic through SubToAPI, usage metadata is returned with every response and visible per API key in the dashboard, so you can track extraction volume and cost by client, document type, or environment without building that instrumentation yourself. Plans start at €9/month for solo use, with team pricing at €19/seat and higher-volume Scale plans at €49/seat — see /pricing for details, or start a free trial at /signup.

Questions

Does Claude read scanned documents without a text layer? Yes — send scanned pages as images (PNG/JPEG) rather than relying on a PDF text layer. Claude processes the visual content directly, which works well for receipts and scanned forms, though very low-resolution scans reduce accuracy.

How do I guarantee valid JSON output every time? No LLM guarantees perfectly formed JSON on every call. Combine explicit "raw JSON only" instructions, client-side parsing that strips stray text, and schema validation with a fallback to manual review for failed cases.

Should I use tool calling or prompt-based JSON for extraction? Tool calling enforces stricter structure and is better for pipelines that need guaranteed field names and types. Prompt-based JSON is simpler to set up and sufficient for lower-stakes or exploratory extraction tasks. See /docs/tools for the tool-calling approach.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →