← Blog

Claude API Structured Data Extraction Example

2026-10-06 · 5 min read · SubToAPI Team

What "structured data extraction" means with Claude

When developers search for a Claude API structured data extraction example, they usually want one thing: a working pattern for turning unstructured text — invoices, emails, support tickets, resumes, product descriptions — into clean JSON that a downstream system can consume without manual parsing. Claude doesn't have a dedicated "extraction endpoint," but you get reliable structured output by combining a well-defined JSON schema, Claude's tool use feature, and a few defensive coding habits around parsing and validation.

This article walks through a complete, copy-pasteable example: extracting structured fields from a messy customer support email, including the schema, the API call, and the code that handles the response safely.

The core pattern: tool use as a schema enforcer

The most reliable way to get structured output from Claude is to define a "tool" whose input schema matches the shape of data you want back, then force Claude to call that tool instead of replying in free text. Claude doesn't execute the tool — you're just using the tool-call mechanism as a strict output contract.

Here's the schema for extracting structured data from a support email:

{
  "name": "extract_ticket_data",
  "description": "Extract structured fields from a customer support email",
  "input_schema": {
    "type": "object",
    "properties": {
      "customer_name": { "type": "string" },
      "issue_category": {
        "type": "string",
        "enum": ["billing", "technical", "account", "other"]
      },
      "urgency": {
        "type": "string",
        "enum": ["low", "medium", "high"]
      },
      "order_id": { "type": ["string", "null"] },
      "summary": { "type": "string" }
    },
    "required": ["customer_name", "issue_category", "urgency", "summary"]
  }
}

Full request example

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 500,
    "tools": [{
      "name": "extract_ticket_data",
      "description": "Extract structured fields from a customer support email",
      "input_schema": {
        "type": "object",
        "properties": {
          "customer_name": { "type": "string" },
          "issue_category": { "type": "string", "enum": ["billing","technical","account","other"] },
          "urgency": { "type": "string", "enum": ["low","medium","high"] },
          "order_id": { "type": ["string","null"] },
          "summary": { "type": "string" }
        },
        "required": ["customer_name","issue_category","urgency","summary"]
      }
    }],
    "tool_choice": { "type": "tool", "name": "extract_ticket_data" },
    "messages": [{
      "role": "user",
      "content": "Hi, this is Maria Delgado. My order #48213 hasnt arrived in 3 weeks and I was charged twice. This is extremely urgent, I need a refund today."
    }]
  }'

Setting tool_choice to force that specific tool removes the chance that Claude replies conversationally instead of returning structured data — this is the single biggest reliability improvement over just asking nicely in a prompt.

Parsing the response

Claude's reply will contain a tool_use content block instead of plain text. In JavaScript:

const response = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
    "content-type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 500,
    tools: [/* schema from above */],
    tool_choice: { type: "tool", name: "extract_ticket_data" },
    messages: [{ role: "user", content: emailText }],
  }),
});

const data = await response.json();
const toolBlock = data.content.find((block) => block.type === "tool_use");
const extracted = toolBlock.input; // already parsed JSON, matches your schema

Because tool_use input is returned as structured JSON (not a string you need to JSON.parse), you skip the fragile step of hoping Claude wraps its answer in matching braces and no trailing commentary.

Handling edge cases

A few things to build in from day one:

Running this behind a simple HTTPS API

The example above talks directly to Anthropic's API with a raw key, which works fine for a prototype. Once extraction logic moves into a product — a support dashboard, an internal tool, a billing pipeline — teams usually want normal REST semantics: scoped API keys, usage tracking per feature, and the ability to swap models without touching every client.

That's what SubToAPI is for: it takes your existing Claude access and exposes it as a standard HTTPS API with its own sub_live_... keys, so the extraction call above becomes a request to https://api.subtoapi.app/v1/messages with your SUBTOAPI_KEY, streaming and tool use supported, plus usage metadata and seats if more than one person on the team needs access. See the quickstart or the tools documentation for the exact request shape, or check pricing if you want this running under a dashboard instead of a bare API key in an env file.

Testing your extraction pipeline

Before shipping, run your schema against a handful of deliberately messy inputs: emails with typos, missing fields, mixed languages, or sarcasm that could confuse category classification. Log the raw tool_use.input for every production call for at least a few weeks — it's the fastest way to catch schema drift when real-world text doesn't match what you tested with.

Questions

Does Claude guarantee valid JSON output for extraction tasks? Not by default in plain text mode. Using tool use with a strict input_schema and forced tool_choice gets you reliably structured, schema-conformant JSON in practice, but you should still validate responses server-side.

Can I extract multiple records from one document in a single call? Yes — define an array property in your schema (e.g. items: { type: "array", items: {...} }) rather than making separate calls per record. It's faster and keeps related fields grouped correctly.

What's the difference between prompting for JSON and using tool use for extraction? Prompting for JSON relies on the model following formatting instructions in free text, which can drift on edge cases. Tool use enforces a schema at the API level and returns already-parsed structured input, making it the more reliable choice for production extraction pipelines.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →