Claude API Image Understanding: Vision API Guide
What "Claude API image understanding" actually means
If you're searching for "Claude API image understanding vision API," you likely want to know one of two things: whether Claude can analyze images through its API (yes), and how to actually send images and get useful results back. Claude's vision capability lets you pass images alongside text in a single message, and the model will describe, analyze, compare, or extract information from what it sees — charts, screenshots, photos, diagrams, handwritten notes, UI mockups, and more.
This isn't a separate "vision API" with its own endpoint. It's the same Messages endpoint you already use for text, with an additional content block type for images. That's the key technical detail most people miss: you don't call a different service, you just structure your request differently.
How image input works in the Claude API
A message to Claude can contain multiple content blocks. For text-only prompts, you send a single text block. For vision tasks, you add one or more image blocks alongside your text, inside the same content array.
Images are sent either as base64-encoded data or as a URL (depending on the provider's support), with a required media type like image/jpeg, image/png, image/gif, or image/webp.
A typical request body looks like this:
{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": "iVBORw0KGgoAAAANSUhEUgAA..."
}
},
{
"type": "text",
"text": "What does this chart show, and what's the trend over time?"
}
]
}
]
}
Claude reads the image and the text together, so you can ask follow-up questions, request structured extraction ("return the data as JSON"), or combine multiple images in one request for comparison.
Common use cases for image understanding
- Document and form extraction — reading scanned invoices, receipts, or ID cards and returning structured fields
- UI/UX review — analyzing screenshots of apps or websites and flagging usability issues
- Chart and graph interpretation — summarizing trends from exported dashboard images
- Accessibility tooling — generating alt text or descriptions for images at scale
- Visual QA — comparing a design mockup against a shipped screenshot
Sending images through SubToAPI
If you're already using SubToAPI to turn your Claude access into a standard HTTPS API, image understanding works through the same Messages endpoint you use for everything else — no separate setup, no extra product to learn.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRgABAQAAAQABAAD..."
}
},
{
"type": "text",
"text": "Extract the total amount and date from this receipt."
}
]
}
]
}'
The response comes back in the same structure as any text completion, with usage metadata included so you can track token consumption per request — useful when you're processing images in bulk and want to keep an eye on cost. See /docs/messages for the full request and response reference, and /docs/quickstart if you're setting up your first API key.
Encoding images in JavaScript
If you're building an app that lets users upload images, you'll usually need to base64-encode the file before sending it:
import fs from "fs";
const imageBuffer = fs.readFileSync("./receipt.jpg");
const base64Image = imageBuffer.toString("base64");
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [
{
role: "user",
content: [
{
type: "image",
source: { type: "base64", media_type: "image/jpeg", data: base64Image }
},
{ type: "text", text: "Describe what's happening in this image." }
]
}
]
})
});
const data = await response.json();
console.log(data);
Practical limits and tips
A few things worth knowing before you build around image understanding:
- Image size matters for both accuracy and cost. Very large images get resized internally, so extremely high-resolution scans don't necessarily improve results — downscaling to a reasonable resolution (under ~1568px on the longest side) before sending often works just as well and reduces payload size.
- Multiple images in one request are supported. You can send several images in the same
contentarray to ask Claude to compare them or extract consistent information across a set. - Text extraction isn't OCR-perfect. Claude is very good at reading printed and handwritten text in images, but for mission-critical structured extraction (legal documents, financial records), validate outputs rather than trusting them blindly.
- Combine vision with tool use carefully. If you want Claude to extract data from an image and then call a function with that data, you can combine image blocks with tool definitions in the same request — see /docs/tools for how tool use is structured.
- Streaming works with image inputs too. If you want partial results as Claude analyzes a complex image, check /docs/streaming for how to handle server-sent events on these requests.
If you're evaluating plans for a project that processes a lot of images — say, a document pipeline or a visual QA tool — check /pricing to see how Solo, Team, and Scale tiers compare, and start with the free trial at /signup to test image requests before committing to a plan.
questions
Does Claude have a separate vision API endpoint? No. Image understanding is handled through the same Messages endpoint used for text. You add an image content block alongside your text block in the same request.
What image formats does Claude support? JPEG, PNG, GIF, and WebP are supported, sent as base64-encoded data with the correct media_type specified in the request.
Can I send multiple images in one request? Yes. You can include several image blocks in the same content array, which lets Claude compare images or extract consistent information across a set.