Claude API Image Input: Vision API Guide for Developers
Claude's vision capability lets you send images alongside text in a single API request and get back analysis, extracted text, descriptions, or structured data. If you're searching for "claude api image input vision api," you likely want to know how to actually format the request — what image formats are supported, how to encode them, and how to combine images with prompts. This guide covers all of that with working examples.
The short answer: Claude accepts images as part of the content array in a message, either as base64-encoded data or as a URL reference, alongside your text prompt. The model supports JPEG, PNG, GIF, and WebP, and you can include multiple images in one request for comparison or multi-page analysis.
How image input works in the Messages API
Instead of sending a plain string as the content field, you send an array of content blocks. Each block has a type — either text or image. This lets you interleave images and text naturally, for example "here's the first screenshot, here's the second, what changed?"
A basic request with a base64-encoded image looks like this:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRg..."
}
},
{
"type": "text",
"text": "What's shown in this image?"
}
]
}]
}'
If you're routing through SubToAPI instead of calling Anthropic directly, the request shape is identical — swap the endpoint and auth header:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo..."}},
{"type": "text", "text": "Extract the total amount from this receipt."}
]
}]
}'
See /docs/messages for the full request reference.
Encoding images for the request
Images must be base64-encoded before they go into the data field. In Node.js:
import fs from "fs";
const imageBuffer = fs.readFileSync("./screenshot.png");
const base64Image = imageBuffer.toString("base64");
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{
role: "user",
content: [
{ type: "image", source: { type: "base64", media_type: "image/png", data: base64Image } },
{ type: "text", text: "Does this UI match the design spec? List any discrepancies." }
]
}]
})
});
const data = await response.json();
console.log(data.content[0].text);
If your images are already hosted somewhere accessible, you can reference them by URL instead of inlining base64 data, which keeps request payloads smaller:
{
"type": "image",
"source": {
"type": "url",
"url": "https://example.com/images/chart.png"
}
}
Multiple images in one request
You can include several images in the same content array, which is useful for comparisons, before/after diffs, or multi-page document analysis. Order matters — Claude reads the content blocks top to bottom, so put each image near the text that refers to it:
{
"role": "user",
"content": [
{"type": "text", "text": "Page 1:"},
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "..."}},
{"type": "text", "text": "Page 2:"},
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "..."}},
{"type": "text", "text": "Summarize both pages into a single set of bullet points."}
]
}
There's a practical limit on how many images you should send per request — beyond 10-20 images, both latency and cost increase meaningfully, and accuracy on any single image can degrade. For bulk document processing, it's usually better to batch images into smaller groups and make multiple requests.
Common use cases
- Document and receipt data extraction — pull structured fields (totals, dates, line items) from photographed or scanned documents
- UI and design review — compare a screenshot against a design mock and flag differences
- Chart and diagram interpretation — describe trends or extract values from a chart image
- Accessibility alt-text generation — generate descriptive captions for images in bulk
- Content moderation — classify or describe uploaded images before they're published
For all of these, the response comes back as regular text in content[0].text, so downstream handling is the same as any other Claude response — you can ask for JSON output directly in your prompt if you need structured data.
Combining vision with streaming and tools
Vision input works with the same streaming and tool-use mechanics as text-only requests. If you want tokens streamed back as the model reasons over an image, set "stream": true and consume the event stream — see /docs/streaming for the format. If you want the model to call a function after analyzing an image (for example, saving extracted invoice data), you can define tools the same way you would for any other request — see /docs/tools.
If you're already using SubToAPI to turn your Claude access into an API with its own keys, streaming, and usage tracking, vision requests work through the same sub_live_ key and show up in the same usage dashboard as your text requests, so you don't need separate billing or monitoring for image-heavy workloads. Check /pricing if you're evaluating plans, or /docs/quickstart to get a key set up in a few minutes.
Image size and format limits
Keep a few practical constraints in mind:
- Supported formats: JPEG, PNG, GIF, WebP
- Very large images are automatically downscaled by the model, so there's little benefit to sending ultra-high-resolution originals — resizing to a reasonable max dimension (around 1568px on the long edge) before upload reduces payload size without hurting accuracy
- Base64 encoding adds roughly 33% overhead to the raw file size, so factor that into request size limits
- Each image consumes tokens as part of the context window — larger or more numerous images mean fewer tokens available for the rest of the conversation
questions
Does Claude API support image URLs or only base64? Both. You can pass "type": "base64" with inline encoded data, or "type": "url" with a publicly accessible image link in the source object.
What image formats can I send to Claude's vision API? JPEG, PNG, GIF, and WebP are supported. Convert other formats (like HEIC or BMP) before sending.
Can I send multiple images in a single Claude API request? Yes, add multiple image content blocks to the content array. Order them near related text for best results, and keep batches small for cost and latency.