Claude API Image Input Support: A Developer Guide
Claude supports image input through its Messages API, letting you send photos, screenshots, diagrams, and scanned documents alongside text in a single request. The model can describe images, extract text, read charts, compare multiple images, and reason about visual content the same way it reasons about text — no separate vision endpoint required.
This guide covers exactly how to structure image input requests, which formats and sizes are supported, common pitfalls, and how to get the same functionality through SubToAPI if you're accessing Claude through an API wrapper instead of a direct Anthropic account.
How Image Input Works in the Messages API
Images are passed as content blocks inside a message, alongside or instead of text blocks. A single message can contain multiple images and text in any order. The model processes them together as one coherent input — so you can ask "what's different between these two screenshots?" and it will compare both.
There are two ways to provide image data:
- Base64-encoded bytes embedded directly in the request body
- A URL pointing to a publicly accessible image (if your provider supports URL fetching)
Base64 is the most reliable option because it doesn't depend on network access from the model provider's side and works with private or local files.
Basic Request Structure
{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": "iVBORw0KGgoAAAANSUhEUgAA..."
}
},
{
"type": "text",
"text": "What's shown in this image? Extract any visible text."
}
]
}
]
}
The order of blocks matters for readability but not for model comprehension — putting the image first and the instruction after (or vice versa) both work. Most developers put the image first, then the question, which mirrors how a human would look at a picture before being asked about it.
Supported Formats and Limits
Claude's image input supports the common web image formats:
image/jpegimage/pngimage/gif(static frame only — not animated)image/webp
Keep these practical limits in mind:
- File size: each image should stay under roughly 5MB after base64 encoding for best performance; larger files may be rejected or truncated depending on the endpoint.
- Image count: you can include multiple images per message, but each one consumes tokens based on its resolution, so very large batches increase cost and latency.
- Resolution: extremely high-resolution images get downscaled internally — there's no benefit to sending a 4000px image over a 1500px one for most use cases, and it costs more in tokens.
- Animated/video content: not supported — only static frames.
Encoding an Image to Base64
import fs from "fs";
const imageBuffer = fs.readFileSync("./screenshot.png");
const base64Image = imageBuffer.toString("base64");
const payload = {
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [
{
role: "user",
content: [
{
type: "image",
source: {
type: "base64",
media_type: "image/png",
data: base64Image,
},
},
{ type: "text", text: "Summarize the key metrics in this dashboard screenshot." },
],
},
],
};
Common Use Cases
- Document OCR and extraction: pull structured data (invoices, receipts, forms) into JSON without a separate OCR pipeline.
- UI review and bug reports: feed in a screenshot and ask Claude to describe layout issues or compare against a design spec.
- Chart and diagram reading: ask questions about values in a graph or flow in a diagram.
- Multi-image comparison: send a "before" and "after" screenshot in the same message and ask what changed.
- Accessibility tooling: generate alt text or descriptions for images at scale.
Combining Images With Tool Use
Image input works alongside tool calling. A common pattern is: send a screenshot, let Claude decide an image shows a bug, and have it call a create_ticket tool with a structured description. The image doesn't need to be treated specially in your tool schema — Claude reasons about the visual content first and then triggers the appropriate tool call, same as with text-only input.
Using Image Input Through SubToAPI
If you're calling Claude through SubToAPI instead of a raw Anthropic key, image input works through the same /v1/messages endpoint you already use for text and streaming — you just add image content blocks the same way shown above.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{ "type": "image", "source": { "type": "base64", "media_type": "image/jpeg", "data": "/9j/4AAQSk..." } },
{ "type": "text", "text": "Describe this image in one sentence." }
]
}
]
}'
This gives you the same usage metadata, team seat tracking, and streaming support for image-containing requests as for text-only ones — useful if multiple people on a team are sending screenshots or scanned documents and you need visibility into who's using what. See the messages docs for the full request schema, or check the quickstart if you're setting up a key for the first time. Pricing for Solo, Team, and Scale plans is on the pricing page.
Troubleshooting Checklist
If image input requests fail or return unexpected results:
- Confirm
media_typematches the actual file format (a PNG saved with a.jpgextension will fail). - Check that base64 data doesn't include a data URI prefix (
data:image/png;base64,) — strip that before sending. - Verify the image isn't corrupted or zero-byte before encoding.
- If using a URL source, confirm the URL is publicly reachable and returns the correct
Content-Typeheader. - Watch total token usage — large or multiple images can push a request close to context limits faster than expected.
FAQ
Can Claude read text inside images (OCR)? Yes. Claude can extract and transcribe text from screenshots, scanned documents, and photos of printed text directly within a normal image content block — no separate OCR service needed.
Can I send a PDF as an image? Not directly as an image block — PDFs need to be converted to image frames (one image per page) or handled through document-specific support if your provider offers it. Multi-page PDFs sent as raw images should be split into individual page images first.
Does image input cost more than text-only requests? Yes, images consume additional tokens proportional to their resolution, which affects billing. Downscaling large images before sending reduces cost without meaningfully affecting Claude's ability to read the content.