← Blog

Claude API Image Understanding: Vision API Guide

2026-10-03 · 4 min read · SubToAPI Team

What "Claude API image understanding" actually means

If you're searching for "Claude API image understanding vision API," you likely want to know one of two things: whether Claude can analyze images through its API (yes), and how to actually send images and get useful results back. Claude's vision capability lets you pass images alongside text in a single message, and the model will describe, analyze, compare, or extract information from what it sees — charts, screenshots, photos, diagrams, handwritten notes, UI mockups, and more.

This isn't a separate "vision API" with its own endpoint. It's the same Messages endpoint you already use for text, with an additional content block type for images. That's the key technical detail most people miss: you don't call a different service, you just structure your request differently.

How image input works in the Claude API

A message to Claude can contain multiple content blocks. For text-only prompts, you send a single text block. For vision tasks, you add one or more image blocks alongside your text, inside the same content array.

Images are sent either as base64-encoded data or as a URL (depending on the provider's support), with a required media type like image/jpeg, image/png, image/gif, or image/webp.

A typical request body looks like this:

{
  "model": "claude-sonnet-4",
  "max_tokens": 1024,
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "image",
          "source": {
            "type": "base64",
            "media_type": "image/png",
            "data": "iVBORw0KGgoAAAANSUhEUgAA..."
          }
        },
        {
          "type": "text",
          "text": "What does this chart show, and what's the trend over time?"
        }
      ]
    }
  ]
}

Claude reads the image and the text together, so you can ask follow-up questions, request structured extraction ("return the data as JSON"), or combine multiple images in one request for comparison.

Common use cases for image understanding

Sending images through SubToAPI

If you're already using SubToAPI to turn your Claude access into a standard HTTPS API, image understanding works through the same Messages endpoint you use for everything else — no separate setup, no extra product to learn.

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 1024,
    "messages": [
      {
        "role": "user",
        "content": [
          {
            "type": "image",
            "source": {
              "type": "base64",
              "media_type": "image/jpeg",
              "data": "/9j/4AAQSkZJRgABAQAAAQABAAD..."
            }
          },
          {
            "type": "text",
            "text": "Extract the total amount and date from this receipt."
          }
        ]
      }
    ]
  }'

The response comes back in the same structure as any text completion, with usage metadata included so you can track token consumption per request — useful when you're processing images in bulk and want to keep an eye on cost. See /docs/messages for the full request and response reference, and /docs/quickstart if you're setting up your first API key.

Encoding images in JavaScript

If you're building an app that lets users upload images, you'll usually need to base64-encode the file before sending it:

import fs from "fs";

const imageBuffer = fs.readFileSync("./receipt.jpg");
const base64Image = imageBuffer.toString("base64");

const response = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-sonnet-4",
    max_tokens: 1024,
    messages: [
      {
        role: "user",
        content: [
          {
            type: "image",
            source: { type: "base64", media_type: "image/jpeg", data: base64Image }
          },
          { type: "text", text: "Describe what's happening in this image." }
        ]
      }
    ]
  })
});

const data = await response.json();
console.log(data);

Practical limits and tips

A few things worth knowing before you build around image understanding:

If you're evaluating plans for a project that processes a lot of images — say, a document pipeline or a visual QA tool — check /pricing to see how Solo, Team, and Scale tiers compare, and start with the free trial at /signup to test image requests before committing to a plan.

questions

Does Claude have a separate vision API endpoint? No. Image understanding is handled through the same Messages endpoint used for text. You add an image content block alongside your text block in the same request.

What image formats does Claude support? JPEG, PNG, GIF, and WebP are supported, sent as base64-encoded data with the correct media_type specified in the request.

Can I send multiple images in one request? Yes. You can include several image blocks in the same content array, which lets Claude compare images or extract consistent information across a set.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →