Claude API GraphQL Wrapper: A Working Example
If you're searching for "claude api graphql api wrapper example," you're probably building a frontend or internal tool that already speaks GraphQL and you want Claude to be just another resolver, not a separate REST call bolted onto the side. This article walks through exactly that: a minimal, working GraphQL wrapper around Claude's API, including schema design, a resolver implementation, streaming over subscriptions, and the tradeoffs you'll hit in production.
The short answer: Claude's API itself is REST/JSON over HTTPS. There's no native GraphQL endpoint from Anthropic, so "wrapping" it means you write a thin GraphQL server (Apollo, GraphQL Yoga, etc.) whose resolvers call Claude's Messages API and shape the response into your schema. This is a common pattern when you already have a GraphQL gateway in front of several services and want Claude to fit the same contract as everything else.
Why wrap Claude in GraphQL at all
A few concrete reasons teams do this instead of calling the REST API directly from the client:
- Single graph for the frontend. If your app already queries users, orders, and documents through GraphQL, adding a
askAssistantmutation keeps the client code consistent instead of mixingfetchcalls to two different API styles. - Field-level shaping. You can expose only the fields you care about (
text,usage.inputTokens,stopReason) and hide the rest of Claude's response payload from the client. - Centralized auth and rate limiting. Your GraphQL gateway already handles JWT auth, so the Claude call inherits that for free instead of needing a separate layer.
- Batching and caching. Tools like DataLoader let you deduplicate identical prompts within a request, which matters if multiple resolvers in the same query happen to need the same completion.
None of this requires Anthropic to expose GraphQL — it's purely a server-side translation layer you own.
Schema design
Start with a schema that models a chat turn, not the raw Claude payload:
type Message {
role: String!
content: String!
}
type AssistantReply {
text: String!
stopReason: String
inputTokens: Int
outputTokens: Int
}
type Query {
_empty: String
}
type Mutation {
askAssistant(
prompt: String!
model: String = "claude-sonnet-4"
history: [Message!]
): AssistantReply!
}
type Subscription {
assistantStream(prompt: String!): String!
}
Keep the mutation input small. Resist the temptation to mirror every Claude parameter (system prompts, tool definitions, stop sequences) in the public schema unless clients actually need to control them — you can hardcode sane defaults server-side and widen the schema later.
Resolver implementation
Here's a resolver using Node and fetch against a generic Claude-compatible Messages endpoint:
const resolvers = {
Mutation: {
askAssistant: async (_, { prompt, model, history = [] }) => {
const messages = [...history.map(m => ({
role: m.role,
content: m.content
})), { role: "user", content: prompt }];
const res = await fetch("https://api.example.com/v1/messages", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${process.env.CLAUDE_API_KEY}`
},
body: JSON.stringify({ model, messages, max_tokens: 1024 })
});
const data = await res.json();
return {
text: data.content?.[0]?.text ?? "",
stopReason: data.stop_reason,
inputTokens: data.usage?.input_tokens,
outputTokens: data.usage?.output_tokens
};
}
}
};
This is the entire wrapper: a GraphQL mutation that validates input, calls the upstream Messages API, and reshapes the JSON into your typed schema. Error handling, retries, and timeout logic belong here too — GraphQL doesn't give you that for free.
Streaming through GraphQL subscriptions
REST streaming (Server-Sent Events) and GraphQL subscriptions don't map 1:1, so this is the part people usually get stuck on. The common approach is to consume the upstream SSE stream server-side and re-publish each token through a GraphQL PubSub:
const resolvers = {
Subscription: {
assistantStream: {
subscribe: async (_, { prompt }, { pubsub }) => {
const channel = `stream-${Date.now()}`;
fetch("https://api.example.com/v1/messages", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${process.env.CLAUDE_API_KEY}`
},
body: JSON.stringify({
model: "claude-sonnet-4",
stream: true,
max_tokens: 1024,
messages: [{ role: "user", content: prompt }]
})
}).then(res => readSSE(res.body, chunk => {
pubsub.publish(channel, { assistantStream: chunk });
}));
return pubsub.asyncIterator(channel);
}
}
}
};
readSSE here is a small helper that parses data: lines from the stream and extracts the text delta. This pattern works with Apollo Server, GraphQL Yoga, or any graphql-js setup that supports subscriptions over WebSockets.
Where SubToAPI fits
If you'd rather not run your own Claude key management, retry logic, and streaming plumbing underneath the wrapper, SubToAPI gives you a clean HTTPS Messages endpoint (sub_live_... keys) with streaming, tool use, and usage metadata already built in — your GraphQL resolver just points at https://api.subtoapi.app/v1/messages instead of managing Claude credentials directly. That keeps the wrapper above exactly the same; you only swap the base URL and auth header.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-4","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'
See the quickstart and messages docs for the full request shape, and streaming docs for the SSE format your readSSE helper needs to parse. If your wrapper needs to expose tool calls through GraphQL fields, the tools guide covers the request/response structure you'd map into your schema.
Practical tradeoffs to know before building this
- Latency stacking. Every hop (client → GraphQL gateway → Claude API) adds round-trip time. For chat UIs, prefer subscriptions/streaming over waiting for a full mutation response.
- Token costs aren't GraphQL's problem. Usage metadata should be returned as schema fields so your billing or analytics layer can read it without a second API call.
- Don't over-model Claude's response. Expose
text,stopReason, and usage fields; avoid leaking Anthropic's internal content-block structure into your public schema, since that structure can change. - Caching completions is risky. Unlike typical GraphQL caching, identical prompts can legitimately return different text. Only cache if your use case is deterministic (e.g., fixed system prompts with temperature 0).
questions
Does Claude's API have a native GraphQL endpoint? No. Claude's API is REST/JSON. A "GraphQL wrapper" means you build your own GraphQL server whose resolvers call the REST Messages API internally.
How do I handle streaming responses in GraphQL? Use GraphQL subscriptions. Consume the upstream Server-Sent Events stream in your resolver and republish each chunk through a PubSub so subscribed clients receive it over a WebSocket.
Is it worth building this wrapper, or should I just call the REST API directly? If your app already has a GraphQL gateway handling auth, caching, and client contracts, wrapping Claude keeps it consistent. If you're starting fresh with no existing graph, calling the Messages API directly is simpler and has less to maintain.