Claude API Ruby on Rails Integration Guide
Integrating the Claude API into a Ruby on Rails application means adding an HTTP client that talks to Anthropic's Messages API (or a compatible gateway), wrapping it in a service object, and deciding how you'll handle streaming, retries, and background processing. There's no official Anthropic gem for Ruby, so most Rails teams either call the REST API directly with Faraday or Net::HTTP, or route requests through a service that exposes a simpler interface.
This guide walks through a practical setup: a ClaudeClient service class, a Rails controller that streams responses to the browser, background job handling for long-running completions, and common pitfalls specific to Rails apps (timeouts, connection pooling, credential storage). If you want the API surface simplified — regular API keys, usage metadata, and team billing in one dashboard instead of managing Anthropic console access per developer — a layer like SubToAPI can sit underneath the same Rails code with no changes beyond the base URL and key.
Setting up credentials
Store your API key in Rails credentials rather than .env files committed to version control:
EDITOR="code --wait" bin/rails credentials:edit
claude:
api_key: sub_live_xxxxxxxxxxxxxxxxxxxx
base_url: https://api.subtoapi.app/v1
Access it anywhere with Rails.application.credentials.claude[:api_key]. Using sub_live_ style keys through SubToAPI means you can also manage per-environment keys (staging vs production) from one dashboard instead of juggling multiple Anthropic console projects — see /docs/quickstart for the exact setup.
A service object for Claude requests
Keep the HTTP logic out of controllers. A simple Faraday-based client looks like this:
# app/services/claude_client.rb
class ClaudeClient
BASE_URL = Rails.application.credentials.claude[:base_url]
API_KEY = Rails.application.credentials.claude[:api_key]
def self.connection
Faraday.new(url: BASE_URL) do |f|
f.request :json
f.response :json
f.adapter Faraday.default_adapter
f.options.timeout = 30
f.options.open_timeout = 5
end
end
def self.messages(model:, messages:, max_tokens: 1024, system: nil)
connection.post("/messages") do |req|
req.headers["Authorization"] = "Bearer #{API_KEY}"
req.headers["Content-Type"] = "application/json"
req.body = {
model: model,
max_tokens: max_tokens,
system: system,
messages: messages
}.compact
end
end
end
Calling it from a controller or service is a single line:
response = ClaudeClient.messages(
model: "claude-sonnet-4",
messages: [{ role: "user", content: "Summarize this ticket: #{ticket.body}" }]
)
response.body["content"].first["text"]
This structure matches the documented request shape at /docs/messages, so swapping between a direct Anthropic integration and a SubToAPI-backed one only requires changing BASE_URL and the key format.
Handling timeouts and retries
Rails apps often hit default Puma or Rack timeouts before a Claude completion finishes, especially for longer prompts. Two practical fixes:
- Increase the Faraday timeout for requests known to be long-running (already set to 30s above; bump to 60-120s for complex prompts).
- Add retry logic for transient 429/5xx responses using
faraday-retry:
Faraday.new(url: BASE_URL) do |f|
f.request :retry, max: 3, interval: 0.5,
exceptions: [Faraday::TimeoutError, Faraday::ConnectionFailed],
retry_statuses: [429, 500, 502, 503]
f.request :json
f.response :json
f.adapter Faraday.default_adapter
end
Streaming responses in a Rails controller
Streaming is the main reason naive Rails integrations feel sluggish — if you wait for the full completion before rendering anything, users stare at a spinner. Rails supports ActionController::Live for server-sent events:
class ChatController < ApplicationController
include ActionController::Live
def stream
response.headers["Content-Type"] = "text/event-stream"
conn = Faraday.new(url: ClaudeClient::BASE_URL)
conn.post("/messages/stream") do |req|
req.headers["Authorization"] = "Bearer #{ClaudeClient::API_KEY}"
req.headers["Content-Type"] = "application/json"
req.body = {
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [{ role: "user", content: params[:prompt] }]
}.to_json
req.options.on_data = Proc.new do |chunk, _|
response.stream.write("data: #{chunk}\n\n")
end
end
rescue IOError
# client disconnected
ensure
response.stream.close
end
end
Make sure Puma is configured with enough threads to handle long-lived streaming connections without starving other requests — a separate thread pool or a dedicated worker tier is common in production. See /docs/streaming for event format details if you're streaming through SubToAPI's SSE endpoint.
Background jobs for non-interactive completions
For anything that doesn't need to render live in the browser — generating report summaries, classifying support tickets, enriching records — push the call into Sidekiq or ActiveJob instead of blocking a web request:
class SummarizeTicketJob < ApplicationJob
queue_as :default
def perform(ticket_id)
ticket = Ticket.find(ticket_id)
response = ClaudeClient.messages(
model: "claude-sonnet-4",
messages: [{ role: "user", content: "Summarize: #{ticket.body}" }]
)
ticket.update!(ai_summary: response.body["content"].first["text"])
end
end
This keeps web dynos responsive and gives you natural retry semantics through Sidekiq's job retry mechanism, separate from the HTTP-level retries in the client.
Tool use from Rails
If your Rails app needs Claude to call back into your own methods — looking up a record, running a calculation — define tools in the request body and handle the tool_use stop reason in your service object:
if response.body["stop_reason"] == "tool_use"
tool_call = response.body["content"].find { |c| c["type"] == "tool_use" }
result = MyApp::ToolRouter.call(tool_call["name"], tool_call["input"])
# send result back as a tool_result message
end
Full request/response shapes for tool definitions are documented at /docs/tools.
Choosing between direct Anthropic access and a gateway
Going straight to Anthropic's API works fine for a single developer or small app. It gets harder once you have a team: each developer needs console access, there's no per-project usage breakdown, and rotating keys means touching Anthropic's dashboard directly. Routing Rails traffic through SubToAPI keeps the same REST shape (/v1/messages, same JSON fields) while adding team seats, per-key usage metadata, and a single dashboard for billing across Solo, Team, and Scale plans. You can start on the free trial at /signup and compare plans at /pricing without changing any of the Rails code above beyond the base URL.
questions
Does Anthropic provide an official Ruby gem for Rails? No. Anthropic maintains official SDKs for Python and TypeScript/JavaScript; Ruby integrations typically use Faraday or Net::HTTP against the REST API directly, as shown above.
How do I stream Claude responses to a Rails view without extra JavaScript libraries? Use ActionController::Live with server-sent events on the backend and a plain EventSource in the browser — no additional gems or frontend frameworks required.
Should Claude API calls run in the web request or a background job? Interactive chat UIs need the web request (ideally streamed); anything that doesn't need live output — summarization, classification, batch enrichment — should run through Sidekiq or ActiveJob to avoid tying up web workers.