Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

Repository files navigation

rails-openrouter

A Ruby client for the OpenRouter API, built around streaming. One API key, 400+ models, and the same ergonomics as OpenRouter's official SDKs — client.chat.completions.create / .stream, chunk iteration, and a final completion assembled for you.

Text, images, PDFs, audio and video in; text and images out. No runtime dependencies: it is net/http and json from the standard library.

client = OpenRouter::Client.new(api_key: ENV["OPENROUTER_API_KEY"])

client.chat.completions.stream(
  model: "anthropic/claude-sonnet-4.5",
  messages: [{ role: "user", content: "Write a haiku about sockets" }]
).each_text { |text| print(text) }

Installation

# Gemfile
gem "rails-openrouter"

Or gem install rails-openrouter. Requires Ruby 3.0+.

The library is namespaced OpenRouter and has no Rails dependency — it works in any Ruby program. Bundler loads it for you; outside Bundler, require "openrouter" and require "rails-openrouter" both work.

Configuration

Pass options per client:

client = OpenRouter::Client.new(
  api_key: ENV["OPENROUTER_API_KEY"], # defaults to ENV["OPENROUTER_API_KEY"]
  site_url: "https://example.com",    # sent as HTTP-Referer, for leaderboards
  app_name: "My App",                 # sent as X-Title
  default_model: "openai/gpt-4o-mini",
  timeout: 600,                       # read timeout, seconds
  open_timeout: 10,
  max_retries: 2
)

…or globally, once, for OpenRouter.client:

OpenRouter.configure do |config|
  config.api_key = ENV["OPENROUTER_API_KEY"]
  config.app_name = "My App"
  config.default_model = "openai/gpt-4o-mini"
end

OpenRouter.chat.completions.create(messages: [{ role: "user", content: "Hi" }])

Clients are thread-safe; a single one can be shared across a web app's threads.

Chat completions

completion = client.chat.completions.create(
  model: "openai/gpt-4o-mini",
  messages: [
    { role: "system", content: "You are terse." },
    { role: "user", content: "Why is the sky blue?" }
  ],
  temperature: 0.2,
  max_tokens: 200
)

completion.choices.first.message.content
completion.usage.total_tokens
completion.to_h # plain Hash with symbol keys

Responses are OpenRouter::Structure objects: dot access, [] with strings or symbols, dig, and to_h. Missing keys return nil rather than raising, so chunk.choices.first.delta.content is safe on chunks that carry no text.

Every unrecognised keyword is forwarded to the API as-is, so OpenRouter-specific and newly shipped parameters work without a gem upgrade:

client.chat.completions.create(
  model: "anthropic/claude-sonnet-4.5",
  models: ["openai/gpt-4o", "google/gemini-2.5-pro"],  # fallbacks
  provider: { order: ["Anthropic"], allow_fallbacks: false },
  reasoning: { effort: "high" },
  transforms: ["middle-out"],
  plugins: [{ id: "web" }],
  usage: { include: true },
  messages: messages
)

Streaming

stream returns an OpenRouter::Stream, a lazy Enumerable over chunks. Nothing is sent until you start iterating.

stream = client.chat.completions.stream(model: model, messages: messages)

stream.each_text { |text| print(text) }        # only the text deltas
stream.each { |chunk| p chunk.choices.first }  # raw chunks
stream.each_reasoning { |text| print(text) }   # reasoning tokens

Chunks are accumulated as they pass through, so once the stream is consumed the assembled response is free:

stream.final_completion  # shaped exactly like a non-streaming response
stream.final_message     # .content, .tool_calls, .reasoning
stream.text              # the full text of choice 0
stream.usage             # prompt/completion/total tokens, and cost

snapshot gives you the partial completion mid-stream without consuming more.

Passing a block streams and returns the finished completion, which is the shortest form when you only want the side effect:

completion = client.chat.completions.create(model: model, messages: messages, stream: true) do |chunk|
  print(chunk.choices.first.delta.content)
end

Stop early and release the socket with close:

stream.each do |chunk|
  break if enough?(chunk)
end
stream.close

Iterating a stream again replays the chunks already seen and then continues from where it stopped — a second pass never issues a second request.

Keep-alive comments (: OPENROUTER PROCESSING, sent while a provider is still queueing) are handled internally and never surface as chunks.

Tool calls

Tool-call arguments arrive as fragments spread across chunks. The accumulator stitches them back together per call index:

stream = client.chat.completions.stream(messages: messages, tools: tools)
stream.each_text { |text| print(text) }

stream.final_message.tool_calls&.each do |call|
  args = JSON.parse(call.function.arguments)
  # dispatch call.function.name with args, append a role: "tool" message, loop
end

See examples/tool_calling.rb for the full round trip.

Multimodal: images, PDFs, audio, video

Attach files by handing them to Message.user(..., attach:), or by putting non-String objects (Pathname, IO, URI, Attachment) in a content array. The right content part is chosen from the file's MIME type: image_url for images, file for documents, input_audio for audio, video_url for video.

client.chat.completions.stream(
  model: "google/gemini-3-flash-preview",
  messages: [
    OpenRouter::Message.user("What changed between these?", attach: [
      "before.png",                          # local file, inlined as a data URL
      Pathname("after.png"),
      "https://example.com/spec.pdf",        # URL: OpenRouter fetches it
      { id: "or_file_abc" }                  # a file already uploaded
    ])
  ]
).each_text { |text| print(text) }

Strings in content are always text, never paths. A path is only read from disk when it arrives through attach:, Content.attach, or as a non-String type — so a user's own words can never be turned into a file read.

# text, not a file read:
{ role: "user", content: "please check ./report.pdf" }

# a file read, because you asked for one:
OpenRouter::Message.user("please check this", attach: ["./report.pdf"])

Build parts explicitly when you want to:

OpenRouter::Content.text("What is this?")
OpenRouter::Content.image("chart.png")                 # or a URL, IO, Pathname
OpenRouter::Content.file("report.pdf")
OpenRouter::Content.file(id: "or_file_abc")            # previously uploaded
OpenRouter::Content.audio("note.m4a")                  # base64 only, per the API
OpenRouter::Content.video("https://example.com/clip.mp4")
OpenRouter::Content.attach("whatever.ext")             # shape picked by MIME type

Supported out of the box: PNG/JPEG/WebP/GIF images; PDF and other documents; wav mp3 aiff aac ogg flac m4a pcm16 audio; mp4 mpeg mov webm video. Types are detected from the file extension, and from magic bytes when an IO or raw bytes arrive without a name. Override either with mime_type: / as:.

PDFs

pdf_engine: is sugar for the file-parser plugin:

client.chat.completions.create(
  model: model,
  messages: [OpenRouter::Message.user("Summarize", attach: ["report.pdf"])],
  pdf_engine: "native"        # model reads the file itself, billed as input tokens
  # "mistral-ocr"             # scanned pages and images, $2 per 1k pages
  # "cloudflare-ai"           # PDF to markdown, free
)

Parsing is charged per request, so for a document you will ask about more than once, append the assistant reply — annotations and all — to your message history and OpenRouter reuses the parse instead of redoing it:

messages << completion.choices.first.message.to_h   # includes annotations
messages << OpenRouter::Message.user("And the risks section?")

Uploading files once

file = client.files.upload("report.pdf")            # path, Pathname, IO or Attachment
file.id                                             # => "or_file_..."

client.files.list
client.files.retrieve(file.id)
client.files.download(file.id, to: "copy.pdf")      # server-created files only
client.files.delete(file.id)

Then reference it by id — no re-encoding on every request:

OpenRouter::Message.user("What is the total?", attach: [{ id: file.id }])

Which models accept what

client.models.list(input_modalities: %w[text image])   # vision models
client.models.list(input_modalities: "file")           # models that read documents
client.models.list(output_modalities: "image")         # models that draw

Media coming back

Image output arrives as data URLs on the message. Content.decode turns one into an Attachment you can save:

completion = client.chat.completions.create(
  model: "google/gemini-3-flash-image",
  messages: [{ role: "user", content: "A cat on a bicycle" }],
  modalities: %w[image text]
)

completion.choices.first.message.images.each_with_index do |image, index|
  OpenRouter::Content.decode(image).save("cat-#{index}.png")
end

This works on streams too — stream.final_message.images is assembled from the chunks like everything else.

Working with attachments directly

attachment = OpenRouter::Attachment.new("clip.mp4")
attachment.mime_type  # => "video/mp4"
attachment.kind       # => :video
attachment.size       # => 2_481_233
attachment.data_url   # => "data:video/mp4;base64,..."
attachment.save("copy.mp4")

Base64 inflates a file by about a third, and the whole thing is held in memory and in the request body. For anything large, prefer an https URL or client.files.upload. config.max_attachment_bytes sets a ceiling if you want one — attachments over it raise OpenRouter::AttachmentError rather than silently building a huge request. Audio is the one modality with no URL form in the API, so it is always inlined.

Other endpoints

client.models.list                                  # every routable model
client.models.list(supported_parameters: "tools")   # filtered
client.models.endpoints("openai/gpt-4o")            # providers, pricing, limits
client.credits.retrieve                             # purchased vs. used
client.key.retrieve                                 # limits for the current key
client.generations.retrieve(completion.id)          # cost accounting for one call
client.files.list                                   # uploaded files

Anything not wrapped yet is reachable through the low-level methods, which return plain hashes:

client.get("some/new/endpoint", query: { foo: "bar" })
client.post("some/new/endpoint", body: { foo: "bar" })
client.stream("some/new/endpoint", body: { foo: "bar", stream: true })

Errors

All errors descend from OpenRouter::Error.

Class Raised on
BadRequestError 400
AuthenticationError 401 — missing, invalid or expired key
InsufficientCreditsError 402
ModerationError 403 — input flagged
NotFoundError 404
RequestTimeoutError 408
RateLimitError 429
BadGatewayError 502 — provider returned garbage
NoProviderAvailableError 503 — no provider matches your routing
InternalServerError other 5xx
APITimeoutError the request timed out locally
APIConnectionError DNS/TCP/TLS failure
AttachmentError a file could not be read, or cannot be sent in that shape

Each APIError carries status, code, body, headers, request_id and metadata — provider failures put the upstream message in error.metadata[:raw], which is also folded into the exception message.

OpenRouter can also report a failure inside a 200 response, and mid-stream after chunks have already arrived. Both are raised as the same typed errors, so a rescue around your iteration is worth having.

Retries

408, 409, 429, 500, 502, 503, 504 and connection failures are retried up to max_retries times (default 2) with exponential backoff and jitter, honouring Retry-After when the server sends it. Streams are only retried while no chunk has reached your code yet — a half-delivered stream is never restarted, because that would duplicate output.

Development

rake test

The suite runs against a real in-process HTTP server (test/test_helper.rb), so chunked transfer encoding, split SSE frames, early disconnects and retries are exercised end to end. No network access and no stubbing library required.

License

MIT.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages