Skip to content
r-uby-devPublic

Repository files navigation

a r.uby.dev project

r.uby.dev project.

Welcome to the canonical llm.rb repository.

llm.rb is a runtime for building agentic AI applications on CRuby. It has zero runtime dependencies by default, supports concurrent and parallel tool execution and has a single coherent API that spans 14+ providers. The README covers a lot of ground and the changelog tracks what has changed between releases.

If you want to see the runtime in action the r.uby.dev website provides a platform where you can - for free. It hosts multiple llm.rb agents, and one of them (bezela) helps manage this repository.

Install

llm.rb requires Ruby 3.4 or later.

gem install llm.rb

Quick start

Agents

The LLM::Agent class is the default high-level interface, and it is recommended for most use-cases. It manages the tool loop and provides configurable features on top of it. For example you can manage the tool loop with a retry budget and a tool call budget - alongside other features.

The runtime is designed to keep the tool loop alive and it will rescue exceptions. When an exception is encountered in a tool it is reported back to the model as an in-band error that allows the model to change course or retry with different parameters.

A tool call requires a tool return (or response), and the lack of one can corrupt the conversation and lead to API-level errors from a provider. But sometimes it is unavoidable (for example, via an interrupt or power loss) so the runtime automatically closes tool calls that fall into that category by telling the model the tool call(s) were cancelled.

Without further ado, a classic "hello world" example:

require "llm"

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: $stdout)
agent.talk "hello world"

Essentials

Stream

A stream can be a simple IO object (eg $stdout) or it can be a subclass of LLM::Stream. An IO object can receive content but it cannot receive the other callbacks that are available to subclasses of LLM::Stream. It is generally a good idea to return as quickly as possible from all of these methods, but especially for the on_content and on_reasoning_content methods: they run inline with the request.

A subclass of LLM::Stream can implement callbacks that the runtime will call throughout an agent's lifetime. All callbacks are optional. The callbacks provide for content, reasoning, tool calls, tool returns, steps in a turn, retries, compaction and more:

class Stream < LLM::Stream
  ##
  # @param [String] content
  #  A chunk of text
  def on_content(content)
    print content
  end

  ##
  # @param [String] content
  #  A chunk of text
  def on_reasoning_content(content)
    warn content
  end

  ##
  # @param [LLM::Function] tool
  #  The tool being called
  def on_tool_call(tool)
    nil
  end

  ##
  # @param [LLM::Function] tool
  #  The tool that returned
  # @param [LLM::Function::Return] result
  #  The return from the tool call
  def on_tool_return(tool, result)
    nil
  end

  ##
  # @note
  #  This method is called _before_ a transformer runs
  # @param [LLM::Transformer] transformer
  #  A transformer
  def on_transform(transformer)
    nil
  end

  ##
  # @note
  #  This method is called _after_ a transformer runs
  # @param [LLM::Transformer] transformer
  #  A transformer
  def on_transform_finish(transformer)
    nil
  end

  ##
  # @note
  #  This method is called _before_ a compactor runs
  # @param [LLM::Compactor] compactor
  #  A compactor
  def on_compaction(compactor)
    nil
  end

  ##
  # @note
  #  This method is called _after_ a compactor runs
  # @param [LLM::Compactor] compactor
  #  A compactor
  def on_compaction_finish(compactor)
    nil
  end

  ##
  # @note
  #  This method is called once per request
  #  in a turn (which can contain multiple
  #  requests)
  # @param [LLM::Context] ctx
  #  The context
  # @param [LLM::Response] res
  #  The response
  def on_step(ctx, res)
    nil
  end

  ##
  # @note
  #  This method is called when a request is
  #  rate limited and retried.
  # @param [LLM::RateLimitError] error
  # @param [Integer] attempt
  def on_retry(error, attempt)
    nil
  end

  ##
  # @note
  #  This method is called _before_ a skill runs
  # @param [LLM::Skill] skill
  #  A skill
  def on_skill_call(skill)
    nil
  end

  ##
  # @note
  #  This method is called _after_ a skill runs
  # @param [LLM::Agent] agent
  #  The agent who ran the skill
  # @param [LLM::Skill] skill
  # @param [LLM::Response] res
  def on_skill_return(agent, skill, res)
    nil
  end
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: Stream.new)
agent.talk "Explain Ruby fibers."
Tools

An agent requires one or more tools to be able to interact with the "outside" world. At a high-level a tool is how a model can access your filesystem, search the internet, post a comment on your behalf and anything else that your own code could do.

The runtime represents a tool as a subclass of LLM::Tool that provides a name, a description, an optional set of parameters and a method that the runtime will call on the model's behalf. A tool is implemented on top of an LLM::Function object, and you might come across it in stream and tracer callbacks. The model decides when and how a tool is called:

class ReadFile < LLM::Tool
  name "read-file"
  description "Read a file"
  parameter :path, String, "The filename or path"
  required %i[path]

  def call(path:)
    {contents: File.read(path)}
  end
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [ReadFile], stream: $stdout)
agent.talk "summarize README.md"
MCP

The Model Context Protocol (MCP) has first-class support in llm.rb. The stdio and http transports work out of the box. MCP tools are translated into subclasses of LLM::Tool that can be used with LLM::Context or LLM::Agent.

require "llm"

llm   = LLM.deepseek(key: ENV["KEY"])
mcp   = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
agent.talk "Run the tool"
Concurrency

The runtime supports six different concurrency strategies that have different attributes. The choice between all of them often depends on the requirements of your application.

IO-bound tools are a good fit for the :async, :thread, and :fiber strategies while true parallelism can be achieved with the :fork and :ractor strategies. The :sequential strategy runs tools one at a time and is the default. The :fork strategy also provides a separate process that offers isolation from its parent.

A couple of concurrency strategies require optional, opt-in dependencies. The async strategy requires the async gem and the fork strategy requires the xchan.rb gem (~> 0.24). The fiber strategy requires a scheduler (Fiber.scheduler) but by default Ruby does not provide one.

The :ractor strategy is the least interchangeable of the six. It runs class-based tools only, and a tool's arguments have to be ractor-shareable.

require "llm"
require "llm/tools"

llm   = LLM.deepseek(key: ENV["KEY"])
tools = LLM::Tool.subclasses
agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
agent.talk "Run the tools in parallel"
Cancellation

It is possible to interrupt a running agent who is between requests or tool calls as long as the cancel request is sent from another thread or fiber running in the same process as the agent. A cancel request can be sent with the LLM::Agent#interrupt! method.

An interrupted tool call is made aware of the interrupt and it can both rescue LLM::Interrupt and/or implement the on_interrupt callback on the tool class. The option to cancel on r.uby.dev is built on top of this feature, and it has a single background process with 24 threads. Each thread can run an agent request that can be interrupted via another thread in the same process.

It is solid and reliable but r.uby.dev had to meet this feature halfway, so expect to build your own infrastructure around it. See cancellation chapter to learn more.

class Search < LLM::Tool
  name "search"
  description "Search many files"
  parameter :pattern, String, "The pattern to search for"
  required %i[pattern]

  ##
  # A raise is delivered here; `on_interrupt` is a notification, and it
  # runs on every strategy - `:sequential` included.
  def call(pattern:)
    search(pattern)
  rescue LLM::Interrupt
    ##
    # A tool can return a value from here, and the turn carries on with
    # it, or re-raise and the fiber that made the request is raised
    # into as well.
    cleanup
    raise
  end

  ##
  # Told on the thread or fiber the call runs on, before the raise on
  # `:fork` and `:ractor` and after the rescue above on the other three.
  def on_interrupt
    cleanup
  end

  private

  def cleanup
    # Release a file, a socket, or a lock here.
  end
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [Search], concurrency: :async)
Thread.new { sleep(1); agent.interrupt! }

begin
  agent.talk "find every TODO in the repository", stream: $stdout
rescue LLM::Interrupt
  puts "cancelled"
end
Cancel by record ID

A common deployment setup is to run your agents in a background process that a web frontend can communicate with (usually via a database). The background process would have one thread per agent, and it could run as many agents as it has threads. This is how the r.uby.dev website is configured, and it is the configuration that the LLM.interrupt method is optimized for: a single process with each agent running in its own thread.

The LLM.interrupt method has access to a process-wide registry that contains every active instance of LLM::Agent, and that includes Sequel and ActiveRecord agents, too. An agent enters the registry when it starts a turn, and it exits the registry afterwards. The method returns true when it sent an interrupt, and otherwise it returns false.

There is often a window between when an agent is queued and when it runs, so a poll approach lets you eventually interrupt the agent, or give up trying:

class InterruptJob
  def call(agent_id:)
    LLM.interrupt(
      id: agent_id,
      attempts: 10,
      interval: 0.1
    )
  end
end
Serialization

Both LLM::Context and LLM::Agent can be serialized to JSON and written to disk. This feature is what supports the ActiveRecord and Sequel integrations too but rather than store the agent directly on disk it is stored in a database column instead.

An agent can be configured to read from and write to a file automatically with the path option. When the file already exists, the agent is restored from the file and continues where he left off. After each turn the agent flushes its state to the file. The text file can be shared like any other text file and it can be used to restore the agent in another process or machine:

require "llm"

path  = File.join(Dir.home, ".agents", "myagent.json")
llm   = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, path:)
agent.talk "remember my name is robert"

##
# Resume the conversation where the agent left off
agent = LLM::Agent.new(llm, path:)
agent.talk "what's my name?"
ActiveRecord

Both LLM::Context and LLM::Agent can be serialized to JSON and stored in a database column. The jsonb column type from PostgreSQL is recommended but it can also be stored as a string on other databases. ActiveRecord support is optimized for the jsonb column type and PostgreSQL.

The column captures everything an agent has done up to that point, it includes tool calls and returns, exchanged messages, and other metadata that carries runtime state. Each agent is an ActiveRecord model that calls acts_as_agent and each row represents an instance of that agent. It can be used with new and existing models alike.

The column should have the name data but this can be changed when the acts_as_agent method is called (eg acts_as_agent(data_column: :my_column)). The column is updated after every request that an agent makes rather than every turn, so an unexpected interrupt can be resumed from from the last request and no progress (or spent tokens) are lost:

require "active_record"
require "llm"
require "llm/active_record"

class Robert < ActiveRecord::Base
  acts_as_agent(format: :jsonb) do |agent|
    agent.set name: "robert",
              description: "an activerecord agent",
              instructions: proc { File.read(File.join(__dir__, "robert", "prompt.md")) },
              tools: :tools,
              concurrency: :async,
              tool_budget: 25,
              tracer: proc { Robert::Tracer::SQL.new(llm, agent: self) }
  end

  ##
  # @return [LLM::MCP]
  def github
    @github ||= LLM::MCP.http(
      url: "https://api.githubcopilot.com/mcp/",
      headers: {"Authorization" => "Bearer #{ENV['GITHUB_RUBYDEV_PAT']}"},
      transport: :net_http_persistent
    )
  end

  ##
  # @return [Array<LLM::Tool>]
  def tools
    github.tools
  end
end

agent = Robert.create!

##
# Every call to `talk` automatically persists
# to the database.
agent.talk "what's new on the llm.rb repository?"

##
# The conversation was persisted to database. A
# fresh instance restores it and continues where
# we left off
agent = Robert.find(agent.id).talk "and what about roda-llm?"

##
# Start an agent console.
# Query agent's state, debug, etc.
# The console does not persist back to the database.
agent.console
SQL optimizations

In a database environment the runtime optimizes for the PostgreSQL database and its builtin support for the jsonb column type. An agent fits in a single column, on a single row, and that column carries everything it has done: messages, tool calls, context usage, and so on. It works well in practice and means you can store an agent almost anywhere.

For scenarios where performance matters most the runtime ships with virtual ActiveRecord classes that never materialize in your database but provide a SQL view into the column where an agent stores its runtime state. They return ActiveRecord::Relation objects, so the filtering happens in the database.

class Agent < ActiveRecord::Base
  acts_as_agent(format: :jsonb) do |agent|
    agent.set name: "activerecord agent"
  end
end

##
# Find an instance of your agent
agent = Agent.find_by(id: 1)

##
# Returns a relation over the agent's messages.
# It is scoped to the agent, and it yields one
# instance of LLM::ActiveRecord::Message per
# message the agent has produced.
messages = LLM::ActiveRecord::Message.for(agent:)

##
# The relation chains like any other
messages.where(role: "assistant")
        .order(position: :desc)
        .limit(10)

##
# Count, too
messages.count

Schema

Each row carries a message, flattened into columns:

column contents
agent_id the agent a message belongs to
id the message id
role the message role
content the message content
tools the tool calls a message carries
position the position of a message in the conversation
data the whole message, as the runtime stores it

Indexes

The queries the view runs are already covered. They expand one agent, found by primary key, so they are index scans. There is nothing to add for LLM::ActiveRecord::Message.for(agent:).

The queries you write on top of it are not. Once a question is asked of every agent, the column is expanded row by row and no index helps the view itself. Index the column for those questions instead:

CREATE INDEX index_agents_on_data
  ON agents USING gin (data jsonb_path_ops);

CREATE INDEX index_agents_on_context_used
  ON agents (((data ->> 'context_used')::int));

The first serves containment (@>) and path queries over the state as a whole. The second serves a scalar key, and the runtime already writes context_used and context_window at the top level, so "sessions over 80% full" becomes cheap. Both assume format: :jsonb.

However: an agent's whole conversation lives in one value, so every save rewrites it, and a GIN index is maintained with it. Prefer an index on a key or two over the whole column.

Structured outputs

LLM::Schema subclasses produce typed, structured output from any model call. Pass a schema to LLM::Context#talk, LLM::Agent#talk, or LLM::Provider#complete to receive validated JSON instead of free text. Schemas work alongside tools and streams.

LLM::Schema can define objects, arrays, enums, nested schemas, and more. It is also used internally by LLM::Tool for parameter definitions, so you already benefit from it when you declare tool parameters.

The LLM::DeepSeek provider includes runtime-level optimisations such as structured output support (despite no official structured outputs API) and SVG image generation. This example uses LLM::Schema with DeepSeek:

class Weather < LLM::Schema
  property :city, String, "The city name"
  property :temperature, Number, "Current temperature"
  property :conditions, String, "Weather conditions"
  required %i[city temperature conditions]
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, schema: Weather)
res = agent.talk "Weather in Paris?"
res.content!  # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}

Debuggers

Console

The LLM::Agent#console method drops you into an interactive console that is built on top of (n)curses. The llm.rb executable packaged with the gem is another way to access the console and ActiveRecord models who have called acts_as_agent can access the console as well (via agent.console).

A console for an ActiveRecord model does not write back to the database. The llm.rb executable automatically associates a session with the current working directory and it can be resumed by calling llm.rb in the same directory at a later point.

The console is not intended to compete with Claude, Codex and friends. It is much more limited, serves an entirely different purpose and is more like a debugger for your agents. The dependencies required by the console are not installed by default, and the easiest way to grab them is via gem install llm-shell.

Demo

llm.rb console demo

Tracer

It is possible to trace what an agent is doing by attaching a tracer. A tracer can hook into requests, tool calls, and other runtime events to debug an agent, provide insights, monitor latency, or export spans to an observability backend. All built-in tracers share one interface, so switching between them means changing a factory method:

It is also possible to create your own tracer by creating a subclass of LLM::Tracer that implements a number of callbacks that cover an agent's lifecycle. The tracer feature provides visibility into what the runtime is doing, and the tracer API lets other code hook into that feature.

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tracer: LLM::Tracer.pretty_logger(llm))
agent.talk "Hello"

Hooks

Guards

LLM::Guard is the hook that sees every tool call before it runs. A guard can let a call through, cancel it, block it with an error, or even answer for it. Because it runs before the tool, anything it intercepts never executes. Policy, validation, quotas, and cost ceilings all live here.

Agents and contexts use LLM::Guard::Null by default, so a guard only runs when you configure one. To write your own guard, subclass LLM::Guard and implement LLM::Guard#call. The pending call arrives as function:. Return a value to close the call, or nil to let it run:

class PolicyGuard < LLM::Guard
  def call(function:)
    if function.name == "exec"
      function.return(error: true, type: "policy_error",
                      message: "exec is disabled")
    end
  end
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [LLM::Tool::Exec, ReadFile], guard: PolicyGuard)
Transformers

It is possible to rewrite outgoing messages before they reach the provider with LLM::Transformer. Create a subclass and implement call(message:) to scrub sensitive data, inject context, or normalize content. The transform runs automatically on every turn, so you never have to change your prompt code.

class RedactEmails < LLM::Transformer
  def call(message:)
    content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
    LLM::Message.new(message.role, content, message.extra)
  end
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, transformer: RedactEmails)
agent.talk "Contact support@example.com for help"
Compactors

Every model has a context window: the finite number of tokens it can consider in a single request. Generally a compactor will drop or summarize older messages to keep the conversation within that window, and it runs automatically before every turn. By default it is disabled so it is a feature you must opt into.

LLM::Compactor::Truncate keeps the most recent messages via an integer count or a percentage like "80%". It preserves tool call and return pairs so the conversation never contains an orphaned result. It is also possible to subclass LLM::Compactor to implement your own compactor with its own logic. Streams can observe the process through the LLM::Stream#on_compaction and LLM::Stream#on_compaction_finish callbacks.

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(
  llm,
  compactor: LLM::Compactor::Truncate,
  compactor_options: {keep: 64}
)
agent.talk "Hello"

Everything else

Automatic retries

Rate-limited requests are retried automatically by default. Agents retry a 429 up to five times with a growing backoff before giving up, so most request failures resolve on their own. Connection and read timeouts are retried the same way. Set retry_budget to change the number of retries, or retry_budget: 0 to disable them.

require "llm"

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, retry_budget: 0)
agent.talk "Hello"
Usage and cost

Every context and agent reports what a conversation has spent and how much room is left, and the numbers answer different questions. A token_usage is the whole conversation, summed as an LLM::Usage, and it is what LLM::Cost prices against the model registry:

require "llm"

llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.talk "Hello"

agent.token_usage  # => LLM::Usage for the whole conversation
agent.cost         # => LLM::Cost, priced from the registry

A context_used is one turn's worth - the live size of the most recent assistant message - so it is what a context window is really being spent on, and context_usage is that as a fraction of the window:

agent.context_used    # => tokens in the latest turn
agent.context_window  # => the model's limit, or nil when unknown
agent.context_usage   # => Rational, eg Rational(100, 10_000)
A2A

The Agent 2 Agent (A2A) protocol has first-class support in llm.rb. The http and jsonrpc transports work out of the box. A2A skills are translated into subclasses of LLM::Tool that can be used with LLM::Context or LLM::Agent.

require "llm"

llm   = LLM.deepseek(key: ENV["KEY"])
a2a   = LLM::A2A.rest(url: "https://remote-agent.example.com")
agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
agent.talk "Run the skill"
Skills

A skill turns a markdown file into a callable tool. When the model calls it, the runtime spawns a subagent with the skill's instructions as its system prompt and the skill's own tool set. The subagent runs one turn and returns the result, then is discarded. Each call is fresh and stateless.

A LLM::Stream can be notified as a skill starts and when it returns. The on_skill_return callback hands back the subagent that ran the skill, so you can inspect its conversation, measure its usage, track costs or add a verification step (eg subagent.talk("verify your work")).

summary.md
---
name: summary
description: Reads recent git history and writes a summary
tools: all
---

Collect the recent git log, analyze each commit,
and write a summary to summary.txt.
agent.rb
require "llm"

llm   = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, skills: ["summary.md"])
agent.talk "Summarize the last week of work"
As a subclass

LLM::Agent.set is a class-level DSL that accepts a Hash of properties. Each key resolves to a corresponding class accessor: name, description, model, tools, instructions, schema, stream, tracer, concurrency, confirm, path, skills, tool_budget, and retry_budget. All options are optional; zero or more can be set. An error is raised for unknown keys so that typos are caught early.

require "llm"
require "llm/tools"

class Agent < LLM::Agent
  set name: "sysadmin",
      description: "system administration agent",
      model: "deepseek-v4-pro",
      tools: [LLM::Tool::Exec]
end

llm = LLM.deepseek(key: ENV["KEY"])
agent = Agent.new(llm)
agent.talk "Run 'date'"

Providers

Each provider is constructed with a class-level factory method on LLM, and the resulting instance is passed to LLM::Context or LLM::Agent. The same API drives every one of them, so switching providers is a one-line change.

What providers does llm.rb support?

  • Anthropic (LLM.anthropic)
  • Google (LLM.google)
  • OpenAI (LLM.openai)
  • DeepSeek (LLM.deepseek)
  • DeepInfra (LLM.deepinfra)
  • xAI (LLM.xai)
  • Z.ai (LLM.zai)
  • Moonshot (Kimi) (LLM.moonshot)
  • OpenRouter (LLM.openrouter)
  • Alibaba (Qwen3) (LLM.alibaba, also LLM.aliyun)
  • Mistral (LLM.mistral)
  • AWS Bedrock (LLM.bedrock)
  • Ollama (LLM.ollama)
  • llama.cpp (LLM.llamacpp)
Implicit

Cloud providers can infer their API key automatically from a set of common defaults that are defined by the models.dev registry that is also distributed with llm.rb.

llm = LLM.openai
llm = LLM.anthropic
llm = LLM.google
llm = LLM.deepseek
llm = LLM.deepinfra
llm = LLM.xai
llm = LLM.zai
llm = LLM.moonshot
llm = LLM.openrouter
llm = LLM.alibaba  # also: LLM.aliyun
llm = LLM.mistral
llm = LLM.bedrock
Explicit

The key option can also be providied explicitly, and certain providers (eg ollama, llamacpp) usually do not require an API key at all.

llm = LLM.openai(key: ENV["OPENAI_API_KEY"])
llm = LLM.anthropic(key: ENV["ANTHROPIC_API_KEY"])
llm = LLM.google(key: ENV["GOOGLE_API_KEY"])
llm = LLM.deepseek(key: ENV["DEEPSEEK_API_KEY"])
llm = LLM.deepinfra(key: ENV["DEEPINFRA_API_KEY"])
llm = LLM.xai(key: ENV["XAI_API_KEY"])
llm = LLM.zai(key: ENV["ZHIPU_API_KEY"])
llm = LLM.moonshot(key: ENV["MOONSHOT_API_KEY"])
llm = LLM.openrouter(key: ENV["OPENROUTER_API_KEY"])
llm = LLM.alibaba(key: ENV["DASHSCOPE_API_KEY"]) # also: LLM.aliyun
llm = LLM.mistral(key: ENV["MISTRAL_API_KEY"])
llm = LLM.bedrock(
  access_key_id: ENV["AWS_ACCESS_KEY_ID"],
  secret_access_key: ENV["AWS_SECRET_ACCESS_KEY"],
  region: ENV["AWS_REGION"]
)
Model Registry

Each provider ships its model catalog, pricing, limits, and modalities with the gem, sourced from models.dev. Reach it from any provider, context, or agent, enumerate models, or sort them by price.

require "llm"

llm      = LLM.openai
registry = llm.registry                # => LLM::Provider#registry
cheapest = registry.models.sort.first  # => LLM::Registry::Model
cheapest.id                            # => "text-embedding-3-small"
cheapest.context_window                # => 8191
cheapest.structured_output?            # => false
Transports

The transport: option selects which HTTP library a provider uses for network communication. Three backends ship out of the box: net/http is always available and the default, net/http/persistent pools connections for many requests to the same host, and curb wraps libcurl. They share one interface, so switching is a one-word change.

llm = LLM.deepseek(
  key: ENV["KEY"],
  transport: :net_http_persistent
)
Timeouts

Providers accept two timeouts:

  • connect_timeout - opening the connection. Defaults to 5 seconds.
  • read_timeout - waiting for a response on an idle connection. Defaults to 600 seconds (10 minutes).

The longer read timeout leaves room for slow reasoning models and local models. The legacy timeout: option remains as a shorthand for read_timeout. Timeouts are retriable: LLM::Agent retries a timed out request up to its retry_budget (five by default), so a dropped connection or a slow first token is often something we can recover from.

llm = LLM.deepseek(
  connect_timeout: 5,   # opening the connection
  read_timeout: 600     # waiting for the next bytes
)
Headers

Providers can accept a custom set of headers with the LLM::Provider#with method. For example, you could set a custom User-Agent header, or provide headers that carry special meaning to certain providers (eg OpenAI, OpenRouter).

llm = LLM.openrouter
llm = llm.with("HTTP-Referer" => "https://example.com")
llm = llm.with("X-OpenRouter-Title" => "Example App")

RAG

Most providers offer an embedding model that can be used for semantic search, or similarity search. An embedding model can generate embeddings that can then be stored in a database that is optimized for storing and querying vectors, such as SQLite's sqlite-vec or PostgreSQL's pg-vector.

llm.rb also includes support for OpenAI's vector store API. It provides a vector database as a HTTP service but we won't cover that here.

require "llm"

llm  = LLM.openai(key: ENV["KEY"])
body = "llm.rb is Ruby's capable AI runtime."
embedding = llm.embed([body]).embeddings.first

# Document is your ActiveRecord or Sequel model
# with a vector column (e.g. sqlite-vec or pgvector)
Document.create!(
  title: "llm.rb",
  body:,
  embedding:,
)

Images

A handful of providers can generate images from a text prompt. OpenAI, Google, xAI, and DeepInfra all support it. The API is the same across providers:

require "llm"

llm = LLM.openai(key: ENV["KEY"])
res = llm.images.create(prompt: "a dog on a rocket to the moon")
IO.copy_stream res.images[0], "rocket.png"
DeepSeek

DeepSeek does not have a dedicated image model, but the runtime generates SVG vector graphics through its text model. Each generation produces a valid SVG document that can be converted to PNG with tools like rsvg-convert. Pass an existing agent to maintain a session across generations:

require "llm"
llm = LLM.deepseek(key: ENV["KEY"])

##
# First generation
res = llm.images.create(prompt: "a rocket on the moon")
IO.copy_stream res.images[0], "rocket.svg"

##
# Refine with follow-up prompts (shares context)
res = llm.images.create(prompt: "add a dog next to the rocket",
                        agent: res.agent)
IO.copy_stream res.images[0], "rocket-with-dog.svg"

FAQ

Where can I see llm.rb in action?

The r.uby.dev website.

What about local LLM support?

The following providers can be run used with models that are running on your own hardware.

  • Ollama
  • Llamacpp
I have a limited budget. What should I do?

There are a few options. The first option is to host your own model, and use the ollama or llamacpp providers. This can be difficult though because a capable model requires hardware that can match it. If you have the ability to self-host, this would be my first option.

The second option is DeepSeek.
The deepseek-v4-flash model costs pennies to use.
And llm.rb has been optimized for deepseek. For example, DeepSeek does not have image generation capabilities but on the llm.rb runtime it does (vector graphics only, though).

The same is true for structured outputs. DeepSeek does not support structured outputs in the same way as OpenAI or Google, but the llm.rb runtime makes it appear as though it does, through the json_object response type.

If you're on a budget, DeepSeek is hard to beat.

Sources other than GitHub?

We are on the radicle.network as well.
Every commit that lands on GitHub also lands on Radicle.
Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
Browse on the web.

Who maintains llm.rb?

The llm.rb project was started more than three years ago by @altruby and @antaz. The primary maintainer is @altruby. Over those three years multiple other contributors have contributed to llm.rb as well, and new contributors are always welcome.

How well tested is llm.rb?

It is battle tested daily.

The console that is distributed with llm.rb is used to build llm.rb so there is a healthy, active feedback loop. It also powers the r.uby.dev website where multiple llm.rb agents are deployed. I'm aware of at least one production Rails deployment at a large-ish company.

And this git repository includes llm.rb agents that help me maintain the documentation and perform other repository maintainence. The feedback loop is constant. Outside of that there is a large test suite that covers live requests (recorded by VCR) and database interactions.

License

This software is released under the terms of the MIT license.
See LICENSE for details.