r.uby.dev project.
Welcome to the canonical llm.rb repository.
llm.rb is a runtime for building agentic AI applications on CRuby. It has zero runtime dependencies by default, supports concurrent and parallel tool execution and has a single coherent API that spans 14+ providers. The README covers a lot of ground and the changelog tracks what has changed between releases.
If you want to see the runtime in action the r.uby.dev website provides a platform where you can - for free. It hosts multiple llm.rb agents, and one of them (bezela) helps manage this repository.
llm.rb requires Ruby 3.4 or later.
gem install llm.rbThe
LLM::Agent
class is the default high-level interface,
and it is recommended for most use-cases. It manages the tool loop
and provides configurable features on top of it. For example you can
manage the tool loop with a retry budget and a tool call budget -
alongside other features.
The runtime is designed to keep the tool loop alive and it will rescue exceptions. When an exception is encountered in a tool it is reported back to the model as an in-band error that allows the model to change course or retry with different parameters.
A tool call requires a tool return (or response), and the lack of one can corrupt the conversation and lead to API-level errors from a provider. But sometimes it is unavoidable (for example, via an interrupt or power loss) so the runtime automatically closes tool calls that fall into that category by telling the model the tool call(s) were cancelled.
Without further ado, a classic "hello world" example:
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: $stdout)
agent.talk "hello world"Stream
A stream can be a simple IO object (eg $stdout)
or it can be a subclass of LLM::Stream.
An IO object can receive content but it cannot
receive the other callbacks that are available
to subclasses of LLM::Stream.
It is generally a good idea to return as
quickly as possible from all of these methods,
but especially for the on_content and
on_reasoning_content methods: they run
inline with the request.
A subclass of LLM::Stream
can implement callbacks that the runtime will call
throughout an agent's lifetime. All callbacks are
optional. The callbacks provide for content, reasoning,
tool calls, tool returns, steps in a turn, retries,
compaction and more:
class Stream < LLM::Stream
##
# @param [String] content
# A chunk of text
def on_content(content)
print content
end
##
# @param [String] content
# A chunk of text
def on_reasoning_content(content)
warn content
end
##
# @param [LLM::Function] tool
# The tool being called
def on_tool_call(tool)
nil
end
##
# @param [LLM::Function] tool
# The tool that returned
# @param [LLM::Function::Return] result
# The return from the tool call
def on_tool_return(tool, result)
nil
end
##
# @note
# This method is called _before_ a transformer runs
# @param [LLM::Transformer] transformer
# A transformer
def on_transform(transformer)
nil
end
##
# @note
# This method is called _after_ a transformer runs
# @param [LLM::Transformer] transformer
# A transformer
def on_transform_finish(transformer)
nil
end
##
# @note
# This method is called _before_ a compactor runs
# @param [LLM::Compactor] compactor
# A compactor
def on_compaction(compactor)
nil
end
##
# @note
# This method is called _after_ a compactor runs
# @param [LLM::Compactor] compactor
# A compactor
def on_compaction_finish(compactor)
nil
end
##
# @note
# This method is called once per request
# in a turn (which can contain multiple
# requests)
# @param [LLM::Context] ctx
# The context
# @param [LLM::Response] res
# The response
def on_step(ctx, res)
nil
end
##
# @note
# This method is called when a request is
# rate limited and retried.
# @param [LLM::RateLimitError] error
# @param [Integer] attempt
def on_retry(error, attempt)
nil
end
##
# @note
# This method is called _before_ a skill runs
# @param [LLM::Skill] skill
# A skill
def on_skill_call(skill)
nil
end
##
# @note
# This method is called _after_ a skill runs
# @param [LLM::Agent] agent
# The agent who ran the skill
# @param [LLM::Skill] skill
# @param [LLM::Response] res
def on_skill_return(agent, skill, res)
nil
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, stream: Stream.new)
agent.talk "Explain Ruby fibers."Tools
An agent requires one or more tools to be able to interact with the "outside" world. At a high-level a tool is how a model can access your filesystem, search the internet, post a comment on your behalf and anything else that your own code could do.
The runtime represents a tool as a subclass
of LLM::Tool
that provides a name, a description, an optional
set of parameters and a method that the runtime
will call on the model's behalf. A tool is implemented
on top of
an LLM::Function
object, and you might come across it in stream and
tracer callbacks. The model decides when and how a
tool is called:
class ReadFile < LLM::Tool
name "read-file"
description "Read a file"
parameter :path, String, "The filename or path"
required %i[path]
def call(path:)
{contents: File.read(path)}
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [ReadFile], stream: $stdout)
agent.talk "summarize README.md"MCP
The Model Context Protocol (MCP) has first-class support
in llm.rb. The stdio and http transports work out of the
box. MCP tools are translated into subclasses of
LLM::Tool that can be
used with
LLM::Context or
LLM::Agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
mcp = LLM::MCP.stdio(argv: ["ruby", "server.rb"])
agent = LLM::Agent.new(llm, stream: $stdout, tools: mcp.tools)
agent.talk "Run the tool"Concurrency
The runtime supports six different concurrency strategies that have different attributes. The choice between all of them often depends on the requirements of your application.
IO-bound tools are a good fit for the :async, :thread,
and :fiber strategies while true parallelism can be achieved
with the :fork and :ractor strategies. The
:sequential strategy runs tools one at a time and is the default.
The :fork strategy also provides a separate process that offers
isolation from its parent.
A couple of concurrency strategies require optional, opt-in dependencies.
The async strategy requires the async
gem and the fork strategy requires the xchan.rb
gem (~> 0.24). The fiber strategy requires a scheduler (Fiber.scheduler) but by
default Ruby does not provide one.
The :ractor strategy is the least interchangeable of the six. It runs
class-based tools only, and a tool's arguments have to be
ractor-shareable.
require "llm"
require "llm/tools"
llm = LLM.deepseek(key: ENV["KEY"])
tools = LLM::Tool.subclasses
agent = LLM::Agent.new(llm, tools:, concurrency: :fork)
agent.talk "Run the tools in parallel"Cancellation
It is possible to interrupt a running agent who
is between requests or tool calls as long as the
cancel request is sent from another thread or fiber
running in the same process as the agent. A cancel
request can be sent with the
LLM::Agent#interrupt!
method.
An interrupted tool call is made aware of the interrupt
and it can both rescue LLM::Interrupt and/or implement
the on_interrupt callback on the tool class. The option to
cancel on r.uby.dev is built on top of
this feature, and it has a single background process
with 24 threads. Each thread can run an agent request that
can be interrupted via another thread in the same process.
It is solid and reliable but r.uby.dev had to meet this feature halfway, so expect to build your own infrastructure around it. See cancellation chapter to learn more.
class Search < LLM::Tool
name "search"
description "Search many files"
parameter :pattern, String, "The pattern to search for"
required %i[pattern]
##
# A raise is delivered here; `on_interrupt` is a notification, and it
# runs on every strategy - `:sequential` included.
def call(pattern:)
search(pattern)
rescue LLM::Interrupt
##
# A tool can return a value from here, and the turn carries on with
# it, or re-raise and the fiber that made the request is raised
# into as well.
cleanup
raise
end
##
# Told on the thread or fiber the call runs on, before the raise on
# `:fork` and `:ractor` and after the rescue above on the other three.
def on_interrupt
cleanup
end
private
def cleanup
# Release a file, a socket, or a lock here.
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [Search], concurrency: :async)
Thread.new { sleep(1); agent.interrupt! }
begin
agent.talk "find every TODO in the repository", stream: $stdout
rescue LLM::Interrupt
puts "cancelled"
endCancel by record ID
A common deployment setup is to run your agents in a
background process that a web frontend can communicate
with (usually via a database). The background process
would have one thread per agent, and it could run as
many agents as it has threads. This is how the
r.uby.dev website is configured,
and it is the configuration that the
LLM.interrupt
method is optimized for: a single process with each agent
running in its own thread.
The
LLM.interrupt
method has access to a process-wide
registry that contains every active instance of
LLM::Agent,
and that includes Sequel and ActiveRecord agents, too. An
agent enters the registry when it starts a turn, and it
exits the registry afterwards. The method returns true
when it sent an interrupt, and otherwise it returns false.
There is often a window between when an agent is queued and when it runs, so a poll approach lets you eventually interrupt the agent, or give up trying:
class InterruptJob
def call(agent_id:)
LLM.interrupt(
id: agent_id,
attempts: 10,
interval: 0.1
)
end
endSerialization
Both LLM::Context
and
LLM::Agent
can be serialized to JSON and written to disk.
This feature is what supports the ActiveRecord
and Sequel integrations too but rather than
store the agent directly on disk it is stored
in a database column instead.
An agent can be configured to read from and
write to a file automatically with the path
option. When the file already exists, the agent
is restored from the file and continues where
he left off. After each turn the agent flushes
its state to the file. The text file can be shared
like any other text file and it can be used to
restore the agent in another process or machine:
require "llm"
path = File.join(Dir.home, ".agents", "myagent.json")
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, path:)
agent.talk "remember my name is robert"
##
# Resume the conversation where the agent left off
agent = LLM::Agent.new(llm, path:)
agent.talk "what's my name?"ActiveRecord
Both
LLM::Context and
LLM::Agent
can be serialized to JSON and stored in a database column. The jsonb
column type from PostgreSQL is recommended but it can also be stored as
a string on other databases. ActiveRecord support is optimized
for the jsonb column type and PostgreSQL.
The column captures everything an agent has done up to that point,
it includes tool calls and returns, exchanged messages, and other
metadata that carries runtime state. Each agent is an ActiveRecord
model that calls acts_as_agent and each row represents an instance
of that agent. It can be used with new and existing models alike.
The column should have the name data but this can be
changed when the acts_as_agent method is called (eg
acts_as_agent(data_column: :my_column)). The column
is updated after every request that an agent makes rather
than every turn, so an unexpected interrupt can be resumed
from from the last request and no progress (or spent tokens)
are lost:
require "active_record"
require "llm"
require "llm/active_record"
class Robert < ActiveRecord::Base
acts_as_agent(format: :jsonb) do |agent|
agent.set name: "robert",
description: "an activerecord agent",
instructions: proc { File.read(File.join(__dir__, "robert", "prompt.md")) },
tools: :tools,
concurrency: :async,
tool_budget: 25,
tracer: proc { Robert::Tracer::SQL.new(llm, agent: self) }
end
##
# @return [LLM::MCP]
def github
@github ||= LLM::MCP.http(
url: "https://api.githubcopilot.com/mcp/",
headers: {"Authorization" => "Bearer #{ENV['GITHUB_RUBYDEV_PAT']}"},
transport: :net_http_persistent
)
end
##
# @return [Array<LLM::Tool>]
def tools
github.tools
end
end
agent = Robert.create!
##
# Every call to `talk` automatically persists
# to the database.
agent.talk "what's new on the llm.rb repository?"
##
# The conversation was persisted to database. A
# fresh instance restores it and continues where
# we left off
agent = Robert.find(agent.id).talk "and what about roda-llm?"
##
# Start an agent console.
# Query agent's state, debug, etc.
# The console does not persist back to the database.
agent.consoleSQL optimizations
In a database environment the runtime optimizes for
the PostgreSQL database and its builtin support for
the jsonb column type. An agent fits in a single
column, on a single row, and that column carries
everything it has done: messages, tool calls,
context usage, and so on. It works well in practice
and means you can store an agent almost anywhere.
For scenarios where performance matters most the runtime
ships with virtual ActiveRecord classes that never materialize
in your database but provide a SQL view into the column where
an agent stores its runtime state. They return
ActiveRecord::Relation
objects, so the filtering happens in the database.
class Agent < ActiveRecord::Base
acts_as_agent(format: :jsonb) do |agent|
agent.set name: "activerecord agent"
end
end
##
# Find an instance of your agent
agent = Agent.find_by(id: 1)
##
# Returns a relation over the agent's messages.
# It is scoped to the agent, and it yields one
# instance of LLM::ActiveRecord::Message per
# message the agent has produced.
messages = LLM::ActiveRecord::Message.for(agent:)
##
# The relation chains like any other
messages.where(role: "assistant")
.order(position: :desc)
.limit(10)
##
# Count, too
messages.countSchema
Each row carries a message, flattened into columns:
| column | contents |
|---|---|
agent_id |
the agent a message belongs to |
id |
the message id |
role |
the message role |
content |
the message content |
tools |
the tool calls a message carries |
position |
the position of a message in the conversation |
data |
the whole message, as the runtime stores it |
Indexes
The queries the view runs are already covered. They expand
one agent, found by primary key, so they are index scans.
There is nothing to add for
LLM::ActiveRecord::Message.for(agent:).
The queries you write on top of it are not. Once a question is asked of every agent, the column is expanded row by row and no index helps the view itself. Index the column for those questions instead:
CREATE INDEX index_agents_on_data
ON agents USING gin (data jsonb_path_ops);
CREATE INDEX index_agents_on_context_used
ON agents (((data ->> 'context_used')::int));The first serves containment (@>) and path queries over
the state as a whole. The second serves a scalar key, and
the runtime already writes context_used and
context_window at the top level, so "sessions over 80%
full" becomes cheap. Both assume format: :jsonb.
However: an agent's whole conversation lives in one value, so every save rewrites it, and a GIN index is maintained with it. Prefer an index on a key or two over the whole column.
Structured outputs
LLM::Schema
subclasses produce typed, structured
output from any model call. Pass a schema to
LLM::Context#talk,
LLM::Agent#talk,
or
LLM::Provider#complete
to receive validated JSON instead of free text. Schemas work alongside tools and streams.
LLM::Schema
can define objects, arrays, enums, nested schemas,
and more. It is also used internally by
LLM::Tool for parameter
definitions, so you already benefit from it when you declare tool
parameters.
The
LLM::DeepSeek
provider includes runtime-level optimisations such as structured
output support (despite no official structured outputs API) and
SVG image generation. This example uses
LLM::Schema with
DeepSeek:
class Weather < LLM::Schema
property :city, String, "The city name"
property :temperature, Number, "Current temperature"
property :conditions, String, "Weather conditions"
required %i[city temperature conditions]
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, schema: Weather)
res = agent.talk "Weather in Paris?"
res.content! # => {city: "Paris", temperature: 15.0, conditions: "Cloudy"}Console
The LLM::Agent#console
method drops you into an interactive console that is built on
top of (n)curses. The llm.rb executable packaged with the gem
is another way to access the console and ActiveRecord models who
have called acts_as_agent can access the console as well
(via agent.console).
A console for an ActiveRecord model does not write back to the
database. The llm.rb executable automatically associates a
session with the current working directory and it can be resumed
by calling llm.rb in the same directory at a later point.
The console is not intended to compete with Claude, Codex and
friends. It is much more limited, serves an entirely different
purpose and is more like a debugger for your agents. The
dependencies required by the console are not installed
by default, and the easiest way to grab them is via
gem install llm-shell.
Tracer
It is possible to trace what an agent is doing by attaching a tracer. A tracer can hook into requests, tool calls, and other runtime events to debug an agent, provide insights, monitor latency, or export spans to an observability backend. All built-in tracers share one interface, so switching between them means changing a factory method:
LLM::Tracer.pretty_logger: human-readable single-line logs to stderr, ideal during development.LLM::Tracer.telemetry: exports spans via OTLP for OpenTelemetry in production.LLM::Tracer.logger: structured JSON to stdout or a file.
It is also possible to create your own tracer by creating a subclass
of LLM::Tracer
that implements a number of callbacks that cover an agent's lifecycle.
The tracer feature provides visibility into what the runtime is doing,
and the tracer API lets other code hook into that feature.
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tracer: LLM::Tracer.pretty_logger(llm))
agent.talk "Hello"Guards
LLM::Guard
is the hook that sees every tool call before it runs. A guard
can let a call through, cancel it, block it with an error, or
even answer for it. Because it runs before the tool, anything
it intercepts never executes. Policy, validation, quotas, and
cost ceilings all live here.
Agents and contexts use
LLM::Guard::Null
by default, so a guard only runs when you configure one. To
write your own guard, subclass
LLM::Guard
and implement
LLM::Guard#call.
The pending call arrives as function:. Return a value to close
the call, or nil to let it run:
class PolicyGuard < LLM::Guard
def call(function:)
if function.name == "exec"
function.return(error: true, type: "policy_error",
message: "exec is disabled")
end
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, tools: [LLM::Tool::Exec, ReadFile], guard: PolicyGuard)Transformers
It is possible to rewrite outgoing messages before they reach the provider with
LLM::Transformer. Create a subclass and implement call(message:) to scrub sensitive data,
inject context, or normalize content. The transform runs automatically
on every turn, so you never have to change your prompt code.
class RedactEmails < LLM::Transformer
def call(message:)
content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
LLM::Message.new(message.role, content, message.extra)
end
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, transformer: RedactEmails)
agent.talk "Contact support@example.com for help"Compactors
Every model has a context window: the finite number of tokens it can consider in a single request. Generally a compactor will drop or summarize older messages to keep the conversation within that window, and it runs automatically before every turn. By default it is disabled so it is a feature you must opt into.
LLM::Compactor::Truncate
keeps the most recent messages via an integer count or a percentage like
"80%". It preserves tool call and return pairs so the conversation
never contains an orphaned result. It is also possible to subclass
LLM::Compactor
to implement your own compactor with its own logic. Streams can observe the
process through the
LLM::Stream#on_compaction
and
LLM::Stream#on_compaction_finish
callbacks.
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(
llm,
compactor: LLM::Compactor::Truncate,
compactor_options: {keep: 64}
)
agent.talk "Hello"Automatic retries
Rate-limited requests are retried automatically by default. Agents
retry a 429 up to five times with a growing backoff before giving
up, so most request failures resolve on their own. Connection and
read timeouts are retried the same way. Set retry_budget
to change the number of retries, or retry_budget: 0 to disable
them.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, retry_budget: 0)
agent.talk "Hello"Usage and cost
Every context and agent reports what a conversation has spent and how
much room is left, and the numbers answer different questions. A
token_usage is the whole conversation, summed as an
LLM::Usage, and it
is what LLM::Cost
prices against the model registry:
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm)
agent.talk "Hello"
agent.token_usage # => LLM::Usage for the whole conversation
agent.cost # => LLM::Cost, priced from the registryA context_used is one turn's worth - the live size of the most recent
assistant message - so it is what a context window is really being
spent on, and context_usage is that as a fraction of the window:
agent.context_used # => tokens in the latest turn
agent.context_window # => the model's limit, or nil when unknown
agent.context_usage # => Rational, eg Rational(100, 10_000)A2A
The Agent 2 Agent (A2A) protocol has first-class support
in llm.rb. The http and jsonrpc transports work out of the
box. A2A skills are translated into subclasses of
LLM::Tool that can be
used with
LLM::Context or
LLM::Agent.
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
a2a = LLM::A2A.rest(url: "https://remote-agent.example.com")
agent = LLM::Agent.new(llm, stream: $stdout, tools: a2a.skills)
agent.talk "Run the skill"Skills
A skill turns a markdown file into a callable tool. When the model calls it, the runtime spawns a subagent with the skill's instructions as its system prompt and the skill's own tool set. The subagent runs one turn and returns the result, then is discarded. Each call is fresh and stateless.
A LLM::Stream
can be notified as a skill starts and when it returns. The on_skill_return
callback hands back the subagent that ran the skill, so you can inspect
its conversation, measure its usage, track costs or add a verification
step (eg subagent.talk("verify your work")).
---
name: summary
description: Reads recent git history and writes a summary
tools: all
---
Collect the recent git log, analyze each commit,
and write a summary to summary.txt.require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
agent = LLM::Agent.new(llm, skills: ["summary.md"])
agent.talk "Summarize the last week of work"As a subclass
LLM::Agent.set
is a class-level DSL that accepts a Hash of properties. Each key resolves to a
corresponding class accessor: name, description, model, tools,
instructions, schema, stream, tracer, concurrency, confirm,
path, skills, tool_budget, and retry_budget. All options are
optional; zero or more can be set.
An error is raised for unknown keys so that typos are caught early.
require "llm"
require "llm/tools"
class Agent < LLM::Agent
set name: "sysadmin",
description: "system administration agent",
model: "deepseek-v4-pro",
tools: [LLM::Tool::Exec]
end
llm = LLM.deepseek(key: ENV["KEY"])
agent = Agent.new(llm)
agent.talk "Run 'date'"Each provider is constructed with a class-level factory method on
LLM, and the resulting instance is passed to
LLM::Context
or
LLM::Agent. The
same API drives every one of them, so switching providers is a one-line
change.
- Anthropic (
LLM.anthropic) - Google (
LLM.google) - OpenAI (
LLM.openai) - DeepSeek (
LLM.deepseek) - DeepInfra (
LLM.deepinfra) - xAI (
LLM.xai) - Z.ai (
LLM.zai) - Moonshot (Kimi) (
LLM.moonshot) - OpenRouter (
LLM.openrouter) - Alibaba (Qwen3) (
LLM.alibaba, alsoLLM.aliyun) - Mistral (
LLM.mistral) - AWS Bedrock (
LLM.bedrock) - Ollama (
LLM.ollama) - llama.cpp (
LLM.llamacpp)
Implicit
Cloud providers can infer their API key automatically from a set of common defaults that are defined by the models.dev registry that is also distributed with llm.rb.
llm = LLM.openai
llm = LLM.anthropic
llm = LLM.google
llm = LLM.deepseek
llm = LLM.deepinfra
llm = LLM.xai
llm = LLM.zai
llm = LLM.moonshot
llm = LLM.openrouter
llm = LLM.alibaba # also: LLM.aliyun
llm = LLM.mistral
llm = LLM.bedrockExplicit
The key option can also be providied explicitly, and certain
providers (eg ollama, llamacpp) usually do not require an API
key at all.
llm = LLM.openai(key: ENV["OPENAI_API_KEY"])
llm = LLM.anthropic(key: ENV["ANTHROPIC_API_KEY"])
llm = LLM.google(key: ENV["GOOGLE_API_KEY"])
llm = LLM.deepseek(key: ENV["DEEPSEEK_API_KEY"])
llm = LLM.deepinfra(key: ENV["DEEPINFRA_API_KEY"])
llm = LLM.xai(key: ENV["XAI_API_KEY"])
llm = LLM.zai(key: ENV["ZHIPU_API_KEY"])
llm = LLM.moonshot(key: ENV["MOONSHOT_API_KEY"])
llm = LLM.openrouter(key: ENV["OPENROUTER_API_KEY"])
llm = LLM.alibaba(key: ENV["DASHSCOPE_API_KEY"]) # also: LLM.aliyun
llm = LLM.mistral(key: ENV["MISTRAL_API_KEY"])
llm = LLM.bedrock(
access_key_id: ENV["AWS_ACCESS_KEY_ID"],
secret_access_key: ENV["AWS_SECRET_ACCESS_KEY"],
region: ENV["AWS_REGION"]
)Model Registry
Each provider ships its model catalog, pricing, limits, and modalities with the gem, sourced from models.dev. Reach it from any provider, context, or agent, enumerate models, or sort them by price.
require "llm"
llm = LLM.openai
registry = llm.registry # => LLM::Provider#registry
cheapest = registry.models.sort.first # => LLM::Registry::Model
cheapest.id # => "text-embedding-3-small"
cheapest.context_window # => 8191
cheapest.structured_output? # => falseTransports
The transport: option selects which HTTP library a provider uses for
network communication. Three backends ship out of the box: net/http
is always available and the default, net/http/persistent pools
connections for many requests to the same host, and curb wraps
libcurl. They share one interface, so switching is a one-word change.
llm = LLM.deepseek(
key: ENV["KEY"],
transport: :net_http_persistent
)Timeouts
Providers accept two timeouts:
connect_timeout- opening the connection. Defaults to 5 seconds.read_timeout- waiting for a response on an idle connection. Defaults to 600 seconds (10 minutes).
The longer read timeout leaves room for slow reasoning models and
local models. The legacy timeout: option remains as a shorthand for
read_timeout. Timeouts are retriable:
LLM::Agent
retries a timed out request up to its
retry_budget
(five by default), so a dropped connection or a slow first token
is often something we can recover from.
llm = LLM.deepseek(
connect_timeout: 5, # opening the connection
read_timeout: 600 # waiting for the next bytes
)Headers
Providers can accept a custom set of headers with
the LLM::Provider#with method.
For example, you could set a custom User-Agent header,
or provide headers that carry special meaning to
certain providers (eg OpenAI, OpenRouter).
llm = LLM.openrouter
llm = llm.with("HTTP-Referer" => "https://example.com")
llm = llm.with("X-OpenRouter-Title" => "Example App")Most providers offer an embedding model that can be used for semantic search, or similarity search. An embedding model can generate embeddings that can then be stored in a database that is optimized for storing and querying vectors, such as SQLite's sqlite-vec or PostgreSQL's pg-vector.
llm.rb also includes support for OpenAI's vector store API. It provides a vector database as a HTTP service but we won't cover that here.
require "llm"
llm = LLM.openai(key: ENV["KEY"])
body = "llm.rb is Ruby's capable AI runtime."
embedding = llm.embed([body]).embeddings.first
# Document is your ActiveRecord or Sequel model
# with a vector column (e.g. sqlite-vec or pgvector)
Document.create!(
title: "llm.rb",
body:,
embedding:,
)A handful of providers can generate images from a text prompt. OpenAI, Google, xAI, and DeepInfra all support it. The API is the same across providers:
require "llm"
llm = LLM.openai(key: ENV["KEY"])
res = llm.images.create(prompt: "a dog on a rocket to the moon")
IO.copy_stream res.images[0], "rocket.png"DeepSeek does not have a dedicated image model, but the runtime
generates SVG vector graphics through its text model. Each
generation produces a valid SVG document that can be converted
to PNG with tools like rsvg-convert. Pass an existing agent
to maintain a session across generations:
require "llm"
llm = LLM.deepseek(key: ENV["KEY"])
##
# First generation
res = llm.images.create(prompt: "a rocket on the moon")
IO.copy_stream res.images[0], "rocket.svg"
##
# Refine with follow-up prompts (shares context)
res = llm.images.create(prompt: "add a dog next to the rocket",
agent: res.agent)
IO.copy_stream res.images[0], "rocket-with-dog.svg"What about local LLM support?
The following providers can be run used with models that are running on your own hardware.
- Ollama
- Llamacpp
I have a limited budget. What should I do?
There are a few options. The first option is to host your own model, and use the ollama or llamacpp providers. This can be difficult though because a capable model requires hardware that can match it. If you have the ability to self-host, this would be my first option.
The second option is DeepSeek.
The deepseek-v4-flash model costs pennies to use.
And llm.rb has been optimized for deepseek. For example,
DeepSeek does not have image generation capabilities
but on the llm.rb runtime it does (vector graphics only,
though).
The same is true for structured outputs. DeepSeek does
not support structured outputs in the same way as OpenAI or
Google, but the llm.rb runtime makes it appear as
though it does, through the json_object response
type.
If you're on a budget, DeepSeek is hard to beat.
Sources other than GitHub?
We are on the radicle.network as well.
Every commit that lands on GitHub also lands on Radicle.
Our repository ID is z2PtfQ6dYwyYaW2aGrztG1sMyDmCE.
Browse on the
web.
Who maintains llm.rb?
The llm.rb project was started more than three years ago by @altruby and @antaz. The primary maintainer is @altruby. Over those three years multiple other contributors have contributed to llm.rb as well, and new contributors are always welcome.
How well tested is llm.rb?
It is battle tested daily.
The console that is distributed with llm.rb is used to build llm.rb so there is a healthy, active feedback loop. It also powers the r.uby.dev website where multiple llm.rb agents are deployed. I'm aware of at least one production Rails deployment at a large-ish company.
And this git repository includes llm.rb agents that help me maintain the documentation and perform other repository maintainence. The feedback loop is constant. Outside of that there is a large test suite that covers live requests (recorded by VCR) and database interactions.
This software is released under the terms of the MIT license.
See LICENSE for details.
