Skip to content

Releases: posit-dev/chatlas

v0.23.0

Choose a tag to compare

@cpsievert cpsievert released this 04 Sep 16:50

New features

  • Chat gains close() and close_async() methods (plus context-manager support) for releasing resources held by the provider -- HTTP connection pools, the Snowflake Snowpark session/connection, and (via close_async()) MCP server sessions. This is useful in long-lived applications like Shiny that create a chat per user session: session.on_ended(chat.close). Providers only close resources they created themselves; caller-supplied clients are left open.

  • ChatSnowflake() gains a session parameter for supplying an existing snowflake.snowpark.Session, mirroring ChatDatabricks()'s workspace_client. This lets one session be shared across multiple chats; Chat.close() only closes sessions that chatlas created itself, leaving caller-supplied sessions open.

  • When running on Posit Connect, chatlas now forwards the Shiny viewer's session token to Connect's LLM gateway (as a Posit-Connect-User-Session-Token header) so gateway usage can be attributed to the viewer. This happens automatically for Shiny content and only affects requests to the gateway.

Changes

  • ChatHuggingFace()'s default model is now Qwen/Qwen3-235B-A22B-Instruct-2507 (previously meta-llama/Llama-3.1-8B-Instruct), matching ellmer's default. (#414)
  • ChatBedrock() now defaults base_url to the official AWS SDKs' endpoint override environment variables when set: AWS_ENDPOINT_URL_BEDROCK_RUNTIME for api="converse", and AWS_ENDPOINT_URL_BEDROCK_MANTLE for api="messages" and api="responses". Similarly, ChatAnthropic() respects the ANTHROPIC_BASE_URL environment variable (via the anthropic SDK). Setting these variables is enough to route requests through a proxy or gateway, so you don't have to pass base_url on every call.

Bug fixes

  • ChatDatabricks() no longer drops the assistant's reply from the conversation when a GPT-OSS endpoint streams typed content. The typed part array was merged into the accumulated completion before it was normalized, so every later text delta was appended to it one character at a time and the finished turn came back empty. (#409)
  • .to_solver() no longer corrupts the system prompt or the prior turns it reads out of Inspect AI's message state. The system prompt was being set to the repr() of the ChatMessageSystem object rather than its text, and message content arriving in Inspect AI's str form (rather than as a list of Content) was iterated one character at a time. (#407)
  • ChatGoogle() no longer raises ValueError: Unknown content type: ContentThinking on the second and later turns when reasoning is enabled; thinking content is now replayed to the model as thought parts, and the thought_signature on thought parts is preserved (previously only tool-call parts kept it). (#403)
  • ChatOllama() now distinguishes a remote endpoint it can't reach from a genuinely missing local install, and validates a supplied model against /api/tags at construction time instead of only when model is omitted. (#393)

v0.22.0

Choose a tag to compare

@cpsievert cpsievert released this 25 Aug 20:09

Changes

  • ChatAnthropic(), ChatBedrock(), and ChatPosit() now require anthropic>=1.0.0. As a result, custom http_clients passed to Anthropic-backed providers must now be httpx2 clients (rather than httpx), matching the anthropic SDK's own requirement.
  • ChatBedrock() and ChatBedrockAnthropic() now raise at construction if no AWS region can be resolved (from the aws_region argument, the AWS_REGION/AWS_DEFAULT_REGION environment variables, or the AWS profile), rather than silently defaulting to us-east-1. This behavior change is inherited from anthropic 1.0.
  • On Anthropic-backed providers, the temperature, top_p, and top_k model parameters are deprecated: anthropic 1.0 removed them from the request schema since current models ignore them. chatlas now forwards them via extra_body (with a DeprecationWarning) so older models that still honor them keep working; this forwarding will be removed in a future release.

Added

  • Chat gains a settable .conversation_id property. When set, the identifier is recorded as the gen_ai.conversation.id attribute on the OpenTelemetry chat and invoke_agent spans, allowing backends to group spans belonging to the same conversation (per the OpenTelemetry GenAI semantic conventions). Developer-facing: intended for frameworks that manage conversation history; chatlas never generates an identifier on its own.

Bug fixes

  • ChatDatabricks() no longer crashes on GPT-OSS endpoints that return message.content / delta.content as a list of typed parts instead of a plain string; text parts are concatenated back into a string and reasoning summaries become ContentThinking/ContentThinkingDelta, both streaming and non-streaming. (#392)

v0.21.2

Choose a tag to compare

@cpsievert cpsievert released this 20 Aug 18:17

Bug fixes

  • Rich tool results containing images or PDFs remain semantic ContentToolResult objects in saved chat history, so restoring a conversation no longer exposes provider-only XML and media as user-authored content.
  • ChatOpenAI() web-search citations no longer report the inline Markdown source link as ContentCitation.grounded_span; OpenAI citation offsets identify that marker rather than the supported answer text. (#388)
  • ChatOpenAI() no longer crashes while streaming web-search citations with current OpenAI SDKs, which emit annotations as model objects rather than dictionaries.
  • ChatAnthropic() now identifies Claude 5 and later models as supporting native structured output, so automatic mode does not fall back to tools. (#390)

v0.21.1

Choose a tag to compare

@cpsievert cpsievert released this 12 Aug 17:13

Chatlas 0.21.1 improves compatibility with the OpenAI 3 SDK and fixes citation and token-cost edge cases.

What is fixed

  • OpenAI-based providers now support native httpx2 clients from the OpenAI 3 SDK. Legacy httpx clients remain supported at runtime during migration. (#387)
  • Anthropic citations backed by document_index now resolve to source URLs correctly, including when prior turns contain document-shaped content. (#382)
  • Token cost lookups now handle output-only models and ellmer's versioned pricing-data format. (#382)

Custom provider compatibility

Custom Provider implementations must accept a turns keyword argument on .stream_content(), .stream_turn(), and .value_turn() so response references can be resolved against the complete request history.

chatlas 0.21.0

Choose a tag to compare

@cpsievert cpsievert released this 04 Aug 17:33

New features

  • New ChatBedrock() gives full access to AWS Bedrock's model catalog — Nova, Llama, Mistral, DeepSeek, Qwen, plus the GPT-5 family, Grok 4.3, and Gemma 4 — not just Claude, none of which were previously available through chatlas. It replaces ChatBedrockAnthropic() as the recommended entrypoint; the right request format ("converse", "responses", or "messages") is picked automatically from the model name, or set api explicitly.

Improvements

  • Updated default models to match the latest generation:
    • Anthropic / BedrockAnthropic / Posit: claude-sonnet-5
    • OpenAI / Completions / OpenRouter: gpt-5.6-terra
  • Echoing turns in the console and notebooks got a round of display improvements:
    • Reasoning/thinking content now actually shows up — it used to silently disappear, since it was wrapped in literal <thinking> tags that a markdown renderer treated as an HTML block and dropped. It renders in a collapsible "Thinking" panel (a <details> block in notebooks) that stays open while streaming and collapses once done, and is capped to the most recent lines when long. (#361)
    • Long tool results no longer flood the screen — they collapse/truncate with a clear count of what's hidden, scrolling internally in notebooks beyond a bounded height. (#361)
    • Images from models or tools now render as compact thumbnails instead of raw base64 data.
    • Web search, fetch, and citation activity is now visible too, grouped into a "Searched the web" / "Read the web" panel that marks which sources were actually cited. (#256)
    • All of these size limits are tunable via Chat.set_echo_options() (tool_result_max_lines, tool_result_max_height, thinking_max_lines, image_max_lines, web_activity_max_sources), and can be turned off entirely with None.
  • Registering a built-in tool (tool_web_search(), tool_web_fetch()) with a provider that can't run it now fails immediately with a clear error naming the tool and provider, instead of silently no-op'ing or dying deep inside a later request. (#367)

Bug fixes

  • echo="all" no longer displays tool results twice — once in full as part of the user turn, and again on their own.
  • Chat.set_echo_options(css_styles=) now actually applies in notebooks.
  • Tool names and argument names are now HTML-escaped in notebook/shiny rendering, closing an HTML-injection hole.
  • A tool that reports progress by yielding more than once, or an MCP server that answers a call with several content parts (text plus an image, say), no longer breaks the request.
  • register_mcp_tools_stdio_async() and register_mcp_tools_http_stream_async() no longer fail when an MCP server leaves some tool annotations unset.
  • ContentCitation, ContentToolRequestFetch, and ContentToolResponseFetch now actually render, instead of silently vanishing due to a markdown link-reference parsing quirk.
  • ChatAnthropic() (and ChatBedrockAnthropic()) now bill refusal-fallback turns at the correct (serving model's) rate rather than the originally requested model's, mirroring ellmer's equivalent fix.

Breaking changes

  • ChatGithub() is now defunct: it always raises RuntimeError. GitHub Models was retired on 2026-07-30, so the underlying API no longer works. Use ChatGoogle() (offers a free tier) or ChatPosit() (offers a free trial) instead.

chatlas 0.20.0

Choose a tag to compare

@cpsievert cpsievert released this 29 Jul 15:58

New features

  • Chat gains a .files accessor for uploading files to a provider once and referencing them across turns without re-sending bytes, plus listing, fetching metadata, downloading, and deleting them. Supported for OpenAI, Anthropic, and Google Gemini. A new ContentUploaded type represents the reference and can be constructed directly to point at a file uploaded out-of-band (e.g. a Vertex gs:// URI). For Google, upload() waits for Gemini to finish processing large media (video, audio) before returning, since the API rejects references to files that aren't yet ACTIVE.
  • Web search and fetch results now surface their citations across all three providers (OpenAI, Anthropic, Google), both progressively during streaming and on the final turn. ContentCitation nests a typed source (a Source subclass — WebSource today, carrying url/title) instead of flat url/title fields, and carries grounded_span (the answer-side span it grounds) plus cited_quote (the source-side quote, populated for ChatAnthropic() web search). source is optional — a citation can ground answer text with no resolvable link. ContentCitation, Source, and WebSource are exported from chatlas.types. A future file/document/RAG source becomes another Source subclass without breaking ContentCitation.source; note that ContentToolResponseSearch.sources is typed narrowly as list[WebSource] and would need widening at that point.
    • When streaming with content="all", ContentCitation objects are emitted as citations arrive — interleaved with text for OpenAI and Anthropic, at stream-end for Google. Its position in the stream (relative to surrounding text) is the placement signal for rendering footnote markers.
    • On the final turn, ContentCitation items appear in the turn's contents list after the ContentText they ground, in the order the provider reported them. Since a turn's text arrives as one accumulated ContentText, position no longer narrows a citation to a span within it — use grounded_span for that.
  • batch_chat() now supports ChatGoogle() (Gemini Developer API batch jobs). Batch is also now documented as supported for ChatGroq(), which already worked via its OpenAI-compatible provider. (Vertex AI is not supported, since its batch API requires GCS bucket URIs instead of inline requests.)
  • ChatOllama() gains a reasoning_effort parameter to enable extended "thinking" for models that support it (e.g. qwen3, gpt-oss).
  • Chat.token_count() gained an include= argument: "new" (default) counts just the given input, while "complete" estimates the total tokens for the next request, including history and system prompt where the provider supports it.

Improvements

  • ChatGoogle() and ChatVertex() now default to gemini-3.5-flash instead of the older gemini-2.5-flash.
  • ChatGroq() now defaults to openai/gpt-oss-20b instead of llama-3.1-8b-instant.
  • Built-in web search and fetch content (ContentToolRequestSearch/ContentToolResponseSearch and ContentToolRequestFetch/ContentToolResponseFetch) is now also emitted while streaming with content="all", for ChatOpenAI(), ChatAnthropic(), and ChatGoogle(). Previously it appeared only on the completed turn, so a UI had no way to show search activity until the whole response had arrived.
  • ContentToolResponseFetch gained a normalized status field ("success", "error", or None when the provider doesn't report an outcome). Providers' finer-grained reasons (Anthropic's url_not_allowed, Google's PAYWALL, …) aren't aligned across providers, so they stay available in extra.
  • ChatOpenAI().token_count() now uses OpenAI's token-counting endpoint for accurate, tool-aware counts instead of a local tiktoken estimate.

Changes

  • MCP support now requires mcp>=2.0.0. The 2.0 release of the mcp SDK renamed its model fields (and removed mcp.server.fastmcp.FastMCP), so older mcp versions are no longer compatible. This only affects users of the optional mcp extra (i.e., register_mcp_tools_*()).
  • Turn.finish_reason is now normalized to a consistent set of values ("success", "tool_use", "max_tokens", "content_filter", "context_window", "stop_sequence") across most providers, so you no longer need provider-specific logic to check why a turn ended. Previously each provider surfaced its own raw string (e.g. Anthropic's "end_turn"/"tool_use" vs. OpenAI Completions' "stop"/"tool_calls" vs. Google's "STOP"/"SAFETY"), so the same outcome could require different checks depending on which Chat*() you used. Reasons chatlas doesn't yet recognize still pass through unchanged.

Bug fixes

  • ChatGoogle() no longer errors when mixing custom tools and built-in tools (e.g. tool_web_search()) on Gemini 3+ models.
  • Turns containing web search/fetch content can now be passed to a different provider (e.g. ChatAnthropic().set_turns(openai_chat.get_turns())). Previously this raised ValueError: Unsupported content type on ChatOpenAI(), and ChatAnthropic() forwarded the other provider's raw payload as if it were its own, producing an invalid request. Each provider now replays only the built-in tool content it produced and drops the rest.
  • ChatOpenAI() web search open_page actions now surface as ContentToolRequestFetch (with the URL) rather than a ContentToolRequestSearch whose "query" was the URL, so renderers no longer show "searched for: https://…". Relatedly, a search action that reports only the plural queries field no longer falls through to the literal string "web search".
  • ChatGoogle() now records its built-in web search and URL-context work in the assistant turn, as ContentToolRequestSearch/ContentToolResponseSearch for grounded searches and ContentToolRequestFetch/ContentToolResponseFetch for fetched URLs. Previously tool_web_search() and tool_web_fetch() worked but reported nothing about what was searched or fetched, unlike ChatAnthropic() and ChatOpenAI(). Google's raw grounding_metadata/url_metadata is kept on each item's extra.
  • .chat_structured() now explains itself when the response is cut short. Previously, extracting a data model large enough to hit the model's output limit failed with a bare JSONDecodeError pointing at a column number in the truncated JSON, giving no hint that max_tokens was the problem (#315). It now raises a ValueError naming max_tokens and suggesting you raise it. Responses truncated by the context window, or stopped by the provider's content filter, are reported the same way. Plain .chat() warns instead of erroring, since a partial response is still usable there — previously it returned truncated text with no indication anything was missing.
  • Streaming two adjacent pieces of same-typed content that define no merge behavior (e.g. two tool requests) no longer raises TypeError. They are now appended as separate content instead.

Breaking changes

  • ContentToolResponseSearch.urls (a list[str]) has been replaced by .sources (a list[WebSource]), each carrying the result's url and title. Code reading .urls should switch to [s.url for s in x.sources].
  • The Provider abstract base class changed shape, which affects third-party Provider subclasses (not users of the built-in Chat*() functions): stream_text() was removed, and stream_content() both returns a Sequence[Content] (subsuming what stream_text() did) and takes a second completion argument holding the merged-so-far completion. Implementations needing state across chunks should read it from completion rather than storing it on self, since one provider instance is shared across forked chats.
  • Provider.token_count()/token_count_async() now take a turns: list[Turn] argument instead of *args: Content | str (affects custom Provider subclasses only).

chatlas 0.19.2

Choose a tag to compare

@cpsievert cpsievert released this 08 Jul 14:08

New features

  • The .app() method now includes latest shinychat features like history, file attachments, etc.

chatlas 0.19.1

Choose a tag to compare

@cpsievert cpsievert released this 01 Jul 22:21

New features

  • Added ChatPosit() for chatting via the Posit AI gateway. (#323)

chatlas 0.19.0

Choose a tag to compare

@cpsievert cpsievert released this 15 Jun 20:30

New features

  • chatlas is now instrumented with OpenTelemetry (OTel) out of the box, making it much easier to see how your app behaves in production — where time goes, how many tokens you're spending, which tools run, and where things fail. Without writing any tracing code, you get spans that capture the full structure of a conversation as one connected trace: an invoke_agent span over the whole chat loop, a chat span per model call, and an execute_tool span per tool invocation, with attributes (token usage, response model/ID, tool errors) that follow the OTel GenAI semantic conventions. Because chatlas keeps its spans active during each call, HTTP spans from provider instrumentors and any spans your own tools emit nest underneath automatically. Point it at any OTel-compatible backend (Logfire, Datadog, Honeycomb, Jaeger, …); message content is omitted by default and opt-in via OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true. See the monitoring guide to get started. (#310)
  • Chat gains a model property to get (or set) the model after the chat is created. Setting it does not validate the model name.
  • ChatGoogle()'s reasoning parameter now accepts a string thinking level ("minimal", "low", "medium", or "high") in addition to an integer token budget.
  • ChatAnthropic()'s reasoning parameter now accepts a string effort level ("low", "medium", "high", "xhigh", or "max") to enable Claude's adaptive thinking, in addition to an integer token budget.

Bug fixes

  • OpenAI-compatible providers (e.g., ChatOllama() with models like qwen3) now capture thinking content returned in a reasoning field, not just reasoning_content. Previously this thinking content was silently dropped.

v0.18.1

Choose a tag to compare

@schloerke schloerke released this 21 May 16:09
3c40bc4

Improvements

  • Content.tagify() implementations (ContentToolRequest, ContentToolResult, ContentThinking) now annotate their return type as htmltools.Tagified and fully tagify their output, complying with htmltools 0.7.0's tightened Tagifiable contract. Embedding these contents inside another .tagify() recursion no longer trips the new boundary check in htmltools 0.7.0. (#311)

Bug fixes

  • ContentPDF is now exported from chatlas.types, matching all other Content subclasses. (#312)

Full changelog: https://github.com/posit-dev/chatlas/blob/main/CHANGELOG.md