Releases: posit-dev/chatlas
Release list
v0.23.0
New features
-
Chatgainsclose()andclose_async()methods (plus context-manager support) for releasing resources held by the provider -- HTTP connection pools, the Snowflake Snowpark session/connection, and (viaclose_async()) MCP server sessions. This is useful in long-lived applications like Shiny that create a chat per user session:session.on_ended(chat.close). Providers only close resources they created themselves; caller-supplied clients are left open. -
ChatSnowflake()gains asessionparameter for supplying an existingsnowflake.snowpark.Session, mirroringChatDatabricks()'sworkspace_client. This lets one session be shared across multiple chats;Chat.close()only closes sessions that chatlas created itself, leaving caller-supplied sessions open. -
When running on Posit Connect, chatlas now forwards the Shiny viewer's session token to Connect's LLM gateway (as a
Posit-Connect-User-Session-Tokenheader) so gateway usage can be attributed to the viewer. This happens automatically for Shiny content and only affects requests to the gateway.
Changes
ChatHuggingFace()'s defaultmodelis nowQwen/Qwen3-235B-A22B-Instruct-2507(previouslymeta-llama/Llama-3.1-8B-Instruct), matching ellmer's default. (#414)ChatBedrock()now defaultsbase_urlto the official AWS SDKs' endpoint override environment variables when set:AWS_ENDPOINT_URL_BEDROCK_RUNTIMEforapi="converse", andAWS_ENDPOINT_URL_BEDROCK_MANTLEforapi="messages"andapi="responses". Similarly,ChatAnthropic()respects theANTHROPIC_BASE_URLenvironment variable (via the anthropic SDK). Setting these variables is enough to route requests through a proxy or gateway, so you don't have to passbase_urlon every call.
Bug fixes
ChatDatabricks()no longer drops the assistant's reply from the conversation when a GPT-OSS endpoint streams typed content. The typed part array was merged into the accumulated completion before it was normalized, so every later text delta was appended to it one character at a time and the finished turn came back empty. (#409).to_solver()no longer corrupts the system prompt or the prior turns it reads out of Inspect AI's message state. The system prompt was being set to therepr()of theChatMessageSystemobject rather than its text, and message content arriving in Inspect AI'sstrform (rather than as a list ofContent) was iterated one character at a time. (#407)ChatGoogle()no longer raisesValueError: Unknown content type: ContentThinkingon the second and later turns whenreasoningis enabled; thinking content is now replayed to the model as thought parts, and thethought_signatureon thought parts is preserved (previously only tool-call parts kept it). (#403)ChatOllama()now distinguishes a remote endpoint it can't reach from a genuinely missing local install, and validates a suppliedmodelagainst/api/tagsat construction time instead of only whenmodelis omitted. (#393)
v0.22.0
Changes
ChatAnthropic(),ChatBedrock(), andChatPosit()now requireanthropic>=1.0.0. As a result, customhttp_clients passed to Anthropic-backed providers must now behttpx2clients (rather thanhttpx), matching the anthropic SDK's own requirement.ChatBedrock()andChatBedrockAnthropic()now raise at construction if no AWS region can be resolved (from theaws_regionargument, theAWS_REGION/AWS_DEFAULT_REGIONenvironment variables, or the AWS profile), rather than silently defaulting tous-east-1. This behavior change is inherited from anthropic 1.0.- On Anthropic-backed providers, the
temperature,top_p, andtop_kmodel parameters are deprecated: anthropic 1.0 removed them from the request schema since current models ignore them. chatlas now forwards them viaextra_body(with aDeprecationWarning) so older models that still honor them keep working; this forwarding will be removed in a future release.
Added
Chatgains a settable.conversation_idproperty. When set, the identifier is recorded as thegen_ai.conversation.idattribute on the OpenTelemetrychatandinvoke_agentspans, allowing backends to group spans belonging to the same conversation (per the OpenTelemetry GenAI semantic conventions). Developer-facing: intended for frameworks that manage conversation history; chatlas never generates an identifier on its own.
Bug fixes
ChatDatabricks()no longer crashes on GPT-OSS endpoints that returnmessage.content/delta.contentas a list of typed parts instead of a plain string; text parts are concatenated back into a string and reasoning summaries becomeContentThinking/ContentThinkingDelta, both streaming and non-streaming. (#392)
v0.21.2
Bug fixes
- Rich tool results containing images or PDFs remain semantic
ContentToolResultobjects in saved chat history, so restoring a conversation no longer exposes provider-only XML and media as user-authored content. ChatOpenAI()web-search citations no longer report the inline Markdown source link asContentCitation.grounded_span; OpenAI citation offsets identify that marker rather than the supported answer text. (#388)ChatOpenAI()no longer crashes while streaming web-search citations with current OpenAI SDKs, which emit annotations as model objects rather than dictionaries.ChatAnthropic()now identifies Claude 5 and later models as supporting native structured output, so automatic mode does not fall back to tools. (#390)
v0.21.1
Chatlas 0.21.1 improves compatibility with the OpenAI 3 SDK and fixes citation and token-cost edge cases.
What is fixed
- OpenAI-based providers now support native
httpx2clients from the OpenAI 3 SDK. Legacyhttpxclients remain supported at runtime during migration. (#387) - Anthropic citations backed by
document_indexnow resolve to source URLs correctly, including when prior turns contain document-shaped content. (#382) - Token cost lookups now handle output-only models and ellmer's versioned pricing-data format. (#382)
Custom provider compatibility
Custom Provider implementations must accept a turns keyword argument on .stream_content(), .stream_turn(), and .value_turn() so response references can be resolved against the complete request history.
chatlas 0.21.0
New features
- New
ChatBedrock()gives full access to AWS Bedrock's model catalog — Nova, Llama, Mistral, DeepSeek, Qwen, plus the GPT-5 family, Grok 4.3, and Gemma 4 — not just Claude, none of which were previously available through chatlas. It replacesChatBedrockAnthropic()as the recommended entrypoint; the right request format ("converse","responses", or"messages") is picked automatically from the model name, or setapiexplicitly.
Improvements
- Updated default models to match the latest generation:
- Anthropic / BedrockAnthropic / Posit:
claude-sonnet-5 - OpenAI / Completions / OpenRouter:
gpt-5.6-terra
- Anthropic / BedrockAnthropic / Posit:
- Echoing turns in the console and notebooks got a round of display improvements:
- Reasoning/thinking content now actually shows up — it used to silently disappear, since it was wrapped in literal
<thinking>tags that a markdown renderer treated as an HTML block and dropped. It renders in a collapsible "Thinking" panel (a<details>block in notebooks) that stays open while streaming and collapses once done, and is capped to the most recent lines when long. (#361) - Long tool results no longer flood the screen — they collapse/truncate with a clear count of what's hidden, scrolling internally in notebooks beyond a bounded height. (#361)
- Images from models or tools now render as compact thumbnails instead of raw base64 data.
- Web search, fetch, and citation activity is now visible too, grouped into a "Searched the web" / "Read the web" panel that marks which sources were actually cited. (#256)
- All of these size limits are tunable via
Chat.set_echo_options()(tool_result_max_lines,tool_result_max_height,thinking_max_lines,image_max_lines,web_activity_max_sources), and can be turned off entirely withNone.
- Reasoning/thinking content now actually shows up — it used to silently disappear, since it was wrapped in literal
- Registering a built-in tool (
tool_web_search(),tool_web_fetch()) with a provider that can't run it now fails immediately with a clear error naming the tool and provider, instead of silently no-op'ing or dying deep inside a later request. (#367)
Bug fixes
echo="all"no longer displays tool results twice — once in full as part of the user turn, and again on their own.Chat.set_echo_options(css_styles=)now actually applies in notebooks.- Tool names and argument names are now HTML-escaped in notebook/shiny rendering, closing an HTML-injection hole.
- A tool that reports progress by yielding more than once, or an MCP server that answers a call with several content parts (text plus an image, say), no longer breaks the request.
register_mcp_tools_stdio_async()andregister_mcp_tools_http_stream_async()no longer fail when an MCP server leaves some tool annotations unset.ContentCitation,ContentToolRequestFetch, andContentToolResponseFetchnow actually render, instead of silently vanishing due to a markdown link-reference parsing quirk.ChatAnthropic()(andChatBedrockAnthropic()) now bill refusal-fallback turns at the correct (serving model's) rate rather than the originally requested model's, mirroring ellmer's equivalent fix.
Breaking changes
ChatGithub()is now defunct: it always raisesRuntimeError. GitHub Models was retired on 2026-07-30, so the underlying API no longer works. UseChatGoogle()(offers a free tier) orChatPosit()(offers a free trial) instead.
chatlas 0.20.0
New features
Chatgains a.filesaccessor for uploading files to a provider once and referencing them across turns without re-sending bytes, plus listing, fetching metadata, downloading, and deleting them. Supported for OpenAI, Anthropic, and Google Gemini. A newContentUploadedtype represents the reference and can be constructed directly to point at a file uploaded out-of-band (e.g. a Vertexgs://URI). For Google,upload()waits for Gemini to finish processing large media (video, audio) before returning, since the API rejects references to files that aren't yetACTIVE.- Web search and fetch results now surface their citations across all three providers (OpenAI, Anthropic, Google), both progressively during streaming and on the final turn.
ContentCitationnests a typedsource(aSourcesubclass —WebSourcetoday, carryingurl/title) instead of flaturl/titlefields, and carriesgrounded_span(the answer-side span it grounds) pluscited_quote(the source-side quote, populated forChatAnthropic()web search).sourceis optional — a citation can ground answer text with no resolvable link.ContentCitation,Source, andWebSourceare exported fromchatlas.types. A future file/document/RAG source becomes anotherSourcesubclass without breakingContentCitation.source; note thatContentToolResponseSearch.sourcesis typed narrowly aslist[WebSource]and would need widening at that point.- When streaming with
content="all",ContentCitationobjects are emitted as citations arrive — interleaved with text for OpenAI and Anthropic, at stream-end for Google. Its position in the stream (relative to surrounding text) is the placement signal for rendering footnote markers. - On the final turn,
ContentCitationitems appear in the turn'scontentslist after theContentTextthey ground, in the order the provider reported them. Since a turn's text arrives as one accumulatedContentText, position no longer narrows a citation to a span within it — usegrounded_spanfor that.
- When streaming with
batch_chat()now supportsChatGoogle()(Gemini Developer API batch jobs). Batch is also now documented as supported forChatGroq(), which already worked via its OpenAI-compatible provider. (Vertex AI is not supported, since its batch API requires GCS bucket URIs instead of inline requests.)ChatOllama()gains areasoning_effortparameter to enable extended "thinking" for models that support it (e.g. qwen3, gpt-oss).Chat.token_count()gained aninclude=argument:"new"(default) counts just the given input, while"complete"estimates the total tokens for the next request, including history and system prompt where the provider supports it.
Improvements
ChatGoogle()andChatVertex()now default togemini-3.5-flashinstead of the oldergemini-2.5-flash.ChatGroq()now defaults toopenai/gpt-oss-20binstead ofllama-3.1-8b-instant.- Built-in web search and fetch content (
ContentToolRequestSearch/ContentToolResponseSearchandContentToolRequestFetch/ContentToolResponseFetch) is now also emitted while streaming withcontent="all", forChatOpenAI(),ChatAnthropic(), andChatGoogle(). Previously it appeared only on the completed turn, so a UI had no way to show search activity until the whole response had arrived. ContentToolResponseFetchgained a normalizedstatusfield ("success","error", orNonewhen the provider doesn't report an outcome). Providers' finer-grained reasons (Anthropic'surl_not_allowed, Google'sPAYWALL, …) aren't aligned across providers, so they stay available inextra.ChatOpenAI().token_count()now uses OpenAI's token-counting endpoint for accurate, tool-aware counts instead of a localtiktokenestimate.
Changes
- MCP support now requires
mcp>=2.0.0. The 2.0 release of themcpSDK renamed its model fields (and removedmcp.server.fastmcp.FastMCP), so oldermcpversions are no longer compatible. This only affects users of the optionalmcpextra (i.e.,register_mcp_tools_*()). Turn.finish_reasonis now normalized to a consistent set of values ("success","tool_use","max_tokens","content_filter","context_window","stop_sequence") across most providers, so you no longer need provider-specific logic to check why a turn ended. Previously each provider surfaced its own raw string (e.g. Anthropic's"end_turn"/"tool_use"vs. OpenAI Completions'"stop"/"tool_calls"vs. Google's"STOP"/"SAFETY"), so the same outcome could require different checks depending on whichChat*()you used. Reasons chatlas doesn't yet recognize still pass through unchanged.
Bug fixes
ChatGoogle()no longer errors when mixing custom tools and built-in tools (e.g.tool_web_search()) on Gemini 3+ models.- Turns containing web search/fetch content can now be passed to a different provider (e.g.
ChatAnthropic().set_turns(openai_chat.get_turns())). Previously this raisedValueError: Unsupported content typeonChatOpenAI(), andChatAnthropic()forwarded the other provider's raw payload as if it were its own, producing an invalid request. Each provider now replays only the built-in tool content it produced and drops the rest. ChatOpenAI()web searchopen_pageactions now surface asContentToolRequestFetch(with the URL) rather than aContentToolRequestSearchwhose "query" was the URL, so renderers no longer show "searched for: https://…". Relatedly, asearchaction that reports only the pluralqueriesfield no longer falls through to the literal string"web search".ChatGoogle()now records its built-in web search and URL-context work in the assistant turn, asContentToolRequestSearch/ContentToolResponseSearchfor grounded searches andContentToolRequestFetch/ContentToolResponseFetchfor fetched URLs. Previouslytool_web_search()andtool_web_fetch()worked but reported nothing about what was searched or fetched, unlikeChatAnthropic()andChatOpenAI(). Google's rawgrounding_metadata/url_metadatais kept on each item'sextra..chat_structured()now explains itself when the response is cut short. Previously, extracting a data model large enough to hit the model's output limit failed with a bareJSONDecodeErrorpointing at a column number in the truncated JSON, giving no hint thatmax_tokenswas the problem (#315). It now raises aValueErrornamingmax_tokensand suggesting you raise it. Responses truncated by the context window, or stopped by the provider's content filter, are reported the same way. Plain.chat()warns instead of erroring, since a partial response is still usable there — previously it returned truncated text with no indication anything was missing.- Streaming two adjacent pieces of same-typed content that define no merge behavior (e.g. two tool requests) no longer raises
TypeError. They are now appended as separate content instead.
Breaking changes
ContentToolResponseSearch.urls(alist[str]) has been replaced by.sources(alist[WebSource]), each carrying the result'surlandtitle. Code reading.urlsshould switch to[s.url for s in x.sources].- The
Providerabstract base class changed shape, which affects third-partyProvidersubclasses (not users of the built-inChat*()functions):stream_text()was removed, andstream_content()both returns aSequence[Content](subsuming whatstream_text()did) and takes a secondcompletionargument holding the merged-so-far completion. Implementations needing state across chunks should read it fromcompletionrather than storing it onself, since one provider instance is shared across forked chats. Provider.token_count()/token_count_async()now take aturns: list[Turn]argument instead of*args: Content | str(affects customProvidersubclasses only).
chatlas 0.19.2
New features
- The
.app()method now includes latest shinychat features like history, file attachments, etc.
chatlas 0.19.1
chatlas 0.19.0
New features
- chatlas is now instrumented with OpenTelemetry (OTel) out of the box, making it much easier to see how your app behaves in production — where time goes, how many tokens you're spending, which tools run, and where things fail. Without writing any tracing code, you get spans that capture the full structure of a conversation as one connected trace: an
invoke_agentspan over the whole chat loop, achatspan per model call, and anexecute_toolspan per tool invocation, with attributes (token usage, response model/ID, tool errors) that follow the OTel GenAI semantic conventions. Because chatlas keeps its spans active during each call, HTTP spans from provider instrumentors and any spans your own tools emit nest underneath automatically. Point it at any OTel-compatible backend (Logfire, Datadog, Honeycomb, Jaeger, …); message content is omitted by default and opt-in viaOTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true. See the monitoring guide to get started. (#310) Chatgains amodelproperty to get (or set) the model after the chat is created. Setting it does not validate the model name.ChatGoogle()'sreasoningparameter now accepts a string thinking level ("minimal","low","medium", or"high") in addition to an integer token budget.ChatAnthropic()'sreasoningparameter now accepts a string effort level ("low","medium","high","xhigh", or"max") to enable Claude's adaptive thinking, in addition to an integer token budget.
Bug fixes
- OpenAI-compatible providers (e.g.,
ChatOllama()with models like qwen3) now capture thinking content returned in areasoningfield, not justreasoning_content. Previously this thinking content was silently dropped.
v0.18.1
Improvements
Content.tagify()implementations (ContentToolRequest,ContentToolResult,ContentThinking) now annotate their return type ashtmltools.Tagifiedand fully tagify their output, complying with htmltools 0.7.0's tightened Tagifiable contract. Embedding these contents inside another.tagify()recursion no longer trips the new boundary check in htmltools 0.7.0. (#311)
Bug fixes
ContentPDFis now exported fromchatlas.types, matching all otherContentsubclasses. (#312)
Full changelog: https://github.com/posit-dev/chatlas/blob/main/CHANGELOG.md