Repository navigation
Route agents to the guides that answer study, cancer-type and subtype questions - #159
Merged
Merged
Conversation
… questions In benchmark run 20260926-1634 the guide fixes from #150 had no effect on the questions they targeted because agents never opened those guides: - Q77 (Minerva viewer for the OHSU HTAN sample): list_studies("OHSU HTAN") returned [] and both models concluded the study doesn't exist; neither read study-resolution-guide, which maps HTAN atlas codes (hta9 = OHSU). - Q107 (what cancer types are in the database): neither model read any guide; faq-guide has the type_of_cancer recipe. - Q103 (adenoid cystic carcinoma drivers): both queried all of acc_2019; sample-filtering-guide has the ONCOTREE_CODE subtype filter. LibreChat agents only see tool and guide descriptions, so the triggers go there: list_guides descriptions for study-resolution, faq and sample-filtering name the questions they answer; the clickhouse_run_select_query guide list adds faq-guide and the triggers; list_studies returns a note pointing to study-resolution-guide when a search matches nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NEwSZhYcMhmj5d8J4o7Jft
ForisKuang
approved these changes
Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The guide fixes in #150 had no effect on the benchmark questions they targeted, because agents never opened those guides. Evidence from cbioportal-mcp-qa run
20260926-1634(claude-code runner, beta prompt, this server at07ffe44), both Haiku and Sonnet:list_studies("OHSU HTAN")→[], thenlist_studies("HTAN")listedbrca_hta9_htan_2022but neither mapped it to OHSU; one read only external-resources-guide. Both said the study doesn't exist.hta9= OHSU)type_of_cancer/ OncoTree directly; read no guide.cancer_study.type_of_cancer_id)top_mutated_genes_in_study('acc_2019')on the whole mixed study.ONCOTREE_CODE = 'ACYC')LibreChat agents only see tool and guide descriptions (the server's
system-prompt.mddoesn't reach them), so the routing goes there:list_guidesdescriptions now name the questions they answer:clickhouse_run_select_queryguide list adds faq-guide and the same triggers for study-resolution and sample-filtering.list_studies: a search with no matches returns a single{"note": …}pointing to study-resolution-guide (HTAN atlas codes, fewer words, other instances) instead of[], matching the tool's existing[{"error": …}]shape.Tests: 128 pass (new: no-match note; guide descriptions and the query tool name the triggers). Ruff: no new findings beyond line length, consistent with the file's existing long description strings.
Targets cbioportal-mcp-qa benchmark Q77, Q107, Q103; verify with the next claude-code run after this deploys.
🤖 Generated with Claude Code
https://claude.ai/code/session_01NEwSZhYcMhmj5d8J4o7Jft