Benchmarks the deployed cBioPortalChat agent by asking it every question in
input/questions.yaml through LibreChat's Agents API, then grading the answers.
Because it calls the real deployment, a run measures what users get: the same system prompt, MCP tools
(cbioportal-database, cbioportal-navigator), prompt caching and model.
Results — one HTML report per run, published with GitHub Pages. Test sets — what each questions file covers, with every question, its reference and rubric.
For every question × model (× repeat):
- Answer from
POST /api/agents/v1/chat/completions, with latency and token usage (input, cache read, cache write, output) → estimated cost at Anthropic list prices. - Execution trace from Langfuse, matched by response id: number of LLM calls, every tool call, and tool errors (e.g. navigator schema errors).
- Grade, pass/fail: an LLM judge (default Sonnet 4.6 on Bedrock, deliberately not one of the models under
test; or through the local Claude Code subscription with
--judge-runner claude-code) checks the answer against the reference answer, expected links and thenotesrubric, using the criterion for the question's track (below). It also records whether the answer declined, so the report can separate precision (right when it answers) from coverage (how often it answers). The judge also sees the answer's tool calls (inputs and truncated results; guide text left out), so statistics the agent actually computed aren't marked as invented. Older claude-code runs are graded from their transcripts. - Objective checks: when the reference is a single number, whether the answer contains it (exact for counts, within rounding for decimals/percentages) — shown where it disagrees with the judge; and whether every study id in the answer's cBioPortal links exists (a hallucination signal).
Every question has a track, and the report shows pass rates per track and model:
| Track | Passes when |
|---|---|
data |
it states the fact in the reference (a count, frequency, list) |
navigation |
it gives a cBioPortal link to the right view: page, study, genes, filters |
analysis |
it uses the right cohort and method and doesn't invent statistics |
out_of_scope |
it clearly declines instead of making something up |
Questions without any reference are still asked and reported, but not graded.
results/ is published with GitHub Pages, so everything the benchmark writes there goes through one
redaction boundary (persist.py). run.json, summary.json, the HTML and Markdown reports, compare outputs,
transcripts and the recorded agent prompt are deep-scrubbed on the way out:
- Known values first: every secret the benchmark loads or can see is masked wherever it appears, also URL-encoded, JSON-escaped, backslash-escaped or HTML-escaped. That covers the LibreChat and Langfuse keys, the cBioAgent Mongo password read from its k8s secret, and every environment variable whose name contains PASS, PWD, SECRET, TOKEN, KEY, AUTH or CREDENTIAL. It also covers the password of any URI in the environment, both as written and percent-decoded. Values under 6 characters are ignored.
- Structure: the value of any key named like a credential (
password,*_token,apiKey,Authorization,Cookie, ...) is masked whole, and a list of strings is also read as a command line. - Patterns for anything else: URI userinfo, bearer tokens, key=value secrets, password flags of database
clients and
--password-style flags. They mask to the end of the line, or of the text when the quoting after them can't be trusted. - Failed commands are recorded as
<executable> exited with status Nand the redacted tail of their output, never with their arguments.
Screenshots are not redacted: they are pixels. The shots/ of a run only hold the public cBioPortal pages
that navigation answers linked to. Pass --no-screenshots to run or render to keep the page text without
them.
uv sync
cp .env.example .env # then fill it inLIBRECHAT_API_KEY: create under Settings → Agent API Keys on beta.chat.cbioportal.org. Beta and prod share a database, so the key works on both. Runs are billed to that account's token balance.LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY: for tool-call stats (optional; runs work without).- AWS credentials for the judge:
AWS_PROFILEwith Bedrock access (e.g.cdsi-imagine-490004633549). Not needed with--judge-runner claude-code(below).
Choosing the model per request requires cbioportal/librechat v0.8.7-custom-v3 or later on the target, where
the Agents API accepts a spec (modelSpec name) alongside the agent id.
# One question, to check the setup
uv run cbioportal-mcp-qa ask "How many studies are in cBioPortal?" --model haiku
# Full benchmark, Haiku vs Sonnet on beta
uv run cbioportal-mcp-qa run --models haiku,sonnet
# A subset, three repeats each to measure consistency
uv run cbioportal-mcp-qa run --questions 1-10 --repeats 3
# Continue an interrupted run (re-asks only missing or failed answers). It keeps the run's --questions,
# --questions-file, --no-grade, judge, concurrency, rendering and --wait-on-limit unless you give them again;
# the --claude-code-* opt-ins must be given again.
uv run cbioportal-mcp-qa run --resume 20260923-1800
# Re-attach traces (Langfuse ingestion can lag), regrade, or re-render
uv run cbioportal-mcp-qa traces 20260923-1800
uv run cbioportal-mcp-qa grade 20260923-1800 --regrade
# After fixing references in input/questions.yaml, regrade an existing run against them (only the expected
# answer, links, notes and track are refreshed; the judge still sees the question text and history that were asked)
uv run cbioportal-mcp-qa grade 20260923-1800 --refresh-questions
uv run cbioportal-mcp-qa report 20260923-1800--target |
Agent | --models |
Request |
|---|---|---|---|
beta |
unified beta agent agent_OHVSJI9Gd6gwsDnFSL-Xl |
haiku, sonnet |
with the cBioPortalChatBeta / cBioPortalChatBetaSonnet spec |
beta-router |
handoff router agent_cbiobeta_router |
router |
no spec: the router and its specialists run their own models |
beta-unified |
unified beta agent | unified |
no spec: the agent's own model |
prod |
prod agent agent_9ZXhcwLIsROBQX0u4JS5F |
haiku, sonnet |
with the prod specs |
Once beta's cBioPortalChatBeta spec points at the router, --target beta returns 400 (the spec no longer
selects a model for the unified agent) and beta-unified is the single-agent baseline. router and unified
stand for whatever models the agents call, so their answers are priced per LLM call from the Langfuse trace.
Each trace records every LLM call (model, agent, start/end, tokens, cost), the handoffs (lc_transfer_to_*),
the agent that answered (routed_to), and tool rounds; summary.json has the routing distribution, p90
latency, LLM calls and tool rounds per answer. run.json records the target agent and every agent it hands
off to (prompt hash, model, last update) and the LibreChat and MCP image tags, where kubectl can read them.
# B against baseline A, per question and per category, repeats pooled (A's model must be named if it ran several)
uv run cbioportal-mcp-qa compare 20260923-1919 20260927-1200 --model-a haikuWrites results/compare/<A>-<model>_vs_<B>-<model>/compare.{html,md,json} and prints the markdown: the
headline metrics below, the same by category and track (with precision, recall, p90, share under 10s and LLM
calls per answer), and every question with its outcome per repeat, pass variance and latency spread,
regressions first. Only questions both runs asked are compared. Runs recorded before per-call traces show
"–" for tool rounds, handoffs and routing.
Comparable runs only. compare refuses, and says why, when the runs used a different judge model (or
either run mixes judge models) or questions file, or when a question's text, conversation history, track,
expected_answer, expected_links or notes differ between the runs (or between one run's repeats) — their
pass/fail would measure different things. Each grade records the judge model that made it and a snapshot of
the question it was graded against, and compare checks those; grades from before per-answer judges fall back
to the run's judge_model. Regrade the older run against the current references (grade <run> --refresh-questions) and compare again, or pass --allow-mismatch to compare anyway: the affected questions
are then marked ⚠ and a warning is printed.
--refresh-questions never changes the question text or history, since the answer was to what was asked. So
even after regrading 20260923-1919, the 9 questions reworded since then still differ from runs that asked the
new wording and can only be compared with --allow-mismatch. For a clean beta comparison, record a fresh
3-repeat baseline (run --target beta-unified --repeats 3) on the current questions file rather than relying on
the regraded 09-23 run.
Every turn is counted. A turn is one question × repeat. Each side reports its expected turns (questions ×
the run's repeats), completed (HTTP 200), failed (an HTTP error or timeout), missing (never recorded,
e.g. an interrupted run), ungraded (completed, has a reference, no grade yet) and no reference turns, in
the headline, per category and per question; an incomplete side gets a warning. Eligible turns are the
graded, failed and missing turns of questions with a reference: a failed or missing turn counts as not passed.
Ungraded turns are left out until graded.
| Metric | Definition |
|---|---|
| Recall | passes / eligible turns — the headline score; failed and missing turns count against it |
| Precision | passes / attempted answers (pass + fail; declines are not attempts) |
| Attempt rate | attempted answers / eligible turns (declines, failures and missing turns are not attempts); recall = precision × attempt rate |
| Pass rate | passes / graded answers — the per-run report's definition, which leaves failed turns out; shown for continuity |
| Latency | median and p90 over completed turns, and again over completed + failed turns with each failure at its elapsed time |
| Under 10s | completed turns under 10s, over completed turns and over completed + failed turns (a failure is never fast) |
The per-run report's coverage is different from the attempt rate: it is the share of graded answers that
weren't declines, so failed requests don't lower it. p90 everywhere (reports, summary.json, compare) is
the upper nearest-rank value, sorted(values)[⌊0.9·n⌋] (clamped to the last value): with 10 or fewer values
it is the maximum, so small samples err high.
Output goes to results/<run-id>/: run.json (every answer, trace and grade), report.html, and
summary.json, plus results/index.html listing all runs. Commit the run directory to publish it.
Keep --concurrency low (default 2, at most ~3): the beta pod is small and shared with real users. A full run
(146 questions × Haiku + Sonnet) takes about 1.5 hours and 35M LibreChat credits ($35 at list prices).
For iterating on prompts, guides and views, --runner claude-code answers each question with headless
Claude Code (claude -p) instead of the deployed agent:
- System prompt: the deployed agent's instructions, read from the cBioAgent MongoDB via
kubectlat the start of the run (--target beta→ the beta agent). Only its hash is stored inrun.json. - Tools: only the two MCP servers (the database MCP and the navigator); Claude Code's built-in tools are disabled and extended thinking is off, matching the deployment.
- Models: the same Haiku 4.5 / Sonnet 5.
- Cost: runs on the Claude subscription of the Claude home it's started with, so point
CLAUDE_CONFIG_DIRat the Claude home you want billed (default~/.claude). Only the judge bills Bedrock, and with--judge-runner claude-codeit doesn't either. If the subscription's usage limit is hit (You've hit your limit,hit your session / weekly limit,usage limit reached), or the connector's login expires, the run stops asking, skips grading, and prints the--resumecommand to continue once the limit resets.
A stop (the usage limit, per-token billing, a plugin, an expired connector login) ends the batch at once:
- answers already in flight finish and are saved as usual (an ordinary failure among them too);
- no queued answer starts a
claudesession, and none gets a record inrun.json; - a reply that is itself a stop isn't recorded either, judged by its own content: the limit reply
(
You've hit your session limit), and a reply in flight that hit a different stop (say a plugin) after it. The first stop is the reason given; any others are listed after it.
So run.json has no junk failures, and run --resume <run> asks exactly the answers that are missing or failed.
The printed command keeps every option the run was started with that isn't a default (--questions,
--no-grade, --questions-file, judge options, --concurrency, --wait-on-limit, the --claude-code-*
opt-ins, …). --resume also restores the run's own options from run.json (options), except the
--claude-code-* and --judge-allow-managed-customizations opt-ins: those guards must be given again.
run.json records the questions the run means to ask (planned_questions, added to when a resume selects
more). Every resume asks the planned answers that are still missing or failed (anything but HTTP 200, the rule
--resume always used) as well as its own selection, so a narrower --questions on one resume never strands
them; a planned question since removed from the questions file is asked as it was planned. A run never exits 0
while a planned answer is missing or failed: it still grades and writes the report, then exits 1 naming the
questions and printing the --resume command.
The report counts a planned turn with no record as a failed request ("not asked"), and compare counts it as
missing, so a stopped run shows as incomplete exactly as it would with failure records. compare checks each
run against its own plan, so it warns about an incomplete run even when the missing questions are ones the
other run didn't ask (and so aren't scored). Runs from before planned_questions are reported from their
records, as before.
Grading works the same way (see
Grading with Claude Code): in-flight grades are saved,
nothing ungraded gets a judge_error from a limit, and no judge call starts once a stop is recorded (the stop
check and the launch are one locked step).
run and grade take --wait-on-limit to wait the usage limit out in the same process instead of stopping:
uv run cbioportal-mcp-qa run --runner claude-code --wait-on-limit --max-wait 300
uv run cbioportal-mcp-qa grade 20261002-0306 --judge-runner claude-code --wait-on-limit- When. The reset time is read from the message:
resets 9:50pm (Pacific/Honolulu),resets 3:40am (UTC),resets 3pm(no zone: the machine's local time). It is the next such time (tomorrow's if today's has passed), plus a minute. The wait is elapsed time, so it is right across a DST change; a time that happens twice (the fall-back hour) is taken as the later one. A message with no time of day (resets Oct 3,resets Mon), an unknown zone, or a time that passed less than an hour ago (the limit is about to clear) waits a fixed 15 minutes instead. - How long.
--max-wait(minutes, default 360) caps the waiting in all, across answering and grading; a wait that would go past it stops as without the flag, with the--resumecommand. - It prints what it waits for and until when (
waiting 51 min, until 2026-10-01 21:51 HST); Ctrl-C stops it, and--resumecontinues later. Only the usage limit is waited out: the other stops need you.
uv run cbioportal-mcp-qa run --runner claude-code --questions 1-20
uv run cbioportal-mcp-qa ask "what is the median age in os target gdc" --runner claude-codeThe prompt the run tested is saved as results/<run>/agent-prompt.md (a record; runs always read the live
agent) and named in the report header with the agent, its hash and when the agent was last updated. Each
answer's transcript (tool calls with their SQL and results, then the answer) is saved under
results/<run>/transcripts/ and linked from the report, in place of the Langfuse trace link. Costs in
claude-code reports are list-price equivalents; nothing is billed per token.
Scores are close to, not identical with, the deployed agent (different harness: no LibreChat recursion limit or eager tool execution). Compare claude-code runs with each other; confirm on beta with the Agents API runner before changing prod. Reports and the results index label the runner.
In claude -p, anything below takes precedence over the subscription login, in this order
(authentication docs): a cloud provider
(CLAUDE_CODE_USE_BEDROCK / _VERTEX / _FOUNDRY / ...), ANTHROPIC_AUTH_TOKEN, ANTHROPIC_API_KEY ("in
non-interactive mode (-p), the key is always used when present"), an apiKeyHelper, then a named Anthropic
profile or federation credentials. The guard fails closed. By default the runner:
-
removes
ANTHROPIC_API_KEY,ANTHROPIC_AUTH_TOKEN,ANTHROPIC_BASE_URL,ANTHROPIC_PROFILE,ANTHROPIC_FEDERATION_*,ANTHROPIC_IDENTITY_TOKEN*and everyCLAUDE_CODE_USE_*from the sessions' environment.CLAUDE_CODE_OAUTH_TOKEN, a subscription token, stays; -
before any model call (the connector probe included), reads every settings source the isolated sessions still load (managed settings):
managed-settings.jsonandmanaged-settings.d/*.jsonin/Library/Application Support/ClaudeCode/or/etc/claude-code/;- the macOS MDM profile (
/Library/Managed Preferences/[<user>/]com.anthropic.claudecode.plist); - the server-managed settings cache
$CLAUDE_CONFIG_DIR/remote-settings.json.
It refuses to start if any of them sets an
apiKeyHelper, apolicyHelper(its output can't be checked), or one of those variables inenv, or can't be read; -
refuses to start unless
claude auth statusconfirms a subscription login. That means a claude.ai login or an OAuth token that reports itssubscriptionType, with noANTHROPIC_AUTH_TOKENin any settings source.auth statusreports a bearer token and a subscription token alike asoauth_token. It also reads the user settings the isolated sessions skip, so anapiKeyHelperor bearer token there refuses too; -
refuses to start on a plan that can receive server-managed settings: only Claude for Teams and Enterprise can (server-managed settings). A
claude -psession fetches and applies its organization's policy without caching it, so anANTHROPIC_AUTH_TOKENor API key the policy sets can't be checked beforehand. OnlysubscriptionTypeproandmaxstart by default. Team, Enterprise, and an unrecognised or missing plan need--claude-code-trust-org-policy(orCLAUDE_CODE_TRUST_ORG_POLICY=1), which trusts the organization's Claude Code policy not to route sessions to per-token billing. A login in an MSK claude.ai organization reportsteam, so it needs this flag (claude auth statusshows your plan assubscriptionType). It doesn't relax the other checks; -
as a backstop, stops the run if a session's
apiKeySourcenames a key, token, helper or bearer. A bearer token reportsnonethere, like the subscription, which is why the checks above run first.
--claude-code-allow-api-billing (or CLAUDE_CODE_ALLOW_API_BILLING=1) turns all of these checks off. run.json
records claude_code.auth_mode (subscription, or with the opt-in api-key, cloud-provider, unconfirmed),
claude_code.account_type (the plan), trust_org_policy and allow_api_billing, the claude auth status
method, provider and whether the login belongs to an organization (no email, org id or name), and the names of
the removed variables.
A CLAUDE_CODE_OAUTH_TOKEN whose auth status doesn't show a subscription type is refused; use /login or
the opt-ins.
Sessions run with --setting-sources "": the Claude home's user settings (and project/local ones) aren't
loaded, so effortLevel, hooks, enabled plugins and an apiKeyHelper in ~/.claude/settings.json stay out of
the benchmark. Managed settings still apply. The login and the claude.ai connectors aren't settings, so the
connector path keeps working; if the connector ever fails to load, the runner stops with an error rather
than answering without the database. --claude-code-user-settings loads the user settings again.
claude 2.1.287 ships plugins inside the binary (<name>@builtin in a session's init event). They load in every
session, even with --setting-sources "" and --safe-mode, and they are not the claude.ai org-synced plugins
(~/.claude/plugins/synced/ holds only a marketplace index). On a Pro login a claude -p session loads
cc-plugin-agents-md (loads AGENTS.md as project instructions), cc-plugin-telemetry (lets plugins log
analytics events) and cc-plugin-plugin-authoring (a skill on writing plugins). On a Team or Enterprise login,
or on a machine with managed settings, it also seats cc-plugin-sec-default outermost. That plugin keeps the
organization's hooks, prompt content and tool policy out of reach of user plugins and adds no policy of its own.
Only managed prependPlugins can unseat it. --bare would skip plugins but never reads the subscription login.
So both the runner and the claude-code judge:
- pass
--settings '{"enabledPlugins": {"<id>@builtin": false, ...}}'for every built-in plugin, which turns all of them off except a seatedcc-plugin-sec-default; - run a preflight before the first model call (the connector probe included). This is a
claude -pon a model that doesn't exist, so the session prints its init event and then fails withmodel_not_foundwithout generating or billing anything. The preflight must prove that: its result has to be themodel_not_founderror for that model, with zero tokens, no model usage andtotal_cost_usd0 (or absent). Anything else refuses to start. A plugin it still finds is added toenabledPlugins: falseand checked again; - refuse to start if a plugin still loads, unless
--claude-code-allow-pluginsis passed (recorded inrun.jsonasallow_plugins, withpluginsanddisabled_plugins, underclaude_codeandclaude_code_judge); - record the plugins each session loaded:
trace.pluginson every answer andjudge_pluginson every grade. A grade made without a judge session (empty answer, no reference) hasjudge_plugins: nulland ajudge_plugins_notesaying so. A session that loads a plugin the preflight didn't find stops the run, the connector probe or the grading. So does an answering or judging session that doesn't report its plugins (no init event, or one without a plugin list), even with--claude-code-allow-plugins. A stopped run prints therun --resumecommand; the answer it stopped on, and any not yet asked, aren't recorded, so resume asks them.
run and grade take --judge-runner claude-code to grade on the local Claude subscription instead of
Bedrock, so a benchmark costs no Bedrock credit. To grade an existing run (e.g. one collected with
--no-grade) locally:
uv run cbioportal-mcp-qa grade <run-id> --judge-runner claude-code --claude-code-trust-org-policy(--claude-code-trust-org-policy is needed on a Team/Enterprise login, e.g. an MSK claude.ai organization; a
Pro or Max login doesn't need it.) --concurrency N grades N answers in parallel (default 1); run takes
--judge-concurrency. --judge-model picks the judge for either runner: a model key (sonnet-4.6), a Bedrock
id or a Claude Code id. By default the claude-code judge is the Bedrock judge's model (JUDGE_MODEL, Sonnet
4.6) through Claude Code, i.e. claude-sonnet-4-6.
-
Same prompt and grade. Each answer goes to
claude -pwith the Bedrock judge's prompt and rubric on stdin, the same JSON schema as structured output (--json-schema), and produces the same grade fields. -
No tools. Built-in tools are off (
--tools ""), and so are MCP servers:--strict-mcp-configwith an empty config, andENABLE_CLAUDEAI_MCP_SERVERS=falsefor the claude.ai connectors. Skills are off (--disable-slash-commands), as are auto memory (CLAUDE_CODE_DISABLE_AUTO_MEMORY=1) and the hooks of every settings source a session can override (--settings '{"disableAllHooks": true, "autoMemoryEnabled": false}'). -
Same billing safeguards as the runner (above): billing variables are removed, the billing guard and plan check run before the first grading call, and a session whose
apiKeySourcenames a key or token stops grading.--claude-code-trust-org-policyand--claude-code-allow-api-billingwork as for the runner. User settings are never loaded (--setting-sources ""):--claude-code-user-settingsonly applies to the runner (gradedoesn't take it). It runs in an empty temp directory. -
Managed customizations refuse the judge. Managed (policy) settings still apply with
--setting-sources "", and a session can't turn them off: managed hooks (aSessionStarthook can add context, an agent hook can run tools that aren't in the session's tool list) run even withdisableAllHooksset outside managed settings; a managedCLAUDE.mdcan't be excluded; managed plugins and MCP servers load.--barewould skip them, but bare mode never reads the subscription login. So before the first call, the judge reads every managed source the runner's billing guard reads:- the managed settings files and drop-ins (
managed-settings.json,managed-settings.d/*.jsonin/Library/Application Support/ClaudeCode/or/etc/claude-code/); - the MDM profile (
/Library/Managed Preferences/[<user>/]com.anthropic.claudecode.plist); - the cached server-managed settings (
$CLAUDE_CONFIG_DIR/remote-settings.json); - the managed
CLAUDE.md,managed-mcp.jsonand.claude/in those directories.
It refuses to grade if any of them sets
hooksorallowManagedHooksOnly, unless managed settings setdisableAllHooks: true. It also refuses onclaudeMd,enabledPlugins,mcpServers,agentoroutputStyle, on any of those managed files, and on a file it can't read.--judge-allow-managed-customizationsgrades anyway: it prints what the sessions inherit and records it underclaude_code_judgeinrun.json, with the judge's auth mode, plan and opt-ins. As a backstop, grading stops if a session's stream shows a hook running (hook_started/hook_response), an MCP server, or any tool besidesStructuredOutput. Plugins are handled as described in Built-in plugins. Residual risk: as for the runner, server-managed settings that a Team or Enterprise organization delivers when aclaude -psession starts aren't cached, so they can't be read beforehand.--claude-code-trust-org-policytrusts that policy not to add hooks or instructions to the judge, as it trusts it not to bill per token; the stream backstop catches hooks that run at session start, plugins and servers, but not a remotely deliveredclaudeMd. - the managed settings files and drop-ins (
-
Failures. A reply that isn't valid JSON or doesn't match the schema is retried once; if it fails again the answer is left ungraded (
judge_errorinrun.jsonsays why) and the nextgradetries it again. The subscription's usage limit stops grading on the call that hit it, with no retry. It is recognized by the same pattern the runner uses:You've hit your limit,hit your session / weekly / 5-hour / … limit,usage limit reached, or a rate limit that names its reset. Grades so far are saved, the report is written, and thegradecommand to resume once the limit resets is printed. That command keeps the run as you named it (id or path) and the judge options you used (--judge-model,--concurrency,--claude-code-trust-org-policy,--claude-code-allow-api-billing,--judge-allow-managed-customizations,--claude-code-allow-plugins). With--wait-on-limitit waits for the reset instead and carries on (above). -
Judge identity. Each grade records its judge as
claude-code:<model>(e.g.claude-code:claude-sonnet-4-6), socomparerefuses to mix them with Bedrock grades (us.anthropic.…), and a run graded by both shows up as mixed. Judge cost isn't estimated: nothing is billed per token.
Less deterministic than Bedrock. Claude Code can't set the temperature, so the claude-code judge can't
grade at temperature 0 like the Bedrock judge, and regrading the same answer can flip a borderline verdict.
Grade the before and after runs of a comparison with the same judge (both claude-code:<model>, or both
Bedrock), and expect a little more noise than with Bedrock. Claude Code grades are not comparable with
the Bedrock-graded 09-23 run (20260923-1919) or any other Bedrock-graded run; regrade that run with
--judge-runner claude-code --regrade if you need it as a baseline.
- Navigator: its public endpoint
https://mcp.cbioportal.org/navigator/mcp(NAVIGATOR_MCP_URL; no login needed). Beta and prod share the navigator deployment. - Database: by default, the claude.ai connector for its public endpoint
https://mcp.cbioportal.org/db/mcp(DATABASE_CONNECTOR_URL), which needs an OAuth login only the connector holds. The runner finds the connector by that URL inclaude mcp list, whatever you named it (orCLAUDE_AI_DATABASE_CONNECTORnames it), and hides every other claude.ai connector from the model. Add the connector in claude.ai once.
That public endpoint is prod's database MCP (cbioagent-clickhouse-mcp, cbioportal/mcp:latest). Beta
runs its own, cbioagent-clickhouse-mcp-beta (cbioportal/mcp:beta on beta's ClickHouse buffers; see
knowledgesystems-k8s-deployment#658, which was still open at the time of writing), with no public endpoint.
So with a beta* target and no DATABASE_MCP_URL, answers come from prod's MCP: the runner prints a
warning and records database_mcp_env: prod in run.json. --require-beta-mcp makes that an error. To
benchmark beta, point DATABASE_MCP_URL at beta's MCP:
# Service and path from k8s-deployment#658 (to be confirmed once it's merged and synced)
kubectl port-forward svc/cbioagent-clickhouse-mcp-beta 18081:80 &
export DATABASE_MCP_URL=http://localhost:18081/db/mcp DATABASE_MCP_ENV=beta
uv run cbioportal-mcp-qa run --runner claude-code --target beta --require-beta-mcp --questions 1-20A port-forward's URL doesn't say what's behind it, so declare it with DATABASE_MCP_ENV (beta, prod or
local); otherwise it's recorded as unknown and --require-beta-mcp refuses it. A declaration that
contradicts the URL's host (e.g. DATABASE_MCP_ENV=beta with the connector or mcp.cbioportal.org, or with
prod's service name) is recorded as conflict: that prints a warning on any target and fails
--require-beta-mcp. With a port-forward,
dropped connections show up as tool errors such as ECONNRESET that the deployed agent wouldn't have hit.
Prod's service works the same way (svc/cbioagent-clickhouse-mcp). run.json's versions records the
deployments it read for the target (versions.deployments) and their image tags and digests.
To test an unmerged cbioportal-mcp branch, or the cbioportal/mcp:beta image, run it locally and point
DATABASE_MCP_URL at it (DATABASE_MCP_ENV=beta for the beta image): docker run --rm -p 18080:8000 --env-file <clickhouse.env> -e CLICKHOUSE_MCP_SERVER_TRANSPORT=http -e CLICKHOUSE_MCP_BIND_HOST=0.0.0.0 -e CLICKHOUSE_MCP_BIND_PORT=8000 <image> (URL http://localhost:18080/mcp). For the same data as beta, its
env file must name beta's ClickHouse database (kubectl get configmap clickhouse-mcp-active-beta -o jsonpath='{.data.CLICKHOUSE_DATABASE}').
Beta's LibreChat puts the database MCP's instructions (from its initialize response) into the agent's
system prompt after the agent's own (serverInstructions: true, k8s-deployment#654). Claude Code already
sends every MCP server's instructions, with --system-prompt too: as a "MCP Server Instructions"
system-reminder at the start of the first user turn, after the system prompt. This was checked against the
request Claude Code 2.1.283 sends, captured by a local stand-in for the API. That's the same order, so the
runner doesn't add them again. For a database MCP reached by URL, run.json's claude_code.prompt records
hashes of the agent instructions, the server instructions and both joined in LibreChat's order (combined).
Through the connector, the server instructions can't be read without its OAuth login, so only the agent
instructions are hashed.
Every run also records versions (both runners): cBioPortal portal / DB schema / gene table versions from
https://www.cbioportal.org/api/info, and the cbioportal-mcp and navigator server versions (with a hash of
the server's instructions) and image digests.
Append to input/questions.yaml with the next unused id (ids are stable; never renumber or reuse):
- id: 147
track: data
type: Clinical Data
study: msk_chord_2024
question: How many patients in MSK-CHORD received immunotherapy?
expected_answer: "4,721" # or leave "" and give links / a rubric
expected_links: []
notes: |
A correct answer must use the treatment table; must not count planned treatments.
checked: 2026-09-23 # when the reference was last verified
source: https://github.com/cBioPortal/cbioportal-mcp/issues/24Where things go:
expected_answer: what a correct answer says — the facts, numbers or conclusion. A single bare number (e.g.548) is also checked automatically against the answer.notes: how to grade —A correct answer must: …/Must not: …. Facts the judge needs but the answer doesn't have to state (e.g. survival statistics the agent can't compute) go here too, labelled as context.expected_links: links that open the right view. Don't pin session-based group comparison links (comparisonId=…); they differ on every run.
Questions from user-feedback issues set source to the issue URL and usually carry a rubric in notes.
Add technical: true when the question asks for code, the schema, or how the agent works: answers are then
expected to be technical and are exempt from the "exposes internals" check.
input/questions-multiturn.yaml is a separate set of follow-up questions modeled on real conversations in
Langfuse (paraphrased; trace_id names the conversation it's based on). Its pass rates are kept apart from
the main 146 questions so those stay comparable across runs:
uv run cbioportal-mcp-qa run --questions-file input/questions-multiturn.yamlA question with history is the user's next message in that conversation; the earlier turns alternate
user / assistant and end with an assistant turn:
- id: 1003
question: And in lung squamous?
history:
- role: user
content: What are the most common KRAS mutations in TCGA lung adenocarcinoma?
- role: assistant
content: In Lung Adenocarcinoma (TCGA, PanCancer Atlas), 168 of 566 profiled samples (29.7%) ...The Agents API runner sends the history as prior messages. claude -p takes a single message, so the
claude-code runner quotes the history ahead of the new message. The judge sees the whole conversation.
Assistant turns are written by hand, not generated, and must be correct: the follow-up is graded, not them.
input/questions-subset.yaml (ids 2001+) asks for statistics over a subset of a study: a clinical subgroup
(sample type, sex, smoking, stage, MSI, HR/HER2, OncoTree code), a pooled set of studies, or a mutation-defined
group. The main set leans on per-study statistics, which precomputed tables and per-study tools answer directly;
these questions measure whether the agent writes the right SQL instead, and doesn't answer with the
whole-study number. Like the multi-turn set, it is its own test set with its own pass rates:
uv run cbioportal-mcp-qa run --runner claude-code --questions-file input/questions-subset.yamlEach question's sql field names the query in input/subset-sql/ that computed its reference, with the
database and date (see its README).
uv run pytest
uv run ruff check src tests && uv run ruff format src testsThe benchmark tests the agent as deployed: its system prompt lives in the agent's instructions in the cBioAgent MongoDB and the LibreChat config, not in this repo.