This is the agent-facing runbook for the Phase 1 shared quorum appliance. The
design lives in
docs/superpowers/specs/2026-06-18-shared-eval-appliance-design.md.
The installed evals-appliance helper described here is the target Phase 1
interface. Until that helper exists on a configured appliance, raw local
bun run quorum ... and scripts/evals-container exec quorum ... commands are
local or break-glass workflows only.
Install the host wrapper from the trusted evals checkout:
scripts/install-evals-appliance /srv/quorumThe installer writes /srv/quorum/bin/evals-appliance and prints that path. It
does not write appliance.json, create credentials, or mutate repositories. The
installed wrapper uses the embedded host config path, normally
/srv/quorum/config/appliance.json, verifies the evals checkout is clean and on
the configured branch, then exports that path to the repo-owned TypeScript CLI.
Direct local bun run appliance ... use can still set EVALS_APPLIANCE_CONFIG
when intentionally running outside the installed wrapper.
Agents operating shared live evals use the appliance helper, not raw quorum commands. The helper owns repo sync, ref resolution, blessed-bundle mounting, preflight, locks, job records, logs, provenance, and cancellation.
Live eval artifacts are sensitive. Do not paste raw transcripts, run homes,
tool-call logs, or credential-bearing files. Prefer status, show, costs,
and reviewed summaries.
Routine appliance operation uses the approved private access path documented in the private ops runbook. Provider-specific break-glass access is for recovery only when the normal private path is unavailable; if it is used, record why.
Do not add real hostnames, account identifiers, access-provider commands, or secret parameter names to this public runbook.
Start with a read-only health check. doctor must not fetch, checkout, build,
start containers, source credentials, remove locks, or mutate job records:
evals-appliance doctor --jsonPrepare the exact Superpowers ref to test:
evals-appliance prepare --json --superpowers-ref <branch-tag-or-sha>The helper must resolve mutable refs to exact SHAs. If prepare returns
lock_busy during an active live job, dirty checkout, ambiguous ref, stale
lock, missing credential bundle, or failed container preflight, stop and report
that result instead of guessing. prepare must not change refs underneath an
active live eval.
Phase 1 shared run-all is Linux-container-only. Windows evals and Antigravity
remain trusted-maintainer break-glass paths until the appliance explicitly
supports them.
Start with the sentinel tier and a narrow target set:
evals-appliance run-all --json --detach \
--superpowers-ref <branch-tag-or-sha> \
-- --tier sentinel \
--coding-agents claude,codex,kimi \
--jobs 4For fragile or single-column targets:
evals-appliance run-all --json --detach \
--superpowers-ref <branch-tag-or-sha> \
-- --tier sentinel \
--coding-agents gemini \
--jobs 1The first JSON response should contain a job_id. Record that id in your work
notes; it is the recovery handle if the SSH session drops.
Use a single-scenario run for a focused RED/GREEN check:
evals-appliance run --json --detach \
--superpowers-ref <branch-tag-or-sha> \
--scenario scenarios/<name> \
--coding-agent <agent>Use --detach by default unless you are deliberately doing a short foreground
smoke and can tolerate the shell owning the lifetime.
Recover or poll a job:
evals-appliance status --json <job-id>Inspect summarized results:
evals-appliance show --json <job-id>
evals-appliance show <job-id>
evals-appliance costs --json <job-id>
evals-appliance costs <job-id>If a batch id is known, the helper may accept it directly. Raw
scripts/evals-container exec quorum show/costs ... remains a local or
break-glass read path because the container's quorum shim may source the live
credential env.
The helper's status should distinguish appliance failure from eval failure. A completed batch with failing cells is a completed job with a failing summary, not an appliance crash.
Cancel through the job record:
evals-appliance cancel --json <job-id>The helper sends SIGINT to the tracked process group and waits for stopped
verdicts or a batch footer. If cancellation returns lost, do not retry a new
live job until doctor --json explains the lock and process state.
Runs produced on a workstation before the appliance existed are moved in two
steps: a local export that scrubs them, and an appliance-side import that
ingests the result. Never copy a raw local results/ tree to the shared box —
run homes contain live agent credentials.
Design:
docs/superpowers/specs/2026-08-09-appliance-results-import-design.md.
On the workstation that holds the runs:
bun run quorum export-runs <results-dir> \
--out <bundle-dir> \
--superpowers-repo <path-to-superpowers-checkout>This copies an allowlist — verdict.json, trajectory.json,
coding-agent-token-usage.json, phase.json, gauntlet-agent/,
coding-agent-workdir/, and the raw session logs lifted to raw-sessions/ —
and drops each run's throwaway $HOME wholesale. auth.json,
credentials.snapshot.yaml, and agent config files never enter the bundle.
--superpowers-repo lets the export resolve the skill tree a run archived back
to an exact commit. Without it, runs whose verdict lacks a rev degrade to
tree_only: the tree hash is still recorded, but no commit is named. It
defaults to $SUPERPOWERS_ROOT.
The summary line reports how each run's superpowers rev was established:
bundle written to /tmp/lane-b-bundle
exported 346, skipped 0
superpowers rev: recorded=196 recovered=128 inferred=22
recorded came from the verdict, recovered was matched exactly from the
archived tree, tree_only means the run used a modified tree that matches no
commit, inferred was borrowed from the nearest co-temporal run in the same
experiment directory and is stored in its own field, and unknown means no
evidence survived.
The bundle is meant to be audited before it crosses to a shared host. Confirm it carries no credentials:
find <bundle-dir> \( -name auth.json -o -name 'credentials.snapshot.yaml' \
-o -name 'config.toml' -o -name '.env*' -o -name '*.pem' -o -name '*.key' \) | headExpect no output. Then transfer the bundle over the approved private access path documented in the private ops runbook.
evals-appliance import --json <bundle-dir>Import verifies every checksum in the manifest and re-runs the credential
denylist against what is actually on disk before anything lands, so a tampered
or mis-built bundle is rejected whole rather than partially applied. It holds
run.lock for the duration; if a live job holds it, import returns lock_busy
and does nothing. Re-running is safe: runs already present are skipped, and
--force replaces them.
Imported runs are visible to the normal read commands:
evals-appliance status --json <run-id>
evals-appliance show <run-id>
evals-appliance costs <run-id>An imported job reports kind: "import" and carries an origin block instead
of refs and a credential bundle, because it was neither built from an
appliance-resolved ref nor run against the blessed bundle. Treat
origin.rev_recovery as the confidence marker: inferred_superpowers_sha is a
neighbour's sha, not evidence about the run itself.
The dashboard is read-only and must not submit or stop jobs:
bun run dashboard --results results --manifest results/grid-manifest.jsonOn the shared box, bind it only to loopback or the approved private network. Use the private ops runbook for operator access and forwarding details.
Raw commands are for local development or trusted-maintainer break-glass only:
scripts/evals-container exec quorum run-all ...
bun run quorum run ...Before using break-glass on the shared box, verify no live job holds run.lock,
record why the helper could not be used, and run doctor --json afterwards.