Scout is an open-source toolkit for designing and operating blinded, anti-Goodhart evaluations for autonomous coding agents. It keeps fast developer feedback in the target repository while scoring hidden cases in a separate, trusted environment.
Important
This public source repository contains only the framework and a synthetic
_example fixture. Real holdouts, canaries, target configuration, and score
histories belong in a separate private operational repository. Never add
real evaluation material to the public repository, even temporarily.
PUBLIC SOURCE PRIVATE OPERATIONAL HUB TARGET REPO
framework + synthetic fixture → real holdouts + scoring runtime ← requests by tag
results by commit status agent works here
The target asks for a holdout check with an immutable annotated
holdout-check-v1-<sha12>-<request-id> tag. The private hub checks out the pinned commit, runs
untrusted target code without network access or holdout visibility, compares
outputs outside the sandbox, and returns only a bounded status result.
The framework includes:
- a target registry and onboarding CLI;
- two-stage, sandboxed scoring orchestration;
- liveness, calibration, and audit activation gates;
- canary, mutation, divergence, and coverage-variance checks;
- terminal status, review, retrospective, and HTML dashboard tools;
- narrow onboard, design, audit, execute, and patch skills with shared science;
- a synthetic fixture that exercises healthy and adversarial signals.
Requirements: Python 3.9 or newer, Git, and Bash. Docker is strongly recommended for scoring untrusted target code. Development checks also use ShellCheck, Ruff, Mypy, and Coverage.
bin/lfd walkthrough
bin/lfd status --json
bin/lfd dashboard
bin/lfd doctor
bin/lfd test
python3 tools/ci_checks.py all
python3 tools/ci_checks.py public-release
python3 tools/check_standards.py allbin/lfd walkthrough executes the synthetic contract → bundle → scoring →
audit → activation lifecycle in a disposable copy, including the expected
activation refusal before independent judgment. See the
walkthrough guide for its evidence and limits.
bin/lfd status and bin/lfd dashboard also work immediately against the
synthetic targets/_example history.
Do not put real eval data in a public GitHub fork. Create a separate private repository from a source release instead:
- Copy a tagged source release into a new directory without its
.gitdirectory. - Initialize a new Git repository and create a private remote.
- Restrict read access to trusted humans. Evaluated agents and their credentials must never be able to read it.
- Run
bin/lfd onboard start <name> <repo-url>and follow the onboarding guide. - Configure the target access token and polling schedule in the private repository only.
Keep the public project as a read-only upstream for framework updates. Review every upstream change before applying it to the private hub, and never send private hub commits back upstream.
| Path | Purpose |
|---|---|
bin/lfd |
Command-line entry point |
tools/ |
Registry, audit, status, dashboard, and policy tools |
ops/ |
Trusted polling, sandboxing, scoring, and status runtime |
templates/target-repo/ |
Files copied into a repository under evaluation |
skills/lfd-*/ |
Narrow workflow skills plus non-invocable shared science and calculators |
targets/_example/ |
Synthetic demonstration data only |
rac/ |
Canonical requirements and architecture decisions |
standards/ |
Requirement-to-control mappings and tool version pins |
docs/ |
Executable walkthrough, architecture, onboarding, and release guidance |
The verified target bundle contains only lfd-execute. Equip writes bounded,
idempotent Codex, Claude, Cursor, and Copilot instruction blocks that point to
that single target-side skill while preserving caller-owned instructions.
Hub-side onboarding, design, audit, and patch skills never enter the target.
Agent-facing commands return a versioned JSON envelope with command, target,
status, stage, artifacts, stable errors, and next actions. Exit 0 means
success, 2 invalid input or contract, 3 pending judgment/external result,
and 4 infrastructure or transport failure.
Scout keeps its canonical standards beside the source. RAC validates the
requirements and accepted decisions, while tools/check_standards.py enforces
deterministic dependency, control-coverage, and CI-contract rules. See
standards/README.md for the control model and local
commands.
The open-source framework is public; an operational hub with real evaluation material is private. That separation is part of the design, not an optional deployment preference. Read SECURITY.md and the architecture guide before onboarding a target.
Contributions are welcome. Start with CONTRIBUTING.md.
Public pull requests must pass the public-release gate and therefore cannot
contain real target directories or generated evaluation artifacts.
Licensed under the Apache License 2.0.