A small, self-contained example of designing a realistic software engineering task for evaluating coding agents — in the style of SWE-bench / Terminal-Bench — including a hidden grading harness and a worked example of red-teaming two candidate fixes.
This exists as a portfolio piece demonstrating the three things "coding RL environment" work actually involves: writing a realistic task, building a test-based rubric that resists gaming, and red-teaming candidate solutions against it.
repo/ the target codebase, with the bug present
ratelimiter/limiter.py a sliding-window rate limiter (buggy)
tests/test_limiter_public.py tests the agent is allowed to see/run
task/ISSUE.md the task prompt, written as a bug report
tests/test_limiter_hidden.py the hidden grading suite (withheld from the agent)
attempts/
attempt_a_rejected/ a plausible but wrong fix
attempt_b_accepted/ the correct fix
writeup/
RUBRIC.md grading criteria beyond pass/fail
FINDINGS.md graded results + why attempt A is the interesting case
Dockerfile pinned, reproducible environment
run_tests.py stdlib-only test runner (no pytest dependency)
SlidingWindowRateLimiter._prune keeps timestamps with t >= cutoff
instead of t > cutoff, so a request exactly window_seconds old is
kept one tick longer than it should be — legitimate requests get
rejected right at the window edge.
python3 run_tests.pyAgainst the unfixed repo/, this fails 2 of 8 tests (both hidden). Copy
in a fix and re-run to grade it.
The headline finding is in writeup/FINDINGS.md: a naive grader that
only checks whether the one named symptom goes away would have accepted
a patch that actually doubles the effective rate limit for every client.
Catching that requires PASS_TO_PASS tests that aren't mentioned anywhere
in the issue text — which is the actual design skill this kind of work
is testing for.
- A second, harder task in the same repo (e.g. a race condition across concurrent clients) to test whether the rubric design generalizes.
- Wiring
attempts/up to an actual agent harness (e.g. mini-SWE-agent or the Claude API with tool use) instead of hand-authored patches, to see whether a real model reproduces the Attempt-A failure mode unprompted.