Skip to content

About

SWE-bench-style coding task and grading harness for evaluating AI agent bug fixes

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Mini SWE-task: sliding-window rate limiter boundary bug

A small, self-contained example of designing a realistic software engineering task for evaluating coding agents — in the style of SWE-bench / Terminal-Bench — including a hidden grading harness and a worked example of red-teaming two candidate fixes.

This exists as a portfolio piece demonstrating the three things "coding RL environment" work actually involves: writing a realistic task, building a test-based rubric that resists gaming, and red-teaming candidate solutions against it.

Structure

repo/                     the target codebase, with the bug present
  ratelimiter/limiter.py  a sliding-window rate limiter (buggy)
  tests/test_limiter_public.py   tests the agent is allowed to see/run
task/ISSUE.md             the task prompt, written as a bug report
tests/test_limiter_hidden.py     the hidden grading suite (withheld from the agent)
attempts/
  attempt_a_rejected/     a plausible but wrong fix
  attempt_b_accepted/     the correct fix
writeup/
  RUBRIC.md               grading criteria beyond pass/fail
  FINDINGS.md             graded results + why attempt A is the interesting case
Dockerfile                pinned, reproducible environment
run_tests.py              stdlib-only test runner (no pytest dependency)

The bug

SlidingWindowRateLimiter._prune keeps timestamps with t >= cutoff instead of t > cutoff, so a request exactly window_seconds old is kept one tick longer than it should be — legitimate requests get rejected right at the window edge.

Running it

python3 run_tests.py

Against the unfixed repo/, this fails 2 of 8 tests (both hidden). Copy in a fix and re-run to grade it.

The interesting part

The headline finding is in writeup/FINDINGS.md: a naive grader that only checks whether the one named symptom goes away would have accepted a patch that actually doubles the effective rate limit for every client. Catching that requires PASS_TO_PASS tests that aren't mentioned anywhere in the issue text — which is the actual design skill this kind of work is testing for.

What I'd extend next

  • A second, harder task in the same repo (e.g. a race condition across concurrent clients) to test whether the rubric design generalizes.
  • Wiring attempts/ up to an actual agent harness (e.g. mini-SWE-agent or the Claude API with tool use) instead of hand-authored patches, to see whether a real model reproduces the Attempt-A failure mode unprompted.

About

SWE-bench-style coding task and grading harness for evaluating AI agent bug fixes

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages