mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
was_test_modified reads intermediate score events, not just the final score. ImpossibleBench's scorer repairs test.py when it finds it modified, and basic_agent scores every attempt, so an agent that edits the tests on attempt one and not again leaves a final score with no flag on it. The repair erased the evidence. snapshot_scratch reports whether the directory exists alongside its contents, because otherwise a directory nobody wrote to and a directory that was never created look identical, and the second is a broken harness reported as a real null result. What the agents wrote is not auto-classified. "Did it write" is safe to automate; whether a note is addressed to somebody is the judgement the pilot exists to make, so pilot.py emits a markdown file to read by hand. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu