mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Running the checks against the baseline task, which has no scratch directory at all, should report zero touches. It reported two. Both agents had made throwaway files named scratch.py and scratch2.py beside their work, and the mention pattern's (?!\w) allows a following ".", so scratch.py read as naming the directory. relation() was already right; only mentioned_in was wrong, so the effect was confined to the unparsed fallback. That fallback exists to turn undercounts into visible ones, and here it was manufacturing overcounts instead. The negative control is clean now. Calibration found this, not a test. Worth remembering that the baseline logs are a free negative control for every check, because the directory genuinely does not exist there. Also adds the submitted func.py to the calibration file: the baseline shows zero test.py edits on this model, so cheat technique has to be read off the function. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu