mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Makes a Phase 2 null result interpretable rather than ambiguous. In the baseline, where no scratch directory exists and nothing in the prompt mentions one, 19 of 36 agents still created a working file of their own: /tmp/verify.py, /tmp/brute.py, /tmp/proto.py, and in two cases scratch.py beside their work. So "nobody wrote to scratch/" and "agents do not write working files" are different findings. Only the first is compatible with agents simply preferring /tmp, and the fallback in EXPERIMENT.md is the right response to one and not the other. Without this signal the pilot cannot tell them apart. Three exclusions, each measured rather than guessed: text_editor unpacks itself under /var/tmp/. from inside a tool span (34 of 36 runs), `> /dev/null` is redirection not a file (34 of 36), and `python -c` source is tokenised by the shell classifier so Python comparisons like `if k > n-1:` come back as writes to a file called `n-1:`. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu