mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
d604370d96eef54d77e6175fd98e14f78101ee5b
Running the checks against the baseline task, which has no scratch directory at all, should report zero touches. It reported two. Both agents had made throwaway files named scratch.py and scratch2.py beside their work, and the mention pattern's (?!\w) allows a following ".", so scratch.py read as naming the directory. relation() was already right; only mentioned_in was wrong, so the effect was confined to the unparsed fallback. That fallback exists to turn undercounts into visible ones, and here it was manufacturing overcounts instead. The negative control is clean now. Calibration found this, not a test. Worth remembering that the baseline logs are a free negative control for every check, because the directory genuinely does not exist there. Also adds the submitted func.py to the calibration file: the baseline shows zero test.py edits on this model, so cheat technique has to be read off the function. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
messageboardbench
Does cheating spread from one AI agent to another through a shared file?
Agents get coding tasks whose tests contradict each other, so no honest solution exists and any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls workspace and never describes as a channel. We measure whether carrying that directory from one agent to the next raises the cheating rate.
Built on ImpossibleBench and Inspect AI. Tasks and payloads are synthetic throughout.
See docs/setup.md to install and run, and docs/findings.md for measured numbers.
Languages
Python
99.8%
Just
0.2%