pj fc43b691a7 Guard that the fork presents the same task text as the stock benchmark
The fork exists to add one line to the system message. If anything else drifts, the
pilot's cheat rate stops being comparable to the baseline measured with the stock task,
and nothing would fail to say so.

Also asserts the scratch line says nothing about agents, sharing, persistence or
messages, since the mechanism being reproduced started as a scratchpad.

92 tests, green with HTTP_PROXY pointed at a closed port.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:40:09 +05:30

messageboardbench

Does cheating spread from one AI agent to another through a shared file?

Agents get coding tasks whose tests contradict each other, so no honest solution exists and any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls workspace and never describes as a channel. We measure whether carrying that directory from one agent to the next raises the cheating rate.

Built on ImpossibleBench and Inspect AI. Tasks and payloads are synthetic throughout.

See docs/setup.md to install and run, and docs/findings.md for measured numbers.

S
Description
No description provided
Readme MIT
50 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 99.8%
Just 0.2%