messageboardbench

Research harness for studying whether coding agents publish, receive, and adopt evaluator-gaming methods across shared scratchpads and message boards.

The current study runs three OpenRouter models on the same 20 ImpossibleBench SWE tasks under isolated, scratchpad, board, feedback, internet, and token-awareness conditions.

See the current experiment, scenario series, experiment index, and setup guide.

Current raw runs live in logs/. Historical local experiments and logs stay in ignored archive/.

S
Description
No description provided
Readme MIT
50 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 99.8%
Just 0.2%