Files
messageboardbench/README.md
T

620 B

messageboardbench

Research harness for studying whether coding agents publish, receive, and adopt evaluator-gaming methods across shared scratchpads and message boards.

The current study runs three OpenRouter models on the same 20 ImpossibleBench SWE tasks under isolated, scratchpad, board, feedback, internet, and token-awareness conditions.

See the current experiment, scenario series, experiment index, and setup guide.

Current raw runs live in logs/. Historical local experiments and logs stay in ignored archive/.