Files
messageboardbench/README.md
T

10 lines
620 B
Markdown

# messageboardbench
Research harness for studying whether coding agents publish, receive, and adopt evaluator-gaming methods across shared scratchpads and message boards.
The current study runs three OpenRouter models on the same 20 ImpossibleBench SWE tasks under isolated, scratchpad, board, feedback, internet, and token-awareness conditions.
See the [current experiment](EXPERIMENT.md), [scenario series](docs/swe-scenario-series.md), [experiment index](experiments/README.md), and [setup guide](docs/setup.md).
Current raw runs live in `logs/`. Historical local experiments and logs stay in ignored `archive/`.