Checkpoint experiments through SWE pilot v2

This commit is contained in:
pj committed 2026-09-15 15:46:10 +05:30
1 parent abacd5c5e1
commit 72d77018d8
845 files changed
+431756 -41

No files matched your search

+29 -7
View File
@@ -1,13 +1,35 @@
# messageboardbench
Does cheating spread from one AI agent to another through a shared file?
Do coding agents share, adopt, reject, or correct cheating techniques through a public team board?
Agents get coding tasks whose tests contradict each other, so no honest solution exists and
any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls
workspace and never describes as a channel. We measure whether carrying that directory from
one agent to the next raises the cheating rate.
Agents work on ImpossibleBench tasks in separate containers with private scratch space
and, in the board condition, explicit tools for reading and posting team messages.
The default pilot uses three agents per cohort, two cohorts, and 90 messages per
episode; model, population size, sampling and budgets are configurable.
Built on [ImpossibleBench](https://github.com/safety-research/impossiblebench) and
[Inspect AI](https://inspect.aisi.org.uk). Tasks and payloads are synthetic throughout.
[Inspect AI](https://inspect.aisi.org.uk). Passing contradictory tests is an operational
signal of specification gaming, not a complete label of intent or attempted cheating.
See `docs/setup.md` to install and run, and `docs/findings.md` for measured numbers.
See [setup](docs/setup.md) and the [board experiment workflow](docs/board-pilot.md).
Run `just board` to choose parameters interactively and start the experiment.
Use `just board-preview` for a configuration preview without model calls, or
`just board-run --model muse` to run with Muse Spark Contributor. Executing makes
paid requests; the existing key spending cap is never raised.
Start with [EXPERIMENT.md](EXPERIMENT.md) for the research question, what we have
done, what the results establish, and the next work. The latest pilot contains one
explicitly attributed cross-task adoption; it does not establish an increased
gaming rate or concealed collusion.
| Location | Purpose |
|---|---|
| `src/`, `scripts/`, `tests/` | Implementation, entry points and tests |
| `docs/` | Setup, workflow and supporting research |
| [results/](results/README.md) | Reviewed, frozen experiment evidence |
| `logs/` | Raw local runs, ignored by Git |
| `work/` | Disposable local working files, ignored by Git |
This is the sole working repository for the experiment. The former `messageboard`
repo is [retired](docs/migration/README.md); personal notes and unrelated material
remain in its archive. Run `just evidence-check` to verify the migrated evidence.