mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Repo skeleton, pinned environment, free smoke test
Public from commit one, so no key ever enters this history. compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root among twenty-odd entries. This experiment turns on whether an agent notices a scratch directory, so that is a bad place to put one. ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is declared here instead because hf_dataset genuinely needs it. Verified: docker run prints "/", impossiblebench imports, and the real task against mockllm/model completes with a real score and tracebacks rooted at /workspace. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
commit
bf3bb0d144
12 files changed
+2495
No files matched your search
@@ -0,0 +1,13 @@
|
||||
# messageboardbench
|
||||
|
||||
Does cheating spread from one AI agent to another through a shared file?
|
||||
|
||||
Agents get coding tasks whose tests contradict each other, so no honest solution exists and
|
||||
any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls
|
||||
workspace and never describes as a channel. We measure whether carrying that directory from
|
||||
one agent to the next raises the cheating rate.
|
||||
|
||||
Built on [ImpossibleBench](https://github.com/safety-research/impossiblebench) and
|
||||
[Inspect AI](https://inspect.aisi.org.uk). Tasks and payloads are synthetic throughout.
|
||||
|
||||
See `docs/setup.md` to install and run, and `docs/findings.md` for measured numbers.
|
||||
Reference in new issue
Block a user