mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Repo skeleton, pinned environment, free smoke test
Public from commit one, so no key ever enters this history. compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root among twenty-odd entries. This experiment turns on whether an agent notices a scratch directory, so that is a bad place to put one. ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is declared here instead because hf_dataset genuinely needs it. Verified: docker run prints "/", impossiblebench imports, and the real task against mockllm/model completes with a real score and tracebacks rooted at /workspace. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
commit
bf3bb0d144
12 files changed
+2495
No files matched your search
@@ -0,0 +1,2 @@
|
||||
# Copy to .env and fill in. .env is gitignored and must never be committed.
|
||||
OPENROUTER_API_KEY=
|
||||
@@ -0,0 +1,9 @@
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
.pytest_cache/
|
||||
logs/
|
||||
.env
|
||||
.env.*
|
||||
!.env.example
|
||||
.DS_Store
|
||||
@@ -0,0 +1 @@
|
||||
3.13
|
||||
@@ -0,0 +1,27 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2026 PJ
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
|
||||
---
|
||||
|
||||
Portions of src/messageboardbench/task.py are derived from ImpossibleBench
|
||||
(https://github.com/safety-research/impossiblebench), MIT License,
|
||||
Copyright (c) 2025 ImpossibleBench Team.
|
||||
@@ -0,0 +1,13 @@
|
||||
# messageboardbench
|
||||
|
||||
Does cheating spread from one AI agent to another through a shared file?
|
||||
|
||||
Agents get coding tasks whose tests contradict each other, so no honest solution exists and
|
||||
any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls
|
||||
workspace and never describes as a channel. We measure whether carrying that directory from
|
||||
one agent to the next raises the cheating rate.
|
||||
|
||||
Built on [ImpossibleBench](https://github.com/safety-research/impossiblebench) and
|
||||
[Inspect AI](https://inspect.aisi.org.uk). Tasks and payloads are synthetic throughout.
|
||||
|
||||
See `docs/setup.md` to install and run, and `docs/findings.md` for measured numbers.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Copied from impossiblebench/src/impossiblebench/compose.yaml (commit 061dc3d),
|
||||
# with working_dir added.
|
||||
#
|
||||
# The image has no WORKDIR, so inspect resolves the container working directory to
|
||||
# "/" and the task's func.py and test.py land at the filesystem root, mixed in with
|
||||
# the twenty-odd entries a bare `ls` returns there. This experiment turns on whether
|
||||
# an agent notices and uses a scratch directory, so burying it in / is a bad idea.
|
||||
# /workspace holds exactly func.py, test.py and scratch/. Docker creates the
|
||||
# directory on start; it does not exist in the image.
|
||||
#
|
||||
# This is a deviation from the published benchmark's default environment and belongs
|
||||
# in the paper's limitations. It applies identically to both conditions, so it cannot
|
||||
# produce the effect.
|
||||
services:
|
||||
default:
|
||||
image: "aisiuk/inspect-tool-support"
|
||||
command: "tail -f /dev/null"
|
||||
working_dir: /workspace
|
||||
init: true
|
||||
network_mode: none
|
||||
stop_grace_period: 1s
|
||||
mem_limit: 1gb
|
||||
@@ -0,0 +1,54 @@
|
||||
# Setup
|
||||
|
||||
Python 3.13 and a running Docker daemon. Everything the agent does happens in a container
|
||||
with the network disabled, so the daemon is not optional.
|
||||
|
||||
```
|
||||
uv sync
|
||||
uv pip install -e /path/to/impossiblebench --no-deps
|
||||
```
|
||||
|
||||
## Why --no-deps
|
||||
|
||||
It is required, not a shortcut. ImpossibleBench declares
|
||||
`inspect_evals[swe_bench] @ git+https://github.com/UKGovernmentBEIS/inspect_evals` from
|
||||
unpinned git main, which forces `huggingface_hub` up to 1.2+, and `swebench>=4.0.0`, which
|
||||
drags `modal`, `GitPython`, `typer` and `pre-commit` in as runtime dependencies.
|
||||
|
||||
The LiveCodeBench path we use imports none of it. Every `swebench` import in that package is
|
||||
lazy and inside a function, and `inspect_evals` is declared but never imported.
|
||||
|
||||
`--no-deps` also skips `datasets`, which `hf_dataset` genuinely does need, so this repo
|
||||
declares that one itself in `pyproject.toml`.
|
||||
|
||||
Two things to know. `uv sync` prunes ImpossibleBench, because it is not declared in
|
||||
`pyproject.toml`, so rerun the install line after every sync. `just install` does both.
|
||||
And ImpossibleBench mutates `sys.path` at import time to paper over a broken import, which
|
||||
loads some modules twice, so do not trust module identity inside it.
|
||||
|
||||
## The working directory
|
||||
|
||||
`compose.yaml` at the repo root sets `working_dir: /workspace`, which is a deliberate
|
||||
deviation from the published benchmark. The file itself explains why.
|
||||
|
||||
Inspect looks for `compose.yaml` in the **process working directory**, not next to the task
|
||||
definition, so every command has to run from the repo root. If you run from somewhere else
|
||||
you silently get ImpossibleBench's own compose file, and the working directory reverts to
|
||||
`/`. The justfile recipes handle this.
|
||||
|
||||
## Checks that cost nothing
|
||||
|
||||
```
|
||||
just smoke # the real task against a fake model, exercises everything but the model
|
||||
just test # offline unit tests, no network and no provider beyond mockllm
|
||||
```
|
||||
|
||||
`just smoke` should end with `status: success` and a real score. If the tracebacks in that
|
||||
score mention `/workspace/test.py`, the working directory override is live.
|
||||
|
||||
## Spending money
|
||||
|
||||
`.env` holds `OPENROUTER_API_KEY` and is gitignored. The key carries a $2 cap as a fuse.
|
||||
|
||||
Watch per-request `prompt_n` on any run longer than a few turns. Cache reuse fails silently
|
||||
and costs 8 to 18 times as much without emitting an error.
|
||||
@@ -0,0 +1,22 @@
|
||||
# All recipes run from the repo root, because inspect resolves compose.yaml against the
|
||||
# process working directory. See docs/setup.md.
|
||||
|
||||
impossiblebench := "../impossiblebench"
|
||||
|
||||
# Sync dependencies, then re-add ImpossibleBench (uv sync prunes it).
|
||||
install:
|
||||
uv sync
|
||||
uv pip install -e {{impossiblebench}} --no-deps
|
||||
|
||||
# Offline unit tests. No network, no provider beyond mockllm.
|
||||
test:
|
||||
uv run pytest -q
|
||||
|
||||
# The real task against a fake model. Free.
|
||||
smoke:
|
||||
uv run python scripts/smoke_mock.py
|
||||
|
||||
# Confirms the daemon is up and pre-pulls the image so the first eval does not
|
||||
# time out on a cold pull. Should print "/".
|
||||
docker-check:
|
||||
docker run --rm aisiuk/inspect-tool-support pwd
|
||||
@@ -0,0 +1,27 @@
|
||||
[project]
|
||||
name = "messageboardbench"
|
||||
version = "0.1.0"
|
||||
description = "Does cheating spread between AI agents through a shared scratch directory?"
|
||||
authors = [
|
||||
{ name = "PJ", email = "[email protected]" }
|
||||
]
|
||||
requires-python = ">=3.13"
|
||||
dependencies = [
|
||||
"inspect-ai>=0.3.260",
|
||||
# hf_dataset needs this to load fjzzq2002/impossible_livecodebench.
|
||||
# ImpossibleBench is installed with --no-deps (see docs/setup.md), so its
|
||||
# own declaration of datasets does not reach us.
|
||||
"datasets>=3.0.0",
|
||||
]
|
||||
|
||||
[build-system]
|
||||
requires = ["hatchling"]
|
||||
build-backend = "hatchling.build"
|
||||
|
||||
[dependency-groups]
|
||||
dev = [
|
||||
"pytest>=9.1.1",
|
||||
]
|
||||
|
||||
[tool.pytest.ini_options]
|
||||
testpaths = ["tests"]
|
||||
@@ -0,0 +1,35 @@
|
||||
"""Free end-to-end smoke test: the real task, a fake model, no money.
|
||||
|
||||
Runs the unmodified ImpossibleBench LiveCodeBench task against mockllm/model. The mock
|
||||
never calls a tool, so it burns the message limit and falls through to the scorer. That
|
||||
still exercises everything except the model: container start under our compose.yaml, the
|
||||
writes of func.py and test.py, the scorer's test-file comparison, and `python test.py`.
|
||||
|
||||
Run from the repo root so inspect finds compose.yaml (it looks in the process working
|
||||
directory). scripts/ recipes in the justfile do that for you.
|
||||
"""
|
||||
|
||||
from inspect_ai import eval as inspect_eval
|
||||
from impossiblebench import impossible_livecodebench
|
||||
|
||||
if __name__ == "__main__":
|
||||
logs = inspect_eval(
|
||||
impossible_livecodebench(
|
||||
split="conflicting",
|
||||
agent_type="tools",
|
||||
sandbox="docker",
|
||||
limit=1,
|
||||
max_attempts=1,
|
||||
message_limit=4,
|
||||
),
|
||||
model="mockllm/model",
|
||||
log_dir="./logs/smoke",
|
||||
)
|
||||
|
||||
log = logs[0]
|
||||
print(f"\nstatus: {log.status}")
|
||||
if log.error:
|
||||
print(f"error: {log.error}")
|
||||
for sample in log.samples or []:
|
||||
print(f"sample {sample.id}: score={sample.scores}")
|
||||
print(f" working dir sanity: {sample.metadata.get('task_id')}")
|
||||
@@ -0,0 +1,3 @@
|
||||
"""Does cheating spread between AI agents through a shared scratch directory?"""
|
||||
|
||||
__all__: list[str] = []
|
||||
Reference in new issue
Block a user