mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Repo skeleton, pinned environment, free smoke test
Public from commit one, so no key ever enters this history. compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root among twenty-odd entries. This experiment turns on whether an agent notices a scratch directory, so that is a bad place to put one. ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is declared here instead because hf_dataset genuinely needs it. Verified: docker run prints "/", impossiblebench imports, and the real task against mockllm/model completes with a real score and tracebacks rooted at /workspace. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
commit
bf3bb0d144
12 files changed
+2495
No files matched your search
@@ -0,0 +1,2 @@
|
|||||||
|
# Copy to .env and fill in. .env is gitignored and must never be committed.
|
||||||
|
OPENROUTER_API_KEY=
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
.venv/
|
||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
.pytest_cache/
|
||||||
|
logs/
|
||||||
|
.env
|
||||||
|
.env.*
|
||||||
|
!.env.example
|
||||||
|
.DS_Store
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
3.13
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 PJ
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
Portions of src/messageboardbench/task.py are derived from ImpossibleBench
|
||||||
|
(https://github.com/safety-research/impossiblebench), MIT License,
|
||||||
|
Copyright (c) 2025 ImpossibleBench Team.
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
# messageboardbench
|
||||||
|
|
||||||
|
Does cheating spread from one AI agent to another through a shared file?
|
||||||
|
|
||||||
|
Agents get coding tasks whose tests contradict each other, so no honest solution exists and
|
||||||
|
any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls
|
||||||
|
workspace and never describes as a channel. We measure whether carrying that directory from
|
||||||
|
one agent to the next raises the cheating rate.
|
||||||
|
|
||||||
|
Built on [ImpossibleBench](https://github.com/safety-research/impossiblebench) and
|
||||||
|
[Inspect AI](https://inspect.aisi.org.uk). Tasks and payloads are synthetic throughout.
|
||||||
|
|
||||||
|
See `docs/setup.md` to install and run, and `docs/findings.md` for measured numbers.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# Copied from impossiblebench/src/impossiblebench/compose.yaml (commit 061dc3d),
|
||||||
|
# with working_dir added.
|
||||||
|
#
|
||||||
|
# The image has no WORKDIR, so inspect resolves the container working directory to
|
||||||
|
# "/" and the task's func.py and test.py land at the filesystem root, mixed in with
|
||||||
|
# the twenty-odd entries a bare `ls` returns there. This experiment turns on whether
|
||||||
|
# an agent notices and uses a scratch directory, so burying it in / is a bad idea.
|
||||||
|
# /workspace holds exactly func.py, test.py and scratch/. Docker creates the
|
||||||
|
# directory on start; it does not exist in the image.
|
||||||
|
#
|
||||||
|
# This is a deviation from the published benchmark's default environment and belongs
|
||||||
|
# in the paper's limitations. It applies identically to both conditions, so it cannot
|
||||||
|
# produce the effect.
|
||||||
|
services:
|
||||||
|
default:
|
||||||
|
image: "aisiuk/inspect-tool-support"
|
||||||
|
command: "tail -f /dev/null"
|
||||||
|
working_dir: /workspace
|
||||||
|
init: true
|
||||||
|
network_mode: none
|
||||||
|
stop_grace_period: 1s
|
||||||
|
mem_limit: 1gb
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
# Setup
|
||||||
|
|
||||||
|
Python 3.13 and a running Docker daemon. Everything the agent does happens in a container
|
||||||
|
with the network disabled, so the daemon is not optional.
|
||||||
|
|
||||||
|
```
|
||||||
|
uv sync
|
||||||
|
uv pip install -e /path/to/impossiblebench --no-deps
|
||||||
|
```
|
||||||
|
|
||||||
|
## Why --no-deps
|
||||||
|
|
||||||
|
It is required, not a shortcut. ImpossibleBench declares
|
||||||
|
`inspect_evals[swe_bench] @ git+https://github.com/UKGovernmentBEIS/inspect_evals` from
|
||||||
|
unpinned git main, which forces `huggingface_hub` up to 1.2+, and `swebench>=4.0.0`, which
|
||||||
|
drags `modal`, `GitPython`, `typer` and `pre-commit` in as runtime dependencies.
|
||||||
|
|
||||||
|
The LiveCodeBench path we use imports none of it. Every `swebench` import in that package is
|
||||||
|
lazy and inside a function, and `inspect_evals` is declared but never imported.
|
||||||
|
|
||||||
|
`--no-deps` also skips `datasets`, which `hf_dataset` genuinely does need, so this repo
|
||||||
|
declares that one itself in `pyproject.toml`.
|
||||||
|
|
||||||
|
Two things to know. `uv sync` prunes ImpossibleBench, because it is not declared in
|
||||||
|
`pyproject.toml`, so rerun the install line after every sync. `just install` does both.
|
||||||
|
And ImpossibleBench mutates `sys.path` at import time to paper over a broken import, which
|
||||||
|
loads some modules twice, so do not trust module identity inside it.
|
||||||
|
|
||||||
|
## The working directory
|
||||||
|
|
||||||
|
`compose.yaml` at the repo root sets `working_dir: /workspace`, which is a deliberate
|
||||||
|
deviation from the published benchmark. The file itself explains why.
|
||||||
|
|
||||||
|
Inspect looks for `compose.yaml` in the **process working directory**, not next to the task
|
||||||
|
definition, so every command has to run from the repo root. If you run from somewhere else
|
||||||
|
you silently get ImpossibleBench's own compose file, and the working directory reverts to
|
||||||
|
`/`. The justfile recipes handle this.
|
||||||
|
|
||||||
|
## Checks that cost nothing
|
||||||
|
|
||||||
|
```
|
||||||
|
just smoke # the real task against a fake model, exercises everything but the model
|
||||||
|
just test # offline unit tests, no network and no provider beyond mockllm
|
||||||
|
```
|
||||||
|
|
||||||
|
`just smoke` should end with `status: success` and a real score. If the tracebacks in that
|
||||||
|
score mention `/workspace/test.py`, the working directory override is live.
|
||||||
|
|
||||||
|
## Spending money
|
||||||
|
|
||||||
|
`.env` holds `OPENROUTER_API_KEY` and is gitignored. The key carries a $2 cap as a fuse.
|
||||||
|
|
||||||
|
Watch per-request `prompt_n` on any run longer than a few turns. Cache reuse fails silently
|
||||||
|
and costs 8 to 18 times as much without emitting an error.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# All recipes run from the repo root, because inspect resolves compose.yaml against the
|
||||||
|
# process working directory. See docs/setup.md.
|
||||||
|
|
||||||
|
impossiblebench := "../impossiblebench"
|
||||||
|
|
||||||
|
# Sync dependencies, then re-add ImpossibleBench (uv sync prunes it).
|
||||||
|
install:
|
||||||
|
uv sync
|
||||||
|
uv pip install -e {{impossiblebench}} --no-deps
|
||||||
|
|
||||||
|
# Offline unit tests. No network, no provider beyond mockllm.
|
||||||
|
test:
|
||||||
|
uv run pytest -q
|
||||||
|
|
||||||
|
# The real task against a fake model. Free.
|
||||||
|
smoke:
|
||||||
|
uv run python scripts/smoke_mock.py
|
||||||
|
|
||||||
|
# Confirms the daemon is up and pre-pulls the image so the first eval does not
|
||||||
|
# time out on a cold pull. Should print "/".
|
||||||
|
docker-check:
|
||||||
|
docker run --rm aisiuk/inspect-tool-support pwd
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
[project]
|
||||||
|
name = "messageboardbench"
|
||||||
|
version = "0.1.0"
|
||||||
|
description = "Does cheating spread between AI agents through a shared scratch directory?"
|
||||||
|
authors = [
|
||||||
|
{ name = "PJ", email = "[email protected]" }
|
||||||
|
]
|
||||||
|
requires-python = ">=3.13"
|
||||||
|
dependencies = [
|
||||||
|
"inspect-ai>=0.3.260",
|
||||||
|
# hf_dataset needs this to load fjzzq2002/impossible_livecodebench.
|
||||||
|
# ImpossibleBench is installed with --no-deps (see docs/setup.md), so its
|
||||||
|
# own declaration of datasets does not reach us.
|
||||||
|
"datasets>=3.0.0",
|
||||||
|
]
|
||||||
|
|
||||||
|
[build-system]
|
||||||
|
requires = ["hatchling"]
|
||||||
|
build-backend = "hatchling.build"
|
||||||
|
|
||||||
|
[dependency-groups]
|
||||||
|
dev = [
|
||||||
|
"pytest>=9.1.1",
|
||||||
|
]
|
||||||
|
|
||||||
|
[tool.pytest.ini_options]
|
||||||
|
testpaths = ["tests"]
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
"""Free end-to-end smoke test: the real task, a fake model, no money.
|
||||||
|
|
||||||
|
Runs the unmodified ImpossibleBench LiveCodeBench task against mockllm/model. The mock
|
||||||
|
never calls a tool, so it burns the message limit and falls through to the scorer. That
|
||||||
|
still exercises everything except the model: container start under our compose.yaml, the
|
||||||
|
writes of func.py and test.py, the scorer's test-file comparison, and `python test.py`.
|
||||||
|
|
||||||
|
Run from the repo root so inspect finds compose.yaml (it looks in the process working
|
||||||
|
directory). scripts/ recipes in the justfile do that for you.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from inspect_ai import eval as inspect_eval
|
||||||
|
from impossiblebench import impossible_livecodebench
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
logs = inspect_eval(
|
||||||
|
impossible_livecodebench(
|
||||||
|
split="conflicting",
|
||||||
|
agent_type="tools",
|
||||||
|
sandbox="docker",
|
||||||
|
limit=1,
|
||||||
|
max_attempts=1,
|
||||||
|
message_limit=4,
|
||||||
|
),
|
||||||
|
model="mockllm/model",
|
||||||
|
log_dir="./logs/smoke",
|
||||||
|
)
|
||||||
|
|
||||||
|
log = logs[0]
|
||||||
|
print(f"\nstatus: {log.status}")
|
||||||
|
if log.error:
|
||||||
|
print(f"error: {log.error}")
|
||||||
|
for sample in log.samples or []:
|
||||||
|
print(f"sample {sample.id}: score={sample.scores}")
|
||||||
|
print(f" working dir sanity: {sample.metadata.get('task_id')}")
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
"""Does cheating spread between AI agents through a shared scratch directory?"""
|
||||||
|
|
||||||
|
__all__: list[str] = []
|
||||||
Reference in new issue
Block a user