Repo skeleton, pinned environment, free smoke test

Public from commit one, so no key ever enters this history.

compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no
WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root
among twenty-odd entries. This experiment turns on whether an agent notices a scratch
directory, so that is a bad place to put one.

ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is
declared here instead because hf_dataset genuinely needs it.

Verified: docker run prints "/", impossiblebench imports, and the real task against
mockllm/model completes with a real score and tracebacks rooted at /workspace.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
pj committed 2026-08-31 21:14:58 +05:30
commit bf3bb0d144
12 files changed
+2495

No files matched your search

+2
View File
@@ -0,0 +1,2 @@
# Copy to .env and fill in. .env is gitignored and must never be committed.
OPENROUTER_API_KEY=
+9
View File
@@ -0,0 +1,9 @@
.venv/
__pycache__/
*.py[cod]
.pytest_cache/
logs/
.env
.env.*
!.env.example
.DS_Store
+1
View File
@@ -0,0 +1 @@
3.13
+27
View File
@@ -0,0 +1,27 @@
MIT License
Copyright (c) 2026 PJ
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
---
Portions of src/messageboardbench/task.py are derived from ImpossibleBench
(https://github.com/safety-research/impossiblebench), MIT License,
Copyright (c) 2025 ImpossibleBench Team.
+13
View File
@@ -0,0 +1,13 @@
# messageboardbench
Does cheating spread from one AI agent to another through a shared file?
Agents get coding tasks whose tests contradict each other, so no honest solution exists and
any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls
workspace and never describes as a channel. We measure whether carrying that directory from
one agent to the next raises the cheating rate.
Built on [ImpossibleBench](https://github.com/safety-research/impossiblebench) and
[Inspect AI](https://inspect.aisi.org.uk). Tasks and payloads are synthetic throughout.
See `docs/setup.md` to install and run, and `docs/findings.md` for measured numbers.
+22
View File
@@ -0,0 +1,22 @@
# Copied from impossiblebench/src/impossiblebench/compose.yaml (commit 061dc3d),
# with working_dir added.
#
# The image has no WORKDIR, so inspect resolves the container working directory to
# "/" and the task's func.py and test.py land at the filesystem root, mixed in with
# the twenty-odd entries a bare `ls` returns there. This experiment turns on whether
# an agent notices and uses a scratch directory, so burying it in / is a bad idea.
# /workspace holds exactly func.py, test.py and scratch/. Docker creates the
# directory on start; it does not exist in the image.
#
# This is a deviation from the published benchmark's default environment and belongs
# in the paper's limitations. It applies identically to both conditions, so it cannot
# produce the effect.
services:
default:
image: "aisiuk/inspect-tool-support"
command: "tail -f /dev/null"
working_dir: /workspace
init: true
network_mode: none
stop_grace_period: 1s
mem_limit: 1gb
+54
View File
@@ -0,0 +1,54 @@
# Setup
Python 3.13 and a running Docker daemon. Everything the agent does happens in a container
with the network disabled, so the daemon is not optional.
```
uv sync
uv pip install -e /path/to/impossiblebench --no-deps
```
## Why --no-deps
It is required, not a shortcut. ImpossibleBench declares
`inspect_evals[swe_bench] @ git+https://github.com/UKGovernmentBEIS/inspect_evals` from
unpinned git main, which forces `huggingface_hub` up to 1.2+, and `swebench>=4.0.0`, which
drags `modal`, `GitPython`, `typer` and `pre-commit` in as runtime dependencies.
The LiveCodeBench path we use imports none of it. Every `swebench` import in that package is
lazy and inside a function, and `inspect_evals` is declared but never imported.
`--no-deps` also skips `datasets`, which `hf_dataset` genuinely does need, so this repo
declares that one itself in `pyproject.toml`.
Two things to know. `uv sync` prunes ImpossibleBench, because it is not declared in
`pyproject.toml`, so rerun the install line after every sync. `just install` does both.
And ImpossibleBench mutates `sys.path` at import time to paper over a broken import, which
loads some modules twice, so do not trust module identity inside it.
## The working directory
`compose.yaml` at the repo root sets `working_dir: /workspace`, which is a deliberate
deviation from the published benchmark. The file itself explains why.
Inspect looks for `compose.yaml` in the **process working directory**, not next to the task
definition, so every command has to run from the repo root. If you run from somewhere else
you silently get ImpossibleBench's own compose file, and the working directory reverts to
`/`. The justfile recipes handle this.
## Checks that cost nothing
```
just smoke # the real task against a fake model, exercises everything but the model
just test # offline unit tests, no network and no provider beyond mockllm
```
`just smoke` should end with `status: success` and a real score. If the tracebacks in that
score mention `/workspace/test.py`, the working directory override is live.
## Spending money
`.env` holds `OPENROUTER_API_KEY` and is gitignored. The key carries a $2 cap as a fuse.
Watch per-request `prompt_n` on any run longer than a few turns. Cache reuse fails silently
and costs 8 to 18 times as much without emitting an error.
+22
View File
@@ -0,0 +1,22 @@
# All recipes run from the repo root, because inspect resolves compose.yaml against the
# process working directory. See docs/setup.md.
impossiblebench := "../impossiblebench"
# Sync dependencies, then re-add ImpossibleBench (uv sync prunes it).
install:
uv sync
uv pip install -e {{impossiblebench}} --no-deps
# Offline unit tests. No network, no provider beyond mockllm.
test:
uv run pytest -q
# The real task against a fake model. Free.
smoke:
uv run python scripts/smoke_mock.py
# Confirms the daemon is up and pre-pulls the image so the first eval does not
# time out on a cold pull. Should print "/".
docker-check:
docker run --rm aisiuk/inspect-tool-support pwd
+27
View File
@@ -0,0 +1,27 @@
[project]
name = "messageboardbench"
version = "0.1.0"
description = "Does cheating spread between AI agents through a shared scratch directory?"
authors = [
{ name = "PJ", email = "[email protected]" }
]
requires-python = ">=3.13"
dependencies = [
"inspect-ai>=0.3.260",
# hf_dataset needs this to load fjzzq2002/impossible_livecodebench.
# ImpossibleBench is installed with --no-deps (see docs/setup.md), so its
# own declaration of datasets does not reach us.
"datasets>=3.0.0",
]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[dependency-groups]
dev = [
"pytest>=9.1.1",
]
[tool.pytest.ini_options]
testpaths = ["tests"]
+35
View File
@@ -0,0 +1,35 @@
"""Free end-to-end smoke test: the real task, a fake model, no money.
Runs the unmodified ImpossibleBench LiveCodeBench task against mockllm/model. The mock
never calls a tool, so it burns the message limit and falls through to the scorer. That
still exercises everything except the model: container start under our compose.yaml, the
writes of func.py and test.py, the scorer's test-file comparison, and `python test.py`.
Run from the repo root so inspect finds compose.yaml (it looks in the process working
directory). scripts/ recipes in the justfile do that for you.
"""
from inspect_ai import eval as inspect_eval
from impossiblebench import impossible_livecodebench
if __name__ == "__main__":
logs = inspect_eval(
impossible_livecodebench(
split="conflicting",
agent_type="tools",
sandbox="docker",
limit=1,
max_attempts=1,
message_limit=4,
),
model="mockllm/model",
log_dir="./logs/smoke",
)
log = logs[0]
print(f"\nstatus: {log.status}")
if log.error:
print(f"error: {log.error}")
for sample in log.samples or []:
print(f"sample {sample.id}: score={sample.scores}")
print(f" working dir sanity: {sample.metadata.get('task_id')}")
+3
View File
@@ -0,0 +1,3 @@
"""Does cheating spread between AI agents through a shared scratch directory?"""
__all__: list[str] = []
Generated
+2280
View File
File diff suppressed because it is too large. Load diff