Files
messageboardbench/justfile
T
pj 0a053c114c Do not attribute the harness's own filesystem actions to the agent
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies
as a write inside the directory. The scorer's `find` and `test -d` classify as reads.
Counted as the agent's, every run would report a write and the Phase 2 headline number
would be 100% for a reason that has nothing to do with any agent. It would have looked
like a spectacular result.

Inspect nests each tool execution in a span of type "tool", so the agent's own actions
are exactly the sandbox events inside one. Analysis now filters on that.

scripts/verify_snapshot.py drives the real task in a real container with a scripted model
that writes a note, and checks the note comes back. Without it, a broken snapshot would
look identical to a true null result: agents simply not writing anything. That is the
failure this pilot exists to avoid, so it is worth a script.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:43:22 +05:30

46 lines
1.4 KiB
Makefile

# All recipes run from the repo root, because inspect resolves compose.yaml against the
# process working directory. See docs/setup.md.
impossiblebench := "../impossiblebench"
# Sync dependencies, then re-add ImpossibleBench (uv sync prunes it).
install:
uv sync
uv pip install -e {{impossiblebench}} --no-deps
# Offline unit tests. No network, no provider beyond mockllm.
test:
uv run pytest -q
# The real task against a fake model. Free.
smoke:
uv run python scripts/smoke_mock.py
# Confirms the daemon is up and pre-pulls the image. Should print "/".
docker-check:
docker run --rm aisiuk/inspect-tool-support pwd
# Confirms the model slug, the key and tool support. Costs about a cent.
model-check:
uv run python scripts/model_check.py
# Phase 1: cheat rate on conflicting, solve rate on original. About $1.
baseline:
uv run python -u scripts/baseline.py conflicting original
# The forked task against a fake model. Free. Checks the scratch wiring.
smoke-scratch:
uv run python scripts/smoke_scratch.py
# Phase 2: thirty agents with a scratch directory. Does anybody write anything?
pilot:
uv run python -u scripts/pilot.py
# Rebuild the CSV and hand-read file from an existing log, no re-run.
report log_dir="logs/pilot":
uv run python scripts/report.py {{log_dir}}
# Proves the scorer captures a file an agent leaves in scratch. Free, needs Docker.
verify-snapshot:
uv run python scripts/verify_snapshot.py