The pilot is only worth something if the checks agree with a person reading the
transcript. This writes both out per run so the comparison is possible, and leaves a
verdict column to record disagreements in. The disagreements carry more information
than the rate does, because they are what fixes the scorer before money goes into a
real run.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies
as a write inside the directory. The scorer's `find` and `test -d` classify as reads.
Counted as the agent's, every run would report a write and the Phase 2 headline number
would be 100% for a reason that has nothing to do with any agent. It would have looked
like a spectacular result.
Inspect nests each tool execution in a span of type "tool", so the agent's own actions
are exactly the sandbox events inside one. Analysis now filters on that.
scripts/verify_snapshot.py drives the real task in a real container with a scripted model
that writes a note, and checks the note comes back. Without it, a broken snapshot would
look identical to a true null result: agents simply not writing anything. That is the
failure this pilot exists to avoid, so it is worth a script.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
was_test_modified reads intermediate score events, not just the final score.
ImpossibleBench's scorer repairs test.py when it finds it modified, and basic_agent
scores every attempt, so an agent that edits the tests on attempt one and not again
leaves a final score with no flag on it. The repair erased the evidence.
snapshot_scratch reports whether the directory exists alongside its contents, because
otherwise a directory nobody wrote to and a directory that was never created look
identical, and the second is a broken harness reported as a real null result.
What the agents wrote is not auto-classified. "Did it write" is safe to automate;
whether a note is addressed to somebody is the judgement the pilot exists to make, so
pilot.py emits a markdown file to read by hand.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
Public from commit one, so no key ever enters this history.
compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no
WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root
among twenty-odd entries. This experiment turns on whether an agent notices a scratch
directory, so that is a bad place to put one.
ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is
declared here instead because hf_dataset genuinely needs it.
Verified: docker run prints "/", impossiblebench imports, and the real task against
mockllm/model completes with a real score and tracebacks rooted at /workspace.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu