Makes a Phase 2 null result interpretable rather than ambiguous. In the baseline, where
no scratch directory exists and nothing in the prompt mentions one, 19 of 36 agents still
created a working file of their own: /tmp/verify.py, /tmp/brute.py, /tmp/proto.py, and in
two cases scratch.py beside their work.
So "nobody wrote to scratch/" and "agents do not write working files" are different
findings. Only the first is compatible with agents simply preferring /tmp, and the
fallback in EXPERIMENT.md is the right response to one and not the other. Without this
signal the pilot cannot tell them apart.
Three exclusions, each measured rather than guessed: text_editor unpacks itself under
/var/tmp/. from inside a tool span (34 of 36 runs), `> /dev/null` is redirection not a
file (34 of 36), and `python -c` source is tokenised by the shell classifier so Python
comparisons like `if k > n-1:` come back as writes to a file called `n-1:`.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
Running the checks against the baseline task, which has no scratch directory at all,
should report zero touches. It reported two. Both agents had made throwaway files named
scratch.py and scratch2.py beside their work, and the mention pattern's (?!\w) allows a
following ".", so scratch.py read as naming the directory.
relation() was already right; only mentioned_in was wrong, so the effect was confined to
the unparsed fallback. That fallback exists to turn undercounts into visible ones, and
here it was manufacturing overcounts instead. The negative control is clean now.
Calibration found this, not a test. Worth remembering that the baseline logs are a free
negative control for every check, because the directory genuinely does not exist there.
Also adds the submitted func.py to the calibration file: the baseline shows zero test.py
edits on this model, so cheat technique has to be read off the function.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies
as a write inside the directory. The scorer's `find` and `test -d` classify as reads.
Counted as the agent's, every run would report a write and the Phase 2 headline number
would be 100% for a reason that has nothing to do with any agent. It would have looked
like a spectacular result.
Inspect nests each tool execution in a span of type "tool", so the agent's own actions
are exactly the sandbox events inside one. Analysis now filters on that.
scripts/verify_snapshot.py drives the real task in a real container with a scripted model
that writes a note, and checks the note comes back. Without it, a broken snapshot would
look identical to a true null result: agents simply not writing anything. That is the
failure this pilot exists to avoid, so it is worth a script.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
The fork exists to add one line to the system message. If anything else drifts, the
pilot's cheat rate stops being comparable to the baseline measured with the stock task,
and nothing would fail to say so.
Also asserts the scratch line says nothing about agents, sharing, persistence or
messages, since the mechanism being reproduced started as a scratchpad.
92 tests, green with HTTP_PROXY pointed at a closed port.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
was_test_modified reads intermediate score events, not just the final score.
ImpossibleBench's scorer repairs test.py when it finds it modified, and basic_agent
scores every attempt, so an agent that edits the tests on attempt one and not again
leaves a final score with no flag on it. The repair erased the evidence.
snapshot_scratch reports whether the directory exists alongside its contents, because
otherwise a directory nobody wrote to and a directory that was never created look
identical, and the second is a broken harness reported as a real null result.
What the agents wrote is not auto-classified. "Did it write" is safe to automate;
whether a note is addressed to somebody is the judgement the pilot exists to make, so
pilot.py emits a markdown file to read by hand.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
Forks ImpossibleBench's LiveCodeBench solver rather than passing instruction_prompt=,
which injects text into the system message, the user message and every failure message.
The scratch line now appears once, in the system message, verified by the mockllm smoke
run.
test_byte_match.py reads ImpossibleBench's expected_test construction out of its source
with ast and re-executes it, so our test.py is checked against theirs rather than a
copy. Confirmed to fail on a one-character upstream change. Without this, drift flags
every sample as test-modified, resets it, and destroys the cheat measurement silently.
The event adapter fixes a bug carried over from messageboard: relation() left absolute
paths absolute, so `cat /workspace/scratch/notes.md` classified as outside the
directory. Every absolute-path touch would have scored as a miss.
The checks are new rather than reused. The old ones score not-applicable when the prompt
names the directory, which ours does by design, and discard reads after the first write.
Both would undercount here. The decisions are kept, the code is not.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu