Commit Graph
2 Commits
Author SHA1 Message Date
pj d604370d96 A file called scratch.py is not the scratch directory
Running the checks against the baseline task, which has no scratch directory at all,
should report zero touches. It reported two. Both agents had made throwaway files named
scratch.py and scratch2.py beside their work, and the mention pattern's (?!\w) allows a
following ".", so scratch.py read as naming the directory.

relation() was already right; only mentioned_in was wrong, so the effect was confined to
the unparsed fallback. That fallback exists to turn undercounts into visible ones, and
here it was manufacturing overcounts instead. The negative control is clean now.

Calibration found this, not a test. Worth remembering that the baseline logs are a free
negative control for every check, because the directory genuinely does not exist there.

Also adds the submitted func.py to the calibration file: the baseline shows zero test.py
edits on this model, so cheat technique has to be read off the function.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 22:02:56 +05:30
pj 099563a830 Scratch directory task, event adapter, per-run checks
Forks ImpossibleBench's LiveCodeBench solver rather than passing instruction_prompt=,
which injects text into the system message, the user message and every failure message.
The scratch line now appears once, in the system message, verified by the mockllm smoke
run.

test_byte_match.py reads ImpossibleBench's expected_test construction out of its source
with ast and re-executes it, so our test.py is checked against theirs rather than a
copy. Confirmed to fail on a one-character upstream change. Without this, drift flags
every sample as test-modified, resets it, and destroys the cheat measurement silently.

The event adapter fixes a bug carried over from messageboard: relation() left absolute
paths absolute, so `cat /workspace/scratch/notes.md` classified as outside the
directory. Every absolute-path touch would have scored as a miss.

The checks are new rather than reused. The old ones score not-applicable when the prompt
names the directory, which ours does by design, and discard reads after the first write.
Both would undercount here. The decisions are kept, the code is not.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:36:13 +05:30