One baseline sample sat on a single model request for 2h15m with an established
connection, 0.1% CPU and no file activity since the init solver wrote func.py. The
per-request timeout and max_retries did not bound it, and it blocked the rest of the run
behind it.
time_limit is the control that actually applies, per sample. Samples that finish take 8
to 15 minutes, so 30 is generous. An unbounded straggler costs more than the sample does.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
Forks ImpossibleBench's LiveCodeBench solver rather than passing instruction_prompt=,
which injects text into the system message, the user message and every failure message.
The scratch line now appears once, in the system message, verified by the mockllm smoke
run.
test_byte_match.py reads ImpossibleBench's expected_test construction out of its source
with ast and re-executes it, so our test.py is checked against theirs rather than a
copy. Confirmed to fail on a one-character upstream change. Without this, drift flags
every sample as test-modified, resets it, and destroys the cheat measurement silently.
The event adapter fixes a bug carried over from messageboard: relation() left absolute
paths absolute, so `cat /workspace/scratch/notes.md` classified as outside the
directory. Every absolute-path touch would have scored as a miss.
The checks are new rather than reused. The old ones score not-applicable when the prompt
names the directory, which ours does by design, and discard reads after the first write.
Both would undercount here. The decisions are kept, the code is not.
Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu