Commit Graph
2 Commits
Author SHA1 Message Date
pj abacd5c5e1 Bound sample wall clock; a request timeout does not stop a hang
One baseline sample sat on a single model request for 2h15m with an established
connection, 0.1% CPU and no file activity since the init solver wrote func.py. The
per-request timeout and max_retries did not bound it, and it blocked the rest of the run
behind it.

time_limit is the control that actually applies, per sample. Samples that finish take 8
to 15 minutes, so 30 is generous. An unbounded straggler costs more than the sample does.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-09-01 00:01:26 +05:30
pj 863c6bd7d8 Analysis rows, pilot and report scripts
was_test_modified reads intermediate score events, not just the final score.
ImpossibleBench's scorer repairs test.py when it finds it modified, and basic_agent
scores every attempt, so an agent that edits the tests on attempt one and not again
leaves a final score with no flag on it. The repair erased the evidence.

snapshot_scratch reports whether the directory exists alongside its contents, because
otherwise a directory nobody wrote to and a directory that was never created look
identical, and the second is a broken harness reported as a real null result.

What the agents wrote is not auto-classified. "Did it write" is safe to automate;
whether a note is addressed to somebody is the judgement the pilot exists to make, so
pilot.py emits a markdown file to read by hand.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:38:52 +05:30