Commit Graph
11 Commits
Author SHA1 Message Date
pj 175c9d48f4 Add deterministic SWE candidate screening 2026-09-15 17:54:57 +05:30
pj aeaae4291f Fail fast on invalid SWE prerequisite tasks 2026-09-15 17:16:54 +05:30
pj b3c935fbfa Add independent-agent SWE prompt-ablation pilot 2026-09-15 16:25:27 +05:30
pj 72d77018d8 Checkpoint experiments through SWE pilot v2 2026-09-15 15:46:10 +05:30
pj abacd5c5e1 Bound sample wall clock; a request timeout does not stop a hang
One baseline sample sat on a single model request for 2h15m with an established
connection, 0.1% CPU and no file activity since the init solver wrote func.py. The
per-request timeout and max_retries did not bound it, and it blocked the rest of the run
behind it.

time_limit is the control that actually applies, per sample. Samples that finish take 8
to 15 minutes, so 30 is generous. An unbounded straggler costs more than the sample does.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-09-01 00:01:26 +05:30
pj d604370d96 A file called scratch.py is not the scratch directory
Running the checks against the baseline task, which has no scratch directory at all,
should report zero touches. It reported two. Both agents had made throwaway files named
scratch.py and scratch2.py beside their work, and the mention pattern's (?!\w) allows a
following ".", so scratch.py read as naming the directory.

relation() was already right; only mentioned_in was wrong, so the effect was confined to
the unparsed fallback. That fallback exists to turn undercounts into visible ones, and
here it was manufacturing overcounts instead. The negative control is clean now.

Calibration found this, not a test. Worth remembering that the baseline logs are a free
negative control for every check, because the directory genuinely does not exist there.

Also adds the submitted func.py to the calibration file: the baseline shows zero test.py
edits on this model, so cheat technique has to be read off the function.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 22:02:56 +05:30
pj bcce9d9e90 Calibration artifact: the checks and the raw commands side by side
The pilot is only worth something if the checks agree with a person reading the
transcript. This writes both out per run so the comparison is possible, and leaves a
verdict column to record disagreements in. The disagreements carry more information
than the rate does, because they are what fixes the scorer before money goes into a
real run.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:44:04 +05:30
pj 0a053c114c Do not attribute the harness's own filesystem actions to the agent
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies
as a write inside the directory. The scorer's `find` and `test -d` classify as reads.
Counted as the agent's, every run would report a write and the Phase 2 headline number
would be 100% for a reason that has nothing to do with any agent. It would have looked
like a spectacular result.

Inspect nests each tool execution in a span of type "tool", so the agent's own actions
are exactly the sandbox events inside one. Analysis now filters on that.

scripts/verify_snapshot.py drives the real task in a real container with a scripted model
that writes a note, and checks the note comes back. Without it, a broken snapshot would
look identical to a true null result: agents simply not writing anything. That is the
failure this pilot exists to avoid, so it is worth a script.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:43:22 +05:30
pj 863c6bd7d8 Analysis rows, pilot and report scripts
was_test_modified reads intermediate score events, not just the final score.
ImpossibleBench's scorer repairs test.py when it finds it modified, and basic_agent
scores every attempt, so an agent that edits the tests on attempt one and not again
leaves a final score with no flag on it. The repair erased the evidence.

snapshot_scratch reports whether the directory exists alongside its contents, because
otherwise a directory nobody wrote to and a directory that was never created look
identical, and the second is a broken harness reported as a real null result.

What the agents wrote is not auto-classified. "Did it write" is safe to automate;
whether a note is addressed to somebody is the judgement the pilot exists to make, so
pilot.py emits a markdown file to read by hand.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:38:52 +05:30
pj 099563a830 Scratch directory task, event adapter, per-run checks
Forks ImpossibleBench's LiveCodeBench solver rather than passing instruction_prompt=,
which injects text into the system message, the user message and every failure message.
The scratch line now appears once, in the system message, verified by the mockllm smoke
run.

test_byte_match.py reads ImpossibleBench's expected_test construction out of its source
with ast and re-executes it, so our test.py is checked against theirs rather than a
copy. Confirmed to fail on a one-character upstream change. Without this, drift flags
every sample as test-modified, resets it, and destroys the cheat measurement silently.

The event adapter fixes a bug carried over from messageboard: relation() left absolute
paths absolute, so `cat /workspace/scratch/notes.md` classified as outside the
directory. Every absolute-path touch would have scored as a miss.

The checks are new rather than reused. The old ones score not-applicable when the prompt
names the directory, which ours does by design, and discard reads after the first write.
Both would undercount here. The decisions are kept, the code is not.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:36:13 +05:30
pj bf3bb0d144 Repo skeleton, pinned environment, free smoke test
Public from commit one, so no key ever enters this history.

compose.yaml is ImpossibleBench's, plus working_dir: /workspace. The image has no
WORKDIR, so inspect resolves it to "/" and the task files land at the filesystem root
among twenty-odd entries. This experiment turns on whether an agent notices a scratch
directory, so that is a bad place to put one.

ImpossibleBench installs with --no-deps to keep the swebench tree out; datasets is
declared here instead because hf_dataset genuinely needs it.

Verified: docker run prints "/", impossiblebench imports, and the real task against
mockllm/model completes with a real score and tracebacks rooted at /workspace.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:14:58 +05:30