mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Do not attribute the harness's own filesystem actions to the agent
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies as a write inside the directory. The scorer's `find` and `test -d` classify as reads. Counted as the agent's, every run would report a write and the Phase 2 headline number would be 100% for a reason that has nothing to do with any agent. It would have looked like a spectacular result. Inspect nests each tool execution in a span of type "tool", so the agent's own actions are exactly the sandbox events inside one. Analysis now filters on that. scripts/verify_snapshot.py drives the real task in a real container with a scripted model that writes a note, and checks the note comes back. Without it, a broken snapshot would look identical to a true null result: agents simply not writing anything. That is the failure this pilot exists to avoid, so it is worth a script. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
1 parent
fc43b691a7
commit
0a053c114c
6 files changed
+225
-11
No files matched your search
@@ -39,3 +39,7 @@ pilot:
|
||||
# Rebuild the CSV and hand-read file from an existing log, no re-run.
|
||||
report log_dir="logs/pilot":
|
||||
uv run python scripts/report.py {{log_dir}}
|
||||
|
||||
# Proves the scorer captures a file an agent leaves in scratch. Free, needs Docker.
|
||||
verify-snapshot:
|
||||
uv run python scripts/verify_snapshot.py
|
||||
Reference in new issue
Block a user