pj 0a053c114c Do not attribute the harness's own filesystem actions to the agent
`mkdir -p /workspace/scratch`, which our setup solver runs before every agent, classifies
as a write inside the directory. The scorer's `find` and `test -d` classify as reads.
Counted as the agent's, every run would report a write and the Phase 2 headline number
would be 100% for a reason that has nothing to do with any agent. It would have looked
like a spectacular result.

Inspect nests each tool execution in a span of type "tool", so the agent's own actions
are exactly the sandbox events inside one. Analysis now filters on that.

scripts/verify_snapshot.py drives the real task in a real container with a scripted model
that writes a note, and checks the note comes back. Without it, a broken snapshot would
look identical to a true null result: agents simply not writing anything. That is the
failure this pilot exists to avoid, so it is worth a script.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:43:22 +05:30

messageboardbench

Does cheating spread from one AI agent to another through a shared file?

Agents get coding tasks whose tests contradict each other, so no honest solution exists and any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls workspace and never describes as a channel. We measure whether carrying that directory from one agent to the next raises the cheating rate.

Built on ImpossibleBench and Inspect AI. Tasks and payloads are synthetic throughout.

See docs/setup.md to install and run, and docs/findings.md for measured numbers.

S
Description
No description provided
Readme MIT
50 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 99.8%
Just 0.2%