pj bcce9d9e90 Calibration artifact: the checks and the raw commands side by side
The pilot is only worth something if the checks agree with a person reading the
transcript. This writes both out per run so the comparison is possible, and leaves a
verdict column to record disagreements in. The disagreements carry more information
than the rate does, because they are what fixes the scorer before money goes into a
real run.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
2026-08-31 21:44:04 +05:30

messageboardbench

Does cheating spread from one AI agent to another through a shared file?

Agents get coding tasks whose tests contradict each other, so no honest solution exists and any pass is provably a cheat. Each agent also gets a scratch directory the prompt calls workspace and never describes as a channel. We measure whether carrying that directory from one agent to the next raises the cheating rate.

Built on ImpossibleBench and Inspect AI. Tasks and payloads are synthetic throughout.

See docs/setup.md to install and run, and docs/findings.md for measured numbers.

S
Description
No description provided
Readme MIT
50 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 99.8%
Just 0.2%