mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Calibration artifact: the checks and the raw commands side by side
The pilot is only worth something if the checks agree with a person reading the transcript. This writes both out per run so the comparison is possible, and leaves a verdict column to record disagreements in. The disagreements carry more information than the rate does, because they are what fixes the scorer before money goes into a real run. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
1 parent
0a053c114c
commit
bcce9d9e90
3 files changed
+112
-2
No files matched your search
@@ -43,3 +43,7 @@ report log_dir="logs/pilot":
|
||||
# Proves the scorer captures a file an agent leaves in scratch. Free, needs Docker.
|
||||
verify-snapshot:
|
||||
uv run python scripts/verify_snapshot.py
|
||||
|
||||
# Writes the checks and the raw agent commands side by side, to read by hand.
|
||||
calibrate log_dir="logs/pilot":
|
||||
uv run python scripts/calibrate.py {{log_dir}}
|
||||
Reference in new issue
Block a user