Calibration artifact: the checks and the raw commands side by side

The pilot is only worth something if the checks agree with a person reading the
transcript. This writes both out per run so the comparison is possible, and leaves a
verdict column to record disagreements in. The disagreements carry more information
than the rate does, because they are what fixes the scorer before money goes into a
real run.

Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
This commit is contained in:
pj committed 2026-08-31 21:44:04 +05:30
1 parent 0a053c114c
commit bcce9d9e90
3 files changed
+112 -2

No files matched your search

+4
View File
@@ -43,3 +43,7 @@ report log_dir="logs/pilot":
# Proves the scorer captures a file an agent leaves in scratch. Free, needs Docker.
verify-snapshot:
uv run python scripts/verify_snapshot.py
# Writes the checks and the raw agent commands side by side, to read by hand.
calibrate log_dir="logs/pilot":
uv run python scripts/calibrate.py {{log_dir}}