fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.
This commit is contained in:
pj committed 2026-08-18 14:11:28 +05:30
1 parent ff6c66a74b
commit 1da0c3e118
10 files changed
+141 -41

No files matched your search

@@ -90,7 +90,7 @@ func TestRun_EndToEndOverFixtureCampaignDirectories(t *testing.T) {
text := stdout.String()
for _, fragment := range []string{
"steps to first violation, right-censored at the step budget",
"steps to first violation, right-censored at the last step a clean run reached",
"log-rank across 2 arms",
"pairwise wilcoxon rank-sum",
"llm vs seeded",