mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-04 20:17:09 +00:00
fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets
A run stops at whichever comes first, the step budget or --duration, so a clean run that reached the wall clock exited with fewer steps than the budget and was still credited with the whole of it. The model arm pays a network call and a screenshot per step, so it reaches the wall sooner and was handed exposure it never had. Nothing checked that two arms shared a budget either. Thirty identical clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14 from the rank-sum while the log-rank in the same report read p 1.0000. groupArms already refused this within one arm. The claims the old convention left in comments and report lines are corrected rather than left standing beside the new behaviour.
This commit is contained in:
1 parent
ff6c66a74b
commit
1da0c3e118
10 files changed
+141
-41
No files matched your search
@@ -90,7 +90,7 @@ func TestRun_EndToEndOverFixtureCampaignDirectories(t *testing.T) {
|
||||
|
||||
text := stdout.String()
|
||||
for _, fragment := range []string{
|
||||
"steps to first violation, right-censored at the step budget",
|
||||
"steps to first violation, right-censored at the last step a clean run reached",
|
||||
"log-rank across 2 arms",
|
||||
"pairwise wilcoxon rank-sum",
|
||||
"llm vs seeded",
|
||||
|
||||
Reference in new issue
Block a user