fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.
This commit is contained in:
pj committed 2026-08-18 14:11:28 +05:30
1 parent ff6c66a74b
commit 1da0c3e118
10 files changed
+141 -41

No files matched your search

+1 -1
View File
@@ -58,7 +58,7 @@ func TestVargaDelaneyA12_BoundaryCases(t *testing.T) {
// The counting definition and the rank-sum route must agree, including when the
// samples are tied against each other, which is the case the evaluation data is
// always in because censored runs are all held at the budget.
// usually in because censored runs pile up on the step they stopped at.
func TestVargaDelaneyA12_AgreesWithRankSumStatistic(t *testing.T) {
cases := [][2][]float64{
{{1, 2, 3}, {2, 3, 4}},