fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.
This commit is contained in:
pj committed 2026-08-18 20:13:12 +05:30
1 parent b7ee23942e
commit f0726a1e61
6 files changed
+144 -52

No files matched your search

-12
View File
@@ -328,15 +328,3 @@ func observationOf(item classifiedRun, budget int) observation {
// the steps it never ran are not exposure it survived.
return observation{Steps: float64(min(item.Steps, budget)), Event: false}
}
// stepTimes is the observations flattened to plain numbers, censored runs held
// at the steps they ran. Holding them there rather than dropping them is
// conservative: it can only understate how much sooner a violating arm finds
// its first defect, never overstate it.
func (a arm) stepTimes() []float64 {
var result []float64
for _, item := range a.observations() {
result = append(result, item.Steps)
}
return result
}