Files
sanderling/cmd/internal-tools/analyze/holm.go
T
pj 019d608f65 feat(analyze): survival analysis over campaign directories
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:03:03 +05:30

37 lines
924 B
Go

package main
import (
"math"
"slices"
)
// holm applies the Holm (1979) step-down correction within one family of
// comparisons, enforcing monotonicity across the sorted p-values the way R's
// p.adjust does. Holm, "A Simple Sequentially Rejective Multiple Test
// Procedure", Scandinavian Journal of Statistics 6(2), 65-70.
func holm(pValues []float64) []float64 {
count := len(pValues)
adjusted := make([]float64, count)
order := make([]int, count)
for index := range order {
order[index] = index
}
slices.SortStableFunc(order, func(left, right int) int {
switch {
case pValues[left] < pValues[right]:
return -1
case pValues[left] > pValues[right]:
return 1
default:
return 0
}
})
running := 0.0
for position, index := range order {
scaled := float64(count-position) * pValues[index]
running = math.Max(running, scaled)
adjusted[index] = math.Min(running, 1)
}
return adjusted
}