fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
This commit is contained in:
pj committed 2026-08-12 23:43:19 +05:30
1 parent f6d562e3dc
commit 2cbe03c3fb
6 files changed
+162 -36

No files matched your search

+6 -1
View File
@@ -23,6 +23,7 @@ type armSummary struct {
MedianStepsToFirstViolation *float64 `json:"median_steps_to_first_violation"`
SurvivalCurve []survivalPoint `json:"survival_curve,omitempty"`
ViolationRate *float64 `json:"violation_rate"`
TotalSteps int `json:"total_steps"`
TotalActions int `json:"total_actions"`
TotalRunHours float64 `json:"total_run_hours"`
Detections int `json:"detections"`
@@ -144,7 +145,11 @@ func summarize(current arm) armSummary {
continue
}
summary.Usable++
summary.TotalActions += item.Steps
summary.TotalSteps += item.Steps
// Steps and actions differ by the steps that chose no action and the
// steps whose action was never dispatched. Only dispatched actions
// exercised the app, so only they belong in a per-action rate.
summary.TotalActions += item.Actions
summary.TotalRunHours += float64(item.DurationMillis) / float64(time.Hour/time.Millisecond)
if item.ClampedToBudget {
summary.EventsHeldAtBudget++