diff --git a/skills/sanderling-run-triage/SKILL.md b/skills/sanderling-run-triage/SKILL.md index d879ab7..c83096a 100644 --- a/skills/sanderling-run-triage/SKILL.md +++ b/skills/sanderling-run-triage/SKILL.md @@ -18,8 +18,9 @@ complete and useful answer. - **0** means the run completed. It does **not** mean no violations. Without `--exit-on-violation` a run that recorded violations still exits 0: measured - on a ten step web run that recorded two, `run complete: 10 steps` and - `2 violation record(s)`, exit code 0. + on a ten step web run of the `throwing` fixture, + `run complete: 10 steps, 10 driven by the generator` and + `1 violation record(s)`, exit code 0. - **2** means the run recorded a violation under `--exit-on-violation` and stopped there. The same ten step run with the flag exits 2 after four steps. - **1** means the harness broke. A bad target gives @@ -27,8 +28,8 @@ complete and useful answer. writes no run directory at all, because the trace is created after the launch succeeds. A run that finished also exits 1 when it holds no verdict to report: a spec with no properties, a run no step of which reached the verifier, or one - whose action generator never drove the app (section 5). Those do leave a full - run directory behind. + whose action generator never drove the app and found nothing (section 5). + Those do leave a full run directory behind. Anything other than 0 and 2 means the run did not complete, and a missing `trace.jsonl` under a 0 or a 2 means there is nothing to judge rather than @@ -193,6 +194,13 @@ of them`, followed by the per reason tally of what never reached the app. `llm-calls.jsonl` carries the cause, one record per step, and the trace's `action_skipped` names it on each step that produced nothing. +Two runs are not refused. One that recorded a violation exits 0 whatever drove +it there, because it holds the verdict the refusal exists to demand, and a +campaign passes no flags, so refusing it would write `exit_code: 1` and lose the +detection to the analysis as missing data. And a sweep measuring where a +generator reaches passes `--allow-no-generator-actions`, for which reaching +nothing on this build is the measurement. + None of the others change the exit code. All of them change what the run proves, which is nothing. @@ -211,7 +219,7 @@ The run says so itself: 7 step(s) judged by nothing: the screen was still moving when it was read ``` -Subtract it. `run complete: 240 steps` with that line is a 233 step run for +Subtract it. A 240 step run with that line is a 233 step run for every purpose that matters, and `replay-ui-summary.sh` reports the pair as "N steps recorded, M verified" for the same reason. A run with many of these is telling you the driver could not get a clean read of your app, which is a diff --git a/skills/sanderling-setup/SKILL.md b/skills/sanderling-setup/SKILL.md index 47d5402..77f5e56 100644 --- a/skills/sanderling-setup/SKILL.md +++ b/skills/sanderling-setup/SKILL.md @@ -205,15 +205,18 @@ step index=10 screen="/index.html" nodes=6 elapsed: 1.715s -run complete: 10 steps +run complete: 10 steps, 10 driven by the generator no violations. ``` `nodes=` is the first number to read and the cheapest lie detector you have. On that page, four elements plus html and body gave `nodes=6`. The same command -against an empty page gives `nodes=2` for every step, and still exits 0 with no -violations. If `nodes` is a handful and never grows, the run is looking at -something that is not your app. +against an empty page gives `nodes=2` for every step, and the run then refuses +itself: the generator has nothing to pick, so the summary reads +`run complete: 10 steps, 0 driven by the generator` and +`10 action(s) never reached the app: no_action_produced 10`, and the process +exits 1 having judged one blank screen ten times. If `nodes` is a handful and +never grows, the run is looking at something that is not your app. `screen=` is web-only and has nothing to do with your spec's screen hooks: the Chrome driver puts the URL hash, or the pathname when there is no hash, on the