Merge branch 'skills-setup-and-triage' into spec-authoring-skills

This commit is contained in:
pj committed 2026-08-15 22:59:47 +05:30
commit 1adbb35305
2 files changed
+470

No files matched your search

+210
View File
@@ -0,0 +1,210 @@
---
name: sanderling-run-triage
description: Work out what a finished sanderling run actually proves. Use before trusting a green run, before filing the bug a red run seems to show, and any time the exit code is the only thing anyone has looked at.
---
# Reading a run honestly
A run produces one number that is easy to read and several that are worth
reading. The easy one says whether a process finished. It does not say whether
anything was checked, whether what was checked was your app, or whether the
violation it reports is about the app at all.
Work through the sections in order and report what you established and what you
could not. "This run is not evidence, and here is the signal that says so" is a
complete and useful answer.
## 1. The exit codes
- **0** means the run completed. It does **not** mean no violations. Without
`--exit-on-violation` a run that recorded violations still exits 0: measured
on a ten step web run that recorded two, `run complete: 10 steps` and
`2 violation record(s)`, exit code 0.
- **2** means the run recorded a violation under `--exit-on-violation` and
stopped there. The same ten step run with the flag exits 2 after four steps.
- **1** means the harness broke. A bad target gives
`error: launch app: page load error net::ERR_UNSAFE_PORT` and exit 1, and
writes no run directory at all, because the trace is created after the launch
succeeds.
Anything other than 0 and 2 means the run did not complete, and a missing
`trace.jsonl` under a 0 or a 2 means there is nothing to judge rather than
nothing to report.
Exit 2 is not a conviction. `.github/scripts/folio-run.sh` is the worked example
worth reading in full: it exists because a thrown predicate reaches exit 2 by the
identical path a real conviction does, and so does a violation of a real but
unrelated property in the same spec. It sorts a trace's violations three ways,
by name and by `is_error`: convictions of the properties the leg gates on,
predicates that threw, and other real violations the leg has nothing to say
about. Do the same sort by hand before you call a 2 a finding.
## 2. Reading a witness
Witnesses live in `trace.jsonl`, one object per step under `witnesses`, keyed by
property name. A real conviction and a real throw from the same run:
```json
{"step": 4, "violations": ["countStaysUnderThree"],
"witnesses": {"countStaysUnderThree": {
"reason": "predicate false", "step": 4, "detected_step": 4,
"extractors": {"count": 3}}}}
{"step": 5, "violations": ["throwsOnceCountIsFour"],
"witnesses": {"throwsOnceCountIsFour": {
"reason": "Error: boom: no reading for this screen at <eval>:501:37(14)",
"is_error": true, "step": 5, "detected_step": 5,
"extractors": {"count": 4}}}}
```
`step` is where the failed obligation was armed and `detected_step` is where the
evaluation produced the violation; for a deferred obligation (a `next`, an
`eventually`) they differ, and `extractors` is `detected_step`'s state, not
`step`'s.
The discipline is one sentence: open the witness and confirm those values could
actually produce that verdict. An iOS witness read `typedAmount = 0`, and
`submitChangesBalanceByTypedAmount` in `examples/folio/sanderling/predicates.ts`
returns true at `typedAmount === 0` before it compares anything. So the trace
appeared to show a conviction that could not have happened. The verdict was real
and the artifact was lying, and until that was resolved neither the bug report
nor the fix could be trusted.
When a witness value looks impossible, suspect the reading before you suspect
the property. Values reach a witness through the driver, and the driver can be
wrong in ways the spec cannot see: erasing a text field used to leave characters
behind, because a backspace only deletes to the left of the cursor and the
runner taps the field's centre, and 7 of 19 measured `InputText` observations
left residue that the spec then reasoned about as if it were the typed value.
Two more things the witness tells you, if the spec extracts them. folio declares
`extract("lastAction", s => s.lastAction)` precisely so they land in the trace:
`applied: true` means the runner saw the dispatch succeed, `applied: null` means
it was dispatched and nobody knows whether it landed, and `relaunched: true`
means the app restarted between the two readings. A property attributing an
effect to an action of unknown fate is unsound; see `sanderling-spec-review`.
Finally, a property violates once. After it fires, its residual stays `false`
(or `{"op": "error", ...}`) for every remaining step and it is never evaluated
again. Measured across steps 4 to 10 of that run, `countStaysUnderThree` reads
`{"op": "false"}` at every step after the first. So the violation count is a
count of distinct properties, not of occurrences, and everything after a
property's first violation is unchecked by that property.
## 3. A green run fails in two ways
Either it checked nothing, or it checked and the fuzzer never reached the bug.
These need opposite responses (fix the spec or the hooks; spend more budget or
better actions) and the exit code distinguishes neither.
The first is not a hypothetical. Against an empty page, six steps, exit 0, `no
violations`, and `countNeverNegative` judged **0 of 6**: its extractor returned
null every step, its guard short-circuited, and its residual read `{"op":
"true"}` at every step, exactly as it reads when it compares real values.
So count, per property, the steps where its guard passed and it compared
something (**judged**) against the steps where it returned true without
comparing anything (**declined**). `.github/scripts/replay-ui-summary.sh` does
this for the replay-ui spec and prints a judged/declined table for exactly this
reason. To do it by hand from a trace:
- fold `extractor_changes` forward per step. Only extractors whose value changed
are recorded, so a step with no entry for an extractor means unchanged, not
absent. Measured: `{"count": {"prev": 0, "curr": 2}}` at one step and no
`count` entry at the next.
- skip steps carrying `skipped_verification` or `transitional`. They advance
nothing.
- apply each property's own guard to the folded values and count.
That script also carries the honest warning about this technique: restating a
property's guard outside the property is a second copy that can drift, so it
checks that the trace's property names still match the ones it counts and that
the spec still declares the extractors it reads, and it fails loudly when either
moves.
Do not try to read judged-versus-declined off `residuals`. `always(p)` residuals
to `{"op": "true"}` whether `p` compared real values or short-circuited, so the
two are indistinguishable there.
## 4. A red run fails in two ways
Either a property was proved false about the app, or a predicate threw.
`is_error` in the witness separates them and they mean opposite things.
A conviction is a claim about the app. A throw is a claim about the spec, and it
is worse than an unhelpful result: the property is violated from that step on
whatever the app does, so it checks nothing for the rest of the run, and under
`--exit-on-violation` the run ended there so nothing past it was checked by
anything. The `reason` carries the JavaScript error and its location, which is
usually enough to find it: `Error: boom: no reading for this screen at
<eval>:501:37(14)`.
The third case is a real violation of a property that is not the one you are
asking about. It is a finding, and it is somebody's bug, but the run has nothing
to say about the question you asked it. Name the property before you claim the
result.
## 5. When a run is not evidence at all
Some runs never got far enough for any of the above to matter, and every one of
them exits 0 and reports no violations.
**It never reached the app.** A launch flake left a fuzzer on the device
launcher for 200 steps in 65 seconds, two nodes per snapshot, exit 0, no
violations (issue #81). The check is that the app's own marker appears in the
trace at all: the folio CI leg greps for `"AddTransactionScreen"` and fails the
run when it is absent, which is more honest than any exit code it could read.
The run's stdout also carries `app left foreground; relaunching` with the
package it found instead.
**The hierarchy is a handful of nodes.** `nodes=` in each step line is the
cheapest signal there is. Measured: 6 on a four-element page, 2 on an empty one.
A run whose `nodes` never leaves single digits is looking at a launcher, a
crash screen, or a page that failed to boot.
**It never left one screen.** `screen=` constant for the whole trace, or a
`route` extractor that never changes value.
**Steps far faster than the run's own median.** Take the per-step deltas from
each step's `timestamp` and compare them against the run's median. A stretch of
steps at a fraction of it is a driver that is not waiting for an app, because
there is no app to wait for: 200 steps in 65 seconds is 325 ms a step, against
seconds a step for a run that is driving something real.
**It spent its budget on one action.** Count `next_action` by kind and selector.
A run whose actions are one selector explored nothing, whatever its step count.
None of these change the exit code. All of them change what the run proves,
which is nothing.
## 6. `skipped_verification`, `transitional`, and the judged count
The runner skips the verifier for a step whose hierarchy was still moving: an
Android NavHost mid cross-fade after the retry budget, or a hierarchy fetch that
failed or came back empty. Pushing such a tree would poison the previous/current
extractor advance and make the next clean step convict a healthy app, so the
step is recorded for replay and judged by nothing. `transitional` marks the
tree; `skipped_verification` is set exactly when the verifier was skipped.
The run says so itself:
```
7 step(s) judged by nothing: the screen was still moving when it was read
```
Subtract it. `run complete: 240 steps` with that line is a 233 step run for
every purpose that matters, and `replay-ui-summary.sh` reports the pair as
"N steps recorded, M verified" for the same reason. A run with many of these is
telling you the driver could not get a clean read of your app, which is a
finding about the setup and worth chasing rather than quietly accepting the
smaller number.
## Reporting
For any run, report: the exit code and whether `--exit-on-violation` was set;
steps recorded against steps verified; per property, judged against declined;
for every violation, its `is_error` and the witness values you actually opened;
and which of the section 5 signals you checked. Name the step behind any claim.
A run is evidence only for the properties that judged, and only for the app it
was actually looking at. Everything else it produced is a log.
+260
View File
@@ -0,0 +1,260 @@
---
name: sanderling-setup
description: Get sanderling running against an app that is not folio. Use before writing a spec for a new app, when deciding what test hooks the app needs, and when a run will not start or starts and sees nothing.
---
# Getting sanderling onto your app
The goal of setup is not a run that finishes. It is a run whose output you can
believe. Two things decide that, and both are usually treated as chores: the
handles the app exposes, and the state the app starts in. Everything else here
is plumbing.
Every flag below is one the binary accepts. `sanderling test -h` is the
authority, not this file and not the manual: the manual currently documents
`--launcher-activity`, which the binary answers with `flag provided but not
defined`. Check before you use a flag you have not seen work.
## 1. Install, then check the host
The CLI installs from the release script, and the spec package from npm:
```sh
curl -fsSL https://raw.githubusercontent.com/priyanshujain/sanderling/master/install.sh | bash
npm install --save-dev @sanderling/spec
```
Both come from the same release tag and the CLI bundles the package's TypeScript
when it evaluates your spec, so they move together.
`sanderling doctor` reports the host's readiness per platform and exits non-zero
if anything is missing. On a Mac with no Android SDK it says:
```
OK adb on PATH
FAIL emulator on PATH or under ANDROID_HOME: not on PATH and ANDROID_HOME is unset
OK java 17+ on PATH
FAIL sidecar JAR is real (not placeholder): placeholder JAR embedded; run `make sidecar && make sanderling` to embed the real fat JAR
error: 2 check(s) failed
```
Scope it with `--platform web|android|ios|ios-device|all` (default `all`). Web
needs a Chromium that launches headless. Android needs `adb`, an emulator on
PATH or under `ANDROID_HOME`, Java 17 or newer, and the embedded sidecar JAR.
iOS needs `xcrun` and `simctl`; `ios-device` adds `devicectl`, the macOS usbmuxd
socket, a connected paired device, and App Store Connect signing credentials.
Read the doctor's Android result as advisory rather than final: its emulator
check today looks only at PATH, `ANDROID_HOME` and `ANDROID_SDK_ROOT`, while a
run also searches `~/Library/Android/sdk`, `~/Android/Sdk` and the Homebrew
command-line-tools paths. The run's own error names every location it tried, so
that is the one to trust. In the other direction, a missing SDK can surface
during a run as `sidecar health check: context deadline exceeded` about thirty
seconds in, which names the symptom and not the cause (issue #69). If you see
it, go back to `sanderling doctor --platform android` before believing anything
about the sidecar.
Two traps if you build from source rather than installing a release. A plain
`go build ./cmd/sanderling` embeds a placeholder sidecar JAR, so every Android
run stops at `sidecar: binary built without -tags withsidecar`; `make sanderling`
(or `make sanderling-android`) embeds the real one. And `go run ./cmd/sanderling
test` collapses the process exit code: a run that exits 2 comes back from
`go run` as 1 with `exit status 2` printed. Use the built binary whenever the
exit code matters, which is always in CI.
## 2. Point it at the app
Android takes the applicationId, boots an AVD with `--avd`, and picks between
attached devices with `--device <serial>` as `adb devices` prints it:
```sh
sanderling test --spec spec.ts --bundle-id com.example.app --avd Pixel_7_API_34
```
iOS takes `--platform ios` and `--ios-device`, which accepts a simulator name or
UDID, or a connected device's name, UDID, or CoreDevice id. `--ios-app-path`
points at the `.app` bundle and is what makes clear-state real; see section 4.
Web takes a URL as the bundle id:
```sh
sanderling test --spec spec.ts --platform web --bundle-id http://127.0.0.1:8799/index.html
```
The web target has to genuinely load. A page that boots to a blank canvas still
produces steps, still exits 0, and proves nothing: folio's own web leg needs
COOP/COEP headers or its sqlite worker never starts, which is why
`.github/scripts/folio-run.sh` serves the build itself instead of using a stock
static server. Confirm the app rendered before you read anything else.
## 3. Test hooks are a prerequisite, not a polish step
This is the part that decides whether a spec is possible at all. The header of
`replay-ui/sanderling/spec.ts` states it as the lesson it is:
> The hooks it drives (data-testid, data-step, ...) were added to the UI for
> this spec. Needing them is the lesson: a UI with no stable handles is a UI
> nothing can assert on, and that is as true for a person writing a test as it
> is for a fuzzer.
A fuzzer is not asking for anything a human test author does not need. It is
only less able to squint at a screenshot and guess. Budget the hooks as part of
adopting sanderling, before the spec, not after the first vacuous run.
`testTag` is the portable name. `internal/hierarchy/hierarchy.go` aliases it to
`resource-id`, `identifier` and `accessibilityIdentifier`, so one selector
matches on every platform. What you have to add differs:
**Compose on Android.** `Modifier.testTag("AddAccountSubmit")` alone does not
reach the accessibility tree. The tree only carries it when a root composable
sets `semantics { testTagsAsResourceId = true }`. folio does this once, at the
app root, through an expect/actual bridge:
`examples/folio/app/shared/src/androidMain/kotlin/app/folio/ui/TestTagBridge.android.kt`.
Without it every `testTag` selector matches nothing, every property over it
declines, and the run goes green having checked nothing.
**Web.** `data-testid` is the hook. Every `data-*` attribute on the element
reaches the spec under `attrs`, camel-cased the way `dataset` does it, so
`data-step-count` reads as `attrs.stepCount`. That is how the replay-ui spec
reads a panel's own claim about which step it is showing rather than re-deriving
it. Hooks that carry a value, not just an identity, are what make cross-panel
agreement properties possible.
**iOS.** `accessibilityIdentifier`, set via `.accessibilityIdentifier` in
SwiftUI or UIKit. Compose Multiplatform maps `testTag` to it for you.
Two rules about the names themselves. A `testTag` selector falls through to a
substring compare, so `{testTag: "Sub"}` matches `AddAccountSubmit`: make each
hook a whole distinct name rather than a fragment of another. And give every
screen a marker of its own, because a route extractor is what lets a property
decline on the screens it has nothing to say about.
The check that a hook exists is not that you added it. It is that you can point
at a step in a real trace where a selector over it resolved to a value.
`sanderling-spec-authoring` covers which hooks a spec needs and in what order to
add them; this section is about what each platform requires before any of that
reaches the tree.
## 4. A run must start from a known state
`--clear-data` defaults to true and is the difference between a repeatable run
and a measurement of your own leftovers. A second run that inherits the first
one's accounts, cache and completed onboarding diverges at step 1: the seed
reproduces nothing, the two runs' step counts are not comparable, and any number
you quote from the pair is noise.
What "clear" reaches depends on the platform, and in two cases it silently
reaches less than you expect:
- Android wipes app data through the sidecar. On OEM builds that deny
`pm clear`, pass `--android-app-path <apk>` and it uninstalls and reinstalls
instead.
- iOS simulator without `--ios-app-path` resets the data container only and
prints `clear-state requested without an app path: resetting the data
container only`. With the path it does a full `simctl` uninstall and install.
The container wipe is a real reset and folio's own iOS leg relies on it; the
reinstall path is the one that races FrontBoard.
- iOS on a physical device without `--ios-app-path` does not clear at all. It
prints `clear-state on a physical device requires --ios-app-path for a
reinstall; skipping (state not cleared)` and carries on. A device run left on
the default flag inherits every previous run's data.
- Web clears cookies and the target origin's storage. It cannot touch your
backend. If your app's state lives on a server, reset it yourself between
runs.
`--clear-data=false` is a legitimate choice in one situation: you have just
installed a fresh build, so the app is already in clear state and an in-run
reinstall would only add a failure mode. Outside that, a run that resumes is a
run you cannot repeat.
## 5. The device does not have to be local
Android talks to whatever adb server the environment names.
`ADB_SERVER_SOCKET=tcp:host:port` (or `tcp:port` for a server on this machine)
is read first, then the older `ANDROID_ADB_SERVER_ADDRESS` and
`ANDROID_ADB_SERVER_PORT` pair, then the loopback default. The CLI shells out to
`adb` and inherits it; the JVM sidecar resolves the same variables when it
attaches to a serial.
Two things to get right. Pass `--device <serial>` exactly as the remote server
reports it: with no serial the sidecar's target is a local `localhost:5555`, not
your remote device. And a serial that already looks like `host:port` is dialled
straight at adbd, bypassing any server, which is a different path with different
failure modes. A value the sidecar cannot parse fails the run rather than
falling back to loopback, and that is deliberate: emulator serials are numbered
per server, so a quiet fallback would drive whatever this machine calls
`emulator-5554` and report the results as the remote device's.
## 6. What a first run prints
A ten step web run, in full:
```
bundled spec: 16532 bytes (sha256=71375ed5bfc7)
bundled web spec: 33131 bytes (sha256=779bae3c8fee)
spec loaded into verifier
trace dir: runs/20260815-172356
running for 1m30s or 10 steps, whichever comes first (seed=7)
step index=1 screen="/index.html" nodes=6
...
step index=10 screen="/index.html" nodes=6
elapsed: 1.715s
run complete: 10 steps
no violations.
```
`nodes=` is the first number to read and the cheapest lie detector you have. On
that page, four elements plus html and body gave `nodes=6`. The same command
against an empty page gives `nodes=2` for every step, and still exits 0 with no
violations. If `nodes` is a handful and never grows, the run is looking at
something that is not your app.
`screen=` is the route marker your spec's screen hooks produce. A run where it
never changes never left one screen.
The summary can carry a third line you should never skim past:
```
7 step(s) judged by nothing: the screen was still moving when it was read
```
Those steps were recorded but no property judged them, so the run's step count
and its checked count are different numbers. `sanderling-run-triage` is about
what to do with that.
Set `--max-steps` whenever you intend to compare two runs: a step budget is what
makes them comparable, since duration alone does not. `--seed` fixes the PRNG,
and seed 0 draws a random one and records it in `meta.json`.
## 7. The run directory
Each run writes `<output>/<UTC timestamp>/`, containing `meta.json`,
`trace.jsonl`, and one PNG per step under `screenshots/`. `--output` defaults to
`./runs`.
`meta.json` is the run's identity: seed, spec path, bundled spec sha256,
platform, bundle id, start and end times, generator, `max_steps`,
`duration_millis`, host, and the `--arm` label if you set one. Two runs that
differ in any of those are different runs and cannot be pooled.
If the app never launched, there is no run directory at all: the launch error
comes before the trace is created. `error: launch app: ...` with nothing under
`./runs` means the run never began, which is a different thing from a run that
began and found nothing.
Open a run with `sanderling replay <dir>`, which accepts either the parent runs
directory or a single run directory.
## Reporting
Say what you actually ran and what came back: the `doctor` output you got rather
than the one you expected, the exact `sanderling test` command, the step count
and the `nodes=` figure from the first run, and for each hook you added, the step
in a real trace where a selector over it resolved. Name what you could not
establish, particularly any platform you did not run on.
Setup is finished when a property can be written that could fail. Write it with
`sanderling-spec-authoring`, review it with `sanderling-spec-review`, and read
the run it produces with `sanderling-run-triage`.