Files
pj 9ab59365b2 redact passwords only, and say what each step did in the log (#93)
* feat(hierarchy): name the route a native tree shows

The screen name was web-only: the Chrome driver stamps sanderling-screen on
the root and nothing else does, so every Android and iOS step recorded and
logged an empty screen. The route marker the tree already carries (the
resource id ending in Screen, the same one Transitional counts) names it.

* feat(runner): say what each step did in the step log

One line per step carried only an index and a node count. It now names the
screen, the action, its target and the typed value, the last through the
same redaction the trace and the prompt use. Emitted after the apply so the
line reports what actually happened, skip reason included.

* fix(sidecar): state on android whether a field is a secure entry

maestro's tree mapper copies a fixed attribute list off the device's XML and
password is not on it, so no android element ever reported the fact and the
conservative rule downstream redacted every typed value in the trace, the
prompt and the log. The XML still carries it: re-read it once per settled
snapshot and state the fact on the text fields it matches. A field it cannot
match stays unstated, which still reads as a credential.

* docs: correct the record that android never reports a secure field

Four places said android reports the fact for nothing and that every typed
value there is redacted. The sidecar now states it, so they described the
old behaviour.

* test(sidecar): fail the build if maestro renames the call the fact comes from

* fix(sidecar): state the fact on a field named by its hint alone

collectTextFields matched on class only, so a node the go side calls editable
off its hintText was left unstated and its typed value redacted.

* docs: record that ios and web state secure:false for compose password fields

Both derive the fact from a widget type a compose app never has, so the
value reaches the trace in the clear. Verified on folio on both targets.

* feat(android): read the application id out of an apk

parses the compiled AndroidManifest.xml rather than shelling out to
aapt2, which lives in the versioned build-tools directory that hosts
with only platform-tools never install.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* feat(cli): let --android-app-path supply the bundle id

--bundle-id stays required everywhere else, and an explicit one still
wins, so the apk can never quietly override what was asked for.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* docs: record that the apk can name the package itself

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning

* feat(folio): ask which android device to run on when none is named

* feat(folio): pin ios recipes to one simulator udid and ask when several match

* docs(ci): say how just ios lands on the simulator the boot step chose

* chore(folio): ignore run output anywhere under examples/folio

* refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade

ScreenName kept its own reading of the route markers and disagreed with
Transitional on a marker repeated by a nested node: it named the screen on
a step the runner was skipping as unsettled. One reading now.

* refactor(sidecar): inline the one attempt passed to callViewHierarchy

A named constant and its own comment for a literal used once.

* test(sidecar): compare the whole tree when checking the annotation changes nothing else

The old assertions checked one id string and one bounds value, and passed
with every other attribute stripped off every node. Now the annotated tree
minus the two facts it stated must equal the input.

* fix(sidecar): match a field to its xml node by class as well as id and bounds

A wrapper drawn to the same bounds as the untagged field inside it shared
the field's key, both were dropped as ambiguous, and every value typed into
an untagged field was redacted. The class tells them apart.

* docs(runs): record what the android hierarchy re-read costs per step

Two 1m runs per binary on folio, same seed, before and after the re-read.
2026-09-05 23:54:12 +05:30

79 lines
5.3 KiB
Markdown

---
title: CLI reference
---
# CLI reference
```
sanderling <command> [flags]
```
## `sanderling test`
Run a spec against an app for a fixed duration.
| Flag | Default | Description |
|---|---|---|
| `--spec` | required | Path to the TypeScript spec. |
| `--bundle-id` | required | Target app bundle ID (Android: applicationId). Optional on Android when `--android-app-path` is given: the applicationId is read from the APK's compiled manifest. |
| `--device` | optional (android) | Android device serial, as `adb devices` reports it. Required when more than one device is attached. |
| `--android-app-path` | optional (android) | Path to the APK. Clear-state reinstalls from it instead of running `pm clear`, and it also supplies `--bundle-id` when that flag is absent. |
| `--platform` | `android` | Target platform: `android`, `ios`, or `web`. |
| `--avd` | optional (android) | Android AVD name to boot if no device is connected. Required only when no device is connected and multiple AVDs exist. |
| `--ios-device` | optional (ios) | iOS target: a simulator name/UDID to boot, or a connected device's name, UDID, or CoreDevice id. |
| `--ios-app-path` | optional (ios) | Path to the `.app` bundle for clear-state reinstall (simulator via `simctl`, device via `devicectl`). |
| `--duration` | `5m` | Total test duration (`30s`, `5m`, `2h`, `1d`). |
| `--max-steps` | `0` | Stop after this many steps (`0` = no cap, the duration governs). A step budget is what makes two generators comparable. |
| `--exit-on-violation` | `false` | Stop the run at the first property violation and exit `2`. |
| `--allow-no-properties` | `false` | Run a spec that registers no properties. Such a run judges nothing and can only report no violations, so it is refused by default; pass this when the run measures what the spec extracts. |
| `--allow-no-generator-actions` | `false` | Finish a run the action generator never drove. Such a run judged whatever screen the spec's setup left it on and explored nothing, so it is refused by default; pass this when the run measures where the generator reaches and reaching nothing is the measurement. A run that recorded a violation is never refused, flag or no flag. |
| `--arm` | optional | Experiment label recorded in the run's metadata. Used by the campaign tool to tell one sweep cell from another. |
| `--seed` | `0` | PRNG seed. `0` uses a random seed and records it in `meta.json`. |
| `--generator` | `seeded` | Who picks each action: `seeded` (the run's PRNG) or `llm` (a vision model). See [the LLM generator](../spec-language/#llm-generator). |
| `--label-source` | `visible-text` | How candidates are named to the `llm` generator: `visible-text` (what a user reads) or `resource-id` (the identifier the app assigned). The seeded generator picks by index and ignores this. |
| `--output` | `./runs` | Output directory for traces. |
| `--clear-data` | `true` | Clear app data before launching so the run starts from a fresh install. Pass `--clear-data=false` to resume prior state. |
Exit codes: `0` the run finished (violations, if any, are in the summary), `2` the run stopped on a violation under `--exit-on-violation`, `1` something went wrong. CI reads the difference between `2` and `1` to tell a found bug from a broken harness.
`1` also covers a run that finished cleanly and holds no verdict, which is not a broken harness but is not evidence either: a spec that registers no properties, a run no step of which reached the verifier, and a run whose action generator never drove the app and found nothing. Each names itself on stderr and leaves its full run directory behind, and each has a flag that says "this is the measurement" when it is.
## `sanderling replay [run-or-runs-dir]`
Serve a local web UI for browsing traces. The positional argument is optional and may point at either a runs directory (the parent of many runs) or a single run directory (auto-detected by the presence of `meta.json`). Defaults to `./runs`.
| Flag | Default | Description |
|---|---|---|
| `--port` | `0` (ephemeral) | TCP port to listen on. |
| `--no-open` | `false` | Skip opening the default browser on startup. |
| `--dev` | `false` | Reverse-proxy non-API requests to the Vite dev server on `127.0.0.1:5173`. |
See [the replay UI page](../replay/) for the panel reference and keyboard shortcuts.
## `sanderling doctor`
Check the host environment for a working sanderling setup.
```
sanderling doctor [--platform web|android|ios|ios-device|all]
```
`--platform` defaults to `all`, which runs every platform's checks (deduped). Pass a specific platform to scope the output.
| Platform | Checks |
|---|---|
| `web` | headless Chromium can launch (the bundled CDP surface boots a real browser). |
| `android` | `adb` and `emulator` on PATH, or under `$ANDROID_HOME`, `$ANDROID_SDK_ROOT` or a standard SDK install location; Java 17+; embedded native sidecar JAR is real. |
| `ios` | `xcrun` on PATH; `simctl` on PATH. The simulator path drives the native companion with no JVM. |
| `ios-device` | the `ios` checks plus `devicectl`; the macOS `usbmuxd` socket; a connected, paired device; App Store Connect signing credentials present. |
## `sanderling version`
Print the CLI version.
## Flags coming in v0.1.0
- `--permissions` to pre-set OS-level permissions (for example `--permissions location=allow,notifications=deny`).
Tracked in the [v0.1.0 milestone](https://github.com/priyanshujain/sanderling/milestone/1).