mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
* feat(hierarchy): name the route a native tree shows The screen name was web-only: the Chrome driver stamps sanderling-screen on the root and nothing else does, so every Android and iOS step recorded and logged an empty screen. The route marker the tree already carries (the resource id ending in Screen, the same one Transitional counts) names it. * feat(runner): say what each step did in the step log One line per step carried only an index and a node count. It now names the screen, the action, its target and the typed value, the last through the same redaction the trace and the prompt use. Emitted after the apply so the line reports what actually happened, skip reason included. * fix(sidecar): state on android whether a field is a secure entry maestro's tree mapper copies a fixed attribute list off the device's XML and password is not on it, so no android element ever reported the fact and the conservative rule downstream redacted every typed value in the trace, the prompt and the log. The XML still carries it: re-read it once per settled snapshot and state the fact on the text fields it matches. A field it cannot match stays unstated, which still reads as a credential. * docs: correct the record that android never reports a secure field Four places said android reports the fact for nothing and that every typed value there is redacted. The sidecar now states it, so they described the old behaviour. * test(sidecar): fail the build if maestro renames the call the fact comes from * fix(sidecar): state the fact on a field named by its hint alone collectTextFields matched on class only, so a node the go side calls editable off its hintText was left unstated and its typed value redacted. * docs: record that ios and web state secure:false for compose password fields Both derive the fact from a widget type a compose app never has, so the value reaches the trace in the clear. Verified on folio on both targets. * feat(android): read the application id out of an apk parses the compiled AndroidManifest.xml rather than shelling out to aapt2, which lives in the versioned build-tools directory that hosts with only platform-tools never install. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * feat(cli): let --android-app-path supply the bundle id --bundle-id stays required everywhere else, and an explicit one still wins, so the apk can never quietly override what was asked for. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * docs: record that the apk can name the package itself Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning * feat(folio): ask which android device to run on when none is named * feat(folio): pin ios recipes to one simulator udid and ask when several match * docs(ci): say how just ios lands on the simulator the boot step chose * chore(folio): ignore run output anywhere under examples/folio * refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade ScreenName kept its own reading of the route markers and disagreed with Transitional on a marker repeated by a nested node: it named the screen on a step the runner was skipping as unsettled. One reading now. * refactor(sidecar): inline the one attempt passed to callViewHierarchy A named constant and its own comment for a literal used once. * test(sidecar): compare the whole tree when checking the annotation changes nothing else The old assertions checked one id string and one bounds value, and passed with every other attribute stripped off every node. Now the annotated tree minus the two facts it stated must equal the input. * fix(sidecar): match a field to its xml node by class as well as id and bounds A wrapper drawn to the same bounds as the untagged field inside it shared the field's key, both were dropped as ambiguous, and every value typed into an untagged field was redacted. The class tells them apart. * docs(runs): record what the android hierarchy re-read costs per step Two 1m runs per binary on folio, same seed, before and after the re-read.
277 lines
13 KiB
Markdown
277 lines
13 KiB
Markdown
---
|
|
name: sanderling-setup
|
|
description: Get sanderling running against an app that is not folio. Use before writing a spec for a new app, when deciding what test hooks the app needs, and when a run will not start or starts and sees nothing.
|
|
---
|
|
|
|
# Getting sanderling onto your app
|
|
|
|
The goal of setup is not a run that finishes. It is a run whose output you can
|
|
believe. Two things decide that, and both are usually treated as chores: the
|
|
handles the app exposes, and the state the app starts in. Everything else here
|
|
is plumbing.
|
|
|
|
Every flag below is one the binary accepts, checked against `sanderling test -h`
|
|
on this revision. That command is the authority, not this file and not the
|
|
manual. Check before you use a flag you have not seen work.
|
|
|
|
## 1. Install, then check the host
|
|
|
|
The CLI installs from the release script, and the spec package from npm:
|
|
|
|
```sh
|
|
curl -fsSL https://raw.githubusercontent.com/priyanshujain/sanderling/master/install.sh | bash
|
|
npm install --save-dev @sanderling/spec
|
|
```
|
|
|
|
Both come from the same release tag and the CLI bundles the package's TypeScript
|
|
when it evaluates your spec, so they move together.
|
|
|
|
`sanderling doctor` reports the host's readiness per platform and exits non-zero
|
|
if anything is missing. Each line names the check and, on a failure, what to do
|
|
about it. On a Mac with the Android SDK installed but a CLI built by a plain
|
|
`go build`, `sanderling doctor --platform android` says:
|
|
|
|
```
|
|
OK adb on PATH or under the Android SDK
|
|
OK emulator on PATH or under the Android SDK
|
|
OK java 17+ on PATH
|
|
FAIL sidecar JAR is real (not placeholder): placeholder JAR embedded; run `make sidecar && make sanderling` to embed the real fat JAR
|
|
error: 1 check(s) failed
|
|
```
|
|
|
|
Scope it with `--platform web|android|ios|ios-device|all` (default `all`). Web
|
|
needs a Chromium that launches headless. Android needs `adb`, an emulator, Java
|
|
17 or newer, and the embedded sidecar JAR. iOS needs `xcrun` and `simctl`;
|
|
`ios-device` adds `devicectl`, the macOS usbmuxd socket, a connected paired
|
|
device, and App Store Connect signing credentials.
|
|
|
|
The `adb` and `emulator` checks resolve through the same helpers a run uses, so
|
|
they search PATH, then `ANDROID_HOME` and `ANDROID_SDK_ROOT`, then
|
|
`~/Library/Android/sdk`, `~/Android/Sdk` and the Homebrew command-line-tools
|
|
paths. A host the doctor passes is a host a run can drive, and a failure names
|
|
every location it tried. What the doctor cannot tell you is the reverse: a
|
|
missing SDK can also surface during a run as `sidecar health check: context
|
|
deadline exceeded` about thirty seconds in, which names the symptom and not the
|
|
cause (issue #69). If you see it, go back to
|
|
`sanderling doctor --platform android` before believing anything about the
|
|
sidecar.
|
|
|
|
Two traps if you build from source rather than installing a release. A plain
|
|
`go build ./cmd/sanderling` embeds a placeholder sidecar JAR, so every Android
|
|
run stops at `sidecar: binary built without -tags withsidecar`; `make sanderling`
|
|
(or `make sanderling-android`) embeds the real one. And `go run ./cmd/sanderling
|
|
test` collapses the process exit code: a run that exits 2 comes back from
|
|
`go run` as 1 with `exit status 2` printed. Use the built binary whenever the
|
|
exit code matters, which is always in CI.
|
|
|
|
## 2. Point it at the app
|
|
|
|
Android takes the applicationId, boots an AVD with `--avd`, and picks between
|
|
attached devices with `--device <serial>` as `adb devices` prints it:
|
|
|
|
```sh
|
|
sanderling test --spec spec.ts --bundle-id com.example.app --avd Pixel_7_API_34
|
|
```
|
|
|
|
`--android-app-path <apk>` stands in for `--bundle-id`: the applicationId is
|
|
read out of the APK's compiled manifest, so the id and the file it came from
|
|
cannot disagree.
|
|
|
|
iOS takes `--platform ios` and `--ios-device`, which accepts a simulator name or
|
|
UDID, or a connected device's name, UDID, or CoreDevice id. `--ios-app-path`
|
|
points at the `.app` bundle and is what makes clear-state real; see section 4.
|
|
|
|
Web takes a URL as the bundle id:
|
|
|
|
```sh
|
|
sanderling test --spec spec.ts --platform web --bundle-id http://127.0.0.1:8799/index.html
|
|
```
|
|
|
|
The web target has to genuinely load. A page that boots to a blank canvas still
|
|
produces steps, still exits 0, and proves nothing: folio's own web leg needs
|
|
COOP/COEP headers or its sqlite worker never starts, which is why
|
|
`.github/scripts/folio-run.sh` serves the build itself instead of using a stock
|
|
static server. Confirm the app rendered before you read anything else.
|
|
|
|
## 3. Test hooks are a prerequisite, not a polish step
|
|
|
|
This is the part that decides whether a spec is possible at all. The header of
|
|
`replay-ui/sanderling/spec.ts` states it as the lesson it is:
|
|
|
|
> The hooks it drives (data-testid, data-step, ...) were added to the UI for
|
|
> this spec. Needing them is the lesson: a UI with no stable handles is a UI
|
|
> nothing can assert on, and that is as true for a person writing a test as it
|
|
> is for a fuzzer.
|
|
|
|
A fuzzer is not asking for anything a human test author does not need. It is
|
|
only less able to squint at a screenshot and guess. Budget the hooks as part of
|
|
adopting sanderling, before the spec, not after the first vacuous run.
|
|
|
|
`testTag` is the portable name. `internal/hierarchy/hierarchy.go` aliases it to
|
|
`resource-id`, `identifier` and `accessibilityIdentifier`, so one selector
|
|
matches on every platform. What you have to add differs:
|
|
|
|
**Compose on Android.** `Modifier.testTag("AddAccountSubmit")` alone does not
|
|
reach the accessibility tree. The tree only carries it when a root composable
|
|
sets `semantics { testTagsAsResourceId = true }`. folio does this once, at the
|
|
app root, through an expect/actual bridge:
|
|
`examples/folio/app/shared/src/androidMain/kotlin/app/folio/ui/TestTagBridge.android.kt`.
|
|
Without it every `testTag` selector matches nothing, every property over it
|
|
declines, and the run goes green having checked nothing.
|
|
|
|
**Web.** `data-testid` is the hook. Every `data-*` attribute on the element
|
|
reaches the spec under `attrs`, camel-cased the way `dataset` does it, so
|
|
`data-step-count` reads as `attrs.stepCount`. That is how the replay-ui spec
|
|
reads a panel's own claim about which step it is showing rather than re-deriving
|
|
it. Hooks that carry a value, not just an identity, are what make cross-panel
|
|
agreement properties possible.
|
|
|
|
**iOS.** `accessibilityIdentifier`, set via `.accessibilityIdentifier` in
|
|
SwiftUI or UIKit. Compose Multiplatform maps `testTag` to it for you.
|
|
|
|
Two rules about the names themselves. On Android and iOS a `testTag` selector
|
|
falls through to a substring compare, so `{testTag: "Sub"}` matches
|
|
`AddAccountSubmit`; on web the same selector compiles to an exact CSS attribute
|
|
match and hits nothing. Make each hook a whole distinct name rather than a
|
|
fragment of another, and you are right on both. And give every screen a marker
|
|
of its own, because a route extractor is what lets a property decline on the
|
|
screens it has nothing to say about.
|
|
|
|
The check that a hook exists is not that you added it. It is that you can point
|
|
at a step in a real trace where a selector over it resolved to a value.
|
|
`sanderling-spec-authoring` covers which hooks a spec needs and in what order to
|
|
add them; this section is about what each platform requires before any of that
|
|
reaches the tree.
|
|
|
|
## 4. A run must start from a known state
|
|
|
|
`--clear-data` defaults to true and is the difference between a repeatable run
|
|
and a measurement of your own leftovers. A second run that inherits the first
|
|
one's accounts, cache and completed onboarding diverges at step 1: the seed
|
|
reproduces nothing, the two runs' step counts are not comparable, and any number
|
|
you quote from the pair is noise.
|
|
|
|
What "clear" reaches depends on the platform, and in two cases it silently
|
|
reaches less than you expect:
|
|
|
|
- Android wipes app data through the sidecar. On OEM builds that deny
|
|
`pm clear`, pass `--android-app-path <apk>` and it uninstalls and reinstalls
|
|
instead.
|
|
- iOS simulator without `--ios-app-path` resets the data container only and
|
|
prints `clear-state requested without an app path: resetting the data
|
|
container only`. With the path it does a full `simctl` uninstall and install.
|
|
The container wipe is a real reset and folio's own iOS leg relies on it; the
|
|
reinstall path is the one that races FrontBoard.
|
|
- iOS on a physical device without `--ios-app-path` does not clear at all. It
|
|
prints `clear-state on a physical device requires --ios-app-path for a
|
|
reinstall; skipping (state not cleared)` and carries on. A device run left on
|
|
the default flag inherits every previous run's data.
|
|
- Web clears cookies and the target origin's storage. It cannot touch your
|
|
backend. If your app's state lives on a server, reset it yourself between
|
|
runs.
|
|
|
|
`--clear-data=false` is a legitimate choice in one situation: you have just
|
|
installed a fresh build, so the app is already in clear state and an in-run
|
|
reinstall would only add a failure mode. Outside that, a run that resumes is a
|
|
run you cannot repeat.
|
|
|
|
## 5. The device does not have to be local
|
|
|
|
Android talks to whatever adb server the environment names.
|
|
`ADB_SERVER_SOCKET=tcp:host:port` (or `tcp:port` for a server on this machine)
|
|
is read first, then the older `ANDROID_ADB_SERVER_ADDRESS` and
|
|
`ANDROID_ADB_SERVER_PORT` pair, then the loopback default. The CLI shells out to
|
|
`adb` and inherits it; the JVM sidecar resolves the same variables when it
|
|
attaches to a serial.
|
|
|
|
Two things to get right. Pass `--device <serial>` exactly as the remote server
|
|
reports it: with no serial the sidecar's target is a local `localhost:5555`, not
|
|
your remote device. And a serial that already looks like `host:port` is dialled
|
|
straight at adbd, bypassing any server, which is a different path with different
|
|
failure modes. A value the sidecar cannot parse fails the run rather than
|
|
falling back to loopback, and that is deliberate: emulator serials are numbered
|
|
per server, so a quiet fallback would drive whatever this machine calls
|
|
`emulator-5554` and report the results as the remote device's.
|
|
|
|
## 6. What a first run prints
|
|
|
|
A ten step web run, in full:
|
|
|
|
```
|
|
bundled spec: 16532 bytes (sha256=71375ed5bfc7)
|
|
bundled web spec: 33131 bytes (sha256=779bae3c8fee)
|
|
spec loaded into verifier
|
|
trace dir: runs/20260815-172356
|
|
running for 1m30s or 10 steps, whichever comes first (seed=7)
|
|
step index=1 screen="/index.html" nodes=6
|
|
...
|
|
step index=10 screen="/index.html" nodes=6
|
|
|
|
elapsed: 1.715s
|
|
|
|
run complete: 10 steps, 10 driven by the generator
|
|
no violations.
|
|
```
|
|
|
|
`nodes=` is the first number to read and the cheapest lie detector you have. On
|
|
that page, four elements plus html and body gave `nodes=6`. The same command
|
|
against an empty page gives `nodes=2` for every step, and the run then refuses
|
|
itself: the generator has nothing to pick, so the summary reads
|
|
`run complete: 10 steps, 0 driven by the generator` and
|
|
`10 action(s) never reached the app: no_action_produced 10`, and the process
|
|
exits 1 having judged one blank screen ten times. If `nodes` is a handful and
|
|
never grows, the run is looking at something that is not your app.
|
|
|
|
`screen=` is web-only and has nothing to do with your spec's screen hooks: the
|
|
Chrome driver puts the URL hash, or the pathname when there is no hash, on the
|
|
root node, and only that driver writes the attribute. On Android and iOS it is
|
|
empty on every step, so an empty `screen=` there is the normal reading and not a
|
|
symptom. On web, a `screen=` that never changes means the run never left one
|
|
URL, which for a single-page app that routes in memory is also normal. Your
|
|
spec's own route extractor is the thing to trust on every platform.
|
|
|
|
The summary can carry a third line you should never skim past:
|
|
|
|
```
|
|
7 step(s) judged by nothing: the screen was still moving when it was read
|
|
```
|
|
|
|
Those steps were recorded but no property judged them, so the run's step count
|
|
and its checked count are different numbers. `sanderling-run-triage` is about
|
|
what to do with that.
|
|
|
|
Set `--max-steps` whenever you intend to compare two runs: a step budget is what
|
|
makes them comparable, since duration alone does not. `--seed` fixes the PRNG,
|
|
and seed 0 draws a random one and records it in `meta.json`.
|
|
|
|
## 7. The run directory
|
|
|
|
Each run writes `<output>/<UTC timestamp>/`, containing `meta.json`,
|
|
`trace.jsonl`, and one PNG per step under `screenshots/`. `--output` defaults to
|
|
`./runs`.
|
|
|
|
`meta.json` is the run's identity: seed, spec path, bundled spec sha256,
|
|
platform, bundle id, start and end times, generator, `max_steps`,
|
|
`duration_millis`, host, and the `--arm` label if you set one. Two runs that
|
|
differ in any of those are different runs and cannot be pooled.
|
|
|
|
If the app never launched, there is no run directory at all: the launch error
|
|
comes before the trace is created. `error: launch app: ...` with nothing under
|
|
`./runs` means the run never began, which is a different thing from a run that
|
|
began and found nothing.
|
|
|
|
Open a run with `sanderling replay <dir>`, which accepts either the parent runs
|
|
directory or a single run directory.
|
|
|
|
## Reporting
|
|
|
|
Say what you actually ran and what came back: the `doctor` output you got rather
|
|
than the one you expected, the exact `sanderling test` command, the step count
|
|
and the `nodes=` figure from the first run, and for each hook you added, the step
|
|
in a real trace where a selector over it resolved. Name what you could not
|
|
establish, particularly any platform you did not run on.
|
|
|
|
Setup is finished when a property can be written that could fail. Write it with
|
|
`sanderling-spec-authoring`, review it with `sanderling-spec-review`, and read
|
|
the run it produces with `sanderling-run-triage`.
|