Files
sanderling/docs/manual/runs.md
pj f572c8ba66 WIP: docs: refresh after iOS + web support (#50)
* docs: README covers iOS + web, surface both example apps

* docs(cli): document --ios-device and per-platform doctor

* docs: tighten README, fold examples into Docs list

* docs(runs): correct --clear-data lifecycle wording

Default behavior no longer wipes app data between runs; --clear-data is now opt-in.

* docs(getting-started): add iOS path, separate folio and folio-web

Document just test-ios under examples/folio, and distinguish the KMP
sample from the React + Vite folio-web sample.

* docs(inspect): document the eight panels

Lists Screenshot, ActionList, Timeline, ViolationsPanel, HierarchyPanel,
SnapshotTable, MetricsChart, ExceptionsPanel. Cross-links HierarchyPanel
to the spec language reference.

* docs(writing-specs): document setup export, flag noLogcatErrors as android-only

Mirrors pkg/spec/README.md so the manual covers the runner's setup-first
fall-through. Marks noLogcatErrors as Android-only so iOS/web spec
authors know it silently no-ops.

* docs(folio): document web target and iOS sanderling test recipe

After the KMP refactor folio also runs on wasmJs and the justfile exposes
just web, just web-build, and just test-ios. Surface all three.

* docs(folio-web): add README

Covers prerequisites, demo credentials, just test recipe, and how the
React + Vite host exposes state to the sanderling spec via stable ids
and data-* attributes.

* docs: scrub driver-implementation name from user docs

Drop the implementation tool name from README, cli.md doctor table, and
spec-language.md. These docs should describe behaviour, not the specific
underlying tool the native sidecar wraps.

* docs(development): scrub driver-implementation name from dev docs

architecture, design-principles, decisions now describe the native
sidecar by role (gRPC surface over OS UI-test pipeline) rather than by
the specific tool it wraps.
2026-05-25 16:17:21 +05:30

3.1 KiB
Raw Permalink Blame History

title
title
Runs

What is a run?

One sanderling test invocation. Fresh install, spec-driven exploration, then the trace lands in runs/<timestamp>/. Typically minutes to hours, not seconds.

A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an hour and see what breaks. Violations are recorded in the trace and exploration continues, so one run can surface many bugs.

Lifecycle

sanderling test --spec spec.ts --bundle-id com.example.app --duration 30m
  │
  ├── launch the app under test (pass --clear-data to wipe app data first)
  ├── boot the sidecar, connect the agent socket
  ├── bundle the spec, load it into goja
  │
  ├── step 0..N:  pause, capture state, evaluate properties, pick action, resume, dispatch
  │
  └── terminate when --duration elapses (or SIGINT)
        └── trace written to ./runs/<timestamp>/
              ├── trace.jsonl
              ├── screenshots/
              └── meta.json

App state across runs

By default the installed app is left in place between runs. Whatever state the previous run left behind (account, cached responses, onboarding completion) carries over. Pass --clear-data to wipe app data before launch and start cold every run. See CLI reference for the flag.

Why runs are long and linear

sanderling does not restart the app every N steps. Each restart throws away two things.

Novelty and coverage signal. The exploration strategy weights actions by whether they reach previously unseen state. Restarting resets that history.

Deep app states. Many screens take many actions to reach: nested settings, a loaded cart, post-checkout flows. A 50-step prefix to reach "cart with 3 items" does not happen if every run starts cold.

Long-linear trajectories find bugs that restart-based testing structurally cannot.

Setup cost amortizes

Preconditions (login, onboarding, consent dialogs) are written as weighted action generators gated on extractors. See writing specs. They fire only when applicable, so login happens once per run, not per step.

Run length Login cost % of run
5 min ~15s 5%
30 min ~15s 0.8%
1 hour ~15s 0.4%
CI: 3 seeds × 10 min ~45s total 2.5%

At any non-trivial run length, preconditions are a rounding error.

Session state

Session tokens, keychain, shared prefs, cookies, and other app-managed persistence survive the full run. If the app logs the user out mid-run, the doLogin generator re-fires automatically because its gating extractor (onLoginScreen) becomes true again. No retry logic. No special-casing.

Termination

A run ends when either of these happens.

  • --duration elapses.
  • The process is interrupted (SIGINT).

The trace is written incrementally, so an interrupted run is still fully inspectable.

Additional termination conditions (--max-steps, --exit-on-violation, hard crash handling) land in v0.1.0.