From e323854fa7416107d5e3a5a11c62603d485faf19 Mon Sep 17 00:00:00 2001 From: PJ Date: Sat, 6 Jun 2026 23:23:57 +0530 Subject: [PATCH] docs(manual): plain-language rewrite of runs page --- docs/manual/runs.md | 49 +++++++++++++++++++++------------------------ 1 file changed, 23 insertions(+), 26 deletions(-) diff --git a/docs/manual/runs.md b/docs/manual/runs.md index 384f30c..eb1b809 100644 --- a/docs/manual/runs.md +++ b/docs/manual/runs.md @@ -2,11 +2,11 @@ title: Runs --- -# What is a run? +# Runs -One `sanderling test` invocation. Fresh install, spec-driven exploration, then the trace lands in `runs//`. Typically minutes to hours, not seconds. +A run is one `sanderling test` invocation: launch the app, explore it under the spec for a fixed duration, write a trace. Runs typically last minutes to hours, not seconds. -A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an hour and see what breaks. Violations are recorded in the trace and exploration continues, so one run can surface many bugs. +A run is not a unit test. The closer picture is: boot a fuzzer for an hour and see what breaks. A violated property is recorded in the trace and exploration continues, so one run can surface many bugs. ## Lifecycle @@ -14,56 +14,53 @@ A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an sanderling test --spec spec.ts --bundle-id com.example.app --duration 30m │ ├── launch the app under test (pass --clear-data to wipe app data first) - ├── boot the sidecar, connect the agent socket - ├── bundle the spec, load it into goja + ├── boot the sidecar (or connect to Chrome on web) + ├── bundle the spec, load it into the JS runtime │ - ├── step 0..N: pause, capture state, evaluate properties, pick action, resume, dispatch + ├── step 0..N: read state, check properties, pick and perform an action │ - └── terminate when --duration elapses (or SIGINT) + └── stop when --duration elapses (or on Ctrl+C) └── trace written to ./runs// ├── trace.jsonl ├── screenshots/ └── meta.json ``` +The trace is written incrementally. An interrupted run is complete up to the step where it stopped. + ## App state across runs -By default the installed app is left in place between runs. Whatever state the previous run left behind (account, cached responses, onboarding completion) carries over. Pass `--clear-data` to wipe app data before launch and start cold every run. See [CLI reference](./cli/#sanderling-test) for the flag. +By default the installed app stays in place between runs, so whatever the previous run left behind (an account, cached responses, completed onboarding) carries over. Pass `--clear-data` to wipe app data before launch and start cold every time. See the [CLI reference](./cli/#sanderling-test). -## Why runs are long and linear +## Why runs are long -sanderling does not restart the app every N steps. Each restart throws away two things. +sanderling does not restart the app every few steps. Restarting throws away two things. -**Novelty and coverage signal.** The exploration strategy weights actions by whether they reach previously unseen state. Restarting resets that history. +**Exploration history.** Action selection favors states the run has not seen. A restart resets that history and the explorer re-treads the same early screens. -**Deep app states.** Many screens take many actions to reach: nested settings, a loaded cart, post-checkout flows. A 50-step prefix to reach "cart with 3 items" does not happen if every run starts cold. +**Deep app states.** Many bugs live in states that take many actions to reach: nested settings, a loaded cart, the screen after the third transaction. A 50-step path to "cart with 3 items" never happens if every run starts cold. -Long-linear trajectories find bugs that restart-based testing structurally cannot. +Long runs reach states that restart-per-test approaches structurally cannot. -## Setup cost amortizes +## Setup cost is paid once -Preconditions (login, onboarding, consent dialogs) are written as weighted action generators gated on extractors. See [writing specs](./writing-specs/#pattern-preconditions-login-onboarding). They fire only when applicable, so login happens once per run, not per step. +Preconditions like login run through the spec's `setup` export (see [writing specs](./writing-specs/#getting-past-login)). They fire when their condition is unmet and go quiet after, so login costs a few seconds once per run, not once per test case. -| Run length | Login cost | % of run | +| Run length | Login cost | Share of run | |---|---|---| | 5 min | ~15s | 5% | | 30 min | ~15s | 0.8% | | 1 hour | ~15s | 0.4% | -| CI: 3 seeds × 10 min | ~45s total | 2.5% | - -At any non-trivial run length, preconditions are a rounding error. ## Session state -Session tokens, keychain, shared prefs, cookies, and other app-managed persistence survive the full run. If the app logs the user out mid-run, the `doLogin` generator re-fires automatically because its gating extractor (`onLoginScreen`) becomes true again. No retry logic. No special-casing. +Session tokens, keychain entries, shared preferences, and cookies survive the whole run. If the app logs the user out mid-run, the gating extractor flips, `setup` re-engages, and the run logs back in. No retry logic needed in the spec. ## Termination -A run ends when either of these happens. +A run ends when: -- `--duration` elapses. -- The process is interrupted (SIGINT). +- `--duration` elapses, or +- the process is interrupted (Ctrl+C). -The trace is written incrementally, so an interrupted run is still fully inspectable. - -Additional termination conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in [v0.1.0](https://github.com/priyanshujain/sanderling/issues/4). +Additional conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in [v0.1.0](https://github.com/priyanshujain/sanderling/issues/4).