mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 11:07:10 +00:00
docs update with case study (#63)
* docs(manual): add introduction page * docs(manual): rewrite getting started as guided first run * docs(manual): rewrite writing specs as a folio tutorial * docs(manual): document missing spec API in reference * docs(manual): plain-language rewrite of runs page * docs: real introductions on index pages and README * fix(docs): sibling links from directory-style pages need ../ * fix(docs): correct sampling and restart-cost claims to match implementation * docs: nav lists Introduction and Case study; roadmap points to milestone * docs(manual): make getting started target the reader's own app, not Folio * docs(manual): add Folio case study page * docs: point manual navigation at the case study * docs(readme): lead with the case study, fix roadmap link * docs: roadmap links to milestone, sync clear-data default and cross-links
This commit is contained in:
13 files changed
+396
-424
No files matched your search
+23
-26
@@ -2,11 +2,11 @@
|
||||
title: Runs
|
||||
---
|
||||
|
||||
# What is a run?
|
||||
# Runs
|
||||
|
||||
One `sanderling test` invocation. Fresh install, spec-driven exploration, then the trace lands in `runs/<timestamp>/`. Typically minutes to hours, not seconds.
|
||||
A run is one `sanderling test` invocation: launch the app, explore it under the spec for a fixed duration, write a trace. Runs typically last minutes to hours, not seconds.
|
||||
|
||||
A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an hour and see what breaks. Violations are recorded in the trace and exploration continues, so one run can surface many bugs.
|
||||
A run is not a unit test. The closer picture is: boot a fuzzer for an hour and see what breaks. A violated property is recorded in the trace and exploration continues, so one run can surface many bugs.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
@@ -14,56 +14,53 @@ A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an
|
||||
sanderling test --spec spec.ts --bundle-id com.example.app --duration 30m
|
||||
│
|
||||
├── launch the app under test (pass --clear-data to wipe app data first)
|
||||
├── boot the sidecar, connect the agent socket
|
||||
├── bundle the spec, load it into goja
|
||||
├── boot the sidecar (or connect to Chrome on web)
|
||||
├── bundle the spec, load it into the JS runtime
|
||||
│
|
||||
├── step 0..N: pause, capture state, evaluate properties, pick action, resume, dispatch
|
||||
├── step 0..N: read state, check properties, pick and perform an action
|
||||
│
|
||||
└── terminate when --duration elapses (or SIGINT)
|
||||
└── stop when --duration elapses (or on Ctrl+C)
|
||||
└── trace written to ./runs/<timestamp>/
|
||||
├── trace.jsonl
|
||||
├── screenshots/
|
||||
└── meta.json
|
||||
```
|
||||
|
||||
The trace is written incrementally. An interrupted run is complete up to the step where it stopped.
|
||||
|
||||
## App state across runs
|
||||
|
||||
By default the installed app is left in place between runs. Whatever state the previous run left behind (account, cached responses, onboarding completion) carries over. Pass `--clear-data` to wipe app data before launch and start cold every run. See [CLI reference](./cli/#sanderling-test) for the flag.
|
||||
By default each run wipes app data before launch and starts cold. Pass `--clear-data=false` to resume whatever the previous run left behind (an account, cached responses, completed onboarding). See the [CLI reference](../cli/#sanderling-test).
|
||||
|
||||
## Why runs are long and linear
|
||||
## Why runs are long
|
||||
|
||||
sanderling does not restart the app every N steps. Each restart throws away two things.
|
||||
sanderling does not restart the app every few steps. Restarting throws away two things.
|
||||
|
||||
**Novelty and coverage signal.** The exploration strategy weights actions by whether they reach previously unseen state. Restarting resets that history.
|
||||
**Accumulated data.** Accounts created, items added, caches warmed, settings changed. Interesting bugs live in apps with history, and a restart wipes it.
|
||||
|
||||
**Deep app states.** Many screens take many actions to reach: nested settings, a loaded cart, post-checkout flows. A 50-step prefix to reach "cart with 3 items" does not happen if every run starts cold.
|
||||
**Deep app states.** Many bugs live in states that take many actions to reach: nested settings, a loaded cart, the screen after the third transaction. A 50-step path to "cart with 3 items" never happens if every run starts cold.
|
||||
|
||||
Long-linear trajectories find bugs that restart-based testing structurally cannot.
|
||||
Long runs reach states that restart-per-test approaches structurally cannot.
|
||||
|
||||
## Setup cost amortizes
|
||||
## Setup cost is paid once
|
||||
|
||||
Preconditions (login, onboarding, consent dialogs) are written as weighted action generators gated on extractors. See [writing specs](./writing-specs/#pattern-preconditions-login-onboarding). They fire only when applicable, so login happens once per run, not per step.
|
||||
Preconditions like login run through the spec's `setup` export (see the [case study](../case-study/#reaching-the-screens-that-matter)). They fire when their condition is unmet and go quiet after, so login costs a few seconds once per run, not once per test case.
|
||||
|
||||
| Run length | Login cost | % of run |
|
||||
| Run length | Login cost | Share of run |
|
||||
|---|---|---|
|
||||
| 5 min | ~15s | 5% |
|
||||
| 30 min | ~15s | 0.8% |
|
||||
| 1 hour | ~15s | 0.4% |
|
||||
| CI: 3 seeds × 10 min | ~45s total | 2.5% |
|
||||
|
||||
At any non-trivial run length, preconditions are a rounding error.
|
||||
|
||||
## Session state
|
||||
|
||||
Session tokens, keychain, shared prefs, cookies, and other app-managed persistence survive the full run. If the app logs the user out mid-run, the `doLogin` generator re-fires automatically because its gating extractor (`onLoginScreen`) becomes true again. No retry logic. No special-casing.
|
||||
Session tokens, keychain entries, shared preferences, and cookies survive the whole run. If the app logs the user out mid-run, the gating extractor flips, `setup` re-engages, and the run logs back in. No retry logic needed in the spec.
|
||||
|
||||
## Termination
|
||||
|
||||
A run ends when either of these happens.
|
||||
A run ends when:
|
||||
|
||||
- `--duration` elapses.
|
||||
- The process is interrupted (SIGINT).
|
||||
- `--duration` elapses, or
|
||||
- the process is interrupted (Ctrl+C).
|
||||
|
||||
The trace is written incrementally, so an interrupted run is still fully inspectable.
|
||||
|
||||
Additional termination conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in [v0.1.0](https://github.com/priyanshujain/sanderling/issues/4).
|
||||
Additional conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in the [v0.1.0 milestone](https://github.com/priyanshujain/sanderling/milestone/1).
|
||||
Reference in new issue
Block a user