docs update with case study (#63)

* docs(manual): add introduction page

* docs(manual): rewrite getting started as guided first run

* docs(manual): rewrite writing specs as a folio tutorial

* docs(manual): document missing spec API in reference

* docs(manual): plain-language rewrite of runs page

* docs: real introductions on index pages and README

* fix(docs): sibling links from directory-style pages need ../

* fix(docs): correct sampling and restart-cost claims to match implementation

* docs: nav lists Introduction and Case study; roadmap points to milestone

* docs(manual): make getting started target the reader's own app, not Folio

* docs(manual): add Folio case study page

* docs: point manual navigation at the case study

* docs(readme): lead with the case study, fix roadmap link

* docs: roadmap links to milestone, sync clear-data default and cross-links
This commit is contained in:
pj authored and GitHub committed 2026-06-09 20:13:54 +05:30
1 parent 90224dfd06
commit 991c583eb9
13 files changed
+396 -424

No files matched your search

+144
View File
@@ -0,0 +1,144 @@
---
title: "Case study: Folio"
---
# Case study: Folio
Folio is a personal-ledger app: log in, create accounts, record credits and debits. It is built once in Kotlin Multiplatform and ships to Android, iOS, and web from a single Compose codebase. The repo carries it under `examples/folio`, with its spec at `examples/folio/sanderling/spec.ts`.
It also carries a bug, the kind that survives manual testing and example-based suites and ships to production. This page follows how sanderling finds it.
## The bug
Folio's add-transaction form saves on submit and navigates home. The submit handler does not disable the button while the save is in flight:
```kotlin
AddTransactionEvent.Submit -> submit()
```
Tap submit twice fast and two transactions post. The balance moves by twice what you typed.
Nobody writes this test. A manual tester taps submit once, sees the right number, moves on. A scripted test encodes the same single tap. The bug lives in a sequence no one thought to script: the double tap, on this screen, with a pending save. That is the class of bug sanderling exists to catch.
## Telling sanderling what is true
You do not script the double tap. You state the invariant and let sanderling find the inputs that break it.
The amount the user types must equal the amount the balance moves:
```ts
const submitMovesBalanceByTypedAmount = always(
next(() => {
if (route.current !== "home") return true;
const action = lastAction.current;
if (action?.kind !== "Tap" && action?.kind !== "DoubleTap") return true;
if (!JSON.stringify(action.on ?? "").includes("TxnSubmit")) return true;
const typed = parseTypedAmount(txnAmountField.previous?.text);
if (typed === 0) return true;
return Math.abs(totalBalance.current - (totalBalance.previous ?? 0)) === typed;
})
);
```
`always` checks the formula at every step; `next` lets it compare the step before a submit to the step after. The guards narrow it to the one transition that matters, a submit that lands back on home, and the last line states the rule: the balance moved by exactly the typed amount. Double-submit moves it by twice that, and the formula is false.
The values it reads come from extractors, which pull state out of the UI tree once per step:
```ts
const route = extract<string | null>("route", s => {
if (s.ax.find({ testTag: "AddTransactionScreen" })) return "add-transaction";
if (s.ax.find({ testTag: "HomeScreen" })) return "home";
// ...other screens
return null;
});
const totalBalance = extract("totalBalance", s =>
s.ax.findAll([{ testTag: "HomeScreen" }, { testTag: "AccountCard" }])
.reduce((sum, c) => sum + parseDollarCents(c.find({ testTag: "AccountBalance" })?.text), 0));
```
Every Folio screen and control carries a `testTag`. Compose exposes it as the resource-id on Android and the accessibility identifier on iOS, so one selector resolves on both. `extract` runs against the live tree each step; properties and actions read `.current` and `.previous`, never the raw state.
A second property states what new accounts must look like: a freshly created account starts at zero.
```ts
const newAccountBalanceIsZero = always(
next(() => {
const before = new Set((accounts.previous ?? []).map(a => a.name));
return accounts.current
.filter(a => !before.has(a.name))
.every(a => a.balance === 0);
})
);
```
## Reaching the screens that matter
A fuzzer that pokes at random never logs in, and never reaches a transaction form. The spec gives sanderling enough to drive the real flows, no more.
Login is a precondition, exported as `setup`. The runner runs it before anything else and falls through to the main pool once it yields nothing:
```ts
const login = actions(() => {
if (loggedIn.current) return [];
const email = loginEmailField.current;
if (email && !email.text) return [InputText({ into: email, text: DEMO_EMAIL })];
const pwd = loginPasswordField.current;
if (pwd && !pwd.text) return [InputText({ into: pwd, text: DEMO_PASSWORD })];
const submit = loginSubmit.current;
return submit ? [Tap({ on: submit })] : [];
});
export const setup = login;
```
It reads the form and offers the next move, never tracking what it did before. If a stray tap logs the user out mid-run, `loggedIn` flips and login re-engages on its own.
The transaction flow spans three screens. `whenRoute` keeps it eligible only where it applies, and returns every reasonable next action, so sometimes the form is filled and submitted, sometimes submitted empty:
```ts
const addTxn = whenRoute(route, ["home", "ledger", "add-transaction"], () => {
if (route.current === "home") {
const cards = accountCards.current;
return cards.length ? [Tap({ on: from(cards).generate() })] : [];
}
if (route.current === "ledger") {
const btn = addTxnButton.current;
return btn ? [Tap({ on: btn })] : [];
}
const field = txnAmountField.current, submit = txnSubmit.current;
const out = [];
if (field) out.push(InputText({ into: field, text: String(amounts.generate()) }));
if (submit) out.push(Tap({ on: submit }));
return out;
});
```
The action pool weights the flows by how much they are worth testing. The transaction chain dominates, account creation stays in the mix so new accounts keep appearing, and `defaultActions` keeps a quarter of the budget on untargeted exploration: random taps, edge-case text in every field, scrolls, swipes, double taps.
```ts
export const actionsRoot = weighted(
[45, addTxn],
[25, addAccount],
[25, defaultActions],
[5, doubleTaps],
);
```
That is the whole input. Two invariants, a way in, and a weighted sense of where to spend time. Nothing here names the bug.
## What the run does
sanderling launches Folio, logs in, and starts exploring. Most steps are unremarkable: open an account, add a transaction, watch the balance move by exactly what was typed, `submitMovesBalanceByTypedAmount` holds.
Then a step lands two taps on submit before the first save settles. Two transactions post. The balance jumps by twice the typed amount. At that step the formula evaluates false and the run records a violation: the step, the screenshot, the offending action, and the residual formula that failed.
The run does not stop. It keeps exploring and keeps checking, so one run surfaces every violation it can reach, not just the first. Open the trace with [`sanderling replay`](../replay/), press `.` to jump to the violation, and step across the boundary to watch the balance double.
sanderling did not know about this bug. It was given what the app guarantees and a realistic way to use it, and it found the sequence that breaks the guarantee. That is the point: you describe what must always be true, and the explorer finds what you would not have thought to try.
## From here
The complete spec is `examples/folio/sanderling/spec.ts`. Folio uses [`just`](https://github.com/casey/just) as a runner: from `examples/folio`, `just install` then `just test` drives it on an Android emulator or device, `just test-ios` on an iOS simulator, and `just web` in Chrome. Then `sanderling replay` and press `.` to land on the violation.
To point sanderling at your own app, see [getting started](../getting-started/). For every selector, operator, action, and sampler, see the [spec language reference](../spec-language/).
+2 -2
View File
@@ -36,7 +36,7 @@ Serve a local web UI for browsing traces. The positional argument is optional an
| `--no-open` | `false` | Skip opening the default browser on startup. |
| `--dev` | `false` | Reverse-proxy non-API requests to the Vite dev server on `127.0.0.1:5173`. |
See [the replay UI page](./replay/) for the panel reference and keyboard shortcuts.
See [the replay UI page](../replay/) for the panel reference and keyboard shortcuts.
## `sanderling doctor`
@@ -65,4 +65,4 @@ Print the CLI version.
- `--max-steps` hard cap on step count.
- `--exit-on-violation` stop the run on the first property violation.
Tracked in [issue #4](https://github.com/priyanshujain/sanderling/issues/4).
Tracked in the [v0.1.0 milestone](https://github.com/priyanshujain/sanderling/milestone/1).
+33 -80
View File
@@ -4,109 +4,62 @@ title: Getting started
# Getting started
Install the CLI, run a spec.
## Prerequisites
**Android / iOS:**
- An Android emulator with API level 30 or newer (or a connected device).
- `adb` on your PATH.
**Web:**
- Chrome installed. sanderling drives it via CDP; no other setup required.
Run `sanderling doctor` to check the host environment.
Install sanderling, write a spec for your app, run it, and open the trace.
## Install
### CLI
The CLI:
```sh
curl -fsSL https://raw.githubusercontent.com/priyanshujain/sanderling/master/install.sh | bash
```
### Spec package ([npm](https://www.npmjs.com/package/@sanderling/spec))
The spec package, in your project:
```sh
npm install --save-dev @sanderling/spec
```
## Your first run
The repo ships two sample apps. `examples/folio` is a Kotlin Multiplatform personal-ledger app that covers Android, iOS, and web (wasmJs) from one shared codebase. `examples/folio-web` is a smaller React + Vite app that covers only the web path. Both carry a TypeScript spec under `sanderling/spec.ts`. Install `just`, then pick a target below.
### Android
From `examples/folio`:
## Check your environment
```sh
just install # build and install the folio APK on a booted emulator or device
just test # run the spec
sanderling doctor
```
With no device connected and multiple AVDs, pick one:
`doctor` reports what the target platform needs and what is missing:
- **Android**: `adb` on your PATH, and an emulator (API 30 or newer) or a connected device.
- **iOS**: Xcode 16 or newer, with a simulator. For a connected iPhone, run `sanderling doctor --platform ios-device`.
- **Web**: Chrome.
## Write a spec
A spec exports two things: `properties` that must always hold, and `actionsRoot`, the actions sanderling may take. The smallest spec that does something useful imports both from the defaults:
```ts
import { defaultActions } from "@sanderling/spec/defaults";
import { noUncaughtExceptions } from "@sanderling/spec/defaults/properties";
export const properties = { noUncaughtExceptions };
export const actionsRoot = defaultActions;
```
This taps, types, scrolls, and swipes at random, and fails the moment your app throws an uncaught exception. From here you add extractors to read your screens, properties that state what your app guarantees, and actions that drive its real flows. The [case study](../case-study/) walks a complete spec, and the [spec language reference](../spec-language/) lists every primitive.
## Run it
Point sanderling at your app:
```sh
AVD=Pixel_7 just test
sanderling test --spec spec.ts --bundle-id com.example.app --platform android
```
Persistent settings can live in a `.env` alongside the justfile (`AVD=Pixel_7`, `DURATION=5m`, and so on).
Use `--platform ios` or `--platform web` for the other targets. By default a run lasts five minutes and starts from a fresh install; `--clear-data=false` resumes prior state, and `--duration 30m` runs longer. The [CLI reference](../cli/) lists every flag.
### iOS
From `examples/folio` (requires Xcode 16+ and `xcodegen`):
## See what it found
```sh
just test-ios # default simulator: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just test-ios # pick a different simulator
sanderling replay
```
`just test-ios` boots the simulator if needed, builds and installs the app, then runs `sanderling test --platform ios`.
#### Physical device
A connected iPhone is driven over a usbmux tunnel by a runner the driver builds and signs at run time. The tunnel talks to macOS's own `usbmuxd`, so nothing extra is installed beyond Xcode. It needs App Store Connect signing credentials in the environment (a gitignored `.env` is loaded by `just`):
```sh
SANDERLING_IOS_TEAM=<10-char team id>
ASC_API_KEY_ID=<key id>
ASC_API_ISSUER_ID=<issuer id>
ASC_API_KEY_PATH=<absolute path to AuthKey_*.p8>
IOS_DEVICE="iPhone" just test-ios-device # name, UDID, or CoreDevice id
```
Run `sanderling doctor --platform ios-device` to check `devicectl`, the `usbmuxd` socket, a connected and paired device, and the signing credentials before a run.
### Web
From either example. For the KMP wasmJs build, use `examples/folio`:
```sh
just web # serve the wasmJs app on a webpack dev server
```
For the React + Vite build, use `examples/folio-web`:
```sh
just test # starts the Vite dev server, then sanderling drives Chrome via CDP
```
No emulator or SDK setup needed for either web path.
### Trace output
When the run ends, the trace lands in `sanderling/runs/<timestamp>/`:
```
runs/2026-04-18T12-34-56/
├── trace.jsonl
├── screenshots/
└── meta.json
```
Browse it with `sanderling replay` (see [replay](./replay/)), or read `trace.jsonl` step by step.
Next: [writing specs](./writing-specs/).
This opens the trace in a local web UI. Step through with `j` and `k`; press `.` to jump to a property violation and see the screenshot, the action, and the failed formula at that step. The [replay page](../replay/) covers the panels and shortcuts.
+7 -5
View File
@@ -4,8 +4,10 @@ title: Manual
# Manual
- [Getting started](./getting-started/)
- [Writing specs](./writing-specs/)
- [Spec language reference](./spec-language/)
- [Runs](./runs/)
- [CLI reference](./cli/)
- [Introduction](./introduction/): what property-based testing is and how sanderling works.
- [Case study: Folio](./case-study/): sanderling finding a real bug in a mobile app, and how the spec is written.
- [Getting started](./getting-started/): install the CLI and run it against Folio.
- [Spec language reference](./spec-language/): every selector, operator, action, and sampler.
- [Runs](./runs/): what happens during a run and why runs are long.
- [Replay](./replay/): the trace browser.
- [CLI reference](./cli/): every command and flag.
+90
View File
@@ -0,0 +1,90 @@
---
title: Introduction
---
# Introduction
sanderling tests mobile and web apps by exploring them on its own and checking rules you write. You describe what must always be true about your app. sanderling drives the app for minutes or hours, performing thousands of taps, swipes, and text inputs, and records every moment a rule breaks.
This is property-based testing, applied to UIs.
## What is property-based testing?
Most UI tests are example-based. You script one path through the app and assert what happens:
1. Open the login screen.
2. Type `[email protected]` and the password.
3. Tap Sign in.
4. Assert the home screen appears.
This proves one path works. It says nothing about the paths you did not script. What happens when a user taps Submit twice in a row? Navigates away mid-form and comes back? Types `999999999` into the amount field? Each scripted test covers exactly one sequence of inputs, so each weird sequence needs its own test. Nobody writes them all, and the bugs live in the ones nobody wrote.
Property-based testing inverts this. Instead of scripting paths, you state rules that must hold on every path:
- An account never shows a negative balance.
- A new account always starts at zero.
- The app never throws an uncaught exception.
These rules are called properties. sanderling explores the app at random, guided by weights you choose, and checks every property at every step. One property covers every path the explorer finds, including the ones you did not think of.
## Why sanderling?
Scripted UI suites are expensive to write and expensive to maintain. Every flow needs its own script, every screen change breaks scripts, and the suite still only covers the flows someone wrote down.
With sanderling you write one spec per app. The spec is small: a handful of rules and a description of what the explorer is allowed to do. The explorer does the rest. A 30-minute run executes thousands of steps through combinations of screens, inputs, and timings that no human would script.
The same spec runs against Android, iOS, and web builds of the same app. Element lookups resolve across platforms, so a spec written once tests all three.
## How it works
You give sanderling two things: an app and a spec.
The spec is a TypeScript file with two exports:
- `properties`: named rules that must hold. For example, "a new account starts with a zero balance".
- `actionsRoot`: a weighted tree of actions sanderling may take. For example, "mostly add transactions, sometimes create accounts, occasionally tap things at random".
sanderling launches the app and runs a loop. Each pass through the loop is one step:
1. **Read the screen.** Capture the UI tree (every element with its text, position, and state), plus logs and exceptions since the last step.
2. **Check every property** against the new state. A broken property is recorded as a violation and the run continues, so one run can surface many bugs.
3. **Pick one action** from the weighted tree and perform it: a tap, a swipe, typed text, a key press.
```
┌────────────────────────────────┐
│ read screen state │
│ (UI tree, logs, exceptions) │
└───────────────┬────────────────┘
▼
┌────────────────────────────────┐
│ check every property │──▶ record violations
└───────────────┬────────────────┘
▼
┌────────────────────────────────┐
│ pick an action by weight, │
│ perform it │
└───────────────┬────────────────┘
▼
repeat
```
The loop runs until the duration you set elapses. Every step is written to a trace: one JSON line and one screenshot per step. After the run, `sanderling replay` opens the trace in a local web UI where you can walk through the run step by step, see exactly what the screen showed when a property broke, and which action caused it.
## What it runs on
| Platform | How sanderling drives the app |
|---|---|
| Android | A native sidecar reads the UI through UIAutomator and injects input. Works on emulators and devices. |
| iOS | A native sidecar reads the UI through XCTest and injects input. Works on simulators. |
| Web | Chrome, driven directly over the Chrome DevTools Protocol. No sidecar. |
Kotlin Multiplatform apps need nothing special: the Android build is tested through the Android driver and the iOS build through the iOS driver.
## Reading this manual
- [Case study: Folio](../case-study/) follows sanderling finding a real bug in the example app, and shows how its spec is written.
- [Getting started](../getting-started/) installs the CLI and runs it against Folio.
- [Spec language reference](../spec-language/) lists every selector, operator, action, and sampler.
- [Runs](../runs/) explains what happens during a run and why runs are long.
- [Replay](../replay/) covers the trace browser.
- [CLI reference](../cli/) lists every command and flag.
+1 -1
View File
@@ -22,7 +22,7 @@ The positional argument can be a runs directory or a single run directory (auto-
| ActionList | Ordered list of steps with the verb and target for each action. |
| Timeline | Per-property lane chart over the run; cells coloured by violated, pending, or holds. |
| ViolationsPanel | Property statuses at the focused step with residual formulas for any that failed. |
| HierarchyPanel | Filterable table of every UI element captured at the focused step; mirrors the selectors documented in [Spec language reference](./spec-language/). |
| HierarchyPanel | Filterable table of every UI element captured at the focused step; mirrors the selectors documented in [Spec language reference](../spec-language/). |
| SnapshotTable | Flattened `snapshots` map for the focused step, with diffs against the previous step. |
| MetricsChart | Heap and other host-side metrics sampled per step. |
| ExceptionsPanel | Uncaught exceptions captured during the run, with a jump-to-first control. |
+23 -26
View File
@@ -2,11 +2,11 @@
title: Runs
---
# What is a run?
# Runs
One `sanderling test` invocation. Fresh install, spec-driven exploration, then the trace lands in `runs/<timestamp>/`. Typically minutes to hours, not seconds.
A run is one `sanderling test` invocation: launch the app, explore it under the spec for a fixed duration, write a trace. Runs typically last minutes to hours, not seconds.
A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an hour and see what breaks. Violations are recorded in the trace and exploration continues, so one run can surface many bugs.
A run is not a unit test. The closer picture is: boot a fuzzer for an hour and see what breaks. A violated property is recorded in the trace and exploration continues, so one run can surface many bugs.
## Lifecycle
@@ -14,56 +14,53 @@ A run is not analogous to a unit test. A closer framing is: boot a fuzzer for an
sanderling test --spec spec.ts --bundle-id com.example.app --duration 30m
│
├── launch the app under test (pass --clear-data to wipe app data first)
├── boot the sidecar, connect the agent socket
├── bundle the spec, load it into goja
├── boot the sidecar (or connect to Chrome on web)
├── bundle the spec, load it into the JS runtime
│
├── step 0..N: pause, capture state, evaluate properties, pick action, resume, dispatch
├── step 0..N: read state, check properties, pick and perform an action
│
└── terminate when --duration elapses (or SIGINT)
└── stop when --duration elapses (or on Ctrl+C)
└── trace written to ./runs/<timestamp>/
├── trace.jsonl
├── screenshots/
└── meta.json
```
The trace is written incrementally. An interrupted run is complete up to the step where it stopped.
## App state across runs
By default the installed app is left in place between runs. Whatever state the previous run left behind (account, cached responses, onboarding completion) carries over. Pass `--clear-data` to wipe app data before launch and start cold every run. See [CLI reference](./cli/#sanderling-test) for the flag.
By default each run wipes app data before launch and starts cold. Pass `--clear-data=false` to resume whatever the previous run left behind (an account, cached responses, completed onboarding). See the [CLI reference](../cli/#sanderling-test).
## Why runs are long and linear
## Why runs are long
sanderling does not restart the app every N steps. Each restart throws away two things.
sanderling does not restart the app every few steps. Restarting throws away two things.
**Novelty and coverage signal.** The exploration strategy weights actions by whether they reach previously unseen state. Restarting resets that history.
**Accumulated data.** Accounts created, items added, caches warmed, settings changed. Interesting bugs live in apps with history, and a restart wipes it.
**Deep app states.** Many screens take many actions to reach: nested settings, a loaded cart, post-checkout flows. A 50-step prefix to reach "cart with 3 items" does not happen if every run starts cold.
**Deep app states.** Many bugs live in states that take many actions to reach: nested settings, a loaded cart, the screen after the third transaction. A 50-step path to "cart with 3 items" never happens if every run starts cold.
Long-linear trajectories find bugs that restart-based testing structurally cannot.
Long runs reach states that restart-per-test approaches structurally cannot.
## Setup cost amortizes
## Setup cost is paid once
Preconditions (login, onboarding, consent dialogs) are written as weighted action generators gated on extractors. See [writing specs](./writing-specs/#pattern-preconditions-login-onboarding). They fire only when applicable, so login happens once per run, not per step.
Preconditions like login run through the spec's `setup` export (see the [case study](../case-study/#reaching-the-screens-that-matter)). They fire when their condition is unmet and go quiet after, so login costs a few seconds once per run, not once per test case.
| Run length | Login cost | % of run |
| Run length | Login cost | Share of run |
|---|---|---|
| 5 min | ~15s | 5% |
| 30 min | ~15s | 0.8% |
| 1 hour | ~15s | 0.4% |
| CI: 3 seeds × 10 min | ~45s total | 2.5% |
At any non-trivial run length, preconditions are a rounding error.
## Session state
Session tokens, keychain, shared prefs, cookies, and other app-managed persistence survive the full run. If the app logs the user out mid-run, the `doLogin` generator re-fires automatically because its gating extractor (`onLoginScreen`) becomes true again. No retry logic. No special-casing.
Session tokens, keychain entries, shared preferences, and cookies survive the whole run. If the app logs the user out mid-run, the gating extractor flips, `setup` re-engages, and the run logs back in. No retry logic needed in the spec.
## Termination
A run ends when either of these happens.
A run ends when:
- `--duration` elapses.
- The process is interrupted (SIGINT).
- `--duration` elapses, or
- the process is interrupted (Ctrl+C).
The trace is written incrementally, so an interrupted run is still fully inspectable.
Additional termination conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in [v0.1.0](https://github.com/priyanshujain/sanderling/issues/4).
Additional conditions (`--max-steps`, `--exit-on-violation`, hard crash handling) land in the [v0.1.0 milestone](https://github.com/priyanshujain/sanderling/milestone/1).
+53 -11
View File
@@ -4,20 +4,22 @@ title: Spec language reference
# Spec language reference
Lookup reference for everything importable from `@sanderling/spec`. For a worked example, read the [case study](../case-study/) first.
## Module structure
A spec is a TypeScript module evaluated by the Go runner each step. It must export `properties` and `actionsRoot` on `globalThis` (the bundler entry point does this automatically via the final two lines):
A spec is a TypeScript module evaluated by the Go runner each step. It exports `properties` and `actionsRoot`, plus an optional `setup`:
```ts
import { ... } from "@sanderling/spec";
export const properties = { ... };
export const actionsRoot = weighted(...);
(globalThis as { actions?: unknown }).actions = actionsRoot;
(globalThis as { properties?: unknown }).properties = properties;
export const setup = login; // optional
```
`setup` is an `ActionGenerator` the runner consults before `actionsRoot` each step. While it returns actions, they run; when it returns an empty list, the runner falls through to `actionsRoot`. Use it for preconditions like login and onboarding. If the app later regresses across the precondition (a logout mid-run), `setup` re-engages on its own.
## State
Every extractor callback receives a `State`:
@@ -71,9 +73,16 @@ Every key-value pair must match. Substring and boolean rules apply per attribute
Known attribute names are typed; you get autocomplete on `testTag`, `text`, `content-desc`, the boolean states (`clickable`, `enabled`, `focused`, `checked`, `selected`), and the cross-platform aliases (`identifier`, `accessibilityIdentifier`, `accessibilityText`, `accessibilityLabel`, `label`, `resource-id`, `class`, `elementType`, `package`, `placeholderValue`, `hintText`). Boolean state attributes accept a native `true` / `false`. Other attribute keys still type-check as a string-valued fallback so raw driver attributes remain reachable.
### Path queries
### Path selectors
Chains of string selectors separated by ` > ` scope each segment to the subtree of the previous match. Path queries are only supported on the tree root (`ax.find`, `ax.findAll`), not on element-scoped `.find`/`.findAll`.
An array of object selectors matches a path: each segment is matched within the subtree of the previous match. Arrays work on the tree root and on element-scoped `.find`/`.findAll`.
```ts
s.ax.find([{ testTag: "LoginScreen" }, { testTag: "LoginEmail" }])
s.ax.findAll([{ testTag: "HomeScreen" }, { testTag: "AccountCard" }])
```
String selectors chain the same way with ` > `, but only on the tree root (`ax.find`, `ax.findAll`):
```ts
s.ax.find("id:HomeScreen > descPrefix:account_card:")
@@ -144,11 +153,13 @@ KMP apps are tested identically to native apps. An Android KMP build uses the An
```ts
const loggedIn = extract((s) => !!s.ax.find("id:home-tab-bar"));
const route = extract("route", (s) => ...); // named form
loggedIn.current // T - value from the current step
loggedIn.previous // T | undefined - value from the previous step, undefined on first step
```
Extractors are evaluated before properties and action generators. Use `.previous` to detect transitions between steps.
Extractors are evaluated before properties and action generators. Use `.previous` to detect transitions between steps. Named extractors appear by name in the replay UI and trace.
## LTL operators
@@ -174,8 +185,11 @@ Extractors are evaluated before properties and action generators. Use `.previous
```ts
Tap({ on: element | string })
DoubleTap({ on: element | string })
LongPress({ on: element | string })
InputText({ into: element | string, text: string })
Swipe({ from: element | Point, to: element | Point, durationMillis?: number })
Scroll({ direction: "up" | "down" | "left" | "right", in?: element | string })
PressKey({ key: Key })
Wait({ durationMillis: number })
```
@@ -189,6 +203,10 @@ On web, `"back"` maps to Backspace and `"home"` is not supported. All other keys
| Generator | Behaviour |
|---|---|
| `taps` | Random tap on a clickable element |
| `doubleTaps` | Random double tap on a clickable element |
| `longPresses` | Random long press on a clickable element |
| `typing` | Types a value from the edge-case corpus into a random editable field |
| `scrolls` | Random scroll gesture |
| `swipes` | Random swipe gesture |
| `waitOnce` | Idles one step |
| `pressKeys` | Presses a random supported key |
@@ -217,23 +235,47 @@ export const actionsRoot = weighted(
);
```
### `from(items)`
### `whenRoute(routeExtractor, routes, body)`
Returns a `Sampler<T>` that cycles through a fixed list. Use `.generate()` to pick an item.
Builds a generator that runs `body` only when the extractor's current value is in `routes` (a string or array of strings). Returns an empty list otherwise.
```ts
const addTxn = whenRoute(route, ["home", "ledger", "add-transaction"], () => {
...
return [Tap({ on: btn })];
});
```
### Samplers
Every sampler has `.generate()`. Draws are seeded by the run's PRNG, so a run replays identically from its seed.
| Sampler | Produces |
|---|---|
| `from(items)` | An item from a fixed list |
| `integers().between(min, max)` | An integer in the range |
| `strings().length(min, max).alpha()` | A random string; `.alpha()` restricts to letters |
| `emails().domain("example.com")` | A random email address |
| `edgeCaseText()` | A value from the adversarial input corpus (empty and whitespace strings, emoji, numeric boundary values, very long strings, injection payloads) |
```ts
const names = from(["Checking", "Savings", "Travel"]);
const amounts = integers().between(1, 500);
// inside an actions() callback:
InputText({ into: nameField, text: names.generate() })
InputText({ into: amountField, text: String(amounts.generate()) })
```
## Default properties
## Defaults
```ts
import { defaultActions, doubleTaps } from "@sanderling/spec/defaults";
import { noUncaughtExceptions, noLogcatErrors } from "@sanderling/spec/defaults/properties";
```
`defaultActions` is a ready-made weighted tree of the built-in generators: taps and typing at weight 100, scrolls 50, swipes 25, double taps 10. Use it as a baseline pool or as one entry in your own tree.
| Property | Fails when |
|---|---|
| `noUncaughtExceptions` | An uncaught exception or `Sanderling.reportError()` call is captured |
| `noLogcatErrors` | Logcat emits any error-level (`E`) lines since the previous step |
| `noLogcatErrors` | Logcat emits any error-level (`E`) lines since the previous step (Android only; holds elsewhere) |
-283
View File
@@ -1,283 +0,0 @@
---
title: Writing specs
---
# Writing specs
A spec has three parts: extractors, properties, and actions.
```ts
import { extract, always, now, actions, weighted, Tap, taps, swipes } from "@sanderling/spec";
const loggedIn = extract((s) => !!s.ax.find("id:home-tab-bar"));
export const properties = {
cartNeverNegative: always(() => cartCount.current >= 0),
};
export const actionsRoot = weighted(
[10, taps],
[2, swipes],
);
```
The Go runner calls into the JS runtime each step. Extractors re-read the current state. Properties re-evaluate with their residual formulas. The action generator returns a tree, and one leaf is sampled by weight and dispatched.
## The `State` object
What extractors receive:
```ts
interface State {
ax: AccessibilityTree;
snapshots: Record<string, unknown>;
lastAction: Action | null;
logs: readonly LogEntry[];
exceptions: readonly ExceptionRecord[];
time: number; // ms since run start
}
```
`ax` is the live UI hierarchy. `snapshots` carries any key-value data pushed by the app SDK. `logs` and `exceptions` contain entries collected since the previous step.
## Extractors
`extract()` wraps a getter that runs against every new state. The returned object exposes `.current` (this step's value) and `.previous` (last step's, or `undefined` on the first step).
```ts
const loggedIn = extract((s) => !!s.ax.find("id:home-tab-bar"));
const balance = extract<number>((s) => s.snapshots["account.balance"] as number ?? 0);
// Inside a property or action:
loggedIn.current // boolean
loggedIn.previous // boolean | undefined
```
Extractors are cheap. Prefer one extractor per concept and reuse it across properties and action generators.
## Finding elements
`ax.find(selector)` returns the first matching `AccessibilityElement`, or `undefined`. `ax.findAll(selector)` returns all matches. Both are available on the tree root and on any element (scoped to its subtree).
**String selectors:**
| Form | Match rule |
|---|---|
| `id:<value>` | Exact match on resource-id, or suffix after `:id/` (Android) |
| `text:<value>` | Substring match on text content |
| `desc:<value>` | Exact match on accessibility description, or starts-with for iOS merged labels |
| `descPrefix:<prefix>` | Starts-with on accessibility description |
| `<attr>:<value>` | Substring match on any raw attribute by name |
**Object selectors** (AND of all given attributes):
```ts
s.ax.find({ accessibilityText: "LoginScreen" })
s.ax.find({ accessibilityText: "login_email" })
```
**Path queries** (global only):
```ts
s.ax.find("id:HomeScreen > descPrefix:account_card:")
```
Each segment is matched within the subtree of the previous match.
**Cross-platform aliases** are resolved automatically. `label` and `accessibilityLabel` both resolve to `accessibilityText`; `content-desc` and `accessibilityText` are interchangeable; `identifier` and `accessibilityIdentifier` resolve to `resource-id`.
See the [Spec language reference](./spec-language/) for the complete selector grammar and per-platform field availability.
## Properties
Properties are named LTL formulas exported from the spec. The verifier evaluates each one every step and fails the run when a formula is violated.
```ts
export const properties = {
balanceNeverNegative: always(() => balance.current >= 0),
loginReachable: eventually(() => loggedIn.current).within(30, "seconds"),
};
```
**Operators:**
- `always(f)` - `f` must hold at every step.
- `eventually(f).within(n, unit)` - `f` must hold at some step within `n` milliseconds, seconds, or steps.
- `now(f)` - evaluates `f` at the current step (used for implication antecedents).
- `next(f)` - evaluates `f` at the next step.
**Combinators** - available on any formula:
```ts
now(() => loggedIn.current).implies(now(() => cartCount.current !== undefined))
formulaA.and(formulaB)
formulaA.or(formulaB)
formulaA.not()
```
`implies`, `and`, `or`, and `not` compose freely.
## Actions
Action generators return a list of actions to perform. The runner samples one from the weighted tree and dispatches it through the driver.
**Built-in generators** (pass directly to `weighted`):
- `taps` - autonomous random taps on clickable elements.
- `swipes` - autonomous random swipe gestures.
- `waitOnce` - idles one step.
- `pressKeys` - presses a random supported key.
**Action constructors:**
```ts
Tap({ on: element }) // tap an element or selector string
InputText({ into: element, text: "hello" }) // clear and type into a field
Swipe({ from: elementOrPoint, to: elementOrPoint, durationMillis?: number })
PressKey({ key: "back" | "home" | "enter" | "tab" | "up" | "down" | "left" | "right" })
Wait({ durationMillis: number })
```
**Samplers** - cycle over a fixed list:
```ts
const names = from(["Checking", "Savings", "Travel"]);
names.generate() // picks from the list
```
**Custom generators:**
```ts
const doLogin = actions(() => {
if (loggedIn.current) return [];
const emailField = loginEmail.current;
const submit = loginSubmit.current;
if (!emailField || !submit) return [];
return [InputText({ into: emailField, text: "[email protected]" }), Tap({ on: submit })];
});
```
**Weighted trees:**
```ts
export const actionsRoot = weighted(
[100, dismissOnboarding],
[50, doLogin],
[10, taps],
[2, swipes],
[1, weighted(
[3, openDeepLink("app://home")],
[1, openDeepLink("app://settings")],
)],
);
```
Weights are relative within each tree. Nested trees get their own local budget.
## Default properties
`@sanderling/spec/defaults/properties` exports ready-made properties:
```ts
import { noUncaughtExceptions, noLogcatErrors } from "@sanderling/spec/defaults/properties";
export const properties = {
noUncaughtExceptions, // fails if the app throws an uncaught exception
noLogcatErrors, // android-only; reads logcat, no-ops on ios/web
};
```
`noLogcatErrors` reads from logcat and only applies on Android. Including it in a spec that targets iOS or web is harmless; it silently holds.
## Pattern: preconditions
sanderling has no setup phase. Preconditions are action generators with high weight that self-disable once the condition is satisfied.
```ts
const onLoginScreen = extract((s) => !!s.ax.find("id:login-form"));
const loginEmailField = extract((s) => s.ax.find("id:email-field"));
const loginSubmit = extract((s) => s.ax.find("id:sign-in-button"));
const doLogin = actions(() => {
if (!onLoginScreen.current) return [];
const email = loginEmailField.current;
const submit = loginSubmit.current;
if (!email || !submit) return [];
return [
InputText({ into: email, text: "[email protected]" }),
Tap({ on: submit }),
];
});
```
Stack these for onboarding, consent dialogs, and cold-start flows:
```ts
export const actionsRoot = weighted(
[100, dismissOnboarding],
[50, doLogin],
[10, taps],
[2, swipes],
);
```
Once `onLoginScreen.current` is false, `doLogin` returns `[]` and drops out of the eligible set automatically.
## Pattern: setup export
Preconditions that drive the app from a fresh state into the surface you actually want to fuzz (login, onboarding, permission grants, seed data) can be exported as `setup` instead of mixing into `actionsRoot`. The runner tries `setup` first; if it yields no action, it falls through to `actionsRoot`. State regressing back across the precondition (logout under fuzz) automatically re-engages setup.
```ts
const login = actions(() => {
if (loggedIn.current) return [];
const email = loginEmailField.current;
const submit = loginSubmit.current;
if (!email || !submit) return [];
return [InputText({ into: email, text: "[email protected]" }), Tap({ on: submit })];
});
export const setup = login;
export const actionsRoot = weighted([60, browse], [40, edit]);
```
`setup` is just an `ActionGenerator`; compose with `actions`, `weighted`, or `whenRoute` exactly like the main pool. Works identically across Android, iOS, and web.
## Pattern: conditional properties
Gate a property so it only applies when a precondition holds:
```ts
export const properties = {
cartPersistsWhenLoggedIn: always(
now(() => loggedIn.current).implies(now(() => cartCount.current !== undefined)),
),
};
```
## Pattern: step-to-step invariants
Use `next()` to express invariants that span two consecutive steps:
```ts
const newAccountBalanceIsZero = always(
next(() => {
const prev = accounts.previous ?? [];
const curr = accounts.current;
if (prev.length === 0 || curr.length === 0) return true;
const prevIds = new Set(prev.map((a) => a.id));
return curr.filter((a) => !prevIds.has(a.id)).every((a) => a.balance === 0);
}),
);
```
## Anti-patterns
**Accessing `state` outside of `extract`.** The `state` argument exists only inside the `extract()` callback. Use extractors and `.current` everywhere else.
**Positional taps.** `Tap({ on: { x: 100, y: 200 } })` breaks on any layout change. Always prefer an `ax.find("id:...")` reference.
**Unbounded `eventually`.** Without `.within(...)`, `eventually` never fails within a finite run. Almost always you want a bound.
**`Wait()` inside generators.** Waiting for a condition belongs in an extractor guard, not inside a generator.
**Retry logic inside generators.** Generators must be pure. Given the same state they produce the same actions. Retry is the runner's responsibility.