Merge branch 'docs-that-are-true' into correctness-and-spec-skills

This commit is contained in:
pj committed 2026-08-16 01:15:30 +05:30
commit 772eb1df5e
10 files changed
+173 -114

No files matched your search

+8 -6
View File
@@ -33,7 +33,7 @@ SEED=9 MAX_STEPS=200 .github/scripts/folio-run.sh android
**web and ios expect the bug.** Folio double-submits a transaction when the
submit button is double-tapped, and two properties catch it:
`submitMovesBalanceByAtMostTypedAmount`, which demands the total balance move by
exactly the amount typed, and `submitCommitsOneTransactionPerAction`, which
no more than the amount typed, and `submitCommitsOneTransactionPerAction`, which
demands no more transactions committed over a window than there were submit
actions in it. A double tap is one action committing two transactions, so it
breaks both.
@@ -186,11 +186,13 @@ from `test/browser/testdata/throwing` (violations and uncaught exceptions, so
every panel has something to render), serves it with `sanderling replay`, and
fuzzes that UI with `replay-ui/sanderling/spec.ts`.
Six of the seven properties there are cross-panel agreements - two panels
deriving the same fact by different paths have to say the same thing - so they
hold for any trace and need no recalibrating when the fixture changes. The
seventh is the stock `noUncaughtExceptions`, which asks nothing of the panels
and only fails if the UI throws. Any violation fails the job.
Three of the seven properties there are cross-panel agreements - two panels
deriving the same fact by different paths have to say the same thing. The other
four are a range invariant on the step in the URL, a count of selected rows
inside the list, a no-effect property across a tab switch, and the stock
`noUncaughtExceptions`, which asks nothing of the panels and only fails if the
UI throws. All seven hold for any trace and need no recalibrating when the
fixture changes. Any violation fails the job.
So does a run that judged nothing. Exit 0 says no property returned false, which
is not the same as any property having been evaluated: each one declines to
+9 -6
View File
@@ -24,7 +24,8 @@ Nobody writes this test. A manual tester taps submit once, sees the right number
You do not script the double tap. You state the invariant and let sanderling find the inputs that break it.
The amount the user types must equal the amount the balance moves:
One submit commits one transaction, so the balance cannot move by more than the
amount that submit typed:
```ts
const submitMovesBalanceByAtMostTypedAmount = always(
@@ -38,16 +39,18 @@ const submitMovesBalanceByAtMostTypedAmount = always(
if (typed === 0) return true;
const before = totalBalance.previous;
if (before === null || totalBalance.current === null) return true;
return Math.abs(totalBalance.current - before) === typed;
return Math.abs(totalBalance.current - before) <= typed;
})
);
```
`always` checks the formula at every step; `next` lets it compare the step before a submit to the step after. The guards narrow it to the one transition that matters, a submit that lands back on home, and the last line states the rule: the balance moved by exactly the typed amount. Double-submit moves it by twice that, and the formula is false.
`always` checks the formula at every step; `next` lets it compare the step before a submit to the step after. The guards narrow it to the one transition that matters, a submit that lands back on home, and the last line states the rule: the balance moved by no more than the typed amount. Double-submit moves it by twice that, and the formula is false.
The window guard is the difference between a property and a false conviction. `totalBalance.previous` is the last total we read, not the total as of the last transaction, so the two numbers being compared can straddle any number of commits: a real run produced a delta of 13000 against a typed 19600, because the window held a double-submit's two 19600 debits and an unrelated 26200 credit. A delta like that is not evidence about the amount typed into any one submit. Exactly one submit action in the window still catches the bug, because the double tap is a single action.
The obvious version of that last line is `=== typed`, and it is the version this spec used to ship. It was wrong. Folio's `createTransaction` runs in a coroutine and Home's total re-renders on the store's own schedule, so a total that has not caught up yet is what a healthy app looks like a frame after a submit, and an equality convicts it. So does a submit the app rejected, and so does a tap that never landed. The bound declines on all three without needing a case for any of them, and it still catches the bug, because twice the typed amount is more than the typed amount. What it gives up is worth naming: a transaction committed for less than the amount typed is a real ledger bug this property no longer sees. It cannot be told apart from a total one frame behind, and a check that fires on both is evidence about neither.
The null guard is not defensive clutter either. Read a balance you could not parse as `0` and the comparison becomes `0 - 0 === typed`, which is false at every healthy submit. A reading you do not have is not evidence, so the property declines to judge. The real spec guards the same way against a balance too large for exact integer arithmetic.
The window guard is what keeps the comparison about one submit. `totalBalance.previous` is the last total we read, not the total as of the last transaction, so the two numbers being compared can straddle any number of commits: a real run produced a delta of 13000 against a typed 19600, because the window held a double-submit's two 19600 debits and an unrelated 26200 credit. A delta over a window like that is not evidence about the amount typed into any one submit, whichever side of the bound it falls on. Exactly one submit action in the window still catches the bug, because the double tap is a single action.
The null guard is not defensive clutter either, and under a bound its failure mode is the quiet one. Read a balance you could not parse as `0` and the comparison becomes `|0 - 0| <= typed`, which is true at every submit: the property stops judging and never says so. A reading you do not have is not evidence, so it has to decline in the open rather than pass by accident. The real spec guards the same way against a balance too large for exact integer arithmetic.
The values it reads come from extractors, which pull state out of the UI tree once per step:
@@ -179,7 +182,7 @@ That is the whole input. Three invariants, a way in, and a weighted sense of whe
## What the run does
sanderling launches Folio, logs in, and starts exploring. Most steps are unremarkable: open an account, add a transaction, watch the balance move by exactly what was typed, `submitMovesBalanceByAtMostTypedAmount` holds.
sanderling launches Folio, logs in, and starts exploring. Most steps are unremarkable: open an account, add a transaction, watch the balance move by the amount that was typed, `submitMovesBalanceByAtMostTypedAmount` holds.
Then a step lands two taps on submit before the first save settles. Two transactions post. The balance jumps by twice the typed amount. At that step the formula evaluates false and the run records a violation: the step, the screenshot, the offending action, and the residual formula that failed.
+4 -4
View File
@@ -30,9 +30,9 @@ sanderling doctor
`doctor` reports what the target platform needs and what is missing:
- **Android**: `adb` on your PATH, and an emulator (API 30 or newer) or a connected device.
- **iOS**: Xcode 16 or newer, with a simulator. For a connected iPhone, run `sanderling doctor --platform ios-device`.
- **Web**: Chrome.
- **Android**: `adb` and `emulator`, on your PATH or under the Android SDK; Java 17 or newer; a real (not placeholder) sidecar JAR in the binary.
- **iOS**: `xcrun` and `simctl`. For a connected iPhone, run `sanderling doctor --platform ios-device`, which also wants `devicectl`, a paired device, and signing credentials.
- **Web**: a Chromium that launches headless.
## Write a spec
@@ -46,7 +46,7 @@ export const properties = { noUncaughtExceptions };
export const actionsRoot = defaultActions;
```
This taps, types, scrolls, and swipes at random, and fails the moment your app throws an uncaught exception. From here you add extractors to read your screens, properties that state what your app guarantees, and actions that drive its real flows. The [case study](../case-study/) walks a complete spec, and the [spec language reference](../spec-language/) lists every primitive.
This taps, types, double-taps, scrolls, and swipes at random. Check the property matches your platform before you trust a green run: `noUncaughtExceptions` reads exceptions the web runtime captures in the page, so it fires on `--platform web` and holds unconditionally on Android and iOS. On Android use `noLogcatErrors` instead, which fails on any error-level log line and so catches an uncaught throwable. iOS has neither today, so an iOS run is only worth as much as the properties you write yourself. From here you add extractors to read your screens, properties that state what your app guarantees, and actions that drive its real flows. The [case study](../case-study/) walks a complete spec, and the [spec language reference](../spec-language/) lists every primitive.
## Run it
+3 -3
View File
@@ -42,7 +42,7 @@ interface State {
| `snapshots` | Key-value data pushed by the app SDK (empty if SDK not integrated) |
| `lastAction` | The action dispatched in the previous step, or `null` on the first step and on any step that dispatched nothing |
| `logs` | Log entries collected since the previous step |
| `exceptions` | Uncaught exceptions or `Sanderling.reportError()` calls since the previous step |
| `exceptions` | Uncaught exceptions and unhandled promise rejections captured in the page. Web only: nothing fills this on Android or iOS, where it is always empty |
| `time` | Milliseconds elapsed since the run started |
`lastAction.applied` is `true` when the runner saw the dispatch succeed and `null` when the apply call failed with the action possibly already delivered: an RPC deadline can fire after the tap reached the app, and nothing can find out afterwards. So there are three states, not two. `state.lastAction === null` means no action ran; `applied === null` means one ran whose fate is unknown. A property that attributes an effect to the action ("this submit must move the balance by the typed amount") has to decline unless `applied` is `true`, or a timeout convicts a healthy app. A property that counts what the app COULD have done should include it: an unconfirmed submit belongs in an upper bound on how many submits a window holds.
@@ -303,9 +303,9 @@ import { defaultActions, doubleTaps } from "@sanderling/spec/defaults";
import { noUncaughtExceptions, noLogcatErrors } from "@sanderling/spec/defaults/properties";
```
`defaultActions` is a ready-made weighted tree of the built-in generators: taps and typing at weight 100, scrolls 50, swipes 25, double taps 10. Use it as a baseline pool or as one entry in your own tree.
`defaultActions` is a ready-made weighted tree of five of the built-in generators: taps and typing at weight 100, scrolls 50, swipes 25, double taps 10. `longPresses`, `pressKeys` and `waitOnce` are not in it; weight them in yourself if you want them. Use it as a baseline pool or as one entry in your own tree.
| Property | Fails when |
|---|---|
| `noUncaughtExceptions` | An uncaught exception or `Sanderling.reportError()` call is captured |
| `noUncaughtExceptions` | The page captured an uncaught exception or an unhandled rejection (web only; holds on Android and iOS, where `state.exceptions` is never populated) |
| `noLogcatErrors` | Logcat emits any error-level (`E`) lines since the previous step (Android only; holds elsewhere) |
+4 -3
View File
@@ -12,9 +12,10 @@ export function next(predicate: () => boolean): Formula {
return globalThis.__sanderling__.next(predicate);
}
// An unbounded `eventually` never forces a violation within a finite run.
// Prefer `.within(n, unit)` when you want the verifier to fail a property
// that stalls.
// An unbounded `eventually` that never fires is violated when the run ends,
// with the reason "eventually never satisfied", so a goal the run does not
// reach is a violation every time. `.within(n, unit)` convicts at the step the
// window closes instead of at run end.
export function eventually(predicate: () => boolean): EventuallyFormula {
return globalThis.__sanderling__.eventually(predicate);
}
+58 -44
View File
@@ -106,34 +106,34 @@ climbed), state a bound on it rather than a prediction of it.
**Prefer an upper bound to an equality.** This is the single most valuable
sentence in this file.
`examples/folio/sanderling/predicates.ts` states one rule about one app both
ways, so the two are worth reading side by side. Each line is the last line of
its predicate, after the guards, at a step where exactly one submit sits in the
window:
Folio shipped one of these both ways and the equality lost, so the two are worth
reading side by side. Each line is the last line of a predicate in
`examples/folio/sanderling/predicates.ts`, after the guards, at a step where
exactly one submit sits in the window:
```ts
// sound, committedAmountExceedsOneSubmit: the violation is moving by MORE
// than the one submit in this window could account for
Math.abs(currAccountBalance - prevAccountBalance) > typedAmount
// tempting, submitChangesBalanceByTypedAmount: it moved by exactly what I typed
// what folio's total-balance property demanded, until 6e8e6d5
Math.abs(currTotalBalance - prevTotalBalance) === typedAmount
// what it demands now
Math.abs(currTotalBalance - prevTotalBalance) <= typedAmount
// and the same bound stated as the violation, over the account's own balance
Math.abs(currAccountBalance - prevAccountBalance) > typedAmount
```
Both catch the bug, because a double submit moves the balance by twice the typed
amount. Only the second also convicts an app that behaved. A balance that has
not moved is a commit still in flight (folio's `createTransaction` runs in a
coroutine), a submit the app rejected, or a tap that never landed, and none of
those is evidence of anything.
All three catch the bug, because a double submit moves the balance by twice the
typed amount and twice x exceeds x. Only the equality also convicts an app that
behaved. A balance that has not moved is a commit still in flight (folio's
`createTransaction` runs in a coroutine, and Home's total re-renders on the
store's own schedule), a submit the app rejected, or a tap that never landed,
and none of those is evidence of anything.
The asymmetry is the point. Moving by more than one submit's worth is not
something a correct app can do, so the bound needs no case for any of the three.
The equality needs a case for each, and every one you forget is a false
conviction. Compare the two guard stacks and the price is exactly legible: the
equality declines on `confirmedApplied` and on `acrossRelaunch`, and the bound
carries neither, because a submit that may not have landed and a restart that
may have eaten the commit both leave the balance under the bound anyway. Those
are the two facts the runner cannot promise (see below), and needing to guard
against both is a cost of the equality, not of the app.
conviction. Two of those cases are facts the runner cannot promise you (see
below): an action it could not confirm was applied, and an action it had to
relaunch the app after. Both leave the balance under the bound and both break an
equality, so a bound counts them and an equality has to decline on them.
You do give something up, so make the trade deliberately. A bound cannot see a
balance that moved by *less* than the typed amount, and for a ledger that is a
@@ -311,7 +311,7 @@ incoming screen together on 425 of 1879 steps measured across 17 runs, better
than one frame in five. Such a frame is evidence about neither screen, and
ranking the markers to pick one is how a spec convicts itself on an animation.
## 6. The ones you get for free
## 6. The ones you get for free, on one platform each
```ts
import { noUncaughtExceptions, noLogcatErrors } from "@sanderling/spec/defaults";
@@ -319,25 +319,38 @@ import { noUncaughtExceptions, noLogcatErrors } from "@sanderling/spec/defaults"
export const properties = { noUncaughtExceptions, /* yours */ };
```
Export `noUncaughtExceptions` before you write anything of your own. It costs a
line, it needs no app knowledge, and a fuzzer typing `'; DROP TABLE--` and a
4096-character string into every field it finds will surface real breakage
through it. `noLogcatErrors` is stricter and Android-only; it holds trivially
elsewhere, so it is worth turning on once you know your app's log hygiene can
support it.
Both read a field the driver fills, and each field is filled on one platform, so
check which one is yours before counting either as coverage. Folio's spec exports
neither, and that is the tell: one spec drives its Android, iOS and web builds,
and neither of these holds anything on all three.
They do not substitute for the shapes above. An app can be thoroughly wrong
about money without throwing once.
`noUncaughtExceptions` fails when `state.exceptions` is non-empty. Only the web
runtime fills it, from `error` and `unhandledrejection` listeners installed in
the page by `pkg/spec/src/web-runtime.ts`. On web it is worth the line: a fuzzer
typing `'; DROP TABLE--` and a 4096-character string into every field it finds
will surface real breakage through it. On Android and iOS the field is never
populated, so the property holds at every step of a run that crashed.
`noLogcatErrors` fails on a log line at level `E`. An uncaught Java or Kotlin
throwable is logged there, so on Android it is the nearest equivalent and worth
turning on once you know your app's log hygiene can support it. It holds
vacuously on web and iOS.
That leaves iOS with neither, and it leaves both platforms uncovered for the
thing that matters most anyway. An app can be thoroughly wrong about money
without throwing once.
## The rules that cut across all of them
**Absence is unknown, never a default.** Extractors return null when the element
is not there, and a property handed null declines. `0`, `""` and `[]` are the
values that turn a property into one that fires on healthy runs: folio's
balances once parsed as `0` on web, so the check became `|0 - 0| === typed` and
was false at every healthy submit. An empty list has the same problem in the
other direction, and it is worse because it looks reasonable. Android renders
Home's own node a frame or two before its list, so `findAll` over the cards
balances once parsed as `0` on web, so the check, an equality at the time,
became `|0 - 0| === typed` and was false at every healthy submit. Under today's
bound the same `0` reads as `|0 - 0| <= typed` and passes at every submit
instead, which is the same defect wearing green. An empty list has the same
problem in the other direction, and it is worse because it looks reasonable.
Android renders Home's own node a frame or two before its list, so `findAll` over the cards
comes back empty while the screen already claims to be Home. That is unknown,
not "no accounts", and reading it as zero accounts killed folio's counting
invariant outright: `countsBefore` was `{}` at every evaluation point of all 17
@@ -378,23 +391,24 @@ of its states is unsound:
One rule covers the last two, and it is the rule that decides shape 2 for you.
An action the runner cannot fully vouch for **still counts toward a bound on
what the app could have done**, and it **never licenses attributing an effect to
it**. So a bound counts it and an equality has to decline on it. That is why
`committedAmountExceedsOneSubmit` needs no `confirmedApplied` guard and no
`acrossRelaunch` guard while `submitChangesBalanceByTypedAmount` needs both: a
property demanding the effect of an action that may never have run, or that a
restart may have swallowed, convicts the app of the runner's own uncertainty.
it**. So a bound counts it and a property demanding an effect has to decline on
it. That is why `committedAmountExceedsOneSubmit`, which only bounds how far the
balance could have moved, needs no `confirmedApplied` guard and no
`acrossRelaunch` guard, while `createdAccountHasNonZeroBalance`, which demands
that a card appear, needs both. Demanding the effect of an action that may never
have run, or that a restart may have swallowed, convicts the app of the runner's
own uncertainty.
`relaunched` is the same shape of fact as `applied`, applied to app state rather
than to dispatch. The action itself did happen. What nobody can promise across
it is that the process ran continuously, that the commit survived, or that the
screen is showing the same slice of the same list it was. So a property assuming
continuous state declines, via `acrossRelaunch(lastAction)`, and folio uses it
in three places: `createdAccountHasNonZeroBalance` declines because Home redraws
from the top and the card carrying the typed name may be an older account laid
out where the new one used to be, the equality property declines because it
demands an effect, and `countSubmitsInWindow` uses it to **stop trusting its own
refusal evidence**, since a relaunch is the one thing that can put a form state
on screen other than the one the tap read.
continuous state declines, via `acrossRelaunch(lastAction)`.
`createdAccountHasNonZeroBalance` declines because Home redraws from the top and
the card carrying the typed name may be an older account laid out where the new
one used to be. `countSubmitsInWindow` uses the same call to **stop trusting its
own refusal evidence**, since a relaunch is the one thing that can put a form
state on screen other than the one the tap read.
Both fields are `true | null` rather than booleans, and that is deliberate: only
the positive report is a fact the runner can vouch for, so `null` is "not
+3 -2
View File
@@ -64,8 +64,9 @@ evaluation produced the violation; for a deferred obligation (a `next`, an
The discipline is one sentence: open the witness and confirm those values could
actually produce that verdict. An iOS witness read `typedAmount = 0`, and
`submitChangesBalanceByTypedAmount` in `examples/folio/sanderling/predicates.ts`
returns true at `typedAmount === 0` before it compares anything. So the trace
`submitChangesBalanceByAtMostTypedAmount` in
`examples/folio/sanderling/predicates.ts` returns true at `typedAmount === 0`
before it compares anything. So the trace
appeared to show a conviction that could not have happened. The verdict was real
and the artifact was lying, and until that was resolved neither the bug report
nor the fix could be trusted.
+37 -28
View File
@@ -10,10 +10,9 @@ believe. Two things decide that, and both are usually treated as chores: the
handles the app exposes, and the state the app starts in. Everything else here
is plumbing.
Every flag below is one the binary accepts. `sanderling test -h` is the
authority, not this file and not the manual: the manual currently documents
`--launcher-activity`, which the binary answers with `flag provided but not
defined`. Check before you use a flag you have not seen work.
Every flag below is one the binary accepts, checked against `sanderling test -h`
on this revision. That command is the authority, not this file and not the
manual. Check before you use a flag you have not seen work.
## 1. Install, then check the host
@@ -28,31 +27,34 @@ Both come from the same release tag and the CLI bundles the package's TypeScript
when it evaluates your spec, so they move together.
`sanderling doctor` reports the host's readiness per platform and exits non-zero
if anything is missing. On a Mac with no Android SDK it says:
if anything is missing. Each line names the check and, on a failure, what to do
about it. On a Mac with the Android SDK installed but a CLI built by a plain
`go build`, `sanderling doctor --platform android` says:
```
OK adb on PATH
FAIL emulator on PATH or under ANDROID_HOME: not on PATH and ANDROID_HOME is unset
OK adb on PATH or under the Android SDK
OK emulator on PATH or under the Android SDK
OK java 17+ on PATH
FAIL sidecar JAR is real (not placeholder): placeholder JAR embedded; run `make sidecar && make sanderling` to embed the real fat JAR
error: 2 check(s) failed
error: 1 check(s) failed
```
Scope it with `--platform web|android|ios|ios-device|all` (default `all`). Web
needs a Chromium that launches headless. Android needs `adb`, an emulator on
PATH or under `ANDROID_HOME`, Java 17 or newer, and the embedded sidecar JAR.
iOS needs `xcrun` and `simctl`; `ios-device` adds `devicectl`, the macOS usbmuxd
socket, a connected paired device, and App Store Connect signing credentials.
needs a Chromium that launches headless. Android needs `adb`, an emulator, Java
17 or newer, and the embedded sidecar JAR. iOS needs `xcrun` and `simctl`;
`ios-device` adds `devicectl`, the macOS usbmuxd socket, a connected paired
device, and App Store Connect signing credentials.
Read the doctor's Android result as advisory rather than final: its emulator
check today looks only at PATH, `ANDROID_HOME` and `ANDROID_SDK_ROOT`, while a
run also searches `~/Library/Android/sdk`, `~/Android/Sdk` and the Homebrew
command-line-tools paths. The run's own error names every location it tried, so
that is the one to trust. In the other direction, a missing SDK can surface
during a run as `sidecar health check: context deadline exceeded` about thirty
seconds in, which names the symptom and not the cause (issue #69). If you see
it, go back to `sanderling doctor --platform android` before believing anything
about the sidecar.
The `adb` and `emulator` checks resolve through the same helpers a run uses, so
they search PATH, then `ANDROID_HOME` and `ANDROID_SDK_ROOT`, then
`~/Library/Android/sdk`, `~/Android/Sdk` and the Homebrew command-line-tools
paths. A host the doctor passes is a host a run can drive, and a failure names
every location it tried. What the doctor cannot tell you is the reverse: a
missing SDK can also surface during a run as `sidecar health check: context
deadline exceeded` about thirty seconds in, which names the symptom and not the
cause (issue #69). If you see it, go back to
`sanderling doctor --platform android` before believing anything about the
sidecar.
Two traps if you build from source rather than installing a release. A plain
`go build ./cmd/sanderling` embeds a placeholder sidecar JAR, so every Android
@@ -123,11 +125,13 @@ agreement properties possible.
**iOS.** `accessibilityIdentifier`, set via `.accessibilityIdentifier` in
SwiftUI or UIKit. Compose Multiplatform maps `testTag` to it for you.
Two rules about the names themselves. A `testTag` selector falls through to a
substring compare, so `{testTag: "Sub"}` matches `AddAccountSubmit`: make each
hook a whole distinct name rather than a fragment of another. And give every
screen a marker of its own, because a route extractor is what lets a property
decline on the screens it has nothing to say about.
Two rules about the names themselves. On Android and iOS a `testTag` selector
falls through to a substring compare, so `{testTag: "Sub"}` matches
`AddAccountSubmit`; on web the same selector compiles to an exact CSS attribute
match and hits nothing. Make each hook a whole distinct name rather than a
fragment of another, and you are right on both. And give every screen a marker
of its own, because a route extractor is what lets a property decline on the
screens it has nothing to say about.
The check that a hook exists is not that you added it. It is that you can point
at a step in a real trace where a selector over it resolved to a value.
@@ -211,8 +215,13 @@ against an empty page gives `nodes=2` for every step, and still exits 0 with no
violations. If `nodes` is a handful and never grows, the run is looking at
something that is not your app.
`screen=` is the route marker your spec's screen hooks produce. A run where it
never changes never left one screen.
`screen=` is web-only and has nothing to do with your spec's screen hooks: the
Chrome driver puts the URL hash, or the pathname when there is no hash, on the
root node, and only that driver writes the attribute. On Android and iOS it is
empty on every step, so an empty `screen=` there is the normal reading and not a
symptom. On web, a `screen=` that never changes means the run never left one
URL, which for a single-page app that routes in memory is also normal. Your
spec's own route extractor is the thing to trust on every platform.
The summary can carry a third line you should never skim past:
+33 -12
View File
@@ -115,18 +115,30 @@ every reading take its answer from there. `routeOfFrame` in folio's
to the next, which is how you state "this action had that effect". `now(f)`
evaluates at the current step inside a formula body.
`eventually(f).within(n, "steps" | "seconds" | "milliseconds")` requires `f`
before the window closes; unbounded, it never fails a finite run. At the top
before the window closes and convicts at the step it does not. Unbounded, it
does not stop being a liveness obligation: one that never fires is violated when
the run ends, with the reason `eventually never satisfied`. So an `eventually`
over a state your run may not reach fires on every run that does not reach it,
and that is the usual way a first spec ends up red for no reason. At the top
level an `eventually` is one goal for the whole run, armed once and discharged
for good the first time it holds; written inside `always` it re-arms at every
step, which asks for the window to be met from everywhere. Every formula has
`.implies`, `.and`, `.or`, `.not`.
The stock properties are in `@sanderling/spec/defaults`. Put
`noUncaughtExceptions` in every spec: it fails when the run captures an uncaught
throwable or a `Sanderling.reportError` call, it needs no hooks and no
calibration, and it is free. `noLogcatErrors` (also exported from
`@sanderling/spec/defaults/properties`) fails on any error-level logcat line and
is Android-only, holding vacuously elsewhere.
The stock properties are in `@sanderling/spec/defaults`. Both are cheap and both
are narrower than their names suggest, so know which platform yours runs on.
`noUncaughtExceptions` fails when `state.exceptions` is non-empty, and today only
the web runtime fills it: `pkg/spec/src/web-runtime.ts` installs `error` and
`unhandledrejection` listeners in the page. On Android and iOS nothing populates
the field, so it holds at every step whatever the app does. Export it on web,
where it is free and real; on native, understand that a green run says nothing
about crashes.
`noLogcatErrors` fails on any log line the driver reports at level `E`, which is
where an uncaught Java or Kotlin throwable lands, so on Android it is the closest
thing to `noUncaughtExceptions`. It holds vacuously on web and iOS. Neither
platform has an equivalent today: an iOS crash is invisible to both properties.
## What makes a good first property
@@ -136,11 +148,17 @@ the list renders rows; a badge counts violation records and the panel counts the
rows it can show for them. Those hold on any run, so they never need
recalibrating against a fixture, and an app that gets the fact wrong in one of
the two places cannot satisfy them however it was driven there.
`replay-ui/sanderling/spec.ts` is six of these plus `noUncaughtExceptions`, and
its header explains the choice.
Three of the seven properties in `replay-ui/sanderling/spec.ts` are this shape:
`stepCountMatchesTheList`, `screenshotShowsTheSelectedStep` and
`badgeCountMatchesThePanel`. The rest of that spec shows what to write when no
second panel derives the fact: a range invariant on user input
(`selectedStepIsInRange`), a counting invariant inside one panel
(`exactlyOneStepIsSelected`), a no-effect property across an action
(`switchingTabsKeepsTheStep`), and the stock `noUncaughtExceptions`. All of them
still hold on any run, which is the property worth keeping.
Contrast a property that needs the fuzzer to reach a specific state, like
folio's "a submit moves the balance by exactly the amount typed". That is where
folio's "a submit moves the balance by no more than the amount typed". That is where
the real bugs are, and it is the harder thing to keep honest: it needs an action
tree that reaches the state, a window that closes often enough to bound what
happened inside it, and attribution that cannot blame the wrong action. Folio's
@@ -159,8 +177,11 @@ on. If you cannot, it is decoration.
picker chooses one. The verbs are `Tap`, `DoubleTap`, `LongPress`, `InputText`,
`Scroll`, `Swipe`, `PressKey`, and `Wait`. The built-in generators are `taps`,
`doubleTaps`, `longPresses`, `typing`, `scrolls`, `swipes`, `pressKeys`, and
`waitOnce`; `defaultActions` bundles them at taps and typing 100, scrolls 50,
swipes 25, double taps 10.
`waitOnce`. `defaultActions` bundles five of them: taps and typing at 100,
scrolls 50, swipes 25, double taps 10. `longPresses`, `pressKeys` and `waitOnce`
are not in it, so a spec that only exports `defaultActions` never presses android
back, never long-presses, and never waits. Weight those in yourself if the app
has behaviour behind them.
`weighted([n, generator], ...)` composes them with relative weights.
`whenRoute(routeExtractor, routes, body)` runs `body` only on the named screens.
+14 -6
View File
@@ -86,12 +86,20 @@ function of that string can separate them.
**Match whole keys, not endings or substrings.** `endsWith` attribution judges an
older account named `Emergency Fund` when the user typed `Fund`.
Selector matching has the same trap and it is easy to miss which keys carry it.
`id`, `text`, `desc` and `descPrefix` resolve by their own rules, so `{id: "Sub"}`
correctly matches nothing. Every other key, `testTag` included, falls through to a
substring compare, so `{testTag: "Sub"}` matches `AddAccountSubmit`. That is not a
match, it is a coincidence, and a property built on it judges whichever element
happens to contain the fragment.
Selector matching has the same trap on Android and iOS, and it is easy to miss
which keys carry it. `id`, `desc` and `descPrefix` resolve by rules of their own
(exact or `:id/`-suffixed, exact or comma-prefixed, starts-with), so
`{id: "Sub"}` correctly matches nothing. Every other key, `text` and `testTag`
included, falls through to a substring compare, so `{testTag: "Sub"}` matches
`AddAccountSubmit`. That is not a match, it is a coincidence, and a property
built on it judges whichever element happens to contain the fragment.
The web path does not share the rule, which is its own trap. `web-runtime.ts`
compiles an object selector to CSS, and every key becomes an exact attribute
match (`descPrefix` alone becomes a `^=` prefix). So the loose selector that
resolved on Android resolves to nothing on web, and every property over it goes
vacuously true rather than red. Reviewing a cross-platform spec means checking
that each selector is exact enough for native and literal enough for web.
**Drop the carrier when the screen changes.** A value carried across a route
change is a value read from a screen that is no longer there.