* ci(folio): run gradle on jdk 21 for the metro plugin the metro gradle plugin folio builds with publishes org.gradle.jvm.version 21 and java 21 class files, so every leg failed at the folio build on a 17 runtime. local builds pass on jdk 25, which is why only ci saw it. * fix(build): clean pkg/spec/dist, not the dead spec-api path * chore: point stale spec-api comments at pkg/spec * fix(spec): publish src so an installed package carries the runtime entries * fix(testrun): alias the installed spec package so one module graph loads * fix(spec): export Direction, ScrollAction and LongPressAction from the entry * docs(spec): cut the package readme to a description and doc links * docs: say how the cli and spec package versions relate * fix(verifier): report whether the last action was confirmed applied Both hosts get applied: true when the runner saw the dispatch succeed and applied: null when it could not, so an unconfirmed action stops arriving at the spec as no action at all. * fix(runner): an apply error leaves the action's fate unknown, not undone A deadline that fires after the tap was dispatched leaves the effect committed. Reporting nil made the spec see an effect with no action to cause it, which is how the counting property convicts a healthy app. * fix(release): stage the sidecar jar at the renamed embed path * test(replay-ui): trace fixtures for the vacuity counts one real green run, one run that rendered nothing, one that judges every property at least once. * ci(replay-ui): count the steps each property judged the exit code says no property returned false; it does not say any property was ever evaluated. this reads the trace and reports judged vs declined per property, and fails when the step page never rendered. * test(replay-ui): cover the summary script from make test * ci(replay-ui): summarise through the vacuity script * docs(ci): explain the replay-ui judged/declined counts * fix(verifier): encode element-valued extractors into the trace An ax element exports with its find/findAll host functions attached, and json.Marshal refuses the whole value over them: json: unsupported type: func(goja.FunctionCall) goja.Value. The encoding failed, curr stayed nil, and the goja hosts (ios, android) recorded null for every element-valued extractor in both the per-step diff and the violation witness. Apply the web host's sanitize rule before marshaling, so one rule encodes an element on both hosts. * test(verifier): pin element encoding to one rule on both hosts * test(runner): assert an element reaches trace.jsonl and its witness * feat(spec): give state.lastAction an applied field Three states, not two: no action is a null lastAction, applied: true is an action the runner confirmed, applied: null is one it dispatched and never learned the fate of. * fix(folio): do not attribute an effect to an unconfirmed action submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both convict by pinning an effect on the last action, so both decline unless the runner saw it applied. The fixtures now say which fate they mean. * test(folio): an unconfirmed submit belongs in the window The count is an upper bound on the submits a window holds, so the tap that may have landed counts and committedTransactionsExceedSubmits has nothing to convict on. * test(runner): a tap that lands under a failed apply is not a double submit Drives the real folio counting predicates through the runner against a device that commits the tap and then times out. The double-submit case is the control: without it a green proves only that the property never fired. * test(verifier): pin the three lastAction states on both hosts The web page is handed the same applied field the goja object exposes, so a property cannot read one thing on native and another on web. * docs(spec-language): document the three lastAction states
Folio
A minimal Kotlin Multiplatform personal-ledger app: login with demo credentials, create accounts, add credits and debits. Shared UI across Android, iOS, and web (wasmJs via Compose for Web). Doubles as the example sanderling runs its property-based specs against.
Stack
- Kotlin Multiplatform + Compose Multiplatform (shared UI)
- SQLDelight for the data layer (unified across platforms)
- kotlinx.coroutines for state flows
- kotlinx.serialization for
@Serializableroute types
Prerequisites
just- JDK 17
- Android SDK (auto-discovered under
$ANDROID_HOME,~/Library/Android/sdk, or the Homebrew cask) - Xcode 16+ and
xcodegen(brew install xcodegen) for iOS
Android
just install # build + install on a booted emulator / device
just uninstall
just clean
iOS
just ios # default device: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just ios # pick a different simulator
just ios regenerates app/iosApp/iosApp.xcodeproj from app/iosApp/project.yml,
builds the KMP framework (Shared.framework from :app:shared), links it
into the SwiftUI host, installs, and launches.
Web
just web # webpack dev server with COOP/COEP headers
just web-build # produce a webpack distributable bundle
just web runs :app:webApp:wasmJsBrowserDevelopmentRun --continuous, so
edits to shared code reload in the browser.
Demo credentials
email: [email protected]
password: ledger123
Run a sanderling test (Android)
just test
If no device is connected, sanderling boots the single AVD it finds. With multiple AVDs, pick one:
AVD=Pixel_7 just test
Persistent settings can live in .env alongside the justfile:
AVD=Pixel_7
DURATION=5m
Traces land in ./sanderling/runs/<timestamp>/.
Run with the LLM action generator
The same sanderling/spec.ts runs under either generator: --generator seeded
(the default weighted fuzzer) or --generator llm, where a vision model picks
from the SAME weighted candidate set (reading the screenshot plus a numbered,
weight-annotated list of concrete actions) and returns one number. The spec's
generator = llm({ model, instructions }) export configures it.
export OPENROUTER_API_KEY=sk-or-... # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio
OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor
prefix from the model id in spec.ts (gpt-5.4-nano, not
openai/gpt-5.4-nano). The model must support image input and strict
json_schema structured outputs. Each step is one multimodal call, so keep the
duration / step budget modest. The trace records the model's reasoning, the
chosen number, and source: "llm" on each action, so the replay UI shows why
each pick was made.
Run a sanderling test (iOS)
just test-ios # default simulator: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just test-ios # pick a different simulator
just test-ios boots the simulator if needed, runs just ios to install
and launch the app, then invokes sanderling test --platform ios. Same
DURATION, SEED, and OUTPUT env vars as the Android target.
How it connects to sanderling
- Each screen sets a stable Compose
testTag(HomeScreen,AccountCard,LedgerRow,TxnAmount, ...). The Sanderling SDK resolvestestTagtoresource-idon Android andaccessibilityIdentifieron iOS. - Identity for list items is the visible text content (account name; txn note + amount). No synthetic IDs encoded in semantics.
contentDescriptionis reserved for real accessibility labels, never as a data carrier.sanderling/spec.tsimports@sanderling/spec, reads state vias.ax.*, asserts properties, and weights the actions the fuzzer picks from.just testinvokessanderling testagainst the installed APK.