Files
pj 11f72a722a follow-ups from the pr #73 review (#77)
* ci(folio): run gradle on jdk 21 for the metro plugin

the metro gradle plugin folio builds with publishes org.gradle.jvm.version 21
and java 21 class files, so every leg failed at the folio build on a 17
runtime. local builds pass on jdk 25, which is why only ci saw it.

* fix(build): clean pkg/spec/dist, not the dead spec-api path

* chore: point stale spec-api comments at pkg/spec

* fix(spec): publish src so an installed package carries the runtime entries

* fix(testrun): alias the installed spec package so one module graph loads

* fix(spec): export Direction, ScrollAction and LongPressAction from the entry

* docs(spec): cut the package readme to a description and doc links

* docs: say how the cli and spec package versions relate

* fix(verifier): report whether the last action was confirmed applied

Both hosts get applied: true when the runner saw the dispatch succeed and
applied: null when it could not, so an unconfirmed action stops arriving at
the spec as no action at all.

* fix(runner): an apply error leaves the action's fate unknown, not undone

A deadline that fires after the tap was dispatched leaves the effect
committed. Reporting nil made the spec see an effect with no action to cause
it, which is how the counting property convicts a healthy app.

* fix(release): stage the sidecar jar at the renamed embed path

* test(replay-ui): trace fixtures for the vacuity counts

one real green run, one run that rendered nothing, one that judges every property at least once.

* ci(replay-ui): count the steps each property judged

the exit code says no property returned false; it does not say any property was ever evaluated. this reads the trace and reports judged vs declined per property, and fails when the step page never rendered.

* test(replay-ui): cover the summary script from make test

* ci(replay-ui): summarise through the vacuity script

* docs(ci): explain the replay-ui judged/declined counts

* fix(verifier): encode element-valued extractors into the trace

An ax element exports with its find/findAll host functions attached, and
json.Marshal refuses the whole value over them: json: unsupported type:
func(goja.FunctionCall) goja.Value. The encoding failed, curr stayed nil,
and the goja hosts (ios, android) recorded null for every element-valued
extractor in both the per-step diff and the violation witness.

Apply the web host's sanitize rule before marshaling, so one rule encodes
an element on both hosts.

* test(verifier): pin element encoding to one rule on both hosts

* test(runner): assert an element reaches trace.jsonl and its witness

* feat(spec): give state.lastAction an applied field

Three states, not two: no action is a null lastAction, applied: true is an
action the runner confirmed, applied: null is one it dispatched and never
learned the fate of.

* fix(folio): do not attribute an effect to an unconfirmed action

submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both
convict by pinning an effect on the last action, so both decline unless the
runner saw it applied. The fixtures now say which fate they mean.

* test(folio): an unconfirmed submit belongs in the window

The count is an upper bound on the submits a window holds, so the tap that may
have landed counts and committedTransactionsExceedSubmits has nothing to
convict on.

* test(runner): a tap that lands under a failed apply is not a double submit

Drives the real folio counting predicates through the runner against a device
that commits the tap and then times out. The double-submit case is the control:
without it a green proves only that the property never fired.

* test(verifier): pin the three lastAction states on both hosts

The web page is handed the same applied field the goja object exposes, so a
property cannot read one thing on native and another on web.

* docs(spec-language): document the three lastAction states
2026-08-15 15:51:33 +05:30
..
2026-04-20 16:04:20 +07:00
2026-07-31 21:12:00 +05:30

Folio

A minimal Kotlin Multiplatform personal-ledger app: login with demo credentials, create accounts, add credits and debits. Shared UI across Android, iOS, and web (wasmJs via Compose for Web). Doubles as the example sanderling runs its property-based specs against.

Stack

  • Kotlin Multiplatform + Compose Multiplatform (shared UI)
  • SQLDelight for the data layer (unified across platforms)
  • kotlinx.coroutines for state flows
  • kotlinx.serialization for @Serializable route types

Prerequisites

  • just
  • JDK 17
  • Android SDK (auto-discovered under $ANDROID_HOME, ~/Library/Android/sdk, or the Homebrew cask)
  • Xcode 16+ and xcodegen (brew install xcodegen) for iOS

Android

just install      # build + install on a booted emulator / device
just uninstall
just clean

iOS

just ios                          # default device: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just ios   # pick a different simulator

just ios regenerates app/iosApp/iosApp.xcodeproj from app/iosApp/project.yml, builds the KMP framework (Shared.framework from :app:shared), links it into the SwiftUI host, installs, and launches.

Web

just web         # webpack dev server with COOP/COEP headers
just web-build   # produce a webpack distributable bundle

just web runs :app:webApp:wasmJsBrowserDevelopmentRun --continuous, so edits to shared code reload in the browser.

Demo credentials

email:    [email protected]
password: ledger123

Run a sanderling test (Android)

just test

If no device is connected, sanderling boots the single AVD it finds. With multiple AVDs, pick one:

AVD=Pixel_7 just test

Persistent settings can live in .env alongside the justfile:

AVD=Pixel_7
DURATION=5m

Traces land in ./sanderling/runs/<timestamp>/.

Run with the LLM action generator

The same sanderling/spec.ts runs under either generator: --generator seeded (the default weighted fuzzer) or --generator llm, where a vision model picks from the SAME weighted candidate set (reading the screenshot plus a numbered, weight-annotated list of concrete actions) and returns one number. The spec's generator = llm({ model, instructions }) export configures it.

export OPENROUTER_API_KEY=sk-or-...   # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm                         # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio

OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor prefix from the model id in spec.ts (gpt-5.4-nano, not openai/gpt-5.4-nano). The model must support image input and strict json_schema structured outputs. Each step is one multimodal call, so keep the duration / step budget modest. The trace records the model's reasoning, the chosen number, and source: "llm" on each action, so the replay UI shows why each pick was made.

Run a sanderling test (iOS)

just test-ios                          # default simulator: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just test-ios   # pick a different simulator

just test-ios boots the simulator if needed, runs just ios to install and launch the app, then invokes sanderling test --platform ios. Same DURATION, SEED, and OUTPUT env vars as the Android target.

How it connects to sanderling

  • Each screen sets a stable Compose testTag (HomeScreen, AccountCard, LedgerRow, TxnAmount, ...). The Sanderling SDK resolves testTag to resource-id on Android and accessibilityIdentifier on iOS.
  • Identity for list items is the visible text content (account name; txn note + amount). No synthetic IDs encoded in semantics.
  • contentDescription is reserved for real accessibility labels, never as a data carrier.
  • sanderling/spec.ts imports @sanderling/spec, reads state via s.ax.*, asserts properties, and weights the actions the fuzzer picks from.
  • just test invokes sanderling test against the installed APK.