Files
sanderling/examples/folio
pj 9ab59365b2 redact passwords only, and say what each step did in the log (#93)
* feat(hierarchy): name the route a native tree shows

The screen name was web-only: the Chrome driver stamps sanderling-screen on
the root and nothing else does, so every Android and iOS step recorded and
logged an empty screen. The route marker the tree already carries (the
resource id ending in Screen, the same one Transitional counts) names it.

* feat(runner): say what each step did in the step log

One line per step carried only an index and a node count. It now names the
screen, the action, its target and the typed value, the last through the
same redaction the trace and the prompt use. Emitted after the apply so the
line reports what actually happened, skip reason included.

* fix(sidecar): state on android whether a field is a secure entry

maestro's tree mapper copies a fixed attribute list off the device's XML and
password is not on it, so no android element ever reported the fact and the
conservative rule downstream redacted every typed value in the trace, the
prompt and the log. The XML still carries it: re-read it once per settled
snapshot and state the fact on the text fields it matches. A field it cannot
match stays unstated, which still reads as a credential.

* docs: correct the record that android never reports a secure field

Four places said android reports the fact for nothing and that every typed
value there is redacted. The sidecar now states it, so they described the
old behaviour.

* test(sidecar): fail the build if maestro renames the call the fact comes from

* fix(sidecar): state the fact on a field named by its hint alone

collectTextFields matched on class only, so a node the go side calls editable
off its hintText was left unstated and its typed value redacted.

* docs: record that ios and web state secure:false for compose password fields

Both derive the fact from a widget type a compose app never has, so the
value reaches the trace in the clear. Verified on folio on both targets.

* feat(android): read the application id out of an apk

parses the compiled AndroidManifest.xml rather than shelling out to
aapt2, which lives in the versioned build-tools directory that hosts
with only platform-tools never install.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* feat(cli): let --android-app-path supply the bundle id

--bundle-id stays required everywhere else, and an explicit one still
wins, so the apk can never quietly override what was asked for.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* docs: record that the apk can name the package itself

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning

* feat(folio): ask which android device to run on when none is named

* feat(folio): pin ios recipes to one simulator udid and ask when several match

* docs(ci): say how just ios lands on the simulator the boot step chose

* chore(folio): ignore run output anywhere under examples/folio

* refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade

ScreenName kept its own reading of the route markers and disagreed with
Transitional on a marker repeated by a nested node: it named the screen on
a step the runner was skipping as unsettled. One reading now.

* refactor(sidecar): inline the one attempt passed to callViewHierarchy

A named constant and its own comment for a literal used once.

* test(sidecar): compare the whole tree when checking the annotation changes nothing else

The old assertions checked one id string and one bounds value, and passed
with every other attribute stripped off every node. Now the annotated tree
minus the two facts it stated must equal the input.

* fix(sidecar): match a field to its xml node by class as well as id and bounds

A wrapper drawn to the same bounds as the untagged field inside it shared
the field's key, both were dropped as ambiguous, and every value typed into
an untagged field was redacted. The class tells them apart.

* docs(runs): record what the android hierarchy re-read costs per step

Two 1m runs per binary on folio, same seed, before and after the re-read.
2026-09-05 23:54:12 +05:30
..
2026-04-20 16:04:20 +07:00

Folio

A minimal Kotlin Multiplatform personal-ledger app: login with demo credentials, create accounts, add credits and debits. Shared UI across Android, iOS, and web (wasmJs via Compose for Web). Doubles as the example sanderling runs its property-based specs against.

Stack

  • Kotlin Multiplatform + Compose Multiplatform (shared UI)
  • SQLDelight for the data layer (unified across platforms)
  • kotlinx.coroutines for state flows
  • kotlinx.serialization for @Serializable route types

Prerequisites

  • just
  • JDK 17
  • Android SDK (auto-discovered under $ANDROID_HOME, ~/Library/Android/sdk, or the Homebrew cask)
  • Xcode 16+ and xcodegen (brew install xcodegen) for iOS

Android

just install                                # asks which device, then builds + installs
ANDROID_DEVICE=emulator-5554 just install   # build + install on that device
ANDROID_DEVICE=emulator-5554 just uninstall
just clean

ANDROID_DEVICE is the serial adb devices reports. Every recipe that installs, uninstalls or fuzzes acts on that serial. Without it they print the online devices and ask which one to use, and pick on their own only when the one device adb can see is an emulator on the local adb server. With no terminal to ask on, a CI job say, the ask becomes a refusal that prints what adb sees.

A run installs the app, clears its state and drives it, which is not something to do to a handset that happens to be the one thing plugged in. An emulator is cheap to rebuild, so a lone local one is the single case worth guessing at.

iOS

just ios                          # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just ios   # name it, by name or UDID

IOS_DEVICE follows the same rule as ANDROID_DEVICE: a lone booted simulator is taken without asking, anything else is asked about, and with no terminal to ask on it refuses and lists what is installed. A name is matched against booted simulators first and available ones second, and it can name several, since the same iPhone exists under every installed runtime. When it does, you pick which, and everything after that addresses the chosen UDID: the build destination, the install, the launch and --ios-device all get the one simulator.

just ios regenerates app/iosApp/iosApp.xcodeproj from app/iosApp/project.yml, builds the KMP framework (Shared.framework from :app:shared), links it into the SwiftUI host, uninstalls any previous copy, installs, and launches. The uninstall matters: folio's signed-in session survives an install over the top, so without it a run opens on the last run's Home screen.

Web

just web         # webpack dev server with COOP/COEP headers
just web-build   # produce a webpack distributable bundle

just web runs :app:webApp:wasmJsBrowserDevelopmentRun --continuous, so edits to shared code reload in the browser.

Demo credentials

email:    [email protected]
password: ledger123

Run a sanderling test (Android)

ANDROID_DEVICE=emulator-5554 just test

The same naming rule as just install applies, and just test asks once for the whole run. If nothing is attached at all, just test boots a bootable AVD and runs against that. With multiple AVDs, pick one:

AVD=Pixel_7 just test

Persistent settings can live in .env alongside the justfile:

ANDROID_DEVICE=emulator-5554
DURATION=5m

The device does not have to be attached to this machine. ADB_SERVER_SOCKET aims adb at another host's adb server, and ANDROID_DEVICE names the serial that server reports. A remote server is shared, so nothing there is ever picked without being named or asked about, one device on it or twenty:

ADB_SERVER_SOCKET=tcp:10.0.0.5:5037
ANDROID_DEVICE=emulator-5556

The older ANDROID_ADB_SERVER_ADDRESS / ANDROID_ADB_SERVER_PORT pair works too, at the same precedence the adb CLI gives it. A remote server is never auto-booted against: when it reports no device, just test says so rather than starting a local emulator that server will never see.

Gradle only assembles the APK. The install goes through adb, which reads those variables, so a remote server needs nothing else. Gradle's own installDebug cannot be used here: its adb client only ever dials loopback.

Traces land in ./sanderling/runs/<timestamp>/.

Run with the LLM action generator

The same sanderling/spec.ts runs under either generator: --generator seeded (the default weighted fuzzer) or --generator llm, where a vision model picks from the SAME weighted candidate set (reading the screenshot plus a numbered, weight-annotated list of concrete actions) and returns one number. The spec's generator = llm({ model, instructions }) export configures it.

export OPENROUTER_API_KEY=sk-or-...   # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm                         # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio

OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor prefix from the model id in spec.ts (gpt-5.4-nano, not openai/gpt-5.4-nano). The model must support image input and strict json_schema structured outputs. Each step is one multimodal call, so keep the duration / step budget modest. The trace records the model's reasoning, the chosen number, and source: "llm" on each action, so the replay UI shows why each pick was made.

Run a sanderling test (iOS)

just test-ios                          # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just test-ios   # name it, by name or UDID

just test-ios settles on a simulator once for the whole run, boots it if needed, runs just ios to install and launch the app, then invokes sanderling test --platform ios. Same DURATION, SEED, and OUTPUT env vars as the Android target.

A physical iPhone is a different target: just test-ios-device requires IOS_DEVICE to name it, and says so rather than running, because an empty one resolves to a booted simulator and would fuzz that instead.

How it connects to sanderling

  • Each screen sets a stable Compose testTag (HomeScreen, AccountCard, LedgerRow, TxnAmount, ...). The Sanderling SDK resolves testTag to resource-id on Android and accessibilityIdentifier on iOS.
  • Identity for list items is the visible text content (account name; txn note + amount). No synthetic IDs encoded in semantics.
  • contentDescription is reserved for real accessibility labels, never as a data carrier.
  • sanderling/spec.ts imports @sanderling/spec, reads state via s.ax.*, asserts properties, and weights the actions the fuzzer picks from.
  • just test invokes sanderling test against the installed APK.