* feat(hierarchy): name the route a native tree shows The screen name was web-only: the Chrome driver stamps sanderling-screen on the root and nothing else does, so every Android and iOS step recorded and logged an empty screen. The route marker the tree already carries (the resource id ending in Screen, the same one Transitional counts) names it. * feat(runner): say what each step did in the step log One line per step carried only an index and a node count. It now names the screen, the action, its target and the typed value, the last through the same redaction the trace and the prompt use. Emitted after the apply so the line reports what actually happened, skip reason included. * fix(sidecar): state on android whether a field is a secure entry maestro's tree mapper copies a fixed attribute list off the device's XML and password is not on it, so no android element ever reported the fact and the conservative rule downstream redacted every typed value in the trace, the prompt and the log. The XML still carries it: re-read it once per settled snapshot and state the fact on the text fields it matches. A field it cannot match stays unstated, which still reads as a credential. * docs: correct the record that android never reports a secure field Four places said android reports the fact for nothing and that every typed value there is redacted. The sidecar now states it, so they described the old behaviour. * test(sidecar): fail the build if maestro renames the call the fact comes from * fix(sidecar): state the fact on a field named by its hint alone collectTextFields matched on class only, so a node the go side calls editable off its hintText was left unstated and its typed value redacted. * docs: record that ios and web state secure:false for compose password fields Both derive the fact from a widget type a compose app never has, so the value reaches the trace in the clear. Verified on folio on both targets. * feat(android): read the application id out of an apk parses the compiled AndroidManifest.xml rather than shelling out to aapt2, which lives in the versioned build-tools directory that hosts with only platform-tools never install. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * feat(cli): let --android-app-path supply the bundle id --bundle-id stays required everywhere else, and an explicit one still wins, so the apk can never quietly override what was asked for. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * docs: record that the apk can name the package itself Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning * feat(folio): ask which android device to run on when none is named * feat(folio): pin ios recipes to one simulator udid and ask when several match * docs(ci): say how just ios lands on the simulator the boot step chose * chore(folio): ignore run output anywhere under examples/folio * refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade ScreenName kept its own reading of the route markers and disagreed with Transitional on a marker repeated by a nested node: it named the screen on a step the runner was skipping as unsettled. One reading now. * refactor(sidecar): inline the one attempt passed to callViewHierarchy A named constant and its own comment for a literal used once. * test(sidecar): compare the whole tree when checking the annotation changes nothing else The old assertions checked one id string and one bounds value, and passed with every other attribute stripped off every node. Now the annotated tree minus the two facts it stated must equal the input. * fix(sidecar): match a field to its xml node by class as well as id and bounds A wrapper drawn to the same bounds as the untagged field inside it shared the field's key, both were dropped as ambiguous, and every value typed into an untagged field was redacted. The class tells them apart. * docs(runs): record what the android hierarchy re-read costs per step Two 1m runs per binary on folio, same seed, before and after the re-read.
6.6 KiB
Folio
A minimal Kotlin Multiplatform personal-ledger app: login with demo credentials, create accounts, add credits and debits. Shared UI across Android, iOS, and web (wasmJs via Compose for Web). Doubles as the example sanderling runs its property-based specs against.
Stack
- Kotlin Multiplatform + Compose Multiplatform (shared UI)
- SQLDelight for the data layer (unified across platforms)
- kotlinx.coroutines for state flows
- kotlinx.serialization for
@Serializableroute types
Prerequisites
just- JDK 17
- Android SDK (auto-discovered under
$ANDROID_HOME,~/Library/Android/sdk, or the Homebrew cask) - Xcode 16+ and
xcodegen(brew install xcodegen) for iOS
Android
just install # asks which device, then builds + installs
ANDROID_DEVICE=emulator-5554 just install # build + install on that device
ANDROID_DEVICE=emulator-5554 just uninstall
just clean
ANDROID_DEVICE is the serial adb devices reports. Every recipe that
installs, uninstalls or fuzzes acts on that serial. Without it they print the
online devices and ask which one to use, and pick on their own only when the
one device adb can see is an emulator on the local adb server. With no terminal
to ask on, a CI job say, the ask becomes a refusal that prints what adb sees.
A run installs the app, clears its state and drives it, which is not something to do to a handset that happens to be the one thing plugged in. An emulator is cheap to rebuild, so a lone local one is the single case worth guessing at.
iOS
just ios # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just ios # name it, by name or UDID
IOS_DEVICE follows the same rule as ANDROID_DEVICE: a lone booted simulator
is taken without asking, anything else is asked about, and with no terminal to
ask on it refuses and lists what is installed. A name is matched against booted
simulators first and available ones second, and it can name several, since the
same iPhone exists under every installed runtime. When it does, you pick which,
and everything after that addresses the chosen UDID: the build destination, the
install, the launch and --ios-device all get the one simulator.
just ios regenerates app/iosApp/iosApp.xcodeproj from app/iosApp/project.yml,
builds the KMP framework (Shared.framework from :app:shared), links it
into the SwiftUI host, uninstalls any previous copy, installs, and launches.
The uninstall matters: folio's signed-in session survives an install over the
top, so without it a run opens on the last run's Home screen.
Web
just web # webpack dev server with COOP/COEP headers
just web-build # produce a webpack distributable bundle
just web runs :app:webApp:wasmJsBrowserDevelopmentRun --continuous, so
edits to shared code reload in the browser.
Demo credentials
email: [email protected]
password: ledger123
Run a sanderling test (Android)
ANDROID_DEVICE=emulator-5554 just test
The same naming rule as just install applies, and just test asks once for
the whole run. If nothing is attached at all, just test boots a bootable AVD
and runs against that. With multiple AVDs, pick one:
AVD=Pixel_7 just test
Persistent settings can live in .env alongside the justfile:
ANDROID_DEVICE=emulator-5554
DURATION=5m
The device does not have to be attached to this machine. ADB_SERVER_SOCKET
aims adb at another host's adb server, and ANDROID_DEVICE names the serial
that server reports. A remote server is shared, so nothing there is ever picked
without being named or asked about, one device on it or twenty:
ADB_SERVER_SOCKET=tcp:10.0.0.5:5037
ANDROID_DEVICE=emulator-5556
The older ANDROID_ADB_SERVER_ADDRESS / ANDROID_ADB_SERVER_PORT pair works
too, at the same precedence the adb CLI gives it. A remote server is never
auto-booted against: when it reports no device, just test says so rather than
starting a local emulator that server will never see.
Gradle only assembles the APK. The install goes through adb, which reads those
variables, so a remote server needs nothing else. Gradle's own installDebug
cannot be used here: its adb client only ever dials loopback.
Traces land in ./sanderling/runs/<timestamp>/.
Run with the LLM action generator
The same sanderling/spec.ts runs under either generator: --generator seeded
(the default weighted fuzzer) or --generator llm, where a vision model picks
from the SAME weighted candidate set (reading the screenshot plus a numbered,
weight-annotated list of concrete actions) and returns one number. The spec's
generator = llm({ model, instructions }) export configures it.
export OPENROUTER_API_KEY=sk-or-... # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio
OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor
prefix from the model id in spec.ts (gpt-5.4-nano, not
openai/gpt-5.4-nano). The model must support image input and strict
json_schema structured outputs. Each step is one multimodal call, so keep the
duration / step budget modest. The trace records the model's reasoning, the
chosen number, and source: "llm" on each action, so the replay UI shows why
each pick was made.
Run a sanderling test (iOS)
just test-ios # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just test-ios # name it, by name or UDID
just test-ios settles on a simulator once for the whole run, boots it if
needed, runs just ios to install and launch the app, then invokes sanderling test --platform ios. Same DURATION, SEED, and OUTPUT env vars as the
Android target.
A physical iPhone is a different target: just test-ios-device requires
IOS_DEVICE to name it, and says so rather than running, because an empty one
resolves to a booted simulator and would fuzz that instead.
How it connects to sanderling
- Each screen sets a stable Compose
testTag(HomeScreen,AccountCard,LedgerRow,TxnAmount, ...). The Sanderling SDK resolvestestTagtoresource-idon Android andaccessibilityIdentifieron iOS. - Identity for list items is the visible text content (account name; txn note + amount). No synthetic IDs encoded in semantics.
contentDescriptionis reserved for real accessibility labels, never as a data carrier.sanderling/spec.tsimports@sanderling/spec, reads state vias.ax.*, asserts properties, and weights the actions the fuzzer picks from.just testinvokessanderling testagainst the installed APK.