mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 11:07:10 +00:00
* feat(hierarchy): name the route a native tree shows The screen name was web-only: the Chrome driver stamps sanderling-screen on the root and nothing else does, so every Android and iOS step recorded and logged an empty screen. The route marker the tree already carries (the resource id ending in Screen, the same one Transitional counts) names it. * feat(runner): say what each step did in the step log One line per step carried only an index and a node count. It now names the screen, the action, its target and the typed value, the last through the same redaction the trace and the prompt use. Emitted after the apply so the line reports what actually happened, skip reason included. * fix(sidecar): state on android whether a field is a secure entry maestro's tree mapper copies a fixed attribute list off the device's XML and password is not on it, so no android element ever reported the fact and the conservative rule downstream redacted every typed value in the trace, the prompt and the log. The XML still carries it: re-read it once per settled snapshot and state the fact on the text fields it matches. A field it cannot match stays unstated, which still reads as a credential. * docs: correct the record that android never reports a secure field Four places said android reports the fact for nothing and that every typed value there is redacted. The sidecar now states it, so they described the old behaviour. * test(sidecar): fail the build if maestro renames the call the fact comes from * fix(sidecar): state the fact on a field named by its hint alone collectTextFields matched on class only, so a node the go side calls editable off its hintText was left unstated and its typed value redacted. * docs: record that ios and web state secure:false for compose password fields Both derive the fact from a widget type a compose app never has, so the value reaches the trace in the clear. Verified on folio on both targets. * feat(android): read the application id out of an apk parses the compiled AndroidManifest.xml rather than shelling out to aapt2, which lives in the versioned build-tools directory that hosts with only platform-tools never install. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * feat(cli): let --android-app-path supply the bundle id --bundle-id stays required everywhere else, and an explicit one still wins, so the apk can never quietly override what was asked for. Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * docs: record that the apk can name the package itself Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc * fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning * feat(folio): ask which android device to run on when none is named * feat(folio): pin ios recipes to one simulator udid and ask when several match * docs(ci): say how just ios lands on the simulator the boot step chose * chore(folio): ignore run output anywhere under examples/folio * refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade ScreenName kept its own reading of the route markers and disagreed with Transitional on a marker repeated by a nested node: it named the screen on a step the runner was skipping as unsettled. One reading now. * refactor(sidecar): inline the one attempt passed to callViewHierarchy A named constant and its own comment for a literal used once. * test(sidecar): compare the whole tree when checking the annotation changes nothing else The old assertions checked one id string and one bounds value, and passed with every other attribute stripped off every node. Now the annotated tree minus the two facts it stated must equal the input. * fix(sidecar): match a field to its xml node by class as well as id and bounds A wrapper drawn to the same bounds as the untagged field inside it shared the field's key, both were dropped as ambiguous, and every value typed into an untagged field was redacted. The class tells them apart. * docs(runs): record what the android hierarchy re-read costs per step Two 1m runs per binary on folio, same seed, before and after the re-read.
171 lines
6.6 KiB
Markdown
171 lines
6.6 KiB
Markdown
# Folio
|
|
|
|
A minimal Kotlin Multiplatform personal-ledger app: login with demo
|
|
credentials, create accounts, add credits and debits. Shared UI across
|
|
Android, iOS, and web (wasmJs via Compose for Web). Doubles as the
|
|
example sanderling runs its property-based specs against.
|
|
|
|
## Stack
|
|
|
|
- Kotlin Multiplatform + Compose Multiplatform (shared UI)
|
|
- SQLDelight for the data layer (unified across platforms)
|
|
- kotlinx.coroutines for state flows
|
|
- kotlinx.serialization for `@Serializable` route types
|
|
|
|
## Prerequisites
|
|
|
|
- `just`
|
|
- JDK 17
|
|
- Android SDK (auto-discovered under `$ANDROID_HOME`, `~/Library/Android/sdk`,
|
|
or the Homebrew cask)
|
|
- Xcode 16+ and `xcodegen` (`brew install xcodegen`) for iOS
|
|
|
|
## Android
|
|
|
|
```sh
|
|
just install # asks which device, then builds + installs
|
|
ANDROID_DEVICE=emulator-5554 just install # build + install on that device
|
|
ANDROID_DEVICE=emulator-5554 just uninstall
|
|
just clean
|
|
```
|
|
|
|
`ANDROID_DEVICE` is the serial `adb devices` reports. Every recipe that
|
|
installs, uninstalls or fuzzes acts on that serial. Without it they print the
|
|
online devices and ask which one to use, and pick on their own only when the
|
|
one device adb can see is an emulator on the local adb server. With no terminal
|
|
to ask on, a CI job say, the ask becomes a refusal that prints what adb sees.
|
|
|
|
A run installs the app, clears its state and drives it, which is not something
|
|
to do to a handset that happens to be the one thing plugged in. An emulator is
|
|
cheap to rebuild, so a lone local one is the single case worth guessing at.
|
|
|
|
## iOS
|
|
|
|
```sh
|
|
just ios # asks which simulator, unless one is booted
|
|
IOS_DEVICE="iPhone 15" just ios # name it, by name or UDID
|
|
```
|
|
|
|
`IOS_DEVICE` follows the same rule as `ANDROID_DEVICE`: a lone booted simulator
|
|
is taken without asking, anything else is asked about, and with no terminal to
|
|
ask on it refuses and lists what is installed. A name is matched against booted
|
|
simulators first and available ones second, and it can name several, since the
|
|
same iPhone exists under every installed runtime. When it does, you pick which,
|
|
and everything after that addresses the chosen UDID: the build destination, the
|
|
install, the launch and `--ios-device` all get the one simulator.
|
|
|
|
`just ios` regenerates `app/iosApp/iosApp.xcodeproj` from `app/iosApp/project.yml`,
|
|
builds the KMP framework (`Shared.framework` from `:app:shared`), links it
|
|
into the SwiftUI host, uninstalls any previous copy, installs, and launches.
|
|
The uninstall matters: folio's signed-in session survives an install over the
|
|
top, so without it a run opens on the last run's Home screen.
|
|
|
|
## Web
|
|
|
|
```sh
|
|
just web # webpack dev server with COOP/COEP headers
|
|
just web-build # produce a webpack distributable bundle
|
|
```
|
|
|
|
`just web` runs `:app:webApp:wasmJsBrowserDevelopmentRun --continuous`, so
|
|
edits to shared code reload in the browser.
|
|
|
|
## Demo credentials
|
|
|
|
```
|
|
email: [email protected]
|
|
password: ledger123
|
|
```
|
|
|
|
## Run a sanderling test (Android)
|
|
|
|
```sh
|
|
ANDROID_DEVICE=emulator-5554 just test
|
|
```
|
|
|
|
The same naming rule as `just install` applies, and `just test` asks once for
|
|
the whole run. If nothing is attached at all, `just test` boots a bootable AVD
|
|
and runs against that. With multiple AVDs, pick one:
|
|
|
|
```sh
|
|
AVD=Pixel_7 just test
|
|
```
|
|
|
|
Persistent settings can live in `.env` alongside the justfile:
|
|
|
|
```
|
|
ANDROID_DEVICE=emulator-5554
|
|
DURATION=5m
|
|
```
|
|
|
|
The device does not have to be attached to this machine. `ADB_SERVER_SOCKET`
|
|
aims adb at another host's adb server, and `ANDROID_DEVICE` names the serial
|
|
that server reports. A remote server is shared, so nothing there is ever picked
|
|
without being named or asked about, one device on it or twenty:
|
|
|
|
```
|
|
ADB_SERVER_SOCKET=tcp:10.0.0.5:5037
|
|
ANDROID_DEVICE=emulator-5556
|
|
```
|
|
|
|
The older `ANDROID_ADB_SERVER_ADDRESS` / `ANDROID_ADB_SERVER_PORT` pair works
|
|
too, at the same precedence the adb CLI gives it. A remote server is never
|
|
auto-booted against: when it reports no device, `just test` says so rather than
|
|
starting a local emulator that server will never see.
|
|
|
|
Gradle only assembles the APK. The install goes through adb, which reads those
|
|
variables, so a remote server needs nothing else. Gradle's own `installDebug`
|
|
cannot be used here: its adb client only ever dials loopback.
|
|
|
|
Traces land in `./sanderling/runs/<timestamp>/`.
|
|
|
|
## Run with the LLM action generator
|
|
|
|
The same `sanderling/spec.ts` runs under either generator: `--generator seeded`
|
|
(the default weighted fuzzer) or `--generator llm`, where a vision model picks
|
|
from the SAME weighted candidate set (reading the screenshot plus a numbered,
|
|
weight-annotated list of concrete actions) and returns one number. The spec's
|
|
`generator = llm({ model, instructions })` export configures it.
|
|
|
|
```sh
|
|
export OPENROUTER_API_KEY=sk-or-... # or OPENAI_API_KEY=sk-... for OpenAI direct
|
|
just test-llm # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio
|
|
```
|
|
|
|
OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor
|
|
prefix from the model id in `spec.ts` (`gpt-5.4-nano`, not
|
|
`openai/gpt-5.4-nano`). The model must support image input **and** strict
|
|
`json_schema` structured outputs. Each step is one multimodal call, so keep the
|
|
duration / step budget modest. The trace records the model's reasoning, the
|
|
chosen number, and `source: "llm"` on each action, so the replay UI shows why
|
|
each pick was made.
|
|
|
|
## Run a sanderling test (iOS)
|
|
|
|
```sh
|
|
just test-ios # asks which simulator, unless one is booted
|
|
IOS_DEVICE="iPhone 15" just test-ios # name it, by name or UDID
|
|
```
|
|
|
|
`just test-ios` settles on a simulator once for the whole run, boots it if
|
|
needed, runs `just ios` to install and launch the app, then invokes `sanderling
|
|
test --platform ios`. Same `DURATION`, `SEED`, and `OUTPUT` env vars as the
|
|
Android target.
|
|
|
|
A physical iPhone is a different target: `just test-ios-device` requires
|
|
`IOS_DEVICE` to name it, and says so rather than running, because an empty one
|
|
resolves to a booted simulator and would fuzz that instead.
|
|
|
|
## How it connects to sanderling
|
|
|
|
- Each screen sets a stable Compose `testTag` (`HomeScreen`, `AccountCard`,
|
|
`LedgerRow`, `TxnAmount`, ...). The Sanderling SDK resolves `testTag` to
|
|
`resource-id` on Android and `accessibilityIdentifier` on iOS.
|
|
- Identity for list items is the visible text content (account name; txn
|
|
note + amount). No synthetic IDs encoded in semantics.
|
|
- `contentDescription` is reserved for real accessibility labels, never as a
|
|
data carrier.
|
|
- `sanderling/spec.ts` imports `@sanderling/spec`, reads state via `s.ax.*`,
|
|
asserts properties, and weights the actions the fuzzer picks from.
|
|
- `just test` invokes `sanderling test` against the installed APK.
|