Files
pj 9ab59365b2 redact passwords only, and say what each step did in the log (#93)
* feat(hierarchy): name the route a native tree shows

The screen name was web-only: the Chrome driver stamps sanderling-screen on
the root and nothing else does, so every Android and iOS step recorded and
logged an empty screen. The route marker the tree already carries (the
resource id ending in Screen, the same one Transitional counts) names it.

* feat(runner): say what each step did in the step log

One line per step carried only an index and a node count. It now names the
screen, the action, its target and the typed value, the last through the
same redaction the trace and the prompt use. Emitted after the apply so the
line reports what actually happened, skip reason included.

* fix(sidecar): state on android whether a field is a secure entry

maestro's tree mapper copies a fixed attribute list off the device's XML and
password is not on it, so no android element ever reported the fact and the
conservative rule downstream redacted every typed value in the trace, the
prompt and the log. The XML still carries it: re-read it once per settled
snapshot and state the fact on the text fields it matches. A field it cannot
match stays unstated, which still reads as a credential.

* docs: correct the record that android never reports a secure field

Four places said android reports the fact for nothing and that every typed
value there is redacted. The sidecar now states it, so they described the
old behaviour.

* test(sidecar): fail the build if maestro renames the call the fact comes from

* fix(sidecar): state the fact on a field named by its hint alone

collectTextFields matched on class only, so a node the go side calls editable
off its hintText was left unstated and its typed value redacted.

* docs: record that ios and web state secure:false for compose password fields

Both derive the fact from a widget type a compose app never has, so the
value reaches the trace in the clear. Verified on folio on both targets.

* feat(android): read the application id out of an apk

parses the compiled AndroidManifest.xml rather than shelling out to
aapt2, which lives in the versioned build-tools directory that hosts
with only platform-tools never install.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* feat(cli): let --android-app-path supply the bundle id

--bundle-id stays required everywhere else, and an explicit one still
wins, so the apk can never quietly override what was asked for.

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* docs: record that the apk can name the package itself

Claude-Session: https://claude.ai/code/session_012PVErdr3ZzyUASeVQDWsUc

* fix(testrun): pass the jvm the flag that silences the jdk 24 unsafe warning

* feat(folio): ask which android device to run on when none is named

* feat(folio): pin ios recipes to one simulator udid and ask when several match

* docs(ci): say how just ios lands on the simulator the boot step chose

* chore(folio): ignore run output anywhere under examples/folio

* refactor(hierarchy): name no screen for a tree Transitional calls a cross-fade

ScreenName kept its own reading of the route markers and disagreed with
Transitional on a marker repeated by a nested node: it named the screen on
a step the runner was skipping as unsettled. One reading now.

* refactor(sidecar): inline the one attempt passed to callViewHierarchy

A named constant and its own comment for a literal used once.

* test(sidecar): compare the whole tree when checking the annotation changes nothing else

The old assertions checked one id string and one bounds value, and passed
with every other attribute stripped off every node. Now the annotated tree
minus the two facts it stated must equal the input.

* fix(sidecar): match a field to its xml node by class as well as id and bounds

A wrapper drawn to the same bounds as the untagged field inside it shared
the field's key, both were dropped as ambiguous, and every value typed into
an untagged field was redacted. The class tells them apart.

* docs(runs): record what the android hierarchy re-read costs per step

Two 1m runs per binary on folio, same seed, before and after the re-read.
2026-09-05 23:54:12 +05:30

171 lines
6.6 KiB
Markdown

# Folio
A minimal Kotlin Multiplatform personal-ledger app: login with demo
credentials, create accounts, add credits and debits. Shared UI across
Android, iOS, and web (wasmJs via Compose for Web). Doubles as the
example sanderling runs its property-based specs against.
## Stack
- Kotlin Multiplatform + Compose Multiplatform (shared UI)
- SQLDelight for the data layer (unified across platforms)
- kotlinx.coroutines for state flows
- kotlinx.serialization for `@Serializable` route types
## Prerequisites
- `just`
- JDK 17
- Android SDK (auto-discovered under `$ANDROID_HOME`, `~/Library/Android/sdk`,
or the Homebrew cask)
- Xcode 16+ and `xcodegen` (`brew install xcodegen`) for iOS
## Android
```sh
just install # asks which device, then builds + installs
ANDROID_DEVICE=emulator-5554 just install # build + install on that device
ANDROID_DEVICE=emulator-5554 just uninstall
just clean
```
`ANDROID_DEVICE` is the serial `adb devices` reports. Every recipe that
installs, uninstalls or fuzzes acts on that serial. Without it they print the
online devices and ask which one to use, and pick on their own only when the
one device adb can see is an emulator on the local adb server. With no terminal
to ask on, a CI job say, the ask becomes a refusal that prints what adb sees.
A run installs the app, clears its state and drives it, which is not something
to do to a handset that happens to be the one thing plugged in. An emulator is
cheap to rebuild, so a lone local one is the single case worth guessing at.
## iOS
```sh
just ios # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just ios # name it, by name or UDID
```
`IOS_DEVICE` follows the same rule as `ANDROID_DEVICE`: a lone booted simulator
is taken without asking, anything else is asked about, and with no terminal to
ask on it refuses and lists what is installed. A name is matched against booted
simulators first and available ones second, and it can name several, since the
same iPhone exists under every installed runtime. When it does, you pick which,
and everything after that addresses the chosen UDID: the build destination, the
install, the launch and `--ios-device` all get the one simulator.
`just ios` regenerates `app/iosApp/iosApp.xcodeproj` from `app/iosApp/project.yml`,
builds the KMP framework (`Shared.framework` from `:app:shared`), links it
into the SwiftUI host, uninstalls any previous copy, installs, and launches.
The uninstall matters: folio's signed-in session survives an install over the
top, so without it a run opens on the last run's Home screen.
## Web
```sh
just web # webpack dev server with COOP/COEP headers
just web-build # produce a webpack distributable bundle
```
`just web` runs `:app:webApp:wasmJsBrowserDevelopmentRun --continuous`, so
edits to shared code reload in the browser.
## Demo credentials
```
email: [email protected]
password: ledger123
```
## Run a sanderling test (Android)
```sh
ANDROID_DEVICE=emulator-5554 just test
```
The same naming rule as `just install` applies, and `just test` asks once for
the whole run. If nothing is attached at all, `just test` boots a bootable AVD
and runs against that. With multiple AVDs, pick one:
```sh
AVD=Pixel_7 just test
```
Persistent settings can live in `.env` alongside the justfile:
```
ANDROID_DEVICE=emulator-5554
DURATION=5m
```
The device does not have to be attached to this machine. `ADB_SERVER_SOCKET`
aims adb at another host's adb server, and `ANDROID_DEVICE` names the serial
that server reports. A remote server is shared, so nothing there is ever picked
without being named or asked about, one device on it or twenty:
```
ADB_SERVER_SOCKET=tcp:10.0.0.5:5037
ANDROID_DEVICE=emulator-5556
```
The older `ANDROID_ADB_SERVER_ADDRESS` / `ANDROID_ADB_SERVER_PORT` pair works
too, at the same precedence the adb CLI gives it. A remote server is never
auto-booted against: when it reports no device, `just test` says so rather than
starting a local emulator that server will never see.
Gradle only assembles the APK. The install goes through adb, which reads those
variables, so a remote server needs nothing else. Gradle's own `installDebug`
cannot be used here: its adb client only ever dials loopback.
Traces land in `./sanderling/runs/<timestamp>/`.
## Run with the LLM action generator
The same `sanderling/spec.ts` runs under either generator: `--generator seeded`
(the default weighted fuzzer) or `--generator llm`, where a vision model picks
from the SAME weighted candidate set (reading the screenshot plus a numbered,
weight-annotated list of concrete actions) and returns one number. The spec's
`generator = llm({ model, instructions })` export configures it.
```sh
export OPENROUTER_API_KEY=sk-or-... # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio
```
OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor
prefix from the model id in `spec.ts` (`gpt-5.4-nano`, not
`openai/gpt-5.4-nano`). The model must support image input **and** strict
`json_schema` structured outputs. Each step is one multimodal call, so keep the
duration / step budget modest. The trace records the model's reasoning, the
chosen number, and `source: "llm"` on each action, so the replay UI shows why
each pick was made.
## Run a sanderling test (iOS)
```sh
just test-ios # asks which simulator, unless one is booted
IOS_DEVICE="iPhone 15" just test-ios # name it, by name or UDID
```
`just test-ios` settles on a simulator once for the whole run, boots it if
needed, runs `just ios` to install and launch the app, then invokes `sanderling
test --platform ios`. Same `DURATION`, `SEED`, and `OUTPUT` env vars as the
Android target.
A physical iPhone is a different target: `just test-ios-device` requires
`IOS_DEVICE` to name it, and says so rather than running, because an empty one
resolves to a booted simulator and would fuzz that instead.
## How it connects to sanderling
- Each screen sets a stable Compose `testTag` (`HomeScreen`, `AccountCard`,
`LedgerRow`, `TxnAmount`, ...). The Sanderling SDK resolves `testTag` to
`resource-id` on Android and `accessibilityIdentifier` on iOS.
- Identity for list items is the visible text content (account name; txn
note + amount). No synthetic IDs encoded in semantics.
- `contentDescription` is reserved for real accessibility labels, never as a
data carrier.
- `sanderling/spec.ts` imports `@sanderling/spec`, reads state via `s.ax.*`,
asserts properties, and weights the actions the fuzzer picks from.
- `just test` invokes `sanderling test` against the installed APK.