Files
pj 7343085614 llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker

* feat(spec): make llm marker inert on the JS picker

* feat(spec): expose __sanderlingSampleInput__ corpus draw

* feat(openrouter): minimal chat-completions client

* test(openrouter): cover request shape, parse, and errors

* feat(verifier): thread screenshot + capture corpus sampler

* feat(verifier): LLM accessors — candidates, config, sampler

* test(verifier): cover AllCandidates, LLMConfig, SampleInput

* feat(trace): record action Source and LLMReasoning

* feat(runner): thread step screenshot into PushSnapshot

* feat(runner): llmSource selects actions via OpenRouter

* feat(runner): wire llmSource selection and trace stamping

* test(runner): cover llmSource selection, mapping, downscale

* docs(folio): add llm action-backend example spec

* docs(folio): document the LLM action backend run

* feat(llmclient): support OPENAI_API_KEY, openrouter wins

* refactor(runner): rename openrouter package to llmclient

* docs: both api keys, example model gpt-5.4-nano

* docs: add pr style rules to claude.md

* fix(runner): explain action kinds in llm prompt to stop swipe loops

* feat(trace): record llm ranked list and chosen rank

* feat(runner): stamp llm ranked list and chosen rank on trace

* fix(runner): tap by selector to survive layout shift after observe

* revert(runner): drop selector-first tap; broke path/testTag selectors

* feat(spec): llm() accepts optional instructions

* feat(verifier): read llm instructions off config

* feat(runner): append spec instructions to llm system prompt

* docs(folio): describe app in llm spec instructions

* feat(bundler): map generator export to globalThis.generator

* feat(verifier): read llm config off globalThis.generator

* feat(runner): gate llm source on --generator flag

* feat(cmd): add --generator llm|seeded flag

* test: cover --generator flag parsing and pickSources gating

* feat(verifier): enumerate llm candidates by walking actionsRoot

collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.

* test(verifier): cover candidate enumeration walk

* feat(verifier): add SetupAction to walk setup without the seeded root

* test(verifier): cover SetupAction setup-only precedence

* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order

* feat(trace): record llm choice number and chosen_action echo

* feat(runner): llm picks one number from weighted candidates

drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).

* test(runner): cover choice schema, strict-skip, and setup precedence

* refactor(verifier): drop the superseded AllCandidates enumeration

* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts

* fix(verifier): label editable fields by hint, not the typed value

an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.

* test(runner): cover weight-suffixed echo and stripWeightSuffix

* fix(runner): accept chosen_action echo that carries the weight suffix

real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).

* fix(verifier): skip llm enumeration on cross-fade frames

a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.

* feat(folio): show current balance on the add-transaction screen

renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.

* fix(replay): derive device space from screen extent, not first node

the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.

* fix(folio): show balance as a compact one-line label

per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.

* fix(folio): move balance into the header, one compact line under the account name

* fix(replay): attribute deferred violations to the causing step, not detection

* fix(replay): show a step's own violations in both panels, no next-step bleed

* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check

* chore: ignore .playwright-mcp scratch output

* docs: document the llm generator and --generator flag

* docs(spec): correct the llm() comment; config reads off globalThis.generator

* docs: add pr description rules
2026-07-31 21:12:00 +05:30
..
2026-07-31 21:12:00 +05:30
2026-04-20 16:04:20 +07:00
2026-07-31 21:12:00 +05:30
2026-07-31 21:12:00 +05:30

Folio

A minimal Kotlin Multiplatform personal-ledger app: login with demo credentials, create accounts, add credits and debits. Shared UI across Android, iOS, and web (wasmJs via Compose for Web). Doubles as the example sanderling runs its property-based specs against.

Stack

  • Kotlin Multiplatform + Compose Multiplatform (shared UI)
  • SQLDelight for the data layer (unified across platforms)
  • kotlinx.coroutines for state flows
  • kotlinx.serialization for @Serializable route types

Prerequisites

  • just
  • JDK 17
  • Android SDK (auto-discovered under $ANDROID_HOME, ~/Library/Android/sdk, or the Homebrew cask)
  • Xcode 16+ and xcodegen (brew install xcodegen) for iOS

Android

just install      # build + install on a booted emulator / device
just uninstall
just clean

iOS

just ios                          # default device: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just ios   # pick a different simulator

just ios regenerates app/iosApp/iosApp.xcodeproj from app/iosApp/project.yml, builds the KMP framework (Shared.framework from :app:shared), links it into the SwiftUI host, installs, and launches.

Web

just web         # webpack dev server with COOP/COEP headers
just web-build   # produce a webpack distributable bundle

just web runs :app:webApp:wasmJsBrowserDevelopmentRun --continuous, so edits to shared code reload in the browser.

Demo credentials

email:    [email protected]
password: ledger123

Run a sanderling test (Android)

just test

If no device is connected, sanderling boots the single AVD it finds. With multiple AVDs, pick one:

AVD=Pixel_7 just test

Persistent settings can live in .env alongside the justfile:

AVD=Pixel_7
DURATION=5m

Traces land in ./sanderling/runs/<timestamp>/.

Run with the LLM action generator

The same sanderling/spec.ts runs under either generator: --generator seeded (the default weighted fuzzer) or --generator llm, where a vision model picks from the SAME weighted candidate set — reading the screenshot plus a numbered, weight-annotated list of concrete actions — and returns one number. The spec's generator = llm({ model, instructions }) export configures it.

export OPENROUTER_API_KEY=sk-or-...   # or OPENAI_API_KEY=sk-... for OpenAI direct
just test-llm                         # or: sanderling test --generator llm --spec sanderling/spec.ts --bundle-id app.folio

OpenRouter wins when both keys are set. With a plain OpenAI key, drop the vendor prefix from the model id in spec.ts (gpt-5.4-nano, not openai/gpt-5.4-nano). The model must support image input and strict json_schema structured outputs. Each step is one multimodal call, so keep the duration / step budget modest. The trace records the model's reasoning, the chosen number, and source: "llm" on each action, so the replay UI shows why each pick was made.

Run a sanderling test (iOS)

just test-ios                          # default simulator: iPhone 17 Pro
IOS_DEVICE="iPhone 15" just test-ios   # pick a different simulator

just test-ios boots the simulator if needed, runs just ios to install and launch the app, then invokes sanderling test --platform ios. Same DURATION, SEED, and OUTPUT env vars as the Android target.

How it connects to sanderling

  • Each screen sets a stable Compose testTag (HomeScreen, AccountCard, LedgerRow, TxnAmount, ...). The Sanderling SDK resolves testTag to resource-id on Android and accessibilityIdentifier on iOS.
  • Identity for list items is the visible text content (account name; txn note + amount). No synthetic IDs encoded in semantics.
  • contentDescription is reserved for real accessibility labels, never as a data carrier.
  • sanderling/spec.ts imports @sanderling/spec, reads state via s.ax.*, asserts properties, and weights the actions the fuzzer picks from.
  • just test invokes sanderling test against the installed APK.