mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
033b17a1028a3c72ba962fc5833bc25dea127350
39
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c60e4ac2f4 |
feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared on the SDK-conditional search path. New ios-device/test-ios-device recipes mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key. |
||
|
|
406b7516b3 |
iOS simulator driver: Go-native companion-backed backend (#62)
* perf(ios): use prebuilt XCTest runner to cut startup * chore(ioscompanion): add companion asset prepare script * feat(ioscompanion): embed and extract simulator companion bundle * test(ioscompanion): cover companion stub and embedded extraction * docs: add third party notices for vendored companion * chore: ignore vendored companion bundle artifact * build(proto): pin simulator companion proto v1.1.8 * build(proto): add dedicated buf module and gen template for pinned proto * build(proto): exclude pinned companion proto from root buf workspace * feat(ioscompanion): commit generated companion gRPC stubs * feat(ioscompanion): map flat companion describe dump to TreeNode JSON * test(ioscompanion): add hierarchy-map golden and unit tests * feat(ioscompanion): port screen-settle stability polling to Go * test(ioscompanion): cover settle transitional, hash, streak, and cap rules * feat(ioscompanion): add USB HID keymap module * test(ioscompanion): cover keymap branches and paste-chord constants * build: embed companion assets via withcompanion tag * feat(ioscompanion): add transport companion interface * feat(ioscompanion): add HID event wrapper and builders * feat(ioscompanion): wire gRPC companion client and Dial * test(ioscompanion): cover HID builders and unit conversions * test(ioscompanion): cover Dial, process-state mapping, and install archive * test(ioscompanion): add gated simulator integration smoke test * feat(ioscompanion): text input and gesture HID composition with pasteboard fallback * test(ioscompanion): cover input composers, paste dialog loop, and pure helpers * feat(ioscompanion): add Describe to companion transport * feat(ioscompanion): implement DeviceDriver with companion supervision * test(ioscompanion): unit tests with fake companion transport * test(ioscompanion): gated companion smoke test * feat(ios): add ResolveTarget for simulator vs physical-device routing * feat(testrun): route iOS simulators through the native companion driver * refactor(testrun): defer the java preflight check to the physical-device path * feat(cli): add --ios-app-path flag * feat(doctor): split iOS checks into simulator and physical-device paths * test(folio): add gate-analyzer fixtures for G1-G5 * feat(folio): add iOS conformance gate script * chore(folio): wire gates recipe, app path, and ignore gate output * style: gofmt struct alignment drift * fix(doctor): probe simctl via xcrun instead of PATH lookup * fix(ioscompanion): spawn companion under driver-lifetime context * test(ioscompanion): prove companion child outlives startup context * fix(ioscompanion): chunk install payload under companion message cap * test(ioscompanion): cover install payload chunking * fix(ioscompanion): reinstall via simctl and sanitize companion env * fix(ioscompanion): wait out unresolved accessibility values after launch * perf(ioscompanion): paste long text for atomic landing * test(ioscompanion): cover paste threshold, retry flow, and sentinel detection * fix(ioscompanion): treat unresolved bridge values as transitional, never as content * fix(ioscompanion): accept masked secure-field values as paste landing * test(ioscompanion): cover sentinel mapping and masked-field landing * fix(ioscompanion): atomic erase and single-send paste to prevent doubling * test(ioscompanion): cover atomic erase, single chord, unverifiable field * fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout * test(ioscompanion): cover bridge-blackout paste verification * fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle * refactor(ioscompanion): name the empty-editable-field sentinel for what it is * perf(ioscompanion): tighten settle streak for the fast companion transport * feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt * refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt * test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests * fix(ioscompanion): retry describe past transient collapsed accessibility dumps * test(ioscompanion): cover collapsed-dump detection * perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses * perf(ioscompanion): tighten settle now that collapses are handled separately * fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text * test(ioscompanion): cover replace-on-input and TextReplacer capability * refactor(ioscompanion): neutralize HID events behind the transport seam * feat(companion): add simulator runner project skeleton * feat(companion): serve accessibility snapshots over the wire protocol * feat(companion): synthesize timestamped touch gestures * feat(companion): type text with replace semantics * feat(companion): serve the wire protocol from a parked runner * feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam * feat(ioscompanion): route text input through a text-editing companion when available * fix(companion): bind listener by port and source screen size from snapshot * feat(ioscompanion): add runner companion JSON transport * test(ioscompanion): cover runner transport protocol mapping * fix(companion): synthesize gestures synchronously to avoid the async completion crash * fix(companion): type on the main thread and recover from focus assertions * fix(companion): keep serving after an automation failure * refactor(companion): tidy snapshot serialization * fix(companion): honor sequential tap gaps and survive synthesis exceptions * feat(ioscompanion): expose native typing with an explicit replace flag * chore(companion): add runner asset prepare script * feat(ioscompanion): embed and extract the runner test bundle * test(ioscompanion): cover runner asset extraction * build(ioscompanion): commit runner asset archive * feat(ioscompanion): pair the legacy companion with the in-simulator runner * test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding * fix(ioscompanion): reconnect after interrupted runner calls instead of restarting * fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts * feat(companion): launch and terminate apps through the automation session * build(ioscompanion): refresh runner asset with session lifecycle * fix(ioscompanion): classify connection deadline expiry as caller budget * fix(companion): capture snapshots on the main thread inside the catch bridge * build(ioscompanion): refresh runner asset with main-thread snapshots * perf(ioscompanion): count read spans toward settle and capture snapshots concurrently * feat(ioscompanion): make the hybrid simulator companion the default * test(folio): cover runner-session orphans in the gate harness * test(ioscompanion): pin the child-lifetime test to the legacy path * fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears * fix(ioscompanion): pause the clear chord so selection applies before the delete * fix(companion): prune the keyboard subtree from snapshots * build(ioscompanion): refresh runner asset without keyboard elements * fix(ioscompanion): capture the screenshot transport before a recovery can reassign it * fix(companion): pin the runner listener to loopback * fix(companion): size the replace delete prefix to cover any focused field * build(ioscompanion): refresh runner asset with loopback bind and replace fix * fix(cli): cancel the run context on SIGINT so spawned children are reaped * fix(testrun): point the device java preflight hint at the ios-device doctor * fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob * test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision * chore: add test-companion target for the withcompanion-tagged suite * chore(ioscompanion): stop tracking the runner archive build artifact * build: produce the runner archive from source like the companion bundle * refactor(conformance): move the gate harness out of examples/folio * chore(folio): drop the gate harness wiring from the example app |
||
|
|
410602d2e1 |
Fix action-system audit findings (#60)
* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
|
||
|
|
19a470121f |
fix: iOS spec driving, InputText replace semantics, native DoubleTap (#58)
* fix(hierarchy): bounds-containment fallback for scoped and path queries Compose on iOS surfaces a testTag node as an empty leaf sibling of the content it labels instead of as an ancestor, so descendant search under the tagged node finds nothing and every path or scoped query returns null. When structural search yields no match, fall back to nodes whose bounds lie inside the scope node's bounds. * feat(sidecar): derive iOS clickable and editable from element type The XCTest hierarchy mapping dropped the element type, leaving no clickable or editable flags on iOS, so the fuzzer's tap and typing verbs never found a candidate inside the app. Map the raw accessibility tree directly and derive clickable, editable, scrollable, and class from the XCUIElementType raw value. * feat(proto): add EraseText RPC for InputText replace semantics * feat(driver): add EraseText to the device driver surface * fix(runner): erase existing field text before InputText InputText appended on native platforms, so repeated draws grew fields without bound. The folio fuzz run wedged on the add-account screen: each draw concatenated another name until the 40-character validation error became permanent. Replace semantics also makes retried typing idempotent. The web driver already replaced via select-all; native now matches. * feat(sidecar): EraseText backend support on android and ios * fix(folio): saturation-gate account creation in the spec The 2-3 step add-account loop outcompeted the 5-step transaction chain at every weighted re-draw, so runs filled with account creation and rarely exercised the balance properties. Stop offering add-account once three accounts exist; the renormalized weights then favor the transaction flow at every step of its chain. * fix(folio): author spec weights to match testing intent Revert the account saturation gate: it starved newAccountBalanceIsZero once it tripped, and a magic account count is app-state tuning, not intent. Instead weight the generators by what the properties need: the transaction chain leads, account creation stays exercised, and doubleTaps gets explicit weight everywhere because double-submission idempotency is what the spec is testing for. * fix(folio): lower doubleTaps weight to 5 * fix(sidecar): surface visible text on iOS static elements Static text and button strings live in the accessibility label on iOS, so the text attribute came through empty and every balance extractor parsed to zero, silently disarming both folio properties. Non-editable elements now fall back title, value, then label; editable fields keep value-only so an empty field's caption does not read as content. * feat(driver): native DoubleTap RPC for a tight inter-tap gap Composing two Tap round trips from the Go client spread the taps by hundreds of milliseconds on iOS, wide enough for the app to navigate between them, so double-submission races could never reproduce. The sidecar now lands both taps back-to-back next to the device transport. * feat(sidecar): pipeline iOS double-tap requests Queue the second tap at the XCTest runner while the first executes. The runner serializes handlers, so this is the tightest gap the transport allows (~350ms per tap round trip); recorded here with measurements for the iOS double-tap limitation. |
||
|
|
c5bb176be8 |
UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed, and surface both in describe() and MarshalJSON. * feat(ltl): negation normal form pass nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error leaf, dualizing Always<->Eventually and preserving bounds. * feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse Apply nnf on construction, reduce bounded Always symmetric to bounded Eventually (vacuous holds once the window closes), add Finalize to resolve undischarged liveness obligations to Violated at run end, and collapse structurally-identical pending obligations. * test(ltl): property-based NNF laws Lock double-negation identity, Always/Eventually duality with bound preservation, leaf pushdown, and not(always true) reaching Violated. * test(ltl): Finalize, bounded eventually, latch, collapse Property tests for monotonic violation latch and eventually-within violating iff n consecutive false, plus Finalize and collapse cases. * feat(inspect): within clause on always residual node A negated bounded eventually serializes as a bounded always; render its bound instead of dropping it. * feat(ltl): witness violations and (bool,error) predicate thunks * test(ltl): migrate thunk call sites to (bool,error) * feat(ltl): flag thrown-predicate witnesses with IsError * refactor(verifier): replace predicate err side-channel with violation witness * test(verifier): witness API for thrown predicates * feat(trace): witnesses map and skipped-verification marker on Step * feat(runner): thread violation witnesses, finalize, skip marker into trace * test(ltl): lock violation witness reason, IsError, and step * test(verifier): finalize surfaces unmet eventually with witness * fix(ltl): eliminate implies and bounded-always false-negatives Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent can no longer defer the whole implication and drop a consequent that was false at the current step. Carry a pending inner past a bounded-Always window close instead of dropping it to holds, so a deferred obligation is resolved by a later step or Finalize. * test(ltl): lock implies and bounded-always false-negative regressions * fix(web-runtime): seed PRNG for reproducible runs and align weighted pick * feat(testrun): inject seed into web bundle via SANDERLING_SEED define * test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring * test(spec): add Go math/rand/v2 PCG oracle and golden fixture * feat(spec): bit-exact PCG port of Go math/rand/v2 * test(spec): assert pcg.ts matches the PCG golden fixture * feat(spec): shared input corpus and press-key pools * feat(spec): action-tree types and Host interface * feat(spec): verb support matrix and warn-once helper * feat(spec): deterministic shared action picker * test(spec): verb matrix and warn-once semantics * test(spec): picker draw-order and determinism * refactor(spec): actions.ts returns pure GeneratorNode data trees * refactor(spec): wire from() sampling through the picker rng * feat(spec): shared runtime-entry installs next-action over pick.ts * feat(spec): export LongPress/Scroll/longPresses/scrolls factories * test(spec): assert data-tree shapes for action factories * test(spec): runtime-entry serializeAction wire-contract round-trip * refactor(spec): bridge data-tree nodes to the legacy goja picker tags * fix(spec): web runtime walks the spec's globalThis.actions data tree * test(spec): tolerate legacy bridge fields on builtin nodes * refactor(spec): installRuntime accepts a lazy root resolver The web bundle imports the runtime before the spec, so the action root on globalThis.actions only exists after the spec evaluates. Accept a function form so the goja and web hosts resolve the root per tick. * refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/ randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG, and the snake_case serializeAction) plus the __sanderling__ action factory binds. web-runtime now implements Host (platform/seedHi/seedLo from the injected 64-bit seed via BigInt, queryCandidates over the live DOM with a per-tick cache, reportUnsupported) and calls installRuntime so both engines run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts matrix instead of silently returning null. Keeps the DOM helpers (selector translation, queryElement, elementHandle, buildState, sanitize, extractors) and the global locking. Net -214 lines (741 -> 527). * test(spec): cover the WEB Host surface and seed precision Replace the deleted-picker tests with Host coverage: platform()==web, seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0, reportUnsupported warning, the installed next-action/extractor globals, and queryCandidates verb routing + per-tick caching over a querySelectorAll stub. * refactor(spec): picker emits native selector + scroll endpoints, setup precedence * feat(spec): goja runtime entry wires the shared picker over the Go host * feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin * feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker * refactor(spec): drop the legacy goja bridge fields from action factories * feat(spec): serialize selector-only string targets for the runner to re-resolve * refactor(verifier): one DecodeAction reads the unified flat wire contract * refactor(verifier): goja host + shared picker replace the duplicate Go picker * refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime * test(verifier): author specs through the shared picker path * test(runner): bundle authored specs with the goja runtime entry * feat(verifier): collect unsupported verbs for the run report * refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource * feat(testrun): surface unsupported verbs in run report * test(verifier): cross-runtime goja/node parity gate on the shared picker * test(verifier): unsupported verbs collected deduped in first-seen order * test(runner): summary reports no unsupported verbs on a clean run * test(spec): golden-fixture cross-runtime parity gate for the node picker Replace the env-driven parity harness with a shared scenario module and a committed golden the node picker asserts independently. The goja side asserts the same golden, so neither runtime invokes the other at test time. * test(verifier): assert goja picker against the same cross-runtime golden Drop the node-subprocess coupling: the goja side now installs a stub __sanderlingHost__ with the fixed candidate list and asserts the committed golden, matching pkg/spec/test/parity.test.ts. * refactor(spec): rename pressKey generator export to pressKeys * refactor(spec): update barrel re-exports for pressKeys * test(spec): update pressKeys generator export name * docs(spec): rename pressKey generator to pressKeys * refactor(spec): extract samplerRng into shared sampler-rng module * feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText) * test(spec): cover fluent value generators determinism and chaining * refactor(bundler): inject globalThis trailer from spec named exports * refactor(bundler): reuse registration trailer in web bundler * test(bundler): cover named-export globalThis registration * feat(spec): add named() to Extracted handle type * feat(web-runtime): named() and cross-extractor read guard * feat(verifier): named() and cross-extractor read guard in goja * test(verifier): cross-extractor read guard and named() * test(web-runtime): export runtime and extractors for tests * test(web-runtime): named() and cross-extractor read guard * refactor(folio): drop manual globalThis trailer (bundler injects it) * refactor(folio): seed txn amounts via integers().between(1,500) * refactor(folio-web): drop manual globalThis trailer (bundler injects it) * fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs * refactor(folio-web): weight valid generators against edgeCaseText for names/amounts * refactor(folio-web): name extractors so violation witnesses are readable * fix(web-runtime): propagate extractor getter throws and unpoison locked global Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake. * test(spec): install fake runtime via defineProperty to survive locked global * test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors * feat(runner): add MaxSteps bound to Options * test(runner): MaxSteps stops after exactly N steps * test(driverpb): drop proto getter round-trip tautology * test(sidecar): drop stub-mode placeholder tautology tests * test(mock): drop default-field-value assertion test * test(ltl): drop Verdict.String tautology tests * refactor(runner): extract RenderSummary for snapshot testing * test(runner): golden snapshots for trace stream and violation summary * feat(web-runtime): capture uncaught errors into state.exceptions * test(integration): add throwing and counter web fixtures * test(integration): add specs for the web fixtures * test(integration): drive web fixtures through the real pipeline in headless Chrome * chore(make): add test-browser target for the Chrome-driven suite * ci: run the Chrome-driven browser suite in a separate job * refactor(test): relocate browser suite to test/browser * refactor(permissions): delete dead internal/permissions package * refactor(test): rename package to browser_test * refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets * chore(make): point test-browser at test/browser * docs(decisions): record internal/permissions deletion * refactor(doctor): use sidecarassets package * refactor(testrun): use sidecarassets package * fix(test): resolve testdata relative to browser_test.go * refactor(verifier): remove dead __sanderlingIndex compat alias * refactor(bundler): use encoding/json for JS string literals * docs(action-space): use vendor-neutral native driver wording * refactor(hierarchy): scrub backend tool name from comments * refactor(driver): scrub backend tool name from comments * refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver * refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap * refactor(chrome): implement DoubleTap as two taps with the gap * refactor(mock): record DoubleTap and DoubleTapSelector actions * refactor(runner): delegate double-tap to driver, drop gesture timing * test(runner): assert double-tap delegates to driver DoubleTap * docs(cmd): add package docs to CLI and developer tools * docs(driver): add package docs to driver interface and chrome backend * docs(driver): add package docs to mock and sidecar backends * docs(platform): add package docs to android and ios device prep * docs: add package docs to bundler and inspect * docs(ltl): add package doc to temporal logic evaluator * docs: add package docs to runner and testrun pipeline * docs: add package docs to trace and verifier * docs(sidecarassets): add package doc for embedded JAR loader * fix(chrome): add disable-dev-shm-usage so Chrome starts in CI * test(chrome): gate real-Chrome driver tests behind the browser tag * chore(make): run chrome driver tests in the browser job * fix(web-runtime): guard global error listeners for non-browser hosts The module registered window error/unhandledrejection listeners at top level, which threw under Node (the spec-api test runner) where globalThis.addEventListener is absent. Register only when the API exists; the real browser run is unaffected. * ci(browser): re-enable unprivileged user namespaces for headless Chrome ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged user namespaces stops headless Chrome from opening its DevTools socket even with --no-sandbox, surfacing as the driver's 'websocket url timeout'. Relax the sysctl for the job and add a direct launch check so a future breakage shows Chrome's own stderr rather than an opaque driver timeout. * ci(browser): pin stable Chrome for the driver tests setup-chrome's default latest pulled a dev Chromium (150) whose remote debugging socket never came up under chromedp, while plain --dump-dom worked. Pin the stable channel, which the driver is tested against. * feat(defaults): add scroll and rebalance action weights Use relative-integer weights (taps/typing co-primary 100, scrolls 50, swipes 25, doubleTaps 10); the picker normalizes by their total. Adds scrolls to defaultActions as a first-class reveal behavior. * feat(defaults): trim scroll action weight wiring * fix(build): point sidecar jar ignore and embed paths at sidecarassets * test(defaults): drop stale longPresses re-export assertion longPresses is opt-in vocabulary, no longer re-exported from defaults/actions.ts since e0d3b20; its builtin resolution is already covered by api.test.ts. Trim the defaults test to scrolls, which is an actual default export. * fix(chrome): raise DevTools websocket read timeout to 60s Chrome cold-start on a loaded CI runner can exceed chromedp's 20s default for reading the DevTools websocket URL, flaking the browser tests with "websocket url timeout reached". Give launch more headroom. |
||
|
|
88db9653e5 |
refactoring default action layer (#51)
* feat(hierarchy): add editable signal with native derivation
* feat(chrome): emit editable flag in hierarchy dump
* feat(verifier): expose editable on ax element objects
* feat(spec): add editable to selector and element types
* feat(verifier): register typing builtin generator
* feat(verifier): typing builtin types edge-case corpus into editable fields
* feat(spec): export typing builtin generator
* feat(spec): add defaultActions bundle
* feat(spec): export @sanderling/spec/defaults subpath
* feat(folio): layer defaultActions breadth over targeted flows
* test(verifier): typing builtin targets editable fields, declines otherwise
* test(hierarchy): editable derivation and selector matching
* test(spec): defaultActions, typing, and defaults barrel resolve
* fix(testrun): alias @sanderling/spec/defaults for the bundler
* test(chrome): editable flag for inputs, textarea, contenteditable
* feat(spec): typing builtin for the web (V8) action path
* chore(folio): auto-boot a bootable AVD in just test/install when none connected
* feat(driver): add ForegroundChecker optional capability
* feat(android): detect foreground package via adb dumpsys
* feat(sidecar): implement ForegroundApp via adb for android
* feat(runner): relaunch app when foreground escapes during exploration
* fix(spec): drop hardware back from defaultActions to stay in-app
* feat(spec): add DoubleTap action type and constructor
* feat(spec): wire DoubleTap through web-runtime serializer
* feat(verifier): bind doubleTap and decode DoubleTap actions
* feat(runner): dispatch DoubleTap as two taps inside one step
* test(doubleTap): cover constructor, verifier round-trip, and runner dispatch
* feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action
* fix(folio): track ledger row count across non-ledger steps; pin reproducer seed
* feat(spec): add doubleTaps random-target builtin to defaultActions
* feat(verifier): add doubleTaps random-target generator
* refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions
* fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives
* feat(verifier): track newly-violated property set per step
Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.
* refactor(runner): emit onset-only violations to trace and summary
Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.
* style(verifier): use maps.Copy for verdict snapshot
* fix(folio): make login spec content-driven (idempotent across re-entries)
* fix(verifier): canonicalize selector strings
Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.
* refactor(folio): replace txn invariants with balanceMatchesAddedTxn
Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.
* refactor(trace): drop WriteScreenshotAfter
Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.
* refactor(runner): one concurrent screenshot per step
Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.
* refactor(inspect-ui): use next step's screenshot for state after
Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.
* feat(sidecar): structural-hash settle poll
Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.
* test(sidecar): cover pollUntilStable and structuralHash
Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.
* feat(spec): accept optional name on extract()
Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.
* test(spec): cover extract name overload
Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.
* feat(verifier): name extractors for diff surfacing
bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.
* chore(folio): name every extract() call
Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.
* feat(verifier): track extractor value transitions
Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.
* test(verifier): cover ChangedExtractors diffs
Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.
* feat(trace): emit extractor_changes per step
Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.
* feat(inspect-ui): render extractor-change breadcrumbs at violations
Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.
* fix(sidecar): cap stability poll independently of settle budget
The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.
* feat(cli): default --clear-data on so runs start fresh
* feat(sidecar): streak-based settle with route-transition detection
Two changes layered into the stability poll:
1. stabilitySnapshot returns null while the tree carries more than one
route-level Screen tag (resource-id / testTag / identifier ending
in "Screen"), so the poll cannot declare a NavHost cross-fade
stable. Apps following the Compose route convention get this
detection for free; apps that don't fall through to the generic
signal below.
2. pollUntilStable now requires an uninterrupted stable streak of at
least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
matches. A late transition that fires after a brief calm window
breaks the streak instead of slipping past. Interval widened to
250ms so UiAutomation isn't hammered under fuzz load.
* test(sidecar): cover streak reset and route-transition rejection
Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.
* feat(runner): re-fetch on transitional hierarchy capture
Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.
fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.
* feat(runner): gate first action on app reaching foreground
* test(runner): cover startup foreground gate and back-press
* feat(verifier): scope random-action targets to app package
Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.
* feat(testrun): pass app package into verifier scope filter
* test(verifier): cover package-scoped target selection
* feat(hierarchy): derive package from resource-id prefix
The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.
* test(hierarchy): cover package derivation from resource-id
* chore: stop tracking inspect-ui/dist build artifacts
* feat(android): detect focused-window package via dumpsys window
* feat(driver): add FocusedWindowChecker capability
* fix(runner): gate first observe on the app window being drawn, not just resumed
* test(mock): add FocusedWindowApp with foreground mirroring
* test(runner): cover startup gate waiting for app window to draw
* feat(proto): add Snapshot RPC for atomic hierarchy+screenshot
Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.
* feat(sidecar): add snapshot default on DriverBackend
Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.
* feat(sidecar): wire Snapshot handler with serialization lock
Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.
* test(sidecar): cover Snapshot wire path and serialization lock
SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.
* feat(driver): expose Snapshot on DeviceDriver and sidecar client
Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.
* feat(driver): add Snapshot to chrome and mock drivers
The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.
* refactor(runner): observe each step via the atomic Snapshot RPC
fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.
* test(runner): assert step uses Snapshot, not raw hierarchy/screenshot
TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.
* test(driver): cover Snapshot in proto descriptor and sidecar client
Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.
* feat(trace): add Transitional flag to Step
* fix(runner): skip verifier for transitional trees after retry budget
When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.
* test(runner): cover transitional step skips verifier and clean control
* refactor(trace): rename Step.Action to Step.NextAction
The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.
* refactor(runner): assign trace action to Step.NextAction field
Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.
* refactor(inspect): decode trace step's next_action JSON field
Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.
* test(inspect): update fixtures to use next_action trace field
Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.
* refactor(inspect-ui): rename Step.action to Step.next_action
Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.
* fix(folio): extract balanceMatchesAddedSum predicate as testable helper
Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.
* fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn
The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.
* test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases
Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.
* fix(build): rebuild sidecar JAR when Kotlin sources change
Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.
* fix(chrome): launch with no-sandbox so headless Chrome starts in CI
* fix(sidecar): type text at cursor instead of clearing the field
InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.
* test(sidecar): assert InputText types at cursor without clearing
Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.
* feat(proto): add LongPress RPC
* chore(proto): regenerate Go stubs for LongPress
* feat(driver): add LongPress to DeviceDriver interface
* feat(sidecar): add LongPress client method
* feat(mock): record LongPress action
* feat(chrome): implement LongPress as press-and-hold
* feat(sidecar): implement longPress across backends
* feat(sidecar): dispatch LongPress RPC to backend
* test(sidecar): cover LongPress dispatch
* test(sidecar): implement longPress in snapshot test backend
* feat(verifier): add LongPress and Scroll action kinds
* feat(folio-spec): predicate that gates balance check on TxnSubmit tap
Replaces the row-sum predicate (which always held by construction since
balance is derived from rows in Folio) with one that compares the typed
amount to the actual balance delta after a tap on TxnSubmit. Catches the
planted double-submit bug.
* feat(folio-spec): wire submitMovesBalanceByTypedAmount property
Adds lastAction and totalBalance extractors and uses them in the new
property. Drops ledgerRows/ledgerBalance extractors since nothing else
referenced them.
* feat(verifier): wire longPresses and scrolls generators
* test(verifier): cover longPresses and scrolls generators
* test(folio-spec): unit tests for submitChangesBalanceByTypedAmount
Covers single vs double submit, the DoubleTap variant, vacuous cases
(null action, wrong kind, wrong target, zero typed), and selector-as-
object coercion.
* feat(spec): add LongPress and Scroll authoring surface
* feat(spec): no-op LongPress and Scroll in web runtime
* feat(spec): re-export longPresses and scrolls as opt-in generators
* test(spec): cover LongPress and Scroll runtime members
* test(proto): expect LongPress in service descriptor
* feat(runner): dispatch LongPress and Scroll actions
* test(runner): cover LongPress and Scroll dispatch
* docs(action-space): move LongPress, Scroll, DoubleTap to current actions
* fix(runner): mark nil/empty hierarchy as transitional
A failed or empty sidecar hierarchy fetch was pushed straight to the
verifier, letting spec extractors crash with "Cannot read property 'map'
of undefined" when findAll returned null. Treat that case like a
transitional capture: skip the verifier push, still record the step, and
keep the loop progressing.
* fix(verifier): populate Action.On when tap chooser picks an element
Coordinate-targeted Taps/DoubleTaps left On empty, so action-gated
properties reading lastAction.on couldn't tell which target was hit and
were vacuously skipped. Resolve the picked element to a stable
key:value selector (resource-id, testTag, text, desc) and validate it
resolves back to the same element so we don't accidentally redirect the
tap to a sibling that shares the identifier.
* fix(folio): add parseTypedAmount helper matching app's parseCents
Raw user input like "50" must become 5000 cents, not 50. The existing
parseDollarCents helper strips non-digits and so reads "50" as 50 cents,
which is correct for formatted balance text but off by 100x for raw
input from the amount field.
* fix(folio): parse raw amount input as cents in submit predicate
txnAmountField holds raw user keystrokes, not formatted balance text.
Route it through parseTypedAmount so "50" reads as $50, matching how
the app commits the transaction.
* fix(folio): carry forward total balance across off-screen transitions
AddTransactionScreen shows neither AccountCard nor LedgerBalance, so the
extractor used to report 0 at the step before submit. That made every
non-zero current balance look like the full delta and tripped the typed
amount property on every honest submit. Remember the last-seen sum and
return it whenever the current snapshot has no balance signal.
* test(folio): cover submit predicate with raw typed-amount inputs
Pipes realistic raw keystrokes through parseTypedAmount + the predicate
so single submits clear and double submits fire as expected.
* feat(folio): add computeHomeTotalBalance helper
Pure helper that tracks Home multi-account total only and carries the last
Home sum across off-Home steps. Ledger's single-account balance is excluded
because mixing it would corrupt cross-screen scale comparisons.
* fix(folio): totalBalance carrier tracks only Home, not Ledger
Home cardSum is a multi-account total; Ledger's LedgerBalance is a single
account on a different scale. Blending them in the carrier produced bogus
cross-screen deltas (prev from Ledger, curr from Home), triggering false
positives in submitMovesBalanceByTypedAmount. Restrict the carrier to
Home AccountCard totals via the computeHomeTotalBalance helper.
* test(spec): cover computeHomeTotalBalance carrier behaviour
Tests Home sums, carrier passthrough on off-Home steps, the Ledger
scale-mismatch case, and a Home > off-Home > Home sequence.
* feat(runner): treat transient apply errors as transitional steps
Sidecar input RPCs occasionally hang with DEADLINE_EXCEEDED or
UNAVAILABLE on long fuzzing runs. The per-step loop previously
propagated any applyAction error and killed the run after a single
flake. Detect transient gRPC failures via status.FromError, mark the
step transitional, skip the post-action idle poll, and continue to the
next step. Fatal errors (outer ctx cancellation, non-transient codes,
verifier crashes) still propagate.
* test(runner): cover transient apply error resilience
TestRunner_TransientApplyErrorMarksTransitional drives the runner
through a wrapper that fails the first TapSelector with a gRPC
DeadlineExceeded then succeeds. Asserts the run does not exit, the
failed step is marked transitional with no violations, and the next
step runs cleanly. TestIsTransientApplyError_Classification covers the
helper's matching rules directly so future code changes don't quietly
drop a transient case.
* fix(folio): gate submit-balance property on Home route landing
totalBalance is only freshly computed when AccountCards are visible on
Home; off-Home landings return the carrier and would false-fire the
property, latching always(next(F)) to false and masking the real
double-submit bug. Skip vacuously when route is not "home".
* test(spec): cover route gate in submit-balance predicate
Adds route arg to existing cases (all use "home") and adds five new
cases: ledger landing with stale carrier, add-transaction with
double-insert delta, null route, plus home-landing positive and
double-insert negative cases anchoring the gate's allow path.
|
||
|
|
f572c8ba66 |
WIP: docs: refresh after iOS + web support (#50)
* docs: README covers iOS + web, surface both example apps * docs(cli): document --ios-device and per-platform doctor * docs: tighten README, fold examples into Docs list * docs(runs): correct --clear-data lifecycle wording Default behavior no longer wipes app data between runs; --clear-data is now opt-in. * docs(getting-started): add iOS path, separate folio and folio-web Document just test-ios under examples/folio, and distinguish the KMP sample from the React + Vite folio-web sample. * docs(inspect): document the eight panels Lists Screenshot, ActionList, Timeline, ViolationsPanel, HierarchyPanel, SnapshotTable, MetricsChart, ExceptionsPanel. Cross-links HierarchyPanel to the spec language reference. * docs(writing-specs): document setup export, flag noLogcatErrors as android-only Mirrors pkg/spec/README.md so the manual covers the runner's setup-first fall-through. Marks noLogcatErrors as Android-only so iOS/web spec authors know it silently no-ops. * docs(folio): document web target and iOS sanderling test recipe After the KMP refactor folio also runs on wasmJs and the justfile exposes just web, just web-build, and just test-ios. Surface all three. * docs(folio-web): add README Covers prerequisites, demo credentials, just test recipe, and how the React + Vite host exposes state to the sanderling spec via stable ids and data-* attributes. * docs: scrub driver-implementation name from user docs Drop the implementation tool name from README, cli.md doctor table, and spec-language.md. These docs should describe behaviour, not the specific underlying tool the native sidecar wraps. * docs(development): scrub driver-implementation name from dev docs architecture, design-principles, decisions now describe the native sidecar by role (gRPC surface over OS UI-test pipeline) rather than by the specific tool it wraps. |
||
|
|
b23fb0c723 |
feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag
Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.
* feat(testrun): add Preflight() before sidecar/driver setup
Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.
* refactor(chrome): split tag (HTML name) from class (CSS classList)
Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.
* feat(chrome): translate legacy string selectors to CSS/XPath
TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.
* feat(trace): add WriteHTML + Step.HTMLAvailable
Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.
* feat(driver): add WebDriver capability + chrome implementation
WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.
* feat(verifier): OverrideExtractorValues for V8-driven extractors
Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.
* feat(spec): add WebState + camelCase attribute aliases
WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.
* feat(runner): per-tick HTML capture for WebDriver-capable drivers
Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.
* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>
Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.
* feat(bundler): BundleWeb + V8-side runtime shim
web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.
* feat(runner): V8 extractor overrides + V8 action source for WebDriver
When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.
* feat(testrun): bundle + install web runtime when platform=web
BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.
* feat(inspect-ui): hierarchy + html panels in run detail
HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.
* fix(folio-web): drop aria-label data-carrier abuse
Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.
* chore: rebuild inspect-ui dist + folio-web .gitignore
Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.
* revert(trace): drop WriteHTML + Step.HTMLAvailable
Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.
* revert(runner): drop per-tick HTML capture
Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.
* revert(driver): drop WebDriver.Document
Document was only consumed by the runner's HTML capture which is gone.
* revert(inspect): drop /html route
Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.
* revert(inspect-ui): drop htmlUrl + html_available type
API surface no longer needs the HTML route; Step.html_available has no
producer.
* revert(inspect-ui): drop HtmlPanel + html tab
Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.
* test(inspect-ui): drop htmlUrl test, add @types/bun
Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.
* chore: rebuild inspect-ui dist without HtmlPanel
Embedded SPA bundle no longer ships the iframe HTML viewer.
* fix(web-runtime): retry action resolution + implement taps/swipes
V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.
* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys
Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.
* chore(folio-web): drop swipes from action root
Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.
* fix(inspect-ui): correct HierarchyPanel CSS variable names
Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.
* fix(chrome): correct PressKey mappings to chromedp/kb constants
Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.
Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.
* fix(cli): -h/--help exits 0 instead of error code
parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.
Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.
* fix(chrome): harden cssEscape for control chars + use [class~=]
Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.
Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.
* fix(web-runtime): use CSS.escape and validate tag-name selectors
The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.
The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.
Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.
* fix(chrome): validate attribute name in unknown-prefix branch
A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.
* fix(selectors): emit valid XPath 1.0 string literals via concat()
Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.
Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.
* fix(runtime): surface unresolved action targets instead of dropping silently
serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.
Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.
* fix(runner): use errgroup-bound ctx so siblings cancel on failure
The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.
* fix(chrome): propagate caller ctx cancellation to CDP calls
InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.
Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.
* fix(verifier): tolerate out-of-range override indices
A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.
Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.
* test(verifier): cover object-shaped extractor overrides
Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.
* fix(web-runtime): lock global runtime hooks against page shadowing
AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.
* perf(web-runtime): cache randomTap candidate DOM scan per tick
The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.
Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.
* fix(web-runtime): cap sanitize recursion to prevent stack overflow
State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.
* fix(web-runtime): enforce pressKey allowlist in factory
The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.
* chore(chrome): drop dead bundleSource/bundleMu
bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.
* fix(chrome): use strconv.Atoi for extractor key parsing
fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.
* fix(doctor): raise per-check timeout to 15s for chromium launch
5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.
* fix(runner): trust V8 coordinates for InputText, even at origin
resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.
Add applyAction tests covering both the typical web case and the (0,0)
edge case.
* test(bundler): lock down deterministic output across builds
The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
|
||
|
|
dd54c24c4e |
feat: --clear-data flag + typed attribute selectors (#48)
* feat(test): add --clear-data flag to clear app data on launch * test+docs: cover --clear-data flag in CLI parser test and reference * feat(spec): type AttrSelector with known attribute names Replace AttrSelector = Record<string, string> with KnownAttrSelectors plus a string|boolean index signature, so authors get autocomplete and type-checking on testTag / focused / clickable / etc. while raw driver attributes still type-check via the fallback. Boolean state attributes accept native booleans; goja stringifies them at the marshal boundary. AccessibilityElement.attrs becomes RawAttrs (typed string-valued shape of the same canonical names) so element.attrs.testTag autocompletes. * test(verifier): native boolean selector value matches focused=true * docs+folio: use native boolean for focused selector and document typed attrs |
||
|
|
c76745e5f1 |
WIP: folio refactor - KotlinConf-style production-app shape (#47)
* feat(hierarchy): testTag alias resolves to resource-id and accessibilityIdentifier
Compose's testTag surfaces as resource-id on Android and as
accessibilityIdentifier on iOS. Selectors written as
{ testTag: "Foo" } now match either, so Sanderling specs can use the
same tag on both platforms.
Also rounds out the iOS identifier aliases so resource-id /
identifier / accessibilityIdentifier all resolve to one another.
* chore(folio): add gradle/libs.versions.toml
Centralises versions for all folio modules ahead of the module split.
Adds new entries for kotlinx-serialization, navigation3, Metro, KSP,
and the JetBrains lifecycle-viewmodel-compose multiplatform artifact.
* refactor(folio): introduce nested KotlinConf-style modules
Split the monolithic :composeApp into :core, :app:shared,
:app:ui-components, and :app:androidApp. The old module is still
present and remains the source of truth until the next commits remove
it; both compile in parallel to keep iOS/Android builds green during
the cut-over.
Highlights:
- :core - SQLDelight schema + LedgerStore + Repository (now an
injectable class, not a singleton object). Methods are suspend to
match generateAsync = true.
- :app:ui-components - design system primitives. IconButton/AppButton
APIs revised: label = real contentDescription, testTag = stable
selector. Drops the data-carrier description argument.
- :app:shared - per-screen ViewModels colocated with screens; pure
composables on (state, onEvent); LocalAppComponent CompositionLocal
for hand-rolled DI; @Serializable Route. Hosts the iOS framework
(baseName Shared).
- :app:androidApp - thin Android entry that constructs the
DriverFactory and hands it to App().
- gradle/libs.versions.toml centralises versions; settings.gradle.kts
enables type-safe project accessors.
Deferred to follow-up PRs (per the design discussion):
- Metro DI: hand-rolled AppComponent for now; Metro graphs are mostly
ceremony for an app this size and add KSP/version risk.
- Navigation3: kept the existing Navigator-as-backstack class,
injected rather than singleton; nav3 isn't shipping a stable
multiplatform artifact for commonMain consumption yet.
- :app:webApp + OPFS sqlite worker: web persistence is real new
wiring (custom worker on @sqlite.org/sqlite-wasm). Master's
WebLedgerStore + Snapshot is being removed by this PR; web stays
buildable as a klib but no app-level wasm binary lands here.
* refactor(folio): delete :composeApp and retarget tooling
Removes the old monolithic module now that :core / :app:shared /
:app:ui-components / :app:androidApp own the source. Updates:
- justfile install/uninstall recipes -> :app:androidApp
- iosApp/project.yml framework path -> ../app/shared/...,
baseName Shared (was ComposeApp); pre-build script invokes
:app:shared:linkDebugFrameworkIosSimulatorArm64
- iosApp/iosApp/iOSApp.swift -> import Shared
- README -> mentions SQLDelight unification, drops the
data-carrier contentDescription notes (now stale), no Layout
section per repo convention
* refactor(folio-spec): query testTag and identify items by visible text
Replaces every accessibilityText / descPrefix data-carrier read with
testTag selectors that resolve to resource-id (Android) or
accessibilityIdentifier (iOS) via the SDK's alias table.
- Routes detected via testTag (LoginScreen, HomeScreen, etc.)
- Account identity = visible account name (no synthetic id encoded
in semantics).
- Ledger row identity = joined text content of the row.
- Active account derived from route alone (not parsed from
contentDescription).
- Focused input read from native focused="true" attribute, not from
a custom focused_input data carrier.
* fix(folio): build green on Android assemble + iOS framework link
- Drop ksp/metro/navigation3 plugin aliases - not actually applied
by any module in this PR (deferred follow-up).
- import awaitAsOne from app.cash.sqldelight.async.coroutines for
the suspend single-row reads enabled by generateAsync = true.
- Drop kotlin.js.ExperimentalWasmJsInterop opt-in from common
compilerOptions (it isn't valid for android/jvm targets).
- :app:shared androidMain pulls in androidx.activity:activity-compose
for the BackHandler actual.
* fix(folio): testTagsAsResourceId at App root + JS-bridge regression test
App.kt sets testTagsAsResourceId=true on the root Box semantics so
Compose's testTag surfaces as Android resource-id (and equivalent on
iOS via accessibilityIdentifier). Without this, testTag stays in the
Compose semantics tree but never reaches the runtime hierarchy that
UIAutomator and Sanderling read.
Also adds TestStateAxObjectSelectorTestTagAlias as a regression
test for the {testTag: ...} object selector resolving through the
SDK alias to resource-id at the JS bridge layer.
* test(verifier): expose PredicateError latching across steps
The runner logs PredicateError once per step. The current implementation
latches the first error per thunk, so the log freezes on step 1 forever
even when later steps would observe different errors. This test fails
today and locks in the contract: PredicateError must reflect the most
recent step.
* fix(verifier): refresh predicate errors per step
EvaluateProperties short-circuits once an Always-property latches to
violated, so the underlying goja predicate stops being called and
formula.err keeps whatever it threw at step 1. The runner logs
PredicateError every step a property is violated, which made every
subsequent log line repeat the step-1 throw. That looks like the spec
runtime is seeing stale state, but it is just stale error reporting.
EvaluateProperties now invokes every registered predicate once per step
purely to refresh formula.err. Verdicts are unaffected. The thunk
itself stops latching so the new value wins on whichever path runs first.
* chore(folio): add Metro DI plugin (1.0.0-RC4) to versions catalog
Adds dev.zacsweers.metro plugin alias and applies it to :core
as a smoke test. Compiler-plugin only, no KSP required.
* chore(folio): apply Metro plugin to :app:shared and :app:androidApp
* feat(folio-core): annotate Repository and SqlLedgerStore with @Inject
* feat(folio-core): scope Repository and SqlLedgerStore as @SingleIn(AppScope)
Both are app-wide singletons so the SqlDelight-backed flows remain
shared across the graph.
* feat(folio): annotate ViewModels with Metro @Inject / @AssistedInject
LedgerViewModel and AddTransactionViewModel use @AssistedInject for
their accountId param plus a nested @AssistedFactory; the rest are
plain @Inject constructor classes.
* feat(folio): introduce Metro AppGraph in commonMain
Single shared @DependencyGraph(AppScope::class) that exposes
Repository, Navigator, and ViewModels. LedgerDatabase enters the
graph via @DependencyGraph.Factory.create(database) so the suspend
DriverFactory.create() can stay outside the DI surface.
@Binds wires SqlLedgerStore to LedgerStore; Navigator is provided
explicitly so its Route.Home start state stays in DI rather than
relying on a default-parameter being honored by the graph.
* fix(folio): expect/actual testTagsAsResourceId so iOS link succeeds
Compose's androidx.compose.ui.semantics.testTagsAsResourceId is
Android-only. Calling it directly from commonMain broke
linkDebugFrameworkIosSimulatorArm64. Replace with an expect Modifier
extension that wires the semantics on Android and is a no-op on
iOS / wasmJs.
* refactor(folio): replace AppComponent with Metro AppGraph in App.kt
App now takes a suspend graph builder; the platform constructs
LedgerDatabase off the suspend DriverFactory.create() before invoking
the Metro graph factory. Routes resolve VMs through LocalAppGraph
instead of the hand-rolled LocalAppComponent.
Drops the loading-state placeholder comment (the empty Box is enough).
* refactor(folio): resolve ViewModels through LocalAppGraph in routes
Each *Route composable now reads the AppGraph from CompositionLocal
and pulls its VM via the appropriate accessor or AssistedFactory.
* refactor(folio): build AppGraph from platform entry points
MainActivity (Android) and MainViewController (iOS) now own the
suspend DriverFactory.create() and feed the resulting LedgerDatabase
into Metro's createGraphFactory<AppGraph.Factory>().
* chore(folio): add navigation-compose 2.9.2 dependency
Adds the JetBrains KMP navigation-compose library to the shared
module. Used in subsequent commits to replace the hand-rolled
Navigator with a typed-route NavHost.
* refactor(folio): replace custom Navigator with NavHost backstack
Wraps androidx.navigation.NavHostController behind the existing
push/replace/back surface so call sites in ViewModels stay unchanged.
App.kt now wires a typed NavHost with @Serializable Route entries
and observes the controller's currentBackStackEntry to drive the
session-based Login/Home redirect.
* fix(folio-core): wire kotlinx-browser so wasmJs DriverFactory compiles
org.w3c.dom.Worker on wasmJs lives in kotlinx-browser, not the stdlib.
Pin 0.5.0 alongside the @sqlite.org/sqlite-wasm 3.53.0-build1 version
that the upcoming web app will depend on, and switch the worker
constructor to the module-worker form that webpack expects.
* feat(folio): scaffold :app:webApp wasmJs module
Compose Multiplatform target that depends on :app:shared and pulls
@sqlite.org/sqlite-wasm 3.53.0-build1 as the npm runtime for the
SQLDelight web worker.
* feat(folio-webApp): add main entrypoint and index.html
main.kt mirrors the iOS entry point: builds DriverFactory + AppGraph
factory, hooks browser back-gesture into WebBackGesture, then mounts
the shared App composable into ComposeViewport.
* feat(folio-webApp): OPFS-backed sqlite worker + webpack config
sqlite.worker.js implements the SQLDelight web-worker protocol
(exec/begin_transaction/end_transaction/rollback_transaction) on top
of @sqlite.org/sqlite-wasm. Prefers the OPFS SAH pool VFS for
persistent storage and falls back to in-memory when OPFS is
unavailable.
webpack.config.d/coopcoep.js sends COOP/COEP headers on the dev
server so cross-origin isolation is available, even though the SAH
pool itself does not require it. webpack.config.d/sqlite-wasm.js
enables asyncWebAssembly so webpack can bundle sqlite3.wasm via the
'new URL("sqlite3.wasm", import.meta.url)' reference inside the
sqlite-wasm package.
* chore(folio): add web/web-build just recipes and refresh yarn lock
Yarn lock picks up @sqlite.org/sqlite-wasm 3.53.0-build1.
* fix(folio-core): probe schema before create on wasmJs
Wasm SqlDriver doesn't auto-track user_version like the Android
driver, so awaitCreate() ran on every page load and tripped over
already-created tables. Read PRAGMA user_version, run
awaitCreate/awaitMigrate based on it, and self-heal pre-existing
tables with version 0 by stamping the current schema version.
* chore(folio-webApp): pin dev-server port and trim worker logging
webpack-dev-server now binds 8088 (or WEBAPP_PORT) so it doesn't
collide with the docs server on 8080. Drop the per-message reply
log; keep only the OPFS init line and error logging.
* chore(folio): nest iosApp under app/ for KotlinConf parity
Match KotlinConf-app's filesystem layout where every entry point (android,
ios, web, shared, ui-components) lives under app/. iosApp is still an Xcode
project, not a Gradle module, so settings.gradle.kts is unchanged.
* feat(folio): testTag identity for AccountName and ledger row cells
Replaces string-heuristic identity in the spec extractors with stable
testTags. AccountCard exposes AccountName; LedgerRow exposes TxnNote
and TxnDate. Spec extractors read those directly instead of filtering
visible text by "starts with $" / "matches digit".
* fix(folio-app): branch start destination on initial session
Read repository.session.value at first composition and pick
Route.Home or Route.Login as the NavHost startDestination. Avoids
the one-frame Home flash on cold start with no persisted session.
* refactor(folio-webApp): hard-fail when OPFS unavailable
Drops the silent in-memory fallback. The README claims OPFS
persistence; falling back without surfacing the degrade made data
loss invisible across reloads. Now the worker errors out and the
Kotlin DriverFactory rejects the create() call instead.
* docs(verifier): document extractor advancement and refresh invariants
Extractor previous/current advance only on PushSnapshot, never per
thunk-call. refreshPredicateErrors depends on this for safe re-entry.
Also flags that re-invoked predicates run outside their LTL gate, so
they must be side-effect-free reads.
* fix(hierarchy): populate ResourceID from accessibilityIdentifier
iOS Compose surfaces testTag as accessibilityIdentifier. Previously
only resource-id and identifier seeded element.ResourceID, leaving
element.id empty for iOS Compose nodes and forcing specs to walk
attrs to recover stable identifiers.
* refactor(folio-spec): use element.id for focused field tag
Now that ResourceID populates uniformly across Android/iOS Compose,
the spec can read element.id directly instead of probing attrs for
each platform's underlying field name.
* fix(folio-spec): pick account card via seeded from(), not Math.random
Math.random() breaks --seed reproducibility. The verifier's seeded
RNG flows through from(), so re-running a seed now produces the
same card pick sequence.
* docs(spec): fix README example to use scoped extractors
The previous snippet referenced `state.ax.find` inside an actions()
body where state is not in scope, and shadowed the imported actions
helper with an export of the same name.
* feat(hierarchy): add FindBySelectorPath for chained object selectors
Each selector in the chain is matched within the descendants of the
previous match. Returns the deepest match (or nil) for FindBySelectorPath
and every deepest match for FindAllBySelectorPath.
* feat(verifier): dispatch JS array selectors to FindBySelectorPath
`state.ax.find([{...}, {...}])` now walks each segment scoped under
the previous match. Strings and single objects keep their existing
single-shot lookup.
* feat(spec): expose SelectorPath in find/findAll signatures
* fix(spec): satisfy AccessibilityElement interface in Tap test fixture
* refactor(folio-spec): collapse chained finds into selector paths
* feat(spec): add keyedBy(element, tags) identity helper
Joins element.find({testTag: tag})?.text per tag with U+001F as the
delimiter so user-visible text can never collide with the separator.
Returns empty string for an undefined element.
* refactor(folio-spec): use keyedBy for ledger row identity
* feat(spec): add whenRoute action gating helper
whenRoute(route, allowedRoutes, body) wraps an actions() generator
that returns [] unless route.current matches one of the allowed
values. Accepts a single route or an array.
* test(spec): cover whenRoute matching, gating, and array routes
* refactor(folio-spec): gate addAccount and addTxn with whenRoute
* refactor(hierarchy): use maps.Copy for attribute merge
Linter flagged the manual loop after recent edits surfaced the hint.
* feat(verifier): dispatch setup generator before actions root
Setup is consulted every step; when it returns ErrNoAction the call falls
through to the existing actionGenerator retry loop. This lets specs split
deterministic preconditions (login, onboarding) out of the weighted action
pool while auto-reengaging if state regresses (e.g. logout under fuzz).
* docs(spec): document setup precondition action generator
* refactor(folio-spec): export login as setup, remove from action pool
Login is deterministic and yields no actions once the app is past the
login screen; sitting at weight 50 in the action pool wasted half of step
picks on a no-op. Promote it to setup so the runner only consults it
while it has work to do, and rebalance remaining weights to round numbers
(addAccount 50, addTxn 40, back 10).
|
||
|
|
34fb73d6cb |
fix(folio): stop abusing contentDescription as data carrier (#45)
* feat(hierarchy): full-attribute selector system
- Add Attributes map to Element (raw platform attrs + serialized booleans)
- Add Selector / AttrFilter types for multi-filter AND matching
- Add matchAttr with alias expansion and substring/boolean semantics
- Add matchSelector (AND of all filters)
- id: and desc: keep exact/suffix/prefix semantics for backward compat
- text: widens to substring via matchAttr
- default: case routes unknown kinds to matchAttr (NEW)
- Add Tree.FindNode / Tree.FindAllNodes returning *Node
- Add Node.Find / Node.FindAll for scoped subtree string search
- Add Node.FindBySelector / Node.FindAllBySelector for object AND search
- Add attributeAliases for cross-platform name expansion
* test(hierarchy): full-attribute selector coverage
- raw resource-id: substring match
- label:/content-desc: alias expansion to accessibilityText on iOS
- scrollable:true/false boolean exact match
- title: iOS-only attribute, graceful nil on Android
- text: substring widening
- Selector AND: both filters must match; single miss returns nil
- Node.Find scoped search: descendants only, not siblings
* feat(verifier): object-form selectors + attrs + element-level find
- ax.find/findAll accept string or {attr:value} JS objects
- Object form builds Selector with AND semantics
- Returned element objects expose attrs sub-object (raw platform attrs)
- Returned element objects expose .find() and .findAll() scoped to subtree
- Element-level .find/.findAll accept string or object selectors
* feat(spec): extend AccessibilityElement and AccessibilityTree types
- AccessibilityElement gains attrs, find(), findAll()
- find/findAll on both Tree and Element accept string | AttrSelector
- AttrSelector = Record<string, string> for object-form AND matching
* feat(folio): migrate to chained object-form selectors
- Replace path queries (desc:X > desc:Y) with chained API
- Screen root lookups use { accessibilityText: "ScreenName" }
- Element-scoped searches use find/findAll with string or object
- Keep string selectors for descPrefix: and desc:Back (shows both forms)
* fix(folio): use account_card:id desc, expose balance via text semantics
* fix(folio): embed accountId in LedgerScreen desc, expose values via text semantics
* fix(spec): replace desc-parsing with text-based parseDollarCents extraction
|
||
|
|
776becdf4b |
Remove in-app SDK (#43)
* chore: delete internal/agent package
* chore(build): remove sdk-android from gradle settings
* chore(makefile): remove sdk-android targets
* chore(ci): remove release-android job from release workflow
* chore(folio): remove sdk-android dependency
* chore(folio): remove SDK initialization from FolioApplication
* chore(folio): delete snapshot extractor files
* feat(folio): add balance to account card content description
* feat(folio): add hierarchy content descriptions to LedgerScreen
* refactor(folio): rewrite spec.ts to use ax extractors
* docs: remove in-app SDK from README
* feat(folio): add focused_input indicator to App
* docs: remove in-app SDK from index
* refactor(runner): remove agent SDK connection and snapshot step
* test(runner): update tests for SDK removal
* docs: remove Android SDK section from getting-started
* refactor(testrun): remove agent SDK connection setup
* docs: remove snapshots from writing-specs
* docs: remove in-app SDK from architecture doc
* docs(folio): update README for SDK removal
* docs: update per-step cycle diagram in architecture doc
* fix(folio): detect screens from unique element presence, not id: selectors
testTag() in Compose is not exposed as resource-id without testTagsAsResourceId.
Use desc: selectors for elements unique to each screen instead of id: path queries.
* feat(folio): add screen root contentDescription for scoped ax selection
Each screen root gets semantics { contentDescription = "ScreenName" } so
sanderling specs can scope element lookups through the screen: desc:LoginScreen > desc:login_submit.
* fix(folio): scope all ax selectors through screen root nodes
Use desc:ScreenName > desc:element path queries so every selector is
rooted at the screen level. focusedInput stays unscoped since it lives
in the app root, outside any screen.
* fix(folio): guard newAccountBalanceIsZero against navigation false positives
Scoped selectors return [] when not on HomeScreen so accounts vanish and
reappear as apparently-new on each visit. Skip the check when prev was empty.
* chore(folio): link @sanderling/spec to local pkg/spec for IDE type checking
* feat(spec): add desc, class, clickable, enabled, checked, focused, selected to AccessibilityElement
Runtime fields set by the verifier were missing from the TypeScript type,
causing linting errors on el.desc and related accesses in specs.
* chore(folio): switch to bun, add tsconfig.json for IDE type checking
- Remove package-lock.json, add bun.lock
- Add tsconfig.json so VSCode resolves @sanderling/spec types
- Fix parseAccount/parseLedgerRow to accept string | undefined
|
||
|
|
6c32fb0e1d |
feat(hierarchy): flat-to-tree + path query selector (#41)
* feat(hierarchy): add Node tree + path query support Preserve parent-child relationships in a Node tree at parse time. Extend Find/FindAll with " > " path operator for scoped queries, e.g. id:LoginScreen > desc:EmailInput. Flat Elements slice and all public signatures unchanged. * test(hierarchy): add path query tests Cover single-level, multi-level, mixed-type, not-found, wrong-subtree, and FindAll across multiple roots. * feat(folio): add modifier param to Screen composable * feat(folio): testTag LoginScreen * feat(folio): testTag HomeScreen * feat(folio): testTag AddAccountScreen * feat(folio): testTag LedgerScreen * feat(folio): testTag AddTransactionScreen * feat(folio): scope spec selectors to screen testTags via path queries * fix(sidecar): align grpc-netty/grpc-okhttp with grpc-core version Maestro 1.40.0 brings grpc-netty:1.50.2 which compiled against AbstractManagedChannelImplBuilder, removed in grpc-core 1.64+. Gradle was upgrading grpc-core to 1.68.0 while leaving grpc-netty at 1.50.2, causing NoClassDefFoundError at startup. Force grpc-netty and grpc-okhttp to match grpcVersion so all gRPC artifacts are binary-compatible. * fix(sidecar): use ephemeral host port for AndroidDriver Hard-coded port 7001 caused TcpForwarder.start() to fail with TimeoutException when a previous sidecar process held the port open. Pick a free port via ServerSocket(0) instead. * chore(sidecar): bump maestro to 2.4.0, exclude graalvm from fat JAR Maestro 2.4.0 still ships grpc-netty:1.50.2 so the grpc transport resolution strategy is kept. GraalVM JS is excluded: unused by our sidecar and causes shadow JAR expansion errors (pom treated as zip). * fix(sidecar): update AndroidDriver calls for Maestro 2.4.0 API launchApp no longer accepts a UUID argument. Third constructor param changed from hostname to emulatorName so drop the "localhost" value. |
||
|
|
2a1b263b8c |
fix: WDA startup flakiness - warmup + connection drop message (#40)
* feat(ios): add simulator management package * feat(testrun): add iOS platform path (simctl launch + direct TCP) * feat(cli): add --ios-device flag and IosDevice option * feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser) * feat(folio-ios): wire SanderlingIos.start() in MainViewController * feat(folio-ios): add test-ios justfile recipe * fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings * fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects * feat(proto): add env map to LaunchRequest * feat(driver): add env param to Launch interface + all implementations * feat(testrun): launch iOS app via XCTest with env vars instead of simctl * feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through * feat(sidecar): wire env map in DriverService + IosDriverBackend in Main * fix(sidecar): use LocalIOSDevice (WDA+simctl) + stop before relaunch * fix(sidecar): include exception type in gRPC error description * fix(sidecar): pick free WDA port instead of hardcoded 9100 Use SocketUtils.nextFreePort to pick a free port in the 22000-23000 range rather than hardcoding 9100, which only worked if a previous WDA session left a listener there. * fix(sdk-ios): check semaphore wait result and throw on snapshot timeout dispatch_semaphore_wait returns nonzero on timeout; ignoring the return value caused pauseAndSnapshot to silently return an empty map, sending a garbage empty STATE frame to the host. Now throws so the agent loop reconnects instead. * fix(folio-ios): register snapshot extractors before starting agent SanderlingIos.start() was called before the snapshot objects were initialized, so a PAUSE message arriving early produced an empty snapshot. Move start() to after all extractors are registered. * chore(ios): remove dead LaunchApp function LaunchApp had no callers since |
||
|
|
97154cf580 |
feat(ios): launch via XCTest with env vars for hierarchy/tap access (#39)
* feat(proto): add env map to LaunchRequest * feat(driver): add env param to Launch interface and all implementations * feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through * feat(testrun): launch iOS app via XCTest with env vars instead of simctl * fix(ios): replace LaunchApp with BootedUDID; simctl launch moved to XCTest path * test(cli): add tests for ios platform flag and ios-device flag parsing * fix(sdk-ios): check semaphore wait result; resolve port from args and env; register extractors before start |
||
|
|
007dcddd69 |
feat(ios): iOS e2e support — Kotlin Native SDK + simulator driver + test-ios (#38)
* feat(ios): add simulator management package * feat(testrun): add iOS platform path (simctl launch + direct TCP) * feat(cli): add --ios-device flag and IosDevice option * feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser) * feat(folio-ios): wire SanderlingIos.start() in MainViewController * feat(folio-ios): add test-ios justfile recipe * fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings * fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects |
||
|
|
af7b7e27b0 |
feat(folio-web): web sample app + CDP spec tests (#36)
* fix(runner): allow nil connection for web platform * fix(testrun): skip SDK handshake for web platform * feat(folio-web): add React/Vite web sample app * feat(folio-web): add sanderling spec * fix(chrome): use InsertText for multi-char text input * feat(hierarchy): add Screen field populated from sanderling-screen attr * fix(chrome): auto-detect viewport from CSS vars, fix InputText accumulation, expose route as screen * fix(runner): fall back to hierarchy root screen when snapshot screen is empty * fix(folio-web): broaden loggedIn extractor to all authenticated pages |
||
|
|
46abb28ed6 |
feat(sdk-android): snapshot property delegate + feature-scoped files + screen fix (#34)
* rename web to inspect ui * feat(sdk-android): add camelToSnakeCase conversion * feat(sdk-android): add snapshot() property delegate with tests * feat(folio): add feature-scoped sanderling snapshot objects * refactor(folio): slim FolioApplication to snapshot object references * fix(folio): rename route snapshot to screen, update spec.ts |
||
|
|
faebfe379d |
test(folio): replace tautological properties with 4 domain invariants (#28)
* test(folio): replace tautological properties with 4 domain invariants Drop properties that can't fail (e.g. List.size >= 0, balances are Long integers) or that just check the extractor itself (account_count equals accounts.length). Keep auth routing liveness, error-clear liveness, and reachability goals. Add four properties that target the real code paths: - balanceMatchesTransactionDelta: a new ledger row shifts the balance by exactly its signed amount (credit +, debit -). - totalEqualsSumOfAccounts: home total equals sum of per-account balances at every state, not only on home. - balanceChangeRequiresActiveAccount: an account's balance can only change while that account is the active one in the navigator. - duplicateAccountNamesRejected: account names are unique under the app's actual dedup rule (case-insensitive after trim). * chore(folio): upgrade sdk-android to io.github.priyanshujain.sanderling:0.0.1-rc4 Group id moved from io.github.priyanshujain to io.github.priyanshujain.sanderling in the rc4 publish. * chore(folio): pin @sanderling/spec to 0.0.1-rc4 |
||
|
|
75780db3fa |
refactor(sdk-android): publish under io.github.priyanshujain.sanderling + ci docs fix (#27)
* refactor(sdk-android): publish under io.github.priyanshujain.sanderling Move the published Maven coordinates to a sanderling sub-namespace so the brand is visible in the dependency line (was io.github.priyanshujain:sdk-android). Sub-namespace is auto-allowed by Sonatype under the verified parent groupId. * chore: update sdk-android coordinates in folio + docs Follow the groupId change to io.github.priyanshujain.sanderling:sdk-android. * ci(docs): use --version flag for d2 install script The d2 installer accepts --version vX.Y.Z, not --tag. The --tag form was rejected as "unrecognized flag" on the docs workflow run after PR #26 merged. |
||
|
|
8ccf95c1cf |
refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling
Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.
* chore(proto): regenerate driverpb after module path rename
The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.
* refactor: rename CLI binary uatu -> sanderling
Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.
* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk
Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.
* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar
Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.
* refactor: rename Uatu API surface -> Sanderling
- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
__uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.
* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling
Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.
* chore(build): rename gradle property + rootProject.name uatu -> sanderling
- Renames the uatu.version gradle property and all its -P references in
Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.
* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1
Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.
* refactor: rename npm package @uatu/spec -> @sanderling/spec
Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.
* docs: rename uatu -> sanderling in README, docs, and URLs
- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
internal/inspect/server.go.
* refactor: rename remaining internal uatu strings -> sanderling
- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
|
||
|
|
a2e96af1af |
WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio Directory-level rename and path references in Go tests, bundle-check, top-level README, and getting-started docs. Package declarations, Gradle config, iOS bundle IDs, and class names follow in later commits. * refactor(folio): rename Kotlin package dev.uatu.sample to app.folio Moves source dirs and sqldelight schema from dev/uatu/sample to app/folio, updates package declarations and imports, and switches Android namespace/applicationId, iOS binaryOption bundleId, and sqldelight database packageName to the new identifier. * refactor(folio): rename SampleApplication to FolioApplication Android manifest now points at .FolioApplication with label 'Folio' instead of 'Uatu Sample'. * refactor(folio): set iOS bundle id and display name to Folio bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio. CFBundleName + CFBundleDisplayName -> 'Folio'. * refactor(folio): update demo email to [email protected] * refactor(folio): point justfile at app.folio bundle id Updates xcrun simctl launch target, uatu test --bundle-id, and the build/uninstall comments to reference folio instead of sample. * test: update fixture package ids to app.folio Sidecar activity-resolver test and verifier spec-integration XML fixtures referenced the old dev.uatu.sample Android package. Updates them to match the folio app's real package id so the tests stay representative of what the CLI sees on-device. * test(verifier): rename SampleApp identifiers to Folio Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and sampleAppHierarchyXML const (now loginHierarchyXML for consistency with the other per-screen fixtures). Updates trailing sample-app mentions in comments and assertion messages. * refactor(folio): rename Gradle/npm/wasm project identifiers to folio settings.gradle.kts rootProject.name, package.json + package-lock.json name, and the WasmJS index.html <title> all still read 'uatu-sample' / 'Uatu Sample'. Realigns them with the Folio brand. * docs(folio): rewrite README title + getting-started bundle id examples/folio/README.md is now titled 'Folio' with the Kotlin source paths corrected to app/folio. Getting-started example uses --bundle-id app.folio. Harness launch message is now generic ('app under test') since uatu-sample-harness is not specific to folio. * chore(folio): drop trailing 'sample' reference in gradle.properties * refactor(folio): rename LoginPage composable to LoginScreen Align with KMP/Android industry convention (NowInAndroid, Cash App, JetBrains samples use Screen, not Page). * refactor(folio): rename HomePage composable to HomeScreen * refactor(folio): rename AddAccountPage composable to AddAccountScreen * refactor(folio): rename LedgerPage composable to LedgerScreen * refactor(folio): rename AddTransactionPage composable to AddTransactionScreen * refactor(folio): split Models.kt into app.folio.data package Account, Transaction (with TxnType), and Session move into their own files under app.folio.data, matching NowInAndroid-style per-type organization. * refactor(folio): move data layer into app.folio.data package Repository, LedgerStore (expect + interface), SqlLedgerStore, WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext, and Snapshot move into app.folio.data. Update all consumer imports. * refactor(folio): move Navigation into app.folio.navigation package Split the former Navigation.kt into Route.kt (sealed interface) and Navigator.kt (singleton). Update consumer imports across screens, App.kt, and FolioApplication. * refactor(folio): move Platform and Format into app.folio.platform Both files carry expect declarations (Platform object, formatDate); grouping them into a dedicated platform package makes the KMP seam obvious and mirrors the structure used by JetBrains samples. * refactor(folio): move login into feature/auth package Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into LoginScreen since it is the sole caller. * refactor(folio): move HomeScreen into feature/home package * refactor(folio): move account creation into feature/account package AddAccountScreen gets its own AddAccountUiState colocated with the screen, replacing the shared UiState.addAccountError. * refactor(folio): move ledger screens into feature/ledger package LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger with AddTransactionUiState (txnError, txnFormType) colocated. The former catch-all UiState.kt is removed now that each screen owns its state alongside its UI. * refactor(folio): split Theme.kt; move theme and icons to subpackages Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and Type.kt (typography) under app.folio.ui.theme. Icons moves to app.folio.ui.icon. Update every consumer's imports to match. * refactor(folio): split ui components into per-file under ui/component Former Widgets.kt and Components.kt become 10 focused files: AppButton, Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/ BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style of one composable per file in a designsystem/component package. * chore(folio): consolidate uatu testing files under uatu/ folder Move spec.ts, package.json, package-lock.json into examples/folio/uatu so all uatu-specific testing artifacts live in one place. runs/ and node_modules/ follow the same convention (both remain gitignored). Update justfile, README, and the two Go consumers (bundle-check tool + verifier/trace tests) that referenced the old path. * refactor(trace): drop folio path in writer test Round-trip only needs a non-empty string; neutralize to keep the library free of folio references. * refactor(sidecar): neutralize ResolveActivity test fixtures Swap app.folio for com.example.app in the fixture strings so the sidecar tests don't reference the example app by name. * refactor(bundle-check): take spec path as argument Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a positional spec path instead so the tool works for any example and leaves no folio reference in the library surface. * test(verifier): add neutral integration spec and hierarchy fixtures Adds testdata/integration_spec.ts with two routes ("list", "form"), an InputText on text_field, a Tap on primary/secondary_action, a safety property (itemCountNonNegative), and a liveness property (submitEventually). Adds hierarchies_test.go with matching XML fixtures. Constants intentionally go in a _test.go at package root rather than testdata/hierarchies.go because go skips .go files under testdata/. * refactor(verifier): replace folio integration tests with neutral ones Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test* entry points to TestIntegrationSpec*. Uses the synthetic spec and hierarchies added in the previous commit so the library's test suite no longer references examples/folio at all. Folio-specific coverage remains covered by examples/folio/justfile's 'just test' which exercises the real spec on device/emulator. * chore: remove cmd/uatu-sample-harness Not referenced by Makefile, docs, CI, or any script. Duplicates the adb reverse helpers already in cmd/uatu/test_run.go, and its name implies ownership by the sample app which violates the library/ example decoupling. If a bare-protocol debugging tool is later needed it belongs inside cmd/uatu/. * docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax The directory-tree Layout section rots faster than the code and duplicates what ls shows for free. The expect/actual paragraph in Stack was the same kind of filler. The README also claimed 'just AVD=Pixel_7 test' but the justfile reads AVD as an env var via env_var_or_default, so the correct invocation is 'AVD=Pixel_7 just test'. |
||
|
|
c1de4bf57a |
fix(sample-app): address PR #18 review findings (#19)
- WebLedgerStore.accountExistsByName case-insensitive (matches SQL COLLATE NOCASE) - drop SampleApplication.maybeInjectDebugError; sample spec now passes clean - drop noLogcatErrors from sample spec (default matches system-wide E logs) - openRandomAccount picks uniformly from findAll instead of first match - clear txnError on any add-transaction interaction, not only valid amount input - escapeForAdbInputText quotes shell metacharacters (quotes, backslash, etc.) - extract buildClearKeyevents helper with empty/normal/cap tests - broaden spec_integration_test.go to cover home, add-account, ledger, add-transaction action generators - document FocusTracker single-focus invariant |
||
|
|
8381a98aaf |
feat(sample-app): ledger spec parity (#18)
* feat(sample-app): add FocusTracker * feat(sample-app): support description on TextInput * feat(sample-app): support description on AppButton and Segmented * feat(sample-app): stable ids on login screen * feat(sample-app): stable ids on add-account screen * feat(sample-app): stable ids on add-transaction screen * feat(sample-app): stable ids on home and ledger screens * feat(sample-app): hoist error + form state into UiState * feat(sample-app): register auth_status + accounts snapshots * feat(sample-app): register ledger_rows + ledger_balance snapshots * feat(sample-app): register error + focused_input snapshots * refactor(sample-app): scaffold spec.ts extractors + safety * feat(sample-app): spec accounting invariants * feat(sample-app): spec state-machine monotonicity * feat(sample-app): spec liveness properties * feat(sample-app): spec auth + account action generators * feat(sample-app): spec transaction action generators * feat(sample-app): spec weighted workflow * feat(sample-app): publish login form input snapshots * feat(sample-app): publish account + txn input snapshots * feat(sample-app): register form input snapshots * fix(sample-app): sequence login + idempotent text inputs in spec * fix(sample-app): clear UiState on page dispose * fix(sample-app): tighten adversarial login, allow retype on error * test(verifier): update sample-app spec integration for new selectors * feat(sidecar): clear focused field before InputText types * refactor(sample-app): simplify loginHelper to focus-driven sequencing * refactor(sample-app): drop input-value UiState mirrors * refactor(sample-app): drop input-value snapshot registrations * test(verifier): adjust sample-app integration for replace-on-type |
||
|
|
7493945251 |
feat: LTL operators, sampling, and default generators (#17)
* feat(ltl): add Now/Next/Eventually/Implies/Or/And/Not formulas Replace the fold-with-latch evaluator with a residual-formula reducer. Each Observe() instantiates a fresh obligation from the root (stripping an outer Always), reduces each pending obligation against current state, latches Violated on first failure, and surfaces Pending verdicts for deferred obligations. Existing Always/Pure/Thunk tests continue to pass. * feat(ltl): support relative duration for eventually().within() * feat(proto): add Swipe, PressKey, RecentLogs RPCs * feat(verifier,runner): formula handles, new action kinds, rich state - verifier: add formula-spec registry; bindNow/bindNext/bindEventually with chainable .implies/.or/.and/.not and .within(n,unit) on eventually; bindFrom for uniform sampling. bindAlways keeps accepting plain predicates. - verifier: store lastTree, lastAction, step time, logs, exceptions on the Verifier; SnapshotInput replaces the (snapshots, tree) pair. stateObject now produces state.lastAction/time/logs/exceptions matching the TS State type. - verifier: make taps/swipes/waitOnce/pressKey built-in generators actually fire; taps picks a clickable, enabled element from the last hierarchy. - agent: add exceptions field to Message wire format. - driver: add Swipe/PressKey/RecentLogs to Driver interface; wire maestro client and mock driver. LogEntry exposed for runner consumption. - runner: apply Swipe/PressKey/Wait actions; collect logcat and exceptions; pass lastAction and step time into PushSnapshot. * feat(spec-api): LTL operators, new actions, richer State - ltl.ts exports now/next/eventually; always overload accepts a Formula - types.ts: Formula gains implies/or/and/not; EventuallyFormula adds .within; State gains lastAction/time/logs/exceptions; Swipe/PressKey/Wait action types - actions.ts: Swipe/PressKey/Wait/from constructors; waitOnce + pressKey default generators - tests exercise the chaining, sampling, and new actions through a recorded fake runtime * feat(sidecar): add swipe, pressKey, recentLogs RPC handlers * feat(sdk-android): capture uncaught exceptions Install a default uncaught handler on Uatu.start, chained with any existing handler so Android's crash reporter still runs. Expose Uatu.reportError for callers to forward caught throwables. A bounded circular buffer (default 50) drains into each STATE message's new exceptions field. Protocol.kt serializes/deserializes the field, matching the Go wire format added to internal/agent/protocol.go. * feat(spec-api): add @uatu/spec/defaults/properties bundle * feat(sample-app): exercise new LTL operators + defaults spec.ts now imports eventually/next/now/from from @uatu/spec and noUncaughtExceptions from @uatu/spec/defaults/properties. It declares three properties that exercise the new surface: - accountCountNonNegative: plain always() safety - addAccountAdvances: always(now(x).implies(next(y))) - eventuallyLoggedIn: eventually(p).within(30, "seconds") - noUncaughtExceptions: imported default The weighted actions root uses from() for random phone/name sampling and entries for taps/swipes/waitOnce/pressKey built-ins. SampleApplication gains a debug hook gated on the system property uatu.inject_error so the e2e run can synthesize an Uatu.reportError and verify noUncaughtExceptions violates. cmd/uatu/test_run.go adds a subpath alias so specs importing "@uatu/spec/defaults/properties" resolve against the in-tree source when running from the uatu checkout. The spec-integration tests swap the old click-counter fixtures for the new login hierarchy. * feat(trace): record swipe/key/wait details + exceptions trace.Step gains an Exceptions array so the trace captures the class/message/stackTrace for each SDK-reported throwable in a step. trace.Action gains FromX/FromY/ToX/ToY/Key/DurationMillis so the full payload of Swipe/PressKey/Wait actions is visible in trace.jsonl. sample-app's debug error hook now gates on ApplicationInfo.DEBUGGABLE instead of a system property (adb setprop fails on non-rooted emulators). |
||
|
|
5b5594ee05 |
Port sample-app to KMP with Android, iOS, Web, and SQLite (#16)
* chore(sample-app): hoist gradle wrapper to sample-app root * chore(sample-app): add KMP root gradle config * chore(sample-app): add composeApp KMP module build config * chore(sample-app): add Android manifest for composeApp * feat(sample-app): add shared domain models and auth constants * feat(sample-app): add shared number and date formatting * feat(sample-app): add cross-platform storage, clock, and id * feat(sample-app): add shared repository with file-backed state * feat(sample-app): add shared in-memory navigator * feat(sample-app): add Compose theme and design tokens * feat(sample-app): add shared UI components (icons, screen, widgets) * feat(sample-app): add Login and Home pages * feat(sample-app): add AddAccount, Ledger, AddTransaction pages * feat(sample-app): add App root composable with routing * feat(sample-app): add Android Application and Activity hosting Compose UI * chore(sample-app): remove legacy android module (replaced by composeApp) * chore(sample-app): bump to latest stable Kotlin/AGP/Compose deps * fix(sample-app): make iOS compile (drop @Volatile, set bundleId) * feat(sample-app): add xcodegen spec, SwiftUI host, and iOS Info.plist * chore(sample-app): ignore build artifacts and generated xcodeproj * chore(sample-app): update justfile for composeApp layout, add ios target * docs(sample-app): rewrite README for KMP + iOS flow * refactor(sample-app): drop in-app status bar and formatClock * feat(sample-app): add SQLDelight schema and per-platform drivers * refactor(sample-app): back Repository with SQLite, drop file serializer * feat(sample-app): wire native back on Android and iOS edge swipe * feat(sample-app): semantic roles, labels, and a11y descriptions * fix(sample-app): link libsqlite3 for iOS target SQLDelight's native driver needs libsqlite3.tbd on iOS; without it the linker fails with undefined _sqlite3_bind_blob and friends. * refactor(sample-app): abstract storage behind LedgerStore interface Platform-specific createLedgerStore() returns a SqlLedgerStore backed by SQLDelight on Android + iOS. Opens the door for a pure in-memory web implementation that does not require a SQLite driver. * feat(sample-app): add wasmJs target with in-memory LedgerStore Wires a Compose Multiplatform browser canvas entry point. The web implementation of LedgerStore is an in-memory model with localStorage persistence, so it does not need a SQLite driver. Back navigation maps the browser back button to the same BackHandler contract Android and iOS use. * chore(sample-app): settings + gitignore for wasmJs dev run Registers the Node.js distributions repository and switches repositoriesMode to PREFER_PROJECT so the Kotlin wasmJs plugin can download its toolchain. Adds kotlin-js-store (lockfile) and ignores runs/, web screenshots, playwright-mcp scratch output. * fix(sample-app): singularize transaction count on Home Shows "1 transaction" not "1 transactions" for accounts with a single transaction; falls back to "$count transactions" otherwise. |
||
|
|
69002b07b1 |
fix(runner,sample-app): surface silent errors and demo a failing property (#15)
* fix(runner): surface non-deadline WaitForIdle errors Previously the WaitForIdle return value was discarded entirely, hiding real driver failures (gRPC transport errors, sidecar crashes) behind the expected deadline-exceeded case. Log non-deadline errors so they are visible without changing control flow. * chore(sample-app): drop unused uptime_millis extractor Registered in SampleApplication but never consumed by spec.ts. * fix(sample-app): drop trivial appIsRunning property app_state was hardcoded to 'running' so the property was a tautology that could never fail. Removing both the extractor and the property is the simplest fix; demo-grade properties that can fail land next. * feat(sample-app): add Reset button that zeroes clickCount Pairs with the next commit's tap-reset action so the fuzzer can violate clickCountNeverDecreases and demonstrate uatu actually finding a property violation. * feat(sample-app): add tap-reset action to exercise Reset button Weighted at 10/122, fuzzer reaches it within a short run. Pairs with the Reset button to demonstrate uatu detecting the clickCountNeverDecreases violation. * fix(runner): filter WaitForIdle errors via context state, not errors.Is errors.Is(err, context.DeadlineExceeded) misses gRPC's wrapped status.DeadlineExceeded, so every step under the maestro driver logged a spurious warning. Check idleCtx.Err() instead — captures both deadline-fired and parent-canceled cases regardless of how the driver wraps them. * chore(sample-app): tune action weights so demo violates in ~30s Prior weights left tap-reset rare enough that short demo runs missed the violation by chance. Bumped to 30/107, with typeUsername reduced since username noise doesn't help exercise clickCount. * refactor(runner): route warnings through slog Adds Options.Logger (defaults to slog.Default()) and converts the three warning sites that were using fmt.Printf. Progress line stays on Printf since it's user-facing UI, not a log. Makes the warnings testable via a capturing handler. * test(runner): assert WaitForIdle driver errors are logged Captures slog output via TextHandler into a buffer and asserts the warning message + injected error text appear when the mock driver returns a non-context error from WaitForIdle. Guards against a regression of the silent-error swallow. |
||
|
|
00f66ba48d |
fix: pre-v0.1 review feedback (#14)
* fix(sdk-android): add @JvmOverloads to Uatu.start Java callers can now invoke start(application) without supplying a Configuration, matching the Kotlin default-arg ergonomics. * feat(agent): add protocol_version to HELLO handshake ProtocolVersion=1 lives on Message and is set by Hello(). Server.Accept rejects mismatches with a clear error. SDK upgrades that don't change the wire format keep the same protocol_version; bump on breaking changes. * test(agent): assert protocol_version in Hello round-trip * feat(sdk-android): send protocol_version=1 in HELLO Mirrors agent.ProtocolVersion on the Go side. Bump in lockstep with the Go constant when the wire format breaks. * chore(sample-app): pull @uatu/spec from npm next tag Replaces the file: dep. Copy-paste users can now npm install against the registry. The release workflow publishes pre-release tags to npm dist-tag 'next', so the sample tracks the latest rc without manual version bumps. Lockfile currently resolves to 0.0.1-rc3. |
||
|
|
d0578dbaaa |
fix(runner): warn on malformed screen snapshot (#13)
* fix(runner): warn on malformed screen snapshot screenFromSnapshot swallowed json.Unmarshal errors, so a non-string screen value silently became "" in the step log and trace while the verifier still saw the raw JSON. Return the error and warn at the call site, matching the hierarchy warning pattern. * docs: clarify --avd is optional for uatu test The CLI accepts --avd as an empty-string default (cmd/uatu/main.go:49) and only requires it when no device is connected and multiple AVDs exist (cmd/uatu/android_env.go:63). Docs and examples that showed it as required or always-passed were misleading. |
||
|
|
16e55086d8 |
fix(runner): surface focus-tap errors in InputText (#12)
* fix(runner): surface focus-tap errors in InputText action A failed Tap/TapSelector before InputText was swallowed, so text typed into the wrong field (or no field) still reported success. Return the error so the step fails explicitly. * feat(sample-app): add username EditText and snapshot Gives the spec a real EditText target (content-desc: username_field) so the InputText action path can be exercised end-to-end. The typed value is mirrored into MainActivity.username and surfaced as the "username" snapshot for spec assertions. * feat(sample-app): exercise InputText action against username field Adds typeUsername action and usernameNeverShrinks property to the sample spec, and extends the integration test to assert the bundled spec emits an InputText(desc:username_field, "alice") action and that the property correctly violates when a snapshot reports a shorter string. |
||
|
|
6b1c17328c |
fix(sample-app): make it standalone, no repo_root assumptions (#8)
* fix(cli): fall back to node_modules for @uatu/spec resolution
Drop the hard failure when the uatu source tree is not reachable from
the spec file. Users integrating uatu in their own app have @uatu/spec
installed via npm; esbuild now resolves it from node_modules.
* build(gradle): drop :sample-app include from root settings
The sample now has its own Gradle project in examples/sample-app/android.
* build(sample-app): vendor gradle wrapper
Users running the sample build the APK via ./gradlew from inside the
sample-app's own android/ directory, no repo-root wrapper required.
* build(sample-app): make gradle project standalone
Drop the project(':sdk-android') dependency in favor of the Maven
Central coordinate io.github.priyanshujain:sdk-android. The sample now
owns its settings.gradle.kts and gradle.properties, so it builds
without any pieces of the uatu source tree.
* chore(sample-app): declare @uatu/spec npm dependency
Mirrors what a downstream user would put in their own package.json.
Uses file: for pre-release development; becomes a normal semver pin
once @uatu/spec ships to npm.
* docs(sample-app): rewrite justfile and add README
Justfile drops repo_root; all recipes run against the local gradle
wrapper and uatu from PATH. README is scoped to what a user needs to
run the sample against their own device.
* docs(manual): update sample install steps to standalone layout
./gradlew :sample-app:installDebug no longer exists; the sample owns
its own wrapper under android/.
* feat(sdk): log when Uatu.start succeeds
Silent SDK start makes the "SDK didn't connect" failure mode
impossible to debug. One INFO line at start time is enough.
* fix(sidecar): launch via am start -W instead of monkey
monkey -p <pkg> -c LAUNCHER 1 is unreliable on API 36+: it reports no
error but silently fails to start the activity, so the SDK never runs
and the CLI times out on the SDK-accept handshake.
Resolve the launcher activity via `cmd package resolve-activity
--brief` and launch it with `am start -W -n`. -W makes the call block
until the activity is up, which also makes the subsequent SDK
accept timing deterministic.
* feat(cli): auto-resolve Android device; boot AVD if none connected
--avd becomes optional. Resolution order:
- use any already-connected adb device;
- else if --avd names an existing AVD, boot it and wait for boot;
- else if --avd is missing or names no AVD, error with a clear message.
Falls back to $ANDROID_HOME/emulator/emulator when the binary is not on
PATH, so a standard Android SDK install works without extra shell setup.
* docs(sample-app): AVD is optional; document both paths
just test runs against any connected device. If none, pass AVD=<name>
to have uatu boot the emulator for you.
* feat(cli): auto-discover Android SDK; auto-pick the lone AVD
adb and emulator are looked up via PATH, then $ANDROID_HOME,
$ANDROID_SDK_ROOT, ~/Library/Android/sdk, ~/Android/Sdk, and the
Homebrew cask path. The discovered platform-tools directory is
prepended to the sidecar's PATH so its adb subprocess calls work too.
When --avd isn't passed and no device is connected, the CLI picks the
sole local AVD and boots it. Multiple AVDs → error listing them.
* build(make): add `make install` that go-installs uatu onto PATH
Puts `uatu` into $GOBIN (or $GOPATH/bin) so the sample and any local
dev flow can call it without PATH= prefixes.
* docs(sample-app): zero-config just test; dotenv-load for persistence
Drop the expectation that users prefix commands with PATH=, ANDROID_HOME=,
or AVD=. `just test` now works as-is; optional knobs can be pinned in a
.env file alongside the justfile.
* fix(sample-app): auto-detect ANDROID_HOME for Gradle tasks
The Go CLI finds the SDK itself, but AGP still needs ANDROID_HOME to
resolve `sdk.dir`. The justfile now resolves it from env or canonical
install paths before invoking ./gradlew, so `just install` works out
of the box on a standard Android SDK setup.
* gitignore runs directory for sample app
|
||
|
|
a74d9fbeae |
build(examples): add justfile for sample-app
Targets: install/uninstall the APK, build the uatu CLI, run 'uatu test' against a named AVD, run the Go verify tests, and clean. Keeps the example self-contained so users can 'just test AVD=pixel_7' after cloning. |
||
|
|
ec3764ad85 |
chore(examples): remove merchant-ledger spec
No longer referenced anywhere — sample-app replaces it as the in-tree example. |
||
|
|
0c775c199d |
feat(examples): add sample-app spec
Introduce examples/sample-app/spec.ts — a minimal property-based spec that taps the sample app's "Click me" button and asserts click_count is monotonic. Bundle-check and the trace writer test now reference the new path. |
||
|
|
c12a3e242f |
refactor(android): move sample app under examples/sample-app/
Rename the gradle sample project from :sdk-android-sample to :sample-app so the isolated example for users lives beside its spec in examples/. |
||
|
|
6a07284e24 | fix(spec): screen-aware sign convention + leaveLedger generator to cycle counterparties | ||
|
|
b19d83fd54 | fix(spec): state-aware enterMobile + globalise properties for evaluator pickup | ||
|
|
6ce4eabe86 | feat(spec): drive merchant onboarding (English, phone, OTP, multi-device, home) | ||
|
|
4dbad9073e |
feat(spec): merchant-ledger property + login/navigation generators
Property ledgerBalanceMatchesTxns asserts displayed balance equals sum(Given) - sum(Received) on either ledger screen. Catches the class of bug where the server-fed CustomerModel.balance diverges from the local-DB-fed CoreDatabaseDao sums (stale cache, partial sync, deleted-txn handling glitch, etc.). Generators are gated by screen state and weighted to push the run through login -> home -> ledger quickly: enterPhone (100), enterOtp (100), openCustomerOrSupplier (80), taps (10), swipes (2). Phone/OTP read from process.env via esbuild defines so credentials never land in source. cmd/internal-tools/bundle-check is a quick sanity tool to confirm the spec bundles before running uatu test. |