collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.
* feat(sidecar): reach USB devices via the adb server by serial
* feat(test): add --device flag to target a specific Android device by serial
* feat(folio): select Android device via ANDROID_DEVICE in justfile
* feat(conformance): add android backend to the gate suite
* feat(android): keep device awake and unlocked so the app stays foreground
* feat(conformance): prep physical android device (autofill/verifier/stayon)
* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run
* fix(verifier): require positive bounds for swipe candidates
A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.
* fix(runner): harden app-scope guard against launcher and overlays
The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.
* feat(android): harden physical-device runs in device prep
Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.
* feat(driver): clear-state via APK reinstall when pm clear is blocked
When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.
* feat(cli): add --android-app-path for clear-state reinstall
Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.
* chore(folio): pass --android-app-path in just test
* fix(runner): clamp swipe/scroll origin out of edge gesture zones
A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.
* perf(sidecar): faster Android text input and drop redundant settle poll
inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.
* fix(verifier): exclude soft-keyboard region from action candidates
The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.
* perf(runner): replace focus-tap settle with a brief wait
The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.
* chore(conformance): platform-aware G5 p95 budget for android
The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.
* fix(sidecar): retry maestro android driver startup
The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.
* chore(conformance): widen android G5 budget to 5500ms
Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.
* web replay fix
* feat(android): force 3-button nav during runs to prevent app drift
On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.
* fix(android): target the selected device in adb reads; don't strand nav mode
Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
foreground/scope guard works when several devices are attached (the --device
path). Previously they ran bare `adb shell`, which errors with multiple
devices, silently disabling app-scope enforcement. The sidecar client passes
its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
the current mode is unknown or already 3-button it leaves nav untouched,
instead of switching and then stranding the device in 3-button. Logic split
into the pure navModeToRestore, now unit tested.
* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds
Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.
* test(verifier): cover keyboardRegionTop, including the decor-view guard
The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.
* style(cli): gofmt testOptions field alignment
* fix(sidecar): keep a leading dash off the fast input path
A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).
* refactor(verifier): scope action candidates by window ownership
Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.
This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.
* fix(runner): re-check foreground at apply time, skip stale actions
ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.
* fix(android): type long ASCII via fast guarded path to stop keystroke escape
A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.
* chore: ignore gate artifacts and local scratch files
* refactor(runner): narrow gesture clamp to the top shade strip
3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.
* chore(format): add .editorconfig enforcing 80-column limit
* chore(format): add prettier config with 80-char printWidth
* chore(deps): add prettier devDependency to replay-ui
* chore(deps): add prettier devDependency to folio-web
* chore(deps): add prettier devDependency to spec package
* chore(format): add swift-format config with 80-char lineLength
* feat(format): add make fmt targets for per-language 80-col formatting
* fix(runner): translate gesture to safe area so near-top scrolls keep direction
Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.
* fix(runner): apply-time guard consults focused window, not just resumed activity
ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.
* test(runner): cover apply-time foreground skip and appIsForeground table
Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.
* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker
- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.
* fix(conformance): pin self-test p95 budget and score install failures as run failures
self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.
* fix(android): require --device when several devices are connected
With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.
* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands
The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.
* perf(verifier): memoize scopedElements per tree
scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.
* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear
Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.
* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms
The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.
* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty
bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.
* perf(sidecar): reuse a single Jackson ObjectMapper
structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.
* refactor(android): remove unused AdbReverse/AdbReverseRemove
No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.
* style(runner): trim non-load-bearing comments from this PR's runner code and tests
* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
* docs(manual): add introduction page
* docs(manual): rewrite getting started as guided first run
* docs(manual): rewrite writing specs as a folio tutorial
* docs(manual): document missing spec API in reference
* docs(manual): plain-language rewrite of runs page
* docs: real introductions on index pages and README
* fix(docs): sibling links from directory-style pages need ../
* fix(docs): correct sampling and restart-cost claims to match implementation
* docs: nav lists Introduction and Case study; roadmap points to milestone
* docs(manual): make getting started target the reader's own app, not Folio
* docs(manual): add Folio case study page
* docs: point manual navigation at the case study
* docs(readme): lead with the case study, fix roadmap link
* docs: roadmap links to milestone, sync clear-data default and cross-links
* feat(companion): add appState, eraseText, pressKey runner handlers
The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.
* feat(ios): resolve physical devices from devicectl
ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.
* feat(ioscompanion): runner-only device driver mode
NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.
* test(ioscompanion): cover device wiring, routing, and shell-out argv
Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.
* feat(testrun): route physical-device iOS runs to the device driver
Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.
* feat(doctor): device prereqs replace java/sidecar for ios-device
iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.
* feat(conformance): device backend uses iphoneos app and tunnel orphan checks
The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).
* feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.
* docs(cli): document ios-device doctor checks and the device flags
The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.
* fix(ioscompanion): resolve signing key path to absolute
xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.
* fix(ioscompanion): re-enable signing for the device runner build
companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.
* fix(ioscompanion): key the device build cache on signing identity
The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.
* docs(getting-started): document physical iOS device setup
Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.
* feat(ios): native usbmux client and in-process tunnel forwarder
Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.
* refactor(ios): drive device tunnel via io.Closer seam
Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.
* refactor(ios): remove iproxy spawn from device runner
* test(ios): cover tunnel close via io.Closer not child process
* feat(doctor): check usbmuxd socket instead of iproxy on PATH
* chore(conformance): drop iproxy orphan check; tunnel is in-process
* docs(ios): device tunnel uses native usbmux, nothing to install
* chore: gitignore the signing keys directory
* feat(folio): add Android launcher icon (black bg, white dot)
* feat(folio): add iOS app icon (black bg, white dot)
* feat(folio): add web favicon (black bg, white dot)
* docs(ioscompanion): fix stale const comments
* refactor(ioscompanion): inline single-use devicectl argv builders
* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests
* refactor(ioscompanion): inline firstNonEmpty
* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam
* test(doctor): trim redundant signing-check test
* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env
* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
* refactor(sidecar): drop IosDriverBackend
* refactor(sidecar): route ios platform off the iOS backend
* chore(sidecar): remove maestro ios dependencies
* feat(testrun): reject physical iOS with a clear message
* refactor(sidecar): drop iOS hierarchy helpers and their test
* build(sidecar): strip iOS runner bundles and classes from the fat jar
* test(testrun): cover physical-iOS rejection
* perf(ios): use prebuilt XCTest runner to cut startup
* chore(ioscompanion): add companion asset prepare script
* feat(ioscompanion): embed and extract simulator companion bundle
* test(ioscompanion): cover companion stub and embedded extraction
* docs: add third party notices for vendored companion
* chore: ignore vendored companion bundle artifact
* build(proto): pin simulator companion proto v1.1.8
* build(proto): add dedicated buf module and gen template for pinned proto
* build(proto): exclude pinned companion proto from root buf workspace
* feat(ioscompanion): commit generated companion gRPC stubs
* feat(ioscompanion): map flat companion describe dump to TreeNode JSON
* test(ioscompanion): add hierarchy-map golden and unit tests
* feat(ioscompanion): port screen-settle stability polling to Go
* test(ioscompanion): cover settle transitional, hash, streak, and cap rules
* feat(ioscompanion): add USB HID keymap module
* test(ioscompanion): cover keymap branches and paste-chord constants
* build: embed companion assets via withcompanion tag
* feat(ioscompanion): add transport companion interface
* feat(ioscompanion): add HID event wrapper and builders
* feat(ioscompanion): wire gRPC companion client and Dial
* test(ioscompanion): cover HID builders and unit conversions
* test(ioscompanion): cover Dial, process-state mapping, and install archive
* test(ioscompanion): add gated simulator integration smoke test
* feat(ioscompanion): text input and gesture HID composition with pasteboard fallback
* test(ioscompanion): cover input composers, paste dialog loop, and pure helpers
* feat(ioscompanion): add Describe to companion transport
* feat(ioscompanion): implement DeviceDriver with companion supervision
* test(ioscompanion): unit tests with fake companion transport
* test(ioscompanion): gated companion smoke test
* feat(ios): add ResolveTarget for simulator vs physical-device routing
* feat(testrun): route iOS simulators through the native companion driver
* refactor(testrun): defer the java preflight check to the physical-device path
* feat(cli): add --ios-app-path flag
* feat(doctor): split iOS checks into simulator and physical-device paths
* test(folio): add gate-analyzer fixtures for G1-G5
* feat(folio): add iOS conformance gate script
* chore(folio): wire gates recipe, app path, and ignore gate output
* style: gofmt struct alignment drift
* fix(doctor): probe simctl via xcrun instead of PATH lookup
* fix(ioscompanion): spawn companion under driver-lifetime context
* test(ioscompanion): prove companion child outlives startup context
* fix(ioscompanion): chunk install payload under companion message cap
* test(ioscompanion): cover install payload chunking
* fix(ioscompanion): reinstall via simctl and sanitize companion env
* fix(ioscompanion): wait out unresolved accessibility values after launch
* perf(ioscompanion): paste long text for atomic landing
* test(ioscompanion): cover paste threshold, retry flow, and sentinel detection
* fix(ioscompanion): treat unresolved bridge values as transitional, never as content
* fix(ioscompanion): accept masked secure-field values as paste landing
* test(ioscompanion): cover sentinel mapping and masked-field landing
* fix(ioscompanion): atomic erase and single-send paste to prevent doubling
* test(ioscompanion): cover atomic erase, single chord, unverifiable field
* fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout
* test(ioscompanion): cover bridge-blackout paste verification
* fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle
* refactor(ioscompanion): name the empty-editable-field sentinel for what it is
* perf(ioscompanion): tighten settle streak for the fast companion transport
* feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt
* refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt
* test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests
* fix(ioscompanion): retry describe past transient collapsed accessibility dumps
* test(ioscompanion): cover collapsed-dump detection
* perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses
* perf(ioscompanion): tighten settle now that collapses are handled separately
* fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text
* test(ioscompanion): cover replace-on-input and TextReplacer capability
* refactor(ioscompanion): neutralize HID events behind the transport seam
* feat(companion): add simulator runner project skeleton
* feat(companion): serve accessibility snapshots over the wire protocol
* feat(companion): synthesize timestamped touch gestures
* feat(companion): type text with replace semantics
* feat(companion): serve the wire protocol from a parked runner
* feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam
* feat(ioscompanion): route text input through a text-editing companion when available
* fix(companion): bind listener by port and source screen size from snapshot
* feat(ioscompanion): add runner companion JSON transport
* test(ioscompanion): cover runner transport protocol mapping
* fix(companion): synthesize gestures synchronously to avoid the async completion crash
* fix(companion): type on the main thread and recover from focus assertions
* fix(companion): keep serving after an automation failure
* refactor(companion): tidy snapshot serialization
* fix(companion): honor sequential tap gaps and survive synthesis exceptions
* feat(ioscompanion): expose native typing with an explicit replace flag
* chore(companion): add runner asset prepare script
* feat(ioscompanion): embed and extract the runner test bundle
* test(ioscompanion): cover runner asset extraction
* build(ioscompanion): commit runner asset archive
* feat(ioscompanion): pair the legacy companion with the in-simulator runner
* test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding
* fix(ioscompanion): reconnect after interrupted runner calls instead of restarting
* fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts
* feat(companion): launch and terminate apps through the automation session
* build(ioscompanion): refresh runner asset with session lifecycle
* fix(ioscompanion): classify connection deadline expiry as caller budget
* fix(companion): capture snapshots on the main thread inside the catch bridge
* build(ioscompanion): refresh runner asset with main-thread snapshots
* perf(ioscompanion): count read spans toward settle and capture snapshots concurrently
* feat(ioscompanion): make the hybrid simulator companion the default
* test(folio): cover runner-session orphans in the gate harness
* test(ioscompanion): pin the child-lifetime test to the legacy path
* fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears
* fix(ioscompanion): pause the clear chord so selection applies before the delete
* fix(companion): prune the keyboard subtree from snapshots
* build(ioscompanion): refresh runner asset without keyboard elements
* fix(ioscompanion): capture the screenshot transport before a recovery can reassign it
* fix(companion): pin the runner listener to loopback
* fix(companion): size the replace delete prefix to cover any focused field
* build(ioscompanion): refresh runner asset with loopback bind and replace fix
* fix(cli): cancel the run context on SIGINT so spawned children are reaped
* fix(testrun): point the device java preflight hint at the ios-device doctor
* fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob
* test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision
* chore: add test-companion target for the withcompanion-tagged suite
* chore(ioscompanion): stop tracking the runner archive build artifact
* build: produce the runner archive from source like the companion bundle
* refactor(conformance): move the gate harness out of examples/folio
* chore(folio): drop the gate harness wiring from the example app
* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
* feat(ltl): attribute violations to the obligation origin step
* feat(verifier): label evaluator observations with the runner step index
* feat(trace): carry the causing step in violation witnesses and summary
* feat(replay): move the violation marker to the causing step
* feat(replay-ui): render witness evidence in the violations panel
* feat(replay-ui): wire witnesses and step jump into violation panels
* fix(ltl): treat next obligations as vacuous at run end
* fix(runner): give the finalize trace record its own step index
* fix(hierarchy): bounds-containment fallback for scoped and path queries
Compose on iOS surfaces a testTag node as an empty leaf sibling of the
content it labels instead of as an ancestor, so descendant search under
the tagged node finds nothing and every path or scoped query returns
null. When structural search yields no match, fall back to nodes whose
bounds lie inside the scope node's bounds.
* feat(sidecar): derive iOS clickable and editable from element type
The XCTest hierarchy mapping dropped the element type, leaving no
clickable or editable flags on iOS, so the fuzzer's tap and typing
verbs never found a candidate inside the app. Map the raw
accessibility tree directly and derive clickable, editable,
scrollable, and class from the XCUIElementType raw value.
* feat(proto): add EraseText RPC for InputText replace semantics
* feat(driver): add EraseText to the device driver surface
* fix(runner): erase existing field text before InputText
InputText appended on native platforms, so repeated draws grew fields
without bound. The folio fuzz run wedged on the add-account screen:
each draw concatenated another name until the 40-character validation
error became permanent. Replace semantics also makes retried typing
idempotent. The web driver already replaced via select-all; native now
matches.
* feat(sidecar): EraseText backend support on android and ios
* fix(folio): saturation-gate account creation in the spec
The 2-3 step add-account loop outcompeted the 5-step transaction chain
at every weighted re-draw, so runs filled with account creation and
rarely exercised the balance properties. Stop offering add-account once
three accounts exist; the renormalized weights then favor the
transaction flow at every step of its chain.
* fix(folio): author spec weights to match testing intent
Revert the account saturation gate: it starved newAccountBalanceIsZero
once it tripped, and a magic account count is app-state tuning, not
intent. Instead weight the generators by what the properties need:
the transaction chain leads, account creation stays exercised, and
doubleTaps gets explicit weight everywhere because double-submission
idempotency is what the spec is testing for.
* fix(folio): lower doubleTaps weight to 5
* fix(sidecar): surface visible text on iOS static elements
Static text and button strings live in the accessibility label on
iOS, so the text attribute came through empty and every balance
extractor parsed to zero, silently disarming both folio properties.
Non-editable elements now fall back title, value, then label;
editable fields keep value-only so an empty field's caption does not
read as content.
* feat(driver): native DoubleTap RPC for a tight inter-tap gap
Composing two Tap round trips from the Go client spread the taps by
hundreds of milliseconds on iOS, wide enough for the app to navigate
between them, so double-submission races could never reproduce. The
sidecar now lands both taps back-to-back next to the device transport.
* feat(sidecar): pipeline iOS double-tap requests
Queue the second tap at the XCTest runner while the first executes.
The runner serializes handlers, so this is the tightest gap the
transport allows (~350ms per tap round trip); recorded here with
measurements for the iOS double-tap limitation.
* refactor: rename inspect to replay across the codebase
Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/,
the CLI subcommand from `sanderling inspect` to `sanderling replay`, and
updates all references in docs, Makefile, README, and Go comments.
* feat(replay-ui): show spec filename with full path on hover
RunList and RunDetail now render the basename of spec_path (e.g.
login.spec.ts) with the full path available as a title tooltip.
* feat(ltl): bound fields on AlwaysFormula and named thunks
Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.
* feat(ltl): negation normal form pass
nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.
* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse
Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.
* test(ltl): property-based NNF laws
Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.
* test(ltl): Finalize, bounded eventually, latch, collapse
Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.
* feat(inspect): within clause on always residual node
A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.
* feat(ltl): witness violations and (bool,error) predicate thunks
* test(ltl): migrate thunk call sites to (bool,error)
* feat(ltl): flag thrown-predicate witnesses with IsError
* refactor(verifier): replace predicate err side-channel with violation witness
* test(verifier): witness API for thrown predicates
* feat(trace): witnesses map and skipped-verification marker on Step
* feat(runner): thread violation witnesses, finalize, skip marker into trace
* test(ltl): lock violation witness reason, IsError, and step
* test(verifier): finalize surfaces unmet eventually with witness
* fix(ltl): eliminate implies and bounded-always false-negatives
Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.
* test(ltl): lock implies and bounded-always false-negative regressions
* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick
* feat(testrun): inject seed into web bundle via SANDERLING_SEED define
* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring
* test(spec): add Go math/rand/v2 PCG oracle and golden fixture
* feat(spec): bit-exact PCG port of Go math/rand/v2
* test(spec): assert pcg.ts matches the PCG golden fixture
* feat(spec): shared input corpus and press-key pools
* feat(spec): action-tree types and Host interface
* feat(spec): verb support matrix and warn-once helper
* feat(spec): deterministic shared action picker
* test(spec): verb matrix and warn-once semantics
* test(spec): picker draw-order and determinism
* refactor(spec): actions.ts returns pure GeneratorNode data trees
* refactor(spec): wire from() sampling through the picker rng
* feat(spec): shared runtime-entry installs next-action over pick.ts
* feat(spec): export LongPress/Scroll/longPresses/scrolls factories
* test(spec): assert data-tree shapes for action factories
* test(spec): runtime-entry serializeAction wire-contract round-trip
* refactor(spec): bridge data-tree nodes to the legacy goja picker tags
* fix(spec): web runtime walks the spec's globalThis.actions data tree
* test(spec): tolerate legacy bridge fields on builtin nodes
* refactor(spec): installRuntime accepts a lazy root resolver
The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.
* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker
Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).
* test(spec): cover the WEB Host surface and seed precision
Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.
* refactor(spec): picker emits native selector + scroll endpoints, setup precedence
* feat(spec): goja runtime entry wires the shared picker over the Go host
* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin
* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker
* refactor(spec): drop the legacy goja bridge fields from action factories
* feat(spec): serialize selector-only string targets for the runner to re-resolve
* refactor(verifier): one DecodeAction reads the unified flat wire contract
* refactor(verifier): goja host + shared picker replace the duplicate Go picker
* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime
* test(verifier): author specs through the shared picker path
* test(runner): bundle authored specs with the goja runtime entry
* feat(verifier): collect unsupported verbs for the run report
* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource
* feat(testrun): surface unsupported verbs in run report
* test(verifier): cross-runtime goja/node parity gate on the shared picker
* test(verifier): unsupported verbs collected deduped in first-seen order
* test(runner): summary reports no unsupported verbs on a clean run
* test(spec): golden-fixture cross-runtime parity gate for the node picker
Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.
* test(verifier): assert goja picker against the same cross-runtime golden
Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.
* refactor(spec): rename pressKey generator export to pressKeys
* refactor(spec): update barrel re-exports for pressKeys
* test(spec): update pressKeys generator export name
* docs(spec): rename pressKey generator to pressKeys
* refactor(spec): extract samplerRng into shared sampler-rng module
* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)
* test(spec): cover fluent value generators determinism and chaining
* refactor(bundler): inject globalThis trailer from spec named exports
* refactor(bundler): reuse registration trailer in web bundler
* test(bundler): cover named-export globalThis registration
* feat(spec): add named() to Extracted handle type
* feat(web-runtime): named() and cross-extractor read guard
* feat(verifier): named() and cross-extractor read guard in goja
* test(verifier): cross-extractor read guard and named()
* test(web-runtime): export runtime and extractors for tests
* test(web-runtime): named() and cross-extractor read guard
* refactor(folio): drop manual globalThis trailer (bundler injects it)
* refactor(folio): seed txn amounts via integers().between(1,500)
* refactor(folio-web): drop manual globalThis trailer (bundler injects it)
* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs
* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts
* refactor(folio-web): name extractors so violation witnesses are readable
* fix(web-runtime): propagate extractor getter throws and unpoison locked global
Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.
* test(spec): install fake runtime via defineProperty to survive locked global
* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors
* feat(runner): add MaxSteps bound to Options
* test(runner): MaxSteps stops after exactly N steps
* test(driverpb): drop proto getter round-trip tautology
* test(sidecar): drop stub-mode placeholder tautology tests
* test(mock): drop default-field-value assertion test
* test(ltl): drop Verdict.String tautology tests
* refactor(runner): extract RenderSummary for snapshot testing
* test(runner): golden snapshots for trace stream and violation summary
* feat(web-runtime): capture uncaught errors into state.exceptions
* test(integration): add throwing and counter web fixtures
* test(integration): add specs for the web fixtures
* test(integration): drive web fixtures through the real pipeline in headless Chrome
* chore(make): add test-browser target for the Chrome-driven suite
* ci: run the Chrome-driven browser suite in a separate job
* refactor(test): relocate browser suite to test/browser
* refactor(permissions): delete dead internal/permissions package
* refactor(test): rename package to browser_test
* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets
* chore(make): point test-browser at test/browser
* docs(decisions): record internal/permissions deletion
* refactor(doctor): use sidecarassets package
* refactor(testrun): use sidecarassets package
* fix(test): resolve testdata relative to browser_test.go
* refactor(verifier): remove dead __sanderlingIndex compat alias
* refactor(bundler): use encoding/json for JS string literals
* docs(action-space): use vendor-neutral native driver wording
* refactor(hierarchy): scrub backend tool name from comments
* refactor(driver): scrub backend tool name from comments
* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver
* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap
* refactor(chrome): implement DoubleTap as two taps with the gap
* refactor(mock): record DoubleTap and DoubleTapSelector actions
* refactor(runner): delegate double-tap to driver, drop gesture timing
* test(runner): assert double-tap delegates to driver DoubleTap
* docs(cmd): add package docs to CLI and developer tools
* docs(driver): add package docs to driver interface and chrome backend
* docs(driver): add package docs to mock and sidecar backends
* docs(platform): add package docs to android and ios device prep
* docs: add package docs to bundler and inspect
* docs(ltl): add package doc to temporal logic evaluator
* docs: add package docs to runner and testrun pipeline
* docs: add package docs to trace and verifier
* docs(sidecarassets): add package doc for embedded JAR loader
* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI
* test(chrome): gate real-Chrome driver tests behind the browser tag
* chore(make): run chrome driver tests in the browser job
* fix(web-runtime): guard global error listeners for non-browser hosts
The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.
* ci(browser): re-enable unprivileged user namespaces for headless Chrome
ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.
* ci(browser): pin stable Chrome for the driver tests
setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.
* feat(defaults): add scroll and rebalance action weights
Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.
* feat(defaults): trim scroll action weight wiring
* fix(build): point sidecar jar ignore and embed paths at sidecarassets
* test(defaults): drop stale longPresses re-export assertion
longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.
* fix(chrome): raise DevTools websocket read timeout to 60s
Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.