* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
* feat(ltl): attribute violations to the obligation origin step
* feat(verifier): label evaluator observations with the runner step index
* feat(trace): carry the causing step in violation witnesses and summary
* feat(replay): move the violation marker to the causing step
* feat(replay-ui): render witness evidence in the violations panel
* feat(replay-ui): wire witnesses and step jump into violation panels
* fix(ltl): treat next obligations as vacuous at run end
* fix(runner): give the finalize trace record its own step index
* fix(hierarchy): bounds-containment fallback for scoped and path queries
Compose on iOS surfaces a testTag node as an empty leaf sibling of the
content it labels instead of as an ancestor, so descendant search under
the tagged node finds nothing and every path or scoped query returns
null. When structural search yields no match, fall back to nodes whose
bounds lie inside the scope node's bounds.
* feat(sidecar): derive iOS clickable and editable from element type
The XCTest hierarchy mapping dropped the element type, leaving no
clickable or editable flags on iOS, so the fuzzer's tap and typing
verbs never found a candidate inside the app. Map the raw
accessibility tree directly and derive clickable, editable,
scrollable, and class from the XCUIElementType raw value.
* feat(proto): add EraseText RPC for InputText replace semantics
* feat(driver): add EraseText to the device driver surface
* fix(runner): erase existing field text before InputText
InputText appended on native platforms, so repeated draws grew fields
without bound. The folio fuzz run wedged on the add-account screen:
each draw concatenated another name until the 40-character validation
error became permanent. Replace semantics also makes retried typing
idempotent. The web driver already replaced via select-all; native now
matches.
* feat(sidecar): EraseText backend support on android and ios
* fix(folio): saturation-gate account creation in the spec
The 2-3 step add-account loop outcompeted the 5-step transaction chain
at every weighted re-draw, so runs filled with account creation and
rarely exercised the balance properties. Stop offering add-account once
three accounts exist; the renormalized weights then favor the
transaction flow at every step of its chain.
* fix(folio): author spec weights to match testing intent
Revert the account saturation gate: it starved newAccountBalanceIsZero
once it tripped, and a magic account count is app-state tuning, not
intent. Instead weight the generators by what the properties need:
the transaction chain leads, account creation stays exercised, and
doubleTaps gets explicit weight everywhere because double-submission
idempotency is what the spec is testing for.
* fix(folio): lower doubleTaps weight to 5
* fix(sidecar): surface visible text on iOS static elements
Static text and button strings live in the accessibility label on
iOS, so the text attribute came through empty and every balance
extractor parsed to zero, silently disarming both folio properties.
Non-editable elements now fall back title, value, then label;
editable fields keep value-only so an empty field's caption does not
read as content.
* feat(driver): native DoubleTap RPC for a tight inter-tap gap
Composing two Tap round trips from the Go client spread the taps by
hundreds of milliseconds on iOS, wide enough for the app to navigate
between them, so double-submission races could never reproduce. The
sidecar now lands both taps back-to-back next to the device transport.
* feat(sidecar): pipeline iOS double-tap requests
Queue the second tap at the XCTest runner while the first executes.
The runner serializes handlers, so this is the tightest gap the
transport allows (~350ms per tap round trip); recorded here with
measurements for the iOS double-tap limitation.
* refactor: rename inspect to replay across the codebase
Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/,
the CLI subcommand from `sanderling inspect` to `sanderling replay`, and
updates all references in docs, Makefile, README, and Go comments.
* feat(replay-ui): show spec filename with full path on hover
RunList and RunDetail now render the basename of spec_path (e.g.
login.spec.ts) with the full path available as a title tooltip.
* feat(ltl): bound fields on AlwaysFormula and named thunks
Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.
* feat(ltl): negation normal form pass
nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.
* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse
Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.
* test(ltl): property-based NNF laws
Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.
* test(ltl): Finalize, bounded eventually, latch, collapse
Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.
* feat(inspect): within clause on always residual node
A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.
* feat(ltl): witness violations and (bool,error) predicate thunks
* test(ltl): migrate thunk call sites to (bool,error)
* feat(ltl): flag thrown-predicate witnesses with IsError
* refactor(verifier): replace predicate err side-channel with violation witness
* test(verifier): witness API for thrown predicates
* feat(trace): witnesses map and skipped-verification marker on Step
* feat(runner): thread violation witnesses, finalize, skip marker into trace
* test(ltl): lock violation witness reason, IsError, and step
* test(verifier): finalize surfaces unmet eventually with witness
* fix(ltl): eliminate implies and bounded-always false-negatives
Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.
* test(ltl): lock implies and bounded-always false-negative regressions
* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick
* feat(testrun): inject seed into web bundle via SANDERLING_SEED define
* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring
* test(spec): add Go math/rand/v2 PCG oracle and golden fixture
* feat(spec): bit-exact PCG port of Go math/rand/v2
* test(spec): assert pcg.ts matches the PCG golden fixture
* feat(spec): shared input corpus and press-key pools
* feat(spec): action-tree types and Host interface
* feat(spec): verb support matrix and warn-once helper
* feat(spec): deterministic shared action picker
* test(spec): verb matrix and warn-once semantics
* test(spec): picker draw-order and determinism
* refactor(spec): actions.ts returns pure GeneratorNode data trees
* refactor(spec): wire from() sampling through the picker rng
* feat(spec): shared runtime-entry installs next-action over pick.ts
* feat(spec): export LongPress/Scroll/longPresses/scrolls factories
* test(spec): assert data-tree shapes for action factories
* test(spec): runtime-entry serializeAction wire-contract round-trip
* refactor(spec): bridge data-tree nodes to the legacy goja picker tags
* fix(spec): web runtime walks the spec's globalThis.actions data tree
* test(spec): tolerate legacy bridge fields on builtin nodes
* refactor(spec): installRuntime accepts a lazy root resolver
The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.
* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker
Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).
* test(spec): cover the WEB Host surface and seed precision
Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.
* refactor(spec): picker emits native selector + scroll endpoints, setup precedence
* feat(spec): goja runtime entry wires the shared picker over the Go host
* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin
* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker
* refactor(spec): drop the legacy goja bridge fields from action factories
* feat(spec): serialize selector-only string targets for the runner to re-resolve
* refactor(verifier): one DecodeAction reads the unified flat wire contract
* refactor(verifier): goja host + shared picker replace the duplicate Go picker
* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime
* test(verifier): author specs through the shared picker path
* test(runner): bundle authored specs with the goja runtime entry
* feat(verifier): collect unsupported verbs for the run report
* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource
* feat(testrun): surface unsupported verbs in run report
* test(verifier): cross-runtime goja/node parity gate on the shared picker
* test(verifier): unsupported verbs collected deduped in first-seen order
* test(runner): summary reports no unsupported verbs on a clean run
* test(spec): golden-fixture cross-runtime parity gate for the node picker
Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.
* test(verifier): assert goja picker against the same cross-runtime golden
Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.
* refactor(spec): rename pressKey generator export to pressKeys
* refactor(spec): update barrel re-exports for pressKeys
* test(spec): update pressKeys generator export name
* docs(spec): rename pressKey generator to pressKeys
* refactor(spec): extract samplerRng into shared sampler-rng module
* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)
* test(spec): cover fluent value generators determinism and chaining
* refactor(bundler): inject globalThis trailer from spec named exports
* refactor(bundler): reuse registration trailer in web bundler
* test(bundler): cover named-export globalThis registration
* feat(spec): add named() to Extracted handle type
* feat(web-runtime): named() and cross-extractor read guard
* feat(verifier): named() and cross-extractor read guard in goja
* test(verifier): cross-extractor read guard and named()
* test(web-runtime): export runtime and extractors for tests
* test(web-runtime): named() and cross-extractor read guard
* refactor(folio): drop manual globalThis trailer (bundler injects it)
* refactor(folio): seed txn amounts via integers().between(1,500)
* refactor(folio-web): drop manual globalThis trailer (bundler injects it)
* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs
* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts
* refactor(folio-web): name extractors so violation witnesses are readable
* fix(web-runtime): propagate extractor getter throws and unpoison locked global
Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.
* test(spec): install fake runtime via defineProperty to survive locked global
* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors
* feat(runner): add MaxSteps bound to Options
* test(runner): MaxSteps stops after exactly N steps
* test(driverpb): drop proto getter round-trip tautology
* test(sidecar): drop stub-mode placeholder tautology tests
* test(mock): drop default-field-value assertion test
* test(ltl): drop Verdict.String tautology tests
* refactor(runner): extract RenderSummary for snapshot testing
* test(runner): golden snapshots for trace stream and violation summary
* feat(web-runtime): capture uncaught errors into state.exceptions
* test(integration): add throwing and counter web fixtures
* test(integration): add specs for the web fixtures
* test(integration): drive web fixtures through the real pipeline in headless Chrome
* chore(make): add test-browser target for the Chrome-driven suite
* ci: run the Chrome-driven browser suite in a separate job
* refactor(test): relocate browser suite to test/browser
* refactor(permissions): delete dead internal/permissions package
* refactor(test): rename package to browser_test
* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets
* chore(make): point test-browser at test/browser
* docs(decisions): record internal/permissions deletion
* refactor(doctor): use sidecarassets package
* refactor(testrun): use sidecarassets package
* fix(test): resolve testdata relative to browser_test.go
* refactor(verifier): remove dead __sanderlingIndex compat alias
* refactor(bundler): use encoding/json for JS string literals
* docs(action-space): use vendor-neutral native driver wording
* refactor(hierarchy): scrub backend tool name from comments
* refactor(driver): scrub backend tool name from comments
* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver
* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap
* refactor(chrome): implement DoubleTap as two taps with the gap
* refactor(mock): record DoubleTap and DoubleTapSelector actions
* refactor(runner): delegate double-tap to driver, drop gesture timing
* test(runner): assert double-tap delegates to driver DoubleTap
* docs(cmd): add package docs to CLI and developer tools
* docs(driver): add package docs to driver interface and chrome backend
* docs(driver): add package docs to mock and sidecar backends
* docs(platform): add package docs to android and ios device prep
* docs: add package docs to bundler and inspect
* docs(ltl): add package doc to temporal logic evaluator
* docs: add package docs to runner and testrun pipeline
* docs: add package docs to trace and verifier
* docs(sidecarassets): add package doc for embedded JAR loader
* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI
* test(chrome): gate real-Chrome driver tests behind the browser tag
* chore(make): run chrome driver tests in the browser job
* fix(web-runtime): guard global error listeners for non-browser hosts
The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.
* ci(browser): re-enable unprivileged user namespaces for headless Chrome
ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.
* ci(browser): pin stable Chrome for the driver tests
setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.
* feat(defaults): add scroll and rebalance action weights
Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.
* feat(defaults): trim scroll action weight wiring
* fix(build): point sidecar jar ignore and embed paths at sidecarassets
* test(defaults): drop stale longPresses re-export assertion
longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.
* fix(chrome): raise DevTools websocket read timeout to 60s
Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.