Commit Graph
100 Commits
Author SHA1 Message Date
pj fc54e86b17 feat(sidecar): add LongPress client method 2026-05-31 18:01:38 +05:30
pj 55352f72c5 feat(driver): add LongPress to DeviceDriver interface 2026-05-31 18:01:24 +05:30
pj 1ffca2e560 chore(proto): regenerate Go stubs for LongPress 2026-05-31 18:01:18 +05:30
pj 26ed8e7760 feat(proto): add LongPress RPC 2026-05-31 18:01:09 +05:30
pj 9f7f2f5b21 test(sidecar): assert InputText types at cursor without clearing
Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.
2026-05-31 17:52:45 +05:30
pj 4d497cce56 fix(sidecar): type text at cursor instead of clearing the field
InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.
2026-05-31 17:52:24 +05:30
pj 17fc698e10 fix(chrome): launch with no-sandbox so headless Chrome starts in CI 2026-05-31 17:09:35 +05:30
pj 726e5e0a20 fix(build): rebuild sidecar JAR when Kotlin sources change
Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.
2026-05-31 15:56:30 +05:30
pj 37079a0d84 test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases
Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.
2026-05-31 15:49:55 +05:30
pj b88b6c239d fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn
The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.
2026-05-31 15:49:51 +05:30
pj 35fce05b65 fix(folio): extract balanceMatchesAddedSum predicate as testable helper
Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.
2026-05-31 15:49:45 +05:30
pj e9c6066ba4 refactor(inspect-ui): rename Step.action to Step.next_action
Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.
2026-05-31 15:41:25 +05:30
pj f237eb2c4d test(inspect): update fixtures to use next_action trace field
Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.
2026-05-31 15:40:34 +05:30
pj 7916123458 refactor(inspect): decode trace step's next_action JSON field
Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.
2026-05-31 15:40:04 +05:30
pj 332e7a034d refactor(runner): assign trace action to Step.NextAction field
Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.
2026-05-31 15:39:43 +05:30
pj 4ee3d97cff refactor(trace): rename Step.Action to Step.NextAction
The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.
2026-05-31 15:39:32 +05:30
pj 8d29b2906f test(runner): cover transitional step skips verifier and clean control 2026-05-31 15:36:04 +05:30
pj d83b2f8da1 fix(runner): skip verifier for transitional trees after retry budget
When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.
2026-05-31 15:35:03 +05:30
pj 2d2c11f830 feat(trace): add Transitional flag to Step 2026-05-31 15:33:25 +05:30
pj 56b762deac test(driver): cover Snapshot in proto descriptor and sidecar client
Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.
2026-05-31 15:28:50 +05:30
pj 5c58610181 test(runner): assert step uses Snapshot, not raw hierarchy/screenshot
TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.
2026-05-31 15:27:45 +05:30
pj 1931c0b57f refactor(runner): observe each step via the atomic Snapshot RPC
fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.
2026-05-31 15:27:39 +05:30
pj c57be03b69 feat(driver): add Snapshot to chrome and mock drivers
The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.
2026-05-31 15:26:03 +05:30
pj c11beb1728 feat(driver): expose Snapshot on DeviceDriver and sidecar client
Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.
2026-05-31 15:25:58 +05:30
pj 8f29c63d12 test(sidecar): cover Snapshot wire path and serialization lock
SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.
2026-05-31 15:25:01 +05:30
pj 389bbf10d1 feat(sidecar): wire Snapshot handler with serialization lock
Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.
2026-05-31 15:23:07 +05:30
pj 4bc3147d2a feat(sidecar): add snapshot default on DriverBackend
Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.
2026-05-31 15:22:40 +05:30
pj 5a19ef885f feat(proto): add Snapshot RPC for atomic hierarchy+screenshot
Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.
2026-05-31 15:22:22 +05:30
pj a1390aecdd test(runner): cover startup gate waiting for app window to draw 2026-05-31 14:35:04 +05:30
pj 5be6eaaad5 test(mock): add FocusedWindowApp with foreground mirroring 2026-05-31 14:34:40 +05:30
pj 06213e444e fix(runner): gate first observe on the app window being drawn, not just resumed 2026-05-31 14:33:09 +05:30
pj 3fe92537be feat(driver): add FocusedWindowChecker capability 2026-05-31 14:30:36 +05:30
pj 2fe01ea254 feat(android): detect focused-window package via dumpsys window 2026-05-31 14:30:10 +05:30
pj 05125b69f5 chore: stop tracking inspect-ui/dist build artifacts 2026-05-31 13:38:40 +05:30
pj d733218c40 test(hierarchy): cover package derivation from resource-id 2026-05-31 12:59:20 +05:30
pj 8401ae908f feat(hierarchy): derive package from resource-id prefix
The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.
2026-05-31 12:59:20 +05:30
pj 6d3561bf6c test(verifier): cover package-scoped target selection 2026-05-31 12:50:30 +05:30
pj 6d0f81ffa2 feat(testrun): pass app package into verifier scope filter 2026-05-31 12:50:30 +05:30
pj 15f3906f1c feat(verifier): scope random-action targets to app package
Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.
2026-05-31 12:50:26 +05:30
pj 16a9136b28 test(runner): cover startup foreground gate and back-press 2026-05-31 12:12:58 +05:30
pj f22e1bada8 feat(runner): gate first action on app reaching foreground 2026-05-31 12:12:58 +05:30
pj ddec95a2c5 feat(runner): re-fetch on transitional hierarchy capture
Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.

fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.
2026-05-31 12:11:29 +05:30
pj 8deef1a425 test(sidecar): cover streak reset and route-transition rejection
Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.
2026-05-31 12:11:20 +05:30
pj f32fbe540b feat(sidecar): streak-based settle with route-transition detection
Two changes layered into the stability poll:

1. stabilitySnapshot returns null while the tree carries more than one
   route-level Screen tag (resource-id / testTag / identifier ending
   in "Screen"), so the poll cannot declare a NavHost cross-fade
   stable. Apps following the Compose route convention get this
   detection for free; apps that don't fall through to the generic
   signal below.

2. pollUntilStable now requires an uninterrupted stable streak of at
   least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
   matches. A late transition that fires after a brief calm window
   breaks the streak instead of slipping past. Interval widened to
   250ms so UiAutomation isn't hammered under fuzz load.
2026-05-31 12:11:14 +05:30
pj 0abbcd7d0c feat(cli): default --clear-data on so runs start fresh 2026-05-31 12:10:05 +05:30
pj a10791e00f fix(sidecar): cap stability poll independently of settle budget
The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.
2026-05-30 22:09:26 +05:30
pj d0cf5a98c7 feat(inspect-ui): render extractor-change breadcrumbs at violations
Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.
2026-05-30 21:58:13 +05:30
pj e44ad56749 feat(trace): emit extractor_changes per step
Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.
2026-05-30 21:58:08 +05:30
pj c66c4e0ad4 test(verifier): cover ChangedExtractors diffs
Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.
2026-05-30 21:58:04 +05:30
pj 419a2d564e feat(verifier): track extractor value transitions
Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.
2026-05-30 21:57:59 +05:30
pj 07c292913c chore(folio): name every extract() call
Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.
2026-05-30 21:54:35 +05:30
pj 1b35e5a6a6 feat(verifier): name extractors for diff surfacing
bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.
2026-05-30 21:54:30 +05:30
pj cbef329cdd test(spec): cover extract name overload
Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.
2026-05-30 21:54:25 +05:30
pj 6fdbf9d2ab feat(spec): accept optional name on extract()
Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.
2026-05-30 21:54:20 +05:30
pj 7291d5965b test(sidecar): cover pollUntilStable and structuralHash
Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.
2026-05-30 21:50:52 +05:30
pj cd4bd5a194 feat(sidecar): structural-hash settle poll
Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.
2026-05-30 21:50:46 +05:30
pj 50e03244d6 refactor(inspect-ui): use next step's screenshot for state after
Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.
2026-05-30 21:48:26 +05:30
pj fb69c1b774 refactor(runner): one concurrent screenshot per step
Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.
2026-05-30 21:48:21 +05:30
pj 095b58afcf refactor(trace): drop WriteScreenshotAfter
Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.
2026-05-30 21:48:17 +05:30
pj f9512354ed refactor(folio): replace txn invariants with balanceMatchesAddedTxn
Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.
2026-05-30 21:44:41 +05:30
pj 352118c199 fix(verifier): canonicalize selector strings
Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.
2026-05-30 21:43:27 +05:30
pj 09c1b8df26 fix(folio): make login spec content-driven (idempotent across re-entries) 2026-05-30 17:04:11 +05:30
pj c69052ae89 style(verifier): use maps.Copy for verdict snapshot 2026-05-30 16:33:23 +05:30
pj 60c4ef7458 refactor(runner): emit onset-only violations to trace and summary
Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.
2026-05-30 16:32:17 +05:30
pj b9fa41553f feat(verifier): track newly-violated property set per step
Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.
2026-05-30 16:30:37 +05:30
pj 3d67e5e3b6 fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives 2026-05-30 16:01:16 +05:30
pj aa42b8e504 refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions 2026-05-30 16:00:14 +05:30
pj 801f09d455 feat(verifier): add doubleTaps random-target generator 2026-05-30 16:00:09 +05:30
pj 177b571b0a feat(spec): add doubleTaps random-target builtin to defaultActions 2026-05-30 16:00:06 +05:30
pj f08d9119b2 fix(folio): track ledger row count across non-ledger steps; pin reproducer seed 2026-05-30 11:44:43 +05:30
pj 9f17afecc2 feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action 2026-05-30 11:07:31 +05:30
pj 87834caf96 test(doubleTap): cover constructor, verifier round-trip, and runner dispatch 2026-05-30 11:05:58 +05:30
pj fb085eeaa0 feat(runner): dispatch DoubleTap as two taps inside one step 2026-05-30 11:05:53 +05:30
pj 6766f1ad58 feat(verifier): bind doubleTap and decode DoubleTap actions 2026-05-30 11:05:50 +05:30
pj 237ead2d33 feat(spec): wire DoubleTap through web-runtime serializer 2026-05-30 11:05:47 +05:30
pj 7006e82ca4 feat(spec): add DoubleTap action type and constructor 2026-05-30 11:05:44 +05:30
pj a9b2b4b4e1 fix(spec): drop hardware back from defaultActions to stay in-app 2026-05-28 11:24:31 +05:30
pj dff91095b8 feat(runner): relaunch app when foreground escapes during exploration 2026-05-28 11:23:50 +05:30
pj dea39d3bac feat(sidecar): implement ForegroundApp via adb for android 2026-05-28 11:21:52 +05:30
pj 09528d84e6 feat(android): detect foreground package via adb dumpsys 2026-05-28 11:21:14 +05:30
pj 4d26bdf00d feat(driver): add ForegroundChecker optional capability 2026-05-28 11:19:49 +05:30
pj 032129cc5f chore(folio): auto-boot a bootable AVD in just test/install when none connected 2026-05-25 19:11:11 +05:30
pj 64d0633bd6 feat(spec): typing builtin for the web (V8) action path 2026-05-25 18:42:30 +05:30
pj 9df104e5b2 test(chrome): editable flag for inputs, textarea, contenteditable 2026-05-25 18:28:32 +05:30
pj b8b183e970 fix(testrun): alias @sanderling/spec/defaults for the bundler 2026-05-25 18:21:12 +05:30
pj 761627dde3 test(spec): defaultActions, typing, and defaults barrel resolve 2026-05-25 18:16:52 +05:30
pj be0d071826 test(hierarchy): editable derivation and selector matching 2026-05-25 18:16:20 +05:30
pj b59a62643d test(verifier): typing builtin targets editable fields, declines otherwise 2026-05-25 18:15:52 +05:30
pj bf235e54c2 feat(folio): layer defaultActions breadth over targeted flows 2026-05-25 18:14:40 +05:30
pj 7fd8b6d347 feat(spec): export @sanderling/spec/defaults subpath 2026-05-25 18:13:36 +05:30
pj c8528af12b feat(spec): add defaultActions bundle 2026-05-25 18:13:36 +05:30
pj 9ca16e7627 feat(spec): export typing builtin generator 2026-05-25 18:13:13 +05:30
pj 923d06813e feat(verifier): typing builtin types edge-case corpus into editable fields 2026-05-25 18:12:31 +05:30
pj e079574571 feat(verifier): register typing builtin generator 2026-05-25 18:12:08 +05:30
pj bf922fcbd5 feat(spec): add editable to selector and element types 2026-05-25 18:11:56 +05:30
pj f2ebedd82a feat(verifier): expose editable on ax element objects 2026-05-25 18:11:56 +05:30
pj 5d21b8c2cf feat(chrome): emit editable flag in hierarchy dump 2026-05-25 18:11:41 +05:30
pj 1d01ca77dc feat(hierarchy): add editable signal with native derivation 2026-05-25 18:11:01 +05:30
pj f572c8ba66 WIP: docs: refresh after iOS + web support (#50)
* docs: README covers iOS + web, surface both example apps

* docs(cli): document --ios-device and per-platform doctor

* docs: tighten README, fold examples into Docs list

* docs(runs): correct --clear-data lifecycle wording

Default behavior no longer wipes app data between runs; --clear-data is now opt-in.

* docs(getting-started): add iOS path, separate folio and folio-web

Document just test-ios under examples/folio, and distinguish the KMP
sample from the React + Vite folio-web sample.

* docs(inspect): document the eight panels

Lists Screenshot, ActionList, Timeline, ViolationsPanel, HierarchyPanel,
SnapshotTable, MetricsChart, ExceptionsPanel. Cross-links HierarchyPanel
to the spec language reference.

* docs(writing-specs): document setup export, flag noLogcatErrors as android-only

Mirrors pkg/spec/README.md so the manual covers the runner's setup-first
fall-through. Marks noLogcatErrors as Android-only so iOS/web spec
authors know it silently no-ops.

* docs(folio): document web target and iOS sanderling test recipe

After the KMP refactor folio also runs on wasmJs and the justfile exposes
just web, just web-build, and just test-ios. Surface all three.

* docs(folio-web): add README

Covers prerequisites, demo credentials, just test recipe, and how the
React + Vite host exposes state to the sanderling spec via stable ids
and data-* attributes.

* docs: scrub driver-implementation name from user docs

Drop the implementation tool name from README, cli.md doctor table, and
spec-language.md. These docs should describe behaviour, not the specific
underlying tool the native sidecar wraps.

* docs(development): scrub driver-implementation name from dev docs

architecture, design-principles, decisions now describe the native
sidecar by role (gRPC surface over OS UI-test pipeline) rather than by
the specific tool it wraps.
2026-05-25 16:17:21 +05:30
pj b23fb0c723 feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag

Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.

* feat(testrun): add Preflight() before sidecar/driver setup

Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.

* refactor(chrome): split tag (HTML name) from class (CSS classList)

Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.

* feat(chrome): translate legacy string selectors to CSS/XPath

TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.

* feat(trace): add WriteHTML + Step.HTMLAvailable

Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.

* feat(driver): add WebDriver capability + chrome implementation

WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.

* feat(verifier): OverrideExtractorValues for V8-driven extractors

Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.

* feat(spec): add WebState + camelCase attribute aliases

WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.

* feat(runner): per-tick HTML capture for WebDriver-capable drivers

Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.

* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>

Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.

* feat(bundler): BundleWeb + V8-side runtime shim

web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.

* feat(runner): V8 extractor overrides + V8 action source for WebDriver

When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.

* feat(testrun): bundle + install web runtime when platform=web

BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.

* feat(inspect-ui): hierarchy + html panels in run detail

HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.

* fix(folio-web): drop aria-label data-carrier abuse

Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.

* chore: rebuild inspect-ui dist + folio-web .gitignore

Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.

* revert(trace): drop WriteHTML + Step.HTMLAvailable

Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.

* revert(runner): drop per-tick HTML capture

Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.

* revert(driver): drop WebDriver.Document

Document was only consumed by the runner's HTML capture which is gone.

* revert(inspect): drop /html route

Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.

* revert(inspect-ui): drop htmlUrl + html_available type

API surface no longer needs the HTML route; Step.html_available has no
producer.

* revert(inspect-ui): drop HtmlPanel + html tab

Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.

* test(inspect-ui): drop htmlUrl test, add @types/bun

Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.

* chore: rebuild inspect-ui dist without HtmlPanel

Embedded SPA bundle no longer ships the iframe HTML viewer.

* fix(web-runtime): retry action resolution + implement taps/swipes

V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.

* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys

Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.

* chore(folio-web): drop swipes from action root

Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.

* fix(inspect-ui): correct HierarchyPanel CSS variable names

Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.

* fix(chrome): correct PressKey mappings to chromedp/kb constants

Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.

Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.

* fix(cli): -h/--help exits 0 instead of error code

parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.

Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.

* fix(chrome): harden cssEscape for control chars + use [class~=]

Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.

Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.

* fix(web-runtime): use CSS.escape and validate tag-name selectors

The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.

The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.

Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.

* fix(chrome): validate attribute name in unknown-prefix branch

A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.

* fix(selectors): emit valid XPath 1.0 string literals via concat()

Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.

Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.

* fix(runtime): surface unresolved action targets instead of dropping silently

serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.

Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.

* fix(runner): use errgroup-bound ctx so siblings cancel on failure

The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.

* fix(chrome): propagate caller ctx cancellation to CDP calls

InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.

Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.

* fix(verifier): tolerate out-of-range override indices

A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.

Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.

* test(verifier): cover object-shaped extractor overrides

Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.

* fix(web-runtime): lock global runtime hooks against page shadowing

AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.

* perf(web-runtime): cache randomTap candidate DOM scan per tick

The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.

Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.

* fix(web-runtime): cap sanitize recursion to prevent stack overflow

State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.

* fix(web-runtime): enforce pressKey allowlist in factory

The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.

* chore(chrome): drop dead bundleSource/bundleMu

bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.

* fix(chrome): use strconv.Atoi for extractor key parsing

fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.

* fix(doctor): raise per-check timeout to 15s for chromium launch

5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.

* fix(runner): trust V8 coordinates for InputText, even at origin

resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.

Add applyAction tests covering both the typical web case and the (0,0)
edge case.

* test(bundler): lock down deterministic output across builds

The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
2026-05-03 11:21:21 +07:00