mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
a8444cc34f95700b077b6cba41d10d1e0b786084
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a8444cc34f |
feat(spec): match idPrefix in the web runtime
The DOM has no package prefix, so the native rule reduces to [id^=]. Both prefix kinds now go through the one key table, which drops the separate descPrefix branch that string and object selectors each carried. |
||
|
|
76dce1a75e |
experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs runner.Options.MaxSteps already worked but was unreachable from the command line. A step budget is what makes two generators comparable: one making a model call per step and one drawing from a PRNG are not comparable per second. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record arm membership and host in meta.json meta.json recorded the seed but not which picker ran, how it was configured, what budget it was given, or which machine produced it. A directory of runs cannot be attributed to an experiment cell without those, which makes any factorial computed from such a directory unanalysable after the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(cli): add --arm and populate run meta from it Model and instructions are recorded only when the LLM picker is the one that will actually run, so a spec declaring generator = llm() that is run under the seeded picker does not label its trace with a model it never called. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): sweep seeds for one experiment cell campaign.json lists the seeds a sweep intended to run and is written before the first run, so a host that dropped runs shows up as missing seeds rather than as a smaller sample. Seed 0 is rejected: sanderling test reads it as "derive a seed from the clock", which is why conformance/gates.sh controls nothing today. Each run contributes one runs.jsonl line carrying steps to first violation by origin step, the step that armed the failed obligation, so the survival analysis never reopens a trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): no silent generator fallback, and llm on web --generator llm against a spec declaring no generator = llm(...) logged a warning and ran the seeded picker. For a comparison campaign that is silent arm corruption: the run completes, the directory looks correct, and the wrong policy drove it. It is now fatal. pickSources also returned the V8 source for both action and extractor on web before it looked at the generator, so the llm policy was unreachable there. The two axes are now independent: the driver picks the extractor source, the flag picks the action source, and llmSource composes with either because the runner populates the candidate list and screenshot on every platform. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): make the hierarchy dump agree with the web runtime Three facts differed between the dump the goja host reads and the DOM the V8 host reads, so the two enumerated different candidates on one page. scrollable was never emitted, and worker.go reads exactly that attribute while targets.ts requires it for scrolls, so the goja host could not offer a single web scroll. clickable tested el.onclick, which React assigns to its root container for event delegation, making the whole viewport a tap target here and in no other enumeration. Both now resolve through the selector sets in pkg/spec/src/web-runtime.ts. The dump also rooted at body while collectTargets walks querySelectorAll("*"), so the goja host never saw html, where page-level scrolling lives. It now roots at documentElement and skips the head subtree, which is all zero-bounds and would otherwise carry script and title text into the trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(conformance): give the gate reproducible seeds SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from the clock", so the tunable controlled nothing and a gate failure could not be re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the results table so a failing row names its stream. The five runs stay on five different streams: a gate that scored one path five times would catch less than one that scores five. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): emit editable as a plain boolean editable was emitted as `isEditable || null`, and an absent field sends internal/hierarchy into the native fallback, which reads any class name containing "EditText" as an Android text widget. On web that is just a CSS class, so a page styling a div with it was editable to the goja host and not to the web runtime, and the model policy could be offered typing into a div. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): leave the head subtree out of the web target walk collectTargets walked querySelectorAll("*") while the hierarchy dump skips head, so the two hosts enumerated different element sets on every page with a <head>. No candidate changes: builtinCandidates pushes only for targets acceptsTarget admits, and head elements have no positive bounds, so the list the draw ranges over is untouched. What changes is that targetIndex now means the same thing on both hosts. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(chrome): compare the facts both hosts derive from one DOM The existing parity harness hand-authors the facts on both sides, so it proves that given identical facts both hosts select identical candidates, and says nothing about the two code paths that derive those facts from a real page. Four divergences lived in that blind spot and it passed throughout. This drives one real page and compares clickable, enabled, editable, scrollable and positiveBounds element by element, plus the element sets themselves, which is what catches a host that omits html or includes head. Reverting any of the four fixes makes it fail naming the element and the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * chore(make): run the browser packages one at a time Both launch Chrome and launching two at once has failed with "Launch: context canceled". Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style: remove every em-dash and en-dash Eighteen occurrences across fourteen files. Each sentence was repunctuated to suit what the dash was doing rather than swapped for a hyphen, which produces comma splices. The minus sign in folio-web's ledger is a minus sign and stays. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): honor the caller context in Launch Launch and clearState ran against d.tabCtx, so a target that accepts the connection and never answers wedged the process past its own --duration and through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for the rest of the sweep with no diagnostic. The browser is still allocated against d.tabCtx first, because chromedp starts Chrome under whichever context calls Run first and allocating under a caller deadline would kill the browser when Launch returns. Everything after allocation goes through runCtx. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(sidecarassets): publish the extracted jar through a rename Extract wrote a 96 MB jar with a plain WriteFile into a temp path every sanderling process on the host shares. On a cold host several concurrent workers all miss the checksum and all write the same path, and O_TRUNC lets one spawn a JVM against another's half-written archive. A fresh experiment host is exactly a cold host. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): kill a run that outlives --run-timeout A wedged run holds its worker for the rest of the sweep, and on an unattended host nothing else will send it a signal. Defaults to three times --duration and must exceed it. A killed run is recorded as timed_out rather than as a generic failure, so the analysis can tell a lost cell from a real crash. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style(test): gofmt browser_test.go Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX |
||
|
|
26b49b379a |
fix ltl semantics and unify action enumeration (#71)
* fix(ltl): give every thunk a construction identity Two distinct unnamed predicates both described as "Thunk(...)", so obligation collapse merged their residuals and could drop a live violation. Identity is assigned at construction and the fields are unexported, so a thunk cannot be built without one. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(ltl): reduce a thrown-predicate residual instead of panicking The verifier substitutes an ErrorFormula for the residual of a property whose predicate threw, and that residual is fed back in on the next step. reduce had no case for it, so the run crashed. It re-reports the same failure now. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(ltl): make a bounded always the dual of a bounded eventually G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still pending when the window closed, so nnf's negation normal form was not semantics preserving. Both sides now range over the observations at which their inner can definitely resolve: the eventually keeps a pending inner as a disjunct, and the always discharges vacuously at window close. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(ltl): arm a one-shot root once per run A root that carries its own horizon is one obligation for the whole run, not one per observation. Re-instantiating a top-level eventually monitored G F<=n(p) instead of F<=n(p) and left one live obligation per step behind; a bounded always restarted its window every step and never closed. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(verifier): stop wrapping a top-level eventually in always `eventually(p).within(300, "seconds")` as a property meant "within 300 seconds of every step", which spawned an obligation per step with its own resolved deadline. A 553-step run carried 553 of them and serialized a 75 KB residual. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(ltl): serialize the resolved deadline of a bounded window Two obligations spawned at different steps from one duration-bounded formula differ only in the deadline the evaluator resolved for them, so they serialized identically and the trace erased a distinction the evaluator makes. The authored window stays in amount/unit; the resolved deadline rides alongside. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(verifier): split a witness's origin step from its detection step A deferred obligation spans two steps: the one that armed it and the one whose reduction failed. They were conflated under one index, so the extractor snapshot (which is the detecting step's state) was reported against the origin step. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(runner): record a witness's detection step in the trace Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * feat(replay-ui): show the step a violation was detected at The witness evidence is the detecting step's state, so say which step that is and let a reader jump to it. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(verifier): record the extractor state the predicates actually read On the web path extractor bodies are evaluated in V8 and injected here, but only the goja value was replaced. The trace diff and the violation witness therefore described a state no property ever saw. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * refactor(spec): one candidate producer over one target-eligibility rule Both hosts routed verbs themselves and both policies enumerated their own actions, and all four drifted. Web sent `swipes` to scrollable containers only, so swipe-to-dismiss on a list row was reachable on native and unreachable on web; the model policy folded gestures its own way and could not reach what the seeded picker drew. A host now reports facts about every element and never decides which verb may act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates is the single enumeration, and the model policy reads it through __sanderlingEnumerateBuiltin__ instead of reimplementing it in Go. Gesture verbs change with it: scrolls stay vertical over scrollable containers, swipes go free-form in all four directions from any element with real bounds. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(runner): name a builtin scroll by its drag origin A builtin gesture carries endpoints and no selector, so every scroll rendered as "Scroll down " in the prompt's recent-action memory and two scrollable regions were indistinguishable. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(chrome): clear storage over cdp instead of scripting an opaque origin Launch runs while the tab is still on about:blank, whose opaque origin denies storage access, so localStorage.clear() threw SecurityError and every web run died at launch. Storage.clearDataForOrigin needs no navigation. The exception helper lands here because "Uncaught" is what hid this for so long. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(chrome): enable the swiftshader webgl fallback Headless Chrome runs with --disable-gpu, and without this flag it refuses the software WebGL backend: getContext returns null, so a canvas-rendered app paints nothing and every screenshot is identical black. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * fix(web): resolve testTag through data-testid or id Compose Multiplatform emits its testTag into the element id, which the native table already accepts via the resource-id alias. The two web selector tables were the only place that rejected it. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * test(spec): type-check the spec api as part of make test The fake runtime in api.test.ts did not return a chainable handle from extract, so the file had not type-checked since named() was added. Wiring the check into make test stops it drifting again. Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J * docs(manual): one-shot eventually and the gesture verbs Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J |
||
|
|
94d9511312 |
test: full test-suite refactor sweep (#61)
* chore(test): start test-suite refactor sweep * test(ltl): pin exact multi-obligation residual AST * test(ltl): table-test finalize Kleene connective combinations * test(ltl): pin reduce over pending inner for bound, Or, Not * test(ltl): marshal bounded Always steps/duration/deadline * test(verifier): cover LTL combinator verdict transitions and within unit panic * test(verifier): table-test DecodeAction kinds and lastAction field exposure * test(verifier): assert WithPlatform(ios) reaches the picker host and key pool * test(verifier): widen weighted-selection assertion to a 5x skew margin * test(verifier): un-skip ax-find round trip with a committed tree fixture * test(runner): pin isWDADrop to sidecar reconnect-failed message origin * test(runner): assert PressKey/Wait trace encoding records kind-specific fields * test(runner): cover RenderSummary unsupported-verbs surfacing branch * test(trace): set Hierarchy in round-trip and lock lossy Tree contract Also add a -race concurrent WriteStep test that asserts N well-formed JSONL lines, catching torn lines if the writer mutex is dropped. * test(trace): round-trip witnesses/changes/metrics/exceptions, pin step-0 witness * test(trace): document ViolationsAreGreppable grep contract and lock-free WriteScreenshot * test(hierarchy): cover invalid-JSON and malformed-bounds parser paths * test(trace): guard writer mutex via WriteStep/Close race on w.file * test(replay): drop unfailable assets and devproxy assertions * test(replay): cache reuses on equal mtime, reparses after append * test(replay): violation marker falls back to detection step when attributed missing * test(replay): corrupt meta/trace dirs return 500 with error body * test(replay): SSE client receives runs.changed after a broadcast * test(replay): Run coalesces creates, ignores write/chmod, closes subs on cancel * fix(sidecar): synchronize health fixture writes and exercise healthError * test(sidecar): cover swipe/longpress/doubletap/erase/presskey/metrics/logs translations * test(sidecar): cover DoubleTapSelector composition and mid-gesture cancel * test(sidecar): assert gRPC error status surfaces from action RPC * fix(chrome): route action methods through runCtx so caller cancellation aborts CDP * fix(chrome): route hierarchy/screenshot/waitidle/metrics through runCtx * refactor(ios): extract pure simctl JSON parsers * refactor(ios): add command-runner seams for EnsureSimulator * test(ios): table-test simctl parsers and EnsureSimulator seams * test(sidecarassets): cover placeholder build path * test(sidecarassets): assert reuse via sentinel bytes not mtime * test(bundler): cover properties-only spec registration * refactor(testrun): extract prepareBundleInputs from Execute * test(testrun): cover prepareBundleInputs aliases and missing-runtime error * test(testrun): table-test resolveRuntimeSibling search edges * test(testrun): exact-output tests for progressHandler line format * fix(cmd): point bundle-check aliases at pkg/spec/src * test(cmd): smoke-test bundle-check resolves spec aliases * test(cmd): table-test hier-check parse and FindAll on fixture * test(cmd): unit-test buildBrowseURL deep-link vs root * test(cmd): drop flaky TestRun_Doctor that launched real Chromium * test(cmd): pin pipeline error to bundle resolution on web platform * test(replay-ui): add bun test script * ci(replay-ui): run bun test via make web-test target * ci(replay-ui): point bun cache key at replay-ui/bun.lock * test(replay-ui): exercise real URL encoding and non-ok throw in getJson * refactor(replay-ui): extract snapshot flatten/getAtPath into lib module * test(replay-ui): pin snapshot flatten/getAtPath path round-trip * refactor(replay-ui): extract action selector/format into lib module * test(replay-ui): pin action selector parse and row formatting * refactor(replay-ui): share one statusFor between panels * refactor(replay-ui): extract run-history derivation into lib module * test(replay-ui): pin shared statusFor precedence and ordering * test(replay-ui): pin run-history derivation alignment * refactor(replay-ui): export clampIndex for testing * refactor(replay-ui): extract keyboard-nav dispatch into pure module * refactor(replay-ui): extract metrics formatters into lib module * test(replay-ui): pin clampIndex step boundaries * test(replay-ui): pin keyboard-nav ownership and key routing * test(replay-ui): pin metrics formatters and path gap handling * refactor(sidecar): expose device-output parsers as internal for testing * test(sidecar): table-test device-output parsers against malformed input * test(sidecar): cover logcat parsing year inference and line skipping * test(sidecar): pin pressKey keycode mapping and unknown-key rejection * test(sidecar): metrics bundleId falls back to launched app and honors override * test(sidecar): loosen deadline upper bound to tolerate slow CI scheduling * test(web-runtime): export selector builders for unit tests * test(web-runtime): guard sanitize cycle, function, and depth limits * test(web-runtime): table-test selector builder quoting and escaping * test(sidecar): collapse scalar-forwarding RPC tests into a table * test(replay-ui): dedup step/summary fixtures into shared module * test(ios): collapse pickSimulator point-tests into a table |
||
|
|
c5bb176be8 |
UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed, and surface both in describe() and MarshalJSON. * feat(ltl): negation normal form pass nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error leaf, dualizing Always<->Eventually and preserving bounds. * feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse Apply nnf on construction, reduce bounded Always symmetric to bounded Eventually (vacuous holds once the window closes), add Finalize to resolve undischarged liveness obligations to Violated at run end, and collapse structurally-identical pending obligations. * test(ltl): property-based NNF laws Lock double-negation identity, Always/Eventually duality with bound preservation, leaf pushdown, and not(always true) reaching Violated. * test(ltl): Finalize, bounded eventually, latch, collapse Property tests for monotonic violation latch and eventually-within violating iff n consecutive false, plus Finalize and collapse cases. * feat(inspect): within clause on always residual node A negated bounded eventually serializes as a bounded always; render its bound instead of dropping it. * feat(ltl): witness violations and (bool,error) predicate thunks * test(ltl): migrate thunk call sites to (bool,error) * feat(ltl): flag thrown-predicate witnesses with IsError * refactor(verifier): replace predicate err side-channel with violation witness * test(verifier): witness API for thrown predicates * feat(trace): witnesses map and skipped-verification marker on Step * feat(runner): thread violation witnesses, finalize, skip marker into trace * test(ltl): lock violation witness reason, IsError, and step * test(verifier): finalize surfaces unmet eventually with witness * fix(ltl): eliminate implies and bounded-always false-negatives Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent can no longer defer the whole implication and drop a consequent that was false at the current step. Carry a pending inner past a bounded-Always window close instead of dropping it to holds, so a deferred obligation is resolved by a later step or Finalize. * test(ltl): lock implies and bounded-always false-negative regressions * fix(web-runtime): seed PRNG for reproducible runs and align weighted pick * feat(testrun): inject seed into web bundle via SANDERLING_SEED define * test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring * test(spec): add Go math/rand/v2 PCG oracle and golden fixture * feat(spec): bit-exact PCG port of Go math/rand/v2 * test(spec): assert pcg.ts matches the PCG golden fixture * feat(spec): shared input corpus and press-key pools * feat(spec): action-tree types and Host interface * feat(spec): verb support matrix and warn-once helper * feat(spec): deterministic shared action picker * test(spec): verb matrix and warn-once semantics * test(spec): picker draw-order and determinism * refactor(spec): actions.ts returns pure GeneratorNode data trees * refactor(spec): wire from() sampling through the picker rng * feat(spec): shared runtime-entry installs next-action over pick.ts * feat(spec): export LongPress/Scroll/longPresses/scrolls factories * test(spec): assert data-tree shapes for action factories * test(spec): runtime-entry serializeAction wire-contract round-trip * refactor(spec): bridge data-tree nodes to the legacy goja picker tags * fix(spec): web runtime walks the spec's globalThis.actions data tree * test(spec): tolerate legacy bridge fields on builtin nodes * refactor(spec): installRuntime accepts a lazy root resolver The web bundle imports the runtime before the spec, so the action root on globalThis.actions only exists after the spec evaluates. Accept a function form so the goja and web hosts resolve the root per tick. * refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/ randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG, and the snake_case serializeAction) plus the __sanderling__ action factory binds. web-runtime now implements Host (platform/seedHi/seedLo from the injected 64-bit seed via BigInt, queryCandidates over the live DOM with a per-tick cache, reportUnsupported) and calls installRuntime so both engines run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts matrix instead of silently returning null. Keeps the DOM helpers (selector translation, queryElement, elementHandle, buildState, sanitize, extractors) and the global locking. Net -214 lines (741 -> 527). * test(spec): cover the WEB Host surface and seed precision Replace the deleted-picker tests with Host coverage: platform()==web, seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0, reportUnsupported warning, the installed next-action/extractor globals, and queryCandidates verb routing + per-tick caching over a querySelectorAll stub. * refactor(spec): picker emits native selector + scroll endpoints, setup precedence * feat(spec): goja runtime entry wires the shared picker over the Go host * feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin * feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker * refactor(spec): drop the legacy goja bridge fields from action factories * feat(spec): serialize selector-only string targets for the runner to re-resolve * refactor(verifier): one DecodeAction reads the unified flat wire contract * refactor(verifier): goja host + shared picker replace the duplicate Go picker * refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime * test(verifier): author specs through the shared picker path * test(runner): bundle authored specs with the goja runtime entry * feat(verifier): collect unsupported verbs for the run report * refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource * feat(testrun): surface unsupported verbs in run report * test(verifier): cross-runtime goja/node parity gate on the shared picker * test(verifier): unsupported verbs collected deduped in first-seen order * test(runner): summary reports no unsupported verbs on a clean run * test(spec): golden-fixture cross-runtime parity gate for the node picker Replace the env-driven parity harness with a shared scenario module and a committed golden the node picker asserts independently. The goja side asserts the same golden, so neither runtime invokes the other at test time. * test(verifier): assert goja picker against the same cross-runtime golden Drop the node-subprocess coupling: the goja side now installs a stub __sanderlingHost__ with the fixed candidate list and asserts the committed golden, matching pkg/spec/test/parity.test.ts. * refactor(spec): rename pressKey generator export to pressKeys * refactor(spec): update barrel re-exports for pressKeys * test(spec): update pressKeys generator export name * docs(spec): rename pressKey generator to pressKeys * refactor(spec): extract samplerRng into shared sampler-rng module * feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText) * test(spec): cover fluent value generators determinism and chaining * refactor(bundler): inject globalThis trailer from spec named exports * refactor(bundler): reuse registration trailer in web bundler * test(bundler): cover named-export globalThis registration * feat(spec): add named() to Extracted handle type * feat(web-runtime): named() and cross-extractor read guard * feat(verifier): named() and cross-extractor read guard in goja * test(verifier): cross-extractor read guard and named() * test(web-runtime): export runtime and extractors for tests * test(web-runtime): named() and cross-extractor read guard * refactor(folio): drop manual globalThis trailer (bundler injects it) * refactor(folio): seed txn amounts via integers().between(1,500) * refactor(folio-web): drop manual globalThis trailer (bundler injects it) * fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs * refactor(folio-web): weight valid generators against edgeCaseText for names/amounts * refactor(folio-web): name extractors so violation witnesses are readable * fix(web-runtime): propagate extractor getter throws and unpoison locked global Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake. * test(spec): install fake runtime via defineProperty to survive locked global * test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors * feat(runner): add MaxSteps bound to Options * test(runner): MaxSteps stops after exactly N steps * test(driverpb): drop proto getter round-trip tautology * test(sidecar): drop stub-mode placeholder tautology tests * test(mock): drop default-field-value assertion test * test(ltl): drop Verdict.String tautology tests * refactor(runner): extract RenderSummary for snapshot testing * test(runner): golden snapshots for trace stream and violation summary * feat(web-runtime): capture uncaught errors into state.exceptions * test(integration): add throwing and counter web fixtures * test(integration): add specs for the web fixtures * test(integration): drive web fixtures through the real pipeline in headless Chrome * chore(make): add test-browser target for the Chrome-driven suite * ci: run the Chrome-driven browser suite in a separate job * refactor(test): relocate browser suite to test/browser * refactor(permissions): delete dead internal/permissions package * refactor(test): rename package to browser_test * refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets * chore(make): point test-browser at test/browser * docs(decisions): record internal/permissions deletion * refactor(doctor): use sidecarassets package * refactor(testrun): use sidecarassets package * fix(test): resolve testdata relative to browser_test.go * refactor(verifier): remove dead __sanderlingIndex compat alias * refactor(bundler): use encoding/json for JS string literals * docs(action-space): use vendor-neutral native driver wording * refactor(hierarchy): scrub backend tool name from comments * refactor(driver): scrub backend tool name from comments * refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver * refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap * refactor(chrome): implement DoubleTap as two taps with the gap * refactor(mock): record DoubleTap and DoubleTapSelector actions * refactor(runner): delegate double-tap to driver, drop gesture timing * test(runner): assert double-tap delegates to driver DoubleTap * docs(cmd): add package docs to CLI and developer tools * docs(driver): add package docs to driver interface and chrome backend * docs(driver): add package docs to mock and sidecar backends * docs(platform): add package docs to android and ios device prep * docs: add package docs to bundler and inspect * docs(ltl): add package doc to temporal logic evaluator * docs: add package docs to runner and testrun pipeline * docs: add package docs to trace and verifier * docs(sidecarassets): add package doc for embedded JAR loader * fix(chrome): add disable-dev-shm-usage so Chrome starts in CI * test(chrome): gate real-Chrome driver tests behind the browser tag * chore(make): run chrome driver tests in the browser job * fix(web-runtime): guard global error listeners for non-browser hosts The module registered window error/unhandledrejection listeners at top level, which threw under Node (the spec-api test runner) where globalThis.addEventListener is absent. Register only when the API exists; the real browser run is unaffected. * ci(browser): re-enable unprivileged user namespaces for headless Chrome ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged user namespaces stops headless Chrome from opening its DevTools socket even with --no-sandbox, surfacing as the driver's 'websocket url timeout'. Relax the sysctl for the job and add a direct launch check so a future breakage shows Chrome's own stderr rather than an opaque driver timeout. * ci(browser): pin stable Chrome for the driver tests setup-chrome's default latest pulled a dev Chromium (150) whose remote debugging socket never came up under chromedp, while plain --dump-dom worked. Pin the stable channel, which the driver is tested against. * feat(defaults): add scroll and rebalance action weights Use relative-integer weights (taps/typing co-primary 100, scrolls 50, swipes 25, doubleTaps 10); the picker normalizes by their total. Adds scrolls to defaultActions as a first-class reveal behavior. * feat(defaults): trim scroll action weight wiring * fix(build): point sidecar jar ignore and embed paths at sidecarassets * test(defaults): drop stale longPresses re-export assertion longPresses is opt-in vocabulary, no longer re-exported from defaults/actions.ts since e0d3b20; its builtin resolution is already covered by api.test.ts. Trim the defaults test to scrolls, which is an actual default export. * fix(chrome): raise DevTools websocket read timeout to 60s Chrome cold-start on a loaded CI runner can exceed chromedp's 20s default for reading the DevTools websocket URL, flaking the browser tests with "websocket url timeout reached". Give launch more headroom. |
||
|
|
88db9653e5 |
refactoring default action layer (#51)
* feat(hierarchy): add editable signal with native derivation
* feat(chrome): emit editable flag in hierarchy dump
* feat(verifier): expose editable on ax element objects
* feat(spec): add editable to selector and element types
* feat(verifier): register typing builtin generator
* feat(verifier): typing builtin types edge-case corpus into editable fields
* feat(spec): export typing builtin generator
* feat(spec): add defaultActions bundle
* feat(spec): export @sanderling/spec/defaults subpath
* feat(folio): layer defaultActions breadth over targeted flows
* test(verifier): typing builtin targets editable fields, declines otherwise
* test(hierarchy): editable derivation and selector matching
* test(spec): defaultActions, typing, and defaults barrel resolve
* fix(testrun): alias @sanderling/spec/defaults for the bundler
* test(chrome): editable flag for inputs, textarea, contenteditable
* feat(spec): typing builtin for the web (V8) action path
* chore(folio): auto-boot a bootable AVD in just test/install when none connected
* feat(driver): add ForegroundChecker optional capability
* feat(android): detect foreground package via adb dumpsys
* feat(sidecar): implement ForegroundApp via adb for android
* feat(runner): relaunch app when foreground escapes during exploration
* fix(spec): drop hardware back from defaultActions to stay in-app
* feat(spec): add DoubleTap action type and constructor
* feat(spec): wire DoubleTap through web-runtime serializer
* feat(verifier): bind doubleTap and decode DoubleTap actions
* feat(runner): dispatch DoubleTap as two taps inside one step
* test(doubleTap): cover constructor, verifier round-trip, and runner dispatch
* feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action
* fix(folio): track ledger row count across non-ledger steps; pin reproducer seed
* feat(spec): add doubleTaps random-target builtin to defaultActions
* feat(verifier): add doubleTaps random-target generator
* refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions
* fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives
* feat(verifier): track newly-violated property set per step
Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.
* refactor(runner): emit onset-only violations to trace and summary
Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.
* style(verifier): use maps.Copy for verdict snapshot
* fix(folio): make login spec content-driven (idempotent across re-entries)
* fix(verifier): canonicalize selector strings
Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.
* refactor(folio): replace txn invariants with balanceMatchesAddedTxn
Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.
* refactor(trace): drop WriteScreenshotAfter
Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.
* refactor(runner): one concurrent screenshot per step
Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.
* refactor(inspect-ui): use next step's screenshot for state after
Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.
* feat(sidecar): structural-hash settle poll
Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.
* test(sidecar): cover pollUntilStable and structuralHash
Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.
* feat(spec): accept optional name on extract()
Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.
* test(spec): cover extract name overload
Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.
* feat(verifier): name extractors for diff surfacing
bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.
* chore(folio): name every extract() call
Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.
* feat(verifier): track extractor value transitions
Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.
* test(verifier): cover ChangedExtractors diffs
Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.
* feat(trace): emit extractor_changes per step
Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.
* feat(inspect-ui): render extractor-change breadcrumbs at violations
Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.
* fix(sidecar): cap stability poll independently of settle budget
The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.
* feat(cli): default --clear-data on so runs start fresh
* feat(sidecar): streak-based settle with route-transition detection
Two changes layered into the stability poll:
1. stabilitySnapshot returns null while the tree carries more than one
route-level Screen tag (resource-id / testTag / identifier ending
in "Screen"), so the poll cannot declare a NavHost cross-fade
stable. Apps following the Compose route convention get this
detection for free; apps that don't fall through to the generic
signal below.
2. pollUntilStable now requires an uninterrupted stable streak of at
least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
matches. A late transition that fires after a brief calm window
breaks the streak instead of slipping past. Interval widened to
250ms so UiAutomation isn't hammered under fuzz load.
* test(sidecar): cover streak reset and route-transition rejection
Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.
* feat(runner): re-fetch on transitional hierarchy capture
Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.
fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.
* feat(runner): gate first action on app reaching foreground
* test(runner): cover startup foreground gate and back-press
* feat(verifier): scope random-action targets to app package
Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.
* feat(testrun): pass app package into verifier scope filter
* test(verifier): cover package-scoped target selection
* feat(hierarchy): derive package from resource-id prefix
The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.
* test(hierarchy): cover package derivation from resource-id
* chore: stop tracking inspect-ui/dist build artifacts
* feat(android): detect focused-window package via dumpsys window
* feat(driver): add FocusedWindowChecker capability
* fix(runner): gate first observe on the app window being drawn, not just resumed
* test(mock): add FocusedWindowApp with foreground mirroring
* test(runner): cover startup gate waiting for app window to draw
* feat(proto): add Snapshot RPC for atomic hierarchy+screenshot
Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.
* feat(sidecar): add snapshot default on DriverBackend
Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.
* feat(sidecar): wire Snapshot handler with serialization lock
Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.
* test(sidecar): cover Snapshot wire path and serialization lock
SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.
* feat(driver): expose Snapshot on DeviceDriver and sidecar client
Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.
* feat(driver): add Snapshot to chrome and mock drivers
The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.
* refactor(runner): observe each step via the atomic Snapshot RPC
fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.
* test(runner): assert step uses Snapshot, not raw hierarchy/screenshot
TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.
* test(driver): cover Snapshot in proto descriptor and sidecar client
Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.
* feat(trace): add Transitional flag to Step
* fix(runner): skip verifier for transitional trees after retry budget
When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.
* test(runner): cover transitional step skips verifier and clean control
* refactor(trace): rename Step.Action to Step.NextAction
The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.
* refactor(runner): assign trace action to Step.NextAction field
Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.
* refactor(inspect): decode trace step's next_action JSON field
Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.
* test(inspect): update fixtures to use next_action trace field
Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.
* refactor(inspect-ui): rename Step.action to Step.next_action
Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.
* fix(folio): extract balanceMatchesAddedSum predicate as testable helper
Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.
* fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn
The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.
* test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases
Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.
* fix(build): rebuild sidecar JAR when Kotlin sources change
Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.
* fix(chrome): launch with no-sandbox so headless Chrome starts in CI
* fix(sidecar): type text at cursor instead of clearing the field
InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.
* test(sidecar): assert InputText types at cursor without clearing
Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.
* feat(proto): add LongPress RPC
* chore(proto): regenerate Go stubs for LongPress
* feat(driver): add LongPress to DeviceDriver interface
* feat(sidecar): add LongPress client method
* feat(mock): record LongPress action
* feat(chrome): implement LongPress as press-and-hold
* feat(sidecar): implement longPress across backends
* feat(sidecar): dispatch LongPress RPC to backend
* test(sidecar): cover LongPress dispatch
* test(sidecar): implement longPress in snapshot test backend
* feat(verifier): add LongPress and Scroll action kinds
* feat(folio-spec): predicate that gates balance check on TxnSubmit tap
Replaces the row-sum predicate (which always held by construction since
balance is derived from rows in Folio) with one that compares the typed
amount to the actual balance delta after a tap on TxnSubmit. Catches the
planted double-submit bug.
* feat(folio-spec): wire submitMovesBalanceByTypedAmount property
Adds lastAction and totalBalance extractors and uses them in the new
property. Drops ledgerRows/ledgerBalance extractors since nothing else
referenced them.
* feat(verifier): wire longPresses and scrolls generators
* test(verifier): cover longPresses and scrolls generators
* test(folio-spec): unit tests for submitChangesBalanceByTypedAmount
Covers single vs double submit, the DoubleTap variant, vacuous cases
(null action, wrong kind, wrong target, zero typed), and selector-as-
object coercion.
* feat(spec): add LongPress and Scroll authoring surface
* feat(spec): no-op LongPress and Scroll in web runtime
* feat(spec): re-export longPresses and scrolls as opt-in generators
* test(spec): cover LongPress and Scroll runtime members
* test(proto): expect LongPress in service descriptor
* feat(runner): dispatch LongPress and Scroll actions
* test(runner): cover LongPress and Scroll dispatch
* docs(action-space): move LongPress, Scroll, DoubleTap to current actions
* fix(runner): mark nil/empty hierarchy as transitional
A failed or empty sidecar hierarchy fetch was pushed straight to the
verifier, letting spec extractors crash with "Cannot read property 'map'
of undefined" when findAll returned null. Treat that case like a
transitional capture: skip the verifier push, still record the step, and
keep the loop progressing.
* fix(verifier): populate Action.On when tap chooser picks an element
Coordinate-targeted Taps/DoubleTaps left On empty, so action-gated
properties reading lastAction.on couldn't tell which target was hit and
were vacuously skipped. Resolve the picked element to a stable
key:value selector (resource-id, testTag, text, desc) and validate it
resolves back to the same element so we don't accidentally redirect the
tap to a sibling that shares the identifier.
* fix(folio): add parseTypedAmount helper matching app's parseCents
Raw user input like "50" must become 5000 cents, not 50. The existing
parseDollarCents helper strips non-digits and so reads "50" as 50 cents,
which is correct for formatted balance text but off by 100x for raw
input from the amount field.
* fix(folio): parse raw amount input as cents in submit predicate
txnAmountField holds raw user keystrokes, not formatted balance text.
Route it through parseTypedAmount so "50" reads as $50, matching how
the app commits the transaction.
* fix(folio): carry forward total balance across off-screen transitions
AddTransactionScreen shows neither AccountCard nor LedgerBalance, so the
extractor used to report 0 at the step before submit. That made every
non-zero current balance look like the full delta and tripped the typed
amount property on every honest submit. Remember the last-seen sum and
return it whenever the current snapshot has no balance signal.
* test(folio): cover submit predicate with raw typed-amount inputs
Pipes realistic raw keystrokes through parseTypedAmount + the predicate
so single submits clear and double submits fire as expected.
* feat(folio): add computeHomeTotalBalance helper
Pure helper that tracks Home multi-account total only and carries the last
Home sum across off-Home steps. Ledger's single-account balance is excluded
because mixing it would corrupt cross-screen scale comparisons.
* fix(folio): totalBalance carrier tracks only Home, not Ledger
Home cardSum is a multi-account total; Ledger's LedgerBalance is a single
account on a different scale. Blending them in the carrier produced bogus
cross-screen deltas (prev from Ledger, curr from Home), triggering false
positives in submitMovesBalanceByTypedAmount. Restrict the carrier to
Home AccountCard totals via the computeHomeTotalBalance helper.
* test(spec): cover computeHomeTotalBalance carrier behaviour
Tests Home sums, carrier passthrough on off-Home steps, the Ledger
scale-mismatch case, and a Home > off-Home > Home sequence.
* feat(runner): treat transient apply errors as transitional steps
Sidecar input RPCs occasionally hang with DEADLINE_EXCEEDED or
UNAVAILABLE on long fuzzing runs. The per-step loop previously
propagated any applyAction error and killed the run after a single
flake. Detect transient gRPC failures via status.FromError, mark the
step transitional, skip the post-action idle poll, and continue to the
next step. Fatal errors (outer ctx cancellation, non-transient codes,
verifier crashes) still propagate.
* test(runner): cover transient apply error resilience
TestRunner_TransientApplyErrorMarksTransitional drives the runner
through a wrapper that fails the first TapSelector with a gRPC
DeadlineExceeded then succeeds. Asserts the run does not exit, the
failed step is marked transitional with no violations, and the next
step runs cleanly. TestIsTransientApplyError_Classification covers the
helper's matching rules directly so future code changes don't quietly
drop a transient case.
* fix(folio): gate submit-balance property on Home route landing
totalBalance is only freshly computed when AccountCards are visible on
Home; off-Home landings return the carrier and would false-fire the
property, latching always(next(F)) to false and masking the real
double-submit bug. Skip vacuously when route is not "home".
* test(spec): cover route gate in submit-balance predicate
Adds route arg to existing cases (all use "home") and adds five new
cases: ledger landing with stale carrier, add-transaction with
double-insert delta, null route, plus home-landing positive and
double-insert negative cases anchoring the gate's allow path.
|
||
|
|
b23fb0c723 |
feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag
Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.
* feat(testrun): add Preflight() before sidecar/driver setup
Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.
* refactor(chrome): split tag (HTML name) from class (CSS classList)
Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.
* feat(chrome): translate legacy string selectors to CSS/XPath
TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.
* feat(trace): add WriteHTML + Step.HTMLAvailable
Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.
* feat(driver): add WebDriver capability + chrome implementation
WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.
* feat(verifier): OverrideExtractorValues for V8-driven extractors
Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.
* feat(spec): add WebState + camelCase attribute aliases
WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.
* feat(runner): per-tick HTML capture for WebDriver-capable drivers
Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.
* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>
Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.
* feat(bundler): BundleWeb + V8-side runtime shim
web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.
* feat(runner): V8 extractor overrides + V8 action source for WebDriver
When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.
* feat(testrun): bundle + install web runtime when platform=web
BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.
* feat(inspect-ui): hierarchy + html panels in run detail
HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.
* fix(folio-web): drop aria-label data-carrier abuse
Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.
* chore: rebuild inspect-ui dist + folio-web .gitignore
Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.
* revert(trace): drop WriteHTML + Step.HTMLAvailable
Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.
* revert(runner): drop per-tick HTML capture
Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.
* revert(driver): drop WebDriver.Document
Document was only consumed by the runner's HTML capture which is gone.
* revert(inspect): drop /html route
Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.
* revert(inspect-ui): drop htmlUrl + html_available type
API surface no longer needs the HTML route; Step.html_available has no
producer.
* revert(inspect-ui): drop HtmlPanel + html tab
Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.
* test(inspect-ui): drop htmlUrl test, add @types/bun
Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.
* chore: rebuild inspect-ui dist without HtmlPanel
Embedded SPA bundle no longer ships the iframe HTML viewer.
* fix(web-runtime): retry action resolution + implement taps/swipes
V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.
* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys
Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.
* chore(folio-web): drop swipes from action root
Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.
* fix(inspect-ui): correct HierarchyPanel CSS variable names
Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.
* fix(chrome): correct PressKey mappings to chromedp/kb constants
Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.
Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.
* fix(cli): -h/--help exits 0 instead of error code
parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.
Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.
* fix(chrome): harden cssEscape for control chars + use [class~=]
Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.
Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.
* fix(web-runtime): use CSS.escape and validate tag-name selectors
The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.
The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.
Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.
* fix(chrome): validate attribute name in unknown-prefix branch
A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.
* fix(selectors): emit valid XPath 1.0 string literals via concat()
Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.
Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.
* fix(runtime): surface unresolved action targets instead of dropping silently
serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.
Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.
* fix(runner): use errgroup-bound ctx so siblings cancel on failure
The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.
* fix(chrome): propagate caller ctx cancellation to CDP calls
InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.
Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.
* fix(verifier): tolerate out-of-range override indices
A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.
Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.
* test(verifier): cover object-shaped extractor overrides
Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.
* fix(web-runtime): lock global runtime hooks against page shadowing
AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.
* perf(web-runtime): cache randomTap candidate DOM scan per tick
The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.
Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.
* fix(web-runtime): cap sanitize recursion to prevent stack overflow
State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.
* fix(web-runtime): enforce pressKey allowlist in factory
The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.
* chore(chrome): drop dead bundleSource/bundleMu
bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.
* fix(chrome): use strconv.Atoi for extractor key parsing
fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.
* fix(doctor): raise per-check timeout to 15s for chromium launch
5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.
* fix(runner): trust V8 coordinates for InputText, even at origin
resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.
Add applyAction tests covering both the typical web case and the (0,0)
edge case.
* test(bundler): lock down deterministic output across builds
The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
|