Commit Graph
3 Commits
Author SHA1 Message Date
pj 808d607eac fix(sidecarassets): publish the extracted jar through a rename
Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 21:38:09 +05:30
pj 94d9511312 test: full test-suite refactor sweep (#61)
* chore(test): start test-suite refactor sweep

* test(ltl): pin exact multi-obligation residual AST

* test(ltl): table-test finalize Kleene connective combinations

* test(ltl): pin reduce over pending inner for bound, Or, Not

* test(ltl): marshal bounded Always steps/duration/deadline

* test(verifier): cover LTL combinator verdict transitions and within unit panic

* test(verifier): table-test DecodeAction kinds and lastAction field exposure

* test(verifier): assert WithPlatform(ios) reaches the picker host and key pool

* test(verifier): widen weighted-selection assertion to a 5x skew margin

* test(verifier): un-skip ax-find round trip with a committed tree fixture

* test(runner): pin isWDADrop to sidecar reconnect-failed message origin

* test(runner): assert PressKey/Wait trace encoding records kind-specific fields

* test(runner): cover RenderSummary unsupported-verbs surfacing branch

* test(trace): set Hierarchy in round-trip and lock lossy Tree contract

Also add a -race concurrent WriteStep test that asserts N well-formed JSONL lines, catching torn lines if the writer mutex is dropped.

* test(trace): round-trip witnesses/changes/metrics/exceptions, pin step-0 witness

* test(trace): document ViolationsAreGreppable grep contract and lock-free WriteScreenshot

* test(hierarchy): cover invalid-JSON and malformed-bounds parser paths

* test(trace): guard writer mutex via WriteStep/Close race on w.file

* test(replay): drop unfailable assets and devproxy assertions

* test(replay): cache reuses on equal mtime, reparses after append

* test(replay): violation marker falls back to detection step when attributed missing

* test(replay): corrupt meta/trace dirs return 500 with error body

* test(replay): SSE client receives runs.changed after a broadcast

* test(replay): Run coalesces creates, ignores write/chmod, closes subs on cancel

* fix(sidecar): synchronize health fixture writes and exercise healthError

* test(sidecar): cover swipe/longpress/doubletap/erase/presskey/metrics/logs translations

* test(sidecar): cover DoubleTapSelector composition and mid-gesture cancel

* test(sidecar): assert gRPC error status surfaces from action RPC

* fix(chrome): route action methods through runCtx so caller cancellation aborts CDP

* fix(chrome): route hierarchy/screenshot/waitidle/metrics through runCtx

* refactor(ios): extract pure simctl JSON parsers

* refactor(ios): add command-runner seams for EnsureSimulator

* test(ios): table-test simctl parsers and EnsureSimulator seams

* test(sidecarassets): cover placeholder build path

* test(sidecarassets): assert reuse via sentinel bytes not mtime

* test(bundler): cover properties-only spec registration

* refactor(testrun): extract prepareBundleInputs from Execute

* test(testrun): cover prepareBundleInputs aliases and missing-runtime error

* test(testrun): table-test resolveRuntimeSibling search edges

* test(testrun): exact-output tests for progressHandler line format

* fix(cmd): point bundle-check aliases at pkg/spec/src

* test(cmd): smoke-test bundle-check resolves spec aliases

* test(cmd): table-test hier-check parse and FindAll on fixture

* test(cmd): unit-test buildBrowseURL deep-link vs root

* test(cmd): drop flaky TestRun_Doctor that launched real Chromium

* test(cmd): pin pipeline error to bundle resolution on web platform

* test(replay-ui): add bun test script

* ci(replay-ui): run bun test via make web-test target

* ci(replay-ui): point bun cache key at replay-ui/bun.lock

* test(replay-ui): exercise real URL encoding and non-ok throw in getJson

* refactor(replay-ui): extract snapshot flatten/getAtPath into lib module

* test(replay-ui): pin snapshot flatten/getAtPath path round-trip

* refactor(replay-ui): extract action selector/format into lib module

* test(replay-ui): pin action selector parse and row formatting

* refactor(replay-ui): share one statusFor between panels

* refactor(replay-ui): extract run-history derivation into lib module

* test(replay-ui): pin shared statusFor precedence and ordering

* test(replay-ui): pin run-history derivation alignment

* refactor(replay-ui): export clampIndex for testing

* refactor(replay-ui): extract keyboard-nav dispatch into pure module

* refactor(replay-ui): extract metrics formatters into lib module

* test(replay-ui): pin clampIndex step boundaries

* test(replay-ui): pin keyboard-nav ownership and key routing

* test(replay-ui): pin metrics formatters and path gap handling

* refactor(sidecar): expose device-output parsers as internal for testing

* test(sidecar): table-test device-output parsers against malformed input

* test(sidecar): cover logcat parsing year inference and line skipping

* test(sidecar): pin pressKey keycode mapping and unknown-key rejection

* test(sidecar): metrics bundleId falls back to launched app and honors override

* test(sidecar): loosen deadline upper bound to tolerate slow CI scheduling

* test(web-runtime): export selector builders for unit tests

* test(web-runtime): guard sanitize cycle, function, and depth limits

* test(web-runtime): table-test selector builder quoting and escaping

* test(sidecar): collapse scalar-forwarding RPC tests into a table

* test(replay-ui): dedup step/summary fixtures into shared module

* test(ios): collapse pickSimulator point-tests into a table
2026-06-06 13:59:08 +05:30
pj c5bb176be8 UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks

Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.

* feat(ltl): negation normal form pass

nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.

* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse

Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.

* test(ltl): property-based NNF laws

Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.

* test(ltl): Finalize, bounded eventually, latch, collapse

Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.

* feat(inspect): within clause on always residual node

A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.

* feat(ltl): witness violations and (bool,error) predicate thunks

* test(ltl): migrate thunk call sites to (bool,error)

* feat(ltl): flag thrown-predicate witnesses with IsError

* refactor(verifier): replace predicate err side-channel with violation witness

* test(verifier): witness API for thrown predicates

* feat(trace): witnesses map and skipped-verification marker on Step

* feat(runner): thread violation witnesses, finalize, skip marker into trace

* test(ltl): lock violation witness reason, IsError, and step

* test(verifier): finalize surfaces unmet eventually with witness

* fix(ltl): eliminate implies and bounded-always false-negatives

Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.

* test(ltl): lock implies and bounded-always false-negative regressions

* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick

* feat(testrun): inject seed into web bundle via SANDERLING_SEED define

* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring

* test(spec): add Go math/rand/v2 PCG oracle and golden fixture

* feat(spec): bit-exact PCG port of Go math/rand/v2

* test(spec): assert pcg.ts matches the PCG golden fixture

* feat(spec): shared input corpus and press-key pools

* feat(spec): action-tree types and Host interface

* feat(spec): verb support matrix and warn-once helper

* feat(spec): deterministic shared action picker

* test(spec): verb matrix and warn-once semantics

* test(spec): picker draw-order and determinism

* refactor(spec): actions.ts returns pure GeneratorNode data trees

* refactor(spec): wire from() sampling through the picker rng

* feat(spec): shared runtime-entry installs next-action over pick.ts

* feat(spec): export LongPress/Scroll/longPresses/scrolls factories

* test(spec): assert data-tree shapes for action factories

* test(spec): runtime-entry serializeAction wire-contract round-trip

* refactor(spec): bridge data-tree nodes to the legacy goja picker tags

* fix(spec): web runtime walks the spec's globalThis.actions data tree

* test(spec): tolerate legacy bridge fields on builtin nodes

* refactor(spec): installRuntime accepts a lazy root resolver

The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.

* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker

Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).

* test(spec): cover the WEB Host surface and seed precision

Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.

* refactor(spec): picker emits native selector + scroll endpoints, setup precedence

* feat(spec): goja runtime entry wires the shared picker over the Go host

* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin

* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker

* refactor(spec): drop the legacy goja bridge fields from action factories

* feat(spec): serialize selector-only string targets for the runner to re-resolve

* refactor(verifier): one DecodeAction reads the unified flat wire contract

* refactor(verifier): goja host + shared picker replace the duplicate Go picker

* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime

* test(verifier): author specs through the shared picker path

* test(runner): bundle authored specs with the goja runtime entry

* feat(verifier): collect unsupported verbs for the run report

* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource

* feat(testrun): surface unsupported verbs in run report

* test(verifier): cross-runtime goja/node parity gate on the shared picker

* test(verifier): unsupported verbs collected deduped in first-seen order

* test(runner): summary reports no unsupported verbs on a clean run

* test(spec): golden-fixture cross-runtime parity gate for the node picker

Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.

* test(verifier): assert goja picker against the same cross-runtime golden

Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.

* refactor(spec): rename pressKey generator export to pressKeys

* refactor(spec): update barrel re-exports for pressKeys

* test(spec): update pressKeys generator export name

* docs(spec): rename pressKey generator to pressKeys

* refactor(spec): extract samplerRng into shared sampler-rng module

* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)

* test(spec): cover fluent value generators determinism and chaining

* refactor(bundler): inject globalThis trailer from spec named exports

* refactor(bundler): reuse registration trailer in web bundler

* test(bundler): cover named-export globalThis registration

* feat(spec): add named() to Extracted handle type

* feat(web-runtime): named() and cross-extractor read guard

* feat(verifier): named() and cross-extractor read guard in goja

* test(verifier): cross-extractor read guard and named()

* test(web-runtime): export runtime and extractors for tests

* test(web-runtime): named() and cross-extractor read guard

* refactor(folio): drop manual globalThis trailer (bundler injects it)

* refactor(folio): seed txn amounts via integers().between(1,500)

* refactor(folio-web): drop manual globalThis trailer (bundler injects it)

* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs

* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts

* refactor(folio-web): name extractors so violation witnesses are readable

* fix(web-runtime): propagate extractor getter throws and unpoison locked global

Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.

* test(spec): install fake runtime via defineProperty to survive locked global

* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors

* feat(runner): add MaxSteps bound to Options

* test(runner): MaxSteps stops after exactly N steps

* test(driverpb): drop proto getter round-trip tautology

* test(sidecar): drop stub-mode placeholder tautology tests

* test(mock): drop default-field-value assertion test

* test(ltl): drop Verdict.String tautology tests

* refactor(runner): extract RenderSummary for snapshot testing

* test(runner): golden snapshots for trace stream and violation summary

* feat(web-runtime): capture uncaught errors into state.exceptions

* test(integration): add throwing and counter web fixtures

* test(integration): add specs for the web fixtures

* test(integration): drive web fixtures through the real pipeline in headless Chrome

* chore(make): add test-browser target for the Chrome-driven suite

* ci: run the Chrome-driven browser suite in a separate job

* refactor(test): relocate browser suite to test/browser

* refactor(permissions): delete dead internal/permissions package

* refactor(test): rename package to browser_test

* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets

* chore(make): point test-browser at test/browser

* docs(decisions): record internal/permissions deletion

* refactor(doctor): use sidecarassets package

* refactor(testrun): use sidecarassets package

* fix(test): resolve testdata relative to browser_test.go

* refactor(verifier): remove dead __sanderlingIndex compat alias

* refactor(bundler): use encoding/json for JS string literals

* docs(action-space): use vendor-neutral native driver wording

* refactor(hierarchy): scrub backend tool name from comments

* refactor(driver): scrub backend tool name from comments

* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver

* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap

* refactor(chrome): implement DoubleTap as two taps with the gap

* refactor(mock): record DoubleTap and DoubleTapSelector actions

* refactor(runner): delegate double-tap to driver, drop gesture timing

* test(runner): assert double-tap delegates to driver DoubleTap

* docs(cmd): add package docs to CLI and developer tools

* docs(driver): add package docs to driver interface and chrome backend

* docs(driver): add package docs to mock and sidecar backends

* docs(platform): add package docs to android and ios device prep

* docs: add package docs to bundler and inspect

* docs(ltl): add package doc to temporal logic evaluator

* docs: add package docs to runner and testrun pipeline

* docs: add package docs to trace and verifier

* docs(sidecarassets): add package doc for embedded JAR loader

* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI

* test(chrome): gate real-Chrome driver tests behind the browser tag

* chore(make): run chrome driver tests in the browser job

* fix(web-runtime): guard global error listeners for non-browser hosts

The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.

* ci(browser): re-enable unprivileged user namespaces for headless Chrome

ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.

* ci(browser): pin stable Chrome for the driver tests

setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.

* feat(defaults): add scroll and rebalance action weights

Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.

* feat(defaults): trim scroll action weight wiring

* fix(build): point sidecar jar ignore and embed paths at sidecarassets

* test(defaults): drop stale longPresses re-export assertion

longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.

* fix(chrome): raise DevTools websocket read timeout to 60s

Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.
2026-06-02 09:52:53 +05:30