Commit Graph
32 Commits
Author SHA1 Message Date
pj cb7936fe36 test(confusion-matrix): reject malformed inputs and keep missing data out of the cells 2026-08-17 23:27:51 +05:30
pj 56a002d87a feat(confusion-matrix): score the checker against a blind reviewer
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.
2026-08-17 23:27:47 +05:30
pj 2aba29a6e3 test(bundle-check): cover the zero-property refusal and pin the reported bundle 2026-08-17 23:27:29 +05:30
pj 26455d5ec4 feat(bundle-check): fail a spec that bundles but registers no properties 2026-08-17 23:27:29 +05:30
pj ddf17129bd feat(campaign): count the runs that were never in the app
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
2026-08-17 14:35:37 +05:30
pj 341d6a0614 feat(corpus-sweep): run one specification against a served corpus of implementations
same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.
2026-08-16 17:45:42 +05:30
pj 146152a3af feat(implementation-sweep): run one campaign against every implementation of a requirement
installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.
2026-08-16 17:45:41 +05:30
pj a45ba76d8e feat(oracle-reduction): replay stored traces under four reduced oracles
re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.
2026-08-16 17:45:41 +05:30
pj a0a8c9c710 feat(defect-identity): count distinct defects across stored runs
a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.
2026-08-16 17:45:35 +05:30
pj 11fca22d36 feat(exploration-reach): count the distinct structural states a stored run visited
the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.
2026-08-16 17:45:35 +05:30
pj 0d1b2cf4e8 feat(label-coverage): report the addressable share of an app's interactive surface
reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.
2026-08-16 17:45:35 +05:30
pj b47509a8e8 test(analyze): recover planted effects through the tool's own entry point
a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.
2026-08-16 17:45:26 +05:30
pj 4dff70b18b feat(analyze): add the seed-paired signed-rank comparison and record the holm family
--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.
2026-08-16 17:45:26 +05:30
pj b99da0be0e feat(analyze): time an event at the step it was detected and report the quartiles
an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.
2026-08-16 17:45:20 +05:30
pj a28178337c refactor(seedspec): move seed spec parsing out of the campaign command
the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.
2026-08-16 17:44:57 +05:30
pj ebed84afc3 feat(campaign): record the label source in the manifest
A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:32:05 +05:30
pj ff2de344e9 feat(campaign): make the label source a cell dimension
A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:32:05 +05:30
pj 0cc20539bc test(campaign): wait for the trap instead of racing it
The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:19 +05:30
pj eacf3fd11f fix(analyze): divide per-hour rates by time actually worked
A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:43:12 +05:30
pj c63a2e4897 fix(campaign): record both clocks a run was measured on
Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:43:12 +05:30
pj dd17f831ef fix(campaign): signal a timed-out run so it reaps its sidecar
CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:27:13 +05:30
pj 2cbe03c3fb fix(analyze): divide by actions that ran
Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:43:19 +05:30
pj f6d562e3dc feat(campaign): count dispatched actions, not steps
A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:43:19 +05:30
pj 019d608f65 feat(analyze): survival analysis over campaign directories
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:03:03 +05:30
pj 76dce1a75e experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs

runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record arm membership and host in meta.json

meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --arm and populate run meta from it

Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): sweep seeds for one experiment cell

campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.

Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): no silent generator fallback, and llm on web

--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.

pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): make the hierarchy dump agree with the web runtime

Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.

scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.

The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(conformance): give the gate reproducible seeds

SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.

The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit editable as a plain boolean

editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): leave the head subtree out of the web target walk

collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(chrome): compare the facts both hosts derive from one DOM

The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.

This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* chore(make): run the browser packages one at a time

Both launch Chrome and launching two at once has failed with "Launch: context
canceled".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style: remove every em-dash and en-dash

Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): honor the caller context in Launch

Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.

The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(sidecarassets): publish the extracted jar through a rename

Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): kill a run that outlives --run-timeout

A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style(test): gofmt browser_test.go

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:20:31 +05:30
pj 94d9511312 test: full test-suite refactor sweep (#61)
* chore(test): start test-suite refactor sweep

* test(ltl): pin exact multi-obligation residual AST

* test(ltl): table-test finalize Kleene connective combinations

* test(ltl): pin reduce over pending inner for bound, Or, Not

* test(ltl): marshal bounded Always steps/duration/deadline

* test(verifier): cover LTL combinator verdict transitions and within unit panic

* test(verifier): table-test DecodeAction kinds and lastAction field exposure

* test(verifier): assert WithPlatform(ios) reaches the picker host and key pool

* test(verifier): widen weighted-selection assertion to a 5x skew margin

* test(verifier): un-skip ax-find round trip with a committed tree fixture

* test(runner): pin isWDADrop to sidecar reconnect-failed message origin

* test(runner): assert PressKey/Wait trace encoding records kind-specific fields

* test(runner): cover RenderSummary unsupported-verbs surfacing branch

* test(trace): set Hierarchy in round-trip and lock lossy Tree contract

Also add a -race concurrent WriteStep test that asserts N well-formed JSONL lines, catching torn lines if the writer mutex is dropped.

* test(trace): round-trip witnesses/changes/metrics/exceptions, pin step-0 witness

* test(trace): document ViolationsAreGreppable grep contract and lock-free WriteScreenshot

* test(hierarchy): cover invalid-JSON and malformed-bounds parser paths

* test(trace): guard writer mutex via WriteStep/Close race on w.file

* test(replay): drop unfailable assets and devproxy assertions

* test(replay): cache reuses on equal mtime, reparses after append

* test(replay): violation marker falls back to detection step when attributed missing

* test(replay): corrupt meta/trace dirs return 500 with error body

* test(replay): SSE client receives runs.changed after a broadcast

* test(replay): Run coalesces creates, ignores write/chmod, closes subs on cancel

* fix(sidecar): synchronize health fixture writes and exercise healthError

* test(sidecar): cover swipe/longpress/doubletap/erase/presskey/metrics/logs translations

* test(sidecar): cover DoubleTapSelector composition and mid-gesture cancel

* test(sidecar): assert gRPC error status surfaces from action RPC

* fix(chrome): route action methods through runCtx so caller cancellation aborts CDP

* fix(chrome): route hierarchy/screenshot/waitidle/metrics through runCtx

* refactor(ios): extract pure simctl JSON parsers

* refactor(ios): add command-runner seams for EnsureSimulator

* test(ios): table-test simctl parsers and EnsureSimulator seams

* test(sidecarassets): cover placeholder build path

* test(sidecarassets): assert reuse via sentinel bytes not mtime

* test(bundler): cover properties-only spec registration

* refactor(testrun): extract prepareBundleInputs from Execute

* test(testrun): cover prepareBundleInputs aliases and missing-runtime error

* test(testrun): table-test resolveRuntimeSibling search edges

* test(testrun): exact-output tests for progressHandler line format

* fix(cmd): point bundle-check aliases at pkg/spec/src

* test(cmd): smoke-test bundle-check resolves spec aliases

* test(cmd): table-test hier-check parse and FindAll on fixture

* test(cmd): unit-test buildBrowseURL deep-link vs root

* test(cmd): drop flaky TestRun_Doctor that launched real Chromium

* test(cmd): pin pipeline error to bundle resolution on web platform

* test(replay-ui): add bun test script

* ci(replay-ui): run bun test via make web-test target

* ci(replay-ui): point bun cache key at replay-ui/bun.lock

* test(replay-ui): exercise real URL encoding and non-ok throw in getJson

* refactor(replay-ui): extract snapshot flatten/getAtPath into lib module

* test(replay-ui): pin snapshot flatten/getAtPath path round-trip

* refactor(replay-ui): extract action selector/format into lib module

* test(replay-ui): pin action selector parse and row formatting

* refactor(replay-ui): share one statusFor between panels

* refactor(replay-ui): extract run-history derivation into lib module

* test(replay-ui): pin shared statusFor precedence and ordering

* test(replay-ui): pin run-history derivation alignment

* refactor(replay-ui): export clampIndex for testing

* refactor(replay-ui): extract keyboard-nav dispatch into pure module

* refactor(replay-ui): extract metrics formatters into lib module

* test(replay-ui): pin clampIndex step boundaries

* test(replay-ui): pin keyboard-nav ownership and key routing

* test(replay-ui): pin metrics formatters and path gap handling

* refactor(sidecar): expose device-output parsers as internal for testing

* test(sidecar): table-test device-output parsers against malformed input

* test(sidecar): cover logcat parsing year inference and line skipping

* test(sidecar): pin pressKey keycode mapping and unknown-key rejection

* test(sidecar): metrics bundleId falls back to launched app and honors override

* test(sidecar): loosen deadline upper bound to tolerate slow CI scheduling

* test(web-runtime): export selector builders for unit tests

* test(web-runtime): guard sanitize cycle, function, and depth limits

* test(web-runtime): table-test selector builder quoting and escaping

* test(sidecar): collapse scalar-forwarding RPC tests into a table

* test(replay-ui): dedup step/summary fixtures into shared module

* test(ios): collapse pickSimulator point-tests into a table
2026-06-06 13:59:08 +05:30
pj c5bb176be8 UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks

Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.

* feat(ltl): negation normal form pass

nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.

* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse

Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.

* test(ltl): property-based NNF laws

Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.

* test(ltl): Finalize, bounded eventually, latch, collapse

Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.

* feat(inspect): within clause on always residual node

A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.

* feat(ltl): witness violations and (bool,error) predicate thunks

* test(ltl): migrate thunk call sites to (bool,error)

* feat(ltl): flag thrown-predicate witnesses with IsError

* refactor(verifier): replace predicate err side-channel with violation witness

* test(verifier): witness API for thrown predicates

* feat(trace): witnesses map and skipped-verification marker on Step

* feat(runner): thread violation witnesses, finalize, skip marker into trace

* test(ltl): lock violation witness reason, IsError, and step

* test(verifier): finalize surfaces unmet eventually with witness

* fix(ltl): eliminate implies and bounded-always false-negatives

Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.

* test(ltl): lock implies and bounded-always false-negative regressions

* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick

* feat(testrun): inject seed into web bundle via SANDERLING_SEED define

* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring

* test(spec): add Go math/rand/v2 PCG oracle and golden fixture

* feat(spec): bit-exact PCG port of Go math/rand/v2

* test(spec): assert pcg.ts matches the PCG golden fixture

* feat(spec): shared input corpus and press-key pools

* feat(spec): action-tree types and Host interface

* feat(spec): verb support matrix and warn-once helper

* feat(spec): deterministic shared action picker

* test(spec): verb matrix and warn-once semantics

* test(spec): picker draw-order and determinism

* refactor(spec): actions.ts returns pure GeneratorNode data trees

* refactor(spec): wire from() sampling through the picker rng

* feat(spec): shared runtime-entry installs next-action over pick.ts

* feat(spec): export LongPress/Scroll/longPresses/scrolls factories

* test(spec): assert data-tree shapes for action factories

* test(spec): runtime-entry serializeAction wire-contract round-trip

* refactor(spec): bridge data-tree nodes to the legacy goja picker tags

* fix(spec): web runtime walks the spec's globalThis.actions data tree

* test(spec): tolerate legacy bridge fields on builtin nodes

* refactor(spec): installRuntime accepts a lazy root resolver

The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.

* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker

Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).

* test(spec): cover the WEB Host surface and seed precision

Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.

* refactor(spec): picker emits native selector + scroll endpoints, setup precedence

* feat(spec): goja runtime entry wires the shared picker over the Go host

* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin

* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker

* refactor(spec): drop the legacy goja bridge fields from action factories

* feat(spec): serialize selector-only string targets for the runner to re-resolve

* refactor(verifier): one DecodeAction reads the unified flat wire contract

* refactor(verifier): goja host + shared picker replace the duplicate Go picker

* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime

* test(verifier): author specs through the shared picker path

* test(runner): bundle authored specs with the goja runtime entry

* feat(verifier): collect unsupported verbs for the run report

* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource

* feat(testrun): surface unsupported verbs in run report

* test(verifier): cross-runtime goja/node parity gate on the shared picker

* test(verifier): unsupported verbs collected deduped in first-seen order

* test(runner): summary reports no unsupported verbs on a clean run

* test(spec): golden-fixture cross-runtime parity gate for the node picker

Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.

* test(verifier): assert goja picker against the same cross-runtime golden

Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.

* refactor(spec): rename pressKey generator export to pressKeys

* refactor(spec): update barrel re-exports for pressKeys

* test(spec): update pressKeys generator export name

* docs(spec): rename pressKey generator to pressKeys

* refactor(spec): extract samplerRng into shared sampler-rng module

* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)

* test(spec): cover fluent value generators determinism and chaining

* refactor(bundler): inject globalThis trailer from spec named exports

* refactor(bundler): reuse registration trailer in web bundler

* test(bundler): cover named-export globalThis registration

* feat(spec): add named() to Extracted handle type

* feat(web-runtime): named() and cross-extractor read guard

* feat(verifier): named() and cross-extractor read guard in goja

* test(verifier): cross-extractor read guard and named()

* test(web-runtime): export runtime and extractors for tests

* test(web-runtime): named() and cross-extractor read guard

* refactor(folio): drop manual globalThis trailer (bundler injects it)

* refactor(folio): seed txn amounts via integers().between(1,500)

* refactor(folio-web): drop manual globalThis trailer (bundler injects it)

* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs

* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts

* refactor(folio-web): name extractors so violation witnesses are readable

* fix(web-runtime): propagate extractor getter throws and unpoison locked global

Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.

* test(spec): install fake runtime via defineProperty to survive locked global

* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors

* feat(runner): add MaxSteps bound to Options

* test(runner): MaxSteps stops after exactly N steps

* test(driverpb): drop proto getter round-trip tautology

* test(sidecar): drop stub-mode placeholder tautology tests

* test(mock): drop default-field-value assertion test

* test(ltl): drop Verdict.String tautology tests

* refactor(runner): extract RenderSummary for snapshot testing

* test(runner): golden snapshots for trace stream and violation summary

* feat(web-runtime): capture uncaught errors into state.exceptions

* test(integration): add throwing and counter web fixtures

* test(integration): add specs for the web fixtures

* test(integration): drive web fixtures through the real pipeline in headless Chrome

* chore(make): add test-browser target for the Chrome-driven suite

* ci: run the Chrome-driven browser suite in a separate job

* refactor(test): relocate browser suite to test/browser

* refactor(permissions): delete dead internal/permissions package

* refactor(test): rename package to browser_test

* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets

* chore(make): point test-browser at test/browser

* docs(decisions): record internal/permissions deletion

* refactor(doctor): use sidecarassets package

* refactor(testrun): use sidecarassets package

* fix(test): resolve testdata relative to browser_test.go

* refactor(verifier): remove dead __sanderlingIndex compat alias

* refactor(bundler): use encoding/json for JS string literals

* docs(action-space): use vendor-neutral native driver wording

* refactor(hierarchy): scrub backend tool name from comments

* refactor(driver): scrub backend tool name from comments

* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver

* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap

* refactor(chrome): implement DoubleTap as two taps with the gap

* refactor(mock): record DoubleTap and DoubleTapSelector actions

* refactor(runner): delegate double-tap to driver, drop gesture timing

* test(runner): assert double-tap delegates to driver DoubleTap

* docs(cmd): add package docs to CLI and developer tools

* docs(driver): add package docs to driver interface and chrome backend

* docs(driver): add package docs to mock and sidecar backends

* docs(platform): add package docs to android and ios device prep

* docs: add package docs to bundler and inspect

* docs(ltl): add package doc to temporal logic evaluator

* docs: add package docs to runner and testrun pipeline

* docs: add package docs to trace and verifier

* docs(sidecarassets): add package doc for embedded JAR loader

* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI

* test(chrome): gate real-Chrome driver tests behind the browser tag

* chore(make): run chrome driver tests in the browser job

* fix(web-runtime): guard global error listeners for non-browser hosts

The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.

* ci(browser): re-enable unprivileged user namespaces for headless Chrome

ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.

* ci(browser): pin stable Chrome for the driver tests

setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.

* feat(defaults): add scroll and rebalance action weights

Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.

* feat(defaults): trim scroll action weight wiring

* fix(build): point sidecar jar ignore and embed paths at sidecarassets

* test(defaults): drop stale longPresses re-export assertion

longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.

* fix(chrome): raise DevTools websocket read timeout to 60s

Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.
2026-06-02 09:52:53 +05:30
pj 8ccf95c1cf refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling

Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.

* chore(proto): regenerate driverpb after module path rename

The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.

* refactor: rename CLI binary uatu -> sanderling

Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.

* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk

Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.

* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar

Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.

* refactor: rename Uatu API surface -> Sanderling

- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
  spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
  __uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.

* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling

Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.

* chore(build): rename gradle property + rootProject.name uatu -> sanderling

- Renames the uatu.version gradle property and all its -P references in
  Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.

* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1

Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.

* refactor: rename npm package @uatu/spec -> @sanderling/spec

Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.

* docs: rename uatu -> sanderling in README, docs, and URLs

- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
  'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
  internal/inspect/server.go.

* refactor: rename remaining internal uatu strings -> sanderling

- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
  localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
2026-04-21 11:57:49 +07:00
pj a2e96af1af WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio

Directory-level rename and path references in Go tests, bundle-check,
top-level README, and getting-started docs. Package declarations,
Gradle config, iOS bundle IDs, and class names follow in later commits.

* refactor(folio): rename Kotlin package dev.uatu.sample to app.folio

Moves source dirs and sqldelight schema from dev/uatu/sample to
app/folio, updates package declarations and imports, and switches
Android namespace/applicationId, iOS binaryOption bundleId, and
sqldelight database packageName to the new identifier.

* refactor(folio): rename SampleApplication to FolioApplication

Android manifest now points at .FolioApplication with label 'Folio'
instead of 'Uatu Sample'.

* refactor(folio): set iOS bundle id and display name to Folio

bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio.
CFBundleName + CFBundleDisplayName -> 'Folio'.

* refactor(folio): update demo email to [email protected]

* refactor(folio): point justfile at app.folio bundle id

Updates xcrun simctl launch target, uatu test --bundle-id, and the
build/uninstall comments to reference folio instead of sample.

* test: update fixture package ids to app.folio

Sidecar activity-resolver test and verifier spec-integration XML
fixtures referenced the old dev.uatu.sample Android package. Updates
them to match the folio app's real package id so the tests stay
representative of what the CLI sees on-device.

* test(verifier): rename SampleApp identifiers to Folio

Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and
sampleAppHierarchyXML const (now loginHierarchyXML for consistency with
the other per-screen fixtures). Updates trailing sample-app mentions in
comments and assertion messages.

* refactor(folio): rename Gradle/npm/wasm project identifiers to folio

settings.gradle.kts rootProject.name, package.json + package-lock.json
name, and the WasmJS index.html <title> all still read 'uatu-sample' /
'Uatu Sample'. Realigns them with the Folio brand.

* docs(folio): rewrite README title + getting-started bundle id

examples/folio/README.md is now titled 'Folio' with the Kotlin source
paths corrected to app/folio. Getting-started example uses --bundle-id
app.folio. Harness launch message is now generic ('app under test')
since uatu-sample-harness is not specific to folio.

* chore(folio): drop trailing 'sample' reference in gradle.properties

* refactor(folio): rename LoginPage composable to LoginScreen

Align with KMP/Android industry convention (NowInAndroid, Cash App,
JetBrains samples use Screen, not Page).

* refactor(folio): rename HomePage composable to HomeScreen

* refactor(folio): rename AddAccountPage composable to AddAccountScreen

* refactor(folio): rename LedgerPage composable to LedgerScreen

* refactor(folio): rename AddTransactionPage composable to AddTransactionScreen

* refactor(folio): split Models.kt into app.folio.data package

Account, Transaction (with TxnType), and Session move into their own
files under app.folio.data, matching NowInAndroid-style per-type
organization.

* refactor(folio): move data layer into app.folio.data package

Repository, LedgerStore (expect + interface), SqlLedgerStore,
WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext,
and Snapshot move into app.folio.data. Update all consumer imports.

* refactor(folio): move Navigation into app.folio.navigation package

Split the former Navigation.kt into Route.kt (sealed interface) and
Navigator.kt (singleton). Update consumer imports across screens,
App.kt, and FolioApplication.

* refactor(folio): move Platform and Format into app.folio.platform

Both files carry expect declarations (Platform object, formatDate);
grouping them into a dedicated platform package makes the KMP seam
obvious and mirrors the structure used by JetBrains samples.

* refactor(folio): move login into feature/auth package

Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline
the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into
LoginScreen since it is the sole caller.

* refactor(folio): move HomeScreen into feature/home package

* refactor(folio): move account creation into feature/account package

AddAccountScreen gets its own AddAccountUiState colocated with the
screen, replacing the shared UiState.addAccountError.

* refactor(folio): move ledger screens into feature/ledger package

LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger
with AddTransactionUiState (txnError, txnFormType) colocated. The
former catch-all UiState.kt is removed now that each screen owns its
state alongside its UI.

* refactor(folio): split Theme.kt; move theme and icons to subpackages

Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and
Type.kt (typography) under app.folio.ui.theme. Icons moves to
app.folio.ui.icon. Update every consumer's imports to match.

* refactor(folio): split ui components into per-file under ui/component

Former Widgets.kt and Components.kt become 10 focused files: AppButton,
Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/
BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style
of one composable per file in a designsystem/component package.

* chore(folio): consolidate uatu testing files under uatu/ folder

Move spec.ts, package.json, package-lock.json into examples/folio/uatu
so all uatu-specific testing artifacts live in one place. runs/ and
node_modules/ follow the same convention (both remain gitignored).
Update justfile, README, and the two Go consumers (bundle-check tool +
verifier/trace tests) that referenced the old path.

* refactor(trace): drop folio path in writer test

Round-trip only needs a non-empty string; neutralize to keep the
library free of folio references.

* refactor(sidecar): neutralize ResolveActivity test fixtures

Swap app.folio for com.example.app in the fixture strings so the
sidecar tests don't reference the example app by name.

* refactor(bundle-check): take spec path as argument

Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a
positional spec path instead so the tool works for any example and
leaves no folio reference in the library surface.

* test(verifier): add neutral integration spec and hierarchy fixtures

Adds testdata/integration_spec.ts with two routes ("list", "form"),
an InputText on text_field, a Tap on primary/secondary_action, a
safety property (itemCountNonNegative), and a liveness property
(submitEventually). Adds hierarchies_test.go with matching XML
fixtures. Constants intentionally go in a _test.go at package root
rather than testdata/hierarchies.go because go skips .go files
under testdata/.

* refactor(verifier): replace folio integration tests with neutral ones

Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test*
entry points to TestIntegrationSpec*. Uses the synthetic spec and
hierarchies added in the previous commit so the library's test suite
no longer references examples/folio at all.

Folio-specific coverage remains covered by examples/folio/justfile's
'just test' which exercises the real spec on device/emulator.

* chore: remove cmd/uatu-sample-harness

Not referenced by Makefile, docs, CI, or any script. Duplicates the
adb reverse helpers already in cmd/uatu/test_run.go, and its name
implies ownership by the sample app which violates the library/
example decoupling. If a bare-protocol debugging tool is later
needed it belongs inside cmd/uatu/.

* docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax

The directory-tree Layout section rots faster than the code and
duplicates what ls shows for free. The expect/actual paragraph in
Stack was the same kind of filler. The README also claimed 'just
AVD=Pixel_7 test' but the justfile reads AVD as an env var via
env_var_or_default, so the correct invocation is 'AVD=Pixel_7
just test'.
2026-04-20 16:04:20 +07:00
pj 0c775c199d feat(examples): add sample-app spec
Introduce examples/sample-app/spec.ts — a minimal property-based spec that
taps the sample app's "Click me" button and asserts click_count is
monotonic. Bundle-check and the trace writer test now reference the new
path.
2026-04-18 10:33:35 +07:00
pj 6ce4eabe86 feat(spec): drive merchant onboarding (English, phone, OTP, multi-device, home) 2026-04-18 01:59:51 +07:00
pj 4dbad9073e feat(spec): merchant-ledger property + login/navigation generators
Property ledgerBalanceMatchesTxns asserts displayed balance equals
sum(Given) - sum(Received) on either ledger screen. Catches the
class of bug where the server-fed CustomerModel.balance diverges
from the local-DB-fed CoreDatabaseDao sums (stale cache, partial
sync, deleted-txn handling glitch, etc.).

Generators are gated by screen state and weighted to push the run
through login -> home -> ledger quickly:
  enterPhone (100), enterOtp (100), openCustomerOrSupplier (80),
  taps (10), swipes (2).

Phone/OTP read from process.env via esbuild defines so credentials
never land in source. cmd/internal-tools/bundle-check is a quick
sanity tool to confirm the spec bundles before running uatu test.
2026-04-18 00:08:40 +07:00