Commit Graph
282 Commits
Author SHA1 Message Date
pj 5892983ff8 fix(corpus-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 2fd67d42f9 fix(campaign): name every missing required flag, in flag order
Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.
2026-08-18 00:17:10 +05:30
pj ab4c42601d fix(web-runtime): focus follows the caret to the field it types into
Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.
2026-08-18 00:16:52 +05:30
pj 04327e63b6 test(chrome): the handle and the enumeration agree on editable too
The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.
2026-08-18 00:15:55 +05:30
pj 93580e076d test(chrome): a hinted field is not named by its css class
The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.
2026-08-18 00:10:53 +05:30
pj bd01a76d70 fix(web-runtime): a handle answers editable for itself, not its container
isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.
2026-08-18 00:05:16 +05:30
pj d5f48266e8 fix(corpus-sweep): name every missing binary, in flag order
Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.
2026-08-18 00:01:26 +05:30
pj 5fc91fdca1 fix(chrome): focus follows the caret to the field it types into
Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.
2026-08-18 00:00:24 +05:30
pj 176b495245 fix(implementation-sweep): name every missing binary, in flag order
Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.
2026-08-17 23:57:54 +05:30
pj 6706461b91 fix(web-runtime): focus descends into the shadow root here too
The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.
2026-08-17 23:48:14 +05:30
pj e432af7828 fix(replay-ui): read data-* attributes by their markup names
The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.
2026-08-17 23:45:33 +05:30
pj 5651adb439 test(implementation-sweep): supply the binaries the missing-binary test does not test
resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.
2026-08-17 23:44:26 +05:30
pj 071f9c7f52 fix(chrome): focus descends into the shadow root
document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.
2026-08-17 23:44:21 +05:30
pj 0bfd194332 test(confusion-matrix): cover cell assignment, precision and recall 2026-08-17 23:27:51 +05:30
pj cb7936fe36 test(confusion-matrix): reject malformed inputs and keep missing data out of the cells 2026-08-17 23:27:51 +05:30
pj 56a002d87a feat(confusion-matrix): score the checker against a blind reviewer
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.
2026-08-17 23:27:47 +05:30
pj 8c8e87761d fix(folio-web): keep submit live for 400ms after saving
Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.
2026-08-17 23:27:37 +05:30
pj bc0f286e29 feat(folio-web): judge one commit per submit over a home-card window
Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.
2026-08-17 23:27:33 +05:30
pj dc2da35e2c feat(folio-web): predicates for counting commits against submit actions 2026-08-17 23:27:33 +05:30
pj 2aba29a6e3 test(bundle-check): cover the zero-property refusal and pin the reported bundle 2026-08-17 23:27:29 +05:30
pj 26455d5ec4 feat(bundle-check): fail a spec that bundles but registers no properties 2026-08-17 23:27:29 +05:30
pj 9523133c59 docs(cli): document --allow-no-properties 2026-08-17 23:27:25 +05:30
pj 1e8db8524d feat(cli): --allow-no-properties opts a run out of the refusal 2026-08-17 23:27:25 +05:30
pj 7bafed8af4 feat(testrun): refuse a run against a spec that registers no properties
A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.
2026-08-17 23:27:21 +05:30
pj 11c1569932 feat(verifier): expose the property names a loaded spec registered 2026-08-17 23:24:31 +05:30
pj 573749f17a fix(make): build the binary instead of matching the build directory
build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.
2026-08-17 23:24:27 +05:30
pj d84975e042 test(chrome): both resolvers agree on tag where a container's name contains its child's 2026-08-17 23:24:10 +05:30
pj 9012ffa304 fix(selectors): tag names the whole tag, not a substring of it
matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.
2026-08-17 23:24:06 +05:30
pj 81fe5f66f2 merge origin/master into llm-recording-and-analysis 2026-08-17 23:21:11 +05:30
pj bf73f41300 docs(triage): name the trace field a run that never started leaves 2026-08-17 14:35:37 +05:30
pj ddf17129bd feat(campaign): count the runs that were never in the app
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
2026-08-17 14:35:37 +05:30
pj ca8b517d98 test(runner): the gate keeps looking until its budget runs out
Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.
2026-08-17 14:35:37 +05:30
pj a7ed433706 fix(runner): budget the foreground gate in time, not in polls
Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.
2026-08-17 14:35:29 +05:30
pj 1ebb8b191f feat(trace): a step can name the precondition it could not meet
A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.
2026-08-17 14:35:19 +05:30
pj 95a19fcc1e ci: pin the first-party actions to commit shas (#86)
third-party actions here were already sha-pinned; the actions/* ones were
on major tags, which are mutable and can be repointed. same treatment and
same trailing version comment, so an upgrade stays a readable diff.
v0.0.3
2026-08-17 10:26:05 +05:30
pj 93c2d2ba74 ci: fix the node 20 warning and pin the protoc plugins (#84)
* ci: replace the archived buf-setup-action with buf-action

buf-setup-action is archived and runs on node20, which the runners now
warn about. buf-action is its supported replacement and runs on node24.
setup_only keeps it an install, since buf lint is its own step.

* ci: bump bun to 1.3.14

* ci: bump setup-chrome to v2.2.0

* ci: move to node 24 and drop the npm oidc workaround

node 22 is in maintenance and ships npm 10, which is why the publish job
had to install npm@latest over it. node 24 is the active lts and bundles
npm 11.17.0, above the 11.5.1 oidc floor, so the extra step goes.

* ci: pin the protoc plugins instead of installing @latest

these generate the committed stubs, so @latest makes codegen depend on
whatever released most recently. pinned to the versions proto/ records:
protoc-gen-go v1.36.11, protoc-gen-go-grpc v1.6.0.

* ci: run the ios leg on macos-26, pinned to a device and a runtime

macos-26 carries no iPhone 16 Pro at all, and on macos-15 that name spanned
iOS 18.5 through 26.2, so the leg could boot a two-major-old runtime. the
pair is now iPhone 17 Pro on iOS 26.2, resolved to a udid before boot, and
an image that drops it fails naming what it does carry.

iPhone 17 Pro is what examples/folio/justfile already defaulted to.

* ci: keep IOS_DEVICE a device name, not the resolved udid

the boot step exported the udid as IOS_DEVICE, and just ios spends that as
xcodebuild's -destination name=, which matches the display name and
rejected it: 'unable to find a device matching { name:6F69910C-... }'.

nothing downstream needed it. install, launch and terminate all address
booted, and sanderling resolves --ios-device against booted simulators
first, so the simulator this step boots is the one they all get.
v0.0.2
2026-08-16 22:57:38 +05:30
pj 6e85cac8b3 merge origin/master into llm-recording-and-analysis
both sides independently fixed the same three bugs, so each one had to pick a
winner rather than keep both implementations.

extractor encoding: master's recordableValue in worker.go wins over ours in
marshal.go, since master's is pinned by extractor_encoding_test.go and ours had
no tests. our error semantics stay: encodeExtractorValue still returns an error
instead of nil, so an extractor cannot vanish from the trace silently.

apply errors: only the residual generic branch takes master's unconfirmed copy,
where the device may have committed the action before the call failed. the
finer branches that know nothing was dispatched keep lastAction = nil, and our
actionSkipReason taxonomy stays alongside master's held/skippedVerification.

selector matching: our matchAttr with matchSelectorKind wins over master's
match, since ours also handles idPrefix. matchSelector now calls it, which git
did not flag as a conflict and left calling a function our side had deleted.

the ltl doc comment takes master's correction: an unbounded eventually that
never fires IS violated at run end.
2026-08-16 18:10:55 +05:30
pj 56f66817f7 test(browser): assert an uncaught page exception reaches the trace
the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.
2026-08-16 17:45:47 +05:30
pj 5967b4c79b docs(manual): document innermost text matching, escape and the web scroll verb
text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.
2026-08-16 17:45:47 +05:30
pj 341d6a0614 feat(corpus-sweep): run one specification against a served corpus of implementations
same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.
2026-08-16 17:45:42 +05:30
pj 146152a3af feat(implementation-sweep): run one campaign against every implementation of a requirement
installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.
2026-08-16 17:45:41 +05:30
pj a45ba76d8e feat(oracle-reduction): replay stored traces under four reduced oracles
re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.
2026-08-16 17:45:41 +05:30
pj a0a8c9c710 feat(defect-identity): count distinct defects across stored runs
a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.
2026-08-16 17:45:35 +05:30
pj 11fca22d36 feat(exploration-reach): count the distinct structural states a stored run visited
the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.
2026-08-16 17:45:35 +05:30
pj 0d1b2cf4e8 feat(label-coverage): report the addressable share of an app's interactive surface
reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.
2026-08-16 17:45:35 +05:30
pj b47509a8e8 test(analyze): recover planted effects through the tool's own entry point
a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.
2026-08-16 17:45:26 +05:30
pj 4dff70b18b feat(analyze): add the seed-paired signed-rank comparison and record the holm family
--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.
2026-08-16 17:45:26 +05:30
pj b99da0be0e feat(analyze): time an event at the step it was detected and report the quartiles
an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.
2026-08-16 17:45:20 +05:30
pj a28178337c refactor(seedspec): move seed spec parsing out of the campaign command
the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.
2026-08-16 17:44:57 +05:30
pj 7823f53453 feat(tracecorpus): load recorded runs for offline measures
reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.
2026-08-16 17:44:52 +05:30