Commit Graph
290 Commits
Author SHA1 Message Date
pj ff6c66a74b fix(confusion-matrix): a campaign that died is missing data, not a true negative
The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.
2026-08-18 14:08:59 +05:30
pj 1d684a98eb ci: let a restored ios bundle survive make's mtime check
The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.
2026-08-18 13:04:31 +05:30
pj 7b6a346443 feat(trace): record the device a run executed on
meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj a7f26e27c4 fix(campaign): quarantine a device that keeps failing fast
Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj bdf5f80842 fix(campaign): refuse to start on a device that is not there
A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj dd9c3a1d9a ci: pin the idb-companion tap to the formula the companion is staged from
The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.
2026-08-18 09:44:05 +05:30
pj badabbaed8 fix(confusion-matrix): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 627f9eeeff fix(implementation-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 5892983ff8 fix(corpus-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 2fd67d42f9 fix(campaign): name every missing required flag, in flag order
Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.
2026-08-18 00:17:10 +05:30
pj ab4c42601d fix(web-runtime): focus follows the caret to the field it types into
Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.
2026-08-18 00:16:52 +05:30
pj 04327e63b6 test(chrome): the handle and the enumeration agree on editable too
The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.
2026-08-18 00:15:55 +05:30
pj 93580e076d test(chrome): a hinted field is not named by its css class
The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.
2026-08-18 00:10:53 +05:30
pj bd01a76d70 fix(web-runtime): a handle answers editable for itself, not its container
isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.
2026-08-18 00:05:16 +05:30
pj d5f48266e8 fix(corpus-sweep): name every missing binary, in flag order
Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.
2026-08-18 00:01:26 +05:30
pj 5fc91fdca1 fix(chrome): focus follows the caret to the field it types into
Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.
2026-08-18 00:00:24 +05:30
pj 176b495245 fix(implementation-sweep): name every missing binary, in flag order
Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.
2026-08-17 23:57:54 +05:30
pj 6706461b91 fix(web-runtime): focus descends into the shadow root here too
The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.
2026-08-17 23:48:14 +05:30
pj e432af7828 fix(replay-ui): read data-* attributes by their markup names
The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.
2026-08-17 23:45:33 +05:30
pj 5651adb439 test(implementation-sweep): supply the binaries the missing-binary test does not test
resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.
2026-08-17 23:44:26 +05:30
pj 071f9c7f52 fix(chrome): focus descends into the shadow root
document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.
2026-08-17 23:44:21 +05:30
pj 0bfd194332 test(confusion-matrix): cover cell assignment, precision and recall 2026-08-17 23:27:51 +05:30
pj cb7936fe36 test(confusion-matrix): reject malformed inputs and keep missing data out of the cells 2026-08-17 23:27:51 +05:30
pj 56a002d87a feat(confusion-matrix): score the checker against a blind reviewer
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.
2026-08-17 23:27:47 +05:30
pj 8c8e87761d fix(folio-web): keep submit live for 400ms after saving
Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.
2026-08-17 23:27:37 +05:30
pj bc0f286e29 feat(folio-web): judge one commit per submit over a home-card window
Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.
2026-08-17 23:27:33 +05:30
pj dc2da35e2c feat(folio-web): predicates for counting commits against submit actions 2026-08-17 23:27:33 +05:30
pj 2aba29a6e3 test(bundle-check): cover the zero-property refusal and pin the reported bundle 2026-08-17 23:27:29 +05:30
pj 26455d5ec4 feat(bundle-check): fail a spec that bundles but registers no properties 2026-08-17 23:27:29 +05:30
pj 9523133c59 docs(cli): document --allow-no-properties 2026-08-17 23:27:25 +05:30
pj 1e8db8524d feat(cli): --allow-no-properties opts a run out of the refusal 2026-08-17 23:27:25 +05:30
pj 7bafed8af4 feat(testrun): refuse a run against a spec that registers no properties
A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.
2026-08-17 23:27:21 +05:30
pj 11c1569932 feat(verifier): expose the property names a loaded spec registered 2026-08-17 23:24:31 +05:30
pj 573749f17a fix(make): build the binary instead of matching the build directory
build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.
2026-08-17 23:24:27 +05:30
pj d84975e042 test(chrome): both resolvers agree on tag where a container's name contains its child's 2026-08-17 23:24:10 +05:30
pj 9012ffa304 fix(selectors): tag names the whole tag, not a substring of it
matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.
2026-08-17 23:24:06 +05:30
pj 81fe5f66f2 merge origin/master into llm-recording-and-analysis 2026-08-17 23:21:11 +05:30
pj bf73f41300 docs(triage): name the trace field a run that never started leaves 2026-08-17 14:35:37 +05:30
pj ddf17129bd feat(campaign): count the runs that were never in the app
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
2026-08-17 14:35:37 +05:30
pj ca8b517d98 test(runner): the gate keeps looking until its budget runs out
Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.
2026-08-17 14:35:37 +05:30
pj a7ed433706 fix(runner): budget the foreground gate in time, not in polls
Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.
2026-08-17 14:35:29 +05:30
pj 1ebb8b191f feat(trace): a step can name the precondition it could not meet
A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.
2026-08-17 14:35:19 +05:30
pj 95a19fcc1e ci: pin the first-party actions to commit shas (#86)
third-party actions here were already sha-pinned; the actions/* ones were
on major tags, which are mutable and can be repointed. same treatment and
same trailing version comment, so an upgrade stays a readable diff.
v0.0.3
2026-08-17 10:26:05 +05:30
pj 93c2d2ba74 ci: fix the node 20 warning and pin the protoc plugins (#84)
* ci: replace the archived buf-setup-action with buf-action

buf-setup-action is archived and runs on node20, which the runners now
warn about. buf-action is its supported replacement and runs on node24.
setup_only keeps it an install, since buf lint is its own step.

* ci: bump bun to 1.3.14

* ci: bump setup-chrome to v2.2.0

* ci: move to node 24 and drop the npm oidc workaround

node 22 is in maintenance and ships npm 10, which is why the publish job
had to install npm@latest over it. node 24 is the active lts and bundles
npm 11.17.0, above the 11.5.1 oidc floor, so the extra step goes.

* ci: pin the protoc plugins instead of installing @latest

these generate the committed stubs, so @latest makes codegen depend on
whatever released most recently. pinned to the versions proto/ records:
protoc-gen-go v1.36.11, protoc-gen-go-grpc v1.6.0.

* ci: run the ios leg on macos-26, pinned to a device and a runtime

macos-26 carries no iPhone 16 Pro at all, and on macos-15 that name spanned
iOS 18.5 through 26.2, so the leg could boot a two-major-old runtime. the
pair is now iPhone 17 Pro on iOS 26.2, resolved to a udid before boot, and
an image that drops it fails naming what it does carry.

iPhone 17 Pro is what examples/folio/justfile already defaulted to.

* ci: keep IOS_DEVICE a device name, not the resolved udid

the boot step exported the udid as IOS_DEVICE, and just ios spends that as
xcodebuild's -destination name=, which matches the display name and
rejected it: 'unable to find a device matching { name:6F69910C-... }'.

nothing downstream needed it. install, launch and terminate all address
booted, and sanderling resolves --ios-device against booted simulators
first, so the simulator this step boots is the one they all get.
v0.0.2
2026-08-16 22:57:38 +05:30
pj 6e85cac8b3 merge origin/master into llm-recording-and-analysis
both sides independently fixed the same three bugs, so each one had to pick a
winner rather than keep both implementations.

extractor encoding: master's recordableValue in worker.go wins over ours in
marshal.go, since master's is pinned by extractor_encoding_test.go and ours had
no tests. our error semantics stay: encodeExtractorValue still returns an error
instead of nil, so an extractor cannot vanish from the trace silently.

apply errors: only the residual generic branch takes master's unconfirmed copy,
where the device may have committed the action before the call failed. the
finer branches that know nothing was dispatched keep lastAction = nil, and our
actionSkipReason taxonomy stays alongside master's held/skippedVerification.

selector matching: our matchAttr with matchSelectorKind wins over master's
match, since ours also handles idPrefix. matchSelector now calls it, which git
did not flag as a conflict and left calling a function our side had deleted.

the ltl doc comment takes master's correction: an unbounded eventually that
never fires IS violated at run end.
2026-08-16 18:10:55 +05:30
pj 56f66817f7 test(browser): assert an uncaught page exception reaches the trace
the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.
2026-08-16 17:45:47 +05:30
pj 5967b4c79b docs(manual): document innermost text matching, escape and the web scroll verb
text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.
2026-08-16 17:45:47 +05:30
pj 341d6a0614 feat(corpus-sweep): run one specification against a served corpus of implementations
same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.
2026-08-16 17:45:42 +05:30
pj 146152a3af feat(implementation-sweep): run one campaign against every implementation of a requirement
installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.
2026-08-16 17:45:41 +05:30
pj a45ba76d8e feat(oracle-reduction): replay stored traces under four reduced oracles
re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.
2026-08-16 17:45:41 +05:30