Commit Graph
350 Commits
Author SHA1 Message Date
pj b278fc869d fix(testrun): an ios run records the simulator it executed on
Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.
2026-08-18 16:38:00 +05:30
pj 6cac99cc32 feat(runner): the summary says how many steps the generator drove
A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.
2026-08-18 16:38:00 +05:30
pj 58168a6a3e fix(testrun): the refusal asks whether the generator drove, not whether anything did
A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.
2026-08-18 16:29:03 +05:30
pj 5a36b583ea test(verifier): an unreadable committed fixture fails, it does not skip
The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.
2026-08-18 16:13:06 +05:30
pj e08e1cc83b feat(bundle-check): --allow-no-properties opts out of the refusal
The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.
2026-08-18 16:12:56 +05:30
pj c8fcc21e6b docs(spec-language): name the hintText selector's host divergence
The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.
2026-08-18 16:11:45 +05:30
pj 15126c291d docs(cli): document --label-source 2026-08-18 16:11:45 +05:30
pj 513b33c0e5 feat(testrun): refuse a run that dispatched none of its actions
Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.
2026-08-18 14:16:47 +05:30
pj 0646611a7f fix(runner): a source that was asked and handed nothing says so
NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.
2026-08-18 14:16:42 +05:30
pj 1da0c3e118 fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets
A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.
2026-08-18 14:11:28 +05:30
pj ff6c66a74b fix(confusion-matrix): a campaign that died is missing data, not a true negative
The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.
2026-08-18 14:08:59 +05:30
pj 1d684a98eb ci: let a restored ios bundle survive make's mtime check
The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.
2026-08-18 13:04:31 +05:30
pj 7b6a346443 feat(trace): record the device a run executed on
meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj a7f26e27c4 fix(campaign): quarantine a device that keeps failing fast
Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj bdf5f80842 fix(campaign): refuse to start on a device that is not there
A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj dd9c3a1d9a ci: pin the idb-companion tap to the formula the companion is staged from
The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.
2026-08-18 09:44:05 +05:30
pj badabbaed8 fix(confusion-matrix): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 627f9eeeff fix(implementation-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 5892983ff8 fix(corpus-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 2fd67d42f9 fix(campaign): name every missing required flag, in flag order
Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.
2026-08-18 00:17:10 +05:30
pj ab4c42601d fix(web-runtime): focus follows the caret to the field it types into
Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.
2026-08-18 00:16:52 +05:30
pj 04327e63b6 test(chrome): the handle and the enumeration agree on editable too
The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.
2026-08-18 00:15:55 +05:30
pj 93580e076d test(chrome): a hinted field is not named by its css class
The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.
2026-08-18 00:10:53 +05:30
pj bd01a76d70 fix(web-runtime): a handle answers editable for itself, not its container
isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.
2026-08-18 00:05:16 +05:30
pj d5f48266e8 fix(corpus-sweep): name every missing binary, in flag order
Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.
2026-08-18 00:01:26 +05:30
pj 5fc91fdca1 fix(chrome): focus follows the caret to the field it types into
Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.
2026-08-18 00:00:24 +05:30
pj 176b495245 fix(implementation-sweep): name every missing binary, in flag order
Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.
2026-08-17 23:57:54 +05:30
pj 6706461b91 fix(web-runtime): focus descends into the shadow root here too
The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.
2026-08-17 23:48:14 +05:30
pj e432af7828 fix(replay-ui): read data-* attributes by their markup names
The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.
2026-08-17 23:45:33 +05:30
pj 5651adb439 test(implementation-sweep): supply the binaries the missing-binary test does not test
resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.
2026-08-17 23:44:26 +05:30
pj 071f9c7f52 fix(chrome): focus descends into the shadow root
document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.
2026-08-17 23:44:21 +05:30
pj 0bfd194332 test(confusion-matrix): cover cell assignment, precision and recall 2026-08-17 23:27:51 +05:30
pj cb7936fe36 test(confusion-matrix): reject malformed inputs and keep missing data out of the cells 2026-08-17 23:27:51 +05:30
pj 56a002d87a feat(confusion-matrix): score the checker against a blind reviewer
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.
2026-08-17 23:27:47 +05:30
pj 8c8e87761d fix(folio-web): keep submit live for 400ms after saving
Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.
2026-08-17 23:27:37 +05:30
pj bc0f286e29 feat(folio-web): judge one commit per submit over a home-card window
Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.
2026-08-17 23:27:33 +05:30
pj dc2da35e2c feat(folio-web): predicates for counting commits against submit actions 2026-08-17 23:27:33 +05:30
pj 2aba29a6e3 test(bundle-check): cover the zero-property refusal and pin the reported bundle 2026-08-17 23:27:29 +05:30
pj 26455d5ec4 feat(bundle-check): fail a spec that bundles but registers no properties 2026-08-17 23:27:29 +05:30
pj 9523133c59 docs(cli): document --allow-no-properties 2026-08-17 23:27:25 +05:30
pj 1e8db8524d feat(cli): --allow-no-properties opts a run out of the refusal 2026-08-17 23:27:25 +05:30
pj 7bafed8af4 feat(testrun): refuse a run against a spec that registers no properties
A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.
2026-08-17 23:27:21 +05:30
pj 11c1569932 feat(verifier): expose the property names a loaded spec registered 2026-08-17 23:24:31 +05:30
pj 573749f17a fix(make): build the binary instead of matching the build directory
build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.
2026-08-17 23:24:27 +05:30
pj d84975e042 test(chrome): both resolvers agree on tag where a container's name contains its child's 2026-08-17 23:24:10 +05:30
pj 9012ffa304 fix(selectors): tag names the whole tag, not a substring of it
matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.
2026-08-17 23:24:06 +05:30
pj 81fe5f66f2 merge origin/master into llm-recording-and-analysis 2026-08-17 23:21:11 +05:30
pj bf73f41300 docs(triage): name the trace field a run that never started leaves 2026-08-17 14:35:37 +05:30
pj ddf17129bd feat(campaign): count the runs that were never in the app
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
2026-08-17 14:35:37 +05:30
pj ca8b517d98 test(runner): the gate keeps looking until its budget runs out
Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.
2026-08-17 14:35:37 +05:30