Commit Graph
100 Commits
Author SHA1 Message Date
pj 8e8eb621c3 refactor(analyze): hoist the sign test's loop bound 2026-08-18 20:20:49 +05:30
pj 37e6298992 test(conformance): g4 sees a doubling appended to what the field held
Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.
2026-08-18 20:20:47 +05:30
pj 132ed25313 docs(skills): quote the summary line the runner prints now
The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.
2026-08-18 20:20:04 +05:30
pj a257a91988 docs(manual): exit 1 also means a run that holds no verdict
And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.
2026-08-18 20:20:04 +05:30
pj dd625de068 docs(analyze): name the tests the tool actually runs
The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.
2026-08-18 20:17:39 +05:30
pj 1008402558 fix(analyze): score a seed pair by which run outlived the other
The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.
2026-08-18 20:16:40 +05:30
pj e2f9864201 fix(spec): keep a secure selector valid beside another key
a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.
2026-08-18 20:16:03 +05:30
pj c8d52cb7d5 test(browser): the exit code a dead run and a violated one actually leave
Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.
2026-08-18 20:14:59 +05:30
pj f4b9ff1468 fix(conformance): g4 skips an input that names no field
jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.
2026-08-18 20:14:28 +05:30
pj faa1a33a68 test(conformance): g4 keeps checking past an input typed at coordinates
An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.
2026-08-18 20:14:14 +05:30
pj b9a48cc36f docs(manual): state what a secure selector matches 2026-08-18 20:14:07 +05:30
pj 9a920444b3 test(chrome): compare the secure fact across both producers
it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.
2026-08-18 20:14:07 +05:30
pj 6aecccd2d6 test(chrome): resolve the secure selector on both matchers
the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.
2026-08-18 20:14:07 +05:30
pj b0895ef179 fix(spec): a secure selector names the password field on web
secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.
2026-08-18 20:14:01 +05:30
pj 8414953135 fix(conformance): g4 reads a doubling off the observed field value
The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.
2026-08-18 20:13:52 +05:30
pj f0726a1e61 fix(analyze): compare arms on censored runs, not on flattened step counts
stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.
2026-08-18 20:13:12 +05:30
pj b7ee23942e feat(analyze): add the gehan generalized wilcoxon test
The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.
2026-08-18 20:13:05 +05:30
pj 3526d7a9dd refactor(analyze): open the log-rank up to a weight on the risk set
The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.
2026-08-18 20:12:59 +05:30
pj 74cae01a5c feat(cli): --allow-no-generator-actions
The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.
2026-08-18 20:10:29 +05:30
pj 93614a48b9 fix(testrun): the dead-run refusal gets its own opt-out
--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.
2026-08-18 20:10:29 +05:30
pj aa49c3bc34 fix(testrun): a recorded violation outranks the dead-run refusal
A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.
2026-08-18 20:08:03 +05:30
pj f631415cb6 test(conformance): the g4 fixture holds what a redacted android trace holds
Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.
2026-08-18 20:07:39 +05:30
pj f7b90a9886 docs(folio): say how to point just test at a remote adb server 2026-08-18 19:54:49 +05:30
pj 34f73b82f2 fix(folio): install through adb so a remote adb server works
Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.
2026-08-18 19:54:45 +05:30
pj ca3519ba3a docs(manual): what an action count with no producer means for a rate 2026-08-18 19:45:58 +05:30
pj ca6a05ca5e fix(analyze): mark an action count of unknown provenance in the report 2026-08-18 19:45:58 +05:30
pj 628eb1949a fix(analyze): refuse to compare attributed and unattributed denominators
One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.
2026-08-18 19:45:55 +05:30
pj 67c9de6365 fix(analyze): read how much of a record's action count names no producer
A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.
2026-08-18 19:45:55 +05:30
pj e0180f891a fix(campaign): a record always says how many actions named no producer
An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.
2026-08-18 19:45:49 +05:30
pj 153f857431 fix(defect-identity): degrade a redacted origin action to its selector
The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.
2026-08-18 19:43:28 +05:30
pj 454988fbc8 fix(trace): an action names the generator that produced it
The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.
2026-08-18 17:58:38 +05:30
pj 66fd5bce5d fix(runner): a secure field's value does not reach state.lastAction either
folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.
2026-08-18 17:32:28 +05:30
pj 38d328df90 fix(verifier): a secure field's typed value never reaches the record
A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.
2026-08-18 17:17:03 +05:30
pj b1e95739ad feat(hierarchy): an element reports whether it masks what is typed into it
ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".
2026-08-18 17:16:55 +05:30
pj 5f1f50c2fb fix(campaign): the action count leaves the setup's login out on a model run
Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.
2026-08-18 16:52:22 +05:30
pj b278fc869d fix(testrun): an ios run records the simulator it executed on
Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.
2026-08-18 16:38:00 +05:30
pj 6cac99cc32 feat(runner): the summary says how many steps the generator drove
A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.
2026-08-18 16:38:00 +05:30
pj 58168a6a3e fix(testrun): the refusal asks whether the generator drove, not whether anything did
A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.
2026-08-18 16:29:03 +05:30
pj 5a36b583ea test(verifier): an unreadable committed fixture fails, it does not skip
The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.
2026-08-18 16:13:06 +05:30
pj e08e1cc83b feat(bundle-check): --allow-no-properties opts out of the refusal
The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.
2026-08-18 16:12:56 +05:30
pj c8fcc21e6b docs(spec-language): name the hintText selector's host divergence
The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.
2026-08-18 16:11:45 +05:30
pj 15126c291d docs(cli): document --label-source 2026-08-18 16:11:45 +05:30
pj 513b33c0e5 feat(testrun): refuse a run that dispatched none of its actions
Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.
2026-08-18 14:16:47 +05:30
pj 0646611a7f fix(runner): a source that was asked and handed nothing says so
NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.
2026-08-18 14:16:42 +05:30
pj 1da0c3e118 fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets
A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.
2026-08-18 14:11:28 +05:30
pj ff6c66a74b fix(confusion-matrix): a campaign that died is missing data, not a true negative
The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.
2026-08-18 14:08:59 +05:30
pj 1d684a98eb ci: let a restored ios bundle survive make's mtime check
The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.
2026-08-18 13:04:31 +05:30
pj 7b6a346443 feat(trace): record the device a run executed on
meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj a7f26e27c4 fix(campaign): quarantine a device that keeps failing fast
Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj bdf5f80842 fix(campaign): refuse to start on a device that is not there
A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-18 13:03:42 +05:30
pj dd9c3a1d9a ci: pin the idb-companion tap to the formula the companion is staged from
The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.
2026-08-18 09:44:05 +05:30
pj badabbaed8 fix(confusion-matrix): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 627f9eeeff fix(implementation-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 5892983ff8 fix(corpus-sweep): name every missing required flag, in flag order 2026-08-18 00:17:10 +05:30
pj 2fd67d42f9 fix(campaign): name every missing required flag, in flag order
Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.
2026-08-18 00:17:10 +05:30
pj ab4c42601d fix(web-runtime): focus follows the caret to the field it types into
Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.
2026-08-18 00:16:52 +05:30
pj 04327e63b6 test(chrome): the handle and the enumeration agree on editable too
The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.
2026-08-18 00:15:55 +05:30
pj 93580e076d test(chrome): a hinted field is not named by its css class
The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.
2026-08-18 00:10:53 +05:30
pj bd01a76d70 fix(web-runtime): a handle answers editable for itself, not its container
isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.
2026-08-18 00:05:16 +05:30
pj d5f48266e8 fix(corpus-sweep): name every missing binary, in flag order
Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.
2026-08-18 00:01:26 +05:30
pj 5fc91fdca1 fix(chrome): focus follows the caret to the field it types into
Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.
2026-08-18 00:00:24 +05:30
pj 176b495245 fix(implementation-sweep): name every missing binary, in flag order
Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.
2026-08-17 23:57:54 +05:30
pj 6706461b91 fix(web-runtime): focus descends into the shadow root here too
The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.
2026-08-17 23:48:14 +05:30
pj e432af7828 fix(replay-ui): read data-* attributes by their markup names
The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.
2026-08-17 23:45:33 +05:30
pj 5651adb439 test(implementation-sweep): supply the binaries the missing-binary test does not test
resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.
2026-08-17 23:44:26 +05:30
pj 071f9c7f52 fix(chrome): focus descends into the shadow root
document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.
2026-08-17 23:44:21 +05:30
pj 0bfd194332 test(confusion-matrix): cover cell assignment, precision and recall 2026-08-17 23:27:51 +05:30
pj cb7936fe36 test(confusion-matrix): reject malformed inputs and keep missing data out of the cells 2026-08-17 23:27:51 +05:30
pj 56a002d87a feat(confusion-matrix): score the checker against a blind reviewer
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.
2026-08-17 23:27:47 +05:30
pj 8c8e87761d fix(folio-web): keep submit live for 400ms after saving
Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.
2026-08-17 23:27:37 +05:30
pj bc0f286e29 feat(folio-web): judge one commit per submit over a home-card window
Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.
2026-08-17 23:27:33 +05:30
pj dc2da35e2c feat(folio-web): predicates for counting commits against submit actions 2026-08-17 23:27:33 +05:30
pj 2aba29a6e3 test(bundle-check): cover the zero-property refusal and pin the reported bundle 2026-08-17 23:27:29 +05:30
pj 26455d5ec4 feat(bundle-check): fail a spec that bundles but registers no properties 2026-08-17 23:27:29 +05:30
pj 9523133c59 docs(cli): document --allow-no-properties 2026-08-17 23:27:25 +05:30
pj 1e8db8524d feat(cli): --allow-no-properties opts a run out of the refusal 2026-08-17 23:27:25 +05:30
pj 7bafed8af4 feat(testrun): refuse a run against a spec that registers no properties
A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.
2026-08-17 23:27:21 +05:30
pj 11c1569932 feat(verifier): expose the property names a loaded spec registered 2026-08-17 23:24:31 +05:30
pj 573749f17a fix(make): build the binary instead of matching the build directory
build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.
2026-08-17 23:24:27 +05:30
pj d84975e042 test(chrome): both resolvers agree on tag where a container's name contains its child's 2026-08-17 23:24:10 +05:30
pj 9012ffa304 fix(selectors): tag names the whole tag, not a substring of it
matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.
2026-08-17 23:24:06 +05:30
pj 81fe5f66f2 merge origin/master into llm-recording-and-analysis 2026-08-17 23:21:11 +05:30
pj bf73f41300 docs(triage): name the trace field a run that never started leaves 2026-08-17 14:35:37 +05:30
pj ddf17129bd feat(campaign): count the runs that were never in the app
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
2026-08-17 14:35:37 +05:30
pj ca8b517d98 test(runner): the gate keeps looking until its budget runs out
Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.
2026-08-17 14:35:37 +05:30
pj a7ed433706 fix(runner): budget the foreground gate in time, not in polls
Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.
2026-08-17 14:35:29 +05:30
pj 1ebb8b191f feat(trace): a step can name the precondition it could not meet
A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.
2026-08-17 14:35:19 +05:30
pj 95a19fcc1e ci: pin the first-party actions to commit shas (#86)
third-party actions here were already sha-pinned; the actions/* ones were
on major tags, which are mutable and can be repointed. same treatment and
same trailing version comment, so an upgrade stays a readable diff.
2026-08-17 10:26:05 +05:30
pj 93c2d2ba74 ci: fix the node 20 warning and pin the protoc plugins (#84)
* ci: replace the archived buf-setup-action with buf-action

buf-setup-action is archived and runs on node20, which the runners now
warn about. buf-action is its supported replacement and runs on node24.
setup_only keeps it an install, since buf lint is its own step.

* ci: bump bun to 1.3.14

* ci: bump setup-chrome to v2.2.0

* ci: move to node 24 and drop the npm oidc workaround

node 22 is in maintenance and ships npm 10, which is why the publish job
had to install npm@latest over it. node 24 is the active lts and bundles
npm 11.17.0, above the 11.5.1 oidc floor, so the extra step goes.

* ci: pin the protoc plugins instead of installing @latest

these generate the committed stubs, so @latest makes codegen depend on
whatever released most recently. pinned to the versions proto/ records:
protoc-gen-go v1.36.11, protoc-gen-go-grpc v1.6.0.

* ci: run the ios leg on macos-26, pinned to a device and a runtime

macos-26 carries no iPhone 16 Pro at all, and on macos-15 that name spanned
iOS 18.5 through 26.2, so the leg could boot a two-major-old runtime. the
pair is now iPhone 17 Pro on iOS 26.2, resolved to a udid before boot, and
an image that drops it fails naming what it does carry.

iPhone 17 Pro is what examples/folio/justfile already defaulted to.

* ci: keep IOS_DEVICE a device name, not the resolved udid

the boot step exported the udid as IOS_DEVICE, and just ios spends that as
xcodebuild's -destination name=, which matches the display name and
rejected it: 'unable to find a device matching { name:6F69910C-... }'.

nothing downstream needed it. install, launch and terminate all address
booted, and sanderling resolves --ios-device against booted simulators
first, so the simulator this step boots is the one they all get.
2026-08-16 22:57:38 +05:30
pj 6e85cac8b3 merge origin/master into llm-recording-and-analysis
both sides independently fixed the same three bugs, so each one had to pick a
winner rather than keep both implementations.

extractor encoding: master's recordableValue in worker.go wins over ours in
marshal.go, since master's is pinned by extractor_encoding_test.go and ours had
no tests. our error semantics stay: encodeExtractorValue still returns an error
instead of nil, so an extractor cannot vanish from the trace silently.

apply errors: only the residual generic branch takes master's unconfirmed copy,
where the device may have committed the action before the call failed. the
finer branches that know nothing was dispatched keep lastAction = nil, and our
actionSkipReason taxonomy stays alongside master's held/skippedVerification.

selector matching: our matchAttr with matchSelectorKind wins over master's
match, since ours also handles idPrefix. matchSelector now calls it, which git
did not flag as a conflict and left calling a function our side had deleted.

the ltl doc comment takes master's correction: an unbounded eventually that
never fires IS violated at run end.
2026-08-16 18:10:55 +05:30
pj 56f66817f7 test(browser): assert an uncaught page exception reaches the trace
the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.
2026-08-16 17:45:47 +05:30
pj 5967b4c79b docs(manual): document innermost text matching, escape and the web scroll verb
text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.
2026-08-16 17:45:47 +05:30
pj 341d6a0614 feat(corpus-sweep): run one specification against a served corpus of implementations
same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.
2026-08-16 17:45:42 +05:30
pj 146152a3af feat(implementation-sweep): run one campaign against every implementation of a requirement
installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.
2026-08-16 17:45:41 +05:30
pj a45ba76d8e feat(oracle-reduction): replay stored traces under four reduced oracles
re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.
2026-08-16 17:45:41 +05:30
pj a0a8c9c710 feat(defect-identity): count distinct defects across stored runs
a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.
2026-08-16 17:45:35 +05:30
pj 11fca22d36 feat(exploration-reach): count the distinct structural states a stored run visited
the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.
2026-08-16 17:45:35 +05:30
pj 0d1b2cf4e8 feat(label-coverage): report the addressable share of an app's interactive surface
reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.
2026-08-16 17:45:35 +05:30
pj b47509a8e8 test(analyze): recover planted effects through the tool's own entry point
a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.
2026-08-16 17:45:26 +05:30
pj 4dff70b18b feat(analyze): add the seed-paired signed-rank comparison and record the holm family
--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.
2026-08-16 17:45:26 +05:30