Commit Graph
354 Commits
Author SHA1 Message Date
pj 5a7c223464 docs(manual): list className among the cross-platform aliases
the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.
2026-08-19 09:55:39 +05:30
pj 2cbd60d70b test(spec): pin className and class on the same elements
this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.
2026-08-19 09:55:35 +05:30
pj 53ee9aa292 test(chrome): compare className across both matchers
one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.
2026-08-19 09:55:25 +05:30
pj 5f66e5fb3e fix(hierarchy): reach the class attribute through className
className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.
2026-08-19 09:55:22 +05:30
pj 31f9c8de11 docs(manual): state how text combines with the key beside it
the object selector section said every pair must match without saying
where the innermost rule then lands.
2026-08-19 09:41:48 +05:30
pj 4b3c2601e7 test(chrome): compare a compound text selector across both matchers
one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.
2026-08-19 09:41:44 +05:30
pj c16cb88d38 test(spec): pin text against the key beside it in either order
object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.
2026-08-19 09:41:39 +05:30
pj 4768b8048b fix(spec): and text with the keys written beside it
a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.
2026-08-19 09:41:31 +05:30
pj f8a025eb83 docs(folio): state that android recipes need ANDROID_DEVICE 2026-08-18 20:47:27 +05:30
pj 8abe2c1ca0 fix(folio): refuse to install and fuzz a device nobody named
adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.
2026-08-18 20:47:24 +05:30
pj 3a79f9f0e7 docs(manual): state what the other boolean state selectors match 2026-08-18 20:44:19 +05:30
pj 84baf18ee9 fix(spec): keep a selector out of the head subtree
the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.
2026-08-18 20:44:15 +05:30
pj 5404909e12 test(chrome): compare checked, selected and focused across both producers
the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.
2026-08-18 20:43:46 +05:30
pj 654c0afd90 test(chrome): resolve the five state selectors on both matchers
the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.
2026-08-18 20:43:40 +05:30
pj c8ca350115 fix(chrome): state every boolean flag the dump can state
internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.
2026-08-18 20:43:29 +05:30
pj d177726c4e fix(spec): keep a tag selector valid beside another key
a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.
2026-08-18 20:43:20 +05:30
pj f7135b81d8 fix(spec): a boolean state selector names what the live element reports
clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.
2026-08-18 20:43:09 +05:30
pj 4bcd82a8a6 fix(analyze): write an undefined paired p-value as null, not as NaN
A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.
2026-08-18 20:24:10 +05:30
pj 8fad1937bb fix(conformance): g4 reads a doubling out of what the field grew by
The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.
2026-08-18 20:22:39 +05:30
pj 8e8eb621c3 refactor(analyze): hoist the sign test's loop bound 2026-08-18 20:20:49 +05:30
pj 37e6298992 test(conformance): g4 sees a doubling appended to what the field held
Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.
2026-08-18 20:20:47 +05:30
pj 132ed25313 docs(skills): quote the summary line the runner prints now
The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.
2026-08-18 20:20:04 +05:30
pj a257a91988 docs(manual): exit 1 also means a run that holds no verdict
And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.
2026-08-18 20:20:04 +05:30
pj dd625de068 docs(analyze): name the tests the tool actually runs
The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.
2026-08-18 20:17:39 +05:30
pj 1008402558 fix(analyze): score a seed pair by which run outlived the other
The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.
2026-08-18 20:16:40 +05:30
pj e2f9864201 fix(spec): keep a secure selector valid beside another key
a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.
2026-08-18 20:16:03 +05:30
pj c8d52cb7d5 test(browser): the exit code a dead run and a violated one actually leave
Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.
2026-08-18 20:14:59 +05:30
pj f4b9ff1468 fix(conformance): g4 skips an input that names no field
jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.
2026-08-18 20:14:28 +05:30
pj faa1a33a68 test(conformance): g4 keeps checking past an input typed at coordinates
An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.
2026-08-18 20:14:14 +05:30
pj b9a48cc36f docs(manual): state what a secure selector matches 2026-08-18 20:14:07 +05:30
pj 9a920444b3 test(chrome): compare the secure fact across both producers
it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.
2026-08-18 20:14:07 +05:30
pj 6aecccd2d6 test(chrome): resolve the secure selector on both matchers
the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.
2026-08-18 20:14:07 +05:30
pj b0895ef179 fix(spec): a secure selector names the password field on web
secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.
2026-08-18 20:14:01 +05:30
pj 8414953135 fix(conformance): g4 reads a doubling off the observed field value
The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.
2026-08-18 20:13:52 +05:30
pj f0726a1e61 fix(analyze): compare arms on censored runs, not on flattened step counts
stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.
2026-08-18 20:13:12 +05:30
pj b7ee23942e feat(analyze): add the gehan generalized wilcoxon test
The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.
2026-08-18 20:13:05 +05:30
pj 3526d7a9dd refactor(analyze): open the log-rank up to a weight on the risk set
The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.
2026-08-18 20:12:59 +05:30
pj 74cae01a5c feat(cli): --allow-no-generator-actions
The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.
2026-08-18 20:10:29 +05:30
pj 93614a48b9 fix(testrun): the dead-run refusal gets its own opt-out
--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.
2026-08-18 20:10:29 +05:30
pj aa49c3bc34 fix(testrun): a recorded violation outranks the dead-run refusal
A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.
2026-08-18 20:08:03 +05:30
pj f631415cb6 test(conformance): the g4 fixture holds what a redacted android trace holds
Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.
2026-08-18 20:07:39 +05:30
pj f7b90a9886 docs(folio): say how to point just test at a remote adb server 2026-08-18 19:54:49 +05:30
pj 34f73b82f2 fix(folio): install through adb so a remote adb server works
Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.
2026-08-18 19:54:45 +05:30
pj ca3519ba3a docs(manual): what an action count with no producer means for a rate 2026-08-18 19:45:58 +05:30
pj ca6a05ca5e fix(analyze): mark an action count of unknown provenance in the report 2026-08-18 19:45:58 +05:30
pj 628eb1949a fix(analyze): refuse to compare attributed and unattributed denominators
One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.
2026-08-18 19:45:55 +05:30
pj 67c9de6365 fix(analyze): read how much of a record's action count names no producer
A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.
2026-08-18 19:45:55 +05:30
pj e0180f891a fix(campaign): a record always says how many actions named no producer
An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.
2026-08-18 19:45:49 +05:30
pj 153f857431 fix(defect-identity): degrade a redacted origin action to its selector
The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.
2026-08-18 19:43:28 +05:30
pj 454988fbc8 fix(trace): an action names the generator that produced it
The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.
2026-08-18 17:58:38 +05:30