record what the model picker did, and make both policies see the same actions (#74)

* feat(llmclient): parse usage and the served model

An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record one typed outcome per model-driven step

llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): expose the step a snapshot was observed at

It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a guard-skipped step is no longer a silent log line

The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): record when a chosen action was never dispatched

A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): document llm-calls.jsonl

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(analyze): survival analysis over campaign directories

Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): select the candidate label source

Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(runner): thread the label source to the model picker

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the label source as arm membership

Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --label-source

Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): dedup candidates by what they execute, not how they read

The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): report every action that was chosen and never dispatched

applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): the echo guard admits a repeated description

Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): count dispatched actions, not steps

A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): lower authored actions the way the seeded arm does

The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): decode a container-only scroll

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): carry the container on an authored scroll

serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): both policies must dispatch the same authored action

Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(hierarchy): match identifiers by role prefix

idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.

* feat(chrome): translate idPrefix to a starts-with id match

* feat(spec): match idPrefix in the web runtime

The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.

* feat(sidecar): match idPrefix in the tap-by-selector path

* docs(manual): document the idPrefix selector

* feat(replay-ui): render idPrefix targets as a prefix tag

* fix(spec): read the injected seed per call

Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.

* test(chrome): compare both selector matchers over one live page

Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.

* fix(hierarchy): give id and desc one meaning in both selector forms

The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.

* test(hierarchy): pin both selector forms and the unknown-key report

* feat(verifier): fail the spec on a selector key that cannot match

An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.

* feat(spec): reject an unknown object-selector key in the web runtime

Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.

* test(spec): pin the unknown-key diagnostic to one text

The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.

* fix(spec): match a merged label by its leading name on web too

The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.

* test(chrome): drive the live-page parity test through both selector forms

* docs(manual): document object-selector key rules

* feat(spec): refuse a multi-item authored sampler while enumerating

from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets

Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): abort on a candidate enumeration that refused

Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(spec): refuse a multi-value generator while enumerating

integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(verifier): setup still draws, and the seeded stream is unmoved

Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): value generators are refused under the model policy too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio): enumerate authored targets and values instead of sampling

Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): enumerate authored targets and values, declare the llm generator

The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: minimal changes, self-documenting code, tests as first-class

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): key web attrs by the names the markup writes

attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name a web handle by the same ladder as a tree element

The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): attrs carries raw attribute names on web too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): confirm focus moved before typing

InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): signal a timed-out run so it reaps its sidecar

CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* perf(runner): confirm focus only when another element holds it

Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): record both clocks a run was measured on

Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide per-hour rates by time actually worked

A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(campaign): wait for the trap instead of racing it

The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(ltl): keep the authored window on a step-bounded obligation

reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(ltl): pin that a slow policy does not fail on time alone

Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(spec): guard the step unit on the authoring surface

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): bound the reachability properties by steps

At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): a step bound counts observations, not runner steps

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): make the label source a cell dimension

A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): record the label source in the manifest

A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): name a web field by its hint, not its CSS class

visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): answer clickable for an element reached through ax

The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: every target runs on this machine, so start one rather than skip it

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a selector the tree cannot resolve is not a focus failure

otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit data-testid so both resolvers name the same element

The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name an element only when the selector names it alone

ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): hold the V8 host to the same naming gate

elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): sibling taps reach the driver at their own coordinates

Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): an ambiguous name loses to the coordinates it was built from

Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): record an element-valued extractor instead of dropping it

An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.

* test(verifier): an unrecordable extractor value is reported, not dropped

* test(runner): element-valued extractors reach the trace

* docs(spec-language): say what a trace records for an element-valued extractor

* docs(claude): add delegation and record-keeping sections

delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.

* feat(driver): declare undelivered-action errors and three optional capabilities

ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.

* feat(driver): add escape to the pressKey surface

escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.

* fix(ios): refuse a gesture the screen has no surface under

the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.

* feat(ios): derive scrollable from the snapshot's tree depth

the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.

* fix(sidecar): stop dropping gestures, selectors and keys in silence

a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.

* fix(sidecar): map the driver's refusals onto the gesture errors

OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.

* fix(selectors): resolve text to the innermost match and scan the root in both forms

an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.

* feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag

a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.

* fix(chrome): emit every markup attribute and read checked and selected off the property

the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.

* fix(chrome): scroll a gesture point into view and dispatch trusted input

getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.

* feat(chrome): read the page's exceptions and navigations, and hold the picker state across them

a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.

* feat(trace): version each step and record its logs, exceptions and navigations

a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.

* feat(runner): bound every device call and record the actions that never reached the app

observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.

* feat(verifier): expose extractor names and rebuilt property formulas

an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.

* feat(testrun): expose the seeded bundle a run loaded

BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity.

* feat(tracecorpus): load recorded runs for offline measures

reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.

* refactor(seedspec): move seed spec parsing out of the campaign command

the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.

* feat(analyze): time an event at the step it was detected and report the quartiles

an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.

* feat(analyze): add the seed-paired signed-rank comparison and record the holm family

--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.

* test(analyze): recover planted effects through the tool's own entry point

a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.

* feat(label-coverage): report the addressable share of an app's interactive surface

reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.

* feat(exploration-reach): count the distinct structural states a stored run visited

the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.

* feat(defect-identity): count distinct defects across stored runs

a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.

* feat(oracle-reduction): replay stored traces under four reduced oracles

re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.

* feat(implementation-sweep): run one campaign against every implementation of a requirement

installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.

* feat(corpus-sweep): run one specification against a served corpus of implementations

same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.

* docs(manual): document innermost text matching, escape and the web scroll verb

text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.

* test(browser): assert an uncaught page exception reaches the trace

the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.

* feat(trace): a step can name the precondition it could not meet

A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.

* fix(runner): budget the foreground gate in time, not in polls

Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.

* test(runner): the gate keeps looking until its budget runs out

Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.

* feat(campaign): count the runs that were never in the app

A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.

* docs(triage): name the trace field a run that never started leaves

* fix(selectors): tag names the whole tag, not a substring of it

matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.

* test(chrome): both resolvers agree on tag where a container's name contains its child's

* fix(make): build the binary instead of matching the build directory

build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.

* feat(verifier): expose the property names a loaded spec registered

* feat(testrun): refuse a run against a spec that registers no properties

A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.

* feat(cli): --allow-no-properties opts a run out of the refusal

* docs(cli): document --allow-no-properties

* feat(bundle-check): fail a spec that bundles but registers no properties

* test(bundle-check): cover the zero-property refusal and pin the reported bundle

* feat(folio-web): predicates for counting commits against submit actions

* feat(folio-web): judge one commit per submit over a home-card window

Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.

* fix(folio-web): keep submit live for 400ms after saving

Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.

* feat(confusion-matrix): score the checker against a blind reviewer

Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.

* test(confusion-matrix): reject malformed inputs and keep missing data out of the cells

* test(confusion-matrix): cover cell assignment, precision and recall

* fix(chrome): focus descends into the shadow root

document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.

* test(implementation-sweep): supply the binaries the missing-binary test does not test

resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.

* fix(replay-ui): read data-* attributes by their markup names

The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.

* fix(web-runtime): focus descends into the shadow root here too

The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.

* fix(implementation-sweep): name every missing binary, in flag order

Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.

* fix(chrome): focus follows the caret to the field it types into

Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.

* fix(corpus-sweep): name every missing binary, in flag order

Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.

* fix(web-runtime): a handle answers editable for itself, not its container

isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.

* test(chrome): a hinted field is not named by its css class

The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.

* test(chrome): the handle and the enumeration agree on editable too

The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.

* fix(web-runtime): focus follows the caret to the field it types into

Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.

* fix(campaign): name every missing required flag, in flag order

Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.

* fix(corpus-sweep): name every missing required flag, in flag order

* fix(implementation-sweep): name every missing required flag, in flag order

* fix(confusion-matrix): name every missing required flag, in flag order

* ci: pin the idb-companion tap to the formula the companion is staged from

The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.

* fix(campaign): refuse to start on a device that is not there

A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): quarantine a device that keeps failing fast

Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the device a run executed on

meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* ci: let a restored ios bundle survive make's mtime check

The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.

* fix(confusion-matrix): a campaign that died is missing data, not a true negative

The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.

* fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.

* fix(runner): a source that was asked and handed nothing says so

NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.

* feat(testrun): refuse a run that dispatched none of its actions

Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.

* docs(cli): document --label-source

* docs(spec-language): name the hintText selector's host divergence

The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.

* feat(bundle-check): --allow-no-properties opts out of the refusal

The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.

* test(verifier): an unreadable committed fixture fails, it does not skip

The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.

* fix(testrun): the refusal asks whether the generator drove, not whether anything did

A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.

* feat(runner): the summary says how many steps the generator drove

A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.

* fix(testrun): an ios run records the simulator it executed on

Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.

* fix(campaign): the action count leaves the setup's login out on a model run

Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.

* feat(hierarchy): an element reports whether it masks what is typed into it

ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".

* fix(verifier): a secure field's typed value never reaches the record

A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.

* fix(runner): a secure field's value does not reach state.lastAction either

folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.

* fix(trace): an action names the generator that produced it

The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.

* fix(defect-identity): degrade a redacted origin action to its selector

The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.

* fix(campaign): a record always says how many actions named no producer

An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.

* fix(analyze): read how much of a record's action count names no producer

A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.

* fix(analyze): refuse to compare attributed and unattributed denominators

One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.

* fix(analyze): mark an action count of unknown provenance in the report

* docs(manual): what an action count with no producer means for a rate

* fix(folio): install through adb so a remote adb server works

Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.

* docs(folio): say how to point just test at a remote adb server

* test(conformance): the g4 fixture holds what a redacted android trace holds

Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.

* fix(testrun): a recorded violation outranks the dead-run refusal

A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.

* fix(testrun): the dead-run refusal gets its own opt-out

--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.

* feat(cli): --allow-no-generator-actions

The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.

* refactor(analyze): open the log-rank up to a weight on the risk set

The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.

* feat(analyze): add the gehan generalized wilcoxon test

The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.

* fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.

* fix(conformance): g4 reads a doubling off the observed field value

The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.

* fix(spec): a secure selector names the password field on web

secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.

* test(chrome): resolve the secure selector on both matchers

the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.

* test(chrome): compare the secure fact across both producers

it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.

* docs(manual): state what a secure selector matches

* test(conformance): g4 keeps checking past an input typed at coordinates

An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.

* fix(conformance): g4 skips an input that names no field

jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.

* test(browser): the exit code a dead run and a violated one actually leave

Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.

* fix(spec): keep a secure selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.

* fix(analyze): score a seed pair by which run outlived the other

The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.

* docs(analyze): name the tests the tool actually runs

The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.

* docs(manual): exit 1 also means a run that holds no verdict

And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.

* docs(skills): quote the summary line the runner prints now

The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.

* test(conformance): g4 sees a doubling appended to what the field held

Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.

* refactor(analyze): hoist the sign test's loop bound

* fix(conformance): g4 reads a doubling out of what the field grew by

The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.

* fix(analyze): write an undefined paired p-value as null, not as NaN

A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.

* fix(spec): a boolean state selector names what the live element reports

clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.

* fix(spec): keep a tag selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.

* fix(chrome): state every boolean flag the dump can state

internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.

* test(chrome): resolve the five state selectors on both matchers

the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.

* test(chrome): compare checked, selected and focused across both producers

the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.

* fix(spec): keep a selector out of the head subtree

the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.

* docs(manual): state what the other boolean state selectors match

* fix(folio): refuse to install and fuzz a device nobody named

adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.

* docs(folio): state that android recipes need ANDROID_DEVICE

* fix(spec): and text with the keys written beside it

a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.

* test(spec): pin text against the key beside it in either order

object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.

* test(chrome): compare a compound text selector across both matchers

one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.

* docs(manual): state how text combines with the key beside it

the object selector section said every pair must match without saying
where the innermost rule then lands.

* fix(hierarchy): reach the class attribute through className

className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.

* test(chrome): compare className across both matchers

one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.

* test(spec): pin className and class on the same elements

this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.

* docs(manual): list className among the cross-platform aliases

the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.

* fix(hierarchy): reach the accessible label through every name for it

label and accessibilityLabel aliased onto accessibilityText alone, which
only the ios sidecar writes, and alias expansion is ONE level: the hop
from accessibilityText to content-desc was never taken, so both keys
matched nothing on android and on the chrome dump, which write the fact
under content-desc. ariaLabel and contentDescription aliased onto
nothing at all and matched nothing anywhere.

The web runtime resolves all four against the live DOM, so a selector
naming a field this way found it on one host and no element at all on
the other. The keys are accepted, so no unknown-key error fires, and a
property over the element that was never found passes having checked
nothing.

Each name lists both keys rather than chaining through accessibilityText:
transitive expansion would silently widen every existing key at once.

* fix(hierarchy): reach a web test tag through testTag and testID

Compose for Web writes a test tag as data-testid, which is what the web
runtime resolves both names against. testTag aliased onto the three
identifier keys and not that one, and testID aliased onto nothing at
all, so a tag the web runtime found on every row of a list named no
element here and every property over it passed vacuously.

* fix(spec): resolve the identifier, label and class aliases against the DOM

identifier, accessibilityIdentifier, accessibilityText and elementType
are the names ios writes four facts under, and internal/hierarchy
aliases each onto the key the other producers write. This table listed
none of them, so each fell through to a raw attribute lookup and built
[accessibilityIdentifier="summary_card"], which no element carries.

Every one of them resolved against the dump on the goja host and named
nothing here. The keys are accepted, so no unknown-key error fires, and
a property over the element that was never found passes having checked
nothing.

* fix(spec): an editable or scrollable selector names what this host derives

Both facts are derived from the live element rather than written by the
markup, and matching them as attributes built [editable="true"], which
no page carries. Both resolve against the dump on the goja host, so a
spec naming a field or a scroll container that way found it there and no
element at all here, with no unknown-key error to say so.

Each reads the same function the fact is derived with, so a selector
cannot name an element this host calls something else: the handle, the
picker's target list and the editable selector all go through
isEditable, and scrollable reads the overflow test collectTargets reads.

scrollable false names nothing rather than every element that does not
scroll: both producers state the fact only where it holds, the way an
element that is no field at all answers to neither value of secure.

* test(chrome): compare the alias keys and the two derived facts

one page, both resolvers, the ten names that resolved on one host only.
each alias is asked beside the key it resolves through, so the pair is
pinned to the same elements rather than each to itself.

the page grows a container that overflows its box and a neighbour that
does not, because scrollable is derived from the box: without one the
only scrolling element on the page is the document root, whose answer
moves with the window.

* fix(hierarchy): bounds is a raw attribute, not a cross-platform key

Every native dump writes the rectangle out as a string under bounds, and
no DOM element carries an attribute of that name, so the key resolved
against the dump and matched nothing on web on every page there is. It
is accepted, so no unknown-key error said so, and no mapping can be
invented for it: there is no DOM fact to map it to.

Off the accepted list the web runtime raises the unknown-key error
instead of matching nothing in silence, and the key still resolves
wherever a producer writes it, through the escape hatch every other raw
attribute already uses: a key some element carries is a key that can
match, on both sides.

* docs(manual): state which attribute each alias reads, and what bounds is

the table listed neither name for the accessible label that a web page
writes, nor the key a web test tag lands on, and said nothing about
elementType. editable and scrollable are boolean states like the rest,
and scrollable is the one of them the platforms state only where it
holds. bounds is a raw driver attribute rather than an accepted key.

* docs(hierarchy): the package doc names every key an alias reaches

it described the alias table as it stood before the label and test-tag
names reached the keys android and web write, and said nothing about
expansion being one level deep, which is why each name has to list every
key rather than hop through another alias.

* fix(spec): a hint selector names the ladder both producers derive

hintText and placeholderValue are the accessible-name ladder, derived
from the live element, and compiling them to [placeholder="..."] made
them name the wrong field or none at all. A field labelled by an
aria-label or a bound <label> carries no placeholder, so it resolved
against the dump on the goja host and reached nothing here; one carrying
both answered to its placeholder here where the dump answers to its
aria-label, which lands a find on an element nobody named.

Both keys read the same fieldHint elementHandle and the hierarchy dump
(internal/driver/chrome/driver.go) derive the fact with, so a selector
cannot name a field this host calls something else. An empty hint names
nothing rather than everything that is no field: both producers write
the fact only where the ladder answered.

placeholder stays the attribute the markup writes, which is what the
dump carries under that name too, so a field whose hint is something
else still answers to it on both hosts.

* fix(chrome): a hint target is not tapped by the placeholder attribute

TapSelector is a third resolver, and it built [placeholder="..."] for
hintText and placeholderValue too. Now that both matchers read the
accessible-name ladder, that CSS names a field whose hint is its
aria-label and whose placeholder happens to carry the value, which is an
element neither matcher named.

No CSS says what the ladder says, so both keys fall through to a match
that reaches nothing and the step fails naming the selector, the way
every other derived key in this file already does. A selector reaches
here only where the dump resolved it to no coordinates at all.

* test(chrome): compare the hint keys and placeholder across both matchers

The page gains four fields that differ one rung at a time: a bound
label, a placeholder, a placeholder an aria-label outranks, and the name
the form gives the field. Only the placeholder rung was reachable
before, so hintText and placeholderValue named a field on the goja host
and no element at all on web for the other three, and named the field
here and nothing there for the rung the ladder passed over.

placeholder was measured empty on both hosts because nothing on the page
carried the attribute, which said nothing about it. It now names the
field the markup wrote it on and not the field whose hint is its
aria-label.

The third resolver reads the same selectors: what TranslateStringSelector
builds for a hint key has to match nothing over CDP rather than the field
carrying the value as a placeholder.

* docs(manual): both hosts read the hint ladder, placeholder is the attribute

The web section said the hintText key does not read the ladder on both
hosts and told authors to select such a field by attrs.hintText instead.
Both hosts read it now, so that instruction is gone rather than left
standing beside a newer sentence.

placeholder is stated as the attribute the markup writes and nothing
more, the tap path is stated as failing by name where no CSS says what
the ladder says, and the alias table gains the row it was missing.
This commit is contained in:
pj authored and GitHub committed 2026-08-19 11:07:49 +05:30
1 parent 95a19fcc1e
commit 9b4ff5f247
243 files changed
+33472 -1062

No files matched your search

+125
View File
@@ -5,9 +5,12 @@ import (
"context"
"fmt"
"io"
"maps"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"time"
"github.com/priyanshujain/sanderling/internal/android"
@@ -62,8 +65,21 @@ type Options struct {
// recorded violations as a ViolationsError, so a caller (CI) can tell
// "the run found the bug" from "the run finished clean".
ExitOnViolation bool
// AllowNoProperties lets a run proceed against a spec that registers no
// properties. The extraction sweeps pass it: they measure what a spec can
// read and report no detection count. Every other run without it is a false
// green.
AllowNoProperties bool
// AllowNoGeneratorActions lets a run finish having never been driven by its
// action generator. The exploration-reach sweeps pass it, because "the
// generator reached nothing on this build" is their measurement rather than
// a broken run.
AllowNoGeneratorActions bool
// Generator selects the action picker: "llm" or the default seeded picker.
Generator string
// LabelSource selects how candidates are named to the model picker, and is
// recorded in meta.json as part of the run's cell.
LabelSource string
// iosUDID, iosIsSimulator, and iosCoreDeviceID are filled by Execute after
// resolving the iOS target, then read by buildDriver to choose the simulator
@@ -79,6 +95,15 @@ type Options struct {
// recorded only when the LLM picker is the one that will actually run, so a
// spec that declares generator = llm() but is run under the seeded picker does
// not label its trace with a model it never called.
// runDevice reads whichever flag named the hardware for this platform. An ios
// run is selected with --ios-device and leaves --device empty.
func runDevice(options Options) string {
if options.Platform == "ios" {
return options.IosDevice
}
return options.Device
}
func buildRunMeta(options Options, bundleSHA256 string, seed int64, host string, llmConfig verifier.LLMConfig, hasLLMConfig bool) trace.Meta {
meta := trace.Meta{
Seed: seed,
@@ -90,9 +115,11 @@ func buildRunMeta(options Options, bundleSHA256 string, seed int64, host string,
SanderlingVersion: "0.0.1",
Arm: options.Arm,
Generator: options.Generator,
LabelSource: options.LabelSource,
MaxSteps: options.MaxSteps,
DurationMillis: options.Duration.Milliseconds(),
Host: host,
Device: runDevice(options),
}
if options.Generator == "llm" && hasLLMConfig {
meta.Model = llmConfig.Model
@@ -186,6 +213,9 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
if err := verifierInstance.Load(string(bundle.JavaScript)); err != nil {
return fmt.Errorf("load spec: %w", err)
}
if !options.AllowNoProperties && len(verifierInstance.PropertyNames()) == 0 {
return NoPropertiesError{Spec: options.Spec}
}
fmt.Fprintln(stdout, "spec loaded into verifier")
runDirectory := filepath.Join(options.Output, time.Now().UTC().Format("20060102-150405"))
@@ -222,6 +252,7 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
TraceWriter: traceWriter,
Logger: newProgressLogger(stdout),
Generator: options.Generator,
LabelSource: options.LabelSource,
StopOnViolation: options.ExitOnViolation,
})
@@ -247,6 +278,19 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
// because it holds no verdict to report. The threshold is every step and not a
// fraction of them: a screen that composes now and then costs a healthy android
// run a step or two, and a check that fired on those would be red on every run.
//
// A run whose generator dispatched no action fails on the same grounds, and the
// threshold is zero for the same reason: a generator with nothing to offer on
// some screens is ordinary, one with nothing to offer on every screen of a whole
// run drove nothing. The count is the generator's alone because a spec's setup
// drives the app before the generator is consulted, so a login that ran leaves
// dispatched actions behind whatever the generator then did.
//
// A recorded violation carries such a run through whether or not
// --exit-on-violation was passed: the refusal exists because a run with no
// verdict must not read as a clean one, and a run that recorded a violation
// holds a verdict. Campaigns pass no flags, so refusing them there would write
// exit_code 1 and lose a real detection to the analysis as missing data.
func runOutcome(options Options, summary runner.Summary) error {
if summary.Steps > 0 && summary.SkippedVerification == summary.Steps {
return VacuousRunError{Steps: summary.Steps}
@@ -254,6 +298,13 @@ func runOutcome(options Options, summary runner.Summary) error {
if options.ExitOnViolation && len(summary.Violations) > 0 {
return ViolationsError{Count: len(summary.Violations)}
}
if !options.AllowNoGeneratorActions && len(summary.Violations) == 0 &&
summary.Steps > 0 && summary.GeneratorActions == 0 {
return NoGeneratorActionsError{
Steps: summary.Steps,
SkippedActions: summary.SkippedActions,
}
}
return nil
}
@@ -269,6 +320,23 @@ func (e ViolationsError) Error() string {
return fmt.Sprintf("%d violation record(s)", e.Count)
}
// BundleSpec produces the goja bundle a run of this spec loaded, seeded as that
// run was. An offline replay of the run's trace has to load the same JavaScript
// the run did, and the seed is one of the bundle's defines, so it is part of
// the bundle's identity.
func BundleSpec(specPath string, seed int64) (bundler.Result, error) {
inputs, err := prepareBundleInputs(Options{Spec: specPath, Seed: seed})
if err != nil {
return bundler.Result{}, err
}
return bundler.Bundle(bundler.Options{
EntryFile: specPath,
RuntimeFile: inputs.gojaRuntimePath,
Defines: inputs.defines,
Aliases: inputs.aliases,
})
}
// VacuousRunError reports a run in which no step reached the verifier, so no
// property ever judged anything. It is not a clean run and it is not a found
// bug: it is a run that produced no evidence either way, and the absence of
@@ -285,6 +353,63 @@ func (e VacuousRunError) Error() string {
e.Steps)
}
// NoPropertiesError reports a spec that bundled and loaded cleanly and holds no
// properties. Nothing is broken: the run would drive the app, fill a trace and
// report no violations having judged nothing, and that green says as much about
// the app as an unplugged meter says about a wire. It stays untyped to the CLI's
// violation path like VacuousRunError, so it exits 1 as a run that cannot
// produce a verdict rather than 2.
type NoPropertiesError struct {
Spec string
}
func (e NoPropertiesError) Error() string {
return fmt.Sprintf(
"%s bundled and loaded into the verifier cleanly and registers no properties: "+
"nothing is wrong with the spec and nothing is wrong with the run, but this run "+
"would check nothing and report no violations. Pass --allow-no-properties for a "+
"run that measures what the spec extracts instead of judging the app",
e.Spec)
}
// NoGeneratorActionsError reports a run not one of whose steps was driven by the
// action generator. Setup can put the app in position, but only the generator
// explores it, so every screen this run judged was one setup left it on and its
// empty violation list says as much about the app as a spec with no properties
// would: the run observed, judged the same state over and over, and exercised
// nothing. It stays untyped to the CLI's violation path like VacuousRunError, so
// it exits 1 as a run that holds no verdict rather than 2.
type NoGeneratorActionsError struct {
Steps int
// SkippedActions is the runner's per-reason count of actions that never
// reached the app, which is where the cause is: a picker with no candidate
// reads differently from one whose every model call failed.
SkippedActions map[string]int
}
func (e NoGeneratorActionsError) Error() string {
return fmt.Sprintf(
"%d step(s) ran and the action generator drove the app in none of them: whatever "+
"the spec's setup did to get the app into position, nothing explored it from "+
"there, so the run judged one screen over and over and its violation count "+
"says nothing about the rest of the app%s. Pass --allow-no-generator-actions "+
"for a run that measures where a generator reaches instead of judging the app",
e.Steps, skipReasonSuffix(e.SkippedActions))
}
// skipReasonSuffix renders the skip-reason tally as a trailing clause, empty
// when the run recorded none.
func skipReasonSuffix(skipped map[string]int) string {
if len(skipped) == 0 {
return ""
}
reasons := make([]string, 0, len(skipped))
for _, reason := range slices.Sorted(maps.Keys(skipped)) {
reasons = append(reasons, fmt.Sprintf("%s %d", reason, skipped[reason]))
}
return ". Actions that never reached the app: " + strings.Join(reasons, ", ")
}
// bundleInputs holds the pre-driver assembly: alias map, seed, esbuild defines,
// and the resolved spec-API/goja-runtime paths the bundler consumes.
type bundleInputs struct {