Files
sanderling/internal/runner/runner_test.go
pj 9b4ff5f247 record what the model picker did, and make both policies see the same actions (#74)
* feat(llmclient): parse usage and the served model

An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record one typed outcome per model-driven step

llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): expose the step a snapshot was observed at

It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a guard-skipped step is no longer a silent log line

The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): record when a chosen action was never dispatched

A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): document llm-calls.jsonl

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(analyze): survival analysis over campaign directories

Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): select the candidate label source

Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(runner): thread the label source to the model picker

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the label source as arm membership

Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --label-source

Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): dedup candidates by what they execute, not how they read

The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): report every action that was chosen and never dispatched

applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): the echo guard admits a repeated description

Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): count dispatched actions, not steps

A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): lower authored actions the way the seeded arm does

The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): decode a container-only scroll

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): carry the container on an authored scroll

serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): both policies must dispatch the same authored action

Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(hierarchy): match identifiers by role prefix

idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.

* feat(chrome): translate idPrefix to a starts-with id match

* feat(spec): match idPrefix in the web runtime

The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.

* feat(sidecar): match idPrefix in the tap-by-selector path

* docs(manual): document the idPrefix selector

* feat(replay-ui): render idPrefix targets as a prefix tag

* fix(spec): read the injected seed per call

Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.

* test(chrome): compare both selector matchers over one live page

Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.

* fix(hierarchy): give id and desc one meaning in both selector forms

The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.

* test(hierarchy): pin both selector forms and the unknown-key report

* feat(verifier): fail the spec on a selector key that cannot match

An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.

* feat(spec): reject an unknown object-selector key in the web runtime

Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.

* test(spec): pin the unknown-key diagnostic to one text

The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.

* fix(spec): match a merged label by its leading name on web too

The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.

* test(chrome): drive the live-page parity test through both selector forms

* docs(manual): document object-selector key rules

* feat(spec): refuse a multi-item authored sampler while enumerating

from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets

Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): abort on a candidate enumeration that refused

Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(spec): refuse a multi-value generator while enumerating

integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(verifier): setup still draws, and the seeded stream is unmoved

Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): value generators are refused under the model policy too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio): enumerate authored targets and values instead of sampling

Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): enumerate authored targets and values, declare the llm generator

The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: minimal changes, self-documenting code, tests as first-class

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): key web attrs by the names the markup writes

attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name a web handle by the same ladder as a tree element

The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): attrs carries raw attribute names on web too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): confirm focus moved before typing

InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): signal a timed-out run so it reaps its sidecar

CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* perf(runner): confirm focus only when another element holds it

Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): record both clocks a run was measured on

Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide per-hour rates by time actually worked

A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(campaign): wait for the trap instead of racing it

The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(ltl): keep the authored window on a step-bounded obligation

reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(ltl): pin that a slow policy does not fail on time alone

Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(spec): guard the step unit on the authoring surface

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): bound the reachability properties by steps

At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): a step bound counts observations, not runner steps

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): make the label source a cell dimension

A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): record the label source in the manifest

A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): name a web field by its hint, not its CSS class

visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): answer clickable for an element reached through ax

The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: every target runs on this machine, so start one rather than skip it

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a selector the tree cannot resolve is not a focus failure

otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit data-testid so both resolvers name the same element

The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name an element only when the selector names it alone

ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): hold the V8 host to the same naming gate

elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): sibling taps reach the driver at their own coordinates

Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): an ambiguous name loses to the coordinates it was built from

Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): record an element-valued extractor instead of dropping it

An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.

* test(verifier): an unrecordable extractor value is reported, not dropped

* test(runner): element-valued extractors reach the trace

* docs(spec-language): say what a trace records for an element-valued extractor

* docs(claude): add delegation and record-keeping sections

delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.

* feat(driver): declare undelivered-action errors and three optional capabilities

ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.

* feat(driver): add escape to the pressKey surface

escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.

* fix(ios): refuse a gesture the screen has no surface under

the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.

* feat(ios): derive scrollable from the snapshot's tree depth

the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.

* fix(sidecar): stop dropping gestures, selectors and keys in silence

a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.

* fix(sidecar): map the driver's refusals onto the gesture errors

OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.

* fix(selectors): resolve text to the innermost match and scan the root in both forms

an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.

* feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag

a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.

* fix(chrome): emit every markup attribute and read checked and selected off the property

the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.

* fix(chrome): scroll a gesture point into view and dispatch trusted input

getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.

* feat(chrome): read the page's exceptions and navigations, and hold the picker state across them

a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.

* feat(trace): version each step and record its logs, exceptions and navigations

a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.

* feat(runner): bound every device call and record the actions that never reached the app

observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.

* feat(verifier): expose extractor names and rebuilt property formulas

an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.

* feat(testrun): expose the seeded bundle a run loaded

BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity.

* feat(tracecorpus): load recorded runs for offline measures

reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.

* refactor(seedspec): move seed spec parsing out of the campaign command

the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.

* feat(analyze): time an event at the step it was detected and report the quartiles

an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.

* feat(analyze): add the seed-paired signed-rank comparison and record the holm family

--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.

* test(analyze): recover planted effects through the tool's own entry point

a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.

* feat(label-coverage): report the addressable share of an app's interactive surface

reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.

* feat(exploration-reach): count the distinct structural states a stored run visited

the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.

* feat(defect-identity): count distinct defects across stored runs

a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.

* feat(oracle-reduction): replay stored traces under four reduced oracles

re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.

* feat(implementation-sweep): run one campaign against every implementation of a requirement

installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.

* feat(corpus-sweep): run one specification against a served corpus of implementations

same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.

* docs(manual): document innermost text matching, escape and the web scroll verb

text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.

* test(browser): assert an uncaught page exception reaches the trace

the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.

* feat(trace): a step can name the precondition it could not meet

A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.

* fix(runner): budget the foreground gate in time, not in polls

Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.

* test(runner): the gate keeps looking until its budget runs out

Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.

* feat(campaign): count the runs that were never in the app

A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.

* docs(triage): name the trace field a run that never started leaves

* fix(selectors): tag names the whole tag, not a substring of it

matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.

* test(chrome): both resolvers agree on tag where a container's name contains its child's

* fix(make): build the binary instead of matching the build directory

build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.

* feat(verifier): expose the property names a loaded spec registered

* feat(testrun): refuse a run against a spec that registers no properties

A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.

* feat(cli): --allow-no-properties opts a run out of the refusal

* docs(cli): document --allow-no-properties

* feat(bundle-check): fail a spec that bundles but registers no properties

* test(bundle-check): cover the zero-property refusal and pin the reported bundle

* feat(folio-web): predicates for counting commits against submit actions

* feat(folio-web): judge one commit per submit over a home-card window

Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.

* fix(folio-web): keep submit live for 400ms after saving

Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.

* feat(confusion-matrix): score the checker against a blind reviewer

Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.

* test(confusion-matrix): reject malformed inputs and keep missing data out of the cells

* test(confusion-matrix): cover cell assignment, precision and recall

* fix(chrome): focus descends into the shadow root

document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.

* test(implementation-sweep): supply the binaries the missing-binary test does not test

resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.

* fix(replay-ui): read data-* attributes by their markup names

The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.

* fix(web-runtime): focus descends into the shadow root here too

The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.

* fix(implementation-sweep): name every missing binary, in flag order

Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.

* fix(chrome): focus follows the caret to the field it types into

Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.

* fix(corpus-sweep): name every missing binary, in flag order

Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.

* fix(web-runtime): a handle answers editable for itself, not its container

isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.

* test(chrome): a hinted field is not named by its css class

The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.

* test(chrome): the handle and the enumeration agree on editable too

The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.

* fix(web-runtime): focus follows the caret to the field it types into

Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.

* fix(campaign): name every missing required flag, in flag order

Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.

* fix(corpus-sweep): name every missing required flag, in flag order

* fix(implementation-sweep): name every missing required flag, in flag order

* fix(confusion-matrix): name every missing required flag, in flag order

* ci: pin the idb-companion tap to the formula the companion is staged from

The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.

* fix(campaign): refuse to start on a device that is not there

A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): quarantine a device that keeps failing fast

Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the device a run executed on

meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* ci: let a restored ios bundle survive make's mtime check

The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.

* fix(confusion-matrix): a campaign that died is missing data, not a true negative

The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.

* fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.

* fix(runner): a source that was asked and handed nothing says so

NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.

* feat(testrun): refuse a run that dispatched none of its actions

Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.

* docs(cli): document --label-source

* docs(spec-language): name the hintText selector's host divergence

The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.

* feat(bundle-check): --allow-no-properties opts out of the refusal

The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.

* test(verifier): an unreadable committed fixture fails, it does not skip

The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.

* fix(testrun): the refusal asks whether the generator drove, not whether anything did

A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.

* feat(runner): the summary says how many steps the generator drove

A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.

* fix(testrun): an ios run records the simulator it executed on

Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.

* fix(campaign): the action count leaves the setup's login out on a model run

Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.

* feat(hierarchy): an element reports whether it masks what is typed into it

ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".

* fix(verifier): a secure field's typed value never reaches the record

A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.

* fix(runner): a secure field's value does not reach state.lastAction either

folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.

* fix(trace): an action names the generator that produced it

The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.

* fix(defect-identity): degrade a redacted origin action to its selector

The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.

* fix(campaign): a record always says how many actions named no producer

An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.

* fix(analyze): read how much of a record's action count names no producer

A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.

* fix(analyze): refuse to compare attributed and unattributed denominators

One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.

* fix(analyze): mark an action count of unknown provenance in the report

* docs(manual): what an action count with no producer means for a rate

* fix(folio): install through adb so a remote adb server works

Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.

* docs(folio): say how to point just test at a remote adb server

* test(conformance): the g4 fixture holds what a redacted android trace holds

Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.

* fix(testrun): a recorded violation outranks the dead-run refusal

A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.

* fix(testrun): the dead-run refusal gets its own opt-out

--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.

* feat(cli): --allow-no-generator-actions

The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.

* refactor(analyze): open the log-rank up to a weight on the risk set

The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.

* feat(analyze): add the gehan generalized wilcoxon test

The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.

* fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.

* fix(conformance): g4 reads a doubling off the observed field value

The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.

* fix(spec): a secure selector names the password field on web

secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.

* test(chrome): resolve the secure selector on both matchers

the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.

* test(chrome): compare the secure fact across both producers

it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.

* docs(manual): state what a secure selector matches

* test(conformance): g4 keeps checking past an input typed at coordinates

An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.

* fix(conformance): g4 skips an input that names no field

jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.

* test(browser): the exit code a dead run and a violated one actually leave

Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.

* fix(spec): keep a secure selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.

* fix(analyze): score a seed pair by which run outlived the other

The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.

* docs(analyze): name the tests the tool actually runs

The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.

* docs(manual): exit 1 also means a run that holds no verdict

And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.

* docs(skills): quote the summary line the runner prints now

The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.

* test(conformance): g4 sees a doubling appended to what the field held

Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.

* refactor(analyze): hoist the sign test's loop bound

* fix(conformance): g4 reads a doubling out of what the field grew by

The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.

* fix(analyze): write an undefined paired p-value as null, not as NaN

A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.

* fix(spec): a boolean state selector names what the live element reports

clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.

* fix(spec): keep a tag selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.

* fix(chrome): state every boolean flag the dump can state

internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.

* test(chrome): resolve the five state selectors on both matchers

the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.

* test(chrome): compare checked, selected and focused across both producers

the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.

* fix(spec): keep a selector out of the head subtree

the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.

* docs(manual): state what the other boolean state selectors match

* fix(folio): refuse to install and fuzz a device nobody named

adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.

* docs(folio): state that android recipes need ANDROID_DEVICE

* fix(spec): and text with the keys written beside it

a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.

* test(spec): pin text against the key beside it in either order

object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.

* test(chrome): compare a compound text selector across both matchers

one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.

* docs(manual): state how text combines with the key beside it

the object selector section said every pair must match without saying
where the innermost rule then lands.

* fix(hierarchy): reach the class attribute through className

className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.

* test(chrome): compare className across both matchers

one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.

* test(spec): pin className and class on the same elements

this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.

* docs(manual): list className among the cross-platform aliases

the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.

* fix(hierarchy): reach the accessible label through every name for it

label and accessibilityLabel aliased onto accessibilityText alone, which
only the ios sidecar writes, and alias expansion is ONE level: the hop
from accessibilityText to content-desc was never taken, so both keys
matched nothing on android and on the chrome dump, which write the fact
under content-desc. ariaLabel and contentDescription aliased onto
nothing at all and matched nothing anywhere.

The web runtime resolves all four against the live DOM, so a selector
naming a field this way found it on one host and no element at all on
the other. The keys are accepted, so no unknown-key error fires, and a
property over the element that was never found passes having checked
nothing.

Each name lists both keys rather than chaining through accessibilityText:
transitive expansion would silently widen every existing key at once.

* fix(hierarchy): reach a web test tag through testTag and testID

Compose for Web writes a test tag as data-testid, which is what the web
runtime resolves both names against. testTag aliased onto the three
identifier keys and not that one, and testID aliased onto nothing at
all, so a tag the web runtime found on every row of a list named no
element here and every property over it passed vacuously.

* fix(spec): resolve the identifier, label and class aliases against the DOM

identifier, accessibilityIdentifier, accessibilityText and elementType
are the names ios writes four facts under, and internal/hierarchy
aliases each onto the key the other producers write. This table listed
none of them, so each fell through to a raw attribute lookup and built
[accessibilityIdentifier="summary_card"], which no element carries.

Every one of them resolved against the dump on the goja host and named
nothing here. The keys are accepted, so no unknown-key error fires, and
a property over the element that was never found passes having checked
nothing.

* fix(spec): an editable or scrollable selector names what this host derives

Both facts are derived from the live element rather than written by the
markup, and matching them as attributes built [editable="true"], which
no page carries. Both resolve against the dump on the goja host, so a
spec naming a field or a scroll container that way found it there and no
element at all here, with no unknown-key error to say so.

Each reads the same function the fact is derived with, so a selector
cannot name an element this host calls something else: the handle, the
picker's target list and the editable selector all go through
isEditable, and scrollable reads the overflow test collectTargets reads.

scrollable false names nothing rather than every element that does not
scroll: both producers state the fact only where it holds, the way an
element that is no field at all answers to neither value of secure.

* test(chrome): compare the alias keys and the two derived facts

one page, both resolvers, the ten names that resolved on one host only.
each alias is asked beside the key it resolves through, so the pair is
pinned to the same elements rather than each to itself.

the page grows a container that overflows its box and a neighbour that
does not, because scrollable is derived from the box: without one the
only scrolling element on the page is the document root, whose answer
moves with the window.

* fix(hierarchy): bounds is a raw attribute, not a cross-platform key

Every native dump writes the rectangle out as a string under bounds, and
no DOM element carries an attribute of that name, so the key resolved
against the dump and matched nothing on web on every page there is. It
is accepted, so no unknown-key error said so, and no mapping can be
invented for it: there is no DOM fact to map it to.

Off the accepted list the web runtime raises the unknown-key error
instead of matching nothing in silence, and the key still resolves
wherever a producer writes it, through the escape hatch every other raw
attribute already uses: a key some element carries is a key that can
match, on both sides.

* docs(manual): state which attribute each alias reads, and what bounds is

the table listed neither name for the accessible label that a web page
writes, nor the key a web test tag lands on, and said nothing about
elementType. editable and scrollable are boolean states like the rest,
and scrollable is the one of them the platforms state only where it
holds. bounds is a raw driver attribute rather than an accepted key.

* docs(hierarchy): the package doc names every key an alias reaches

it described the alias table as it stood before the label and test-tag
names reached the keys android and web write, and said nothing about
expansion being one level deep, which is why each name has to list every
key rather than hop through another alias.

* fix(spec): a hint selector names the ladder both producers derive

hintText and placeholderValue are the accessible-name ladder, derived
from the live element, and compiling them to [placeholder="..."] made
them name the wrong field or none at all. A field labelled by an
aria-label or a bound <label> carries no placeholder, so it resolved
against the dump on the goja host and reached nothing here; one carrying
both answered to its placeholder here where the dump answers to its
aria-label, which lands a find on an element nobody named.

Both keys read the same fieldHint elementHandle and the hierarchy dump
(internal/driver/chrome/driver.go) derive the fact with, so a selector
cannot name a field this host calls something else. An empty hint names
nothing rather than everything that is no field: both producers write
the fact only where the ladder answered.

placeholder stays the attribute the markup writes, which is what the
dump carries under that name too, so a field whose hint is something
else still answers to it on both hosts.

* fix(chrome): a hint target is not tapped by the placeholder attribute

TapSelector is a third resolver, and it built [placeholder="..."] for
hintText and placeholderValue too. Now that both matchers read the
accessible-name ladder, that CSS names a field whose hint is its
aria-label and whose placeholder happens to carry the value, which is an
element neither matcher named.

No CSS says what the ladder says, so both keys fall through to a match
that reaches nothing and the step fails naming the selector, the way
every other derived key in this file already does. A selector reaches
here only where the dump resolved it to no coordinates at all.

* test(chrome): compare the hint keys and placeholder across both matchers

The page gains four fields that differ one rung at a time: a bound
label, a placeholder, a placeholder an aria-label outranks, and the name
the form gives the field. Only the placeholder rung was reachable
before, so hintText and placeholderValue named a field on the goja host
and no element at all on web for the other three, and named the field
here and nothing there for the rung the ladder passed over.

placeholder was measured empty on both hosts because nothing on the page
carried the attribute, which said nothing about it. It now names the
field the markup wrote it on and not the field whose hint is its
aria-label.

The third resolver reads the same selectors: what TranslateStringSelector
builds for a hint key has to match nothing over CDP rather than the field
carrying the value as a placeholder.

* docs(manual): both hosts read the hint ladder, placeholder is the attribute

The web section said the hintText key does not read the ladder on both
hosts and told authors to select such a field by attrs.hintText instead.
Both hosts read it now, so that instruction is gone rather than left
standing beside a newer sentence.

placeholder is stated as the attribute the markup writes and nothing
more, the tap path is stated as failing by name where no CSS says what
the ladder says, and the alias table gains the row it was missing.
2026-08-19 11:07:49 +05:30

3182 lines
110 KiB
Go

package runner
import (
"bufio"
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log/slog"
"maps"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/bundler"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
const fixtureSpec = `
import { actions, always, extract, Tap } from "@sanderling/spec";
const balance = extract(state => state.snapshots.balance ?? 0);
globalThis.properties = {
balanceNonNegative: always(() => balance.current >= 0),
};
globalThis.actions = actions(() => [Tap({ on: "id:next" })]);
`
// zeroWaitSpec's only action is a Wait the runner cannot perform, so every step
// chooses an action that never reaches the device.
const zeroWaitSpec = `
import { actions, always, Wait } from "@sanderling/spec";
globalThis.properties = {
alwaysHolds: always(() => true),
};
globalThis.actions = actions(() => [Wait({ durationMillis: 0 })]);
`
// absentSelectorSpec names an element the tree never holds, so every step
// dispatches by selector and the driver is the layer that finds nothing.
const absentSelectorSpec = `
import { actions, always, Tap } from "@sanderling/spec";
globalThis.properties = {
alwaysHolds: always(() => true),
};
globalThis.actions = actions(() => [Tap({ on: "id:absent" })]);
`
// noActionSpec's generator offers nothing on any screen, so every step asks the
// source for an action and is handed none.
const noActionSpec = `
import { actions, always } from "@sanderling/spec";
globalThis.properties = {
alwaysHolds: always(() => true),
};
globalThis.actions = actions(() => []);
`
const violationSpec = `
import { actions, always } from "@sanderling/spec";
globalThis.properties = {
balanceNonNegative: always(() => false),
};
globalThis.actions = actions(() => []);
`
type harness struct {
mock *mockdriver.Driver
verifier *verifier.Verifier
writer *trace.Writer
directory string
}
func newHarness(t *testing.T) *harness {
return newHarnessWithSpec(t, fixtureSpec)
}
func fastFocusSettle(t *testing.T) {
prev := focusTapSettle
focusTapSettle = time.Millisecond
t.Cleanup(func() { focusTapSettle = prev })
}
// fastForegroundGate shrinks the startup gate's wall-clock budget so a test that
// drives it to exhaustion takes milliseconds. The budget stays a duration, which
// is the property under test.
func fastForegroundGate(t *testing.T) {
budget, interval := foregroundReadyBudget, foregroundPollInterval
foregroundReadyBudget = 200 * time.Millisecond
foregroundPollInterval = time.Millisecond
t.Cleanup(func() {
foregroundReadyBudget, foregroundPollInterval = budget, interval
})
}
func mustDispatch(t *testing.T, drv driver.DeviceDriver, action verifier.Action, tree *hierarchy.Tree) {
t.Helper()
skipped, err := applyAction(context.Background(), drv, action, tree)
if err != nil {
t.Fatalf("applyAction: %v", err)
}
if skipped != "" {
t.Fatalf("applyAction reported %q: the action never reached the driver", skipped)
}
}
// bundleSpec compiles an authored TS spec with the goja runtime entry so the
// loaded bundle installs __sanderlingNextAction__ (the shared picker).
func bundleSpec(t *testing.T, specSource string) string {
t.Helper()
dir := t.TempDir()
specPath := filepath.Join(dir, "spec.ts")
if err := os.WriteFile(specPath, []byte(specSource), 0o600); err != nil {
t.Fatal(err)
}
apiPath, err := filepath.Abs("../../pkg/spec/src/index.ts")
if err != nil {
t.Fatal(err)
}
runtimePath, err := filepath.Abs("../../pkg/spec/src/goja-runtime.ts")
if err != nil {
t.Fatal(err)
}
bundle, err := bundler.Bundle(bundler.Options{
EntryFile: specPath,
RuntimeFile: runtimePath,
Aliases: map[string]string{"@sanderling/spec": apiPath},
})
if err != nil {
t.Fatal(err)
}
return string(bundle.JavaScript)
}
func newHarnessWithSpec(t *testing.T, spec string) *harness {
t.Helper()
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
verifierInstance, err := verifier.New()
if err != nil {
t.Fatal(err)
}
if err := verifierInstance.Load(bundleSpec(t, spec)); err != nil {
t.Fatal(err)
}
state := &harness{
mock: mockdriver.New(),
verifier: verifierInstance,
writer: writer,
directory: directory,
}
t.Cleanup(func() { _ = writer.Close() })
return state
}
func TestRunner_HappyPathStepsAndTraces(t *testing.T) {
state := newHarness(t)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Errorf("expected at least one step, got 0")
}
if len(summary.Violations) != 0 {
t.Errorf("no violations expected, got %v", summary.Violations)
}
// Every builtin verb is supported on every platform, so a clean run must
// report no unsupported verbs (the runner still wires the field through).
if len(summary.UnsupportedVerbs) != 0 {
t.Errorf("expected no unsupported verbs, got %v", summary.UnsupportedVerbs)
}
actions := state.mock.Actions()
if !containsAction(actions, mockdriver.ActionTapSelector, "id:next") {
t.Errorf("expected TapSelector with id:next, got %v", actions)
}
}
// TestRunner_SeededRunRecordsNoModelCalls keeps the arms distinguishable: the
// seeded picker consults nothing, so its run directory must carry no model-call
// output at all rather than a file of empty records.
func TestRunner_SeededRunRecordsNoModelCalls(t *testing.T) {
state := newHarness(t)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 10 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
if _, err := os.Stat(filepath.Join(state.directory, trace.LLMCallFileName)); !os.IsNotExist(err) {
t.Errorf("stat %s = %v, want no model-call file for a seeded run", trace.LLMCallFileName, err)
}
}
// seededLoginSetupSpec drives the first two steps from setup, the way a
// login-fronted spec does, and leaves the rest to the seeded action root.
const seededLoginSetupSpec = `
import { always, actions, taps, typing, weighted, Tap } from "@sanderling/spec";
globalThis.properties = { ok: always(() => true) };
let setupTapsLeft = 2;
globalThis.setup = actions(() => (setupTapsLeft-- > 0 ? [Tap({ on: "id:Submit" })] : []));
globalThis.actions = weighted([1, taps], [1, typing]);
`
// TestRunner_SeededSetupActionsAreNotTheGeneratorDrivingTheApp: a seeded run's
// login steps used to be indistinguishable from its exploration, so the arm
// divided its defect rate by every action it dispatched while the model arm
// divided by the ones its policy chose. The two rates were then compared.
func TestRunner_SeededSetupActionsAreNotTheGeneratorDrivingTheApp(t *testing.T) {
state := newHarnessWithSpec(t, seededLoginSetupSpec)
state.mock.HierarchyJSON = llmTreeJSON
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 30 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 4,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.DispatchedActions != 4 {
t.Errorf("DispatchedActions = %d, want 4: every step drove the app",
summary.DispatchedActions)
}
if summary.GeneratorActions != 2 {
t.Errorf("GeneratorActions = %d, want 2: the picker drove the two steps setup left it",
summary.GeneratorActions)
}
lines := readTraceLines(t, state.writer.Directory())
if len(lines) != 4 {
t.Fatalf("wrote %d trace lines, want 4", len(lines))
}
for _, line := range lines[:2] {
if line.NextAction == nil || line.NextAction.Source != trace.ActionSourceSetup {
t.Errorf("step %d action = %+v, want one named setup", line.Step, line.NextAction)
}
}
for _, line := range lines[2:] {
if line.NextAction == nil || line.NextAction.Source != trace.ActionSourceSeeded {
t.Errorf("step %d action = %+v, want one named seeded", line.Step, line.NextAction)
}
}
}
func TestRunner_MaxStepsStopsAfterExactlyNSteps(t *testing.T) {
state := newHarness(t)
const maxSteps = 3
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
// A long duration ensures MaxSteps, not the deadline, ends the run.
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 10 * time.Millisecond,
MaxSteps: maxSteps,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != maxSteps {
t.Errorf("expected exactly %d steps, got %d", maxSteps, summary.Steps)
}
}
// TestRenderSummary_SurfacesUnsupportedVerbs exercises the real path that takes
// verbs the picker requested but the platform cannot dispatch and puts them in
// front of the operator. TestRunner_HappyPath only ever asserts the field stays
// empty (every builtin is supported, so its non-empty arm can never fire), so
// without this the "unsupported on %s: ..." branch could be deleted and every
// unsupported-verb regression would pass silently.
func TestRenderSummary_SurfacesUnsupportedVerbs(t *testing.T) {
summary := Summary{Steps: 3, UnsupportedVerbs: []string{"longPresses", "scrolls"}}
var out bytes.Buffer
RenderSummary(&out, summary, "ios")
if !strings.Contains(out.String(), "unsupported on ios: longPresses, scrolls") {
t.Errorf("expected unsupported verbs line for ios, got:\n%s", out.String())
}
}
// TestRenderSummary_OmitsUnsupportedLineWhenNone guards the inverse: a clean run
// must not print a stray "unsupported on" line, so a future refactor cannot
// start emitting an empty list and alarm the operator on every run.
func TestRenderSummary_OmitsUnsupportedLineWhenNone(t *testing.T) {
var out bytes.Buffer
RenderSummary(&out, Summary{Steps: 2}, "android")
if strings.Contains(out.String(), "unsupported on") {
t.Errorf("did not expect an unsupported-verbs line, got:\n%s", out.String())
}
}
// A step nothing judged is not a step that passed. The run prints its count so
// a green summary cannot hide a run that skipped most of its steps, which is
// what a screen that keeps moving under the reads would produce.
func TestRenderSummary_CountsTheStepsNothingJudged(t *testing.T) {
var out bytes.Buffer
RenderSummary(&out, Summary{Steps: 10, SkippedVerification: 4}, "android")
if !strings.Contains(out.String(), "4 step(s) judged by nothing") {
t.Errorf("expected the unjudged-step count, got:\n%s", out.String())
}
out.Reset()
RenderSummary(&out, Summary{Steps: 10}, "android")
if strings.Contains(out.String(), "judged by nothing") {
t.Errorf("a run that judged every step must not print the line, got:\n%s", out.String())
}
}
func TestRunner_ViolationSurfacesInSummary(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if len(summary.Violations) == 0 {
t.Errorf("expected at least one violation, got %v", summary.Violations)
}
if !containsProperty(summary.Violations, "balanceNonNegative") {
t.Errorf("expected balanceNonNegative in violations: %v", summary.Violations)
}
}
func TestRunner_ViolationSurfacesOnlyOnOnsetStep(t *testing.T) {
// violationSpec uses always(() => false): onset fires on step 1 and the
// residual stays violated forever. The runner must record the violation
// exactly once (at the onset step) in both summary.Violations and trace
// lines, not on every subsequent step.
state := newHarnessWithSpec(t, violationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to prove onset-only behavior, got %d", summary.Steps)
}
if len(summary.Violations) != 1 {
t.Fatalf("expected exactly one ViolationRecord (onset only), got %d: %v",
len(summary.Violations), summary.Violations)
}
if summary.Violations[0].StepIndex != 1 {
t.Errorf("onset step: got %d, want 1 (always(()=>false) violates immediately)",
summary.Violations[0].StepIndex)
}
if !slices.Equal(summary.Violations[0].Properties, []string{"balanceNonNegative"}) {
t.Errorf("onset properties: got %v, want [balanceNonNegative]",
summary.Violations[0].Properties)
}
file, err := os.Open(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
type traceLine struct {
Step int `json:"step"`
Violations []string `json:"violations"`
}
linesWithViolations := 0
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var line traceLine
if err := json.Unmarshal(scanner.Bytes(), &line); err != nil {
t.Fatalf("trace line decode: %v", err)
}
if len(line.Violations) == 0 {
continue
}
linesWithViolations++
if line.Step != 1 {
t.Errorf("step %d unexpectedly emitted violations %v (should be onset-only at step 1)",
line.Step, line.Violations)
}
}
if err := scanner.Err(); err != nil {
t.Fatalf("scan trace: %v", err)
}
if linesWithViolations != 1 {
t.Errorf("expected exactly 1 trace line with violations, got %d", linesWithViolations)
}
}
func TestRunner_NextViolationAttributedToCausingStep(t *testing.T) {
// always(next(p)): the obligation spawned at step 2 fails against step 3's
// state. The summary record and the trace witness must attribute the
// violation to step 2 (the causing step); the trace line that carries it is
// still step 3, where the failure was detected.
const nextViolationSpec = `
import { actions, always, next, extract } from "@sanderling/spec";
let observed = 0;
const tick = extract(() => ++observed);
globalThis.properties = {
nextHolds: always(next(() => tick.current < 3)),
};
globalThis.actions = actions(() => []);
`
state := newHarnessWithSpec(t, nextViolationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 5,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if len(summary.Violations) != 1 {
t.Fatalf("expected exactly one ViolationRecord, got %d: %v",
len(summary.Violations), summary.Violations)
}
if summary.Violations[0].StepIndex != 2 {
t.Errorf("summary step: got %d, want 2 (the step that spawned the next obligation)",
summary.Violations[0].StepIndex)
}
if !slices.Equal(summary.Violations[0].Properties, []string{"nextHolds"}) {
t.Errorf("properties: got %v, want [nextHolds]", summary.Violations[0].Properties)
}
file, err := os.Open(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
type traceLine struct {
Step int `json:"step"`
Violations []string `json:"violations"`
Witnesses map[string]trace.Witness `json:"witnesses"`
}
found := false
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var line traceLine
if err := json.Unmarshal(scanner.Bytes(), &line); err != nil {
t.Fatalf("trace line decode: %v", err)
}
if len(line.Violations) == 0 {
continue
}
found = true
if line.Step != 3 {
t.Errorf("violation detected on step %d, want 3", line.Step)
}
witness, ok := line.Witnesses["nextHolds"]
if !ok {
t.Fatalf("step %d carries no witness for nextHolds", line.Step)
}
if witness.Step != 2 {
t.Errorf("witness step: got %d, want 2 (causing step)", witness.Step)
}
// The two indices the witness spans are recorded separately: the
// extractor snapshot it carries is step 3's state, not step 2's.
if witness.DetectedStep != 3 {
t.Errorf("witness detected step: got %d, want 3", witness.DetectedStep)
}
}
if err := scanner.Err(); err != nil {
t.Fatalf("scan trace: %v", err)
}
if !found {
t.Error("no trace line carried the violation")
}
}
func TestRunner_AlwaysNextLeavesNoEndOfRunViolation(t *testing.T) {
// always(next(p)) ends every run with a pending deferred check. That
// residue is vacuous (no successor state to check), so neither the
// summary nor the trace may report an end-of-run violation.
const spec = `
import { actions, always, next } from "@sanderling/spec";
globalThis.properties = {
nextHolds: always(next(() => true)),
};
globalThis.actions = actions(() => []);
`
state := newHarnessWithSpec(t, spec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if len(summary.Violations) != 0 {
t.Errorf("expected no violations, got %v", summary.Violations)
}
}
func TestRunner_FinalizeRecordUsesDistinctStepIndex(t *testing.T) {
// An eventually that never fires is reported at run end through a
// synthetic trace record. That record must carry its own step index so no
// two trace lines share one (duplicate indices made the replay UI select
// two rows at once).
const spec = `
import { actions, eventually } from "@sanderling/spec";
globalThis.properties = {
neverFires: eventually(() => false),
};
globalThis.actions = actions(() => []);
`
state := newHarnessWithSpec(t, spec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if !containsProperty(summary.Violations, "neverFires") {
t.Fatalf("expected neverFires in violations: %v", summary.Violations)
}
file, err := os.Open(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
type traceLine struct {
Step int `json:"step"`
Violations []string `json:"violations"`
}
seen := map[int]bool{}
finalizeStep := 0
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var line traceLine
if err := json.Unmarshal(scanner.Bytes(), &line); err != nil {
t.Fatalf("trace line decode: %v", err)
}
if seen[line.Step] {
t.Errorf("duplicate step index %d in trace", line.Step)
}
seen[line.Step] = true
if slices.Contains(line.Violations, "neverFires") {
finalizeStep = line.Step
}
}
if err := scanner.Err(); err != nil {
t.Fatalf("scan trace: %v", err)
}
if finalizeStep != summary.Steps+1 {
t.Errorf("finalize record step = %d, want %d (steps+1)", finalizeStep, summary.Steps+1)
}
}
func TestRunner_ThrowingPredicateIsLoggedNotPanic(t *testing.T) {
const throwingSpec = `
import { actions, always, Tap } from "@sanderling/spec";
globalThis.properties = {
broken: always(() => { throw new Error("bad predicate"); }),
};
globalThis.actions = actions(() => [Tap({ on: "id:next" })]);
`
state := newHarnessWithSpec(t, throwingSpec)
var buffer bytes.Buffer
logger := slog.New(slog.NewTextHandler(&buffer, &slog.HandlerOptions{Level: slog.LevelWarn}))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: logger,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if !containsProperty(summary.Violations, "broken") {
t.Errorf("expected broken in violations: %v", summary.Violations)
}
if !strings.Contains(buffer.String(), "bad predicate") {
t.Errorf("expected predicate error in log, got %q", buffer.String())
}
}
func TestRunner_RejectsMissingFields(t *testing.T) {
_, err := Run(context.Background(), Options{Duration: time.Second})
if err == nil || !strings.Contains(err.Error(), "Driver") {
t.Errorf("expected Driver-required error, got %v", err)
}
}
func TestRunner_RejectsZeroDuration(t *testing.T) {
_, err := Run(context.Background(), Options{
Driver: mockdriver.New(),
Verifier: mustNewVerifier(t),
TraceWriter: mustNewTraceWriter(t),
})
if err == nil || !strings.Contains(err.Error(), "Duration") {
t.Errorf("expected Duration-required error, got %v", err)
}
}
func TestRunner_StampsHierarchyResolvedBoundsAndResiduals(t *testing.T) {
state := newHarness(t)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"com.fixture:id/next","bounds":"[40,80,240,160]"},"children":[],"clickable":true,"enabled":true}`
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
text := string(body)
if !strings.Contains(text, `"selector":"id:next"`) {
t.Errorf("expected selector in trace: %s", text)
}
if !strings.Contains(text, `"resolved_bounds":{"x":40,"y":80,"width":200,"height":80}`) {
t.Errorf("expected resolved_bounds in trace: %s", text)
}
if !strings.Contains(text, `"tap_point":{"x":140,"y":120}`) {
t.Errorf("expected tap_point in trace: %s", text)
}
if !strings.Contains(text, `"hierarchy":{"elements":`) {
t.Errorf("expected hierarchy in trace: %s", text)
}
if !strings.Contains(text, `"residuals":{`) {
t.Errorf("expected residuals in trace: %s", text)
}
}
// TestTraceActionFor_StaleCoordinatesDoNotOverrideTreeCenter pins the stamp
// to applyAction's resolution rule: when On resolves in the tree, the trace
// tap point must be the tree center even if the action carries stale X/Y
// from an earlier tick.
func TestTraceActionFor_StaleCoordinatesDoNotOverrideTreeCenter(t *testing.T) {
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"next","bounds":"[100,200,300,400]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
action := verifier.Action{Kind: verifier.ActionKindTap, On: "id:next", X: 50, Y: 60}
traceAction := traceActionFor(action, tree)
if traceAction.TapPoint == nil {
t.Fatal("expected a tap point")
}
if traceAction.TapPoint.X != 200 || traceAction.TapPoint.Y != 300 {
t.Errorf("tap point = (%d,%d), want tree center (200,300)",
traceAction.TapPoint.X, traceAction.TapPoint.Y)
}
if traceAction.ResolvedBounds == nil {
t.Fatal("expected resolved bounds")
}
if traceAction.ResolvedBounds.X != 100 || traceAction.ResolvedBounds.Y != 200 {
t.Errorf("resolved bounds origin = (%d,%d), want (100,200)",
traceAction.ResolvedBounds.X, traceAction.ResolvedBounds.Y)
}
}
// TestTraceActionFor_RecordsKindSpecificFields locks each action kind's trace
// encoding. PressKey must carry its Key and Wait its DurationMillis; if either
// branch of traceActionFor drops the field (or a field rename desyncs from the
// trace.Action struct) the replay UI silently renders a key-less PressKey or a
// zero-duration Wait. Swipe's endpoint encoding is covered separately via
// applyAction (TestApplyAction_ScrollWithPrecomputedEndpointsSwipes).
func TestTraceActionFor_RecordsKindSpecificFields(t *testing.T) {
cases := []struct {
name string
action verifier.Action
check func(*testing.T, *trace.Action)
}{
{
"PressKey records key",
verifier.Action{Kind: verifier.ActionKindPressKey, Key: "back"},
func(t *testing.T, a *trace.Action) {
if a.Key != "back" {
t.Errorf("Key = %q, want %q", a.Key, "back")
}
},
},
{
"Wait records duration",
verifier.Action{Kind: verifier.ActionKindWait, DurationMillis: 250},
func(t *testing.T, a *trace.Action) {
if a.DurationMillis != 250 {
t.Errorf("DurationMillis = %d, want 250", a.DurationMillis)
}
},
},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
traceAction := traceActionFor(testCase.action, nil)
if traceAction.Kind != string(testCase.action.Kind) {
t.Errorf("Kind = %q, want %q", traceAction.Kind, testCase.action.Kind)
}
testCase.check(t, traceAction)
})
}
}
func TestRunner_LogsWaitForIdleDriverErrors(t *testing.T) {
state := newHarness(t)
state.mock.Failures[mockdriver.ActionWaitForIdle] = errors.New("sidecar lost gRPC stream")
var logBuf bytes.Buffer
logger := slog.New(slog.NewTextHandler(&logBuf, &slog.HandlerOptions{Level: slog.LevelWarn}))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: logger,
}); err != nil {
t.Fatalf("Run: %v", err)
}
output := logBuf.String()
if !strings.Contains(output, "wait_for_idle failed") {
t.Errorf("expected wait_for_idle warning, got: %q", output)
}
if !strings.Contains(output, "sidecar lost gRPC stream") {
t.Errorf("expected driver error message in warning, got: %q", output)
}
}
func TestApplyAction_InputTextErasesExistingTextBeforeTyping(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"username","text":"stale-value","bounds":"[10,10,500,100]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindInputText, On: "id:username", Text: "alice"}
mustDispatch(t, driverMock, action, tree)
// The post-tap settle is now a brief internal sleep, not a WaitForIdle RPC,
// so the recorded driver actions are tap, erase, input.
actions := driverMock.Actions()
if len(actions) != 3 {
t.Fatalf("want tap, erase, input; got %v", actions)
}
if actions[0].Kind != mockdriver.ActionTap {
t.Errorf("first action = %v, want tap", actions[0].Kind)
}
if actions[1].Kind != mockdriver.ActionEraseText || actions[1].CharacterCount != len("stale-value") {
t.Errorf("second action = %+v, want erase_text of %d characters", actions[1], len("stale-value"))
}
if actions[2].Kind != mockdriver.ActionInputText || actions[2].Text != "alice" {
t.Errorf("third action = %+v, want input_text alice", actions[2])
}
}
// TestApplyAction_InputTextSkipsEraseForReplacingDriver pins that a driver
// asserting the TextReplacer capability never pays the pre-erase round-trip:
// its InputText already replaces the field's content.
func TestApplyAction_InputTextSkipsEraseForReplacingDriver(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"username","text":"stale-value","bounds":"[10,10,500,100]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.ReplacesText = true
action := verifier.Action{Kind: verifier.ActionKindInputText, On: "id:username", Text: "alice"}
mustDispatch(t, driverMock, action, tree)
if containsAction(driverMock.Actions(), mockdriver.ActionEraseText, "") {
t.Errorf("replacing driver must not be asked to erase: %v", driverMock.Actions())
}
if !containsAction(driverMock.Actions(), mockdriver.ActionInputText, "") {
t.Errorf("expected InputText, got %v", driverMock.Actions())
}
}
func TestApplyAction_InputTextSkipsEraseWhenTargetEmpty(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"username","bounds":"[10,10,500,100]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindInputText, On: "id:username", Text: "alice"}
mustDispatch(t, driverMock, action, tree)
if containsAction(driverMock.Actions(), mockdriver.ActionEraseText, "") {
t.Errorf("empty field must not be erased: %v", driverMock.Actions())
}
}
func TestApplyAction_InputTextSurfacesFocusTapError(t *testing.T) {
t.Run("selector focus tap fails", func(t *testing.T) {
driverMock := mockdriver.New()
driverMock.Failures[mockdriver.ActionTapSelector] = errors.New("adb unreachable")
action := verifier.Action{Kind: verifier.ActionKindInputText, On: "id:username", Text: "alice"}
_, err := applyAction(context.Background(), driverMock, action, nil)
if err == nil {
t.Fatalf("expected focus tap failure to surface, got nil")
}
if containsAction(driverMock.Actions(), mockdriver.ActionInputText, "") {
t.Errorf("InputText must not run after focus tap failed: %v", driverMock.Actions())
}
})
t.Run("coordinate focus tap fails", func(t *testing.T) {
driverMock := mockdriver.New()
driverMock.Failures[mockdriver.ActionTap] = errors.New("tap driver error")
action := verifier.Action{Kind: verifier.ActionKindInputText, X: 10, Y: 20, Text: "alice"}
_, err := applyAction(context.Background(), driverMock, action, nil)
if err == nil {
t.Fatalf("expected focus tap failure to surface, got nil")
}
if containsAction(driverMock.Actions(), mockdriver.ActionInputText, "") {
t.Errorf("InputText must not run after focus tap failed: %v", driverMock.Actions())
}
})
}
// loginFocusOnEmail is the folio login screen as Android reports it once the
// email field has been typed into: email holds focus, password does not.
const loginFocusOnEmail = `{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"LoginEmail","text":"[email protected]","bounds":"[94,240,986,372]"},"focused":true,"children":[]},
{"attributes":{"resource-id":"LoginPassword","bounds":"[94,461,986,593]"},"focused":false,"children":[]}
]}`
const loginFocusOnPassword = `{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"LoginEmail","text":"[email protected]","bounds":"[94,240,986,372]"},"focused":false,"children":[]},
{"attributes":{"resource-id":"LoginPassword","bounds":"[94,461,986,593]"},"focused":true,"children":[]}
]}`
// A keyboard overlay window can sit over the field the focus tap aims at, so
// the tap never reaches it and focus stays where it was. Typing then appends to
// the previously focused field: on folio the password ran into the email field
// and the login setup leaf retried forever.
func TestApplyAction_InputTextStopsWhenAnotherFieldHoldsFocus(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(loginFocusOnEmail)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.HierarchyJSON = loginFocusOnEmail
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:LoginPassword",
Text: "ledger123",
}
_, err = applyAction(context.Background(), driverMock, action, tree)
if err == nil {
t.Fatal("a focus tap that never focused the target must be reported, not typed through")
}
if !strings.Contains(err.Error(), "LoginEmail") {
t.Errorf("error must name the field holding focus, got: %v", err)
}
if containsAction(driverMock.Actions(), mockdriver.ActionInputText, "") {
t.Errorf("password text must not be typed into the focused email field: %v", driverMock.Actions())
}
}
func TestApplyAction_InputTextTypesWhenTargetTakesFocus(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(loginFocusOnEmail)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.HierarchyJSON = loginFocusOnPassword
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:LoginPassword",
Text: "ledger123",
}
mustDispatch(t, driverMock, action, tree)
if !typedText(driverMock.Actions(), "ledger123") {
t.Errorf("expected InputText once the target holds focus, got %v", driverMock.Actions())
}
}
// Platforms whose hierarchy omits focus entirely (iOS) have nothing to compare,
// so they must not pay a hierarchy read per InputText.
func TestApplyAction_InputTextSkipsFocusCheckWhenHierarchyOmitsFocus(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"LoginPassword","bounds":"[94,461,986,593]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:LoginPassword",
Text: "ledger123",
}
mustDispatch(t, driverMock, action, tree)
if containsAction(driverMock.Actions(), mockdriver.ActionHierarchy, "") {
t.Errorf("no focus to compare: expected no hierarchy read, got %v", driverMock.Actions())
}
if !typedText(driverMock.Actions(), "ledger123") {
t.Errorf("expected InputText, got %v", driverMock.Actions())
}
}
const loginFocusOnNothing = `{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"resource-id":"LoginEmail","bounds":"[94,240,986,372]"},"focused":false,"children":[]},
{"attributes":{"resource-id":"LoginPassword","bounds":"[94,461,986,593]"},"focused":false,"children":[]}
]}`
// Typed text can only be corrupted into a field that already holds focus, so
// a pre-tap hierarchy showing focus elsewhere is the one class worth the
// confirming read.
func TestApplyAction_InputTextConfirmsFocusWhenAnotherFieldHeldItBeforeTheTap(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(loginFocusOnEmail)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.HierarchyJSON = loginFocusOnPassword
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:LoginPassword",
Text: "ledger123",
}
mustDispatch(t, driverMock, action, tree)
if !containsAction(driverMock.Actions(), mockdriver.ActionHierarchy, "") {
t.Errorf("another field held focus before the tap: expected the confirming read, got %v", driverMock.Actions())
}
if !typedText(driverMock.Actions(), "ledger123") {
t.Errorf("expected InputText once the target took focus, got %v", driverMock.Actions())
}
}
// The confirming read is a device round-trip on every InputText step. Where no
// other element holds focus before the tap there is no field for the text to
// be corrupted into, so the read buys nothing and must not be paid for.
func TestApplyAction_InputTextSkipsFocusCheckWhenNoOtherFieldHoldsFocus(t *testing.T) {
fastFocusSettle(t)
for name, beforeTap := range map[string]string{
"target already holds focus": loginFocusOnPassword,
"nothing holds focus": loginFocusOnNothing,
} {
t.Run(name, func(t *testing.T) {
tree, err := hierarchy.Parse(beforeTap)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.HierarchyJSON = loginFocusOnEmail
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:LoginPassword",
Text: "ledger123",
}
mustDispatch(t, driverMock, action, tree)
if containsAction(driverMock.Actions(), mockdriver.ActionHierarchy, "") {
t.Errorf("no other field held focus: expected no confirming read, got %v", driverMock.Actions())
}
if !typedText(driverMock.Actions(), "ledger123") {
t.Errorf("expected InputText, got %v", driverMock.Actions())
}
})
}
}
// A selector the hierarchy cannot resolve is the guard saying "I cannot tell",
// not "another element holds focus". Answering the second manufactures an apply
// error on every InputText the tree has no node for, and three in a row abort
// the run.
func TestApplyAction_InputTextTypesWhenTheTargetIsNotInTheHierarchy(t *testing.T) {
fastFocusSettle(t)
tree, err := hierarchy.Parse(loginFocusOnEmail)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
driverMock.HierarchyJSON = loginFocusOnEmail
action := verifier.Action{
Kind: verifier.ActionKindInputText,
On: "data-testid:LoginPassword",
Text: "ledger123",
}
mustDispatch(t, driverMock, action, tree)
if !typedText(driverMock.Actions(), "ledger123") {
t.Errorf("an unresolvable target must not block typing, got %v", driverMock.Actions())
}
}
func typedText(actions []mockdriver.Action, text string) bool {
for _, action := range actions {
if action.Kind == mockdriver.ActionInputText && action.Text == text {
return true
}
}
return false
}
func TestApplyAction_V8InputTextTapsAtCoordinates(t *testing.T) {
fastFocusSettle(t)
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindInputText, X: 50, Y: 100, Text: "alice"}
mustDispatch(t, driverMock, action, nil)
actions := driverMock.Actions()
if !containsAction(actions, mockdriver.ActionTap, "") {
t.Errorf("expected focus Tap before InputText, got %v", actions)
}
if !containsAction(actions, mockdriver.ActionInputText, "") {
t.Errorf("expected InputText after focus Tap, got %v", actions)
}
}
func TestApplyAction_V8InputTextAtOriginStillTaps(t *testing.T) {
fastFocusSettle(t)
driverMock := mockdriver.New()
// V8 emits real (0,0) coordinates for an element at viewport top-left
// (post-#15 the runtime nullifies unresolved actions, so a non-null
// InputText with (0,0) is a deliberate edge tap, not a sentinel).
action := verifier.Action{Kind: verifier.ActionKindInputText, X: 0, Y: 0, Text: "alice"}
mustDispatch(t, driverMock, action, nil)
if !containsAction(driverMock.Actions(), mockdriver.ActionTap, "") {
t.Errorf("expected focus Tap at (0,0), got %v", driverMock.Actions())
}
}
// Attribute values match by substring, so "data-testid:card" answers to
// "card-1" and "card-10" alike. A candidate built where the selector named one
// element can execute where it names several, and the tree lookup would send
// every one of them to the first match.
func TestApplyAction_AmbiguousSelectorTapsTheActionsOwnCoordinates(t *testing.T) {
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"data-testid":"card-1","bounds":"[0,100,200,200]"},"children":[]},
{"attributes":{"data-testid":"card-10","bounds":"[0,300,200,400]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindTap, On: "data-testid:card-1", X: 100, Y: 350}
mustDispatch(t, driverMock, action, tree)
for _, dispatched := range driverMock.Actions() {
if dispatched.Kind == mockdriver.ActionTap && dispatched.X == 100 && dispatched.Y == 350 {
return
}
}
t.Errorf("tap reached the driver at %v, want the action's own (100,350)", driverMock.Actions())
}
// A bare-string target carries no coordinates of its own, so an ambiguous name
// is all there is to act on and the first match stays the answer. Refusing it
// would drop an authored action.
func TestApplyAction_AmbiguousSelectorWithoutCoordinatesTapsTheFirstMatch(t *testing.T) {
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,1080,2340]"},"children":[
{"attributes":{"data-testid":"card-1","bounds":"[0,100,200,200]"},"children":[]},
{"attributes":{"data-testid":"card-10","bounds":"[0,300,200,400]"},"children":[]}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindTap, On: "data-testid:card-1"}
mustDispatch(t, driverMock, action, tree)
for _, dispatched := range driverMock.Actions() {
if dispatched.Kind == mockdriver.ActionTap && dispatched.X == 100 && dispatched.Y == 150 {
return
}
}
t.Errorf("tap reached the driver at %v, want the first match's centre (100,150)", driverMock.Actions())
}
func TestApplyAction_DoubleTapDispatchesDoubleTapAtCoordinates(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindDoubleTap, X: 100, Y: 200}
mustDispatch(t, driverMock, action, nil)
taps := 0
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionDoubleTap && a.X == 100 && a.Y == 200 {
taps++
}
}
if taps != 1 {
t.Errorf("expected 1 DoubleTap call at (100,200), got %d in %v", taps, driverMock.Actions())
}
}
func TestApplyAction_DoubleTapDispatchesDoubleTapSelector(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindDoubleTap, On: "id:save"}
mustDispatch(t, driverMock, action, nil)
taps := 0
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionDoubleTapSelector && a.Selector == "id:save" {
taps++
}
}
if taps != 1 {
t.Errorf("expected 1 DoubleTapSelector call with id:save, got %d in %v", taps, driverMock.Actions())
}
}
func TestApplyAction_LongPressDispatchesAtResolvedCoordinates(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindLongPress, X: 120, Y: 240}
mustDispatch(t, driverMock, action, nil)
found := false
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionLongPress && a.X == 120 && a.Y == 240 {
found = true
}
}
if !found {
t.Errorf("expected LongPress at (120,240), got %v", driverMock.Actions())
}
}
func TestApplyAction_ScrollWithPrecomputedEndpointsSwipes(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{
Kind: verifier.ActionKindScroll,
Direction: "down",
FromX: 100,
FromY: 500,
ToX: 100,
ToY: 300,
DurationMillis: 300,
}
mustDispatch(t, driverMock, action, nil)
found := false
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionSwipe && a.FromX == 100 && a.FromY == 500 && a.ToX == 100 && a.ToY == 300 {
found = true
}
}
if !found {
t.Errorf("expected Swipe with precomputed endpoints, got %v", driverMock.Actions())
}
}
// scrollingDriver is a driver that scrolls by something other than a drag,
// which is what the web driver is: a browser scrolls on wheel input and only
// ever treats a drag as a drag.
type scrollingDriver struct {
*mockdriver.Driver
scrolls [][4]int
}
func (d *scrollingDriver) Scroll(
_ context.Context,
fromX, fromY, toX, toY int,
_ time.Duration,
) error {
d.scrolls = append(d.scrolls, [4]int{fromX, fromY, toX, toY})
return nil
}
func TestApplyAction_ScrollPrefersTheScrollCapabilityOverASwipe(t *testing.T) {
drv := &scrollingDriver{Driver: mockdriver.New()}
action := verifier.Action{
Kind: verifier.ActionKindScroll,
Direction: "down",
FromX: 100,
FromY: 500,
ToX: 100,
ToY: 300,
DurationMillis: 300,
}
mustDispatch(t, drv, action, nil)
want := [4]int{100, 500, 100, 300}
if len(drv.scrolls) != 1 || drv.scrolls[0] != want {
t.Fatalf("scrolls = %v, want one %v", drv.scrolls, want)
}
for _, a := range drv.Actions() {
if a.Kind == mockdriver.ActionSwipe {
t.Errorf("the scroll was dispatched as a swipe: %v", drv.Actions())
}
}
}
func TestApplyAction_ScrollDirectionUsesInversion(t *testing.T) {
driverMock := mockdriver.New()
treeJSON := `{"attributes":{"resource-id":"com.fixture:id/list","bounds":"[0,0,400,800]"},"children":[],"enabled":true}`
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
action := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "down", On: "id:list"}
mustDispatch(t, driverMock, action, tree)
var swipe *mockdriver.Action
for i := range driverMock.Actions() {
if driverMock.Actions()[i].Kind == mockdriver.ActionSwipe {
a := driverMock.Actions()[i]
swipe = &a
}
}
if swipe == nil {
t.Fatalf("expected a Swipe, got %v", driverMock.Actions())
}
// "down" reveals lower content by dragging the finger up, so toY < fromY.
if swipe.ToY >= swipe.FromY {
t.Errorf("expected toY < fromY for scroll down, got from=%d to=%d", swipe.FromY, swipe.ToY)
}
}
func TestApplyAction_ScrollNearTopKeepsDirectionAfterClamp(t *testing.T) {
driverMock := mockdriver.New()
// Full-screen root sets the 1080x2400 screen (marginY=200); the scrollable
// list sits inside the top margin (y 20..180), where the clamp must fire.
treeJSON := `{"attributes":{"bounds":"[0,0,1080,2400]"},"children":[
{"attributes":{"resource-id":"com.fixture:id/toplist","scrollable":"true","bounds":"[0,20,1080,180]"},"children":[],"enabled":true}
]}`
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
action := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "up", On: "id:toplist"}
mustDispatch(t, driverMock, action, tree)
var swipe *mockdriver.Action
for i := range driverMock.Actions() {
if driverMock.Actions()[i].Kind == mockdriver.ActionSwipe {
a := driverMock.Actions()[i]
swipe = &a
}
}
if swipe == nil {
t.Fatalf("expected a Swipe, got %v", driverMock.Actions())
}
if swipe.FromY != 200 {
t.Errorf("origin not pushed below the shade strip, got fromY=%d want 200", swipe.FromY)
}
if swipe.ToY <= swipe.FromY {
t.Errorf("scroll up reversed by the clamp: from=%d to=%d (want toY > fromY)", swipe.FromY, swipe.ToY)
}
}
func TestApplyAction_ScrollScreenFallback(t *testing.T) {
driverMock := mockdriver.New()
treeJSON := `{"attributes":{"bounds":"[0,0,400,800]"},"children":[],"enabled":true}`
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
// On unset: container falls back to whole-screen (root) bounds.
action := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "up"}
mustDispatch(t, driverMock, action, tree)
var swipe *mockdriver.Action
for i := range driverMock.Actions() {
if driverMock.Actions()[i].Kind == mockdriver.ActionSwipe {
a := driverMock.Actions()[i]
swipe = &a
}
}
if swipe == nil {
t.Fatalf("expected a Swipe, got %v", driverMock.Actions())
}
if swipe.FromX != 200 || swipe.FromY != 400 {
t.Errorf("expected swipe to start at screen center (200,400), got (%d,%d)", swipe.FromX, swipe.FromY)
}
// "up" reveals upper content by dragging the finger down, so toY > fromY.
if swipe.ToY <= swipe.FromY {
t.Errorf("expected toY > fromY for scroll up, got from=%d to=%d", swipe.FromY, swipe.ToY)
}
}
// Every shape applyAction cannot dispatch has to name why. Returning a bare nil
// leaves the step recording a next_action that reached no driver at all, which
// an executed-action count then reads as work done.
func TestApplyAction_NonDispatchPathsReportWhy(t *testing.T) {
tree, err := hierarchy.Parse(`{"attributes":{"resource-id":"root","bounds":"[0,0,400,800]"},"children":[]}`)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
cases := []struct {
name string
action verifier.Action
tree *hierarchy.Tree
want actionSkipReason
}{
{
name: "long press whose selector is not on screen",
action: verifier.Action{Kind: verifier.ActionKindLongPress, On: "id:gone"},
tree: tree,
want: actionSkippedUnresolvedSelector,
},
{
name: "press key without a key",
action: verifier.Action{Kind: verifier.ActionKindPressKey},
want: actionSkippedMissingKey,
},
{
name: "wait without a duration",
action: verifier.Action{Kind: verifier.ActionKindWait},
want: actionSkippedZeroDurationWait,
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
driverMock := mockdriver.New()
skipped, err := applyAction(context.Background(), driverMock, tc.action, tc.tree)
if err != nil {
t.Fatalf("applyAction: %v", err)
}
if skipped != tc.want {
t.Errorf("skip reason = %q, want %q", skipped, tc.want)
}
if len(driverMock.Actions()) != 0 {
t.Errorf("nothing must reach the driver, got %v", driverMock.Actions())
}
})
}
}
// TestRunner_RecordsWhyAChosenActionNeverRan drives a spec whose only action is
// undispatchable and pins that each step says so on its trace line. The step is
// left non-transitional: nothing was dispatched, so the verified screen still
// describes the device.
func TestRunner_RecordsWhyAChosenActionNeverRan(t *testing.T) {
state := newHarnessWithSpec(t, zeroWaitSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 2 {
t.Fatalf("Steps = %d, want 2", summary.Steps)
}
lines := readTraceLines(t, state.writer.Directory())
acted := 0
for _, line := range lines {
if line.NextAction == nil {
continue
}
acted++
if line.ActionSkipped != string(actionSkippedZeroDurationWait) {
t.Errorf("step %d action_skipped = %q, want %q",
line.Step, line.ActionSkipped, actionSkippedZeroDurationWait)
}
if line.Transitional {
t.Errorf("step %d marked transitional; nothing was dispatched, so the screen is unchanged", line.Step)
}
}
if acted != 2 {
t.Fatalf("want 2 steps carrying a next_action, got %d", acted)
}
}
// The other half of the contract: a step whose action really was dispatched
// must carry no reason at all, or every step looks skipped.
func TestRunner_DispatchedActionRecordsNoSkipReason(t *testing.T) {
state := newHarness(t)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
if !containsAction(state.mock.Actions(), mockdriver.ActionTapSelector, "id:next") {
t.Fatalf("fixture tap never dispatched, got %v", state.mock.Actions())
}
for _, line := range readTraceLines(t, state.writer.Directory()) {
if line.NextAction != nil && line.ActionSkipped != "" {
t.Errorf("step %d recorded action_skipped=%q for a dispatched action",
line.Step, line.ActionSkipped)
}
}
}
// TestRunner_ASourceAskedAndHandedNothingSaysSo covers the run that touched the
// app zero times: every step asked the source for an action and got none, which
// is what a picker whose every model call fails does. Without a reason on the
// step the trace, the summary and the exit status of that run are the ones a run
// that exercised all three steps produces.
func TestRunner_ASourceAskedAndHandedNothingSaysSo(t *testing.T) {
state := newHarnessWithSpec(t, noActionSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 3 {
t.Fatalf("Steps = %d, want 3", summary.Steps)
}
if summary.DispatchedActions != 0 {
t.Fatalf("DispatchedActions = %d, want 0: the fixture offers no action to dispatch",
summary.DispatchedActions)
}
if got := summary.SkippedActions[string(actionSkippedNoActionProduced)]; got != 3 {
t.Errorf("summary counted %d step(s) as %q, want 3: %v",
got, actionSkippedNoActionProduced, summary.SkippedActions)
}
for _, line := range readTraceLines(t, state.writer.Directory()) {
if line.NextAction != nil {
t.Errorf("step %d carries a next_action the source never produced", line.Step)
}
if line.ActionSkipped != string(actionSkippedNoActionProduced) {
t.Errorf("step %d action_skipped = %q, want %q",
line.Step, line.ActionSkipped, actionSkippedNoActionProduced)
}
if line.SkippedVerification {
t.Errorf("step %d reads as held; the source was asked on it", line.Step)
}
}
var rendered bytes.Buffer
RenderSummary(&rendered, summary, "android")
if !strings.Contains(rendered.String(), "3 action(s) never reached the app") {
t.Errorf("the run reports as a clean one:\n%s", rendered.String())
}
}
// The other half: a held step never asked the source for anything, so it must
// stay distinguishable from one that asked and was handed nothing. The fixture
// here has an action to offer, and no step of this run gets to hear it.
func TestRunner_AHeldStepIsNotRecordedAsASourceThatDeclined(t *testing.T) {
state := newHarness(t)
state.mock.Failures[mockdriver.ActionSnapshot] = errors.New("adb: device offline")
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.SkippedVerification != 3 {
t.Fatalf("SkippedVerification = %d, want 3", summary.SkippedVerification)
}
if len(summary.SkippedActions) != 0 {
t.Errorf("held steps counted as skipped actions: %v", summary.SkippedActions)
}
for _, line := range readTraceLines(t, state.writer.Directory()) {
if !line.SkippedVerification {
t.Errorf("step %d was not recorded as held", line.Step)
}
if line.ActionSkipped != "" {
t.Errorf("step %d action_skipped = %q; the source was never asked",
line.Step, line.ActionSkipped)
}
}
}
type traceStepLine struct {
Step int `json:"step"`
NextAction *trace.Action `json:"next_action"`
ActionSkipped string `json:"action_skipped"`
Transitional bool `json:"transitional"`
ObservationError string `json:"observation_error"`
SkippedVerification bool `json:"skipped_verification"`
}
func readTraceLines(t *testing.T, directory string) []traceStepLine {
t.Helper()
body, err := os.ReadFile(filepath.Join(directory, "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
var lines []traceStepLine
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceStepLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
lines = append(lines, line)
}
return lines
}
func TestRunner_ParallelFetchCallsAllDriverMethods(t *testing.T) {
state := newHarness(t)
state.mock.MetricsData = driver.Metrics{CPUPercent: 5.0, HeapBytes: 1024, TotalMemoryBytes: 4096}
state.mock.LogEntries = []driver.LogEntry{
{UnixMillis: 1000, Level: "E", Tag: "test", Message: "boom"},
}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
BundleID: "com.fixture",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
actions := state.mock.Actions()
var hasSnapshot, hasMetrics, hasLogs bool
for _, a := range actions {
switch a.Kind {
case mockdriver.ActionSnapshot:
hasSnapshot = true
case mockdriver.ActionMetrics:
hasMetrics = true
case mockdriver.ActionRecentLogs:
hasLogs = true
}
}
if !hasSnapshot {
t.Error("expected Snapshot call in mock actions")
}
if !hasMetrics {
t.Error("expected Metrics call in mock actions")
}
if !hasLogs {
t.Error("expected RecentLogs call in mock actions")
}
}
// TestRunner_UsesAtomicSnapshot ensures the runner observes a step's UI
// through the paired Snapshot RPC instead of racing two independent
// hierarchy + screenshot reads. The pair must come from one on-device
// frame; a regression to separate calls is what this test catches.
func TestRunner_UsesAtomicSnapshot(t *testing.T) {
state := newHarness(t)
state.mock.ImageData = driver.Image{PNG: []byte("png"), Width: 1, Height: 1}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected at least one step")
}
var snapshotCalls, hierarchyCalls, screenshotCalls int
for _, action := range state.mock.Actions() {
switch action.Kind {
case mockdriver.ActionSnapshot:
snapshotCalls++
case mockdriver.ActionHierarchy:
hierarchyCalls++
case mockdriver.ActionScreenshot:
screenshotCalls++
}
}
if snapshotCalls == 0 {
t.Errorf("expected at least one Snapshot call, got %d", snapshotCalls)
}
// The recorded pair still comes from Snapshot. The standalone hierarchy
// reads are the composition detector (changedOnReread), one per step at
// most, and they are never the source of what the step records.
if hierarchyCalls > summary.Steps {
t.Errorf("expected at most one standalone Hierarchy call per step (runner must observe through Snapshot), got %d over %d steps",
hierarchyCalls, summary.Steps)
}
if snapshotCalls < summary.Steps {
t.Errorf("expected a Snapshot per step, got %d over %d steps", snapshotCalls, summary.Steps)
}
if screenshotCalls != 0 {
t.Errorf("expected zero standalone Screenshot calls (runner must use Snapshot), got %d", screenshotCalls)
}
}
// TestRunner_OneScreenshotPerStep verifies the runner writes a single
// screenshot per step, captured concurrently with hierarchy so the two
// observations describe the same UI moment.
func TestRunner_OneScreenshotPerStep(t *testing.T) {
state := newHarness(t)
state.mock.ImageData = driver.Image{PNG: []byte("fakepng"), Width: 100, Height: 200}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps for screenshot test, got %d", summary.Steps)
}
screenshotDir := filepath.Join(state.writer.Directory(), "screenshots")
for step := 1; step <= summary.Steps; step++ {
path := filepath.Join(screenshotDir, fmt.Sprintf("step-%05d.png", step))
if _, err := os.Stat(path); os.IsNotExist(err) {
t.Errorf("expected screenshot for step %d at %s", step, path)
}
}
entries, err := os.ReadDir(screenshotDir)
if err != nil {
t.Fatal(err)
}
for _, entry := range entries {
if strings.Contains(entry.Name(), "-after") {
t.Errorf("unexpected -after screenshot remains: %s", entry.Name())
}
}
}
// TestRunner_StableTransitionalTreeIsVerified feeds a driver whose hierarchy
// constantly carries two route-level *Screen ids but never changes between
// retry attempts. Such a tree is a settled state that merely matches the
// transitional heuristic, so the runner must verify it (the always-false
// predicate's violation surfaces) instead of skipping the verifier forever.
func TestRunner_StableTransitionalTreeIsVerified(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"root"},"children":[
{"attributes":{"resource-id":"AddAccountScreen"},"children":[]},
{"attributes":{"resource-id":"HomeScreen"},"children":[]}
]}`
state.mock.ImageData = driver.Image{PNG: []byte("fakepng"), Width: 100, Height: 200}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected at least one step")
}
if !containsProperty(summary.Violations, "balanceNonNegative") {
t.Fatalf("expected verifier to run on a stable two-screen tree, got %v", summary.Violations)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if line.Transitional {
t.Errorf("step %d: stable tree must not be marked transitional", line.Step)
}
}
// The early break must still persist the step's screenshot.
screenshotPath := filepath.Join(state.writer.Directory(), "screenshots", "step-00001.png")
if _, err := os.Stat(screenshotPath); err != nil {
t.Errorf("expected screenshot for step 1 at %s: %v", screenshotPath, err)
}
}
// snapshotCrossFade wraps a mock driver so every Snapshot call returns a
// transitional two-screen tree whose JSON differs from the previous call,
// mimicking a genuine cross-fade in flight.
type snapshotCrossFade struct {
*mockdriver.Driver
calls int
}
func (d *snapshotCrossFade) Snapshot(ctx context.Context) (string, driver.Image, error) {
d.calls++
_, image, err := d.Driver.Snapshot(ctx)
hierarchyJSON := fmt.Sprintf(`{"attributes":{"resource-id":"root"},"children":[
{"attributes":{"resource-id":"AddAccountScreen","text":"frame-%d"},"children":[]},
{"attributes":{"resource-id":"HomeScreen"},"children":[]}
]}`, d.calls)
return hierarchyJSON, image, err
}
// TestRunner_GenuineCrossFadeStillRetried pins the existing behavior for real
// transitions: a tree that keeps changing between retry attempts exhausts the
// budget, stays transitional, and the verifier is skipped for the step.
func TestRunner_GenuineCrossFadeStillRetried(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
wrapped := &snapshotCrossFade{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected at least one step")
}
if len(summary.Violations) != 0 {
t.Fatalf("verifier must be skipped on cross-fade steps; got %v", summary.Violations)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
lines := 0
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
lines++
if !line.Transitional {
t.Errorf("step %d: expected transitional=true on every step, got false", line.Step)
}
if len(line.Violations) != 0 {
t.Errorf("step %d: verifier must be skipped, got violations %v", line.Step, line.Violations)
}
}
if lines == 0 {
t.Fatal("expected trace lines, got none")
}
}
// TestRunner_CleanTreeStillVerified is the control: a single-screen hierarchy
// must not be marked transitional and the verifier must still run, surfacing
// the always-false predicate's violation on the onset step.
func TestRunner_CleanTreeStillVerified(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"HomeScreen"},"children":[]}`
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if !containsProperty(summary.Violations, "balanceNonNegative") {
t.Fatalf("expected verifier to surface balanceNonNegative on a clean tree, got %v", summary.Violations)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if line.Transitional {
t.Errorf("step %d: clean tree must not be marked transitional", line.Step)
}
}
}
// snapshotFailFirst wraps a mock driver so the first Snapshot call returns an
// error (mimicking a sidecar timeout while fetching view hierarchy), then
// delegates every subsequent call back to the mock.
type snapshotFailFirst struct {
*mockdriver.Driver
calls int
}
func (d *snapshotFailFirst) Snapshot(ctx context.Context) (string, driver.Image, error) {
d.calls++
if d.calls == 1 {
return "", driver.Image{}, errors.New("Timeout while fetching view hierarchy")
}
return d.Driver.Snapshot(ctx)
}
// TestRunner_NilHierarchyMarksTransitional verifies that when the sidecar's
// hierarchy fetch fails (nil tree), the runner marks the step transitional and
// skips the verifier instead of pushing a nil tree that would crash the spec.
// Subsequent steps with a clean tree still drive the verifier normally.
func TestRunner_NilHierarchyMarksTransitional(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"HomeScreen"},"children":[]}`
wrapped := &snapshotFailFirst{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to verify the first is skipped and the second runs, got %d", summary.Steps)
}
// violationSpec always() => false fires on the first verifier push. With
// step 1's verifier skipped, onset moves to step 2.
if len(summary.Violations) != 1 {
t.Fatalf("expected exactly one onset record, got %d: %v", len(summary.Violations), summary.Violations)
}
if summary.Violations[0].StepIndex != 2 {
t.Errorf("onset step: got %d, want 2 (step 1 verifier skipped due to nil tree)", summary.Violations[0].StepIndex)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
var first traceLine
if err := json.Unmarshal(bytes.SplitN(bytes.TrimSpace(body), []byte("\n"), 2)[0], &first); err != nil {
t.Fatalf("decode first trace line: %v", err)
}
if first.Step != 1 {
t.Fatalf("first trace line step: got %d, want 1", first.Step)
}
if !first.Transitional {
t.Error("first step must be marked transitional when the hierarchy fetch failed")
}
if len(first.Violations) != 0 {
t.Errorf("step 1 must skip the verifier; got violations %v", first.Violations)
}
}
// tapSelectorFailFirst wraps a mock driver so the first TapSelector call
// returns a gRPC DeadlineExceeded error (mimicking a sidecar-side RPC hang),
// then delegates every subsequent call back to the mock.
type tapSelectorFailFirst struct {
*mockdriver.Driver
calls int
}
func (d *tapSelectorFailFirst) TapSelector(ctx context.Context, selector string) error {
d.calls++
if d.calls == 1 {
return status.Error(codes.DeadlineExceeded, "boom")
}
return d.Driver.TapSelector(ctx, selector)
}
// TestRunner_TransientApplyErrorMarksTransitional verifies that a transient
// gRPC error from applyAction (e.g. sidecar RPC deadline) does not kill the
// run: the step is marked transitional, the verifier is skipped for it, and
// the loop continues with the next step running cleanly.
func TestRunner_TransientApplyErrorMarksTransitional(t *testing.T) {
state := newHarness(t)
wrapped := &tapSelectorFailFirst{Driver: state.mock}
var logBuf bytes.Buffer
logger := slog.New(slog.NewTextHandler(&logBuf, &slog.HandlerOptions{Level: slog.LevelWarn}))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 300 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: logger,
})
if err != nil {
t.Fatalf("Run must not return on transient apply error, got %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to prove the loop continued past the failed apply, got %d", summary.Steps)
}
if len(summary.Violations) != 0 {
t.Errorf("transient apply error must not surface as a violation, got %v", summary.Violations)
}
if !strings.Contains(logBuf.String(), "apply error; marking step transitional") {
t.Errorf("expected apply-error WARN log, got %q", logBuf.String())
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
lines := bytes.Split(bytes.TrimSpace(body), []byte("\n"))
var first, second traceLine
if err := json.Unmarshal(lines[0], &first); err != nil {
t.Fatalf("decode first trace line: %v", err)
}
if err := json.Unmarshal(lines[1], &second); err != nil {
t.Fatalf("decode second trace line: %v", err)
}
if first.Step != 1 || !first.Transitional {
t.Errorf("step 1 must be transitional after transient apply error, got step=%d transitional=%v", first.Step, first.Transitional)
}
// The step still records a next_action it never dispatched, so the reason
// has to be on the line or an executed-action count includes it.
var firstSkip struct {
ActionSkipped string `json:"action_skipped"`
}
if err := json.Unmarshal(lines[0], &firstSkip); err != nil {
t.Fatalf("decode first trace line: %v", err)
}
if firstSkip.ActionSkipped != string(actionSkippedApplyError) {
t.Errorf("step 1 action_skipped = %q, want %q", firstSkip.ActionSkipped, actionSkippedApplyError)
}
if len(first.Violations) != 0 {
t.Errorf("transient apply step must have no violations, got %v", first.Violations)
}
if second.Step != 2 || second.Transitional {
t.Errorf("step 2 must run cleanly after the transient step, got step=%d transitional=%v", second.Step, second.Transitional)
}
}
// internalApplyErrorFailFirst wraps a mock driver so the first InputText call
// fails with the bare Internal error the iOS runner's input handler emits
// when it chokes (HTTP 500 with an empty body), then recovers.
type internalApplyErrorFailFirst struct {
*mockdriver.Driver
calls int
}
func (d *internalApplyErrorFailFirst) TapSelector(ctx context.Context, selector string) error {
d.calls++
if d.calls == 1 {
return status.Error(codes.Internal, "UnknownFailure(errorResponse=Request for inputText failed, code: 500, body: )")
}
return d.Driver.TapSelector(ctx, selector)
}
// TestRunner_InternalApplyErrorMarksTransitional pins the policy that a
// one-off device-side failure (e.g. the iOS input handler's bare 500) is
// absorbed as a transitional step instead of killing the run. Persistent
// failure is covered by the consecutive-failure cap.
func TestRunner_InternalApplyErrorMarksTransitional(t *testing.T) {
state := newHarness(t)
wrapped := &internalApplyErrorFailFirst{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 300 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run must not return on a one-off internal apply error, got %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to prove the loop continued, got %d", summary.Steps)
}
}
// TestIsWDADrop_Classification pins the fatal-vs-recoverable boundary: only
// the sidecar's explicit reconnect-failure signal is a drop. A structured
// UNAVAILABLE that happens to embed raw exception text (e.g. ConnectException
// from the original failure) means the sidecar already recovered and the run
// must continue.
// TestIsWDADrop_Classification pins isWDADrop to the exact phrase the sidecar
// throws at its reconnect site (sidecar DriverBackend.kt):
//
// throw IllegalStateException("WDA reconnect failed: $restartErr", cause)
//
// If that message is reworded on the Kotlin side without updating the Go
// matcher, a fatal, unrecoverable WDA drop is misclassified as transient and
// the run burns its budget retrying a dead channel instead of aborting.
func TestIsWDADrop_Classification(t *testing.T) {
// sidecarReconnectFailedMessage mirrors the literal the sidecar emits; the
// matcher's contract is keyed on this exact prefix.
const sidecarReconnectFailedMessage = "WDA reconnect failed"
cases := []struct {
name string
err error
want bool
}{
{
"unavailable with embedded ConnectException is recovered",
status.Error(codes.Unavailable, "connection dropped mid-action; the action may have applied: java.net.ConnectException: Failed to connect to /127.0.0.1:22161"),
false,
},
{
"sidecar reconnect-failed message is a drop",
status.Error(codes.Internal, "java.lang.IllegalStateException: "+sidecarReconnectFailedMessage+": IOSDriverTimeoutException"),
true,
},
{"generic internal is not a drop", status.Error(codes.Internal, "boom"), false},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
if got := isWDADrop(testCase.err); got != testCase.want {
t.Errorf("got %v, want %v", got, testCase.want)
}
})
}
}
// tapSelectorAlwaysUnavailable wraps a mock driver so every TapSelector call
// fails with a transient Unavailable error, mimicking a device whose channel
// never recovers between steps.
type tapSelectorAlwaysUnavailable struct {
*mockdriver.Driver
}
func (d *tapSelectorAlwaysUnavailable) TapSelector(ctx context.Context, selector string) error {
return status.Error(codes.Unavailable, "connection dropped mid-action; the action may have applied")
}
// TestRunner_ConsecutiveTransientApplyFailuresAbort verifies the run fails
// fast once transient apply errors form an unbroken streak instead of burning
// the whole budget on a wedged device.
func TestRunner_ConsecutiveTransientApplyFailuresAbort(t *testing.T) {
state := newHarness(t)
wrapped := &tapSelectorAlwaysUnavailable{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err == nil {
t.Fatal("Run must abort after consecutive transient apply failures")
}
if !strings.Contains(err.Error(), "consecutive failures") {
t.Errorf("expected consecutive-failure abort, got %v", err)
}
// The aborting step returns before it is recorded, so the summary holds
// the steps before the cap-hitting one.
if summary.Steps != maxConsecutiveApplyFailures-1 {
t.Errorf("run must stop at the failure cap, got %d recorded steps", summary.Steps)
}
}
// TestRunner_WaitActionSkipsIdle ensures the runner does not call WaitForIdle
// after a Wait action - the action already provides settling time.
func TestRunner_WaitActionSkipsIdle(t *testing.T) {
const waitSpec = `
import { actions, Wait } from "@sanderling/spec";
globalThis.actions = actions(() => [Wait({ durationMillis: 5 })]);
`
state := newHarnessWithSpec(t, waitSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 150 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
for _, action := range state.mock.Actions() {
if action.Kind == mockdriver.ActionWaitForIdle {
t.Fatalf("Wait action must skip WaitForIdle, got: %v", action)
}
}
}
func mustNewVerifier(t *testing.T) *verifier.Verifier {
t.Helper()
verifierInstance, err := verifier.New()
if err != nil {
t.Fatal(err)
}
return verifierInstance
}
func mustNewTraceWriter(t *testing.T) *trace.Writer {
t.Helper()
writer, err := trace.NewWriter(t.TempDir())
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = writer.Close() })
return writer
}
func containsAction(actions []mockdriver.Action, kind mockdriver.ActionKind, payload string) bool {
for _, action := range actions {
if action.Kind != kind {
continue
}
switch kind {
case mockdriver.ActionLaunch:
if action.BundleID == payload {
return true
}
case mockdriver.ActionTapSelector:
if action.Selector == payload {
return true
}
case mockdriver.ActionTerminate:
return true
default:
return true
}
}
return false
}
func containsProperty(records []ViolationRecord, property string) bool {
for _, record := range records {
if slices.Contains(record.Properties, property) {
return true
}
}
return false
}
func TestRunner_RelaunchesWhenAppLeavesForeground(t *testing.T) {
fastForegroundGate(t)
state := newHarness(t)
// The app is in front when the run starts and a foreign app every time a
// step looks, so the startup gate passes and every step's guard must
// relaunch.
state.mock.ForegroundResults = []string{"app.folio", "com.android.chrome"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
relaunches := 0
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionLaunch && a.BundleID == "app.folio" && !a.ClearState {
relaunches++
}
}
if relaunches == 0 {
t.Fatal("expected runner to relaunch app.folio when foreground escaped, got none")
}
}
func TestRunner_NoRelaunchWhenAppInForeground(t *testing.T) {
state := newHarness(t)
state.mock.ForegroundResults = []string{"app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionLaunch {
t.Fatalf("expected no relaunch while app in foreground, got %v", a)
}
}
}
// TestRunner_WaitsForForegroundBeforeFirstAction verifies the startup gate
// brings the app forward (back-press + relaunch) before any tap fires when the
// device boots showing a system dialog.
func TestRunner_WaitsForForegroundBeforeFirstAction(t *testing.T) {
fastForegroundGate(t)
state := newHarness(t)
// First the device shows a system setup screen, then the app is on top.
state.mock.ForegroundResults = []string{"com.google.android.setupwizard", "app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
actions := state.mock.Actions()
firstLaunch, firstTap := -1, -1
backPressed := false
for i, a := range actions {
switch {
case a.Kind == mockdriver.ActionLaunch && firstLaunch < 0:
firstLaunch = i
case a.Kind == mockdriver.ActionTap && firstTap < 0:
firstTap = i
case a.Kind == mockdriver.ActionPressKey && a.Key == "back":
backPressed = true
}
}
if firstLaunch < 0 {
t.Fatal("expected a relaunch to bring the app forward, got none")
}
if !backPressed {
t.Fatal("expected a back-press to dismiss the system dialog, got none")
}
if firstTap >= 0 && firstLaunch > firstTap {
t.Fatalf("expected the foreground gate (launch at %d) before the first tap (at %d)", firstLaunch, firstTap)
}
}
// TestRunner_WaitsForWindowDrawnBeforeFirstAction verifies the startup gate
// keeps waiting while the app is the resumed activity but its window has not
// drawn yet (a leftover screen still focused). It must poll the focused-window
// signal rather than relaunching, and only proceed once the window names the
// app.
func TestRunner_WaitsForWindowDrawnBeforeFirstAction(t *testing.T) {
fastForegroundGate(t)
state := newHarness(t)
// The app is resumed immediately, but its window lags: the outgoing
// settings screen stays focused for two checks before the app draws.
state.mock.ForegroundResults = []string{"app.folio"}
state.mock.FocusedWindowResults = []string{"com.android.settings", "com.android.settings", "app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
// The gate must have polled the focused window until it named the app,
// i.e. at least the three queued results were consumed.
if calls := state.mock.FocusedWindowCalls(); calls < 3 {
t.Fatalf("expected the gate to poll the focused window until drawn (>=3 calls), got %d", calls)
}
// The resumed app was never a foreign app, so the gate must not relaunch
// or back-press to "fix" a window that simply had not drawn yet.
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionPressKey && a.Key == "back" {
t.Fatal("expected no back-press while waiting for the window to draw")
}
}
}
// TestAwaitForeground_RelaunchesThenWaitsForWindow locks the per-step scope
// guard's recovery: after the app leaves to the launcher, it must relaunch AND
// keep polling the focused window until it names the app, so the step never
// observes or acts while the launcher is on screen (where InputText would land
// in the launcher's type-to-search filter). A single fire-and-forget relaunch,
// which returns before the window draws on a slow physical device, is the bug
// this guards against.
func TestAwaitForeground_RelaunchesThenWaitsForWindow(t *testing.T) {
fastForegroundGate(t)
m := mockdriver.New()
// Foreground: launcher on the first poll (still gone), then the app. Focus:
// the launcher window lingers one extra poll before the app's window draws.
m.ForegroundResults = []string{"com.android.launcher", "app.folio"}
m.FocusedWindowResults = []string{"com.android.launcher", "app.folio"}
logger := slog.New(slog.NewTextHandler(io.Discard, &slog.HandlerOptions{Level: slog.LevelWarn}))
options := Options{BundleID: "app.folio", Driver: m, IdleTimeout: 10 * time.Millisecond}
awaitForeground(context.Background(), options, logger, 7)
relaunches, backs := 0, 0
for _, a := range m.Actions() {
switch {
case a.Kind == mockdriver.ActionLaunch && a.BundleID == "app.folio" && !a.ClearState:
relaunches++
case a.Kind == mockdriver.ActionPressKey && a.Key == "back":
backs++
}
}
if relaunches != 1 {
t.Fatalf("expected exactly one relaunch while the app was gone, got %d", relaunches)
}
if backs != 1 {
t.Fatalf("expected one back-press to dismiss a possible dialog before relaunch, got %d", backs)
}
// The window lagged one poll behind the resumed activity, so the focused
// window must have been queried at least twice before the gate returned.
if calls := m.FocusedWindowCalls(); calls < 2 {
t.Fatalf("expected the guard to poll the focused window until drawn (>=2), got %d", calls)
}
}
func TestClampGestureToSafeArea_KeepsOriginBelowShadeStrip(t *testing.T) {
screen := hierarchy.Bounds{Left: 0, Top: 0, Right: 1080, Bottom: 2400} // marginY = 200
// Origin in the shade strip: the whole segment shifts down by 138, so the
// downward gesture stays downward (447 -> 585) instead of reversing.
fromX, fromY, toX, toY := clampGestureToSafeArea(802, 62, 802, 447, screen)
if fromX != 802 || fromY != 200 || toX != 802 || toY != 585 {
t.Errorf("segment not translated below the shade strip: from=(%d,%d) to=(%d,%d), want from=(802,200) to=(802,585)", fromX, fromY, toX, toY)
}
fromX, fromY, _, _ = clampGestureToSafeArea(5, 1200, 540, 1200, screen)
if fromX != 5 || fromY != 1200 {
t.Errorf("side origin must pass through, got (%d,%d), want (5,1200)", fromX, fromY)
}
fromX, fromY, _, _ = clampGestureToSafeArea(540, 2399, 540, 1200, screen)
if fromX != 540 || fromY != 2399 {
t.Errorf("bottom origin must pass through, got (%d,%d), want (540,2399)", fromX, fromY)
}
_, _, toX, toY = clampGestureToSafeArea(540, 1200, -50, 9999, screen)
if toX != 0 || toY != 2400 {
t.Errorf("off-screen destination not clamped to screen edges: got (%d,%d), want (0,2400)", toX, toY)
}
fromX, _, _, _ = clampGestureToSafeArea(-30, 1200, 540, 1200, screen)
if fromX != 0 {
t.Errorf("off-screen origin x must clamp to 0, got %d", fromX)
}
fromX, fromY, toX, toY = clampGestureToSafeArea(802, 62, 802, 447, hierarchy.Bounds{})
if fromX != 802 || fromY != 62 || toX != 802 || toY != 447 {
t.Error("coordinates must pass through unchanged when screen size is unknown")
}
}
// TestScreenBounds_UsesMaxExtentNotRoot guards the screen-size source: the
// Android hierarchy root reports zero bounds, so the screen rectangle must come
// from the maximum element extent or the gesture clamp silently no-ops.
func TestScreenBounds_UsesMaxExtentNotRoot(t *testing.T) {
tree := &hierarchy.Tree{
Root: &hierarchy.Node{Element: hierarchy.Element{Bounds: hierarchy.Bounds{}}},
Elements: []*hierarchy.Element{
{Bounds: hierarchy.Bounds{}},
{Bounds: hierarchy.Bounds{Left: 0, Top: 0, Right: 1080, Bottom: 2160}},
{Bounds: hierarchy.Bounds{Left: 0, Top: 2268, Right: 1080, Bottom: 2400}},
},
}
got := screenBounds(tree)
if got.Right != 1080 || got.Bottom != 2400 {
t.Fatalf("screenBounds = %+v, want right=1080 bottom=2400", got)
}
}
// TestEnsureForeground_DismissesSystemOverlay locks the shade fix: when the app
// is still the resumed activity but a system overlay (notification shade) holds
// the focused window, the guard must dismiss it with back rather than relaunch
// or act on the obscured app.
func TestEnsureForeground_DismissesSystemOverlay(t *testing.T) {
m := mockdriver.New()
// Resumed activity stays the app; the focused window is the shade.
m.ForegroundResults = []string{"app.folio"}
m.FocusedWindowResults = []string{"com.android.systemui"}
logger := slog.New(slog.NewTextHandler(io.Discard, &slog.HandlerOptions{Level: slog.LevelWarn}))
options := Options{BundleID: "app.folio", Driver: m, IdleTimeout: 10 * time.Millisecond}
got, inScope := ensureForeground(context.Background(), options, logger, 5)
if got != foregroundOverlayDismissed {
t.Fatalf("the guard reported %v, want foregroundOverlayDismissed; "+
"an obscured app is not a relaunched one", got)
}
if !inScope {
t.Fatal("the guard reported the app out of scope; a dismissed overlay leaves " +
"the app resumed, and marking the step unmet would hide the real ones")
}
backs, relaunches := 0, 0
for _, a := range m.Actions() {
switch {
case a.Kind == mockdriver.ActionPressKey && a.Key == "back":
backs++
case a.Kind == mockdriver.ActionLaunch:
relaunches++
}
}
if backs != 1 {
t.Fatalf("expected one back-press to collapse the shade, got %d", backs)
}
if relaunches != 0 {
t.Fatalf("a resumed-but-obscured app must not be relaunched, got %d relaunches", relaunches)
}
}
func TestAppIsForeground(t *testing.T) {
readErr := errors.New("adb read failed")
cases := []struct {
name string
bundleID string
foreground []string
foregErr error
focused []string
focusErr error
want bool
}{
{name: "no bundle id", bundleID: "", foreground: []string{"app.folio"}, want: true},
{name: "foreground unknown", bundleID: "app.folio", foreground: nil, want: true},
{name: "foreground read error", bundleID: "app.folio", foregErr: readErr, want: true},
{name: "foreign foreground", bundleID: "app.folio", foreground: []string{"com.android.chrome"}, want: false},
{name: "app resumed and focused", bundleID: "app.folio", foreground: []string{"app.folio"}, focused: []string{"app.folio"}, want: true},
{name: "app resumed but overlay focused", bundleID: "app.folio", foreground: []string{"app.folio"}, focused: []string{"com.android.systemui"}, want: false},
{name: "app resumed, focus unknown", bundleID: "app.folio", foreground: []string{"app.folio"}, focused: []string{""}, want: true},
{name: "app resumed, focus read error", bundleID: "app.folio", foreground: []string{"app.folio"}, focusErr: readErr, want: true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
m := mockdriver.New()
m.ForegroundResults = tc.foreground
m.ForegroundErr = tc.foregErr
m.FocusedWindowResults = tc.focused
m.FocusedWindowErr = tc.focusErr
options := Options{BundleID: tc.bundleID, Driver: m}
if got := appIsForeground(context.Background(), options); got != tc.want {
t.Errorf("appIsForeground = %v, want %v", got, tc.want)
}
})
}
}
// A system overlay holds focus while the app stays resumed every step, so the
// fixture's id:next tap must never reach the driver.
func TestRunner_SkipsActionWhenOverlayStealsFocusAtApplyTime(t *testing.T) {
state := newHarness(t)
state.mock.ForegroundResults = []string{"app.folio"}
state.mock.FocusedWindowResults = []string{"app.folio", "com.android.systemui"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected the loop to run steps")
}
if containsAction(state.mock.Actions(), mockdriver.ActionTapSelector, "id:next") {
t.Error("apply-time guard failed: a tap fired while a system overlay held focus")
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
var skipped bool
for _, line := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var step struct {
ActionSkipped string `json:"action_skipped"`
}
if err := json.Unmarshal(line, &step); err != nil {
t.Fatalf("decode trace line: %v", err)
}
skipped = skipped || step.ActionSkipped == string(actionSkippedForeground)
}
if !skipped {
t.Errorf("no step recorded action_skipped=%q, so the undispatched action looks executed", actionSkippedForeground)
}
}
// TestRunner_StopOnViolationEndsAtTheFirstViolation pins the gate CI runs on:
// a step budget of 8 against a spec that only violates on the third step must
// end on step 3 and write nothing after it, so the trace's last state is the
// one that produced the violation.
func TestRunner_StopOnViolationEndsAtTheFirstViolation(t *testing.T) {
const thirdStepViolationSpec = `
import { actions, always, extract } from "@sanderling/spec";
let observed = 0;
const tick = extract(() => ++observed);
globalThis.properties = {
staysUnderThree: always(() => tick.current < 3),
};
globalThis.actions = actions(() => []);
`
state := newHarnessWithSpec(t, thirdStepViolationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 8,
StopOnViolation: true,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if !containsProperty(summary.Violations, "staysUnderThree") {
t.Fatalf("expected staysUnderThree to fire, got %v", summary.Violations)
}
if summary.Steps != 3 {
t.Errorf("steps: got %d, want 3 (the run must stop at the violating step, not run the 8-step budget)",
summary.Steps)
}
for _, step := range traceStepIndices(t, state.writer.Directory()) {
if step > summary.Steps {
t.Errorf("trace kept stepping after the violation: found step %d past step %d",
step, summary.Steps)
}
}
}
// TestRunner_WithoutStopOnViolationRunsTheWholeBudget is the other half: the
// default must stay a full-budget fuzz run, so turning the flag on is the only
// thing that shortens a run.
func TestRunner_WithoutStopOnViolationRunsTheWholeBudget(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 4,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 4 {
t.Errorf("steps: got %d, want 4; a violation must not shorten a default run", summary.Steps)
}
}
// traceStepIndices reads every step index the trace recorded, so a test can
// assert on what the run actually wrote rather than on the summary alone.
func traceStepIndices(t *testing.T, directory string) []int {
t.Helper()
file, err := os.Open(filepath.Join(directory, "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
var steps []int
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var line struct {
Step int `json:"step"`
}
if err := json.Unmarshal(scanner.Bytes(), &line); err != nil {
t.Fatalf("trace line decode: %v", err)
}
steps = append(steps, line.Step)
}
if err := scanner.Err(); err != nil {
t.Fatalf("scan trace: %v", err)
}
return steps
}
// siblingCardsSpec is the shape folio's Home screen authors: every account card
// carries the same testTag, and each one is its own tap target.
const siblingCardsSpec = `
import { actions, always, extract, Tap } from "@sanderling/spec";
const cards = extract(state => state.ax.findAll({ testTag: "AccountCard" }));
globalThis.properties = {
alwaysHolds: always(() => true),
};
globalThis.actions = actions(() => cards.current.map(card => Tap({ on: card })));
`
const siblingCardsTree = `{
"attributes": {"resource-id": "root", "bounds": "[0,0,100,300]"},
"children": [
{"attributes": {"testTag": "AccountCard", "text": "Alpha", "bounds": "[0,0,100,100]"}, "clickable": true, "children": []},
{"attributes": {"testTag": "AccountCard", "text": "Beta", "bounds": "[0,100,100,200]"}, "clickable": true, "children": []},
{"attributes": {"testTag": "AccountCard", "text": "Gamma", "bounds": "[0,200,100,300]"}, "clickable": true, "children": []}
]
}`
// TestRunner_SiblingTapsReachTheDriverAtTheirOwnCoordinates is the check the
// whole selector-uniqueness rule exists for. Three cards sharing one testTag
// each produce their own tap; if they reach the driver naming a selector all
// three answer to, resolveCoordinates re-resolves every one of them onto the
// first card and the fuzzer can never open the other two.
func TestRunner_SiblingTapsReachTheDriverAtTheirOwnCoordinates(t *testing.T) {
state := newHarnessWithSpec(t, siblingCardsSpec)
tree, err := hierarchy.Parse(siblingCardsTree)
if err != nil {
t.Fatal(err)
}
if err := state.verifier.PushSnapshot(verifier.SnapshotInput{Tree: tree}); err != nil {
t.Fatal(err)
}
for range 40 {
action, err := state.verifier.NextAction()
if err != nil {
t.Fatalf("NextAction: %v", err)
}
if action.Kind != verifier.ActionKindTap {
t.Fatalf("spec offers taps only, got %q", action.Kind)
}
mustDispatch(t, state.mock, action, tree)
}
tapped := map[string]bool{}
for _, dispatched := range state.mock.Actions() {
if dispatched.Kind != mockdriver.ActionTap {
t.Fatalf("expected coordinate taps only, got %v", dispatched)
}
tapped[fmt.Sprintf("%d,%d", dispatched.X, dispatched.Y)] = true
}
want := []string{"50,50", "50,150", "50,250"}
for _, center := range want {
if !tapped[center] {
t.Errorf("no tap reached the driver at (%s); the driver saw %v", center, slices.Sorted(maps.Keys(tapped)))
}
}
if len(tapped) != len(want) {
t.Errorf("driver saw %d distinct tap points, want %d: %v", len(tapped), len(want), slices.Sorted(maps.Keys(tapped)))
}
}
// tapReachesNoElement wraps a mock driver so every tap reports what the chrome
// driver reports when the action's point holds no element: nothing was
// dispatched, so the app cannot have responded.
type tapReachesNoElement struct {
*mockdriver.Driver
}
func (d *tapReachesNoElement) TapSelector(
_ context.Context,
selector string,
) error {
return fmt.Errorf("%w: %s", driver.ErrGestureUndelivered, selector)
}
func (d *tapReachesNoElement) Tap(_ context.Context, x, y int) error {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
// TestRunner_UndeliveredGestureIsRecordedNotSilentlyClean is the runner half of
// the silent-actuation bug: a run whose every tap reached nothing used to look
// exactly like a run that exercised the app and found no violations. The step
// now names the reason, and the run says how many actions did nothing.
func TestRunner_UndeliveredGestureIsRecordedNotSilentlyClean(t *testing.T) {
state := newHarness(t)
wrapped := &tapReachesNoElement{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 10 * time.Second,
MaxSteps: maxConsecutiveApplyFailures + 2,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf(
"a gesture that reached nothing is not a device fault: %v",
err,
)
}
if summary.Steps <= maxConsecutiveApplyFailures {
t.Fatalf(
"the run must outlive the apply-failure cap, got %d steps",
summary.Steps,
)
}
undelivered := summary.SkippedActions[string(actionSkippedGestureUndelivered)]
if undelivered != summary.Steps {
t.Errorf(
"gesture_undelivered count = %d, want %d (every tap reached nothing)",
undelivered,
summary.Steps,
)
}
var rendered bytes.Buffer
RenderSummary(&rendered, summary, "web")
if !strings.Contains(rendered.String(), "gesture_undelivered") {
t.Errorf(
"the summary must not read clean while nothing was actuated, got:\n%s",
rendered.String(),
)
}
type traceLine struct {
Step int `json:"step"`
ActionSkipped string `json:"action_skipped"`
Transitional bool `json:"transitional"`
}
body, err := os.ReadFile(
filepath.Join(state.writer.Directory(), "trace.jsonl"),
)
if err != nil {
t.Fatal(err)
}
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if line.ActionSkipped != "gesture_undelivered" {
t.Errorf("step %d: action_skipped = %q, want %q",
line.Step, line.ActionSkipped, "gesture_undelivered")
}
if line.Transitional {
t.Errorf(
"step %d: an undelivered gesture leaves the verified screen intact, "+
"so the step must not be transitional",
line.Step,
)
}
}
}
// selectorMatchesNothing wraps a mock driver so every by-selector tap reports
// what the drivers report when the selector names no element on the screen.
type selectorMatchesNothing struct {
*mockdriver.Driver
}
func (d *selectorMatchesNothing) TapSelector(_ context.Context, selector string) error {
return fmt.Errorf("%w: %q", driver.ErrSelectorMatchedNothing, selector)
}
// TestRunner_SelectorThatMatchesNothingIsRecordedApartFromAnUndeliveredGesture
// keeps the two silent paths distinguishable. A selector that named no element
// is a resolution failure, so the step records unresolved_selector, keeps the
// verified screen (not transitional), and does not spend the apply-failure
// budget that a wedged device is meant to exhaust.
func TestRunner_SelectorThatMatchesNothingIsRecordedApartFromAnUndeliveredGesture(t *testing.T) {
state := newHarnessWithSpec(t, absentSelectorSpec)
wrapped := &selectorMatchesNothing{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 10 * time.Second,
MaxSteps: maxConsecutiveApplyFailures + 2,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("a selector that matched nothing is not a device fault: %v", err)
}
if summary.Steps <= maxConsecutiveApplyFailures {
t.Fatalf("the run must outlive the apply-failure cap, got %d steps", summary.Steps)
}
if got := summary.SkippedActions[string(actionSkippedUnresolvedSelector)]; got != summary.Steps {
t.Errorf("unresolved_selector count = %d, want %d", got, summary.Steps)
}
if got := summary.SkippedActions[string(actionSkippedGestureUndelivered)]; got != 0 {
t.Errorf("gesture_undelivered count = %d, want 0: nothing was dispatched to a point", got)
}
var rendered bytes.Buffer
RenderSummary(&rendered, summary, "android")
if !strings.Contains(rendered.String(), "unresolved_selector") {
t.Errorf("the summary must name the actions that found no target, got:\n%s", rendered.String())
}
type traceLine struct {
Step int `json:"step"`
ActionSkipped string `json:"action_skipped"`
Transitional bool `json:"transitional"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if line.ActionSkipped != "unresolved_selector" {
t.Errorf("step %d: action_skipped = %q, want %q",
line.Step, line.ActionSkipped, "unresolved_selector")
}
if line.Transitional {
t.Errorf("step %d: nothing was dispatched, so the step must not be transitional", line.Step)
}
}
}
// TestApplyAction_TapAboveTheViewportReachesTheDriver is the other half of the
// off-screen reach fix. The web host names an unnamed candidate by coordinates
// alone, and a candidate the growing document pushed above the fold carries a
// negative y; dropping it here denies the driver the scroll that would bring it
// back, so reach below the fold works and reach above it does not.
func TestApplyAction_TapAboveTheViewportReachesTheDriver(t *testing.T) {
mock := mockdriver.New()
mustDispatch(t, mock, verifier.Action{
Kind: verifier.ActionKindTap,
X: 622,
Y: -208,
}, nil)
dispatched := mock.Actions()
if len(dispatched) != 1 {
t.Fatalf("driver saw %d actions, want 1: %v", len(dispatched), dispatched)
}
if dispatched[0].Kind != mockdriver.ActionTap ||
dispatched[0].X != 622 || dispatched[0].Y != -208 {
t.Errorf("driver saw %v, want a tap at (622,-208)", dispatched[0])
}
}
// TestApplyAction_TapAboveTheViewportKeepsTheUndeliveredReport holds the
// distinction the fix must not collapse: a point the driver cannot put an
// element under is still a failure, reported as ErrGestureUndelivered rather
// than as an action the runner declined to try.
func TestApplyAction_TapAboveTheViewportKeepsTheUndeliveredReport(t *testing.T) {
drv := &tapReachesNoElement{Driver: mockdriver.New()}
skipped, err := applyAction(context.Background(), drv, verifier.Action{
Kind: verifier.ActionKindTap,
X: 622,
Y: -208,
}, nil)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Errorf("err = %v, want ErrGestureUndelivered", err)
}
if skipped != "" {
t.Errorf("skip reason = %q, want none: the driver was called", skipped)
}
}
// TestRenderSummary_NamesTheActionsThatNeverReachedTheApp is the report half of
// the silent-actuation class. A run that chose an action every step and dropped
// every one of them printed the same "no violations" as a run that exercised
// the app, because the reasons lived only in the trace and a warn line.
func TestRenderSummary_NamesTheActionsThatNeverReachedTheApp(t *testing.T) {
state := newHarnessWithSpec(t, zeroWaitSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 5 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
reason := string(actionSkippedZeroDurationWait)
if summary.SkippedActions[reason] != 3 {
t.Errorf("SkippedActions[%s] = %d, want 3", reason, summary.SkippedActions[reason])
}
var rendered bytes.Buffer
RenderSummary(&rendered, summary, "web")
want := "3 action(s) never reached the app: " + reason + " 3"
if !strings.Contains(rendered.String(), want) {
t.Errorf("summary must carry %q, got:\n%s", want, rendered.String())
}
var clean bytes.Buffer
RenderSummary(&clean, Summary{Steps: 3}, "web")
if strings.Contains(clean.String(), "never reached the app") {
t.Errorf("a run that dropped nothing must not carry the line, got:\n%s", clean.String())
}
}
// wedgedTapSelector never answers the first by-selector tap, which is what a
// driver call that has stopped returning looks like from the step loop.
type wedgedTapSelector struct {
*mockdriver.Driver
calls int
}
func (d *wedgedTapSelector) TapSelector(ctx context.Context, selector string) error {
d.calls++
if d.calls == 1 {
<-ctx.Done()
return ctx.Err()
}
return d.Driver.TapSelector(ctx, selector)
}
// Duration is a loop condition checked between steps, so an action that never
// returns held the run for as long as the process lived and only a kill from
// outside ended it. The step is what must fail, under a reason of its own: a
// wedge is not the gesture that reached nothing and not the action the runner
// declined to dispatch, and an analysis that cannot tell them apart cannot say
// whether the device answered at all.
func TestRunner_AnActionThatNeverReturnsFailsTheStepNotTheRun(t *testing.T) {
state := newHarness(t)
previousApplyTimeout := applyTimeout
applyTimeout = 50 * time.Millisecond
t.Cleanup(func() { applyTimeout = previousApplyTimeout })
wrapped := &wedgedTapSelector{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run must outlive an action that never returns, got %v", err)
}
if summary.Steps != 3 {
t.Fatalf("Steps = %d, want 3: the wedge costs one step, not the run", summary.Steps)
}
lines := readTraceLines(t, state.writer.Directory())
if lines[0].ActionSkipped != string(actionSkippedApplyTimeout) {
t.Errorf("step 1 action_skipped = %q, want %q",
lines[0].ActionSkipped, actionSkippedApplyTimeout)
}
if !lines[0].Transitional {
t.Error("step 1 must be transitional: the action was dispatched and its effect is unknown")
}
for _, line := range lines[1:] {
if line.ActionSkipped != "" {
t.Errorf("step %d action_skipped = %q, want the wedge confined to the step that wedged",
line.Step, line.ActionSkipped)
}
}
if summary.SkippedActions[string(actionSkippedApplyTimeout)] != 1 {
t.Errorf("SkippedActions = %v, want one %s", summary.SkippedActions, actionSkippedApplyTimeout)
}
}
// snapshotFailThenEmpty fails the first observation outright and answers the
// second with a dump holding no elements at all.
type snapshotFailThenEmpty struct {
*mockdriver.Driver
calls int
}
func (d *snapshotFailThenEmpty) Snapshot(ctx context.Context) (string, driver.Image, error) {
d.calls++
switch d.calls {
case 1:
return "", driver.Image{}, errors.New("Timeout while fetching view hierarchy")
case 2:
return "", driver.Image{}, nil
}
return d.Driver.Snapshot(ctx)
}
// A step that read nothing because the read failed and a step that read a
// screen with nothing on it are recorded identically as a nil tree, so a run
// that observed nothing at all reports exactly like a run that observed an
// empty app and found no violation in it.
func TestRunner_AFailedObservationIsNotAnObservationOfAnEmptyScreen(t *testing.T) {
state := newHarness(t)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"HomeScreen"},"children":[]}`
wrapped := &snapshotFailThenEmpty{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 3 {
t.Fatalf("Steps = %d, want 3", summary.Steps)
}
lines := readTraceLines(t, state.writer.Directory())
if !strings.Contains(lines[0].ObservationError, "Timeout while fetching view hierarchy") {
t.Errorf("step 1 observation_error = %q, want the driver's own failure",
lines[0].ObservationError)
}
if lines[1].ObservationError != "" {
t.Errorf("step 2 observation_error = %q, want none: the screen was read and held no elements",
lines[1].ObservationError)
}
if lines[2].ObservationError != "" {
t.Errorf("step 3 observation_error = %q, want none", lines[2].ObservationError)
}
if summary.FailedObservations != 1 {
t.Errorf("FailedObservations = %d, want 1", summary.FailedObservations)
}
var rendered bytes.Buffer
RenderSummary(&rendered, summary, "android")
if !strings.Contains(rendered.String(), "1 step(s) observed nothing") {
t.Errorf("summary hides the failed observation: %q", rendered.String())
}
}