record what the model picker did, and make both policies see the same actions (#74)

* feat(llmclient): parse usage and the served model

An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record one typed outcome per model-driven step

llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): expose the step a snapshot was observed at

It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a guard-skipped step is no longer a silent log line

The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): record when a chosen action was never dispatched

A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): document llm-calls.jsonl

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(analyze): survival analysis over campaign directories

Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): select the candidate label source

Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(runner): thread the label source to the model picker

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the label source as arm membership

Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --label-source

Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): dedup candidates by what they execute, not how they read

The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): report every action that was chosen and never dispatched

applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): the echo guard admits a repeated description

Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): count dispatched actions, not steps

A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): lower authored actions the way the seeded arm does

The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): decode a container-only scroll

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): carry the container on an authored scroll

serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): both policies must dispatch the same authored action

Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(hierarchy): match identifiers by role prefix

idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.

* feat(chrome): translate idPrefix to a starts-with id match

* feat(spec): match idPrefix in the web runtime

The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.

* feat(sidecar): match idPrefix in the tap-by-selector path

* docs(manual): document the idPrefix selector

* feat(replay-ui): render idPrefix targets as a prefix tag

* fix(spec): read the injected seed per call

Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.

* test(chrome): compare both selector matchers over one live page

Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.

* fix(hierarchy): give id and desc one meaning in both selector forms

The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.

* test(hierarchy): pin both selector forms and the unknown-key report

* feat(verifier): fail the spec on a selector key that cannot match

An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.

* feat(spec): reject an unknown object-selector key in the web runtime

Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.

* test(spec): pin the unknown-key diagnostic to one text

The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.

* fix(spec): match a merged label by its leading name on web too

The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.

* test(chrome): drive the live-page parity test through both selector forms

* docs(manual): document object-selector key rules

* feat(spec): refuse a multi-item authored sampler while enumerating

from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets

Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): abort on a candidate enumeration that refused

Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(spec): refuse a multi-value generator while enumerating

integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(verifier): setup still draws, and the seeded stream is unmoved

Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): value generators are refused under the model policy too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio): enumerate authored targets and values instead of sampling

Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): enumerate authored targets and values, declare the llm generator

The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: minimal changes, self-documenting code, tests as first-class

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): key web attrs by the names the markup writes

attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name a web handle by the same ladder as a tree element

The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): attrs carries raw attribute names on web too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): confirm focus moved before typing

InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): signal a timed-out run so it reaps its sidecar

CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* perf(runner): confirm focus only when another element holds it

Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): record both clocks a run was measured on

Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide per-hour rates by time actually worked

A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(campaign): wait for the trap instead of racing it

The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(ltl): keep the authored window on a step-bounded obligation

reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(ltl): pin that a slow policy does not fail on time alone

Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(spec): guard the step unit on the authoring surface

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): bound the reachability properties by steps

At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): a step bound counts observations, not runner steps

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): make the label source a cell dimension

A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): record the label source in the manifest

A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): name a web field by its hint, not its CSS class

visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): answer clickable for an element reached through ax

The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: every target runs on this machine, so start one rather than skip it

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a selector the tree cannot resolve is not a focus failure

otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit data-testid so both resolvers name the same element

The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name an element only when the selector names it alone

ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): hold the V8 host to the same naming gate

elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): sibling taps reach the driver at their own coordinates

Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): an ambiguous name loses to the coordinates it was built from

Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): record an element-valued extractor instead of dropping it

An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.

* test(verifier): an unrecordable extractor value is reported, not dropped

* test(runner): element-valued extractors reach the trace

* docs(spec-language): say what a trace records for an element-valued extractor

* docs(claude): add delegation and record-keeping sections

delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.

* feat(driver): declare undelivered-action errors and three optional capabilities

ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.

* feat(driver): add escape to the pressKey surface

escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.

* fix(ios): refuse a gesture the screen has no surface under

the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.

* feat(ios): derive scrollable from the snapshot's tree depth

the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.

* fix(sidecar): stop dropping gestures, selectors and keys in silence

a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.

* fix(sidecar): map the driver's refusals onto the gesture errors

OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.

* fix(selectors): resolve text to the innermost match and scan the root in both forms

an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.

* feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag

a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.

* fix(chrome): emit every markup attribute and read checked and selected off the property

the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.

* fix(chrome): scroll a gesture point into view and dispatch trusted input

getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.

* feat(chrome): read the page's exceptions and navigations, and hold the picker state across them

a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.

* feat(trace): version each step and record its logs, exceptions and navigations

a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.

* feat(runner): bound every device call and record the actions that never reached the app

observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.

* feat(verifier): expose extractor names and rebuilt property formulas

an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.

* feat(testrun): expose the seeded bundle a run loaded

BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity.

* feat(tracecorpus): load recorded runs for offline measures

reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.

* refactor(seedspec): move seed spec parsing out of the campaign command

the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.

* feat(analyze): time an event at the step it was detected and report the quartiles

an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.

* feat(analyze): add the seed-paired signed-rank comparison and record the holm family

--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.

* test(analyze): recover planted effects through the tool's own entry point

a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.

* feat(label-coverage): report the addressable share of an app's interactive surface

reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.

* feat(exploration-reach): count the distinct structural states a stored run visited

the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.

* feat(defect-identity): count distinct defects across stored runs

a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.

* feat(oracle-reduction): replay stored traces under four reduced oracles

re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.

* feat(implementation-sweep): run one campaign against every implementation of a requirement

installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.

* feat(corpus-sweep): run one specification against a served corpus of implementations

same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.

* docs(manual): document innermost text matching, escape and the web scroll verb

text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.

* test(browser): assert an uncaught page exception reaches the trace

the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.

* feat(trace): a step can name the precondition it could not meet

A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.

* fix(runner): budget the foreground gate in time, not in polls

Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.

* test(runner): the gate keeps looking until its budget runs out

Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.

* feat(campaign): count the runs that were never in the app

A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.

* docs(triage): name the trace field a run that never started leaves

* fix(selectors): tag names the whole tag, not a substring of it

matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.

* test(chrome): both resolvers agree on tag where a container's name contains its child's

* fix(make): build the binary instead of matching the build directory

build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.

* feat(verifier): expose the property names a loaded spec registered

* feat(testrun): refuse a run against a spec that registers no properties

A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.

* feat(cli): --allow-no-properties opts a run out of the refusal

* docs(cli): document --allow-no-properties

* feat(bundle-check): fail a spec that bundles but registers no properties

* test(bundle-check): cover the zero-property refusal and pin the reported bundle

* feat(folio-web): predicates for counting commits against submit actions

* feat(folio-web): judge one commit per submit over a home-card window

Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.

* fix(folio-web): keep submit live for 400ms after saving

Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.

* feat(confusion-matrix): score the checker against a blind reviewer

Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.

* test(confusion-matrix): reject malformed inputs and keep missing data out of the cells

* test(confusion-matrix): cover cell assignment, precision and recall

* fix(chrome): focus descends into the shadow root

document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.

* test(implementation-sweep): supply the binaries the missing-binary test does not test

resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.

* fix(replay-ui): read data-* attributes by their markup names

The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.

* fix(web-runtime): focus descends into the shadow root here too

The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.

* fix(implementation-sweep): name every missing binary, in flag order

Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.

* fix(chrome): focus follows the caret to the field it types into

Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.

* fix(corpus-sweep): name every missing binary, in flag order

Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.

* fix(web-runtime): a handle answers editable for itself, not its container

isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.

* test(chrome): a hinted field is not named by its css class

The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.

* test(chrome): the handle and the enumeration agree on editable too

The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.

* fix(web-runtime): focus follows the caret to the field it types into

Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.

* fix(campaign): name every missing required flag, in flag order

Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.

* fix(corpus-sweep): name every missing required flag, in flag order

* fix(implementation-sweep): name every missing required flag, in flag order

* fix(confusion-matrix): name every missing required flag, in flag order

* ci: pin the idb-companion tap to the formula the companion is staged from

The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.

* fix(campaign): refuse to start on a device that is not there

A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): quarantine a device that keeps failing fast

Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the device a run executed on

meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* ci: let a restored ios bundle survive make's mtime check

The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.

* fix(confusion-matrix): a campaign that died is missing data, not a true negative

The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.

* fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.

* fix(runner): a source that was asked and handed nothing says so

NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.

* feat(testrun): refuse a run that dispatched none of its actions

Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.

* docs(cli): document --label-source

* docs(spec-language): name the hintText selector's host divergence

The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.

* feat(bundle-check): --allow-no-properties opts out of the refusal

The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.

* test(verifier): an unreadable committed fixture fails, it does not skip

The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.

* fix(testrun): the refusal asks whether the generator drove, not whether anything did

A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.

* feat(runner): the summary says how many steps the generator drove

A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.

* fix(testrun): an ios run records the simulator it executed on

Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.

* fix(campaign): the action count leaves the setup's login out on a model run

Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.

* feat(hierarchy): an element reports whether it masks what is typed into it

ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".

* fix(verifier): a secure field's typed value never reaches the record

A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.

* fix(runner): a secure field's value does not reach state.lastAction either

folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.

* fix(trace): an action names the generator that produced it

The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.

* fix(defect-identity): degrade a redacted origin action to its selector

The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.

* fix(campaign): a record always says how many actions named no producer

An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.

* fix(analyze): read how much of a record's action count names no producer

A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.

* fix(analyze): refuse to compare attributed and unattributed denominators

One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.

* fix(analyze): mark an action count of unknown provenance in the report

* docs(manual): what an action count with no producer means for a rate

* fix(folio): install through adb so a remote adb server works

Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.

* docs(folio): say how to point just test at a remote adb server

* test(conformance): the g4 fixture holds what a redacted android trace holds

Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.

* fix(testrun): a recorded violation outranks the dead-run refusal

A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.

* fix(testrun): the dead-run refusal gets its own opt-out

--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.

* feat(cli): --allow-no-generator-actions

The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.

* refactor(analyze): open the log-rank up to a weight on the risk set

The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.

* feat(analyze): add the gehan generalized wilcoxon test

The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.

* fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.

* fix(conformance): g4 reads a doubling off the observed field value

The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.

* fix(spec): a secure selector names the password field on web

secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.

* test(chrome): resolve the secure selector on both matchers

the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.

* test(chrome): compare the secure fact across both producers

it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.

* docs(manual): state what a secure selector matches

* test(conformance): g4 keeps checking past an input typed at coordinates

An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.

* fix(conformance): g4 skips an input that names no field

jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.

* test(browser): the exit code a dead run and a violated one actually leave

Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.

* fix(spec): keep a secure selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.

* fix(analyze): score a seed pair by which run outlived the other

The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.

* docs(analyze): name the tests the tool actually runs

The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.

* docs(manual): exit 1 also means a run that holds no verdict

And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.

* docs(skills): quote the summary line the runner prints now

The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.

* test(conformance): g4 sees a doubling appended to what the field held

Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.

* refactor(analyze): hoist the sign test's loop bound

* fix(conformance): g4 reads a doubling out of what the field grew by

The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.

* fix(analyze): write an undefined paired p-value as null, not as NaN

A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.

* fix(spec): a boolean state selector names what the live element reports

clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.

* fix(spec): keep a tag selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.

* fix(chrome): state every boolean flag the dump can state

internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.

* test(chrome): resolve the five state selectors on both matchers

the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.

* test(chrome): compare checked, selected and focused across both producers

the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.

* fix(spec): keep a selector out of the head subtree

the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.

* docs(manual): state what the other boolean state selectors match

* fix(folio): refuse to install and fuzz a device nobody named

adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.

* docs(folio): state that android recipes need ANDROID_DEVICE

* fix(spec): and text with the keys written beside it

a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.

* test(spec): pin text against the key beside it in either order

object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.

* test(chrome): compare a compound text selector across both matchers

one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.

* docs(manual): state how text combines with the key beside it

the object selector section said every pair must match without saying
where the innermost rule then lands.

* fix(hierarchy): reach the class attribute through className

className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.

* test(chrome): compare className across both matchers

one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.

* test(spec): pin className and class on the same elements

this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.

* docs(manual): list className among the cross-platform aliases

the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.

* fix(hierarchy): reach the accessible label through every name for it

label and accessibilityLabel aliased onto accessibilityText alone, which
only the ios sidecar writes, and alias expansion is ONE level: the hop
from accessibilityText to content-desc was never taken, so both keys
matched nothing on android and on the chrome dump, which write the fact
under content-desc. ariaLabel and contentDescription aliased onto
nothing at all and matched nothing anywhere.

The web runtime resolves all four against the live DOM, so a selector
naming a field this way found it on one host and no element at all on
the other. The keys are accepted, so no unknown-key error fires, and a
property over the element that was never found passes having checked
nothing.

Each name lists both keys rather than chaining through accessibilityText:
transitive expansion would silently widen every existing key at once.

* fix(hierarchy): reach a web test tag through testTag and testID

Compose for Web writes a test tag as data-testid, which is what the web
runtime resolves both names against. testTag aliased onto the three
identifier keys and not that one, and testID aliased onto nothing at
all, so a tag the web runtime found on every row of a list named no
element here and every property over it passed vacuously.

* fix(spec): resolve the identifier, label and class aliases against the DOM

identifier, accessibilityIdentifier, accessibilityText and elementType
are the names ios writes four facts under, and internal/hierarchy
aliases each onto the key the other producers write. This table listed
none of them, so each fell through to a raw attribute lookup and built
[accessibilityIdentifier="summary_card"], which no element carries.

Every one of them resolved against the dump on the goja host and named
nothing here. The keys are accepted, so no unknown-key error fires, and
a property over the element that was never found passes having checked
nothing.

* fix(spec): an editable or scrollable selector names what this host derives

Both facts are derived from the live element rather than written by the
markup, and matching them as attributes built [editable="true"], which
no page carries. Both resolve against the dump on the goja host, so a
spec naming a field or a scroll container that way found it there and no
element at all here, with no unknown-key error to say so.

Each reads the same function the fact is derived with, so a selector
cannot name an element this host calls something else: the handle, the
picker's target list and the editable selector all go through
isEditable, and scrollable reads the overflow test collectTargets reads.

scrollable false names nothing rather than every element that does not
scroll: both producers state the fact only where it holds, the way an
element that is no field at all answers to neither value of secure.

* test(chrome): compare the alias keys and the two derived facts

one page, both resolvers, the ten names that resolved on one host only.
each alias is asked beside the key it resolves through, so the pair is
pinned to the same elements rather than each to itself.

the page grows a container that overflows its box and a neighbour that
does not, because scrollable is derived from the box: without one the
only scrolling element on the page is the document root, whose answer
moves with the window.

* fix(hierarchy): bounds is a raw attribute, not a cross-platform key

Every native dump writes the rectangle out as a string under bounds, and
no DOM element carries an attribute of that name, so the key resolved
against the dump and matched nothing on web on every page there is. It
is accepted, so no unknown-key error said so, and no mapping can be
invented for it: there is no DOM fact to map it to.

Off the accepted list the web runtime raises the unknown-key error
instead of matching nothing in silence, and the key still resolves
wherever a producer writes it, through the escape hatch every other raw
attribute already uses: a key some element carries is a key that can
match, on both sides.

* docs(manual): state which attribute each alias reads, and what bounds is

the table listed neither name for the accessible label that a web page
writes, nor the key a web test tag lands on, and said nothing about
elementType. editable and scrollable are boolean states like the rest,
and scrollable is the one of them the platforms state only where it
holds. bounds is a raw driver attribute rather than an accepted key.

* docs(hierarchy): the package doc names every key an alias reaches

it described the alias table as it stood before the label and test-tag
names reached the keys android and web write, and said nothing about
expansion being one level deep, which is why each name has to list every
key rather than hop through another alias.

* fix(spec): a hint selector names the ladder both producers derive

hintText and placeholderValue are the accessible-name ladder, derived
from the live element, and compiling them to [placeholder="..."] made
them name the wrong field or none at all. A field labelled by an
aria-label or a bound <label> carries no placeholder, so it resolved
against the dump on the goja host and reached nothing here; one carrying
both answered to its placeholder here where the dump answers to its
aria-label, which lands a find on an element nobody named.

Both keys read the same fieldHint elementHandle and the hierarchy dump
(internal/driver/chrome/driver.go) derive the fact with, so a selector
cannot name a field this host calls something else. An empty hint names
nothing rather than everything that is no field: both producers write
the fact only where the ladder answered.

placeholder stays the attribute the markup writes, which is what the
dump carries under that name too, so a field whose hint is something
else still answers to it on both hosts.

* fix(chrome): a hint target is not tapped by the placeholder attribute

TapSelector is a third resolver, and it built [placeholder="..."] for
hintText and placeholderValue too. Now that both matchers read the
accessible-name ladder, that CSS names a field whose hint is its
aria-label and whose placeholder happens to carry the value, which is an
element neither matcher named.

No CSS says what the ladder says, so both keys fall through to a match
that reaches nothing and the step fails naming the selector, the way
every other derived key in this file already does. A selector reaches
here only where the dump resolved it to no coordinates at all.

* test(chrome): compare the hint keys and placeholder across both matchers

The page gains four fields that differ one rung at a time: a bound
label, a placeholder, a placeholder an aria-label outranks, and the name
the form gives the field. Only the placeholder rung was reachable
before, so hintText and placeholderValue named a field on the goja host
and no element at all on web for the other three, and named the field
here and nothing there for the rung the ladder passed over.

placeholder was measured empty on both hosts because nothing on the page
carried the attribute, which said nothing about it. It now names the
field the markup wrote it on and not the field whose hint is its
aria-label.

The third resolver reads the same selectors: what TranslateStringSelector
builds for a hint key has to match nothing over CDP rather than the field
carrying the value as a placeholder.

* docs(manual): both hosts read the hint ladder, placeholder is the attribute

The web section said the hintText key does not read the ladder on both
hosts and told authors to select such a field by attrs.hintText instead.
Both hosts read it now, so that instruction is gone rather than left
standing beside a newer sentence.

placeholder is stated as the attribute the markup writes and nothing
more, the tap path is stated as failing by name where no CSS says what
the ladder says, and the alias table gains the row it was missing.
This commit is contained in:
pj authored and GitHub committed 2026-08-19 11:07:49 +05:30
1 parent 95a19fcc1e
commit 9b4ff5f247
243 files changed
+33472 -1062

No files matched your search

+6 -2
View File
@@ -23,7 +23,7 @@ import (
// - else if exactly one AVD exists locally, boot it;
// - else fail with a helpful message listing the available AVDs.
func EnsureDevice(ctx context.Context, serial, avdName string, stdout io.Writer) error {
devices, err := listAdbDevices(ctx)
devices, err := ConnectedDevices(ctx)
if err != nil {
return fmt.Errorf("list adb devices: %w", err)
}
@@ -358,7 +358,11 @@ var standardSDKRoots = []string{
"/usr/local/share/android-commandlinetools",
}
func listAdbDevices(ctx context.Context) ([]string, error) {
// ConnectedDevices lists the serials adb reports as online. It goes through
// the adb CLI so the ADB_SERVER_SOCKET / ANDROID_ADB_SERVER_ADDRESS pair the
// process was started with selects the same server every other adb call in the
// run talks to, rather than assuming a server on this machine.
func ConnectedDevices(ctx context.Context) ([]string, error) {
adb, err := AdbBinary()
if err != nil {
return nil, err
+377 -35
View File
@@ -11,6 +11,7 @@ import (
"sync"
"time"
"github.com/chromedp/cdproto/cdp"
"github.com/chromedp/cdproto/input"
"github.com/chromedp/cdproto/network"
"github.com/chromedp/cdproto/page"
@@ -31,6 +32,14 @@ type Driver struct {
logsMu sync.Mutex
logs []driver.LogEntry
navigationsMu sync.Mutex
navigations []driver.Navigation
// pickerState is the seeded picker's draw position, held here rather than
// in the page: a navigation replaces the page's runtime, and a runtime that
// starts over restarts the seed's stream at its first draw.
pickerState string
}
// New creates a new ChromeDriver. Call Terminate when done.
@@ -97,6 +106,19 @@ func New() *Driver {
d.logsMu.Unlock()
})
chromedp.ListenTarget(tabCtx, func(ev any) {
e, ok := ev.(*page.EventFrameNavigated)
if !ok || e.Frame == nil || e.Frame.ParentID != "" {
return
}
d.navigationsMu.Lock()
d.navigations = append(d.navigations, driver.Navigation{
URL: e.Frame.URL,
UnixMillis: time.Now().UnixMilli(),
})
d.navigationsMu.Unlock()
})
return d
}
@@ -136,9 +158,22 @@ func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, _
})()`, &dims)); err == nil && dims[0] > 0 && dims[1] > 0 {
_ = chromedp.Run(runCtx, chromedp.EmulateViewport(dims[0], dims[1]))
}
// The opening navigation is the harness arriving, not the app navigating.
_, _ = d.Navigations(ctx)
return nil
}
// Navigations returns the document-replacing main-frame navigations seen since
// the last call and forgets them. Each one replaced the page's runtime, which
// is what separates "the app reloaded" from "the picker repeated itself".
func (d *Driver) Navigations(context.Context) ([]driver.Navigation, error) {
d.navigationsMu.Lock()
defer d.navigationsMu.Unlock()
drained := d.navigations
d.navigations = nil
return drained, nil
}
// clearState wipes the target's stored data before the application loads.
// Script cannot do it: the tab still sits on about:blank, whose opaque origin
// denies storage access, so `localStorage.clear()` throws SecurityError and
@@ -205,14 +240,86 @@ func (d *Driver) Terminate(_ context.Context) error {
return nil
}
// pointInViewScript scrolls a point the caller took from the hierarchy back
// inside the viewport and reports where to dispatch at, plus whether anything
// is there to receive it.
//
// The emulated viewport is sized once at launch, but getBoundingClientRect goes
// on reporting elements the growing document has pushed below it, so the two
// disagree the moment an app adds content. Input coordinates are
// viewport-relative: a click below the fold is hit-tested to the document root,
// which delivers it to <html> and never to the element the caller named. No
// error is raised on any layer, so the step reads as an action that landed and
// changed nothing.
//
// Only a point outside the viewport is moved, so a gesture that already had a
// reachable target dispatches exactly where it did before.
const pointInViewScript = `
(function(x, y) {
const root = document.scrollingElement || document.documentElement;
let shiftX = 0, shiftY = 0;
if (x < 0 || x >= window.innerWidth) shiftX = Math.round(x - window.innerWidth / 2);
if (y < 0 || y >= window.innerHeight) shiftY = Math.round(y - window.innerHeight / 2);
if (shiftX || shiftY) {
const fromX = root.scrollLeft, fromY = root.scrollTop;
root.scrollLeft = fromX + shiftX;
root.scrollTop = fromY + shiftY;
shiftX = root.scrollLeft - fromX;
shiftY = root.scrollTop - fromY;
}
const atX = x - shiftX, atY = y - shiftY;
return [atX, atY, document.elementFromPoint(atX, atY) ? 1 : 0];
})(%d, %d)`
// pointInView returns the point to dispatch a gesture at for the point the
// caller named, having scrolled it into view. It fails with
// driver.ErrGestureUndelivered when no scroll can put an element under it.
func pointInView(runCtx context.Context, x, y int) (int, int, error) {
var point [3]int
script := fmt.Sprintf(pointInViewScript, x, y)
if err := chromedp.Run(runCtx, chromedp.Evaluate(script, &point)); err != nil {
return 0, 0, err
}
if point[2] == 0 {
return 0, 0, fmt.Errorf(
"%w: (%d,%d)",
driver.ErrGestureUndelivered,
x,
y,
)
}
return point[0], point[1], nil
}
func (d *Driver) Tap(ctx context.Context, x, y int) error {
runCtx, cancel := d.runCtx(ctx)
defer cancel()
atX, atY, err := pointInView(runCtx, x, y)
if err != nil {
return err
}
return chromedp.Run(runCtx,
chromedp.MouseClickXY(float64(x), float64(y)),
chromedp.MouseClickXY(float64(atX), float64(atY)),
)
}
// requireSelectorMatch reports driver.ErrSelectorMatchedNothing when the
// selector names no node on the page right now. chromedp.Click waits instead,
// so without this the caller hears a deadline (or nothing at all) for an action
// that had no target.
func requireSelectorMatch(runCtx context.Context, target, selector string) error {
var nodes []*cdp.Node
if err := chromedp.Run(runCtx,
chromedp.Nodes(target, &nodes, chromedp.BySearch, chromedp.AtLeast(0)),
); err != nil {
return err
}
if len(nodes) == 0 {
return fmt.Errorf("%w: %q", driver.ErrSelectorMatchedNothing, selector)
}
return nil
}
func (d *Driver) TapSelector(ctx context.Context, selector string) error {
runCtx, cancel := d.runCtx(ctx)
defer cancel()
@@ -222,6 +329,9 @@ func (d *Driver) TapSelector(ctx context.Context, selector string) error {
// reject it loudly if it isn't a valid CSS selector.
target = selector
}
if err := requireSelectorMatch(runCtx, target, selector); err != nil {
return err
}
if isXPath {
return chromedp.Run(runCtx, chromedp.Click(target, chromedp.NodeVisible, chromedp.BySearch))
}
@@ -233,16 +343,55 @@ func (d *Driver) TapSelector(ctx context.Context, selector string) error {
// primitive, so the gesture is two taps with this gap.
const doubleTapGap = 50 * time.Millisecond
// DoubleTap resolves the point once and dispatches both taps there: resolving
// per tap would scroll the second one away from the element the first hit.
func (d *Driver) DoubleTap(ctx context.Context, x, y int) error {
return webDoubleTap(ctx, func() error { return d.Tap(ctx, x, y) })
runCtx, cancel := d.runCtx(ctx)
defer cancel()
atX, atY, err := pointInView(runCtx, x, y)
if err != nil {
return err
}
return webDoubleTap(ctx, func(clickCount int) error {
return chromedp.Run(
runCtx,
chromedp.MouseClickXY(
float64(atX),
float64(atY),
chromedp.ClickCount(clickCount),
),
)
})
}
func (d *Driver) DoubleTapSelector(ctx context.Context, selector string) error {
return webDoubleTap(ctx, func() error { return d.TapSelector(ctx, selector) })
runCtx, cancel := d.runCtx(ctx)
defer cancel()
target, isXPath, err := TranslateStringSelector(selector)
if err != nil {
target = selector
}
if err := requireSelectorMatch(runCtx, target, selector); err != nil {
return err
}
options := []chromedp.QueryOption{chromedp.NodeVisible}
if isXPath {
options = append(options, chromedp.BySearch)
}
return webDoubleTap(ctx, func(clickCount int) error {
if clickCount < 2 {
return chromedp.Run(runCtx, chromedp.Click(target, options...))
}
return chromedp.Run(runCtx, chromedp.DoubleClick(target, options...))
})
}
func webDoubleTap(ctx context.Context, tap func() error) error {
if err := tap(); err != nil {
// webDoubleTap dispatches the pair a browser reads as one double click. Blink
// raises dblclick off the click count the second event carries, so two taps
// that both say "first click" arrive at a dblclick handler as two ordinary
// clicks and the gesture never happens at all.
func webDoubleTap(ctx context.Context, tap func(clickCount int) error) error {
if err := tap(1); err != nil {
return err
}
timer := time.NewTimer(doubleTapGap)
@@ -252,7 +401,7 @@ func webDoubleTap(ctx context.Context, tap func() error) error {
return ctx.Err()
case <-timer.C:
}
return tap()
return tap(2)
}
func (d *Driver) InputText(callerCtx context.Context, text string) error {
@@ -311,31 +460,69 @@ func (d *Driver) EraseText(callerCtx context.Context, _ int) error {
)
}
// Swipe drags a finger across the page as a trusted touch stream. Events
// synthesized in the page carry isTrusted false: they reach a handler that
// happens to listen, but never enter the input pipeline that scrolls, honours
// touch-action or resolves a gesture.
func (d *Driver) Swipe(ctx context.Context, fromX, fromY, toX, toY int, duration time.Duration) error {
runCtx, cancel := d.runCtx(ctx)
defer cancel()
millis := max(duration.Milliseconds(), 50)
script := fmt.Sprintf(`
(function() {
const el = document.elementFromPoint(%d, %d);
if (!el) return;
const steps = Math.max(1, Math.floor(%d / 16));
const dx = (%d - %d) / steps;
const dy = (%d - %d) / steps;
el.dispatchEvent(new PointerEvent('pointerdown', {clientX: %d, clientY: %d, bubbles: true}));
for (let i = 1; i <= steps; i++) {
el.dispatchEvent(new PointerEvent('pointermove', {clientX: %d + dx*i, clientY: %d + dy*i, bubbles: true}));
}
el.dispatchEvent(new PointerEvent('pointerup', {clientX: %d, clientY: %d, bubbles: true}));
})();`,
fromX, fromY,
millis,
toX, fromX, toY, fromY,
fromX, fromY,
fromX, fromY,
toX, toY,
atX, atY, err := pointInView(runCtx, fromX, fromY)
if err != nil {
return err
}
toX, toY = toX-(fromX-atX), toY-(fromY-atY)
steps := max(int(duration.Milliseconds())/16, 1)
actions := []chromedp.Action{touchAt(input.TouchStart, atX, atY)}
for i := 1; i <= steps; i++ {
actions = append(actions, touchAt(input.TouchMove,
atX+(toX-atX)*i/steps, atY+(toY-atY)*i/steps))
}
actions = append(
actions,
input.DispatchTouchEvent(input.TouchEnd, []*input.TouchPoint{}),
)
return chromedp.Run(runCtx, chromedp.Evaluate(script, nil))
return chromedp.Run(runCtx, actions...)
}
func touchAt(kind input.TouchType, x, y int) *input.DispatchTouchEventParams {
return input.DispatchTouchEvent(
kind,
[]*input.TouchPoint{{X: float64(x), Y: float64(y)}},
)
}
// Scroll moves the content under the point with a trusted wheel, which is how a
// browser scrolls.
//
// A finger drag scrolls too, but it ends in a fling whose distance follows the
// release velocity: five identical 240 px drags moved the page 354 to 616 px,
// so two runs of one seed would explore different screens. A wheel delta lands
// exactly, and chains from the element under the point out to its scrollable
// ancestors, which is what scrolling a named container means. The drag stays as
// Swipe, the verb for the gestures only a finger reaches.
func (d *Driver) Scroll(
ctx context.Context,
fromX, fromY, toX, toY int,
_ time.Duration,
) error {
runCtx, cancel := d.runCtx(ctx)
defer cancel()
atX, atY, err := pointInView(runCtx, fromX, fromY)
if err != nil {
return err
}
wheel := input.DispatchMouseEvent(input.MouseWheel, float64(atX), float64(atY)).
WithDeltaX(float64(fromX - toX)).
WithDeltaY(float64(fromY - toY))
// The wheel is applied off the CDP round trip, so without this the next
// read races it: a step could observe the page before its own scroll, and
// the pending scroll then lands during the following one.
return chromedp.Run(runCtx, wheel, chromedp.Evaluate(
`new Promise(done => requestAnimationFrame(() => requestAnimationFrame(done)))`,
nil,
awaitPromise,
))
}
func (d *Driver) PressKey(ctx context.Context, key string) error {
@@ -351,6 +538,10 @@ func (d *Driver) PressKey(ctx context.Context, key string) error {
func (d *Driver) LongPress(ctx context.Context, x, y int) error {
runCtx, cancel := d.runCtx(ctx)
defer cancel()
x, y, err := pointInView(runCtx, x, y)
if err != nil {
return err
}
script := fmt.Sprintf(`
(function() {
const el = document.elementFromPoint(%d, %d);
@@ -412,6 +603,23 @@ func (d *Driver) Hierarchy(ctx context.Context) (string, error) {
if (tag === 'input') return !NON_TEXT_INPUT_TYPES.includes((el.type || '').toLowerCase());
return false;
}
// An editable field's own text is the transient typed value; its hint names
// its purpose, which is the rung visibleLabel (internal/verifier/llm.go) reads
// first for such an element. Without it a web field reached the model named by
// its CSS class, an identifier no user can read. Same ladder as fieldHint in
// pkg/spec/src/web-runtime.ts, so one field is named one way on both hosts.
function fieldHint(el) {
if (!isEditableElement(el)) return '';
const ariaLabel = el.getAttribute('aria-label');
if (ariaLabel) return ariaLabel;
for (const label of el.labels || []) {
const text = (label.textContent || '').trim();
if (text) return text;
}
const placeholder = el.getAttribute('placeholder');
if (placeholder) return placeholder;
return el.getAttribute('name') || '';
}
// Shadow roots are part of the page a user sees, so they are part of the page
// we enumerate. Compose for Web mounts its canvas AND its accessibility tree
// inside a shadow root on the mount element, so a light-DOM-only walk reports
@@ -434,15 +642,68 @@ func (d *Driver) Hierarchy(ctx context.Context) (string, error) {
', [onclick]'));
const editableSet = new Set(deepQuery(
'input, textarea, [contenteditable]').filter(isEditableElement));
// Descended once for the whole dump, for the reason selectAllScript above
// descends: document.activeElement names the shadow host, so a Compose for
// Web app reported focus on its mount element and never on the field.
let focusedElement = document.activeElement;
while (focusedElement && focusedElement.shadowRoot && focusedElement.shadowRoot.activeElement) {
focusedElement = focusedElement.shadowRoot.activeElement;
}
// Descending is still not enough on Compose for Web: it takes keystrokes on a
// 1px transparent input pinned to the caret, and that input is a SIBLING of
// the accessibility tree rather than a node in it. DOM focus therefore never
// reaches the semantics element carrying the test tag, so confirmFocus in
// internal/runner/runner.go saw an unnamed element hold focus after every
// focus tap and refused to type. Compose declares the caret's box in these
// custom properties, which the input inherits from the container that
// positions it, so the field being typed into is the innermost editable box
// that caret sits in.
const CARET_ORIGIN_PROPERTY = '--compose-internal-web-backing-input-left';
function fieldBehindTheCaret(caretInput) {
if (!caretInput || caretInput.tagName !== 'INPUT') return null;
if (!getComputedStyle(caretInput).getPropertyValue(CARET_ORIGIN_PROPERTY).trim()) return null;
const caret = caretInput.getBoundingClientRect();
const x = (caret.left + caret.right) / 2;
const y = (caret.top + caret.bottom) / 2;
let field = null;
let fieldArea = Infinity;
for (const candidate of editableSet) {
if (candidate === caretInput) continue;
const box = candidate.getBoundingClientRect();
const area = box.width * box.height;
if (area <= 0 || area >= fieldArea) continue;
if (x < box.left || x > box.right || y < box.top || y > box.bottom) continue;
field = candidate;
fieldArea = area;
}
return field;
}
focusedElement = fieldBehindTheCaret(focusedElement) || focusedElement;
function buildTree(el, isRoot) {
const rect = el.getBoundingClientRect();
// Every attribute the markup wrote, keyed as written, which is what attrs
// means on the native hosts and what rawAttributes in
// pkg/spec/src/web-runtime.ts already gives the page-side handle. Emitting
// only the standard set left a spec's data-* reads (folio-web's data-cents,
// data-account-id, data-balance) undefined on the goja host and absent from
// the trace, so an offline replay of the same step could not see them at
// all. The derived keys below overwrite anything of the same name.
const attrs = {};
for (const attribute of el.attributes || []) {
attrs[attribute.name] = attribute.value;
}
const bounds = '[' + Math.round(rect.left) + ',' + Math.round(rect.top) + ',' +
Math.round(rect.right) + ',' + Math.round(rect.bottom) + ']';
if (rect.width > 0 || rect.height > 0) attrs.bounds = bounds;
const text = (el.textContent || '').trim().slice(0, 200);
if (text) attrs.text = text;
if (el.id) attrs['resource-id'] = el.id;
// The V8 host names a target by data-testid (IDENTITY_KEYS in
// pkg/spec/src/web-runtime.ts) and TapSelector translates the selector into
// a CSS attribute match, so a dump without this attribute leaves the goja
// host unable to resolve a target the other two resolve fine.
const testid = el.getAttribute('data-testid');
if (testid) attrs['data-testid'] = testid;
const label = el.getAttribute('aria-label') || el.getAttribute('alt') || el.getAttribute('title') || '';
if (label) attrs['content-desc'] = label;
const tag = (el.tagName || '').toLowerCase();
@@ -450,6 +711,8 @@ func (d *Driver) Hierarchy(ctx context.Context) (string, error) {
if (el.className && typeof el.className === 'string' && el.className.trim()) {
attrs['class'] = el.className.trim();
}
const hint = fieldHint(el);
if (hint) attrs['hintText'] = hint;
// The goja host reads scrollable off this attribute (internal/verifier
// worker.go targets). Without it every web element looks unscrollable there,
// so the goja-side enumeration offers no scroll while the V8 picker, which
@@ -476,11 +739,22 @@ func (d *Driver) Hierarchy(ctx context.Context) (string, error) {
return {
attributes: attrs,
children: children,
clickable: isClickable || null,
enabled: isEnabled(el) || null,
focused: document.activeElement === el || null,
checked: el.checked || null,
selected: el.selected || null,
// Emitted as plain booleans, never null: internal/hierarchy writes the
// attribute a selector matches on only where the producer stated the
// flag, so a state that arrives as null is one no selector can ask about.
// {clickable: false} and {enabled: false} matched nothing at all here
// while matching on android, which states every flag both ways.
clickable: isClickable,
enabled: isEnabled(el),
focused: focusedElement === el,
// A component keeps what it likes in these two properties, so what is
// emitted is the flag the field declares and not the property's value.
checked: el.checked === true,
selected: el.selected === true,
// Emitted as a plain boolean, never null, on every editable field: a
// consumer deciding what a typed value may be recorded as has to tell
// "not a secure entry" apart from "nobody said", and android says nothing.
secure: isEditable ? el.type === 'password' : null,
// Emitted as a plain boolean, never null: internal/hierarchy falls back to
// the native heuristic when the field is absent, which reads any class
// name containing "EditText" as an Android text widget. On web that is a
@@ -936,19 +1210,87 @@ new Promise((resolve, reject) => {
read();
})`
// Exceptions returns the uncaught errors and unhandled rejections the page
// runtime has buffered so far. The buffer is cumulative, which is what
// state.exceptions means inside the page (buildState in
// pkg/spec/src/web-runtime.ts), so the host and the page read one list.
func (d *Driver) Exceptions(ctx context.Context) ([]driver.Exception, error) {
const script = `JSON.stringify(window.__sanderlingExceptions__ ? window.__sanderlingExceptions__() : [])`
var encoded string
runCtx, cancel := d.runCtx(ctx)
defer cancel()
if err := chromedp.Run(runCtx, chromedp.Evaluate(script, &encoded)); err != nil {
return nil, fmt.Errorf("evaluate exceptions: %w", err)
}
if encoded == "" || encoded == "[]" {
return nil, nil
}
var captured []struct {
Class string `json:"class"`
Message string `json:"message"`
StackTrace string `json:"stackTrace"`
UnixMillis int64 `json:"unixMillis"`
}
if err := json.Unmarshal([]byte(encoded), &captured); err != nil {
return nil, fmt.Errorf("decode exceptions %s: %w", encoded, err)
}
result := make([]driver.Exception, 0, len(captured))
for _, entry := range captured {
result = append(result, driver.Exception{
Class: entry.Class,
Message: entry.Message,
StackTrace: entry.StackTrace,
UnixMillis: entry.UnixMillis,
})
}
return result, nil
}
// nextActionScript puts the carried draw position back before the picker
// decides and reads the new one out afterwards, in the one evaluation, so no
// navigation can land between the restore and the draw.
const nextActionScript = `((carried) => {
if (!window.__sanderlingNextAction__) return "{}";
if (carried !== "" && window.__sanderlingRestorePickerState__) {
window.__sanderlingRestorePickerState__(carried);
}
const action = window.__sanderlingNextAction__();
const state = window.__sanderlingPickerState__ ? window.__sanderlingPickerState__() : "";
return JSON.stringify({action, state});
})(%s)`
// NextActionFromV8 invokes the bundle-installed action generator and returns
// the resulting Action JSON. Returns an empty json.RawMessage when the
// generator declines to act this tick.
//
// The picker's draw position rides along: it lives here rather than in the
// page, because a page that navigates gets a fresh runtime whose picker would
// otherwise start the seed's stream over at its first draw on every reload.
func (d *Driver) NextActionFromV8(ctx context.Context) (json.RawMessage, error) {
const script = `JSON.stringify(window.__sanderlingNextAction__ ? window.__sanderlingNextAction__() : null)`
script := fmt.Sprintf(nextActionScript, strconv.Quote(d.pickerState))
var encoded string
runCtx, cancel := d.runCtx(ctx)
defer cancel()
if err := chromedp.Run(runCtx, chromedp.Evaluate(script, &encoded)); err != nil {
return nil, fmt.Errorf("evaluate next action: %w", err)
}
if encoded == "" || encoded == "null" {
if encoded == "" {
return nil, nil
}
return json.RawMessage(encoded), nil
var decoded struct {
Action json.RawMessage `json:"action"`
State string `json:"state"`
}
if err := json.Unmarshal([]byte(encoded), &decoded); err != nil {
return nil, fmt.Errorf("decode next action %s: %w", encoded, err)
}
// An empty state means the page had no runtime to ask, so the position we
// already hold is still the run's position.
if decoded.State != "" {
d.pickerState = decoded.State
}
if len(decoded.Action) == 0 || string(decoded.Action) == "null" {
return nil, nil
}
return decoded.Action, nil
}
+557
View File
@@ -9,11 +9,14 @@ import (
"net"
"net/http"
"net/http/httptest"
"os"
"strings"
"testing"
"time"
"github.com/chromedp/chromedp"
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// TestLaunch_ClearStateWipesStorageForTheTargetOrigin covers the CLI's default
@@ -306,6 +309,91 @@ func TestHierarchy_ScrollableAttribute(t *testing.T) {
}
}
// TestHierarchy_HintTextNamesAnEditableField covers the attribute visibleLabel
// (internal/verifier/llm.go) reads FIRST for an editable element. Without it a
// web field reached the model named by its CSS class, an identifier no user can
// read, on exactly the channel the label-source experiment varies. The ladder is
// fieldHint's in pkg/spec/src/web-runtime.ts, rung for rung.
func TestHierarchy_HintTextNamesAnEditableField(t *testing.T) {
const html = `<body>` +
`<label id="amount-label" for="amount">Amount</label>` +
`<input id="amount" class="input amount-input" placeholder="0.00" name="amount-field">` +
`<input id="search" class="input search-input" aria-label="Search" placeholder="Type here" name="q">` +
`<label id="note-label" for="note"> </label>` +
`<input id="note" class="input note-input" placeholder="What's this for?" name="note-field">` +
`<input id="reference" class="input" name="reference-field">` +
`<input id="unnamed" class="input">` +
`<input id="agree" class="checkbox" type="checkbox" placeholder="ignored">` +
`<button id="go" class="button" placeholder="ignored">go</button>` +
`</body>`
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(html))
}))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
type node struct {
Attributes map[string]string `json:"attributes"`
Editable bool `json:"editable"`
Children []node `json:"children"`
}
var root node
if err := json.Unmarshal([]byte(dump), &root); err != nil {
t.Fatalf("unmarshal hierarchy: %v", err)
}
fieldByID := map[string]node{}
var walk func(n node)
walk = func(n node) {
if id := n.Attributes["resource-id"]; id != "" {
fieldByID[id] = n
}
for _, c := range n.Children {
walk(c)
}
}
walk(root)
for _, tc := range []struct {
id string
want string
editable bool
}{
{"search", "Search", true},
{"amount", "Amount", true},
{"note", "What's this for?", true},
{"reference", "reference-field", true},
{"unnamed", "", true},
{"agree", "", false},
{"go", "", false},
} {
field := fieldByID[tc.id]
if field.Attributes["hintText"] != tc.want {
t.Errorf("%q: hintText = %q, want %q", tc.id, field.Attributes["hintText"], tc.want)
}
// visibleLabel reaches the hint only for an element the dump calls
// editable, so a field named right and marked wrong is still named by
// its class downstream.
if field.Editable != tc.editable {
t.Errorf("%q: editable = %v, want %v", tc.id, field.Editable, tc.editable)
}
if tc.want != "" && field.Attributes["hintText"] == field.Attributes["class"] {
t.Errorf("%q: named by its CSS class %q", tc.id, field.Attributes["class"])
}
}
}
// TestRunCtx_CallerCancelPropagates confirms that cancelling the caller's
// context cancels the chromedp-bound context returned by runCtx. This is the
// channel by which step deadlines and Ctrl-C reach in-flight CDP calls.
@@ -1013,3 +1101,472 @@ func TestEvaluateExtractors_RejectsAnUnenvelopedReading(t *testing.T) {
t.Errorf("EvaluateExtractors failed with %q, want it to name the bundle mismatch", err)
}
}
// TestHierarchy_CarriesEveryMarkupAttribute covers the data the spec actually
// reads. folio-web's extractors read data-cents, data-account-id and
// data-balance off the elements they find; the dump used to emit a fixed
// standard set, so those values were absent from the goja host and from every
// stored trace, and a selector over them resolved nothing offline.
func TestHierarchy_CarriesEveryMarkupAttribute(t *testing.T) {
const html = `<body>` +
`<div id="total-balance" data-cents="125000">$1,250.00</div>` +
`<div id="card" data-testid="account-card" data-account-id="acct-7" data-balance="4200">Tim</div>` +
`</body>`
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, "data:text/html,"+html, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("parse hierarchy: %v", err)
}
total := tree.Find("id:total-balance")
if total == nil {
t.Fatal("total-balance not in the dump")
}
if got := total.Attributes["data-cents"]; got != "125000" {
t.Errorf(`attrs["data-cents"] = %q, want "125000"`, got)
}
card := tree.Find(`data-account-id:acct-7`)
if card == nil {
t.Fatal("no element resolves by a data attribute the markup carries")
}
if got := card.Attributes["data-balance"]; got != "4200" {
t.Errorf(`attrs["data-balance"] = %q, want "4200"`, got)
}
if got := card.Attributes["data-testid"]; got != "account-card" {
t.Errorf(`attrs["data-testid"] = %q, want "account-card"`, got)
}
}
// growingPage serves a page whose content starts shorter than one screen and
// grows past it when its first button is tapped, which is what a chat, feed or
// ledger does as a run drives it.
const growingPage = `<body style="margin:0">
<button id="grow" style="height:80px">grow</button>
<div id="status">idle</div>
<div id="rest"></div>
<script>
document.getElementById('grow').addEventListener('click', function() {
document.getElementById('rest').innerHTML =
'<div style="height:900px"></div>' +
'<button id="below" style="height:60px">below</button>';
document.getElementById('below').addEventListener('click', function() {
document.getElementById('status').textContent = 'below tapped';
});
});
</script></body>`
// TestTap_ActuatesAnElementBelowTheLaunchViewport pins the invariant the whole
// web path rests on: an element the driver reports as present and clickable can
// be acted on. The emulated viewport is sized once at launch, so content the
// app adds afterwards lies below it while getBoundingClientRect keeps reporting
// where it is; a click dispatched there is hit-tested to the document root and
// the element never sees it, with no error anywhere.
func TestTap_ActuatesAnElementBelowTheLaunchViewport(t *testing.T) {
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(growingPage))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
tapByID := func(id string) {
t.Helper()
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("Parse: %v", err)
}
element := tree.Find("id:" + id)
if element == nil {
t.Fatalf("%s is not in the dump", id)
}
if !element.Clickable {
t.Fatalf("%s is not reported clickable", id)
}
x, y := element.Bounds.Center()
if err := d.Tap(ctx, x, y); err != nil {
t.Fatalf("Tap %s at (%d,%d): %v", id, x, y, err)
}
}
tapByID("grow")
tapByID("below")
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("Parse: %v", err)
}
status := tree.Find("id:status")
if status == nil {
t.Fatal("status is not in the dump")
}
if status.Text != "below tapped" {
t.Errorf(
"status = %q, want %q: the tap reached no element",
status.Text,
"below tapped",
)
}
}
// TestTap_ReportsAGestureThatReachesNoElement covers the half of the same bug
// that no scrolling can fix: a page that cannot scroll leaves the point out of
// reach, and the caller has to hear about it rather than read a clean run.
func TestTap_ReportsAGestureThatReachesNoElement(t *testing.T) {
const page = `<body style="margin:0;overflow:hidden">
<div style="height:80px">top</div>
<div id="rest" style="height:900px;overflow:hidden"></div>
<style>html{overflow:hidden}</style></body>`
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(page))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
err := d.Tap(ctx, 100, 5000)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf(
"Tap far below an unscrollable page: err = %v, want ErrGestureUndelivered",
err,
)
}
}
// TestTapSelector_ReportsASelectorThatMatchesNothing covers the by-selector
// half of a step that reads as dispatched and did nothing: the node the
// selector names is not on the page, so the click has no target at all.
func TestTapSelector_ReportsASelectorThatMatchesNothing(t *testing.T) {
const page = `<body style="margin:0">
<button id="present">here</button>
<div id="status">none</div>
<script>
document.getElementById('present').addEventListener('click', function () {
document.getElementById('status').textContent = 'present tapped';
});
</script></body>`
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(page))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
missCtx, missCancel := context.WithTimeout(ctx, 10*time.Second)
defer missCancel()
if err := d.TapSelector(missCtx, "id:absent"); !errors.Is(err, driver.ErrSelectorMatchedNothing) {
t.Fatalf("TapSelector on an absent element: err = %v, want ErrSelectorMatchedNothing", err)
}
if err := d.DoubleTapSelector(missCtx, "id:absent"); !errors.Is(err, driver.ErrSelectorMatchedNothing) {
t.Fatalf("DoubleTapSelector on an absent element: err = %v, want ErrSelectorMatchedNothing", err)
}
if err := d.TapSelector(ctx, "id:present"); err != nil {
t.Fatalf("TapSelector on the element that is there: %v", err)
}
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("Parse: %v", err)
}
if status := tree.Find("id:status"); status == nil || status.Text != "present tapped" {
t.Fatalf("status = %+v, want the tap on the present element to have landed", status)
}
}
const doubleClickPage = `<body style="margin:0">
<div id="target" style="width:200px;height:60px">edit me</div>
<div id="status">none</div>
<script>
var box = document.getElementById('target');
var report = document.getElementById('status');
var clicks = 0;
box.addEventListener('click', function () { clicks++; });
box.addEventListener('dblclick', function () {
report.textContent = 'edited after ' + clicks + ' clicks';
});
</script></body>`
func doubleClickStatus(t *testing.T, d *Driver, ctx context.Context) string {
t.Helper()
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("Parse: %v", err)
}
status := tree.Find("id:status")
if status == nil {
t.Fatal("status is not in the dump")
}
return status.Text
}
// TestDoubleTap_ReachesADoubleClickHandler pins a gesture the web driver had no
// way to deliver. Blink raises dblclick off the click count the second event
// carries, so a pair that both said "first click" arrived as two ordinary
// clicks: every double-click affordance on the web (an editable list row, a
// canvas, a table cell) was unreachable, with no error on any layer.
func TestDoubleTap_ReachesADoubleClickHandler(t *testing.T) {
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(doubleClickPage))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
if err := d.DoubleTap(ctx, 100, 30); err != nil {
t.Fatalf("DoubleTap: %v", err)
}
if status := doubleClickStatus(t, d, ctx); status != "edited after 2 clicks" {
t.Errorf(
"status = %q, want %q: the pair never read as one double click",
status,
"edited after 2 clicks",
)
}
}
// TestDoubleTapSelector_ReachesADoubleClickHandler covers the same gesture on
// the path the runner takes when the action names its target rather than a
// point.
func TestDoubleTapSelector_ReachesADoubleClickHandler(t *testing.T) {
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(doubleClickPage))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
if err := d.DoubleTapSelector(ctx, "id:target"); err != nil {
t.Fatalf("DoubleTapSelector: %v", err)
}
if status := doubleClickStatus(t, d, ctx); status != "edited after 2 clicks" {
t.Errorf(
"status = %q, want %q: the pair never read as one double click",
status,
"edited after 2 clicks",
)
}
}
// gesturesServer serves the fixture both gesture tests measure against: a
// document taller than the emulated viewport, a scrollable container inside it,
// and a row that dismisses on a horizontal drag.
func gesturesServer(t *testing.T) *httptest.Server {
t.Helper()
body, err := os.ReadFile("testdata/gestures.html")
if err != nil {
t.Fatal(err)
}
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write(body)
}),
)
t.Cleanup(server.Close)
return server
}
// TestScroll_MovesThePageAndAScrollableContainer covers the verb the runner
// lowers every Scroll action onto. Script-dispatched pointer events are
// untrusted and a browser never scrolls on them, so the web Scroll used to
// leave scrollY and every scrollTop exactly where they were while reporting a
// step that ran. The repeat also pins the distance: a run that scrolls a
// different amount each time explores differently on the same seed.
func TestScroll_MovesThePageAndAScrollableContainer(t *testing.T) {
server := gesturesServer(t)
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
read := func(expression string) int {
t.Helper()
var value int
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(expression, &value)); err != nil {
t.Fatalf("evaluate %s: %v", expression, err)
}
return value
}
const pageScroll = `Math.round(window.scrollY)`
const containerScroll = `Math.round(document.getElementById("inner").scrollTop)`
if before := read(pageScroll); before != 0 {
t.Fatalf("scrollY before = %d, want 0", before)
}
var pageDistances []int
for range 3 {
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(`window.scrollTo(0, 0)`, nil)); err != nil {
t.Fatalf("reset: %v", err)
}
if err := d.Scroll(ctx, 195, 500, 195, 260, 300*time.Millisecond); err != nil {
t.Fatalf("Scroll: %v", err)
}
pageDistances = append(pageDistances, read(pageScroll))
}
if pageDistances[0] <= 0 {
t.Errorf(
"scrollY after = %d, want > 0: the page never scrolled",
pageDistances[0],
)
}
if pageDistances[0] != pageDistances[1] ||
pageDistances[1] != pageDistances[2] {
t.Errorf(
"scrollY over three identical scrolls = %v, want one distance",
pageDistances,
)
}
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(`window.scrollTo(0, 0)`, nil)); err != nil {
t.Fatalf("reset: %v", err)
}
containerY := read(
`Math.round(document.getElementById("inner").getBoundingClientRect().top + 100)`,
)
if before := read(containerScroll); before != 0 {
t.Fatalf("container scrollTop before = %d, want 0", before)
}
if err := d.Scroll(ctx, 195, containerY, 195, containerY-120, 300*time.Millisecond); err != nil {
t.Fatalf("Scroll in the container: %v", err)
}
if after := read(containerScroll); after <= 0 {
t.Errorf(
"container scrollTop after = %d, want > 0: the container never scrolled",
after,
)
}
if after := read(pageScroll); after != 0 {
t.Errorf(
"scrollY = %d, want 0: a scroll inside a container moved the page instead",
after,
)
}
}
// TestScroll_ReportsAGestureThatReachesNoElement keeps the scroll path on the
// same footing as the tap path: a page that cannot bring the point into the
// viewport has to say the gesture reached nothing rather than read as a step
// that scrolled.
func TestScroll_ReportsAGestureThatReachesNoElement(t *testing.T) {
const page = `<body style="margin:0;overflow:hidden">
<div style="height:80px">top</div>
<style>html{overflow:hidden}</style></body>`
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(page))
}),
)
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
err := d.Scroll(ctx, 100, 5000, 100, 4800, 300*time.Millisecond)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf(
"Scroll far below an unscrollable page: err = %v, want ErrGestureUndelivered",
err,
)
}
}
// TestSwipe_DeliversATrustedDragToARowHandler covers what the manual says
// sideways swipes are for. Script-dispatched pointer events carry isTrusted
// false, which is the mark of a gesture the browser never routed: nothing in
// the page's own input pipeline saw it, so scrolling, touch-action and any
// handler that filters on trust behave as if the finger never moved.
func TestSwipe_DeliversATrustedDragToARowHandler(t *testing.T) {
server := gesturesServer(t)
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
if err := d.Swipe(ctx, 300, 40, 100, 40, 300*time.Millisecond); err != nil {
t.Fatalf("Swipe: %v", err)
}
var status string
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(
`document.getElementById("status").textContent`, &status)); err != nil {
t.Fatalf("evaluate: %v", err)
}
if status != "dismissed left trusted" {
t.Errorf("row status = %q, want %q", status, "dismissed left trusted")
}
}
@@ -0,0 +1,582 @@
//go:build browser
package chrome
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"path/filepath"
"slices"
"testing"
"time"
"github.com/chromedp/chromedp"
"github.com/priyanshujain/sanderling/internal/bundler"
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// One page, one checkbox, two readers of its state.
//
// docs/manual/spec-language.md lists `checked` on every element find returns.
// The goja host reads it off the hierarchy dump this driver builds; the V8 host
// reads it off the live DOM through elementHandle in
// pkg/spec/src/web-runtime.ts. A field one host does not expose is silent: the
// property reading it compares undefined and holds on every screen.
//
// The state is read before and after a real click, because HTML keeps checkbox
// state in the DOM property and not in the markup attribute: an implementation
// reading element.getAttribute("checked") reports the starting value forever and
// passes any test that only reads a freshly loaded page.
func TestElementState_ChecksTrackTheLiveDOM(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/element-state.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
installStateProbe(ctx, t, d)
requireChecked(ctx, t, d, "toggle-all", false)
requireChecked(ctx, t, d, "toggle-done", true)
clickElement(ctx, t, d, "id:toggle-all")
clickElement(ctx, t, d, "id:toggle-done")
requireChecked(ctx, t, d, "toggle-all", true)
requireChecked(ctx, t, d, "toggle-done", false)
}
// requireChecked holds both hosts to one answer. The dump is re-read per call so
// the goja side is compared at the same page state as the handle.
func requireChecked(
ctx context.Context,
t *testing.T,
d *Driver,
id string,
want bool,
) {
t.Helper()
state := elementStateFromWebRuntime(ctx, t, d, id)
if state.Checked == nil {
t.Fatalf(
"the ax handle for %q exposes no `checked` field; "+
"docs/manual/spec-language.md lists it on every element find returns",
id,
)
}
if *state.Checked != want {
t.Errorf(
"the ax handle reports %q checked=%v, want %v (its markup attribute reads %q)",
id,
*state.Checked,
want,
state.AttrChecked,
)
}
if got := checkedInHierarchyDump(ctx, t, d, id); got != want {
t.Errorf(
"the hierarchy dump reports %q checked=%v, want %v",
id,
got,
want,
)
}
}
// The rest of the boolean state the manual lists, on the same page.
//
// `selected` is read after the selection is moved off the markup's option, for
// the same reason `checked` is: the attribute records only where the page
// started.
func TestElementState_ReportsTheOtherDocumentedBooleans(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/element-state.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
installStateProbe(ctx, t, d)
requireBoolean(
ctx,
t,
d,
"save",
"enabled",
func(s elementState) *bool { return s.Enabled },
true,
)
requireBoolean(
ctx,
t,
d,
"cancel",
"enabled",
func(s elementState) *bool { return s.Enabled },
false,
)
requireBoolean(
ctx,
t,
d,
"filter-active",
"selected",
func(s elementState) *bool { return s.Selected },
true,
)
requireBoolean(
ctx,
t,
d,
"filter-all",
"selected",
func(s elementState) *bool { return s.Selected },
false,
)
if err := chromedp.Run(
d.tabCtx,
chromedp.Evaluate(`document.getElementById('filter').selectedIndex = 0`, nil),
); err != nil {
t.Fatalf("move the selection: %v", err)
}
requireBoolean(
ctx,
t,
d,
"filter-active",
"selected",
func(s elementState) *bool { return s.Selected },
false,
)
requireBoolean(
ctx,
t,
d,
"filter-all",
"selected",
func(s elementState) *bool { return s.Selected },
true,
)
clickElement(ctx, t, d, "id:editing")
requireBoolean(
ctx,
t,
d,
"editing",
"focused",
func(s elementState) *bool { return s.Focused },
true,
)
requireBoolean(
ctx,
t,
d,
"save",
"focused",
func(s elementState) *bool { return s.Focused },
false,
)
}
// Every editable field states `secure`, false included. internal/verifier
// redacts a typed value whenever the target does not positively report "not a
// secure entry", so a dump that omitted the key on an ordinary text field would
// redact the whole recent-action memory on web.
func TestElementState_SecureStatesEveryEditableField(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/element-state.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
for _, testCase := range []struct {
selector string
want bool
}{
{"id:secret", true},
{"id:editing", false},
} {
element := elementInHierarchyDump(ctx, t, d, testCase.selector)
if !element.SecureReported() {
t.Errorf("the dump states no secure fact for %s", testCase.selector)
}
if element.Secure != testCase.want {
t.Errorf("%s secure = %v, want %v", testCase.selector, element.Secure, testCase.want)
}
}
if button := elementInHierarchyDump(ctx, t, d, "id:save"); button.SecureReported() {
t.Error("a button is not a text entry and states nothing")
}
}
// Focus belongs to the node the user is typing into, not to the element the
// shadow tree is mounted on.
//
// document.activeElement stops at a shadow boundary and names the host, so a
// Compose for Web app reports focus on its mount element forever. confirmFocus
// in internal/runner/runner.go re-reads the dump after a focus tap and refuses
// to type when the field it tapped is not the one holding focus, so every
// InputText step on such an app failed and the run aborted.
func TestElementState_FocusDescendsIntoTheShadowRoot(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/shadow-focus.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
field := elementInHierarchyDump(ctx, t, d, "id:shadow-field")
x, y := field.Bounds.Center()
if err := d.Tap(ctx, x, y); err != nil {
t.Fatalf("Tap the field: %v", err)
}
if tapped := elementInHierarchyDump(ctx, t, d, "id:shadow-field"); !tapped.Focused {
t.Error("the field inside the shadow root reports no focus after being tapped")
}
if host := elementInHierarchyDump(ctx, t, d, "id:app"); host.Focused {
t.Error("the shadow host reports focus, so the text would land there")
}
}
// Focus belongs to the field the caret sits in, not to the input the caret is.
//
// Compose for Web takes keystrokes on a 1px transparent input pinned to the
// caret, and that input is a sibling of the accessibility tree rather than a
// node in it. Descending activeElement through the shadow roots therefore lands
// on a node no selector can name, and every semantics element reads unfocused,
// so confirmFocus in internal/runner/runner.go rejected each focus tap with "an
// unnamed element holds focus" and no InputText step ever ran.
//
// Both fields are tapped, because reporting the first editable in the tree
// would satisfy the email half of this and still type into the wrong field.
func TestElementState_FocusFollowsTheCaretToItsField(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/compose-backing-input.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
for _, field := range []string{"EmailField", "PasswordField"} {
tapped := elementInHierarchyDump(ctx, t, d, "id:"+field)
x, y := tapped.Bounds.Center()
if err := d.Tap(ctx, x, y); err != nil {
t.Fatalf("Tap %s: %v", field, err)
}
if focused := elementInHierarchyDump(ctx, t, d, "id:"+field); !focused.Focused {
t.Errorf("%s reports no focus after being tapped", field)
}
if caret := elementInHierarchyDump(ctx, t, d, "id:caret-input"); caret.Focused {
t.Errorf("the hidden caret input reports focus after tapping %s, "+
"and no selector can name it", field)
}
if held := focusedElements(ctx, t, d); len(held) != 1 {
t.Errorf("after tapping %s the dump reports %d focused elements %v, want 1",
field, len(held), held)
}
}
}
func focusedElements(ctx context.Context, t *testing.T, d *Driver) []string {
t.Helper()
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("parse hierarchy: %v", err)
}
var held []string
for _, element := range tree.Elements {
if element.Focused {
held = append(held, element.ResourceID+"/"+element.Attributes["tag"])
}
}
return held
}
func elementInHierarchyDump(
ctx context.Context,
t *testing.T,
d *Driver,
selector string,
) *hierarchy.Element {
t.Helper()
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("parse hierarchy: %v", err)
}
element := tree.Find(selector)
if element == nil {
t.Fatalf("the hierarchy dump holds no element matching %q", selector)
}
return element
}
func requireBoolean(
ctx context.Context,
t *testing.T,
d *Driver,
id string,
field string,
read func(elementState) *bool,
want bool,
) {
t.Helper()
got := read(elementStateFromWebRuntime(ctx, t, d, id))
if got == nil {
t.Fatalf(
"the ax handle for %q exposes no `%s` field; "+
"docs/manual/spec-language.md lists it on every element find returns",
id, field,
)
}
if *got != want {
t.Errorf(
"the ax handle reports %q %s=%v, want %v",
id,
field,
*got,
want,
)
}
}
// elementState is the boolean state one element reports, as pointers: a field
// the handle does not expose at all decodes as absent rather than as false.
type elementState struct {
Checked *bool `json:"checked"`
Enabled *bool `json:"enabled"`
Focused *bool `json:"focused"`
Selected *bool `json:"selected"`
AttrChecked string `json:"attrChecked"`
}
func elementStateFromWebRuntime(
ctx context.Context,
t *testing.T,
d *Driver,
id string,
) elementState {
t.Helper()
var encoded string
script := `JSON.stringify(window.__sanderlingElementState__(` + jsArgument(
id,
) + `))`
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(script, &encoded)); err != nil {
t.Fatalf("read web runtime state for %q: %v", id, err)
}
if encoded == "null" {
t.Fatalf("the web runtime resolved no element for id %q", id)
}
var state elementState
if err := json.Unmarshal([]byte(encoded), &state); err != nil {
t.Fatalf("decode web runtime state: %v", err)
}
return state
}
func checkedInHierarchyDump(
ctx context.Context,
t *testing.T,
d *Driver,
id string,
) bool {
t.Helper()
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("parse hierarchy: %v", err)
}
element := tree.Find("id:" + id)
if element == nil {
t.Fatalf("the hierarchy dump holds no element with id %q", id)
}
return element.Checked
}
func clickElement(
ctx context.Context,
t *testing.T,
d *Driver,
selector string,
) {
t.Helper()
if err := d.TapSelector(ctx, selector); err != nil {
t.Fatalf("TapSelector(%q): %v", selector, err)
}
}
func installStateProbe(ctx context.Context, t *testing.T, d *Driver) {
t.Helper()
specSource := filepath.Join(repoRootDir(t), "pkg", "spec")
probe, err := bundler.BundleWeb(bundler.WebOptions{
EntryFile: filepath.Join(specSource, "test", "dom-state-probe.ts"),
WebRuntimeFile: filepath.Join(specSource, "src", "web-runtime.ts"),
})
if err != nil {
t.Fatalf("bundle dom state probe: %v", err)
}
if err := d.InstallBundle(ctx, probe.JavaScript); err != nil {
t.Fatalf("install dom state probe: %v", err)
}
}
// Every key the manual offers, measured at the page.
//
// A key name the driver does not map presses nothing, and a spec clause written
// over it ("escape discards the edit in progress") can never fail: the run stays
// green having actuated nothing. The page records its own keydown events, so
// what is asserted here is what the DOM received, not what the driver sent.
func TestPressKey_ArrivesAtThePage(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/element-state.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
for _, keyCase := range []struct{ key, want string }{
{"enter", "Enter"},
{"tab", "Tab"},
{"escape", "Escape"},
{"up", "ArrowUp"},
{"down", "ArrowDown"},
{"left", "ArrowLeft"},
{"right", "ArrowRight"},
} {
t.Run(keyCase.key, func(t *testing.T) {
forgetKeys(ctx, t, d)
if err := d.PressKey(ctx, keyCase.key); err != nil {
t.Fatalf("PressKey(%q): %v", keyCase.key, err)
}
got := keysSeenByThePage(ctx, t, d)
if !slices.Equal(got, []string{keyCase.want}) {
t.Errorf("PressKey(%q) reached the page as %v, want [%s]",
keyCase.key, got, keyCase.want)
}
})
}
// back and home have no browser meaning, and reporting that is the whole
// point: a key that quietly presses nothing is indistinguishable from a
// requirement that holds.
for _, key := range []string{"back", "home"} {
t.Run(key+" is reported unsupported", func(t *testing.T) {
forgetKeys(ctx, t, d)
if err := d.PressKey(ctx, key); err == nil {
t.Errorf("PressKey(%q) reported no error on web", key)
}
if got := keysSeenByThePage(ctx, t, d); len(got) != 0 {
t.Errorf("PressKey(%q) reached the page as %v", key, got)
}
})
}
}
func forgetKeys(ctx context.Context, t *testing.T, d *Driver) {
t.Helper()
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(`window.__keys__ = []`, nil)); err != nil {
t.Fatalf("reset recorded keys: %v", err)
}
}
func keysSeenByThePage(ctx context.Context, t *testing.T, d *Driver) []string {
t.Helper()
var encoded string
script := `JSON.stringify(window.__keys__)`
if err := chromedp.Run(d.tabCtx, chromedp.Evaluate(script, &encoded)); err != nil {
t.Fatalf("read recorded keys: %v", err)
}
var keys []string
if err := json.Unmarshal([]byte(encoded), &keys); err != nil {
t.Fatalf("decode recorded keys: %v", err)
}
return keys
}
// The dump declares clickable, enabled, checked, selected and editable as
// flags, and a component keeps whatever it likes in the properties two of them
// are read from: a selector element names the selected item, not a boolean.
// Emitting the property raw cost the whole observation, because a dump is
// decoded as one document and one string in it fails all of it.
// pkg/spec/src/web-runtime.ts already answers `state.selected === true`, so the
// two hosts also disagreed about the same fact on the same page.
func TestElementState_AComponentPropertyDoesNotBlankTheTree(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/custom-element-flags.html", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
dump, err := d.Hierarchy(ctx)
if err != nil {
t.Fatalf("Hierarchy: %v", err)
}
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatalf("Parse: %v", err)
}
if tree.UnreadableFlags != 0 {
t.Errorf("UnreadableFlags = %d, want 0: the dump must send booleans for the fields it declares as flags", tree.UnreadableFlags)
}
tabs := tree.Find("id:tabs")
if tabs == nil {
t.Fatalf("the element holding the string property is missing from a tree of %d elements", len(tree.Elements))
}
if tabs.Selected {
t.Error("selected must be false: the property holds an item name, not a flag")
}
if picker := tree.Find("id:picker"); picker == nil || picker.Checked {
t.Errorf("checked must be false for a property holding a string, got %+v", picker)
}
if toggle := tree.Find("id:toggle"); toggle == nil || !toggle.Checked {
t.Errorf("a real checkbox must still read checked, got %+v", toggle)
}
}
+166 -20
View File
@@ -10,6 +10,7 @@ import (
"path/filepath"
"runtime"
"slices"
"strconv"
"testing"
"time"
@@ -45,12 +46,33 @@ import (
// elementFacts is one element as a producer reports it: the tag, for readable
// failures, and every fact acceptsTarget consults.
type elementFacts struct {
tag string
clickable bool
enabled bool
editable bool
scrollable bool
positiveBounds bool
tag string
clickable bool
enabled bool
editable bool
scrollable bool
hintText string
// secure is three-valued: "true", "false", or "" where the producer states
// nothing, which android states for every element. It decides whether a
// typed value may be written into the shared record, so the two producers
// disagreeing about it writes a credential into a run's trace.
secure string
// checked, selected and focused reach a spec through the ax handle alone:
// they decide nothing about which action is offered, so the target
// enumeration does not carry them, and a selector naming one of them
// resolves against this same reading.
checked bool
selected bool
focused bool
// handleClickable and handleEditable are the same facts on the ax element a
// spec reaches through state.ax.find, a third place they are computed and the
// one that has twice been the odd one out: clickable answered a hardcoded
// true while the other two resolved a selector, and editable read the
// INHERITED isContentEditable, which makes every span inside a contenteditable
// container typeable.
handleClickable bool
handleEditable bool
positiveBounds bool
}
// factRow pairs an element's facts with the id both producers key on, kept in
@@ -91,6 +113,8 @@ func TestHierarchy_DerivesTheSameFactsAsTheWebRuntime(t *testing.T) {
requireEveryElementNamed(t, "the hierarchy dump", fromDump)
requireEveryElementNamed(t, "the web runtime", fromWebRuntime)
requireBothPolarities(t, fromWebRuntime)
requireEverySecureState(t, fromWebRuntime)
requireTheHandleAgreesWithTheEnumeration(t, fromWebRuntime)
compareEnumeratedElements(t, fromDump, fromWebRuntime)
compareDerivedFacts(t, fromDump, fromWebRuntime)
})
@@ -116,6 +140,11 @@ func factsFromHierarchyDump(t *testing.T, dump string) []factRow {
enabled: element.Enabled,
editable: element.Editable,
scrollable: element.Attributes["scrollable"] == "true",
hintText: element.Attributes["hintText"],
secure: secureFromDump(element),
checked: element.Checked,
selected: element.Selected,
focused: element.Focused,
positiveBounds: hasPositiveBounds(
element.Bounds.Width(),
element.Bounds.Height(),
@@ -155,14 +184,21 @@ func factsFromWebRuntime(
t.Fatalf("read web runtime facts: %v", err)
}
var wire []struct {
ID string `json:"id"`
Tag string `json:"tag"`
Clickable bool `json:"clickable"`
Enabled bool `json:"enabled"`
Editable bool `json:"editable"`
Scrollable bool `json:"scrollable"`
Width int `json:"width"`
Height int `json:"height"`
ID string `json:"id"`
Tag string `json:"tag"`
Clickable bool `json:"clickable"`
Enabled bool `json:"enabled"`
Editable bool `json:"editable"`
Scrollable bool `json:"scrollable"`
HintText string `json:"hintText"`
HandleClickable bool `json:"handleClickable"`
HandleEditable bool `json:"handleEditable"`
Secure *bool `json:"secure"`
Checked bool `json:"checked"`
Selected bool `json:"selected"`
Focused bool `json:"focused"`
Width int `json:"width"`
Height int `json:"height"`
}
if err := json.Unmarshal([]byte(encoded), &wire); err != nil {
t.Fatalf("decode web runtime facts: %v", err)
@@ -172,18 +208,43 @@ func factsFromWebRuntime(
rows = append(rows, factRow{
id: item.ID,
facts: elementFacts{
tag: item.Tag,
clickable: item.Clickable,
enabled: item.Enabled,
editable: item.Editable,
scrollable: item.Scrollable,
positiveBounds: hasPositiveBounds(item.Width, item.Height),
tag: item.Tag,
clickable: item.Clickable,
enabled: item.Enabled,
editable: item.Editable,
scrollable: item.Scrollable,
hintText: item.HintText,
handleClickable: item.HandleClickable,
handleEditable: item.HandleEditable,
secure: secureFromWebRuntime(item.Secure),
checked: item.Checked,
selected: item.Selected,
focused: item.Focused,
positiveBounds: hasPositiveBounds(item.Width, item.Height),
},
})
}
return rows
}
// secureFromDump and secureFromWebRuntime read the same three-valued fact off
// the two producers. The dump leaves the field out entirely for an element it
// states nothing about, the web runtime reports null for it, and both mean the
// same thing: unknown, not "not a secure entry".
func secureFromDump(element *hierarchy.Element) string {
if !element.SecureReported() {
return ""
}
return strconv.FormatBool(element.Secure)
}
func secureFromWebRuntime(secure *bool) string {
if secure == nil {
return ""
}
return strconv.FormatBool(*secure)
}
// hasPositiveBounds is the positiveBounds fact of pkg/spec/src/targets.ts,
// applied to both producers so the geometry comparison cannot drift from the
// rule it stands in for.
@@ -222,6 +283,10 @@ func requireBothPolarities(t *testing.T, rows []factRow) {
{"enabled", func(f elementFacts) bool { return f.enabled }},
{"editable", func(f elementFacts) bool { return f.editable }},
{"scrollable", func(f elementFacts) bool { return f.scrollable }},
{"hintText", func(f elementFacts) bool { return f.hintText != "" }},
{"checked", func(f elementFacts) bool { return f.checked }},
{"selected", func(f elementFacts) bool { return f.selected }},
{"focused", func(f elementFacts) bool { return f.focused }},
{"positiveBounds", func(f elementFacts) bool { return f.positiveBounds }},
} {
var sawTrue, sawFalse bool
@@ -244,6 +309,63 @@ func requireBothPolarities(t *testing.T, rows []factRow) {
}
}
// requireEverySecureState is requireBothPolarities for the one fact that is not
// a boolean. A fixture holding no password entry would compare "false" against
// "false" over the whole page and pass while proving nothing about the state
// that decides whether a typed value may be written down.
func requireEverySecureState(t *testing.T, rows []factRow) {
t.Helper()
seen := map[string]bool{}
for _, row := range rows {
seen[row.facts.secure] = true
}
for _, state := range []string{"true", "false", ""} {
if !seen[state] {
t.Errorf(
"the fixture no longer holds an element the web runtime reports "+
"secure=%q for, so comparing that state proves nothing",
state,
)
}
}
}
// requireTheHandleAgreesWithTheEnumeration compares the V8 host against itself.
// An element a spec reaches through state.ax and the same element in the
// enumeration must be clickable, and typeable, to the same degree, or a spec
// taps a container the picker calls inert and types into a box the picker calls
// read-only. The handle resolves each fact by element.matches over a selector
// while the enumeration resolves it by membership of the set that selector
// queried, and this is where the two answers are held together over a real page:
// an [onclick] attribute, an onclick property that is not one, a <span> whose
// contenteditable is inherited from its container, elements inside a shadow root.
func requireTheHandleAgreesWithTheEnumeration(t *testing.T, rows []factRow) {
t.Helper()
for _, row := range rows {
for _, fact := range []struct {
name string
handle bool
enumeration bool
}{
{"clickable", row.facts.handleClickable, row.facts.clickable},
{"editable", row.facts.handleEditable, row.facts.editable},
} {
if fact.handle != fact.enumeration {
t.Errorf(
"%q (<%s>): the ax handle reports %s=%v, the enumeration reports "+
"%s=%v",
row.id,
row.facts.tag,
fact.name,
fact.handle,
fact.name,
fact.enumeration,
)
}
}
}
}
// compareEnumeratedElements is the check that the two producers walk the same
// document. It is what notices a producer that roots at body and never sees
// `html`, or one that enumerates the head subtree the other drops.
@@ -307,6 +429,27 @@ func compareDerivedFacts(t *testing.T, fromDump, fromWebRuntime []factRow) {
web.tag,
)
}
if dump.hintText != web.hintText {
t.Errorf(
"%q (<%s>): the hierarchy dump names the field %q, the web runtime "+
"names it %q; the model is shown a different control on each host",
row.id,
dump.tag,
dump.hintText,
web.hintText,
)
}
if dump.secure != web.secure {
t.Errorf(
"%q (<%s>): the hierarchy dump derives secure=%q, the web runtime "+
"derives secure=%q; one host would write into the record a value "+
"the other redacts",
row.id,
dump.tag,
dump.secure,
web.secure,
)
}
for _, fact := range []struct {
name string
dump bool
@@ -316,6 +459,9 @@ func compareDerivedFacts(t *testing.T, fromDump, fromWebRuntime []factRow) {
{"enabled", dump.enabled, web.enabled},
{"editable", dump.editable, web.editable},
{"scrollable", dump.scrollable, web.scrollable},
{"checked", dump.checked, web.checked},
{"selected", dump.selected, web.selected},
{"focused", dump.focused, web.focused},
{"positiveBounds", dump.positiveBounds, web.positiveBounds},
} {
if fact.dump != fact.web {
@@ -0,0 +1,154 @@
//go:build browser
package chrome
import (
"context"
"net/http"
"net/http/httptest"
"path/filepath"
"strconv"
"testing"
"time"
"github.com/chromedp/chromedp"
"github.com/priyanshujain/sanderling/internal/bundler"
)
const streamSeed = 1
// A seed describes one stream of actions, and a web page can end it.
//
// The picker's draw position lived in the page, and the bundle is registered to
// run at every freshly-navigated document, so any navigation built a new picker
// at the seed's first draw. A page that submits a form, follows a link or
// reloads therefore replayed draw one forever: on the angular-dart TodoMVC
// implementation, whose form GET-submits on Enter, seed 1 chose PressKey enter
// on 200 of 200 steps and created no todo at all. The run reported clean, so
// nothing but this distinguishes it from a seed that chose badly.
func TestNextActionFromV8_ReloadDoesNotRestartTheSeedStream(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
url := server.URL + "/picker-stream.html"
const calls = 8
uninterrupted := pickerStream(t, url, calls, calls)
if distinctActions(uninterrupted) < 2 {
t.Fatalf("the uninterrupted stream never varies (%v); a restart would be invisible", uninterrupted)
}
reloaded := pickerStream(t, url, calls, calls/2)
for index := range uninterrupted {
if reloaded[index] != uninterrupted[index] {
t.Fatalf("action %d after a reload is %s, uninterrupted the seed chose %s\nreloaded: %v\nuninterrupted: %v",
index, reloaded[index], uninterrupted[index], reloaded, uninterrupted)
}
}
}
// A navigation the trace cannot see is a run nobody can read: an analysis has
// no way to tell an app that reloaded from a generator that repeated itself.
func TestNavigations_ReportTheDocumentThatReplacedThePage(t *testing.T) {
server := httptest.NewServer(http.FileServer(http.Dir("testdata")))
defer server.Close()
url := server.URL + "/picker-stream.html"
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
if err := d.Launch(ctx, url, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
opening, err := d.Navigations(ctx)
if err != nil {
t.Fatalf("Navigations: %v", err)
}
if len(opening) != 0 {
t.Fatalf("the run's own opening navigation was reported as the app navigating: %v", opening)
}
if err := chromedp.Run(d.tabCtx, chromedp.Reload()); err != nil {
t.Fatalf("reload: %v", err)
}
reported, err := d.Navigations(ctx)
if err != nil {
t.Fatalf("Navigations: %v", err)
}
if len(reported) != 1 {
t.Fatalf("a reload reported %d navigations, want 1: %v", len(reported), reported)
}
if reported[0].URL != url {
t.Errorf("navigation URL = %q, want %q", reported[0].URL, url)
}
if reported[0].UnixMillis == 0 {
t.Error("navigation carries no timestamp")
}
drained, err := d.Navigations(ctx)
if err != nil {
t.Fatalf("Navigations: %v", err)
}
if len(drained) != 0 {
t.Errorf("navigations were reported twice: %v", drained)
}
}
// pickerStream drives the real seeded picker over the page and returns the
// action it chose on each call, reloading the page once after reloadAfter
// calls. Reloading past the call count leaves the stream uninterrupted.
func pickerStream(t *testing.T, url string, calls, reloadAfter int) []string {
t.Helper()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
if err := d.Launch(ctx, url, false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
installPickerStreamProbe(ctx, t, d)
chosen := make([]string, 0, calls)
for call := 0; call < calls; call++ {
if call == reloadAfter {
if err := chromedp.Run(d.tabCtx, chromedp.Reload()); err != nil {
t.Fatalf("reload before call %d: %v", call, err)
}
}
action, err := d.NextActionFromV8(ctx)
if err != nil {
t.Fatalf("NextActionFromV8 call %d: %v", call, err)
}
if len(action) == 0 {
t.Fatalf("call %d chose nothing; the page has tappable targets", call)
}
chosen = append(chosen, string(action))
}
return chosen
}
func installPickerStreamProbe(ctx context.Context, t *testing.T, d *Driver) {
t.Helper()
specSource := filepath.Join(repoRootDir(t), "pkg", "spec")
probe, err := bundler.BundleWeb(bundler.WebOptions{
EntryFile: filepath.Join(specSource, "test", "picker-stream-probe.ts"),
WebRuntimeFile: filepath.Join(specSource, "src", "web-runtime.ts"),
Defines: map[string]string{"SANDERLING_SEED": strconv.Itoa(streamSeed)},
})
if err != nil {
t.Fatalf("bundle picker stream probe: %v", err)
}
if err := d.InstallBundle(ctx, probe.JavaScript); err != nil {
t.Fatalf("install picker stream probe: %v", err)
}
}
func distinctActions(actions []string) int {
seen := map[string]struct{}{}
for _, action := range actions {
seen[action] = struct{}{}
}
return len(seen)
}
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,80 @@
<!doctype html>
<html id="page">
<head id="page-head">
<meta id="page-charset" charset="utf-8" />
<title id="page-title">compose backing input</title>
</head>
<body id="page-body">
<!-- Shaped like Compose for Web: keystrokes go through a 1px transparent
input pinned to the caret, a sibling of the accessibility tree rather
than a node in it, so DOM focus never lands on the semantics element
carrying the test tag. The custom properties are declared on the
container and read by the input, exactly as Compose writes them. -->
<div id="app"></div>
<script id="page-script">
const root = document.getElementById("app").attachShadow({ mode: "open" });
root.innerHTML = `
<div
id="caret-holder"
style="
position: absolute;
top: 0px;
left: 0px;
--compose-internal-web-backing-input-left: 0;
--compose-internal-web-backing-input-top: 0;
--compose-internal-web-backing-input-width: 1;
--compose-internal-web-backing-input-height: 17.578125;
"
>
<input
id="caret-input"
style="
position: absolute;
top: calc(var(--compose-internal-web-backing-input-top) * 1px);
left: calc(var(--compose-internal-web-backing-input-left) * 1px);
width: calc(var(--compose-internal-web-backing-input-width) * 1px);
height: calc(var(--compose-internal-web-backing-input-height) * 1px);
padding: 0;
color: transparent;
background: transparent;
caret-color: transparent;
border: none;
outline: none;
z-index: -1;
"
/>
</div>
<div id="a11y-root" role="presentation">
<div
id="EmailField"
role="textbox"
contenteditable="true"
style="position: absolute; left: 34px; top: 78px; width: 688px; height: 18px;"
></div>
<div
id="PasswordField"
role="textbox"
contenteditable="true"
style="position: absolute; left: 34px; top: 158px; width: 688px; height: 18px;"
></div>
</div>`;
const holder = root.getElementById("caret-holder");
const caret = root.getElementById("caret-input");
for (const field of root.querySelectorAll('[role="textbox"]')) {
field.addEventListener("mousedown", (event) => {
event.preventDefault();
const box = field.getBoundingClientRect();
holder.style.setProperty(
"--compose-internal-web-backing-input-left",
String(Math.round(event.clientX)),
);
holder.style.setProperty(
"--compose-internal-web-backing-input-top",
String(Math.round(box.top)),
);
caret.focus();
});
}
</script>
</body>
</html>
@@ -0,0 +1,22 @@
<!doctype html>
<html id="page">
<head id="page-head">
<meta id="page-charset" charset="utf-8" />
<title id="page-title">custom element flags</title>
</head>
<body id="page-body">
<div id="tabs">
<a id="tab-all" href="#/">All</a>
<a id="tab-active" href="#/active">Active</a>
</div>
<div id="picker"></div>
<input id="toggle" type="checkbox" checked />
<script id="page-script">
// What a component keeps in these properties is whatever the component
// means by them: a selector element names the selected item here rather
// than saying whether it is itself selected.
document.getElementById('tabs').selected = 'active';
document.getElementById('picker').checked = 'partial';
</script>
</body>
</html>
+28
View File
@@ -0,0 +1,28 @@
<!doctype html>
<html id="page">
<head id="page-head">
<meta id="page-charset" charset="utf-8" />
<title id="page-title">element state</title>
</head>
<body id="page-body">
<!-- toggle-all starts unchecked in the markup and toggle-done starts
checked, so a reader of the `checked` ATTRIBUTE reports each one's
starting value forever and disagrees with both after one click. -->
<input id="toggle-all" type="checkbox" />
<input id="toggle-done" type="checkbox" checked />
<input id="editing" type="text" value="buy milk" />
<input id="secret" type="password" />
<button id="save">save</button>
<button id="cancel" disabled>cancel</button>
<select id="filter">
<option id="filter-all">all</option>
<option id="filter-active" selected>active</option>
</select>
<script id="page-script">
window.__keys__ = [];
document.addEventListener('keydown', function (event) {
window.__keys__.push(event.key);
});
</script>
</body>
</html>
+12 -1
View File
@@ -25,10 +25,21 @@
<div id="shadow-overlay">
<button id="shadow-save">save</button>
<button id="shadow-cancel" disabled>cancel</button>
<input id="shadow-amount" type="text" value="10" />
<input id="shadow-amount" type="text" value="10" placeholder="0.00" />
<input id="shadow-password" type="password" placeholder="Password" />
<input id="shadow-agree" type="checkbox" />
<select id="shadow-choice">
<option id="shadow-choice-first">first</option>
</select>
<div id="shadow-plain">plain</div>
</div>
<div id="shadow-scroller"><div id="shadow-scroller-content"></div></div>`;
// document.activeElement names the HOST while focus is inside its shadow
// root, so a producer that does not descend reports focus on the mount
// element and never on the control, which is what a Compose for Web page
// looks like from the outside.
root.getElementById("shadow-agree").checked = true;
root.getElementById("shadow-save").focus();
</script>
</body>
</html>
+23 -2
View File
@@ -52,9 +52,22 @@
<option id="choice-first">first</option>
</select>
<input id="amount" type="text" value="10" />
<input id="agree" type="checkbox" />
<!-- One field per rung of the hint ladder both producers name an editable
field by. A rung read on one side only sends the model a different name
for the same field on the two hosts, which is the observation channel
the label-source arms vary. -->
<label id="hint-label-text" for="hint-label">Amount</label>
<input id="hint-label" type="text" placeholder="0.00" name="amount-field" />
<input id="hint-aria" type="text" aria-label="Search" placeholder="Type here" name="q" />
<input id="hint-placeholder" type="text" placeholder="What's this for?" name="note-field" />
<input id="hint-name" type="text" name="reference-field" />
<input id="password" type="password" placeholder="Password" />
<input id="agree" type="checkbox" placeholder="not a hint" />
<textarea id="notes"></textarea>
<div id="bio" contenteditable="true">bio</div>
<!-- contenteditable is INHERITED, so #bio-word answers isContentEditable
true while matching neither producer's editable selector itself. It is
the element the ax handle and the enumeration disagreed on. -->
<div id="bio" contenteditable="true">bio <span id="bio-word">word</span></div>
<div id="menu" role="button">menu</div>
<!-- One element per ARIA role both producers resolve as clickable. A role
covered on one side only makes that control reachable for one host,
@@ -84,6 +97,14 @@
<div id="filler"></div>
<script id="page-script">
document.getElementById("delegating-root").onclick = function () {};
// checked and focused are live state: the markup attribute records only
// what the page loaded with, and nothing writes focus down at all. Both
// producers read them off the element, so the fixture has to hold one of
// each or the comparison is false against false over the whole page.
// #choice-first is the third: an option a select made current with no
// attribute anywhere saying so.
document.getElementById("agree").checked = true;
document.getElementById("save").focus();
</script>
</body>
</html>
+52
View File
@@ -0,0 +1,52 @@
<!doctype html>
<html>
<head>
<style>
:root { --frame-w: 390; --frame-h: 640; }
body { margin: 0; font: 16px sans-serif; }
#row { height: 80px; background: #eef; touch-action: pan-y; }
#inner { height: 200px; overflow-y: scroll; border: 2px solid #333; }
#inner div { height: 60px; }
#filler div { height: 60px; border-bottom: 1px solid #ccc; }
</style>
</head>
<body>
<div id="row" data-testid="row">swipe me</div>
<div id="status" data-testid="status">idle</div>
<div id="inner" data-testid="inner"></div>
<div id="filler"></div>
<script>
const inner = document.getElementById('inner');
for (let i = 0; i < 30; i++) {
const item = document.createElement('div');
item.textContent = 'inner ' + i;
inner.appendChild(item);
}
const filler = document.getElementById('filler');
for (let i = 0; i < 40; i++) {
const item = document.createElement('div');
item.textContent = 'page ' + i;
filler.appendChild(item);
}
const row = document.getElementById('row');
const status = document.getElementById('status');
let startX = 0, startY = 0, tracking = false;
row.addEventListener('pointerdown', function(e) {
startX = e.clientX;
startY = e.clientY;
tracking = true;
row.setPointerCapture(e.pointerId);
});
row.addEventListener('pointermove', function(e) {
if (!tracking) return;
const dx = e.clientX - startX, dy = e.clientY - startY;
if (Math.abs(dx) <= 40 || Math.abs(dx) <= Math.abs(dy)) return;
tracking = false;
status.textContent =
(dx < 0 ? 'dismissed left' : 'dismissed right') +
(e.isTrusted ? ' trusted' : ' untrusted');
});
row.addEventListener('pointerup', function() { tracking = false; });
</script>
</body>
</html>
+14
View File
@@ -0,0 +1,14 @@
<!doctype html>
<meta charset="utf-8">
<title>picker stream</title>
<style>
button { display: block; width: 120px; height: 30px; margin: 10px; }
</style>
<form action="" method="get">
<input id="entry" placeholder="what needs to be done">
</form>
<button id="one">one</button>
<button id="two">two</button>
<button id="three">three</button>
<button id="four">four</button>
<button id="five">five</button>
+79
View File
@@ -0,0 +1,79 @@
<!doctype html>
<html id="page" lang="en">
<head id="page-head">
<meta id="page-charset" charset="utf-8" />
<title id="page-title">selector parity</title>
</head>
<body id="page-body">
<div id="summary_card" data-testid="summary" aria-label="summary_card, 3 customers">summary</div>
<div id="customer_list" aria-label="customer_list">
<div id="customer_row_a1" data-testid="customer-row" aria-label="customer_row_a1, Alice">Alice</div>
<div id="customer_row_b2" data-testid="customer-row" aria-label="customer_row_b2, Bob">Bob</div>
<div id="supplier_row_c3" aria-label="supplier_row_c3, Carol">Carol</div>
</div>
<div id="status_row"><span id="status_badge" data-state="sent-badge">Sent ✓</span></div>
<div id="draft_row"><span id="draft_badge">Sent ✓</span></div>
<div id="unsent_row"><span id="unsent_badge">3<!-- --> unsent</span></div>
<div id="split_row">Sen<span id="split_tail">t here</span></div>
<div id="nested_row" class="status">Sent <span id="nested_badge" class="status">Sent</span></div>
<!-- secure is derived from the field type, so the fields here have to cover
every state it answers with: the password entry, the three shapes of
editable field that are not one, and a control that is no field at all
and answers to neither value. -->
<form id="login_form">
<input id="login_email" type="email" aria-label="login_email" />
<input id="login_password" type="password" aria-label="login_password" />
<textarea id="login_note" aria-label="login_note"></textarea>
<div id="login_terms" contenteditable="true">terms</div>
<input id="login_remember" type="checkbox" aria-label="login_remember" />
</form>
<!-- hintText is the accessible-name ladder rather than one attribute, so
these fields differ one rung at a time: a bound label, a placeholder,
a placeholder an aria-label outranks, and the name the form gives the
field. -->
<div id="hint_row">
<label id="hint_phone_label" for="hint_phone">Phone number</label>
<input id="hint_phone" />
<input id="hint_search" placeholder="Search customers" />
<input id="hint_amount" aria-label="Amount in rupees" placeholder="0.00" />
<input id="hint_code" name="verification_code" />
</div>
<!-- The boolean states are derived from the live element, never written by
the markup, so these controls differ one state at a time: a disabled
button and an aria-disabled role control, a box ticked by script with
no checked attribute beside one cleared by script that has the
attribute, and a select whose first option is selected without one. -->
<div id="state_row">
<button id="state_save">save</button>
<button id="state_cancel" disabled>cancel</button>
<div id="state_submit" role="button" aria-disabled="true">submit</div>
<input id="state_remember" type="checkbox" aria-label="state_remember" />
<input id="state_agree" type="checkbox" checked aria-label="state_agree" />
<select id="state_month" aria-label="state_month">
<option id="state_january">January</option>
<option id="state_february">February</option>
</select>
</div>
<!-- scrollable is derived from the box rather than written by the markup,
and both producers state it only where it holds, so one container here
overflows its box and its neighbour does not. -->
<div id="scroll_row">
<div id="scroll_box" style="height: 20px; overflow: auto">
<div id="scroll_content" style="height: 200px">scrolls</div>
</div>
<div id="scroll_still">still</div>
</div>
<script id="state_script">
document.getElementById("state_remember").checked = true;
document.getElementById("state_agree").checked = false;
document.getElementById("state_save").focus();
</script>
<todo-app id="todo_app">
<todo-list id="todo_list">
<li id="todo_1">Buy milk</li>
<li id="todo_2">Pay rent</li>
</todo-list>
<a id="todo_link" href="#/all">All</a>
</todo-app>
</body>
</html>
+17
View File
@@ -0,0 +1,17 @@
<!doctype html>
<html id="page">
<head id="page-head">
<meta id="page-charset" charset="utf-8" />
<title id="page-title">shadow focus</title>
</head>
<body id="page-body">
<!-- Shaped like Compose for Web: the app mounts its accessibility tree
inside a shadow root on this element, so document.activeElement answers
with the host for every field the user is typing into. -->
<div id="app"></div>
<script id="page-script">
const root = document.getElementById("app").attachShadow({ mode: "open" });
root.innerHTML = `<input id="shadow-field" type="text" />`;
</script>
</body>
</html>
+41 -4
View File
@@ -32,13 +32,27 @@ func TranslateStringSelector(selector string) (string, bool, error) {
switch kind {
case "id", "resource-id":
return `[id="` + cssEscape(value) + `"]`, false, nil
case "idPrefix":
return `[id^="` + cssEscape(value) + `"]`, false, nil
case "class":
return `[class~="` + cssEscape(value) + `"]`, false, nil
case "tag":
return cssEscape(value), false, nil
case "text":
return `//*[normalize-space(text())=` + xpathStringLiteral(value) + `]`, true, nil
case "desc", "label", "content-desc", "accessibilityLabel", "accessibilityText", "ariaLabel", "aria-label":
// Substring of the element's whole text, the way internal/hierarchy
// reads the same selector: an element reading "Sent ✓" answers to
// text:Sent on every platform, and one React wrote as `{count} unsent`
// answers to text:unsent though its text arrives as two text nodes.
// normalize-space(text()) would read only the first of them. The
// not() clause is what keeps the badge's ancestors, up to <html>, from
// answering for it.
return `//*[` + innermostTextPredicate(value) + `]`, true, nil
case "desc":
// Mirrors the native rule: the label itself, or the label at the head of
// an iOS merged label ("account_card:7, Tim, $100").
escaped := cssEscape(value)
return `:is([aria-label="` + escaped + `"], [aria-label^="` + escaped + `, "])`, false, nil
case "label", "content-desc", "accessibilityLabel", "accessibilityText", "ariaLabel", "aria-label":
return `[aria-label="` + cssEscape(value) + `"]`, false, nil
case "descPrefix":
return `[aria-label^="` + cssEscape(value) + `"]`, false, nil
@@ -50,13 +64,25 @@ func TranslateStringSelector(selector string) (string, bool, error) {
return `:is([data-testid="` + escaped + `"], [id="` + escaped + `"])`, false, nil
case "testID", "testid", "data-testid":
return `[data-testid="` + cssEscape(value) + `"]`, false, nil
case "placeholder", "placeholderValue", "hintText":
case "placeholder":
// The attribute the markup writes. hintText and placeholderValue name
// the accessible-name ladder above it instead (fieldHint in driver.go),
// which no CSS says, so they fall through to a match that reaches
// nothing and the step fails by name. Building this selector for them
// tapped a field whose hint is its aria-label and whose placeholder
// happens to carry the value, which is an element neither matcher
// names: a selector reaches here only where the dump resolved it to no
// coordinates at all.
return `[placeholder="` + cssEscape(value) + `"]`, false, nil
default:
if !attrNamePattern.MatchString(kind) {
return "", false, fmt.Errorf("unsafe selector prefix %q", kind)
}
return `[` + kind + `="` + cssEscape(value) + `"]`, false, nil
operator := `*=`
if value == "true" || value == "false" {
operator = `=`
}
return `[` + kind + operator + `"` + cssEscape(value) + `"]`, false, nil
}
}
@@ -88,6 +114,17 @@ func cssEscape(value string) string {
return builder.String()
}
// innermostTextPredicate matches an element whose text contains value and whose
// descendants do not, which is the innermost match internal/hierarchy resolves
// the same selector to. The same predicate appears in
// pkg/spec/src/web-runtime.ts.
func innermostTextPredicate(value string) string {
contains := `contains(normalize-space(.), ` + xpathStringLiteral(
value,
) + `)`
return contains + ` and not(.//*[` + contains + `])`
}
// xpathStringLiteral wraps the value in a valid XPath 1.0 string literal.
// XPath 1.0 has no escape syntax, so a value containing both ' and " must be
// composed via concat(). The output already includes the surrounding quotes
+47 -7
View File
@@ -12,18 +12,45 @@ func TestTranslateStringSelector_KnownKeys(t *testing.T) {
{"resource-id:account-name", `[id="account-name"]`, false},
{"class:btn-primary", `[class~="btn-primary"]`, false},
{"tag:button", `button`, false},
{"text:Sign in", `//*[normalize-space(text())="Sign in"]`, true},
{`text:Say "hi"`, `//*[normalize-space(text())='Say "hi"']`, true},
{`text:it's`, `//*[normalize-space(text())="it's"]`, true},
{`text:it's "fine"`, `//*[normalize-space(text())=concat("it's ", '"', "fine", '"', "")]`, true},
{"desc:logout", `[aria-label="logout"]`, false},
{
"text:Sign in",
`//*[contains(normalize-space(.), "Sign in") and not(.//*[contains(normalize-space(.), "Sign in")])]`,
true,
},
{
`text:Say "hi"`,
`//*[contains(normalize-space(.), 'Say "hi"') and not(.//*[contains(normalize-space(.), 'Say "hi"')])]`,
true,
},
{
`text:it's`,
`//*[contains(normalize-space(.), "it's") and not(.//*[contains(normalize-space(.), "it's")])]`,
true,
},
{
`text:it's "fine"`,
`//*[contains(normalize-space(.), concat("it's ", '"', "fine", '"', "")) and not(.//*[contains(normalize-space(.), concat("it's ", '"', "fine", '"', ""))])]`,
true,
},
// desc also accepts an iOS merged label, the way internal/hierarchy does.
{"desc:logout", `:is([aria-label="logout"], [aria-label^="logout, "])`, false},
{"label:logout", `[aria-label="logout"]`, false},
{"accessibilityLabel:logout", `[aria-label="logout"]`, false},
{"aria-label:Sign in", `[aria-label="Sign in"]`, false},
{"descPrefix:account:", `[aria-label^="account:"]`, false},
// Must stay the string pkg/spec/src/web-runtime.ts builds for the same
// selector; the two translators feed the same page.
{"idPrefix:customer_row_", `[id^="customer_row_"]`, false},
{"testTag:submit", `:is([data-testid="submit"], [id="submit"])`, false},
{"testID:submit", `[data-testid="submit"]`, false},
{"placeholder:Email", `[placeholder="Email"]`, false},
// hintText and placeholderValue name the accessible-name ladder, which
// no CSS says, so they reach nothing here rather than the field whose
// placeholder happens to carry the value and whose hint is its
// aria-label. Both matchers name that field by its aria-label alone, so
// a tap by placeholder acts on an element nobody selected.
{"hintText:Email", `[hintText*="Email"]`, false},
{"placeholderValue:Email", `[placeholderValue*="Email"]`, false},
}
for _, testCase := range cases {
got, isXPath, err := TranslateStringSelector(testCase.selector)
@@ -43,8 +70,21 @@ func TestTranslateStringSelector_UnknownPrefixPassesThrough(t *testing.T) {
if err != nil {
t.Fatal(err)
}
if got != `[role="button"]` {
t.Errorf("unknown prefix should map to attribute selector, got %q", got)
if got != `[role*="button"]` {
t.Errorf(
"unknown prefix should map to a substring attribute selector, got %q",
got,
)
}
got, _, err = TranslateStringSelector("aria-expanded:true")
if err != nil {
t.Fatal(err)
}
if got != `[aria-expanded="true"]` {
t.Errorf(
"a boolean value should map to an exact attribute selector, got %q",
got,
)
}
}
+59
View File
@@ -4,9 +4,23 @@ package driver
import (
"context"
"encoding/json"
"errors"
"time"
)
// ErrGestureUndelivered reports a coordinate gesture that reached no element at
// all, so the app cannot have responded to it. It is not a device fault and
// says nothing about the device's health: the runner records it on the step
// rather than counting it toward the apply-failure streak, which is what makes
// a gesture that did nothing distinguishable from one the app ignored.
var ErrGestureUndelivered = errors.New("gesture reached no element")
// ErrSelectorMatchedNothing reports an action dispatched by selector whose
// selector named nothing on the current screen. It is a resolution failure, not
// a delivery one: no point was ever computed, so it stays separate from
// ErrGestureUndelivered and the runner records it as an unresolved selector.
var ErrSelectorMatchedNothing = errors.New("selector matched no element")
// DeviceDriver abstracts the platform-specific UI automation backend. v0.1
// surface matches proto/driverpb/driver.proto. The sidecar implementation
// lives under driver/sidecar; the web implementation under driver/chrome;
@@ -60,6 +74,20 @@ type ForegroundChecker interface {
ForegroundApp(ctx context.Context) (string, error)
}
// Scroller is the optional capability for drivers whose scroll interaction is
// not a finger drag. On a touch device the two are the same gesture, so a
// driver that does not implement this gets its Scroll actions as a Swipe. A
// browser scrolls on wheel input instead, and treats a drag as a drag.
type Scroller interface {
// Scroll moves the content under (fromX, fromY) by the vector to the
// destination point, the same endpoints Swipe takes.
Scroll(
ctx context.Context,
fromX, fromY, toX, toY int,
duration time.Duration,
) error
}
// TextReplacer is the optional capability for drivers whose InputText already
// replaces the field's content instead of appending to it. The runner must
// skip its pre-erase for such drivers: the erase would be a redundant
@@ -95,6 +123,37 @@ type LogEntry struct {
Message string
}
// ExceptionReporter is the optional capability for reporting the uncaught
// errors an app has captured so far. The runner feeds them to state.exceptions,
// which the default noUncaughtExceptions property reads. Drivers with no way to
// observe them simply do not implement it and the property stays vacuous there.
type ExceptionReporter interface {
Exceptions(ctx context.Context) ([]Exception, error)
}
// NavigationReporter is the optional capability for reporting the
// document-replacing navigations seen since the last call. A navigation
// restarts the app's own runtime, so a trace without them cannot separate an
// app that reloaded from a generator that repeated itself.
type NavigationReporter interface {
Navigations(ctx context.Context) ([]Navigation, error)
}
// Navigation is one document-replacing navigation: a reload, a form submit, a
// route change that swapped the document.
type Navigation struct {
URL string
UnixMillis int64
}
// Exception is one uncaught throwable the app captured.
type Exception struct {
Class string
Message string
StackTrace string
UnixMillis int64
}
type Image struct {
PNG []byte
Width int
+27 -3
View File
@@ -737,7 +737,19 @@ func (d *Driver) Terminate(ctx context.Context) error {
})
}
// offScreen reports a point the device has no surface under. The hierarchy
// reaches past the screen wherever a scroll container holds content beyond the
// fold, so an action derived from it can name a point no touch can land on.
// The far edge is exclusive: a touch at x == screenWidth arrives at
// screenWidth-1, which is a point the action never named.
func (d *Driver) offScreen(x, y int) bool {
return x < 0 || y < 0 || x >= d.screenWidth || y >= d.screenHeight
}
func (d *Driver) Tap(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
d.mu.Lock()
d.lastTap.x = float64(x)
d.lastTap.y = float64(y)
@@ -749,18 +761,27 @@ func (d *Driver) Tap(ctx context.Context, x, y int) error {
}
func (d *Driver) DoubleTap(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, doubleTapEvents(float64(x), float64(y), d.doubleTapGapMilliseconds)...)
})
}
func (d *Driver) LongPress(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, longPressEvents(float64(x), float64(y), longPressHoldMilliseconds)...)
})
}
func (d *Driver) Swipe(ctx context.Context, fromX, fromY, toX, toY int, duration time.Duration) error {
if d.offScreen(fromX, fromY) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, fromX, fromY)
}
seconds := duration.Seconds()
if seconds <= 0 {
seconds = 0.25
@@ -807,12 +828,15 @@ func (d *Driver) runnerTyper() transport.TextTyper {
}
// pressKeyUsage maps the logical key names mobile runs emit to a HID usage.
// Only Return/Enter has a hardware-keyboard equivalent on the simulator; other
// names (notably "back" and "home") have no HID key and report unsupported.
// Return/Enter and Escape are the hardware-keyboard keys the simulator has;
// other names (notably "back" and "home") have no HID key and report
// unsupported.
func pressKeyUsage(key string) (uint32, bool) {
switch key {
case "enter", "return", "Enter", "Return":
return usageReturn, true
case "escape", "Escape":
return usageEscape, true
default:
return 0, false
}
@@ -847,7 +871,7 @@ func (d *Driver) resolveSelectorCenter(ctx context.Context, selector string) (in
}
element := tree.Find(selector)
if element == nil {
return 0, 0, fmt.Errorf("selector %q matched no element", selector)
return 0, 0, fmt.Errorf("%w: %q", driver.ErrSelectorMatchedNothing, selector)
}
x, y := element.Bounds.Center()
return x, y, nil
+118
View File
@@ -1554,3 +1554,121 @@ func TestAcquireDeviceLockKeepsDistinctDevicesIndependent(t *testing.T) {
}
defer second.Close()
}
func TestSwipeReportsAGestureThatStartsOffScreen(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
// The hierarchy extends past the screen whenever a scroll container holds
// content below the fold, so the runner's clamp to the tree's extent can
// place a gesture where the device has no surface to receive it.
err := d.Swipe(
context.Background(),
200,
1200,
200,
800,
300*time.Millisecond,
)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf(
"Swipe starting below the screen: err = %v, want ErrGestureUndelivered",
err,
)
}
if slices.Contains(companion.recorded(), "hid") {
t.Fatal("an undeliverable swipe must not reach the companion")
}
}
func TestSwipeOnScreenReachesTheCompanion(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
if err := d.Swipe(context.Background(), 200, 600, 200, 200, 300*time.Millisecond); err != nil {
t.Fatalf("Swipe: %v", err)
}
if !slices.Contains(companion.recorded(), "hid") {
t.Fatal("an on-screen swipe should reach the companion")
}
}
// TestTapSelectorReportsASelectorThatMatchesNothing keeps a by-selector tap
// that resolved to no element out of the steps that read as dispatched.
func TestTapSelectorReportsASelectorThatMatchesNothing(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
err := d.TapSelector(context.Background(), "id:absent")
if !errors.Is(err, driver.ErrSelectorMatchedNothing) {
t.Fatalf("TapSelector on an absent element: err = %v, want ErrSelectorMatchedNothing", err)
}
if slices.Contains(companion.recorded(), "hid") {
t.Fatal("a selector that matched nothing must not reach the companion")
}
}
// TestGesturesRefuseTheFarEdgeOfTheScreen holds iOS to the extent the device
// enforces. A tap at x == screenWidth is not delivered at that point on the
// simulator, so admitting it dispatches a gesture the app receives somewhere
// other than where the trace says it landed.
func TestGesturesRefuseTheFarEdgeOfTheScreen(t *testing.T) {
gestures := map[string]func(*Driver, int, int) error{
"Tap": func(d *Driver, x, y int) error { return d.Tap(context.Background(), x, y) },
"DoubleTap": func(d *Driver, x, y int) error { return d.DoubleTap(context.Background(), x, y) },
"LongPress": func(d *Driver, x, y int) error { return d.LongPress(context.Background(), x, y) },
"Swipe": func(d *Driver, x, y int) error {
return d.Swipe(context.Background(), x, y, 200, 400, 300*time.Millisecond)
},
}
for name, gesture := range gestures {
t.Run(name, func(t *testing.T) {
for _, point := range []struct{ x, y int }{{390, 400}, {200, 844}} {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
err := gesture(d, point.x, point.y)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf("%s at (%d,%d): err = %v, want ErrGestureUndelivered",
name, point.x, point.y, err)
}
if slices.Contains(companion.recorded(), "hid") {
t.Fatalf("%s at (%d,%d) reached the companion", name, point.x, point.y)
}
}
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
if err := gesture(d, 389, 843); err != nil {
t.Fatalf("%s at the last on-screen point: %v", name, err)
}
})
}
}
// TestTapReportsAPointOffScreen is the iOS half of letting the runner hand off
// every point it resolved. The runner no longer refuses a point outside the
// screen, because a driver that can scroll reaches it; this one cannot, and a
// point with no surface under it has to be reported rather than synthesised
// into nothing.
func TestTapReportsAPointOffScreen(t *testing.T) {
for _, point := range []struct {
name string
x, y int
}{
{"above the screen", 200, -208},
{"below the screen", 200, 1200},
} {
t.Run(point.name, func(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
err := d.Tap(context.Background(), point.x, point.y)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf("Tap at (%d,%d): err = %v, want ErrGestureUndelivered",
point.x, point.y, err)
}
if slices.Contains(companion.recorded(), "hid") {
t.Fatal("an undeliverable tap must not reach the companion")
}
})
}
}
+98 -3
View File
@@ -42,6 +42,7 @@ type rawElement struct {
AXLabel *string `json:"AXLabel"`
AXValue *string `json:"AXValue"`
Type string `json:"type"`
Depth int `json:"depth"`
Enabled bool `json:"enabled"`
}
@@ -53,6 +54,7 @@ type treeNode struct {
Clickable *bool `json:"clickable,omitempty"`
Enabled *bool `json:"enabled,omitempty"`
Editable *bool `json:"editable,omitempty"`
Secure *bool `json:"secure,omitempty"`
}
// MapHierarchy converts a flat describe-all dump from the simulator companion
@@ -76,12 +78,18 @@ func MapHierarchy(dump []byte, screenWidth, screenHeight int) ([]byte, error) {
if len(dump) > 0 {
_ = json.Unmarshal(dump, &rawElements)
}
elements := make([]rawElement, 0, len(rawElements))
for _, raw := range rawElements {
var element rawElement
if err := json.Unmarshal(raw, &element); err != nil {
continue
}
if child, ok := mapElement(&element); ok {
elements = append(elements, element)
}
scrollable := scrollableElements(elements, screenWidth, screenHeight)
for index := range elements {
if child, ok := mapElement(&elements[index], scrollable[index]); ok {
root.Children = append(root.Children, child)
}
}
@@ -89,7 +97,84 @@ func MapHierarchy(dump []byte, screenWidth, screenHeight int) ([]byte, error) {
return json.Marshal(root)
}
func mapElement(element *rawElement) (treeNode, bool) {
// frameTolerance absorbs the sub-point rounding in companion frames, so a child
// that sits flush against its container's edge does not read as escaping it.
const frameTolerance = 0.5
// scrollableElements reports, per element, whether it is a container that clips
// content reaching past its own frame. That is the same fact Android reads off
// uiautomator's scrollable attribute and the web driver derives from overflow:
// the container can actually scroll, because there is content it is not showing.
//
// Three conditions together, because the snapshot has no clipping flag. The
// element must sit strictly inside its parent on at least one edge, which
// separates a real container from the stack of full-screen wrappers that
// inherit its overflow; some element in its subtree must lie outside it; and it
// must be on the screen, since a dismissed keyboard is reported below the screen
// and clips a much taller child without any gesture being able to reach it.
// A dump without depth (the legacy accessibility bridge) makes every element a
// root, and roots are never marked, so that path reports no scroll rather than
// a guessed one.
func scrollableElements(elements []rawElement, screenWidth, screenHeight int) []bool {
screen := rawFrame{Width: float64(screenWidth), Height: float64(screenHeight)}
scrollable := make([]bool, len(elements))
var ancestors []int
for index, element := range elements {
for len(ancestors) > 0 && elements[ancestors[len(ancestors)-1]].Depth >= element.Depth {
ancestors = ancestors[:len(ancestors)-1]
}
if len(ancestors) > 0 && hasArea(element.Frame) &&
overlaps(element.Frame, screen) &&
sitsInside(element.Frame, elements[ancestors[len(ancestors)-1]].Frame) &&
subtreeEscapes(elements, index) {
scrollable[index] = true
}
ancestors = append(ancestors, index)
}
return scrollable
}
// overlaps reports whether two frames share any area.
func overlaps(frame, other rawFrame) bool {
return frame.X < other.X+other.Width && other.X < frame.X+frame.Width &&
frame.Y < other.Y+other.Height && other.Y < frame.Y+frame.Height
}
// sitsInside reports whether frame is strictly smaller than container on at
// least one edge.
func sitsInside(frame, container rawFrame) bool {
return frame.X > container.X+frameTolerance ||
frame.Y > container.Y+frameTolerance ||
frame.X+frame.Width < container.X+container.Width-frameTolerance ||
frame.Y+frame.Height < container.Y+container.Height-frameTolerance
}
// subtreeEscapes reports whether any descendant of the element at index is
// positioned outside its frame. Descendants are the run that follows it while
// the depth stays greater, which is the pre-order walk the companion emits.
func subtreeEscapes(elements []rawElement, index int) bool {
frame := elements[index].Frame
for next := index + 1; next < len(elements) && elements[next].Depth > elements[index].Depth; next++ {
child := elements[next].Frame
if !hasArea(child) {
continue
}
if child.X < frame.X-frameTolerance ||
child.Y < frame.Y-frameTolerance ||
child.X+child.Width > frame.X+frame.Width+frameTolerance ||
child.Y+child.Height > frame.Y+frame.Height+frameTolerance {
return true
}
}
return false
}
func hasArea(frame rawFrame) bool {
return finite(frame.X) && finite(frame.Y) && finite(frame.Width) &&
finite(frame.Height) && frame.Width > 0 && frame.Height > 0
}
func mapElement(element *rawElement, scrollable bool) (treeNode, bool) {
if element.Type == "" {
return treeNode{}, false
}
@@ -106,6 +191,10 @@ func mapElement(element *rawElement) (treeNode, bool) {
"class": element.Type,
}
if scrollable {
attributes["scrollable"] = "true"
}
if id := stringValue(element.AXUniqueID); id != "" {
attributes["identifier"] = id
}
@@ -140,6 +229,11 @@ func mapElement(element *rawElement) (treeNode, bool) {
if editable {
yes := true
node.Editable = &yes
// Stated on every editable field, false included: a consumer deciding
// what a typed value may be recorded as has to tell "not a secure
// entry" apart from "nobody said".
secure := element.Type == "SecureTextField"
node.Secure = &secure
}
if element.Type == "Button" {
yes := true
@@ -150,7 +244,8 @@ func mapElement(element *rawElement) (treeNode, bool) {
}
func isEditable(elementType string) bool {
return elementType == "TextArea" || elementType == "TextField"
return elementType == "TextArea" || elementType == "TextField" ||
elementType == "SecureTextField"
}
// labelIsDisplayedText reports whether an element type's AXLabel is the string
@@ -4,6 +4,7 @@ import (
"encoding/json"
"os"
"path/filepath"
"slices"
"strings"
"testing"
@@ -331,3 +332,97 @@ func TestDumpIsCollapsed(t *testing.T) {
})
}
}
func scrollableBounds(tree *hierarchy.Tree) []hierarchy.Bounds {
var bounds []hierarchy.Bounds
for _, element := range tree.FindAllNodes("scrollable:true") {
bounds = append(bounds, element.Bounds)
}
return bounds
}
func TestScrollableMarksTheContainerThatClipsOverflowingContent(t *testing.T) {
tree := mapAndParse(
t,
readDump(t, "home-scrolling-describe.json"),
402,
874,
)
got := scrollableBounds(tree)
want := []hierarchy.Bounds{{Left: 0, Top: 122, Right: 402, Bottom: 699}}
if !slices.Equal(got, want) {
t.Fatalf("scrollable containers = %+v, want %+v", got, want)
}
}
func TestScrollableIsAbsentWhenNothingOverflows(t *testing.T) {
tree := mapAndParse(t, readDump(t, "home-fixed-describe.json"), 402, 874)
if got := scrollableBounds(tree); len(got) != 0 {
t.Fatalf(
"scrollable containers = %+v, want none on a screen that does not scroll",
got,
)
}
}
func TestScrollableIsAbsentWithoutTreeDepth(t *testing.T) {
tree := mapAndParse(t, readDump(t, "accounts-describe.json"), 402, 874)
if got := scrollableBounds(tree); len(got) != 0 {
t.Fatalf(
"scrollable containers = %+v, want none from a dump that carries no depth",
got,
)
}
}
func TestScrollableIgnoresAContainerOffTheScreen(t *testing.T) {
// The dismissed keyboard is reported below the screen, and its prediction
// bar clips a much taller child, so it satisfies every other condition.
dump := `[
{"type":"Application","depth":0,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true},
{"type":"Other","depth":1,"frame":{"x":0,"y":874,"width":402,"height":54},"enabled":true},
{"type":"Other","depth":2,"frame":{"x":0,"y":274,"width":402,"height":1254},"enabled":true}
]`
tree := mapAndParse(t, []byte(dump), 402, 874)
if got := scrollableBounds(tree); len(got) != 0 {
t.Fatalf("scrollable containers = %+v, want none off the screen", got)
}
}
// A secure text entry is what XCUITest reports a password field as
// (companion/Sources/ElementTypeName.swift). It has to reach the tree as a
// fact, because that is what lets a typed value be redacted from the record
// without redacting every other field's.
func TestSecureFactPerEditableType(t *testing.T) {
cases := []struct {
elementType string
wantReported bool
wantSecure bool
}{
{"SecureTextField", true, true},
{"TextField", true, false},
{"TextArea", true, false},
{"Button", false, false},
{"StaticText", false, false},
}
for _, testCase := range cases {
dump := `[{"type":"` + testCase.elementType + `","frame":{"x":0,"y":0,"width":10,"height":10},"enabled":true}]`
element := parseSingle(t, dump)
if element.SecureReported() != testCase.wantReported {
t.Errorf("%s reported secure = %v, want %v",
testCase.elementType, element.SecureReported(), testCase.wantReported)
}
if element.Secure != testCase.wantSecure {
t.Errorf("%s secure = %v, want %v", testCase.elementType, element.Secure, testCase.wantSecure)
}
}
}
// A secure text entry is a field a run must be able to type into, or the login
// screens every real app opens on are unreachable.
func TestSecureTextFieldIsEditable(t *testing.T) {
dump := `[{"type":"SecureTextField","frame":{"x":0,"y":0,"width":10,"height":10},"enabled":true}]`
if element := parseSingle(t, dump); !element.Editable {
t.Error("a secure text entry must be editable")
}
}
+1
View File
@@ -9,6 +9,7 @@ const (
usage1 = 30
usage0 = 39
usageReturn = 40
usageEscape = 41
usageTab = 43
usageSpace = 44
usageBackspace = 42
@@ -0,0 +1,47 @@
package ioscompanion
import (
"context"
"testing"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/transport"
)
// keyRecordingCompanion keeps the HID events a press produced, so the assertion
// is over what reached the transport rather than over the lookup that built it.
type keyRecordingCompanion struct {
fakeCompanion
events []transport.HIDEvent
}
func (c *keyRecordingCompanion) SendHID(
_ context.Context,
events ...transport.HIDEvent,
) error {
c.events = append(c.events, events...)
return nil
}
// docs/manual/spec-language.md documents escape, and the simulator's HID stream
// carries it: usage 41 is the keyboard escape. Without it a spec clause over
// escape reports unsupported on iOS while the same clause runs on web.
func TestPressKeyEscapeReachesTheHIDStream(t *testing.T) {
companion := &keyRecordingCompanion{}
d := newTestDriver(companion)
if err := d.PressKey(context.Background(), "escape"); err != nil {
t.Fatalf("PressKey escape: %v", err)
}
// 41 is the USB HID keyboard escape usage, stated here rather than read
// from the production table so a wrong table entry cannot agree with itself.
want := []transport.HIDEvent{transport.KeyDown(41), transport.KeyUp(41)}
if len(companion.events) != len(want) {
t.Fatalf("sent %v, want %v", companion.events, want)
}
for index, event := range companion.events {
if event != want[index] {
t.Fatalf("sent %v, want %v", companion.events, want)
}
}
}
@@ -0,0 +1,35 @@
[
{"type":"Application","depth":0,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":"Folio","AXValue":null,"AXUniqueId":null},
{"type":"Window","depth":1,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":2,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":3,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":4,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":5,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":6,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"ScrollView","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":8,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":9,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":"HomeScreen"},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":76,"width":97.33,"height":24},"enabled":true,"AXLabel":"Accounts","AXValue":null,"AXUniqueId":null},
{"type":"Button","depth":10,"frame":{"x":340,"y":71.33,"width":48,"height":48},"enabled":true,"AXLabel":"Log out","AXValue":null,"AXUniqueId":"LogoutButton"},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":100,"width":104,"height":14.33},"enabled":true,"AXLabel":"[email protected]","AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":122.33,"width":402,"height":577},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":11,"frame":{"x":0,"y":122.33,"width":402,"height":577},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":11,"frame":{"x":173,"y":178.33,"width":56,"height":56},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":11,"frame":{"x":131.33,"y":248.33,"width":139.33,"height":18.0},"enabled":true,"AXLabel":"No accounts yet","AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":11,"frame":{"x":36,"y":272.33,"width":330,"height":28.67},"enabled":true,"AXLabel":"Create your first account to start tracking transactions.","AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":676,"width":402,"height":48},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":716.33,"width":105.67,"height":14.33},"enabled":true,"AXLabel":"TOTAL BALANCE","AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":20,"y":716.33,"width":362,"height":107.67},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":730.67,"width":85.33,"height":33.33},"enabled":true,"AXLabel":"$0.00","AXValue":null,"AXUniqueId":"TotalBalance"},
{"type":"StaticText","depth":10,"frame":{"x":307.67,"y":749.67,"width":74.33,"height":14.33},"enabled":true,"AXLabel":"0 accounts","AXValue":null,"AXUniqueId":null},
{"type":"Button","depth":10,"frame":{"x":20,"y":777,"width":362,"height":48},"enabled":true,"AXLabel":"+ Add account","AXValue":null,"AXUniqueId":"AddAccountButton"},
{"type":"StaticText","depth":11,"frame":{"x":140.67,"y":792,"width":120.67,"height":18},"enabled":true,"AXLabel":"+ Add account","AXValue":null,"AXUniqueId":null},
{"type":"Window","depth":1,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":2,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":3,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null}
]
@@ -0,0 +1,35 @@
[
{"type":"Application","depth":0,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":"Folio","AXValue":null,"AXUniqueId":null},
{"type":"Window","depth":1,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":2,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":3,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":4,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":5,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":6,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"ScrollView","depth":7,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":8,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":9,"frame":{"x":0,"y":0,"width":402,"height":874},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":"HomeScreen"},
{"type":"Other","depth":10,"frame":{"x":0,"y":62,"width":402,"height":778},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":76,"width":97.33,"height":24},"enabled":true,"AXLabel":"Accounts","AXValue":null,"AXUniqueId":null},
{"type":"Button","depth":10,"frame":{"x":340,"y":71.33,"width":48,"height":48},"enabled":true,"AXLabel":"Log out","AXValue":null,"AXUniqueId":"LogoutButton"},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":100,"width":104,"height":14.33},"enabled":true,"AXLabel":"[email protected]","AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":10,"frame":{"x":0,"y":122.33,"width":402,"height":577},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Other","depth":11,"frame":{"x":0,"y":122.33,"width":402,"height":577},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"Button","depth":11,"frame":{"x":20,"y":704.33,"width":362,"height":19},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":"AccountCard"},
{"type":"Button","depth":11,"frame":{"x":20,"y":130.33,"width":362,"height":72},"enabled":true,"AXLabel":"A0, Account 0, $0.00, 0 transactions","AXValue":null,"AXUniqueId":"AccountCard"},
{"type":"StaticText","depth":12,"frame":{"x":90,"y":150.33,"width":83.67,"height":18},"enabled":true,"AXLabel":"Account 0","AXValue":null,"AXUniqueId":"AccountName"},
{"type":"StaticText","depth":12,"frame":{"x":90,"y":168.33,"width":104,"height":14.33},"enabled":true,"AXLabel":"0 transactions","AXValue":null,"AXUniqueId":"AccountTxnCount"},
{"type":"Button","depth":11,"frame":{"x":20,"y":212.33,"width":362,"height":72.0},"enabled":true,"AXLabel":"A1, Account 1, $0.00, 0 transactions","AXValue":null,"AXUniqueId":"AccountCard"},
{"type":"StaticText","depth":12,"frame":{"x":90,"y":232.33,"width":83.67,"height":18},"enabled":true,"AXLabel":"Account 1","AXValue":null,"AXUniqueId":"AccountName"},
{"type":"Button","depth":11,"frame":{"x":20,"y":786.33,"width":362,"height":72},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":"AccountCard"},
{"type":"Button","depth":11,"frame":{"x":20,"y":1278.33,"width":362,"height":72},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":"AccountCard"},
{"type":"Other","depth":10,"frame":{"x":0,"y":676,"width":402,"height":48},"enabled":true,"AXLabel":null,"AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":716.33,"width":105.67,"height":14.33},"enabled":true,"AXLabel":"TOTAL BALANCE","AXValue":null,"AXUniqueId":null},
{"type":"StaticText","depth":10,"frame":{"x":20,"y":730.67,"width":85.33,"height":33.33},"enabled":true,"AXLabel":"$0.00","AXValue":null,"AXUniqueId":"TotalBalance"},
{"type":"Button","depth":10,"frame":{"x":20,"y":777,"width":362,"height":48},"enabled":true,"AXLabel":"+ Add account","AXValue":null,"AXUniqueId":"AddAccountButton"},
{"type":"StaticText","depth":11,"frame":{"x":140.67,"y":792,"width":120.67,"height":18},"enabled":true,"AXLabel":"+ Add account","AXValue":null,"AXUniqueId":null}
]
@@ -336,7 +336,13 @@ func (c *runnerCompanion) PressKey(ctx context.Context, key string) error {
case "enter", "return", "Enter", "Return":
_, err := c.call(ctx, "pressKey", map[string]any{"key": "return"})
return err
case "escape", "Escape":
_, err := c.call(ctx, "pressKey", map[string]any{"key": "escape"})
return err
default:
return fmt.Errorf("runner companion cannot press key %q; only return is supported", key)
return fmt.Errorf(
"runner companion cannot press key %q; only return and escape are supported",
key,
)
}
}
@@ -58,7 +58,7 @@ type TextEditor interface {
// EraseText deletes characterCount characters from the focused field.
EraseText(ctx context.Context, characterCount int) error
// PressKey presses the named logical key (currently only return/enter).
// PressKey presses the named logical key (return/enter and escape).
PressKey(ctx context.Context, key string) error
}
+21 -5
View File
@@ -8,7 +8,9 @@ import (
"time"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/credentials/insecure"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/android"
"github.com/priyanshujain/sanderling/internal/driver"
@@ -129,19 +131,33 @@ func (c *Client) Terminate(ctx context.Context) error {
return err
}
// asGestureError translates the sidecar's OUT_OF_RANGE refusal of a point the
// device has no surface under into driver.ErrGestureUndelivered. Only the
// sidecar knows the screen extent, and the platform silently drops such a
// gesture, so without this the step reads as an action that landed.
func asGestureError(err error) error {
if status.Code(err) != codes.OutOfRange {
return err
}
return fmt.Errorf("%w: %s", driver.ErrGestureUndelivered, status.Convert(err).Message())
}
func (c *Client) Tap(ctx context.Context, x, y int) error {
_, err := c.stub.Tap(ctx, &driverpb.Point{X: int32(x), Y: int32(y)})
return err
return asGestureError(err)
}
func (c *Client) LongPress(ctx context.Context, x, y int) error {
_, err := c.stub.LongPress(ctx, &driverpb.Point{X: int32(x), Y: int32(y)})
return err
return asGestureError(err)
}
func (c *Client) TapSelector(ctx context.Context, selector string) error {
_, err := c.stub.TapSelector(ctx, &driverpb.Selector{Value: selector})
return err
if status.Code(err) == codes.NotFound {
return fmt.Errorf("%w: %s", driver.ErrSelectorMatchedNothing, status.Convert(err).Message())
}
return asGestureError(err)
}
// doubleTapGap is the inter-tap delay for the selector fallback: short enough
@@ -155,7 +171,7 @@ const doubleTapGap = 50 * time.Millisecond
// navigation to interleave between the taps.
func (c *Client) DoubleTap(ctx context.Context, x, y int) error {
_, err := c.stub.DoubleTap(ctx, &driverpb.Point{X: int32(x), Y: int32(y)})
return err
return asGestureError(err)
}
func (c *Client) DoubleTapSelector(ctx context.Context, selector string) error {
@@ -192,7 +208,7 @@ func (c *Client) Swipe(ctx context.Context, fromX, fromY, toX, toY int, duration
To: &driverpb.Point{X: int32(toX), Y: int32(toY)},
DurationMillis: duration.Milliseconds(),
})
return err
return asGestureError(err)
}
func (c *Client) PressKey(ctx context.Context, key string) error {
+82 -2
View File
@@ -2,6 +2,7 @@ package sidecar
import (
"context"
"errors"
"io"
"net"
"strings"
@@ -13,6 +14,7 @@ import (
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/driver"
driverpb "github.com/priyanshujain/sanderling/proto/driverpb"
)
@@ -46,8 +48,10 @@ type fakeServer struct {
logEntries []*driverpb.LogEntry
metrics *driverpb.MetricsResponse
healthError error
tapError error
healthError error
tapError error
gestureError error
selectorError error
}
func (s *fakeServer) Health(_ context.Context, _ *driverpb.Empty) (*driverpb.HealthStatus, error) {
@@ -85,6 +89,9 @@ func (s *fakeServer) Tap(_ context.Context, point *driverpb.Point) (*driverpb.Em
if s.tapError != nil {
return nil, s.tapError
}
if s.gestureError != nil {
return nil, s.gestureError
}
s.taps = append(s.taps, point.GetX(), point.GetY())
return &driverpb.Empty{}, nil
}
@@ -92,6 +99,9 @@ func (s *fakeServer) Tap(_ context.Context, point *driverpb.Point) (*driverpb.Em
func (s *fakeServer) TapSelector(_ context.Context, selector *driverpb.Selector) (*driverpb.Empty, error) {
s.mutex.Lock()
defer s.mutex.Unlock()
if s.selectorError != nil {
return nil, s.selectorError
}
s.tapSelectors = append(s.tapSelectors, selector.GetValue())
return &driverpb.Empty{}, nil
}
@@ -113,6 +123,9 @@ func (s *fakeServer) WaitForIdle(_ context.Context, duration *driverpb.Duration)
func (s *fakeServer) LongPress(_ context.Context, point *driverpb.Point) (*driverpb.Empty, error) {
s.mutex.Lock()
defer s.mutex.Unlock()
if s.gestureError != nil {
return nil, s.gestureError
}
s.longPresses = append(s.longPresses, point.GetX(), point.GetY())
return &driverpb.Empty{}, nil
}
@@ -120,6 +133,9 @@ func (s *fakeServer) LongPress(_ context.Context, point *driverpb.Point) (*drive
func (s *fakeServer) DoubleTap(_ context.Context, point *driverpb.Point) (*driverpb.Empty, error) {
s.mutex.Lock()
defer s.mutex.Unlock()
if s.gestureError != nil {
return nil, s.gestureError
}
s.doubleTaps = append(s.doubleTaps, point.GetX(), point.GetY())
return &driverpb.Empty{}, nil
}
@@ -127,6 +143,9 @@ func (s *fakeServer) DoubleTap(_ context.Context, point *driverpb.Point) (*drive
func (s *fakeServer) Swipe(_ context.Context, request *driverpb.SwipeRequest) (*driverpb.Empty, error) {
s.mutex.Lock()
defer s.mutex.Unlock()
if s.gestureError != nil {
return nil, s.gestureError
}
s.swipes = append(s.swipes, request)
return &driverpb.Empty{}, nil
}
@@ -682,3 +701,64 @@ func TestClient_RecentLogsSinceBranches(t *testing.T) {
})
}
}
// TestClient_SelectorThatMatchesNothingReportsIt keeps a by-selector tap that
// named no element distinguishable from a gesture that reached no point: the
// sidecar refuses it with NOT_FOUND and the client names the resolution
// failure rather than the delivery one.
func TestClient_SelectorThatMatchesNothingReportsIt(t *testing.T) {
taps := map[string]func(*Client) error{
"TapSelector": func(c *Client) error { return c.TapSelector(context.Background(), "id:absent") },
"DoubleTapSelector": func(c *Client) error { return c.DoubleTapSelector(context.Background(), "id:absent") },
}
for name, tap := range taps {
t.Run(name, func(t *testing.T) {
state := newHarness(t)
state.fake.mutex.Lock()
state.fake.selectorError = status.Error(
codes.NotFound, "selector id:absent matched no element")
state.fake.mutex.Unlock()
client, _ := Dial(state.address)
defer client.Close()
err := tap(client)
if !errors.Is(err, driver.ErrSelectorMatchedNothing) {
t.Fatalf("err = %v, want ErrSelectorMatchedNothing", err)
}
if errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatal("a selector that matched nothing must not read as an undelivered gesture")
}
})
}
}
// TestClient_OffScreenGestureReportsUndelivered holds the Android client to the
// contract Chrome and iOS already meet: a gesture the device had no surface
// under reports driver.ErrGestureUndelivered rather than returning nil and
// letting the step read as an action that landed.
func TestClient_OffScreenGestureReportsUndelivered(t *testing.T) {
gestures := map[string]func(*Client) error{
"Tap": func(c *Client) error { return c.Tap(context.Background(), 160, 900) },
"DoubleTap": func(c *Client) error { return c.DoubleTap(context.Background(), 160, 900) },
"LongPress": func(c *Client) error { return c.LongPress(context.Background(), 160, 900) },
"Swipe": func(c *Client) error {
return c.Swipe(context.Background(), 160, 900, 160, 700, time.Second)
},
}
for name, gesture := range gestures {
t.Run(name, func(t *testing.T) {
state := newHarness(t)
state.fake.mutex.Lock()
state.fake.gestureError = status.Error(
codes.OutOfRange, "gesture point (160,900) is outside the 320x640 screen")
state.fake.mutex.Unlock()
client, _ := Dial(state.address)
defer client.Close()
err := gesture(client)
if !errors.Is(err, driver.ErrGestureUndelivered) {
t.Fatalf("err = %v, want ErrGestureUndelivered", err)
}
})
}
}
+488 -86
View File
@@ -6,9 +6,11 @@
// String selectors (global scan or element-scoped):
// attribute:value - substring match; exact for "true"/"false" booleans
// id:<suffix> - substring on resource-id / identifier (backward compat)
// text:<value> - substring on text attribute
// idPrefix:<prefix> - starts-with on resource-id / identifier, package prefix skipped
// text:<value> - substring on text attribute, innermost match only
// desc:<value> - substring on content-desc / accessibilityText
// descPrefix:<prefix> - starts-with on content-desc / accessibilityText
// tag:<value> - exact match on the element's tag name (web)
//
// Object selectors (multi-attribute AND, element-scoped or global):
// { attr: value, ... } - all key/value pairs must match, each key resolved by
@@ -17,11 +19,14 @@
// Path queries (global scan only, string form):
// <sel> > <sel> > ... - each segment matched within subtree of previous match
//
// Cross-platform aliases are expanded automatically: "label" / "accessibilityLabel"
// resolve to accessibilityText; "content-desc" also checks accessibilityText and
// vice-versa; "identifier" / "accessibilityIdentifier" / "testTag" resolve to
// resource-id (and to each other) so a Compose testTag matches whether the
// underlying platform exposes it as resource-id (Android) or accessibilityIdentifier (iOS).
// Cross-platform aliases are expanded automatically, one level deep: every name
// for a fact lists every key a producer writes it under rather than hopping
// through another alias. "label" / "accessibilityLabel" / "ariaLabel" /
// "contentDescription" resolve to accessibilityText and content-desc, which also
// check each other; "identifier" / "accessibilityIdentifier" / "testTag" /
// "testID" resolve to resource-id, to each other and to data-testid, so a
// Compose testTag matches whether the platform exposes it as resource-id
// (Android), accessibilityIdentifier (iOS) or data-testid (web).
package hierarchy
import (
@@ -29,6 +34,7 @@ import (
"fmt"
"maps"
"regexp"
"slices"
"sort"
"strconv"
"strings"
@@ -69,10 +75,20 @@ type Element struct {
Focused bool `json:"focused,omitempty"`
Selected bool `json:"selected,omitempty"`
Editable bool `json:"editable,omitempty"`
Secure bool `json:"secure,omitempty"`
Bounds Bounds `json:"bounds"`
Attributes map[string]string `json:"attrs,omitempty"`
}
// SecureReported reports whether the producer stated this element's secure
// fact at all. Android never does, so an element without it is unknown rather
// than known not to be a secure entry, and a caller deciding what may be
// written down has to tell those two apart.
func (e *Element) SecureReported() bool {
_, reported := e.Attributes["secure"]
return reported
}
// Node is one node in the hierarchy tree.
type Node struct {
Element
@@ -84,18 +100,149 @@ type Node struct {
type Tree struct {
Root *Node `json:"-"`
Elements []*Element `json:"elements"`
// UnreadableFlags counts the boolean fields the producer sent as something
// other than a boolean. They are dropped rather than failing the dump, so
// the count is what keeps the drop from being silent.
UnreadableFlags int `json:"unreadableFlags,omitempty"`
}
// treeJSON is the stored form of a Tree. `depths` is the pre-order depth of
// each element, which is what turns the flat array back into Root: a stored
// tree without it (every trace written before the field existed) decodes with
// a nil Root and resolves no selector, exactly as it did before.
//
// A depth per element rather than a parent index per element: the numbers are
// one digit deep into most hierarchies where a parent index is three, and a
// step already costs 86 KB on android.
type treeJSON struct {
Elements []*Element `json:"elements"`
Depths []int `json:"depths,omitempty"`
UnreadableFlags int `json:"unreadable_flags,omitempty"`
}
func (t Tree) MarshalJSON() ([]byte, error) {
return json.Marshal(treeJSON{
Elements: t.Elements,
Depths: t.depths(),
UnreadableFlags: t.UnreadableFlags,
})
}
// depths walks Root, and yields nothing unless the walk covers exactly the
// elements the flat array holds: a hand-built Tree whose Root and Elements
// disagree would otherwise store a shape that rebuilds into a different tree.
func (t Tree) depths() []int {
if t.Root == nil {
return nil
}
depths := make([]int, 0, len(t.Elements))
var walk func(node *Node, depth int)
walk = func(node *Node, depth int) {
depths = append(depths, depth)
for _, child := range node.Children {
walk(child, depth+1)
}
}
walk(t.Root, 0)
if len(depths) != len(t.Elements) {
return nil
}
return depths
}
func (t *Tree) UnmarshalJSON(data []byte) error {
var stored treeJSON
if err := json.Unmarshal(data, &stored); err != nil {
return err
}
t.Elements = stored.Elements
t.UnreadableFlags = stored.UnreadableFlags
t.Root = t.rebuild(stored.Depths)
return nil
}
// rebuild re-parents the flat pre-order array from the stored depths. Every
// element is re-seated inside its Node so Tree.Elements and &node.Element stay
// the same pointer, which is the identity the verifier's element scope and the
// picker's target list are keyed on.
func (t *Tree) rebuild(depths []int) *Node {
if !wellFormedDepths(depths, len(t.Elements)) {
return nil
}
stack := make([]*Node, 0, 32)
for index, depth := range depths {
node := &Node{Element: *t.Elements[index], tree: t}
t.Elements[index] = &node.Element
stack = stack[:depth]
if depth > 0 {
parent := stack[depth-1]
parent.Children = append(parent.Children, node)
}
stack = append(stack, node)
}
return stack[0]
}
// wellFormedDepths accepts only a single-rooted pre-order sequence: one root at
// the head and no child deeper than one level below its predecessor.
func wellFormedDepths(depths []int, elementCount int) bool {
if len(depths) == 0 || len(depths) != elementCount || depths[0] != 0 {
return false
}
for index := 1; index < len(depths); index++ {
if depths[index] < 1 || depths[index] > depths[index-1]+1 {
return false
}
}
return true
}
// treeNodeJSON mirrors the sidecar TreeNode JSON structure.
type treeNodeJSON struct {
Attributes map[string]string `json:"attributes"`
Children []treeNodeJSON `json:"children"`
Clickable *bool `json:"clickable"`
Enabled *bool `json:"enabled"`
Focused *bool `json:"focused"`
Checked *bool `json:"checked"`
Selected *bool `json:"selected"`
Editable *bool `json:"editable"`
Clickable flagJSON `json:"clickable"`
Enabled flagJSON `json:"enabled"`
Focused flagJSON `json:"focused"`
Checked flagJSON `json:"checked"`
Selected flagJSON `json:"selected"`
Editable flagJSON `json:"editable"`
Secure flagJSON `json:"secure"`
}
// flagJSON is one boolean field of a node. A value that is not a boolean
// leaves the flag unset and marks itself unreadable rather than failing the
// document: a dump is one observation of a whole screen, and one node's bit is
// no reason to discard every element on it. Malformed bounds are already
// treated this way.
type flagJSON struct {
set bool
value bool
unreadable bool
}
func (f *flagJSON) UnmarshalJSON(data []byte) error {
if string(data) == "null" {
return nil
}
var value bool
if err := json.Unmarshal(data, &value); err != nil {
f.unreadable = true
return nil
}
f.set = true
f.value = value
return nil
}
func (n *treeNodeJSON) unreadableFlags() int {
count := 0
for _, flag := range []flagJSON{n.Clickable, n.Enabled, n.Focused, n.Checked, n.Selected, n.Editable, n.Secure} {
if flag.unreadable {
count++
}
}
return count
}
// Selector describes a multi-attribute AND match.
@@ -115,9 +262,14 @@ type AttrFilter struct {
var attributeAliases = map[string][]string{
// Android XML legacy name; web driver uses content-desc; the sidecar normalises to accessibilityText
"content-desc": {"accessibilityText"},
// iOS AXElement / UIKit names
"label": {"accessibilityText"},
"accessibilityLabel": {"accessibilityText"},
// Every other name for the accessible label. Alias expansion is one level,
// so each name lists both keys a producer writes the fact under rather than
// hopping through accessibilityText: android and the chrome dump write
// content-desc, the ios sidecar writes accessibilityText.
"label": {"accessibilityText", "content-desc"},
"accessibilityLabel": {"accessibilityText", "content-desc"},
"ariaLabel": {"accessibilityText", "content-desc"},
"contentDescription": {"accessibilityText", "content-desc"},
// accessibilityText is the canonical key; also check content-desc for Android/web
"accessibilityText": {"content-desc"},
// resource-id canonical key; also check identifier (iOS AXElement raw field)
@@ -125,19 +277,188 @@ var attributeAliases = map[string][]string{
// iOS identifier names
"identifier": {"resource-id", "accessibilityIdentifier"},
"accessibilityIdentifier": {"resource-id", "identifier"},
// Compose testTag surfaces as resource-id on Android, accessibilityIdentifier on iOS
"testTag": {"resource-id", "identifier", "accessibilityIdentifier"},
// Compose testTag surfaces as resource-id on Android, accessibilityIdentifier
// on iOS and data-testid on web, which is the key the web runtime resolves
// both names against.
"testTag": {"resource-id", "identifier", "accessibilityIdentifier", "data-testid"},
"testID": {"data-testid"},
// iOS AXElement raw name for hintText
"placeholderValue": {"hintText"},
// iOS AXElement raw name for class
"elementType": {"class"},
// DOM property name for class; every producer writes the attribute as class
"className": {"class"},
}
// matchAttr returns true when the element has an attribute matching attr:value.
// Alias expansion is applied so cross-platform names resolve correctly.
// Boolean values ("true"/"false") use exact comparison; all others use substring.
// Returns false gracefully when no candidate attribute has data.
// selectorKeys is every key an object selector may use. It is the union of the
// selector kinds, the attribute names the drivers emit on some platform, and
// the cross-platform aliases, so a key that is meaningful on ONE platform stays
// silently empty on the others rather than failing the run there.
//
// pkg/spec/test/fixtures/selector-keys.json holds the same list for the web
// runtime; a test on each side asserts its own list against that file, which is
// what keeps one spec from being accepted by one runtime and rejected by the
// other.
var selectorKeys = []string{
"accessibilityIdentifier",
"accessibilityLabel",
"accessibilityText",
"aria-label",
"ariaLabel",
"checked",
"class",
"className",
"clickable",
"content-desc",
"contentDescription",
"data-testid",
"desc",
"descPrefix",
"editable",
"elementType",
"enabled",
"focused",
"hintText",
"id",
"idPrefix",
"identifier",
"label",
"package",
"placeholder",
"placeholderValue",
"resource-id",
"scrollable",
"secure",
"selected",
"tag",
"testID",
"testTag",
"text",
"title",
"value",
}
var selectorKeySet = func() map[string]bool {
set := make(map[string]bool, len(selectorKeys))
for _, key := range selectorKeys {
set[key] = true
}
return set
}()
// SelectorKeys returns the accepted object-selector keys, sorted.
func SelectorKeys() []string {
return slices.Clone(selectorKeys)
}
// UnknownSelectorKeys returns the keys in sel that name neither an accepted
// selector key nor an attribute some element in the tree carries. Such a key
// can never match: the caller gets an empty result on every screen, which reads
// exactly like a screen that has no matching element.
func (t *Tree) UnknownSelectorKeys(sel Selector) []string {
if t == nil {
return nil
}
var unknown []string
for _, filter := range sel.Filters {
if selectorKeySet[filter.Attr] || t.carriesAttribute(filter.Attr) {
continue
}
if !slices.Contains(unknown, filter.Attr) {
unknown = append(unknown, filter.Attr)
}
}
return unknown
}
// carriesAttribute is the escape hatch for raw driver attributes this package
// does not enumerate: a key some element actually has is a key that can match.
func (t *Tree) carriesAttribute(key string) bool {
for _, element := range t.Elements {
if _, ok := element.Attributes[key]; ok {
return true
}
}
return false
}
// UnknownSelectorKeyMessage is the diagnostic for keys UnknownSelectorKeys
// returned. pkg/spec/src/web-runtime.ts raises the identical text, so one
// mistake reads the same whichever runtime the spec ran on.
func UnknownSelectorKeyMessage(keys []string) string {
quoted := make([]string, len(keys))
for i, key := range keys {
quoted[i] = strconv.Quote(key)
}
return fmt.Sprintf(
"selector key %s cannot match: no element carries that attribute, and it is not one of the accepted keys: %s",
strings.Join(quoted, ", "),
strings.Join(selectorKeys, ", "),
)
}
// matchSelectorKind resolves the selector keys that name a matching rule rather
// than an attribute: they read a derived field and compare it their own way,
// where an ordinary key does a substring test against the raw attribute map.
// The second return is false when kind names an ordinary attribute.
//
// The string form and the object form both come through here, so one key cannot
// mean one thing in "id:save" and another in {id: "save"}. It used to: the
// object form fell through to the attribute map, which carries no `id` or
// `desc` key on any platform, so those selectors matched nothing at all and
// said nothing about it.
func matchSelectorKind(element *Element, kind, value string) (bool, bool) {
switch kind {
case "id":
return element.ResourceID == value ||
strings.HasSuffix(element.ResourceID, ":id/"+value), true
case "idPrefix":
return matchIDPrefix(element.ResourceID, value), true
case "desc":
return element.Description == value ||
strings.HasPrefix(element.Description, value+", "), true
case "descPrefix":
return strings.HasPrefix(element.Description, value), true
case "tag":
// Both DOM resolvers compile this to a CSS type selector, which is the
// whole tag name. A substring rule here made {tag: "li"} name
// <todo-list> and {tag: "a"} name <todo-app>, so a selector meant for a
// row resolved to the container holding it.
tag, ok := element.Attributes["tag"]
return ok && tag == value, true
default:
return false, false
}
}
// matchIDPrefix is the id: rule with starts-with in place of equality: the
// whole identifier, or the local name after Android's "<package>:id/". Without
// the second form a role prefix would only match when the caller wrote the
// package out, which is exactly the string that varies between build variants.
func matchIDPrefix(resourceID, value string) bool {
if strings.HasPrefix(resourceID, value) {
return true
}
const marker = ":id/"
if index := strings.Index(resourceID, marker); index >= 0 {
return strings.HasPrefix(resourceID[index+len(marker):], value)
}
return false
}
// matchAttr returns true when the element matches key:value. It is the one
// entry point for both selector forms: the string form's kind and the object
// form's key are the same name and get the same rule.
//
// Keys naming a rule (id, desc and the prefix forms) resolve in
// matchSelectorKind; everything else is an attribute name, with alias expansion
// so cross-platform names resolve correctly. Boolean values ("true"/"false")
// use exact comparison; all others use substring. Returns false gracefully when
// no candidate attribute has data.
func matchAttr(element *Element, attr, value string) bool {
if matched, handled := matchSelectorKind(element, attr, value); handled {
return matched
}
candidates := append([]string{attr}, attributeAliases[attr]...)
for _, key := range candidates {
attrVal, ok := element.Attributes[key]
@@ -158,22 +479,62 @@ func matchAttr(element *Element, attr, value string) bool {
}
// matchSelector returns true when all filters in sel match the element (AND
// semantics). Each filter goes through match, the same rule the string form
// semantics). Each filter goes through matchAttr, the same rule the string form
// resolves a "kind:value" segment by, so {id: "Submit"} and "id:Submit" can
// never resolve to different elements. Applying matchAttr directly here made
// the object form skip the kind arms entirely: id, desc and descPrefix name no
// attribute any producer writes, so those keys matched NOTHING through an
// object selector while the string form matched, and every property over the
// never resolve to different elements. Reaching the attribute map directly here
// made the object form skip the kind arms entirely: id, desc and descPrefix
// name no attribute any producer writes, so those keys matched NOTHING through
// an object selector while the string form matched, and every property over the
// missing element passed vacuously.
func matchSelector(element *Element, sel Selector) bool {
for _, f := range sel.Filters {
if !match(element, f.Attr, f.Value) {
if !matchAttr(element, f.Attr, f.Value) {
return false
}
}
return true
}
func selectorReadsText(sel Selector) bool {
for _, f := range sel.Filters {
if f.Attr == "text" {
return true
}
}
return false
}
// innermostMatches drops a match a descendant of it also makes. An element's
// text is its whole subtree's text on web and on iOS, so every ancestor of a
// matching element matches too, up to the root, and the deepest match is the
// element the author named. An ancestor whose own text carries the value where
// no descendant of it does keeps its match.
func innermostMatches(nodes []*Node) []*Node {
if len(nodes) == 0 {
return nodes
}
matched := make(map[*Node]bool, len(nodes))
for _, node := range nodes {
matched[node] = true
}
var kept []*Node
for _, node := range nodes {
if !hasMatchingDescendant(node, matched) {
kept = append(kept, node)
}
}
return kept
}
func hasMatchingDescendant(node *Node, matched map[*Node]bool) bool {
for _, child := range node.Children {
if matched[child] || hasMatchingDescendant(child, matched) {
return true
}
}
return false
}
// Parse parses a sidecar TreeNode JSON hierarchy.
func Parse(text string) (*Tree, error) {
text = strings.TrimSpace(text)
@@ -190,6 +551,7 @@ func Parse(text string) (*Tree, error) {
}
func walkNode(node *treeNodeJSON, tree *Tree) *Node {
tree.UnreadableFlags += node.unreadableFlags()
n := &Node{Element: *elementFromNode(node), tree: tree}
tree.Elements = append(tree.Elements, &n.Element)
for i := range node.Children {
@@ -230,23 +592,26 @@ func elementFromNode(node *treeNodeJSON) *Element {
}
element.Screen = attrs["sanderling-screen"]
if node.Clickable != nil {
element.Clickable = *node.Clickable
if node.Clickable.set {
element.Clickable = node.Clickable.value
}
if node.Enabled != nil {
element.Enabled = *node.Enabled
if node.Enabled.set {
element.Enabled = node.Enabled.value
}
if node.Focused != nil {
element.Focused = *node.Focused
if node.Focused.set {
element.Focused = node.Focused.value
}
if node.Checked != nil {
element.Checked = *node.Checked
if node.Checked.set {
element.Checked = node.Checked.value
}
if node.Selected != nil {
element.Selected = *node.Selected
if node.Selected.set {
element.Selected = node.Selected.value
}
if node.Editable != nil {
element.Editable = *node.Editable
if node.Secure.set {
element.Secure = node.Secure.value
}
if node.Editable.set {
element.Editable = node.Editable.value
} else {
element.Editable = strings.Contains(element.Class, "EditText") || attrs["hintText"] != ""
}
@@ -260,20 +625,23 @@ func elementFromNode(node *treeNodeJSON) *Element {
element.Attributes = make(map[string]string, len(attrs)+5)
maps.Copy(element.Attributes, attrs)
if node.Clickable != nil {
element.Attributes["clickable"] = strconv.FormatBool(*node.Clickable)
if node.Clickable.set {
element.Attributes["clickable"] = strconv.FormatBool(node.Clickable.value)
}
if node.Enabled != nil {
element.Attributes["enabled"] = strconv.FormatBool(*node.Enabled)
if node.Enabled.set {
element.Attributes["enabled"] = strconv.FormatBool(node.Enabled.value)
}
if node.Focused != nil {
element.Attributes["focused"] = strconv.FormatBool(*node.Focused)
if node.Focused.set {
element.Attributes["focused"] = strconv.FormatBool(node.Focused.value)
}
if node.Checked != nil {
element.Attributes["checked"] = strconv.FormatBool(*node.Checked)
if node.Checked.set {
element.Attributes["checked"] = strconv.FormatBool(node.Checked.value)
}
if node.Selected != nil {
element.Attributes["selected"] = strconv.FormatBool(*node.Selected)
if node.Selected.set {
element.Attributes["selected"] = strconv.FormatBool(node.Selected.value)
}
if node.Secure.set {
element.Attributes["secure"] = strconv.FormatBool(node.Secure.value)
}
element.Attributes["editable"] = strconv.FormatBool(element.Editable)
@@ -346,20 +714,54 @@ func (t *Tree) FindAllNodes(selector string) []*Node {
return searchSubtree(t.Root, kind, value)
}
// FindBySelectorPath walks the selector chain starting from the tree root.
func (t *Tree) FindBySelectorPath(path []Selector) *Node {
if t == nil || t.Root == nil {
// FindBySelector returns the first Node in the tree matching sel, or nil. The
// root is a candidate, the way it is for the string form: one selector cannot
// mean one thing written "id:page" and another written {id: "page"}.
func (t *Tree) FindBySelector(sel Selector) *Node {
if t == nil {
return nil
}
return t.Root.FindBySelectorPath(path)
return firstNode(searchSubtreeBySelector(t.Root, sel))
}
// FindAllBySelector returns every Node in the tree matching sel, root included.
func (t *Tree) FindAllBySelector(sel Selector) []*Node {
if t == nil {
return nil
}
return searchSubtreeBySelector(t.Root, sel)
}
// FindBySelectorPath walks the selector chain starting from the tree root.
func (t *Tree) FindBySelectorPath(path []Selector) *Node {
if t == nil || t.Root == nil || len(path) == 0 {
return nil
}
for _, candidate := range t.FindAllBySelector(path[0]) {
if len(path) == 1 {
return candidate
}
if deeper := candidate.FindBySelectorPath(path[1:]); deeper != nil {
return deeper
}
}
return nil
}
// FindAllBySelectorPath walks the selector chain starting from the tree root.
func (t *Tree) FindAllBySelectorPath(path []Selector) []*Node {
if t == nil || t.Root == nil {
if t == nil || t.Root == nil || len(path) == 0 {
return nil
}
return t.Root.FindAllBySelectorPath(path)
var result []*Node
for _, candidate := range t.FindAllBySelector(path[0]) {
if len(path) == 1 {
result = append(result, candidate)
continue
}
result = append(result, candidate.FindAllBySelectorPath(path[1:])...)
}
return result
}
// Find returns the first Node scoped to this node (descendants, with spatial
@@ -376,9 +778,13 @@ func (n *Node) FindAll(selector string) []*Node {
if !ok {
return nil
}
return n.scopedNodes(func(element *Element) bool {
return match(element, kind, value)
nodes := n.scopedNodes(func(element *Element) bool {
return matchAttr(element, kind, value)
})
if kind == "text" {
return innermostMatches(nodes)
}
return nodes
}
// FindBySelector returns the first Node scoped to this node matching sel (AND semantics).
@@ -388,9 +794,13 @@ func (n *Node) FindBySelector(sel Selector) *Node {
// FindAllBySelector returns all Nodes scoped to this node matching sel (AND semantics).
func (n *Node) FindAllBySelector(sel Selector) []*Node {
return n.scopedNodes(func(element *Element) bool {
nodes := n.scopedNodes(func(element *Element) bool {
return matchSelector(element, sel)
})
if selectorReadsText(sel) {
return innermostMatches(nodes)
}
return nodes
}
// FindBySelectorPath walks a chain of selectors. The first selector is matched
@@ -578,11 +988,11 @@ func searchSubtree(root *Node, kind, value string) []*Node {
return nil
}
var result []*Node
if match(&root.Element, kind, value) {
result = append(result, root)
}
for _, child := range root.Children {
result = append(result, searchSubtree(child, kind, value)...)
collectMatches(root, func(element *Element) bool {
return matchAttr(element, kind, value)
}, &result)
if kind == "text" {
return innermostMatches(result)
}
return result
}
@@ -593,11 +1003,11 @@ func searchSubtreeBySelector(root *Node, sel Selector) []*Node {
return nil
}
var result []*Node
if matchSelector(&root.Element, sel) {
result = append(result, root)
}
for _, child := range root.Children {
result = append(result, searchSubtreeBySelector(child, sel)...)
collectMatches(root, func(element *Element) bool {
return matchSelector(element, sel)
}, &result)
if selectorReadsText(sel) {
return innermostMatches(result)
}
return result
}
@@ -610,24 +1020,6 @@ func parseSelector(selector string) (string, string, bool) {
return selector[:index], selector[index+1:], true
}
func match(element *Element, kind, value string) bool {
switch kind {
case "id":
if element.ResourceID == value {
return true
}
return strings.HasSuffix(element.ResourceID, ":id/"+value)
case "text":
return matchAttr(element, "text", value)
case "desc":
return element.Description == value || strings.HasPrefix(element.Description, value+", ")
case "descPrefix":
return strings.HasPrefix(element.Description, value)
default:
return matchAttr(element, kind, value)
}
}
// boundsPattern matches "[l,t,r,b]" (4-value Android/sidecar format).
var boundsPattern = regexp.MustCompile(`^\[(-?\d+),(-?\d+),(-?\d+),(-?\d+)\]$`)
@@ -659,3 +1051,13 @@ func parseBounds(text string) (Bounds, error) {
}
return Bounds{}, fmt.Errorf("bounds %q: not in [L,T,R,B] or [x1,y1][x2,y2] form", text)
}
// Tree returns the tree this node belongs to, or nil for a node built outside
// Parse. Selector validation needs the whole tree: a key absent from one
// subtree but present elsewhere is a key that can match.
func (n *Node) Tree() *Tree {
if n == nil {
return nil
}
return n.tree
}
+710 -4
View File
@@ -1,6 +1,11 @@
package hierarchy
import "testing"
import (
"encoding/json"
"slices"
"strings"
"testing"
)
// sampleDump is a sidecar TreeNode JSON equivalent of the old XML fixture.
const sampleDump = `{
@@ -106,6 +111,95 @@ func TestDescPrefix(t *testing.T) {
}
}
// idPrefixDump is a list whose rows carry a durable role prefix followed by the
// record's identifier, the convention that makes every row's full id unique and
// unwritable in a spec.
const idPrefixDump = `{
"attributes": {"resource-id": "com.example:id/customer_list", "bounds": "[0,0,100,400]"},
"children": [
{"attributes": {"resource-id": "com.example:id/customer_row_abc-123", "bounds": "[0,0,100,100]"}, "children": []},
{"attributes": {"resource-id": "com.example:id/customer_row_def-456", "bounds": "[0,100,100,200]"}, "children": []},
{"attributes": {"resource-id": "com.example:id/supplier_row_xyz", "bounds": "[0,200,100,300]"}, "children": []},
{"attributes": {"identifier": "customer_row_ghi-789", "bounds": "[0,300,100,400]"}, "children": []}
]
}`
func TestIDPrefixMatchesEveryRowSharingTheRole(t *testing.T) {
tree, _ := Parse(idPrefixDump)
rows := tree.FindAll("idPrefix:customer_row_")
if len(rows) != 3 {
t.Fatalf("want 3 customer rows, got %d", len(rows))
}
}
func TestIDPrefixDoesNotRequireThePackagePrefix(t *testing.T) {
tree, _ := Parse(idPrefixDump)
el := tree.Find("idPrefix:customer_row_abc")
if el == nil {
t.Fatal("expected the local name after :id/ to match on its own")
}
if el.ResourceID != "com.example:id/customer_row_abc-123" {
t.Fatalf("got %q", el.ResourceID)
}
}
func TestIDPrefixAlsoMatchesTheWholeIdentifier(t *testing.T) {
tree, _ := Parse(idPrefixDump)
if tree.Find("idPrefix:com.example:id/customer_row_") == nil {
t.Fatal("expected a package-qualified prefix to match")
}
}
func TestIDPrefixMatchesIOSAccessibilityIdentifier(t *testing.T) {
tree, _ := Parse(idPrefixDump)
el := tree.Find("idPrefix:customer_row_ghi")
if el == nil {
t.Fatal("expected identifier to match on a node with no resource-id")
}
if el.Bounds.Top != 300 {
t.Fatalf("matched the wrong node: %+v", el.Bounds)
}
}
func TestIDPrefixMatchesNothingWhenNoIDStartsWithIt(t *testing.T) {
tree, _ := Parse(idPrefixDump)
if rows := tree.FindAll("idPrefix:invoice_row_"); len(rows) != 0 {
t.Fatalf("want no matches, got %d", len(rows))
}
}
func TestIDPrefixIsNotASubstringMatch(t *testing.T) {
tree, _ := Parse(idPrefixDump)
if tree.Find("idPrefix:row_") != nil {
t.Fatal("expected starts-with, not substring")
}
}
// The string and object forms are one rule, so a prefix filter combined with a
// second attribute has to keep the same meaning it has on its own.
func TestIDPrefixInObjectSelector(t *testing.T) {
tree, _ := Parse(idPrefixDump)
sel := Selector{Filters: []AttrFilter{{Attr: "idPrefix", Value: "customer_row_"}}}
if nodes := tree.Root.FindAllBySelector(sel); len(nodes) != 3 {
t.Fatalf("want 3 customer rows, got %d", len(nodes))
}
}
func TestDescPrefixInObjectSelector(t *testing.T) {
input := `{
"attributes": {},
"children": [
{"attributes": {"content-desc": "customer_row_abc-123", "bounds": "[0,0,100,100]"}, "children": []},
{"attributes": {"content-desc": "supplier_row_xyz", "bounds": "[0,100,100,200]"}, "children": []}
]
}`
tree, _ := Parse(input)
sel := Selector{Filters: []AttrFilter{{Attr: "descPrefix", Value: "customer_row_"}}}
if nodes := tree.Root.FindAllBySelector(sel); len(nodes) != 1 {
t.Fatalf("want 1 customer row, got %d", len(nodes))
}
}
func TestBoolFieldsFromNode(t *testing.T) {
input := `{
"attributes": {"resource-id": "x", "bounds": "[0,0,100,100]"},
@@ -141,6 +235,53 @@ func TestBoolFieldsFromNode(t *testing.T) {
}
}
// secure is the one state flag with three answers: a producer that reports
// nothing leaves the element unknown rather than known-not-secure, and a
// consumer deciding what a typed value may be written into a record reads the
// difference.
func TestSecureIsUnknownUntilAProducerReportsIt(t *testing.T) {
cases := []struct {
name string
node string
wantReported bool
wantSecure bool
}{
{"reported secure", `{"attributes": {"bounds": "[0,0,10,10]"}, "secure": true}`, true, true},
{"reported not secure", `{"attributes": {"bounds": "[0,0,10,10]"}, "secure": false}`, true, false},
{"never reported", `{"attributes": {"bounds": "[0,0,10,10]"}}`, false, false},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
tree, err := Parse(testCase.node)
if err != nil {
t.Fatalf("Parse: %v", err)
}
element := tree.Elements[0]
if element.SecureReported() != testCase.wantReported {
t.Errorf("SecureReported = %v, want %v", element.SecureReported(), testCase.wantReported)
}
if element.Secure != testCase.wantSecure {
t.Errorf("Secure = %v, want %v", element.Secure, testCase.wantSecure)
}
})
}
}
// A selector reaches the fact by the same route every other boolean state does.
func TestSecureIsSelectable(t *testing.T) {
tree, err := Parse(`{"attributes": {"resource-id": "root", "bounds": "[0,0,10,10]"}, "children": [
{"attributes": {"resource-id": "pwd", "bounds": "[0,0,10,5]"}, "secure": true, "children": []},
{"attributes": {"resource-id": "email", "bounds": "[0,5,10,10]"}, "secure": false, "children": []}
]}`)
if err != nil {
t.Fatalf("Parse: %v", err)
}
element := tree.Find("secure:true")
if element == nil || element.ResourceID != "pwd" {
t.Errorf("secure:true resolved to %+v, want the pwd element", element)
}
}
func TestEditableDerivation(t *testing.T) {
cases := []struct {
name string
@@ -390,6 +531,41 @@ const iosAttrDump = `{
"attributes": {"accessibilityText": "Close", "title": "Settings", "bounds": "[0,0,100,50]"},
"children": [],
"enabled": true
},
{
"attributes": {"identifier": "Feed", "scrollable": "true", "bounds": "[0,120,390,700]"},
"children": [],
"enabled": true
}
]
}`
const classAttrDump = `{
"attributes": {"resource-id": "com.app:id/list", "class": "android.widget.FrameLayout", "bounds": "[0,0,1080,2340]"},
"children": [
{
"attributes": {"resource-id": "com.app:id/row1", "class": "android.widget.Button", "bounds": "[0,0,1080,200]"},
"children": [],
"clickable": true,
"enabled": true
}
]
}`
const webAttrDump = `{
"attributes": {"resource-id": "page", "tag": "html", "bounds": "[0,0,1280,720]"},
"children": [
{
"attributes": {"resource-id": "login_email", "content-desc": "login_email", "tag": "input", "bounds": "[0,0,300,40]"},
"children": [],
"editable": true,
"enabled": true
},
{
"attributes": {"resource-id": "customer_row_a1", "data-testid": "customer-row", "tag": "div", "bounds": "[0,40,300,80]"},
"children": [],
"clickable": true,
"enabled": true
}
]
}`
@@ -410,6 +586,135 @@ func TestLabelAliasMatchesAccessibilityText(t *testing.T) {
}
}
// className is a spec key no producer writes: android reports the view class,
// ios the element type and the chrome dump el.className, all under `class`. The
// key matched nothing at all here while the web runtime resolved it against the
// live DOM, and the key being accepted meant no unknown-key error said so.
func TestClassNameAliasMatchesClass(t *testing.T) {
tree, _ := Parse(classAttrDump)
el := tree.Find("className:Button")
if el == nil {
t.Fatal("expected className: to match the class attribute via alias")
}
if el.ResourceID != "com.app:id/row1" {
t.Fatalf("got %q, want row1", el.ResourceID)
}
object := tree.FindBySelector(Selector{Filters: []AttrFilter{
{Attr: "className", Value: "Button"},
}})
if object == nil {
t.Fatal("expected the object form to match the class attribute via alias")
}
if object.ResourceID != el.ResourceID {
t.Fatalf("the object form matched %q, want %q", object.ResourceID, el.ResourceID)
}
}
// One fact, four names, and only two of them reached it. Android and the chrome
// dump write the accessible name under content-desc; ios writes it under
// accessibilityText. label and accessibilityLabel aliased onto accessibilityText
// alone, and alias expansion is ONE level, so the hop from there to content-desc
// was never taken: both keys matched nothing on the two platforms that write
// content-desc. ariaLabel and contentDescription aliased onto nothing at all and
// matched nothing anywhere. The web runtime resolves all four against the live
// DOM, so a selector naming a field this way found it on one host and no element
// at all on the other, with no unknown-key error to say so.
func TestAccessibilityLabelAliasesReachContentDesc(t *testing.T) {
tree, _ := Parse(webAttrDump)
for _, key := range []string{"label", "accessibilityLabel", "ariaLabel", "contentDescription"} {
element := tree.Find(key + ":login_email")
if element == nil {
t.Fatalf("expected %s: to match the content-desc attribute via alias", key)
}
if element.ResourceID != "login_email" {
t.Fatalf("%s: matched %q, want login_email", key, element.ResourceID)
}
object := tree.FindBySelector(Selector{Filters: []AttrFilter{
{Attr: key, Value: "login_email"},
}})
if object == nil {
t.Fatalf("expected the object form of %s to match content-desc via alias", key)
}
if object.ResourceID != element.ResourceID {
t.Fatalf("the object form of %s matched %q, want %q",
key, object.ResourceID, element.ResourceID)
}
}
}
// The iOS sidecar writes the same fact under accessibilityText, which the two
// names already reached and have to keep reaching.
func TestAccessibilityLabelAliasesStillReachAccessibilityText(t *testing.T) {
tree, _ := Parse(iosAttrDump)
for _, key := range []string{"label", "accessibilityLabel", "ariaLabel", "contentDescription"} {
if tree.Find(key+":Close") == nil {
t.Fatalf("expected %s: to match the accessibilityText attribute via alias", key)
}
}
}
// Compose for Web writes a test tag as data-testid, the key the web runtime
// resolves testTag and testID against. testTag reached the three identifier
// keys and not that one, and testID aliased onto nothing at all, so a tag the
// web runtime found on both rows named no element here.
func TestTestTagAliasesReachDataTestID(t *testing.T) {
tree, _ := Parse(webAttrDump)
for _, key := range []string{"testTag", "testID"} {
element := tree.Find(key + ":customer-row")
if element == nil {
t.Fatalf("expected %s: to match the data-testid attribute via alias", key)
}
if element.ResourceID != "customer_row_a1" {
t.Fatalf("%s: matched %q, want customer_row_a1", key, element.ResourceID)
}
object := tree.FindBySelector(Selector{Filters: []AttrFilter{
{Attr: key, Value: "customer-row"},
}})
if object == nil {
t.Fatalf("expected the object form of %s to match data-testid via alias", key)
}
if object.ResourceID != element.ResourceID {
t.Fatalf("the object form of %s matched %q, want %q",
key, object.ResourceID, element.ResourceID)
}
}
}
// testTag keeps reaching the identifier keys android and ios write it under.
func TestTestTagStillReachesTheIdentifierKeys(t *testing.T) {
if tree, _ := Parse(androidAttrDump); tree.Find("testTag:row1") == nil {
t.Fatal("expected testTag: to match the resource-id attribute via alias")
}
if tree, _ := Parse(iosAttrDump); tree.Find("testTag:Feed") == nil {
t.Fatal("expected testTag: to match the identifier attribute via alias")
}
}
// bounds is a raw driver attribute rather than a cross-platform key: every dump
// writes the rectangle out as a string and no DOM element carries an attribute
// of that name, so the key resolved here and matched nothing on web on every
// page there is, with no unknown-key error to say so and no mapping to invent
// for it. Off the accepted list the web runtime raises that error, and the
// escape hatch for an attribute the tree carries is what keeps it resolving
// where a producer writes it.
func TestBoundsIsAReachableRawAttributeAndNotAnAcceptedKey(t *testing.T) {
if slices.Contains(SelectorKeys(), "bounds") {
t.Error("bounds names no fact a DOM carries, so it cannot be a cross-platform key")
}
tree, _ := Parse(androidAttrDump)
selector := Selector{Filters: []AttrFilter{{Attr: "bounds", Value: "[0,0,1080,200]"}}}
if unknown := tree.UnknownSelectorKeys(selector); len(unknown) != 0 {
t.Errorf("bounds is an attribute this dump carries, got unknown %v", unknown)
}
node := tree.FindBySelector(selector)
if node == nil {
t.Fatal("expected bounds to match the raw attribute the dump writes")
}
if node.ResourceID != "com.app:id/row1" {
t.Fatalf("bounds matched %q, want row1", node.ResourceID)
}
}
func TestContentDescAliasOnIOS(t *testing.T) {
tree, _ := Parse(iosAttrDump)
el := tree.Find("content-desc:Close")
@@ -456,11 +761,14 @@ func TestTitleReturnsNilForAndroid(t *testing.T) {
}
}
func TestScrollableGracefulIgnoreOnIOS(t *testing.T) {
func TestScrollableMatchesOnIOS(t *testing.T) {
tree, _ := Parse(iosAttrDump)
el := tree.Find("scrollable:true")
if el != nil {
t.Fatal("expected scrollable:true to return nil on iOS hierarchy (graceful ignore)")
if el == nil {
t.Fatal("expected scrollable:true to match the iOS scroll container")
}
if el.Attributes["identifier"] != "Feed" {
t.Fatalf("got %q, want Feed", el.Attributes["identifier"])
}
}
@@ -475,6 +783,113 @@ func TestTextIsNowSubstring(t *testing.T) {
}
}
// subtreeTextDump is the shape web and iOS report: an element's text is its
// whole subtree's text, so a badge's ancestors carry the badge's words.
// split_row is the ancestor whose own text carries the value where no
// descendant of it does; nested_row is the one whose badge carries it too.
const subtreeTextDump = `{
"attributes": {"resource-id": "page", "text": "Sent Sent here Sent Sent", "bounds": "[0,0,1080,2340]"},
"children": [
{
"attributes": {"resource-id": "status_row", "text": "Sent", "bounds": "[0,0,1080,100]"},
"children": [
{"attributes": {"resource-id": "status_badge", "text": "Sent", "bounds": "[0,0,200,100]"}, "children": []}
]
},
{
"attributes": {"resource-id": "split_row", "text": "Sent here", "bounds": "[0,100,1080,200]"},
"children": [
{"attributes": {"resource-id": "split_tail", "text": "t here", "bounds": "[0,100,200,200]"}, "children": []}
]
},
{
"attributes": {"resource-id": "nested_row", "text": "Sent Sent", "bounds": "[0,200,1080,300]"},
"children": [
{"attributes": {"resource-id": "nested_badge", "text": "Sent", "bounds": "[0,200,200,300]"}, "children": []}
]
}
]
}`
func TestTextNamesTheInnermostMatch(t *testing.T) {
tree, _ := Parse(subtreeTextDump)
want := []string{"status_badge", "split_row", "nested_badge"}
if got := resourceIDsOf(tree.FindAllNodes("text:Sent")); !slices.Equal(
got,
want,
) {
t.Errorf("text:Sent matched %v, want %v", got, want)
}
sel := Selector{Filters: []AttrFilter{{Attr: "text", Value: "Sent"}}}
if got := resourceIDsOf(tree.FindAllBySelector(sel)); !slices.Equal(
got,
want,
) {
t.Errorf("{text: Sent} matched %v, want %v", got, want)
}
node := tree.FindNode("text:Sent")
if node == nil || node.ResourceID != "status_badge" {
t.Errorf("find named %v, want the deepest match status_badge", node)
}
}
func TestScopedTextNamesTheInnermostMatch(t *testing.T) {
tree, _ := Parse(subtreeTextDump)
want := []string{"status_badge", "split_row", "nested_badge"}
if got := resourceIDsOf(tree.Root.FindAll("text:Sent")); !slices.Equal(
got,
want,
) {
t.Errorf("scoped text:Sent matched %v, want %v", got, want)
}
sel := Selector{Filters: []AttrFilter{{Attr: "text", Value: "Sent"}}}
if got := resourceIDsOf(tree.Root.FindAllBySelector(sel)); !slices.Equal(
got,
want,
) {
t.Errorf("scoped {text: Sent} matched %v, want %v", got, want)
}
}
// The root answers a selector whichever form the selector is written in: the
// string form scans the tree from the root down, and the object form used to
// start at the root's children and lose it.
func TestRootMatchesInBothSelectorForms(t *testing.T) {
tree, _ := Parse(subtreeTextDump)
sel := Selector{Filters: []AttrFilter{{Attr: "id", Value: "page"}}}
want := []string{"page"}
if got := resourceIDsOf(tree.FindAllNodes("id:page")); !slices.Equal(
got,
want,
) {
t.Errorf("id:page matched %v, want %v", got, want)
}
if got := resourceIDsOf(tree.FindAllBySelector(sel)); !slices.Equal(
got,
want,
) {
t.Errorf("{id: page} matched %v, want %v", got, want)
}
if got := resourceIDsOf(tree.FindAllBySelectorPath([]Selector{sel})); !slices.Equal(
got,
want,
) {
t.Errorf("[{id: page}] matched %v, want %v", got, want)
}
if node := tree.FindBySelector(sel); node == nil ||
node.ResourceID != "page" {
t.Errorf("find({id: page}) named %v, want page", node)
}
}
func resourceIDsOf(nodes []*Node) []string {
var ids []string
for _, node := range nodes {
ids = append(ids, node.ResourceID)
}
return ids
}
func TestMultiFilterSelectorAND(t *testing.T) {
tree, _ := Parse(androidAttrDump)
sel := Selector{Filters: []AttrFilter{
@@ -1036,6 +1451,233 @@ func TestTreeTransitional(t *testing.T) {
}
}
// A key means the same thing whichever form the author writes it in. The object
// form used to fall through to the raw attribute map, which carries neither
// "id" nor "desc" on any platform.
func TestObjectSelectorIDMatchesTheSameElementsAsTheStringForm(t *testing.T) {
tree, _ := Parse(sampleDump)
sel := Selector{Filters: []AttrFilter{{Attr: "id", Value: "row"}}}
object := tree.FindAllBySelector(sel)
if len(object) != len(tree.FindAll("id:row")) {
t.Fatalf("object form matched %d, string form %d", len(object), len(tree.FindAll("id:row")))
}
if len(object) != 2 {
t.Fatalf("want 2 rows, got %d", len(object))
}
}
func TestObjectSelectorDescMatchesTheSameElementsAsTheStringForm(t *testing.T) {
tree, _ := Parse(sampleDump)
sel := Selector{Filters: []AttrFilter{{Attr: "desc", Value: "row"}}}
if len(tree.FindAllBySelector(sel)) != len(tree.FindAll("desc:row")) {
t.Fatal("object and string form disagree on desc")
}
}
func TestUnknownSelectorKeyIsReported(t *testing.T) {
tree, _ := Parse(sampleDump)
sel := Selector{Filters: []AttrFilter{{Attr: "descripton", Value: "row"}}}
unknown := tree.UnknownSelectorKeys(sel)
if len(unknown) != 1 || unknown[0] != "descripton" {
t.Fatalf("got %v, want [descripton]", unknown)
}
message := UnknownSelectorKeyMessage(unknown)
if !strings.Contains(message, `"descripton"`) || !strings.Contains(message, "resource-id") {
t.Fatalf("message names neither the key nor the accepted list: %s", message)
}
}
func TestAcceptedSelectorKeyAbsentFromTheScreenIsNotUnknown(t *testing.T) {
tree, _ := Parse(sampleDump)
sel := Selector{Filters: []AttrFilter{{Attr: "title", Value: "Settings"}}}
if unknown := tree.UnknownSelectorKeys(sel); len(unknown) != 0 {
t.Fatalf("a platform-specific key must stay silent, got %v", unknown)
}
}
// Raw driver attributes stay reachable: a key some element carries can match,
// whether or not this package enumerates it.
func TestRawDriverAttributeIsNotUnknown(t *testing.T) {
input := `{
"attributes": {"resource-id": "root", "important-for-accessibility": "true"},
"children": []
}`
tree, _ := Parse(input)
sel := Selector{Filters: []AttrFilter{{Attr: "important-for-accessibility", Value: "true"}}}
if unknown := tree.UnknownSelectorKeys(sel); len(unknown) != 0 {
t.Fatalf("got %v, want none", unknown)
}
}
// TestStoredTreeRebuildsRootAndResolvesSelectors is the offline half of every
// trace: a tree that only survives as a flat pre-order array resolves nothing,
// because every lookup walks Root.
func TestStoredTreeRebuildsRootAndResolvesSelectors(t *testing.T) {
tree, err := Parse(sampleDump)
if err != nil {
t.Fatal(err)
}
stored, err := json.Marshal(tree)
if err != nil {
t.Fatal(err)
}
var decoded Tree
if err := json.Unmarshal(stored, &decoded); err != nil {
t.Fatal(err)
}
if decoded.Root == nil {
t.Fatal("Root not rebuilt from the stored tree")
}
if len(decoded.Elements) != len(tree.Elements) {
t.Fatalf(
"elements = %d, want %d",
len(decoded.Elements),
len(tree.Elements),
)
}
for index, element := range decoded.Elements {
if element.ResourceID != tree.Elements[index].ResourceID {
t.Fatalf(
"element %d = %q, want %q",
index,
element.ResourceID,
tree.Elements[index].ResourceID,
)
}
}
online := tree.Find("id:app:id/title")
offline := decoded.Find("id:app:id/title")
if online == nil {
t.Fatal("the selector does not resolve online; the fixture is wrong")
}
if offline == nil || offline.Text != online.Text {
t.Fatalf(
"selector resolves online to %+v, offline to %+v",
online,
offline,
)
}
if decoded.Find("id:app:id/title") != decoded.Elements[1] {
t.Error(
"the rebuilt nodes and the decoded element list are different pointers",
)
}
childCount := 0
var walk func(node *Node)
walk = func(node *Node) {
childCount++
for _, child := range node.Children {
walk(child)
}
}
walk(decoded.Root)
if childCount != len(decoded.Elements) {
t.Errorf(
"the rebuilt tree holds %d nodes, the element list %d",
childCount,
len(decoded.Elements),
)
}
}
// TestStoredTreeWithoutDepthsKeepsTheOldShape: traces written before the
// depths field must still load, and must say "no structure" rather than
// inventing one.
func TestStoredTreeWithoutDepthsKeepsTheOldShape(t *testing.T) {
var decoded Tree
if err := json.Unmarshal([]byte(`{"elements":[{"resourceId":"root"},{"resourceId":"child"}]}`), &decoded); err != nil {
t.Fatal(err)
}
if len(decoded.Elements) != 2 {
t.Fatalf("elements = %d, want 2", len(decoded.Elements))
}
if decoded.Root != nil {
t.Errorf(
"Root = %+v, want nil for a stored tree that carries no depths",
decoded.Root,
)
}
}
func TestStoredTreeRejectsDepthsItCannotRebuild(t *testing.T) {
for name, stored := range map[string]string{
"count mismatch": `{"elements":[{"resourceId":"a"},{"resourceId":"b"}],"depths":[0]}`,
"rootless": `{"elements":[{"resourceId":"a"}],"depths":[1]}`,
"second root": `{"elements":[{"resourceId":"a"},{"resourceId":"b"}],"depths":[0,0]}`,
"skipped level": `{"elements":[{"resourceId":"a"},{"resourceId":"b"}],"depths":[0,2]}`,
} {
var decoded Tree
if err := json.Unmarshal([]byte(stored), &decoded); err != nil {
t.Fatalf("%s: %v", name, err)
}
if decoded.Root != nil {
t.Errorf("%s: Root = %+v, want nil", name, decoded.Root)
}
}
}
// TestTreeStoresNoDepthsForAnUnwalkableRoot keeps the stored shape honest:
// a hand-built Tree whose Root and Elements disagree must not claim a
// structure that would rebuild into a different tree.
func TestTreeStoresNoDepthsForAnUnwalkableRoot(t *testing.T) {
tree := &Tree{
Root: &Node{Element: Element{ResourceID: "root"}},
Elements: []*Element{{ResourceID: "root"}, {ResourceID: "orphan"}},
}
stored, err := json.Marshal(tree)
if err != nil {
t.Fatal(err)
}
if strings.Contains(string(stored), "depths") {
t.Errorf("stored a shape that cannot be rebuilt: %s", stored)
}
}
// A producer that puts a string where a flag belongs must cost that flag and
// nothing else. Failing the document instead blanks the whole tree, and every
// extractor then reads a screen with no elements on it.
func TestParseUnreadableFlagKeepsTheRestOfTheTree(t *testing.T) {
input := `{"attributes":{"resource-id":"app:id/root","bounds":"[0,0,400,800]"},"children":[
{"attributes":{"resource-id":"app:id/tabs","text":"keep"},"selected":"active","children":[]},
{"attributes":{"resource-id":"app:id/leaf"},"clickable":true,"children":[]}
]}`
tree, err := Parse(input)
if err != nil {
t.Fatalf("Parse: %v", err)
}
if len(tree.Elements) != 3 {
t.Fatalf("elements = %d, want 3: one unreadable flag must not cost the tree", len(tree.Elements))
}
tabs := tree.Find("id:tabs")
if tabs == nil {
t.Fatal("the element carrying the unreadable flag was dropped")
}
if tabs.Text != "keep" {
t.Errorf("neighbouring field corrupted: text=%q", tabs.Text)
}
if tabs.Selected {
t.Error("selected must stay unset when the value the producer sent is not a boolean")
}
leaf := tree.Find("id:leaf")
if leaf == nil || !leaf.Clickable {
t.Errorf("a readable flag elsewhere in the tree must survive, got %+v", leaf)
}
if tree.UnreadableFlags != 1 {
t.Errorf("UnreadableFlags = %d, want 1: a dropped flag has to be countable", tree.UnreadableFlags)
}
stored, err := json.Marshal(tree)
if err != nil {
t.Fatal(err)
}
var decoded Tree
if err := json.Unmarshal(stored, &decoded); err != nil {
t.Fatal(err)
}
if decoded.UnreadableFlags != 1 {
t.Errorf("stored UnreadableFlags = %d, want 1: the trace has to carry it", decoded.UnreadableFlags)
}
}
// selectorFormsDump carries one node per id shape a real dump produces, plus
// nodes carrying a description in the ", " form the desc rule knows about and a
// text the text rule matches on a substring.
@@ -1115,3 +1757,67 @@ func TestSelectorFormsResolveTheSameElement(t *testing.T) {
})
}
}
// customElementDump is the shape a page built from custom elements reports: a
// container's tag name contains the tag name of what it holds, so "todo-list"
// carries "li" and "todo-app" carries "a".
const customElementDump = `{
"attributes": {"tag": "todo-app", "resource-id": "app", "bounds": "[0,0,800,600]"},
"children": [
{
"attributes": {"tag": "todo-list", "resource-id": "list", "bounds": "[0,0,800,400]"},
"children": [
{"attributes": {"tag": "li", "resource-id": "todo_1", "bounds": "[0,0,800,50]"}, "children": []},
{"attributes": {"tag": "li", "resource-id": "todo_2", "bounds": "[0,50,800,100]"}, "children": []}
]
},
{
"attributes": {"tag": "a", "resource-id": "filter_all", "bounds": "[0,400,800,450]"},
"children": []
}
]
}`
func TestTagNamesTheWholeTagNotASubstringOfIt(t *testing.T) {
tree, _ := Parse(customElementDump)
for _, test := range []struct {
value string
want []string
}{
{"li", []string{"todo_1", "todo_2"}},
{"a", []string{"filter_all"}},
} {
t.Run(test.value, func(t *testing.T) {
sel := Selector{Filters: []AttrFilter{{Attr: "tag", Value: test.value}}}
if got := resourceIDsOf(tree.FindAllBySelector(sel)); !slices.Equal(got, test.want) {
t.Errorf("{tag: %q} matched %v, want %v", test.value, got, test.want)
}
stringForm := "tag:" + test.value
if got := resourceIDsOf(tree.FindAllNodes(stringForm)); !slices.Equal(got, test.want) {
t.Errorf("%q matched %v, want %v", stringForm, got, test.want)
}
})
}
}
func TestTagMatchesTheElementItNames(t *testing.T) {
tree, _ := Parse(customElementDump)
for _, test := range []struct {
value string
want []string
}{
{"todo-app", []string{"app"}},
{"todo-list", []string{"list"}},
} {
t.Run(test.value, func(t *testing.T) {
sel := Selector{Filters: []AttrFilter{{Attr: "tag", Value: test.value}}}
if got := resourceIDsOf(tree.FindAllBySelector(sel)); !slices.Equal(got, test.want) {
t.Errorf("{tag: %q} matched %v, want %v", test.value, got, test.want)
}
stringForm := "tag:" + test.value
if got := resourceIDsOf(tree.FindAllNodes(stringForm)); !slices.Equal(got, test.want) {
t.Errorf("%q matched %v, want %v", stringForm, got, test.want)
}
})
}
}
+63
View File
@@ -0,0 +1,63 @@
package hierarchy
import (
"encoding/json"
"os"
"path/filepath"
"slices"
"testing"
)
// selectorKeysGolden is the cross-runtime contract for object selectors: the
// keys both runtimes accept, and the diagnostic both raise for a key neither
// can match.
type selectorKeysGolden struct {
Keys []string `json:"keys"`
UnknownKeyExample []string `json:"unknownKeyExample"`
UnknownKeyMessage string `json:"unknownKeyMessage"`
}
// Both runtimes reject an object-selector key they do not know, so the two key
// lists have to be one list. Were they to drift, a spec would be accepted by
// the runtime that lists the key and fail the run on the one that does not, and
// the difference would only show on the platform nobody ran first.
// pkg/spec/test/selector-keys.test.ts asserts the SAME file from the web side.
func TestSelectorKeysMatchTheCrossRuntimeList(t *testing.T) {
golden := loadSelectorKeysGolden(t)
if got := SelectorKeys(); !slices.Equal(got, golden.Keys) {
t.Errorf("native key list\n got=%v\nwant=%v", got, golden.Keys)
}
}
// An author who hits this on Android and again on web must read one sentence,
// not two dialects of it.
func TestUnknownSelectorKeyMessageMatchesTheCrossRuntimeText(t *testing.T) {
golden := loadSelectorKeysGolden(t)
if got := UnknownSelectorKeyMessage(golden.UnknownKeyExample); got != golden.UnknownKeyMessage {
t.Errorf("native message\n got=%q\nwant=%q", got, golden.UnknownKeyMessage)
}
}
func TestSelectorKeysAreSorted(t *testing.T) {
keys := SelectorKeys()
if !slices.IsSorted(keys) {
t.Errorf("keys must stay sorted so the two lists compare readably: %v", keys)
}
}
func loadSelectorKeysGolden(t *testing.T) selectorKeysGolden {
t.Helper()
path, err := filepath.Abs("../../pkg/spec/test/fixtures/selector-keys.json")
if err != nil {
t.Fatal(err)
}
body, err := os.ReadFile(path)
if err != nil {
t.Fatalf("read selector keys: %v", err)
}
var golden selectorKeysGolden
if err := json.Unmarshal(body, &golden); err != nil {
t.Fatalf("decode selector keys: %v", err)
}
return golden
}
+14
View File
@@ -119,7 +119,21 @@ type JSONSchema struct {
// Response is the slice of a chat-completions response we read.
type Response struct {
// Model is the model the provider actually served. A router can satisfy one
// requested id with a differently-priced variant, so cost accounting reads
// this rather than the requested id.
Model string `json:"model"`
Choices []Choice `json:"choices"`
Usage Usage `json:"usage"`
}
// Usage is the provider's token accounting for one call, present on every
// non-streaming OpenAI-compatible chat completion. Zero values mean the
// provider omitted the object.
type Usage struct {
PromptTokens int `json:"prompt_tokens"`
CompletionTokens int `json:"completion_tokens"`
TotalTokens int `json:"total_tokens"`
}
// Choice is one completion choice.
+55
View File
@@ -106,6 +106,61 @@ func TestChatCompletionRequestShapeAndParse(t *testing.T) {
}
}
// TestChatCompletionParsesUsageAndServedModel pins the accounting fields: cost
// per defect and tokens per action are computed from them, and a router can
// serve a request with a differently-priced model than the one asked for.
func TestChatCompletionParsesUsageAndServedModel(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write([]byte(`{
"model": "vendor/model-2026-05",
"choices": [{"message": {"content": "{}"}}],
"usage": {"prompt_tokens": 1200, "completion_tokens": 34, "total_tokens": 1234}
}`))
}))
defer server.Close()
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", server.URL)
client, err := New()
if err != nil {
t.Fatalf("New: %v", err)
}
response, err := client.ChatCompletion(context.Background(), Request{Model: "vendor/model"})
if err != nil {
t.Fatalf("ChatCompletion: %v", err)
}
if response.Model != "vendor/model-2026-05" {
t.Errorf("served model = %q, want vendor/model-2026-05", response.Model)
}
want := Usage{PromptTokens: 1200, CompletionTokens: 34, TotalTokens: 1234}
if response.Usage != want {
t.Errorf("usage = %+v, want %+v", response.Usage, want)
}
}
// TestChatCompletionToleratesMissingUsage keeps a provider that omits the usage
// object from failing the call; the record simply carries zero tokens.
func TestChatCompletionToleratesMissingUsage(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write([]byte(`{"choices":[{"message":{"content":"{}"}}]}`))
}))
defer server.Close()
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", server.URL)
client, err := New()
if err != nil {
t.Fatalf("New: %v", err)
}
response, err := client.ChatCompletion(context.Background(), Request{Model: "m"})
if err != nil {
t.Fatalf("ChatCompletion: %v", err)
}
if (response.Usage != Usage{}) {
t.Errorf("usage = %+v, want the zero value when the provider omits it", response.Usage)
}
}
func TestNewRequiresAPIKey(t *testing.T) {
t.Setenv("OPENROUTER_API_KEY", "")
t.Setenv("OPENAI_API_KEY", "")
+31 -19
View File
@@ -37,6 +37,11 @@ type Evaluator struct {
violated bool
steps int
violation *Violation
// observations counts the states this evaluator actually reduced, which is
// what a `within(n, "steps")` window is measured in. It differs from steps
// whenever the caller's numbering skipped an observation, and the two are
// told apart in the serialized AST by expiresAtObservation.
observations int
// oneShot marks a root that is armed once at the first observation rather
// than re-asserted at every one; armed records that it has been.
oneShot bool
@@ -93,6 +98,7 @@ func (e *Evaluator) ObserveAtStep(now time.Time, step int) Verdict {
return VerdictViolated
}
e.steps = step
e.observations++
obligations := make([]obligation, 0, len(e.pending)+1)
obligations = append(obligations, e.pending...)
@@ -102,7 +108,7 @@ func (e *Evaluator) ObserveAtStep(now time.Time, step int) Verdict {
e.pending = e.pending[:0]
for _, entry := range obligations {
result := reduce(entry.formula, now)
result := reduce(entry.formula, now, e.observations)
switch result.status {
case statusHolds:
// drop
@@ -374,7 +380,11 @@ func pending(f Formula) reduceResult {
return reduceResult{status: statusPending, formula: f}
}
func reduce(formula Formula, now time.Time) reduceResult {
// reduce advances one obligation against the current state. `now` is the
// observation's wall clock and `observation` its index in the sequence of
// states this evaluator reduced; the two are the clocks a duration-bounded and
// a step-bounded window are resolved against.
func reduce(formula Formula, now time.Time, observation int) reduceResult {
switch concrete := formula.(type) {
case PureFormula:
if concrete.Value {
@@ -399,7 +409,7 @@ func reduce(formula Formula, now time.Time) reduceResult {
return violatedWith(concrete, "predicate false")
case NowFormula:
return reduce(concrete.Inner, now)
return reduce(concrete.Inner, now, observation)
case NextFormula:
// Next defers the inner obligation to the following step without
@@ -414,23 +424,24 @@ func reduce(formula Formula, now time.Time) reduceResult {
concrete.Deadline = now.Add(concrete.Duration)
concrete.HasDeadline = true
}
innerResult := reduce(concrete.Inner, now)
if concrete.HasStepBound && !concrete.HasExpiryObservation {
concrete.ExpiryObservation = observation + concrete.StepBound - 1
concrete.HasExpiryObservation = true
}
innerResult := reduce(concrete.Inner, now, observation)
if innerResult.status == statusHolds {
return holds()
}
// The window is measured in observations at which the inner could have
// discharged, so an inner that is merely pending has not discharged and
// the window closing on it is a violation.
if concrete.HasStepBound && concrete.StepBound <= 1 {
if concrete.HasExpiryObservation && observation >= concrete.ExpiryObservation {
return violatedFrom(innerResult, concrete, "eventually bound exhausted")
}
if concrete.HasDeadline && !now.Before(concrete.Deadline) {
return violatedFrom(innerResult, concrete, "eventually deadline reached")
}
next := concrete
if concrete.HasStepBound {
next.StepBound = concrete.StepBound - 1
}
// F(inner) unrolls to inner or X F(inner). A pending inner is a
// deferred way of satisfying the promise, so it is kept as a disjunct
// rather than dropped; dropping it is what made an inner that only
@@ -449,11 +460,11 @@ func reduce(formula Formula, now time.Time) reduceResult {
return reduce(OrFormula{
Left: pushNot(concrete.Antecedent),
Right: nnf(concrete.Consequent),
}, now)
}, now, observation)
case OrFormula:
left := reduce(concrete.Left, now)
right := reduce(concrete.Right, now)
left := reduce(concrete.Left, now, observation)
right := reduce(concrete.Right, now, observation)
if left.status == statusHolds || right.status == statusHolds {
return holds()
}
@@ -469,8 +480,8 @@ func reduce(formula Formula, now time.Time) reduceResult {
return pending(OrFormula{Left: left.formula, Right: right.formula})
case AndFormula:
left := reduce(concrete.Left, now)
right := reduce(concrete.Right, now)
left := reduce(concrete.Left, now, observation)
right := reduce(concrete.Right, now, observation)
if left.status == statusViolated {
return violatedFrom(left, concrete, "conjunct violated")
}
@@ -489,7 +500,7 @@ func reduce(formula Formula, now time.Time) reduceResult {
return pending(AndFormula{Left: left.formula, Right: right.formula})
case NotFormula:
inner := reduce(concrete.Inner, now)
inner := reduce(concrete.Inner, now, observation)
switch inner.status {
case statusHolds:
return violatedWith(concrete, "negated formula held")
@@ -506,7 +517,11 @@ func reduce(formula Formula, now time.Time) reduceResult {
concrete.Deadline = now.Add(concrete.Duration)
concrete.HasDeadline = true
}
innerResult := reduce(concrete.Inner, now)
if concrete.HasStepBound && !concrete.HasExpiryObservation {
concrete.ExpiryObservation = observation + concrete.StepBound - 1
concrete.HasExpiryObservation = true
}
innerResult := reduce(concrete.Inner, now, observation)
if innerResult.status == statusViolated {
return violatedFrom(innerResult, concrete, "always inner violated")
}
@@ -517,16 +532,13 @@ func reduce(formula Formula, now time.Time) reduceResult {
// hold". A pending inner has not been breached inside the window, so it
// discharges vacuously here exactly as its negation violates on the
// Eventually side.
if concrete.HasStepBound && concrete.StepBound <= 1 {
if concrete.HasExpiryObservation && observation >= concrete.ExpiryObservation {
return holds()
}
if concrete.HasDeadline && !now.Before(concrete.Deadline) {
return holds()
}
next := concrete
if concrete.HasStepBound {
next.StepBound = concrete.StepBound - 1
}
if innerResult.status == statusHolds {
return pending(next)
}
+1 -1
View File
@@ -153,5 +153,5 @@ func TestObserve_PanicsOnUnknownFormulaType(t *testing.T) {
t.Errorf("expected panic on unsupported formula type")
}
}()
reduce(unsupportedFormula{}, time.Now())
reduce(unsupportedFormula{}, time.Now(), 1)
}
+87 -36
View File
@@ -39,12 +39,14 @@ func (e ErrorFormula) describe() string {
// window and is vacuously satisfied once the window closes. An unbounded
// Always carries no bound fields and is checked at every observed step.
type AlwaysFormula struct {
Inner Formula
StepBound int
HasStepBound bool
Duration time.Duration
Deadline time.Time
HasDeadline bool
Inner Formula
StepBound int
HasStepBound bool
ExpiryObservation int
HasExpiryObservation bool
Duration time.Duration
Deadline time.Time
HasDeadline bool
}
type PureFormula struct {
@@ -92,13 +94,20 @@ type NextFormula struct {
// resolves the absolute deadline on first reduction using the observation
// time. This matches the "within N seconds of obligation instantiation"
// semantics used by nested Always(Eventually(...).within(...)) formulas.
//
// StepBound is the step-domain counterpart: the window counts observations the
// evaluator reduced, and ExpiryObservation is the absolute closing observation
// the evaluator resolves on first reduction, exactly as Deadline is for
// Duration.
type EventuallyFormula struct {
Inner Formula
StepBound int
HasStepBound bool
Duration time.Duration
Deadline time.Time
HasDeadline bool
Inner Formula
StepBound int
HasStepBound bool
ExpiryObservation int
HasExpiryObservation bool
Duration time.Duration
Deadline time.Time
HasDeadline bool
}
type ImpliesFormula struct {
@@ -180,6 +189,9 @@ func (a AlwaysFormula) describe() string {
if a.HasStepBound {
parts = append(parts, fmt.Sprintf("steps=%d", a.StepBound))
}
if a.HasExpiryObservation {
parts = append(parts, fmt.Sprintf("expiresAtObservation=%d", a.ExpiryObservation))
}
if a.HasDeadline {
parts = append(parts, "deadline="+a.Deadline.Format(time.RFC3339Nano))
} else if a.Duration > 0 {
@@ -198,6 +210,9 @@ func (e EventuallyFormula) describe() string {
if e.HasStepBound {
parts = append(parts, fmt.Sprintf("steps=%d", e.StepBound))
}
if e.HasExpiryObservation {
parts = append(parts, fmt.Sprintf("expiresAtObservation=%d", e.ExpiryObservation))
}
if e.HasDeadline {
parts = append(parts, "deadline="+e.Deadline.Format(time.RFC3339Nano))
} else if e.Duration > 0 {
@@ -230,33 +245,73 @@ type withinNode struct {
// the same duration differ only here, so without it they serialize
// identically and the trace erases the distinction the evaluator makes.
Deadline int64 `json:"deadline,omitempty"`
// ExpiresAtObservation is the step-domain counterpart of Deadline: the
// index of the observation the window closes at. It names observations
// rather than runner steps because a step the verifier skipped never
// reached the evaluator and so cannot close a window; the pair of fields
// is what lets a reader tell the two numberings apart.
ExpiresAtObservation int `json:"expiresAtObservation,omitempty"`
}
// boundWindow is the optional window shared by AlwaysFormula and
// EventuallyFormula: the window the spec authored plus the absolute close the
// evaluator resolved for this obligation.
type boundWindow struct {
hasStepBound bool
stepBound int
hasExpiryObservation bool
expiryObservation int
duration time.Duration
hasDeadline bool
deadline time.Time
}
// withinFor renders the bound clause of a bounded Always or Eventually. The
// authored window (steps or duration) stays in amount/unit so readers keep
// seeing what the spec asked for; the resolved deadline rides alongside.
func withinFor(
hasStepBound bool,
stepBound int,
duration time.Duration,
hasDeadline bool,
deadline time.Time,
) *withinNode {
var node *withinNode
// seeing what the spec asked for; the resolved close rides alongside.
func withinFor(window boundWindow) *withinNode {
switch {
case hasStepBound:
node = &withinNode{Amount: int64(stepBound), Unit: "steps"}
case duration > 0:
node = &withinNode{Amount: duration.Milliseconds(), Unit: "milliseconds"}
case hasDeadline:
return &withinNode{Amount: deadline.UnixMilli(), Unit: "deadline"}
case window.hasStepBound:
node := &withinNode{Amount: int64(window.stepBound), Unit: "steps"}
if window.hasExpiryObservation {
node.ExpiresAtObservation = window.expiryObservation
}
return node
case window.duration > 0:
node := &withinNode{Amount: window.duration.Milliseconds(), Unit: "milliseconds"}
if window.hasDeadline {
node.Deadline = window.deadline.UnixMilli()
}
return node
case window.hasDeadline:
return &withinNode{Amount: window.deadline.UnixMilli(), Unit: "deadline"}
default:
return nil
}
if hasDeadline {
node.Deadline = deadline.UnixMilli()
}
func (a AlwaysFormula) boundWindow() boundWindow {
return boundWindow{
hasStepBound: a.HasStepBound,
stepBound: a.StepBound,
hasExpiryObservation: a.HasExpiryObservation,
expiryObservation: a.ExpiryObservation,
duration: a.Duration,
hasDeadline: a.HasDeadline,
deadline: a.Deadline,
}
}
func (e EventuallyFormula) boundWindow() boundWindow {
return boundWindow{
hasStepBound: e.HasStepBound,
stepBound: e.StepBound,
hasExpiryObservation: e.HasExpiryObservation,
expiryObservation: e.ExpiryObservation,
duration: e.Duration,
hasDeadline: e.HasDeadline,
deadline: e.Deadline,
}
return node
}
func (a AlwaysFormula) MarshalJSON() ([]byte, error) {
@@ -265,9 +320,7 @@ func (a AlwaysFormula) MarshalJSON() ([]byte, error) {
Arg Formula `json:"arg"`
Within *withinNode `json:"within,omitempty"`
}{Op: "always", Arg: a.Inner}
payload.Within = withinFor(
a.HasStepBound, a.StepBound, a.Duration, a.HasDeadline, a.Deadline,
)
payload.Within = withinFor(a.boundWindow())
return json.Marshal(payload)
}
@@ -298,9 +351,7 @@ func (e EventuallyFormula) MarshalJSON() ([]byte, error) {
Arg Formula `json:"arg"`
Within *withinNode `json:"within,omitempty"`
}{Op: "eventually", Arg: e.Inner}
payload.Within = withinFor(
e.HasStepBound, e.StepBound, e.Duration, e.HasDeadline, e.Deadline,
)
payload.Within = withinFor(e.boundWindow())
return json.Marshal(payload)
}
+16 -12
View File
@@ -65,21 +65,25 @@ func pushNot(formula Formula) Formula {
return NextFormula{Inner: pushNot(concrete.Inner)}
case AlwaysFormula:
return EventuallyFormula{
Inner: pushNot(concrete.Inner),
StepBound: concrete.StepBound,
HasStepBound: concrete.HasStepBound,
Duration: concrete.Duration,
Deadline: concrete.Deadline,
HasDeadline: concrete.HasDeadline,
Inner: pushNot(concrete.Inner),
StepBound: concrete.StepBound,
HasStepBound: concrete.HasStepBound,
ExpiryObservation: concrete.ExpiryObservation,
HasExpiryObservation: concrete.HasExpiryObservation,
Duration: concrete.Duration,
Deadline: concrete.Deadline,
HasDeadline: concrete.HasDeadline,
}
case EventuallyFormula:
return AlwaysFormula{
Inner: pushNot(concrete.Inner),
StepBound: concrete.StepBound,
HasStepBound: concrete.HasStepBound,
Duration: concrete.Duration,
Deadline: concrete.Deadline,
HasDeadline: concrete.HasDeadline,
Inner: pushNot(concrete.Inner),
StepBound: concrete.StepBound,
HasStepBound: concrete.HasStepBound,
ExpiryObservation: concrete.ExpiryObservation,
HasExpiryObservation: concrete.HasExpiryObservation,
Duration: concrete.Duration,
Deadline: concrete.Deadline,
HasDeadline: concrete.HasDeadline,
}
default:
return NotFormula{Inner: formula}
+195
View File
@@ -0,0 +1,195 @@
package ltl
import (
"encoding/json"
"strings"
"testing"
"time"
)
func alwaysFalse() func() (bool, error) {
return func() (bool, error) { return false, nil }
}
// A step-bounded window counts the observations the evaluator reduced, and the
// residual has to keep saying which window the spec authored rather than the
// part of it that is left. Bug class: the replay UI renders "within N steps"
// straight off the residual, so a shrinking N tells the reader the spec asked
// for a window it never asked for.
func TestStepBoundedEventually_ResidualKeepsAuthoredWindow(t *testing.T) {
evaluator := NewEvaluator(EventuallyWithinSteps(ThunkNamed("p", alwaysFalse()), 5))
for index := range 3 {
if got := evaluator.ObserveAt(time.Unix(int64(index), 0)); got != VerdictPending {
t.Fatalf("observation %d: got %v, want pending", index+1, got)
}
}
body, err := json.Marshal(evaluator.Residual())
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(body), `"unit":"steps"`) || !strings.Contains(string(body), `"amount":5`) {
t.Errorf("authored window lost after reduction: %s", body)
}
if !strings.Contains(string(body), `"expiresAtObservation":5`) {
t.Errorf("resolved expiry missing: %s", body)
}
}
// Two obligations spawned at different observations from one `within(n,
// "steps")` window close at different observations, and the serialized AST has
// to keep them apart the way a resolved deadline keeps two duration-bounded
// ones apart. Bug class: the trace shows one node where the evaluator holds
// several distinct obligations.
func TestStepBoundedEventually_ObligationsSerializeApart(t *testing.T) {
evaluator := NewEvaluator(Always(EventuallyWithinSteps(ThunkNamed("p", alwaysFalse()), 3)))
for index := range 2 {
if got := evaluator.ObserveAt(time.Unix(int64(index), 0)); got != VerdictPending {
t.Fatalf("observation %d: got %v, want pending", index+1, got)
}
}
body, err := json.Marshal(evaluator.Residual())
if err != nil {
t.Fatal(err)
}
text := string(body)
if strings.Count(text, `"amount":3`) != 2 {
t.Errorf("both obligations should report the authored window of 3: %s", text)
}
if !strings.Contains(text, `"expiresAtObservation":3`) || !strings.Contains(text, `"expiresAtObservation":4`) {
t.Errorf("obligations armed at different observations share a closing observation: %s", text)
}
}
// A bounded Always is the dual of a bounded Eventually, so its window resolves
// and serializes the same way.
func TestStepBoundedAlways_ResidualKeepsAuthoredWindow(t *testing.T) {
formula := AlwaysFormula{
Inner: ThunkNamed("p", func() (bool, error) { return true, nil }),
StepBound: 4,
HasStepBound: true,
}
evaluator := NewEvaluator(formula)
for index := range 2 {
if got := evaluator.ObserveAt(time.Unix(int64(index), 0)); got != VerdictPending {
t.Fatalf("observation %d: got %v, want pending", index+1, got)
}
}
body, err := json.Marshal(evaluator.Residual())
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(body), `"amount":4`) {
t.Errorf("authored window lost after reduction: %s", body)
}
if !strings.Contains(string(body), `"expiresAtObservation":4`) {
t.Errorf("resolved expiry missing: %s", body)
}
}
// A step the verifier skipped (a transitional tree, an empty hierarchy) never
// reached the evaluator, so the property was given no chance to discharge
// there and the window must not charge for it. The runner's step numbering
// only labels the witness; it does not drive the window.
func TestStepBoundedEventually_SkippedRunnerStepsDoNotConsumeWindow(t *testing.T) {
observed := 0
inner := ThunkNamed("p", func() (bool, error) {
observed++
return observed == 3, nil
})
evaluator := NewEvaluator(EventuallyWithinSteps(inner, 3))
var verdict Verdict
for _, runnerStep := range []int{1, 7, 19} {
verdict = evaluator.ObserveAtStep(time.Unix(int64(runnerStep), 0), runnerStep)
}
if verdict != VerdictHolds {
t.Errorf("three observations inside a three-observation window: got %v, want holds", verdict)
}
}
// The witness still carries the runner's numbering, so a report names the step
// that armed the obligation even though the window counted observations.
func TestStepBoundedEventually_WitnessCarriesRunnerStep(t *testing.T) {
evaluator := NewEvaluator(EventuallyWithinSteps(ThunkNamed("p", alwaysFalse()), 2))
for _, runnerStep := range []int{4, 11} {
evaluator.ObserveAtStep(time.Unix(int64(runnerStep), 0), runnerStep)
}
witness := evaluator.Violation()
if witness == nil {
t.Fatal("no violation recorded")
}
if witness.Step != 4 {
t.Errorf("witness Step = %d, want the runner step that armed the obligation (4)", witness.Step)
}
}
// An undischarged bounded eventually is a broken liveness promise at run end
// whatever unit bounded it. Bug class: choosing "steps" over "seconds" quietly
// turning an unmet obligation into a vacuous pass.
func TestFinalize_StepBoundedEventuallyMatchesWallClock(t *testing.T) {
byStep, stepEvaluator := runAndFinalize(EventuallyWithinSteps(ThunkNamed("p", alwaysFalse()), 50), 3)
byClock, clockEvaluator := runAndFinalize(EventuallyWithin(ThunkNamed("p", alwaysFalse()), time.Hour), 3)
if byStep != VerdictViolated || byClock != VerdictViolated {
t.Fatalf("step bound = %v, wall clock = %v, want both violated", byStep, byClock)
}
stepWitness, clockWitness := stepEvaluator.Violation(), clockEvaluator.Violation()
if stepWitness == nil || clockWitness == nil {
t.Fatal("both undischarged obligations must carry a witness")
}
if stepWitness.Reason != clockWitness.Reason {
t.Errorf("reasons diverge: step %q, wall clock %q", stepWitness.Reason, clockWitness.Reason)
}
if stepWitness.Step != clockWitness.Step {
t.Errorf("origin steps diverge: step %d, wall clock %d", stepWitness.Step, clockWitness.Step)
}
}
// The reason the unit exists. Two action-selection policies get the same
// 300-step budget and reach the same state at the same step, but the model
// policy takes 359 seconds where the seeded policy takes 47 because it makes a
// provider call per step. A wall-clock bound fails the slow policy on elapsed
// time alone; the same window written in steps decides both policies alike.
func TestStepBound_SlowPolicyDoesNotFailOnTimeAlone(t *testing.T) {
const budget = 300
const satisfiedAtObservation = 260
seededCadence := 47 * time.Second / budget
modelCadence := 359 * time.Second / budget
run := func(cadence time.Duration, bound func(Formula) Formula) Verdict {
observed := 0
inner := ThunkNamed("someTransactionExists", func() (bool, error) {
observed++
return observed >= satisfiedAtObservation, nil
})
evaluator := NewEvaluator(bound(inner))
base := time.Unix(1780000000, 0)
for index := range budget {
verdict := evaluator.ObserveAtStep(base.Add(time.Duration(index)*cadence), index+1)
if verdict != VerdictPending {
return verdict
}
}
return evaluator.Finalize()
}
byClock := func(inner Formula) Formula { return EventuallyWithin(inner, 300*time.Second) }
bySteps := func(inner Formula) Formula { return EventuallyWithinSteps(inner, 1915) }
if got := run(seededCadence, byClock); got != VerdictHolds {
t.Errorf("wall-clock bound under the seeded policy: got %v, want holds", got)
}
if got := run(modelCadence, byClock); got != VerdictViolated {
t.Errorf("wall-clock bound under the model policy: got %v, want violated (the false positive this unit removes)", got)
}
if got := run(seededCadence, bySteps); got != VerdictHolds {
t.Errorf("step bound under the seeded policy: got %v, want holds", got)
}
if got := run(modelCadence, bySteps); got != VerdictHolds {
t.Errorf("step bound under the model policy: got %v, want holds", got)
}
}
@@ -0,0 +1,158 @@
package runner
import (
"context"
"fmt"
"testing"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// authoredParityTreeJSON holds one target per authored action shape: a button to
// tap, a field to type into, a scrollable container to scroll, and a disabled
// button, which is a target like any other here.
const authoredParityTreeJSON = `{
"attributes": {"bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "Save", "text": "Save", "bounds": "[0,0,200,60]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"resource-id": "Amount", "class": "EditText", "hintText": "Amount", "bounds": "[0,100,400,160]"}, "enabled": true, "children": []},
{"attributes": {"resource-id": "List", "scrollable": "true", "bounds": "[0,300,400,700]"}, "children": []},
{"attributes": {"resource-id": "Off", "text": "Off", "bounds": "[0,700,200,760]"}, "clickable": true, "enabled": false, "children": []}
]
}`
// TestPoliciesDispatchTheSameAuthoredAction is the authored-leaf half of the
// policy-parity claim (policy_parity_test.go in the verifier package covers the
// builtin verbs): for one authored action, the seeded picker and the model's
// candidate must reach the DRIVER as the same call. The two carry it
// differently on purpose (a selector target keeps its coordinates on one side
// and re-resolves on the other), so the comparison is what executes, never the
// struct.
func TestPoliciesDispatchTheSameAuthoredAction(t *testing.T) {
fastFocusSettle(t)
cases := []struct {
name string
leaf string
}{
{"tap an element", `const e = state.ax.find("id:Save"); return e ? [Tap({on: e})] : [];`},
{"tap a selector", `return [Tap({on: "id:Save"})];`},
// Attempting a disabled control is a legitimate thing for a UI fuzzer to
// do and is exactly where boundary defects live, so neither policy may
// quietly refuse to offer it.
{"tap a disabled element", `const e = state.ax.find("id:Off"); return e ? [Tap({on: e})] : [];`},
{"tap a disabled selector", `return [Tap({on: "id:Off"})];`},
{"double-tap an element", `const e = state.ax.find("id:Save"); return e ? [DoubleTap({on: e})] : [];`},
{"long-press an element", `const e = state.ax.find("id:Save"); return e ? [LongPress({on: e})] : [];`},
{"type into an element", `const e = state.ax.find("id:Amount"); return e ? [InputText({into: e, text: "42"})] : [];`},
{"type into a selector", `return [InputText({into: "id:Amount", text: "42"})];`},
{"scroll an element", `const e = state.ax.find("id:List"); return e ? [Scroll({direction: "down", in: e})] : [];`},
{"scroll a selector", `return [Scroll({direction: "down", in: "id:List"})];`},
{"scroll with no container", `return [Scroll({direction: "up"})];`},
{"swipe with a duration", `return [Swipe({from: {x: 10, y: 600}, to: {x: 10, y: 100}, durationMillis: 300})];`},
{"swipe with no duration", `return [Swipe({from: {x: 10, y: 600}, to: {x: 10, y: 100}})];`},
{"press a key", `return [PressKey({key: "back"})];`},
{"wait", `return [Wait({durationMillis: 5})];`},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
spec := authoredSpec(testCase.leaf)
tree := mustParseTree(t, authoredParityTreeJSON)
seededVerifier := loadAuthoredSpec(t, spec, tree)
seededAction, err := seededVerifier.NextAction()
if err != nil {
t.Fatalf("seeded picker declined the authored action: %v", err)
}
modelVerifier := loadAuthoredSpec(t, spec, tree)
candidates := mustCandidates(t, modelVerifier, verifier.LabelSourceVisibleText)
if len(candidates) != 1 {
t.Fatalf("model was offered %d candidates, want the one authored action", len(candidates))
}
seeded := dispatchToMock(t, seededAction, tree)
model := dispatchToMock(t, candidates[0].Action, tree)
if len(seeded) != len(model) {
t.Fatalf("policies dispatched different calls\n seeded=%v\n model=%v", seeded, model)
}
for i := range seeded {
if seeded[i] != model[i] {
t.Errorf("call %d differs\n seeded=%+v\n model=%+v", i, seeded[i], model[i])
}
}
})
}
}
// TestSeededAuthoredScrollDragsInsideItsContainer is the half of the scroll
// contract parity alone cannot check: both policies agreeing on a drag from a
// point to itself would still be two policies scrolling nothing. The list sits
// at y 300..700, so the gesture has to start inside it and travel upward to
// reveal what is below.
func TestSeededAuthoredScrollDragsInsideItsContainer(t *testing.T) {
tree := mustParseTree(t, authoredParityTreeJSON)
spec := authoredSpec(`const e = state.ax.find("id:List"); return e ? [Scroll({direction: "down", in: e})] : [];`)
action, err := loadAuthoredSpec(t, spec, tree).NextAction()
if err != nil {
t.Fatalf("seeded picker declined the authored scroll: %v", err)
}
calls := dispatchToMock(t, action, tree)
if len(calls) != 1 {
t.Fatalf("want one driver call, got %v", calls)
}
swipe := calls[0]
if swipe.FromY < 300 || swipe.FromY > 700 {
t.Errorf("gesture starts at y=%d, outside the list (300..700)", swipe.FromY)
}
if swipe.ToY >= swipe.FromY {
t.Errorf("scroll down must drag upward, got from y=%d to y=%d", swipe.FromY, swipe.ToY)
}
}
func authoredSpec(leaf string) string {
return fmt.Sprintf(`
import { actions, DoubleTap, InputText, LongPress, PressKey, Scroll, Swipe, Tap, Wait } from "@sanderling/spec";
globalThis.actions = actions(() => { %s });
`, leaf)
}
func loadAuthoredSpec(t *testing.T, spec string, tree *hierarchy.Tree) *verifier.Verifier {
t.Helper()
loaded, err := verifier.New()
if err != nil {
t.Fatal(err)
}
if err := loaded.Load(bundleSpec(t, spec)); err != nil {
t.Fatal(err)
}
if err := loaded.PushSnapshot(verifier.SnapshotInput{Snapshots: verifier.Snapshots{}, Tree: tree}); err != nil {
t.Fatal(err)
}
return loaded
}
func mustParseTree(t *testing.T, treeJSON string) *hierarchy.Tree {
t.Helper()
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
return tree
}
// dispatchToMock runs one action through applyAction and returns what the
// driver was asked to do.
func dispatchToMock(t *testing.T, action verifier.Action, tree *hierarchy.Tree) []mockdriver.Action {
t.Helper()
drv := mockdriver.New()
skipped, err := applyAction(context.Background(), drv, action, tree)
if err != nil {
t.Fatalf("applyAction: %v", err)
}
if skipped != "" {
t.Fatalf("action %+v never reached the driver: %s", action, skipped)
}
return drv.Actions()
}
@@ -0,0 +1,96 @@
package runner
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/trace"
)
// throwingDriver reports one captured uncaught error, the way the chrome
// driver reports the page's buffer.
type throwingDriver struct {
*mockdriver.Driver
}
func (d *throwingDriver) Exceptions(
context.Context,
) ([]driver.Exception, error) {
return []driver.Exception{{
Class: "TypeError",
Message: "cannot read balance of null",
StackTrace: "at render (app.js:12)",
UnixMillis: 1700000000000,
}}, nil
}
// TestRunner_LogsAndExceptionsLandInTheTrace covers the error surface an
// offline oracle has no other source for: the default properties read
// state.logs and state.exceptions, and neither used to survive the step.
func TestRunner_LogsAndExceptionsLandInTheTrace(t *testing.T) {
state := newHarness(t)
state.mock.LogEntries = []driver.LogEntry{{
UnixMillis: 1700000000123,
Level: "E",
Tag: "AndroidRuntime",
Message: "FATAL EXCEPTION: main",
}}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 1,
Driver: &throwingDriver{Driver: state.mock},
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
body, err := os.ReadFile(
filepath.Join(state.writer.Directory(), "trace.jsonl"),
)
if err != nil {
t.Fatal(err)
}
line := strings.SplitN(strings.TrimSpace(string(body)), "\n", 2)[0]
var stored struct {
TraceVersion int `json:"trace_version"`
Logs []trace.LogEntry `json:"logs"`
Exceptions []trace.Exception `json:"exceptions"`
}
if err := json.Unmarshal([]byte(line), &stored); err != nil {
t.Fatalf("decode trace line: %v\n%s", err, line)
}
if stored.TraceVersion != trace.TraceVersion {
t.Errorf(
"trace_version = %d, want %d; an old trace could not be told apart",
stored.TraceVersion,
trace.TraceVersion,
)
}
want := trace.LogEntry{
UnixMillis: 1700000000123,
Level: "E",
Tag: "AndroidRuntime",
Message: "FATAL EXCEPTION: main",
}
if len(stored.Logs) != 1 || stored.Logs[0] != want {
t.Errorf("logs on disk = %+v, want [%+v]", stored.Logs, want)
}
if len(stored.Exceptions) != 1 ||
stored.Exceptions[0].Class != "TypeError" ||
stored.Exceptions[0].Message != "cannot read balance of null" ||
stored.Exceptions[0].StackTrace != "at render (app.js:12)" {
t.Errorf("exceptions on disk = %+v", stored.Exceptions)
}
}
@@ -0,0 +1,225 @@
package runner
import (
"bytes"
"context"
"encoding/json"
"errors"
"io"
"log/slog"
"os"
"path/filepath"
"testing"
"time"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/trace"
)
// slowToDrawDriver is an app that is the resumed activity immediately and whose
// window takes drawsAfter focus polls to appear: the shape of a real cold start,
// where ResumedActivity flips before the first frame. WaitForIdle returns at
// once, as it does on a device whose UI thread is quiet between frames.
type slowToDrawDriver struct {
*mockdriver.Driver
drawsAfter int
polls int
}
func (d *slowToDrawDriver) ForegroundApp(context.Context) (string, error) {
return guardedBundleID, nil
}
func (d *slowToDrawDriver) FocusedWindowApp(context.Context) (string, error) {
d.polls++
if d.polls > d.drawsAfter {
return guardedBundleID, nil
}
return "", nil
}
func (d *slowToDrawDriver) WaitForIdle(context.Context, time.Duration) error { return nil }
// The gate's budget has to be a duration, not a count of polls. A count is not a
// budget: each poll costs whatever the driver's idle wait happens to take, so the
// same launch clears the gate on one device and exhausts it on another. That is
// what an 80-run campaign against one app measured: on API 34 the idle wait
// returned in ~100ms, eight polls gave up 1.2s in, and the app's window drew at
// ~1.9s; on API 36 the same eight polls spanned 3s and cleared the same launch.
// Same app, same harness, a verdict that came from the device.
func TestAwaitForeground_BudgetIsTimeNotPolls(t *testing.T) {
fastForegroundGate(t)
const drawsAfter = 40
device := &slowToDrawDriver{Driver: mockdriver.New(), drawsAfter: drawsAfter}
options := Options{
BundleID: guardedBundleID,
Driver: device,
IdleTimeout: time.Millisecond,
}
ready := awaitForeground(context.Background(), options, discardLogger(), 0)
if !ready {
t.Fatalf("the gate gave up after %d poll(s) while the app was the resumed activity "+
"and its window drew on poll %d; a budget counted in polls expires at a "+
"different wall-clock time on every device", device.polls, drawsAfter+1)
}
if device.polls <= drawsAfter {
t.Fatalf("the gate polled %d time(s), want more than %d: it has to keep looking "+
"until its budget runs out, not stop at a fixed count", device.polls, drawsAfter)
}
}
// neverDrawsDriver never brings the app forward: the app under test is not the
// foreground app and no relaunch changes that, which is what a genuinely unmet
// precondition looks like.
type neverDrawsDriver struct {
*mockdriver.Driver
}
func (d *neverDrawsDriver) ForegroundApp(context.Context) (string, error) {
return "com.android.launcher", nil
}
func (d *neverDrawsDriver) FocusedWindowApp(context.Context) (string, error) {
return "com.android.launcher", nil
}
func (d *neverDrawsDriver) WaitForIdle(context.Context, time.Duration) error { return nil }
// A run whose app never came to the foreground explored nothing, and reporting it
// as a run that found no violations counts a harness failure as evidence about
// the app. It has to end the run and say so where a campaign can count it, not
// warn once and carry on into steps that observe some other app.
func TestRun_AppNeverReachesForegroundEndsTheRun(t *testing.T) {
fastForegroundGate(t)
state := newHarness(t)
device := &neverDrawsDriver{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: time.Millisecond,
MaxSteps: 3,
BundleID: guardedBundleID,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
var notReached ForegroundNotReachedError
if !errors.As(err, &notReached) {
t.Fatalf("Run returned %v, want a ForegroundNotReachedError: a run that never got "+
"the app on screen is not a run that explored it", err)
}
if notReached.BundleID != guardedBundleID {
t.Errorf("the error names %q, want %q", notReached.BundleID, guardedBundleID)
}
if summary.Steps != 0 {
t.Errorf("the run took %d step(s) after its precondition failed, want 0", summary.Steps)
}
records := preconditionRecords(t, state.writer.Directory())
if len(records) != 1 {
t.Fatalf("the trace holds %d precondition record(s), want 1: a campaign has to be "+
"able to count this without grepping logs", len(records))
}
if records[0].Index != 0 {
t.Errorf("the record is on step %d, want 0 (no step ever ran)", records[0].Index)
}
if records[0].PreconditionFailure != preconditionAppNotForeground {
t.Errorf("the trace records %q, want %q",
records[0].PreconditionFailure, preconditionAppNotForeground)
}
}
// leavesForegroundForeverDriver walks out of the app after the first step and
// never comes back, so the per-step scope guard exhausts its budget mid-run.
type leavesForegroundForeverDriver struct {
*mockdriver.Driver
checks int
}
func (d *leavesForegroundForeverDriver) ForegroundApp(context.Context) (string, error) {
d.checks++
if d.checks <= 2 {
return guardedBundleID, nil
}
return "com.android.launcher", nil
}
func (d *leavesForegroundForeverDriver) FocusedWindowApp(ctx context.Context) (string, error) {
return d.ForegroundApp(ctx)
}
func (d *leavesForegroundForeverDriver) WaitForIdle(context.Context, time.Duration) error {
return nil
}
// The same fact mid-run: a step the guard could not return to the app observed
// something that is not the app under test, and until now nothing in the trace
// said so.
func TestRun_StepsOutsideTheAppAreRecordedInTheTrace(t *testing.T) {
fastForegroundGate(t)
state := newHarness(t)
device := &leavesForegroundForeverDriver{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: time.Millisecond,
MaxSteps: 2,
BundleID: guardedBundleID,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
records := preconditionRecords(t, state.writer.Directory())
if len(records) == 0 {
t.Fatal("no step recorded that the guard never got the app back, so a run that " +
"spent its steps outside the app reads exactly like one that explored it")
}
for _, record := range records {
if record.Index == 0 {
t.Error("a mid-run failure was recorded as the startup gate's verdict (step 0)")
}
if record.PreconditionFailure != preconditionAppNotForeground {
t.Errorf("step %d records %q, want %q",
record.Index, record.PreconditionFailure, preconditionAppNotForeground)
}
}
}
func discardLogger() *slog.Logger {
return slog.New(slog.NewTextHandler(io.Discard, &slog.HandlerOptions{Level: slog.LevelError}))
}
// preconditionRecords reads the trace back off disk and returns the steps naming
// an unmet precondition, decoded through trace.Step so the test reads the same
// field a campaign would.
func preconditionRecords(t *testing.T, directory string) []trace.Step {
t.Helper()
body, err := os.ReadFile(filepath.Join(directory, "trace.jsonl"))
if err != nil {
t.Fatalf("read trace: %v", err)
}
var records []trace.Step
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
if len(raw) == 0 {
continue
}
var step trace.Step
if err := json.Unmarshal(raw, &step); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if step.PreconditionFailure != "" {
records = append(records, step)
}
}
return records
}
+171 -32
View File
@@ -13,7 +13,9 @@ import (
"log/slog"
"regexp"
"strings"
"time"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/llmclient"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
@@ -49,8 +51,15 @@ type llmSource struct {
// instructions is optional spec-level guidance appended to the system prompt
// to steer the model's bug-hunting (empty when unset).
instructions string
logger *slog.Logger
history *actionHistory
// labelSource is the channel each candidate's target is named by in the
// numbered list. It lives here rather than on the verifier because only this
// policy reads labels at all.
labelSource string
logger *slog.Logger
history *actionHistory
// recorder persists one record per step: what was sent, what came back, and
// how the step ended. Nil only in unit tests that never select.
recorder llmCallRecorder
// lastSource/lastReasoning describe the most recent NextAction so the runner
// can stamp the trace. lastSource is "llm" only when the LLM (not setup)
@@ -71,10 +80,18 @@ type llmSelection struct {
chosenAction string
}
// llmCallRecorder persists one selection record per step. *trace.Writer
// implements it.
type llmCallRecorder interface {
WriteLLMCall(call trace.LLMCall) error
}
// NextAction returns the step's action. Setup precedence is preserved by
// running the JS path first (the llm marker is inert there, so a null result
// means setup yielded nothing); the LLM selection then takes over.
func (s *llmSource) NextAction(ctx context.Context) (verifier.Action, error) {
// means setup yielded nothing); the LLM selection then takes over. Every way a
// step can end lands one record keyed to stepIndex, so a step the guards threw
// away is never confused with one the picker declined.
func (s *llmSource) NextAction(ctx context.Context, stepIndex int) (verifier.Action, error) {
s.lastSource = ""
s.lastReasoning = ""
s.lastChoice = 0
@@ -85,55 +102,94 @@ func (s *llmSource) NextAction(ctx context.Context) (verifier.Action, error) {
// setup (e.g. login) first but never the weighted picker.
action, err := s.verifier.SetupAction()
if err == nil {
s.history.add(describeAction(action))
s.record(stepIndex, trace.LLMCall{Outcome: trace.LLMOutcomeSetupAction})
s.history.add(describeAction(action, s.verifier.Tree()))
return action, nil
}
if !errors.Is(err, verifier.ErrNoAction) {
s.record(stepIndex, trace.LLMCall{Outcome: trace.LLMOutcomeSetupFailed, Error: err.Error()})
return verifier.Action{}, err
}
selection, ok := s.selectViaLLM(ctx)
if !ok {
// Any failure (HTTP error, unusable output, invalid choice, echo
// mismatch) skips the step; the next step re-observes and tries again.
selection, call, err := s.selectViaLLM(ctx)
s.record(stepIndex, call)
if err != nil {
return verifier.Action{}, err
}
if call.Outcome != trace.LLMOutcomeSelected {
// Every other outcome skips the step; the next step re-observes and
// tries again. The record says which one it was.
return verifier.Action{}, verifier.ErrNoAction
}
s.lastSource = "llm"
s.lastReasoning = selection.reasoning
s.lastChoice = selection.choice
s.lastChosenAction = selection.chosenAction
s.history.add(describeAction(selection.action))
s.history.add(describeAction(selection.action, s.verifier.Tree()))
return selection.action, nil
}
// selectViaLLM runs one multimodal call and maps the chosen number to an action.
// It returns ok=false on any error/empty/invalid output, logging the cause; the
// caller turns that into a skipped step.
func (s *llmSource) selectViaLLM(ctx context.Context) (llmSelection, bool) {
candidates := s.verifier.Candidates()
if len(candidates) == 0 {
return llmSelection{}, false
// The returned record is complete whichever way the selection ended: its
// Outcome is trace.LLMOutcomeSelected exactly when the returned selection is
// usable.
//
// The error is the spec refusing this policy rather than a step going nowhere:
// every later step would refuse identically, so the run stops instead of
// recording two hundred skipped steps.
func (s *llmSource) selectViaLLM(ctx context.Context) (llmSelection, trace.LLMCall, error) {
call := trace.LLMCall{Timestamp: time.Now(), Model: s.model}
candidates, err := s.verifier.Candidates(s.labelSource)
if err != nil {
call.Outcome = trace.LLMOutcomeCandidatesFailed
call.Error = err.Error()
return llmSelection{}, call, err
}
if len(candidates) == 0 {
call.Outcome = trace.LLMOutcomeNoCandidates
return llmSelection{}, call, nil
}
call.Candidates = recordCandidates(candidates)
response, err := s.client.ChatCompletion(ctx, s.buildRequest(candidates))
request, screenshotReference := s.buildRequest(candidates)
call.Screenshot = screenshotReference
call.SystemPrompt, call.UserPrompt = promptTexts(request)
requestStart := time.Now()
response, err := s.client.ChatCompletion(ctx, request)
call.LatencyMillis = time.Since(requestStart).Milliseconds()
if err != nil {
s.logger.Warn("llm action selection failed", "err", err)
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeRequestFailed
call.Error = err.Error()
return llmSelection{}, call, nil
}
call.ServedModel = response.Model
call.PromptTokens = response.Usage.PromptTokens
call.CompletionTokens = response.Usage.CompletionTokens
call.TotalTokens = response.Usage.TotalTokens
if len(response.Choices) == 0 {
s.logger.Warn("llm returned no choices")
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeNoChoices
return llmSelection{}, call, nil
}
call.RawResponse = response.Choices[0].Message.Content
output, err := parseChoice(response.Choices[0].Message.Content)
output, err := parseChoice(call.RawResponse)
if err != nil {
s.logger.Warn("llm output unusable", "err", err)
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeUnparsableResponse
call.Error = err.Error()
return llmSelection{}, call, nil
}
call.Choice = output.Choice
call.EchoedAction = output.ChosenAction
call.Reasoning = output.Reasoning
// choice is 1-based into the numbered list.
if output.Choice < 1 || output.Choice > len(candidates) {
s.logger.Warn("llm choice out of range", "choice", output.Choice, "candidates", len(candidates))
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeChoiceOutOfRange
return llmSelection{}, call, nil
}
candidate := candidates[output.Choice-1]
// Strict skip: the echoed action must match the numbered entry, so a model
@@ -143,19 +199,79 @@ func (s *llmSource) selectViaLLM(ctx context.Context) (llmSelection, bool) {
if stripWeightSuffix(output.ChosenAction) != candidate.Description {
s.logger.Warn("llm chosen_action mismatch; skipping",
"choice", output.Choice, "echoed", output.ChosenAction, "candidate", candidate.Description)
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeEchoMismatch
return llmSelection{}, call, nil
}
action, err := s.actionForCandidate(candidate, output.Text)
if err != nil {
s.logger.Warn("building action from candidate failed", "choice", output.Choice, "err", err)
return llmSelection{}, false
call.Outcome = trace.LLMOutcomeActionBuildFailed
call.Error = err.Error()
return llmSelection{}, call, nil
}
call.Outcome = trace.LLMOutcomeSelected
return llmSelection{
action: action,
reasoning: output.Reasoning,
choice: output.Choice,
chosenAction: candidate.Description,
}, true
}, call, nil
}
// record stamps the step this selection belongs to and appends the record. A
// write failure must not kill the run, but it does mean the step is
// unattributable, so it is logged.
func (s *llmSource) record(stepIndex int, call trace.LLMCall) {
if s.recorder == nil {
return
}
call.Step = stepIndex
if call.Timestamp.IsZero() {
call.Timestamp = time.Now()
}
if err := s.recorder.WriteLLMCall(call); err != nil {
s.logger.Warn("llm call record failed", "step", stepIndex, "err", err)
}
}
// recordCandidates snapshots the numbered list as the prompt rendered it, so a
// campaign that varies candidate labelling can recover the labels each call saw.
func recordCandidates(candidates []verifier.ActionCandidate) []trace.LLMCandidate {
recorded := make([]trace.LLMCandidate, 0, len(candidates))
for _, candidate := range candidates {
entry := trace.LLMCandidate{
Index: candidate.Index,
Kind: string(candidate.Kind),
Description: candidate.Description,
Label: candidate.Label,
}
if candidate.Weighted {
entry.Weight = candidate.Weight
}
recorded = append(recorded, entry)
}
return recorded
}
// promptTexts reads the text parts back off the built request, so the record is
// what went on the wire rather than a second rendering of the same templates.
func promptTexts(request llmclient.Request) (systemPrompt, userPrompt string) {
for _, message := range request.Messages {
var parts []string
for _, part := range message.Content {
if part.Type == "text" {
parts = append(parts, part.Text)
}
}
text := strings.Join(parts, "\n")
switch message.Role {
case "system":
systemPrompt = text
case "user":
userPrompt = text
}
}
return systemPrompt, userPrompt
}
// actionForCandidate turns a chosen candidate into the executable action. The
@@ -182,14 +298,23 @@ func (s *llmSource) actionForCandidate(candidate verifier.ActionCandidate, text
// buildRequest assembles the one-shot multimodal request: a system frame, the
// numbered candidate list plus recent-action memory, and the downscaled
// screenshot. The strict json_schema response format pins the ranked output.
func (s *llmSource) buildRequest(candidates []verifier.ActionCandidate) llmclient.Request {
//
// The second return is the run-relative path of the screenshot attached, empty
// when none was, so a record can never claim an image the call did not carry.
// It names the step the image was captured at, which is the last observed step
// rather than the current one whenever an observation was skipped.
func (s *llmSource) buildRequest(candidates []verifier.ActionCandidate) (llmclient.Request, string) {
userParts := []llmclient.ContentPart{llmclient.TextPart(s.userPrompt(candidates))}
screenshotReference := ""
if screenshot := s.verifier.Screenshot(); len(screenshot) > 0 {
if dataURL, ok := screenshotDataURL(screenshot, llmMaxImageEdge); ok {
userParts = append(userParts, llmclient.ImagePart(dataURL))
if step := s.verifier.SnapshotStep(); step > 0 {
screenshotReference = trace.ScreenshotReference(step)
}
}
}
return llmclient.Request{
request := llmclient.Request{
Model: s.model,
Messages: []llmclient.Message{
{Role: "system", Content: []llmclient.ContentPart{llmclient.TextPart(s.systemPrompt())}},
@@ -197,6 +322,7 @@ func (s *llmSource) buildRequest(candidates []verifier.ActionCandidate) llmclien
},
ResponseFormat: choiceResponseFormat(len(candidates)),
}
return request, screenshotReference
}
// systemPrompt is the base framing plus any spec-level instructions, appended as
@@ -297,10 +423,13 @@ func parseChoice(content string) (choiceOutput, error) {
}
// describeAction renders a short action summary for the recent-action memory.
func describeAction(action verifier.Action) string {
// The tree is the one the action was chosen against, which is what says whether
// a typed value may be written into the memory at all.
func describeAction(action verifier.Action, tree *hierarchy.Tree) string {
switch action.Kind {
case verifier.ActionKindInputText:
return fmt.Sprintf("InputText %s = %q", actionTarget(action), action.Text)
return fmt.Sprintf("InputText %s = %q",
actionTarget(action), verifier.RecordedActionText(action, tree))
case verifier.ActionKindScroll:
// A builtin gesture carries endpoints rather than a selector, so name the
// container by where the drag starts; that is what tells two scrollable
@@ -367,9 +496,10 @@ func (h *actionHistory) recent() []historyEntry {
return h.entries
}
// stampActionSource records the backend that chose an action on the trace.
// Only an LLM-selected action (not a setup action the JS path produced) carries
// source="llm" and the model's reasoning.
// stampActionSource names the model as the producer of an action it chose. The
// spec's two generators name themselves on the wire (runtime-entry.ts) and
// traceActionFor has already carried that over; a model pick is built here from
// the candidate list, so nothing on the wire could have named it.
func stampActionSource(traceAction *trace.Action, source ActionSource) {
if traceAction == nil {
return
@@ -384,6 +514,15 @@ func stampActionSource(traceAction *trace.Action, source ActionSource) {
traceAction.LLMChosenAction = llm.lastChosenAction
}
// generatorChoseAction reports whether this step drove the app on the action
// generator's behalf rather than on setup's, which is the exposure a per-action
// rate divides by. Both policies are read the same way, off the stamped trace
// action: an unstamped one is a source the runner cannot name, and counting it
// as setup would report a run that explored as having explored nothing.
func generatorChoseAction(traceAction *trace.Action) bool {
return traceAction != nil && traceAction.Source != trace.ActionSourceSetup
}
// screenshotDataURL downscales the PNG and encodes it as a data URL for the
// image content part.
func screenshotDataURL(pngBytes []byte, maxEdge int) (string, bool) {
+738 -13
View File
@@ -5,21 +5,27 @@ import (
"context"
"encoding/json"
"errors"
"fmt"
"image"
"image/png"
"io"
"log/slog"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"regexp"
"slices"
"strconv"
"strings"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/llmclient"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
@@ -69,11 +75,11 @@ func TestDescribeActionNamesGesturesByOrigin(t *testing.T) {
Kind: verifier.ActionKindScroll, Direction: "down",
FromX: 200, FromY: 500, ToX: 200, ToY: 340,
}
if got := describeAction(builtin); got != "Scroll down (200,500)" {
if got := describeAction(builtin, nil); got != "Scroll down (200,500)" {
t.Errorf("builtin gesture described as %q, want the drag origin", got)
}
authored := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "up", On: "id:List"}
if got := describeAction(authored); got != "Scroll up id:List" {
if got := describeAction(authored, nil); got != "Scroll up id:List" {
t.Errorf("authored scroll described as %q, want its selector", got)
}
}
@@ -168,6 +174,9 @@ type fakeOpenRouter struct {
chosenAction string
text string
reasoning string
usage llmclient.Usage
servedModel string
delay time.Duration
lastRequest map[string]any
}
@@ -184,8 +193,11 @@ func newFakeOpenRouter(t *testing.T) *fakeOpenRouter {
"text": fake.text,
})
response, _ := json.Marshal(llmclient.Response{
Model: fake.servedModel,
Choices: []llmclient.Choice{{Message: llmclient.ResponseMessage{Content: string(content)}}},
Usage: fake.usage,
})
time.Sleep(fake.delay)
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write(response)
}))
@@ -193,7 +205,52 @@ func newFakeOpenRouter(t *testing.T) *fakeOpenRouter {
return fake
}
// fakeCallRecorder keeps the per-step selection records in memory so a test can
// assert on them without opening the run directory.
type fakeCallRecorder struct {
calls []trace.LLMCall
}
func (r *fakeCallRecorder) WriteLLMCall(call trace.LLMCall) error {
r.calls = append(r.calls, call)
return nil
}
func recordedCalls(t *testing.T, source *llmSource) []trace.LLMCall {
t.Helper()
recorder, ok := source.recorder.(*fakeCallRecorder)
if !ok {
t.Fatalf("recorder = %T, want *fakeCallRecorder", source.recorder)
}
return recorder.calls
}
func lastCall(t *testing.T, source *llmSource) trace.LLMCall {
t.Helper()
calls := recordedCalls(t, source)
if len(calls) == 0 {
t.Fatal("no selection record written")
}
return calls[len(calls)-1]
}
// mustCandidates enumerates the model policy's list, failing the test on the
// refusal an authored multi-item sampler raises.
func mustCandidates(t *testing.T, v *verifier.Verifier, labelSource string) []verifier.ActionCandidate {
t.Helper()
candidates, err := v.Candidates(labelSource)
if err != nil {
t.Fatalf("Candidates: %v", err)
}
return candidates
}
func newLLMSource(t *testing.T, fake *fakeOpenRouter) (*llmSource, *verifier.Verifier) {
t.Helper()
return newLLMSourceWithSpec(t, fake, llmFixtureSpec)
}
func newLLMSourceWithSpec(t *testing.T, fake *fakeOpenRouter, spec string) (*llmSource, *verifier.Verifier) {
t.Helper()
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", fake.server.URL)
@@ -206,7 +263,7 @@ func newLLMSource(t *testing.T, fake *fakeOpenRouter) (*llmSource, *verifier.Ver
if err != nil {
t.Fatal(err)
}
if err := verifierInstance.Load(bundleSpec(t, llmFixtureSpec)); err != nil {
if err := verifierInstance.Load(bundleSpec(t, spec)); err != nil {
t.Fatal(err)
}
if _, ok := verifierInstance.LLMConfig(); !ok {
@@ -219,21 +276,62 @@ func newLLMSource(t *testing.T, fake *fakeOpenRouter) (*llmSource, *verifier.Ver
model: "test/model",
logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
history: newActionHistory(llmHistorySize),
recorder: &fakeCallRecorder{},
}
return source, verifierInstance
}
func pushLLMSnapshot(t *testing.T, v *verifier.Verifier) {
t.Helper()
pushLLMSnapshotAtStep(t, v, 0)
}
func pushSnapshotTree(t *testing.T, v *verifier.Verifier, treeJSON string) {
t.Helper()
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatal(err)
}
if err := v.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
ScreenshotPNG: tinyPNG(t),
}); err != nil {
t.Fatalf("PushSnapshot: %v", err)
}
}
func pushLLMSnapshotAtStep(t *testing.T, v *verifier.Verifier, stepIndex int) {
t.Helper()
tree, err := hierarchy.Parse(llmTreeJSON)
if err != nil {
t.Fatal(err)
}
if err := v.PushSnapshot(verifier.SnapshotInput{Tree: tree, ScreenshotPNG: tinyPNG(t)}); err != nil {
if err := v.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
ScreenshotPNG: tinyPNG(t),
StepIndex: stepIndex,
}); err != nil {
t.Fatalf("PushSnapshot: %v", err)
}
}
func readLLMCalls(t *testing.T, directory string) []trace.LLMCall {
t.Helper()
body, err := os.ReadFile(filepath.Join(directory, trace.LLMCallFileName))
if err != nil {
t.Fatalf("read %s: %v", trace.LLMCallFileName, err)
}
var calls []trace.LLMCall
for _, line := range strings.Split(strings.TrimSpace(string(body)), "\n") {
var call trace.LLMCall
if err := json.Unmarshal([]byte(line), &call); err != nil {
t.Fatalf("decode record %q: %v", line, err)
}
calls = append(calls, call)
}
return calls
}
func candidateByKind(t *testing.T, candidates []verifier.ActionCandidate, kind verifier.ActionKind) verifier.ActionCandidate {
t.Helper()
for _, candidate := range candidates {
@@ -278,6 +376,93 @@ func TestPickSourcesSeededByDefault(t *testing.T) {
}
}
// TestPickSourcesGivesTheLabelSourceToTheModelPickerOnly pins the asymmetry the
// labelling factorial depends on: the label channel reaches the model picker,
// and the seeded picker is handed a source that has nowhere to put one.
func TestPickSourcesGivesTheLabelSourceToTheModelPickerOnly(t *testing.T) {
fake := newFakeOpenRouter(t)
_, verifierInstance := newLLMSource(t, fake)
logger := slog.New(slog.NewTextHandler(io.Discard, nil))
action, _, err := pickSources(Options{
Verifier: verifierInstance,
Generator: "llm",
LabelSource: verifier.LabelSourceResourceID,
Logger: logger,
})
if err != nil {
t.Fatalf("pickSources: %v", err)
}
model, ok := action.(*llmSource)
if !ok {
t.Fatalf("action source = %T, want *llmSource", action)
}
if model.labelSource != verifier.LabelSourceResourceID {
t.Errorf("labelSource = %q, want %q", model.labelSource, verifier.LabelSourceResourceID)
}
seeded, _, err := pickSources(Options{
Verifier: verifierInstance,
Generator: "seeded",
LabelSource: verifier.LabelSourceResourceID,
Logger: logger,
})
if err != nil {
t.Fatalf("pickSources: %v", err)
}
if _, ok := seeded.(gojaSource); !ok {
t.Errorf("action source = %T, want gojaSource, which carries no label channel", seeded)
}
}
// labelSplitTreeJSON names one control two ways, so a record can say which
// channel the model was reading.
const labelSplitTreeJSON = `{
"attributes": {"bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "add_credit_button", "text": "Add credit", "bounds": "[0,0,400,100]"}, "clickable": true, "enabled": true, "children": []}
]
}`
// TestLLMSourceRecordsTheLabelsTheModelSaw closes the loop from the flag to the
// artifact: llm-calls.jsonl carries the candidate list as rendered, so a
// directory of runs can be checked against the cell it claims rather than
// trusted.
func TestLLMSourceRecordsTheLabelsTheModelSaw(t *testing.T) {
for _, want := range []struct{ labelSource, label string }{
{verifier.LabelSourceVisibleText, "Add credit"},
{verifier.LabelSourceResourceID, "add_credit_button"},
} {
t.Run(want.labelSource, func(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
source.labelSource = want.labelSource
pushSnapshotTree(t, verifierInstance, labelSplitTreeJSON)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, want.labelSource), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
if _, err := source.NextAction(context.Background(), 0); err != nil {
t.Fatalf("NextAction: %v", err)
}
call := lastCall(t, source)
if call.Outcome != trace.LLMOutcomeSelected {
t.Fatalf("outcome = %q, want selected", call.Outcome)
}
if len(call.Candidates) == 0 {
t.Fatal("no candidates recorded")
}
if call.Candidates[0].Label != want.label {
t.Errorf("recorded label = %q, want %q", call.Candidates[0].Label, want.label)
}
if !strings.Contains(call.UserPrompt, want.label) {
t.Errorf("prompt does not carry %q:\n%s", want.label, call.UserPrompt)
}
})
}
}
// seededFixtureSpec declares no generator = llm(...), so --generator llm has
// nothing to build a picker from.
const seededFixtureSpec = `
@@ -376,7 +561,7 @@ func TestLLMSourceDrivesExecutedActions(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshot(t, verifierInstance)
candidates := verifierInstance.Candidates()
candidates := mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText)
// Step 1: the model picks the Tap on Submit by its number, echoing its
// description.
@@ -385,7 +570,7 @@ func TestLLMSourceDrivesExecutedActions(t *testing.T) {
fake.chosenAction = tap.Description
fake.reasoning = "tap submit"
fake.text = ""
action, err := source.NextAction(context.Background())
action, err := source.NextAction(context.Background(), 1)
if err != nil {
t.Fatalf("NextAction: %v", err)
}
@@ -411,7 +596,7 @@ func TestLLMSourceDrivesExecutedActions(t *testing.T) {
fake.chosenAction = typing.Description
fake.reasoning = "type a name"
fake.text = "Priya"
action, err = source.NextAction(context.Background())
action, err = source.NextAction(context.Background(), 1)
if err != nil {
t.Fatalf("NextAction: %v", err)
}
@@ -440,7 +625,7 @@ func TestLLMSourceSkipsOnOutOfRangeChoice(t *testing.T) {
fake.choice = 9999
fake.chosenAction = "whatever"
_, err := source.NextAction(context.Background())
_, err := source.NextAction(context.Background(), 1)
if !errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want ErrNoAction for an out-of-range choice", err)
}
@@ -455,12 +640,12 @@ func TestLLMSourceAcceptsEchoWithWeightSuffix(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshot(t, verifierInstance)
candidates := verifierInstance.Candidates()
candidates := mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText)
tap := candidateByKind(t, candidates, verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description + " (w" + strconv.Itoa(tap.Weight) + ")"
action, err := source.NextAction(context.Background())
action, err := source.NextAction(context.Background(), 1)
if err != nil {
t.Fatalf("NextAction: %v", err)
}
@@ -490,14 +675,14 @@ func TestLLMSourceStrictSkipsOnEchoMismatch(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshot(t, verifierInstance)
candidates := verifierInstance.Candidates()
candidates := mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText)
// A valid number, but the echoed action disagrees with that numbered entry:
// the model reasoned about one control and picked another's number.
tap := candidateByKind(t, candidates, verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = "Tap \"Something Else\""
_, err := source.NextAction(context.Background())
_, err := source.NextAction(context.Background(), 1)
if !errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want ErrNoAction on chosen_action mismatch", err)
}
@@ -506,6 +691,51 @@ func TestLLMSourceStrictSkipsOnEchoMismatch(t *testing.T) {
}
}
// llmSharedLabelTreeJSON has two rows a user reads as the same word, so the
// numbered list holds two entries rendering identically.
const llmSharedLabelTreeJSON = `{
"attributes": {"bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "delete_alpha", "text": "Delete", "bounds": "[0,0,400,100]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"resource-id": "delete_beta", "text": "Delete", "bounds": "[0,100,400,200]"}, "clickable": true, "enabled": true, "children": []}
]
}`
// TestLLMSourceEchoGuardAdmitsARepeatedDescription is the other half of the
// strict skip: it compares the echo against the entry the model NUMBERED, so a
// description shared by two entries still selects the one whose number came
// back. A guard that looked the echo up by description instead would run the
// first row for both numbers.
func TestLLMSourceEchoGuardAdmitsARepeatedDescription(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushSnapshotTree(t, verifierInstance, llmSharedLabelTreeJSON)
var repeated []verifier.ActionCandidate
for _, candidate := range mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText) {
if candidate.Description == `Tap "Delete"` {
repeated = append(repeated, candidate)
}
}
if len(repeated) != 2 {
t.Fatalf("want two entries sharing one description, got %d", len(repeated))
}
second := repeated[1]
fake.choice = second.Index
fake.chosenAction = second.Description
action, err := source.NextAction(context.Background(), 1)
if err != nil {
t.Fatalf("NextAction err = %v, want the second row's tap", err)
}
if action.On != "id:delete_beta" {
t.Errorf("action targets %q, want id:delete_beta", action.On)
}
if source.lastSource != "llm" {
t.Errorf("lastSource = %q, want llm; a repeated description must not strict-skip", source.lastSource)
}
}
func TestLLMSourceSkipsOnHTTPError(t *testing.T) {
fake := newFakeOpenRouter(t)
// Replace the handler with one that always errors.
@@ -516,12 +746,475 @@ func TestLLMSourceSkipsOnHTTPError(t *testing.T) {
pushLLMSnapshot(t, verifierInstance)
fake.choice = 1
_, err := source.NextAction(context.Background())
_, err := source.NextAction(context.Background(), 1)
if !errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want ErrNoAction on HTTP failure", err)
}
}
// TestRunner_EveryModelCallFailingIsNotACleanRun drives the whole loop against a
// provider that answers every call with a rate limit. The picker hands back no
// action on every step, so the run observes the app, judges the same launch
// screen over and over and never touches it. Only llm-calls.jsonl used to know:
// the trace, the summary and the exit status were a healthy run's.
func TestRunner_EveryModelCallFailingIsNotACleanRun(t *testing.T) {
fake := newFakeOpenRouter(t)
fake.server.Config.Handler = http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusTooManyRequests)
})
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", fake.server.URL)
state := newHarnessWithSpec(t, llmFixtureSpec)
state.mock.HierarchyJSON = llmTreeJSON
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 30 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Generator: "llm",
LabelSource: verifier.LabelSourceVisibleText,
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("Run: %v", err)
}
calls := readLLMCalls(t, state.writer.Directory())
if len(calls) != 3 {
t.Fatalf("recorded %d model calls, want 3", len(calls))
}
for _, call := range calls {
if call.Outcome != trace.LLMOutcomeRequestFailed {
t.Fatalf("call at step %d ended %q, want %q: the fixture must fail every call",
call.Step, call.Outcome, trace.LLMOutcomeRequestFailed)
}
}
if summary.DispatchedActions != 0 {
t.Fatalf("DispatchedActions = %d, want 0: every model call failed",
summary.DispatchedActions)
}
if got := summary.SkippedActions[string(actionSkippedNoActionProduced)]; got != 3 {
t.Errorf("summary counted %d step(s) as %q, want 3: %v",
got, actionSkippedNoActionProduced, summary.SkippedActions)
}
for _, line := range readTraceLines(t, state.writer.Directory()) {
if line.ActionSkipped != string(actionSkippedNoActionProduced) {
t.Errorf("step %d action_skipped = %q, want %q",
line.Step, line.ActionSkipped, actionSkippedNoActionProduced)
}
}
for _, action := range state.mock.Actions() {
switch action.Kind {
case mockdriver.ActionTap, mockdriver.ActionTapSelector, mockdriver.ActionInputText:
t.Errorf("the run drove the app with %v after every model call failed", action)
}
}
}
// llmLoginSetupSpec is the shape every spec with a login has: setup drives the
// app for its first steps and then yields nothing, leaving the rest of the run
// to the generator.
const llmLoginSetupSpec = `
import { llm, always, actions, taps, typing, weighted, Tap } from "@sanderling/spec";
globalThis.properties = { ok: always(() => true) };
let setupTapsLeft = 2;
globalThis.setup = actions(() => (setupTapsLeft-- > 0 ? [Tap({ on: "id:Submit" })] : []));
globalThis.actions = weighted([1, taps], [1, typing]);
globalThis.generator = llm({ model: "test/model" });
`
// TestRunner_SetupActionsAreNotTheGeneratorDrivingTheApp is the folio run: the
// spec's login setup dispatches the first steps, then every model call fails.
// Two actions reached the app and none of them explored it, and a run counted
// by dispatched actions alone reports that as a clean run.
func TestRunner_SetupActionsAreNotTheGeneratorDrivingTheApp(t *testing.T) {
fake := newFakeOpenRouter(t)
fake.server.Config.Handler = http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusTooManyRequests)
})
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", fake.server.URL)
state := newHarnessWithSpec(t, llmLoginSetupSpec)
state.mock.HierarchyJSON = llmTreeJSON
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 30 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 4,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Generator: "llm",
LabelSource: verifier.LabelSourceVisibleText,
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("Run: %v", err)
}
wantOutcomes := []string{
trace.LLMOutcomeSetupAction, trace.LLMOutcomeSetupAction,
trace.LLMOutcomeRequestFailed, trace.LLMOutcomeRequestFailed,
}
calls := readLLMCalls(t, state.writer.Directory())
if len(calls) != len(wantOutcomes) {
t.Fatalf("recorded %d selection records, want %d", len(calls), len(wantOutcomes))
}
for index, call := range calls {
if call.Outcome != wantOutcomes[index] {
t.Fatalf("call at step %d ended %q, want %q", call.Step, call.Outcome, wantOutcomes[index])
}
}
lines := readTraceLines(t, state.writer.Directory())
if len(lines) != 4 {
t.Fatalf("wrote %d trace lines, want 4", len(lines))
}
for _, line := range lines[:2] {
if line.NextAction == nil || line.ActionSkipped != "" {
t.Errorf("step %d = %+v, want setup's action dispatched", line.Step, line)
}
}
for _, line := range lines[2:] {
if line.ActionSkipped != string(actionSkippedNoActionProduced) {
t.Errorf("step %d action_skipped = %q, want %q",
line.Step, line.ActionSkipped, actionSkippedNoActionProduced)
}
}
if summary.DispatchedActions != 2 {
t.Errorf("DispatchedActions = %d, want 2: setup drove the app twice",
summary.DispatchedActions)
}
if summary.GeneratorActions != 0 {
t.Errorf("GeneratorActions = %d, want 0: every model call failed",
summary.GeneratorActions)
}
if got := summary.SkippedActions[string(actionSkippedNoActionProduced)]; got != 2 {
t.Errorf("summary counted %d step(s) as %q, want 2: %v",
got, actionSkippedNoActionProduced, summary.SkippedActions)
}
taps := 0
for _, action := range state.mock.Actions() {
if action.Kind == mockdriver.ActionTap || action.Kind == mockdriver.ActionTapSelector {
taps++
}
}
if taps != 2 {
t.Errorf("the run drove the app %d time(s), want the 2 setup taps only", taps)
}
}
// TestRunner_SetupAndGeneratorBothDrivingIsAHealthyRun is the same spec with a
// provider that answers: setup drives its steps and the model drives the rest,
// which is what an ordinary login-fronted run looks like. Counting only the
// generator's actions must not turn it red.
func TestRunner_SetupAndGeneratorBothDrivingIsAHealthyRun(t *testing.T) {
fake := newFakeOpenRouter(t)
t.Setenv("OPENROUTER_API_KEY", "test-key")
t.Setenv("OPENROUTER_BASE_URL", fake.server.URL)
state := newHarnessWithSpec(t, llmLoginSetupSpec)
state.mock.HierarchyJSON = llmTreeJSON
pushSnapshotTree(t, state.verifier, llmTreeJSON)
tap := candidateByKind(t,
mustCandidates(t, state.verifier, verifier.LabelSourceVisibleText),
verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 30 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 4,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Generator: "llm",
LabelSource: verifier.LabelSourceVisibleText,
Logger: slog.New(slog.NewTextHandler(io.Discard, nil)),
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.DispatchedActions != 4 {
t.Errorf("DispatchedActions = %d, want 4: every step drove the app",
summary.DispatchedActions)
}
if summary.GeneratorActions != 2 {
t.Errorf("GeneratorActions = %d, want 2: the model drove the two steps setup left it",
summary.GeneratorActions)
}
lines := readTraceLines(t, state.writer.Directory())
if len(lines) != 4 {
t.Fatalf("wrote %d trace lines, want 4", len(lines))
}
for _, line := range lines {
if line.NextAction == nil || line.ActionSkipped != "" {
t.Fatalf("step %d = %+v, want an action dispatched", line.Step, line)
}
}
for _, line := range lines[:2] {
if line.NextAction.Source != trace.ActionSourceSetup {
t.Errorf("step %d action source = %q, want setup: setup chose it",
line.Step, line.NextAction.Source)
}
}
for _, line := range lines[2:] {
if line.NextAction.Source != trace.ActionSourceModel {
t.Errorf("step %d action source = %q, want llm", line.Step, line.NextAction.Source)
}
}
}
// llmSetupFixtureSpec drives the first action from setup, so the model is never
// consulted for that step.
const llmSetupFixtureSpec = `
import { llm, always, actions, taps, typing, weighted, Tap } from "@sanderling/spec";
globalThis.properties = { ok: always(() => true) };
globalThis.setup = actions(() => [Tap({ on: "id:Submit" })]);
globalThis.actions = weighted([1, taps], [1, typing]);
globalThis.generator = llm({ model: "test/model" });
`
// TestLLMCallRecordSeparatesGuardSkipFromDecline pins the reason these records
// exist. A step the echo guard threw away and a step where the picker had
// nothing to choose both used to leave nothing behind but a log line, so any
// defect-yield or actions-per-hour figure computed from a model run silently
// mixed the two with each other and with steps that acted.
func TestLLMCallRecordSeparatesGuardSkipFromDecline(t *testing.T) {
const stepIndex = 7
guardSkipped := func(t *testing.T) trace.LLMCall {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshot(t, verifierInstance)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = `Tap "Something Else"`
if _, err := source.NextAction(context.Background(), stepIndex); !errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want ErrNoAction on echo mismatch", err)
}
return lastCall(t, source)
}
declined := func(t *testing.T) trace.LLMCall {
fake := newFakeOpenRouter(t)
// No snapshot pushed, so the action tree yields nothing to pick from.
source, _ := newLLMSource(t, fake)
if _, err := source.NextAction(context.Background(), stepIndex); !errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want ErrNoAction with no candidates", err)
}
return lastCall(t, source)
}
acted := func(t *testing.T) trace.LLMCall {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshot(t, verifierInstance)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
if _, err := source.NextAction(context.Background(), stepIndex); err != nil {
t.Fatalf("NextAction: %v", err)
}
return lastCall(t, source)
}
skip, decline, pick := guardSkipped(t), declined(t), acted(t)
if skip.Outcome != trace.LLMOutcomeEchoMismatch {
t.Errorf("guard-skipped outcome = %q, want %q", skip.Outcome, trace.LLMOutcomeEchoMismatch)
}
if decline.Outcome != trace.LLMOutcomeNoCandidates {
t.Errorf("declined outcome = %q, want %q", decline.Outcome, trace.LLMOutcomeNoCandidates)
}
if pick.Outcome != trace.LLMOutcomeSelected {
t.Errorf("executed outcome = %q, want %q", pick.Outcome, trace.LLMOutcomeSelected)
}
for _, call := range []trace.LLMCall{skip, decline, pick} {
if call.Step != stepIndex {
t.Errorf("record step = %d, want %d so it joins its trace line", call.Step, stepIndex)
}
if call.Timestamp.IsZero() {
t.Error("record carries no timestamp")
}
}
// The guard skip must carry what only the dropped log line used to hold.
if skip.Choice == 0 || skip.EchoedAction != `Tap "Something Else"` || skip.RawResponse == "" {
t.Errorf("guard-skip record = %+v, want the choice, the echo, and the raw response", skip)
}
if len(skip.Candidates) == 0 {
t.Error("guard-skip record must keep the candidate list the mismatch is judged against")
}
// A decline never reached the provider, so it must not look like a call.
if decline.RawResponse != "" || decline.UserPrompt != "" || len(decline.Candidates) != 0 {
t.Errorf("decline record = %+v, want no prompt, candidates, or response", decline)
}
}
// TestLLMCallRecordsCandidateListAsShown covers the companion experiment that
// varies how candidates are labelled: the labels are an independent variable, so
// each call must be recoverable with the exact numbered list it saw.
func TestLLMCallRecordsCandidateListAsShown(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
source.instructions = "hunt for double submits"
pushLLMSnapshot(t, verifierInstance)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
if _, err := source.NextAction(context.Background(), 1); err != nil {
t.Fatalf("NextAction: %v", err)
}
call := lastCall(t, source)
if len(call.Candidates) == 0 {
t.Fatal("no candidate list recorded")
}
for _, candidate := range call.Candidates {
line := fmt.Sprintf("%d. %s", candidate.Index, candidate.Description)
if !strings.Contains(call.UserPrompt, line) {
t.Errorf("recorded candidate %q is not a line of the prompt:\n%s", line, call.UserPrompt)
}
if candidate.Weight > 0 && !strings.Contains(call.UserPrompt, fmt.Sprintf("%s (w%d)", candidate.Description, candidate.Weight)) {
t.Errorf("recorded weight %d for %q is not the weight the prompt showed:\n%s",
candidate.Weight, candidate.Description, call.UserPrompt)
}
}
numberedLine := regexp.MustCompile(`(?m)^\d+\. `)
if shown := len(numberedLine.FindAllString(call.UserPrompt, -1)); shown != len(call.Candidates) {
t.Errorf("prompt showed %d numbered lines but %d candidates were recorded", shown, len(call.Candidates))
}
if got := candidateLabels(call.Candidates); !slices.Contains(got, tap.Label) {
t.Errorf("recorded labels = %v, want the target label %q among them", got, tap.Label)
}
// The system prompt is assembled from a constant plus spec instructions, so
// the record has to be the assembled text, not the constant.
if call.SystemPrompt != source.systemPrompt() {
t.Errorf("recorded system prompt = %q, want the assembled prompt", call.SystemPrompt)
}
if !strings.Contains(call.SystemPrompt, "hunt for double submits") {
t.Error("recorded system prompt dropped the spec instructions")
}
if call.Model != "test/model" {
t.Errorf("recorded model = %q, want test/model", call.Model)
}
}
func candidateLabels(candidates []trace.LLMCandidate) []string {
labels := make([]string, 0, len(candidates))
for _, candidate := range candidates {
labels = append(labels, candidate.Label)
}
return labels
}
func TestLLMCallRecordsSetupDrivenStep(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSourceWithSpec(t, fake, llmSetupFixtureSpec)
pushLLMSnapshot(t, verifierInstance)
action, err := source.NextAction(context.Background(), 2)
if err != nil {
t.Fatalf("NextAction: %v", err)
}
if action.On != "id:Submit" {
t.Fatalf("action = %+v, want the setup tap on id:Submit", action)
}
call := lastCall(t, source)
if call.Outcome != trace.LLMOutcomeSetupAction {
t.Errorf("outcome = %q, want %q", call.Outcome, trace.LLMOutcomeSetupAction)
}
if call.UserPrompt != "" || call.RawResponse != "" || call.TotalTokens != 0 {
t.Errorf("setup-driven step = %+v, want no model call recorded", call)
}
if source.lastSource != "" {
t.Errorf("lastSource = %q, want empty: setup chose the action, not the model", source.lastSource)
}
}
func TestLLMCallFileRecordsUsageLatencyAndScreenshot(t *testing.T) {
fake := newFakeOpenRouter(t)
fake.usage = llmclient.Usage{PromptTokens: 1200, CompletionTokens: 34, TotalTokens: 1234}
fake.servedModel = "vendor/model-2026-05"
fake.delay = 15 * time.Millisecond
source, verifierInstance := newLLMSource(t, fake)
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
source.recorder = writer
pushLLMSnapshotAtStep(t, verifierInstance, 4)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
if _, err := source.NextAction(context.Background(), 4); err != nil {
t.Fatalf("NextAction: %v", err)
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
calls := readLLMCalls(t, directory)
if len(calls) != 1 {
t.Fatalf("recorded %d calls, want 1", len(calls))
}
call := calls[0]
if call.PromptTokens != 1200 || call.CompletionTokens != 34 || call.TotalTokens != 1234 {
t.Errorf("tokens = %d/%d/%d, want 1200/34/1234",
call.PromptTokens, call.CompletionTokens, call.TotalTokens)
}
if call.LatencyMillis < 15 {
t.Errorf("latency = %dms, want at least the server's 15ms", call.LatencyMillis)
}
if call.ServedModel != "vendor/model-2026-05" {
t.Errorf("served model = %q, want the id the provider reported", call.ServedModel)
}
if want := trace.ScreenshotReference(4); call.Screenshot != want {
t.Errorf("screenshot = %q, want %q", call.Screenshot, want)
}
if call.Reasoning == "" || call.EchoedAction != tap.Description {
t.Errorf("record = %+v, want the parsed reasoning and echo", call)
}
}
// TestLLMCallScreenshotNamesObservedStep guards the one case where the image
// sent is not the current step's: the runner skips PushSnapshot on a
// transitional observation, so the model still sees the last observed screen.
func TestLLMCallScreenshotNamesObservedStep(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSource(t, fake)
pushLLMSnapshotAtStep(t, verifierInstance, 4)
tap := candidateByKind(t, mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText), verifier.ActionKindTap)
fake.choice = tap.Index
fake.chosenAction = tap.Description
if _, err := source.NextAction(context.Background(), 6); err != nil {
t.Fatalf("NextAction: %v", err)
}
call := lastCall(t, source)
if call.Step != 6 {
t.Errorf("record step = %d, want the current step 6", call.Step)
}
if want := trace.ScreenshotReference(4); call.Screenshot != want {
t.Errorf("screenshot = %q, want %q: the image sent was step 4's", call.Screenshot, want)
}
}
func TestDownscalePNGShrinksLongEdge(t *testing.T) {
large := image.NewRGBA(image.Rect(0, 0, 2048, 1024))
var buffer bytes.Buffer
@@ -576,3 +1269,35 @@ func tinyPNG(t *testing.T) []byte {
}
return buffer.Bytes()
}
// llmSamplerFixtureSpec drives the model policy over an authored leaf that
// samples one of three targets, which is the shape the seeded picker draws from
// and the model policy cannot.
const llmSamplerFixtureSpec = `
import { actions, from, llm, Tap, always } from "@sanderling/spec";
globalThis.properties = { ok: always(() => true) };
const targets = from(["id:Submit", "id:Name"]);
globalThis.actions = actions(() => [Tap({ on: targets.generate() })]);
globalThis.generator = llm({ model: "test/model" });
`
// TestLLMSourceRefusesAMultiItemAuthoredSampler: the step must fail the run, not
// skip. A skip would leave the model quietly fuzzing a spec whose authored
// targets it can never reach past the first, which is the comparison the seeded
// arm is measured against.
func TestLLMSourceRefusesAMultiItemAuthoredSampler(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSourceWithSpec(t, fake, llmSamplerFixtureSpec)
pushLLMSnapshot(t, verifierInstance)
_, err := source.NextAction(context.Background(), 1)
if err == nil || errors.Is(err, verifier.ErrNoAction) {
t.Fatalf("NextAction err = %v, want the run to stop on a sampler the model cannot draw", err)
}
if !strings.Contains(err.Error(), "targets.generate()") {
t.Errorf("error does not name the offending leaf: %v", err)
}
if outcome := lastCall(t, source).Outcome; outcome != trace.LLMOutcomeCandidatesFailed {
t.Errorf("recorded outcome = %q, want %q", outcome, trace.LLMOutcomeCandidatesFailed)
}
}
+82
View File
@@ -0,0 +1,82 @@
package runner
import (
"bufio"
"context"
"encoding/json"
"os"
"path/filepath"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/trace"
)
// navigatingDriver reports one navigation per step, the way a page that submits
// a form or reloads does.
type navigatingDriver struct {
*mockdriver.Driver
url string
}
func (d *navigatingDriver) Navigations(context.Context) ([]driver.Navigation, error) {
return []driver.Navigation{{URL: d.url, UnixMillis: 1700000000000}}, nil
}
// A run whose app replaced its own document has to say so on the step it
// happened. Without it the trace shows a generator repeating one action and no
// reason for it, and an analysis cannot tell that from a seed that chose badly.
func TestRunner_TheTraceRecordsThatThePageNavigated(t *testing.T) {
state := newHarness(t)
const url = "http://127.0.0.1/index.html?"
navigating := &navigatingDriver{Driver: state.mock, url: url}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: navigating,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
if err := state.writer.Close(); err != nil {
t.Fatalf("close trace: %v", err)
}
file, err := os.Open(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
recorded := 0
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var step struct {
Index int `json:"step"`
Navigations []trace.Navigation `json:"navigations"`
}
if err := json.Unmarshal(scanner.Bytes(), &step); err != nil {
t.Fatalf("trace line decode: %v", err)
}
for _, navigation := range step.Navigations {
recorded++
if navigation.URL != url {
t.Errorf("step %d records navigation to %q, want %q", step.Index, navigation.URL, url)
}
if navigation.UnixMillis == 0 {
t.Errorf("step %d records a navigation with no timestamp", step.Index)
}
}
}
if recorded == 0 {
t.Fatal("the page navigated on every step and the trace holds no record of it")
}
}
+243
View File
@@ -0,0 +1,243 @@
package runner
import (
"context"
"fmt"
"strings"
"testing"
"time"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// typedCredential stands in for what a login setup types. Nothing the run
// records or sends may carry it.
const typedCredential = "fixture-passphrase-9f21"
// iosLoginTreeJSON is the shape internal/driver/ioscompanion produces for a
// login form: every editable field reports `secure`, so a field carrying
// `secure:false` is positively known not to be a credential entry.
const iosLoginTreeJSON = `{
"attributes": {"bounds": "[0,0,390,844]", "class": "Window"},
"children": [
{"attributes": {"identifier": "LoginEmail", "class": "TextField", "hintText": "Email", "bounds": "[20,200,370,244]"},
"editable": true, "enabled": true, "secure": false, "children": []},
{"attributes": {"identifier": "LoginPassword", "class": "SecureTextField", "hintText": "Password", "bounds": "[20,260,370,304]"},
"editable": true, "enabled": true, "secure": true, "children": []}
]
}`
// webLoginTreeJSON is the same form as the chrome dump reports it.
const webLoginTreeJSON = `{
"attributes": {"bounds": "[0,0,1280,800]", "tag": "html"},
"children": [
{"attributes": {"resource-id": "login-email", "tag": "input", "hintText": "Email", "bounds": "[20,200,400,240]"},
"editable": true, "enabled": true, "secure": false, "children": []},
{"attributes": {"resource-id": "login-password", "tag": "input", "hintText": "Password", "bounds": "[20,260,400,300]"},
"editable": true, "enabled": true, "secure": true, "children": []}
]
}`
// androidLoginTreeJSON is the same form on Android, where the native tree
// mapper drops uiautomator's password attribute and neither field carries the
// fact.
const androidLoginTreeJSON = `{
"attributes": {"bounds": "[0,0,1080,2340]"},
"children": [
{"attributes": {"resource-id": "login_email", "class": "android.widget.EditText", "hintText": "Email", "bounds": "[20,200,1060,300]"},
"enabled": true, "children": []},
{"attributes": {"resource-id": "login_password", "class": "android.widget.EditText", "hintText": "Password", "bounds": "[20,320,1060,420]"},
"enabled": true, "children": []}
]
}`
func typeInto(selector string) verifier.Action {
return verifier.Action{Kind: verifier.ActionKindInputText, On: selector, Text: typedCredential}
}
func TestTraceActionForRedactsTypedValuesTheTargetCannotClear(t *testing.T) {
for _, testCase := range []struct {
name string
treeJSON string
selector string
}{
{"ios secure field", iosLoginTreeJSON, "id:LoginPassword"},
{"web secure field", webLoginTreeJSON, "id:login-password"},
{"android field reported as neither", androidLoginTreeJSON, "id:login_email"},
{"android password field", androidLoginTreeJSON, "id:login_password"},
} {
t.Run(testCase.name, func(t *testing.T) {
traceAction := traceActionFor(typeInto(testCase.selector), mustParseTree(t, testCase.treeJSON))
if traceAction.Text != verifier.RedactedInputText {
t.Errorf("trace action text = %q, want it redacted", traceAction.Text)
}
if traceAction.Selector != testCase.selector {
t.Errorf("trace selector = %q, want %q so the record still shows which field was typed into",
traceAction.Selector, testCase.selector)
}
})
}
}
func TestTraceActionForKeepsTypedValuesForAFieldReportedNotSecure(t *testing.T) {
for _, testCase := range []struct {
name string
treeJSON string
selector string
}{
{"ios", iosLoginTreeJSON, "id:LoginEmail"},
{"web", webLoginTreeJSON, "id:login-email"},
} {
t.Run(testCase.name, func(t *testing.T) {
const address = "[email protected]"
action := verifier.Action{Kind: verifier.ActionKindInputText, On: testCase.selector, Text: address}
traceAction := traceActionFor(action, mustParseTree(t, testCase.treeJSON))
if traceAction.Text != address {
t.Errorf("trace action text = %q, want the typed value on a field reported not secure", traceAction.Text)
}
})
}
}
// The redaction is a rendering rule, never a change to what is dispatched: the
// app has to receive the keystrokes a user would have produced.
func TestApplyActionTypesTheRealValueIntoEveryField(t *testing.T) {
for _, testCase := range []struct {
name string
treeJSON string
selector string
}{
{"ios secure field", iosLoginTreeJSON, "id:LoginPassword"},
{"web secure field", webLoginTreeJSON, "id:login-password"},
{"android field", androidLoginTreeJSON, "id:login_password"},
} {
t.Run(testCase.name, func(t *testing.T) {
fastFocusSettle(t)
driverMock := mockdriver.New()
mustDispatch(t, driverMock, typeInto(testCase.selector), mustParseTree(t, testCase.treeJSON))
typed := ""
for _, recorded := range driverMock.Actions() {
if recorded.Kind == mockdriver.ActionInputText {
typed = recorded.Text
}
}
if typed != typedCredential {
t.Errorf("driver received %q, want the real typed value", typed)
}
})
}
}
// lastActionExtractorSpec is what examples/folio/sanderling/spec.ts does with
// the previous step's action: it extracts state.lastAction whole. Extractor
// values are written to the trace as extractor_changes, so a typed value that
// reaches state.lastAction.text lands in the run directory through the spec
// rather than through the runner.
const lastActionExtractorSpec = `
import { actions, always, extract, InputText } from "@sanderling/spec";
const reported = extract("lastAction", state => state.lastAction);
globalThis.properties = {
theActionReachesTheSpec: always(() => reported.current !== undefined),
};
globalThis.actions = actions(() => [InputText({ into: "%s", text: "` + typedCredential + `" })]);
`
func TestTheTraceNeverCarriesATypedSecretThroughALastActionExtractor(t *testing.T) {
for _, testCase := range []struct {
name string
treeJSON string
selector string
}{
{"ios secure field", iosLoginTreeJSON, "id:LoginPassword"},
{"web secure field", webLoginTreeJSON, "id:login-password"},
{"android field reported as neither", androidLoginTreeJSON, "id:login_email"},
} {
t.Run(testCase.name, func(t *testing.T) {
fastFocusSettle(t)
state := newHarnessWithSpec(t, fmt.Sprintf(lastActionExtractorSpec, testCase.selector))
state.mock.HierarchyJSON = testCase.treeJSON
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
steps := traceSteps(t, state.writer.Directory())
if len(steps) < 2 {
t.Fatalf("trace holds %d step(s), want the step that reports the action back", len(steps))
}
change, ok := steps[1].ExtractorChanges["lastAction"]
if !ok {
t.Fatalf("step 2 recorded no reading of state.lastAction: %+v", steps[1])
}
if strings.Contains(string(change.Curr), typedCredential) {
t.Errorf("the trace carries the typed value through state.lastAction: %s", change.Curr)
}
if !strings.Contains(string(change.Curr), verifier.RedactedInputText) {
t.Errorf("state.lastAction reported %s, want the typed value redacted in place",
change.Curr)
}
})
}
}
const secureLoginSpec = `
import { llm, always, actions, InputText, taps, typing, weighted } from "@sanderling/spec";
globalThis.properties = { ok: always(() => true) };
globalThis.setup = actions(() => {
if (globalThis.__typed) return [];
globalThis.__typed = true;
return [InputText({ into: "id:LoginPassword", text: "` + typedCredential + `" })];
});
globalThis.actions = weighted([1, taps], [1, typing]);
globalThis.generator = llm({ model: "test/model" });
`
// The recorded call is llm-calls.jsonl: its user prompt carries the
// recent-action memory and its candidates carry the numbered list, which is
// where a login run put the account password in cleartext beside a screenshot
// of the same screen.
func TestLLMCallRecordKeepsTypedSecretsOutOfThePromptAndCandidates(t *testing.T) {
fake := newFakeOpenRouter(t)
source, verifierInstance := newLLMSourceWithSpec(t, fake, secureLoginSpec)
pushSnapshotTree(t, verifierInstance, iosLoginTreeJSON)
action, err := source.NextAction(context.Background(), 0)
if err != nil {
t.Fatalf("setup NextAction: %v", err)
}
if action.Text != typedCredential {
t.Fatalf("setup action text = %q, want the real value dispatched to the app", action.Text)
}
pushSnapshotTree(t, verifierInstance, iosLoginTreeJSON)
fake.choice = 1
fake.chosenAction = mustCandidates(t, verifierInstance, verifier.LabelSourceVisibleText)[0].Description
if _, err := source.NextAction(context.Background(), 1); err != nil {
t.Fatalf("llm NextAction: %v", err)
}
call := lastCall(t, source)
if strings.Contains(call.UserPrompt, typedCredential) {
t.Error("the recorded user prompt carries the typed value")
}
if !strings.Contains(call.UserPrompt, "InputText") {
t.Errorf("the recorded user prompt lost the recent-action memory entirely: %q", call.UserPrompt)
}
for _, candidate := range call.Candidates {
if strings.Contains(candidate.Description, typedCredential) {
t.Errorf("recorded candidate %d carries the typed value: %q", candidate.Index, candidate.Description)
}
}
}
+552 -98
View File
@@ -45,6 +45,11 @@ type Options struct {
// spec's generator = llm({...}) config; anything else (the default) uses the
// seeded weighted picker. Both draw from the same actionsRoot candidate set.
Generator string
// LabelSource selects how candidates are named to the model picker
// (verifier.LabelSourceVisibleText or verifier.LabelSourceResourceID). The
// seeded picker selects by index and never reads a label, so this reaches
// the model picker only.
LabelSource string
}
type Summary struct {
@@ -61,6 +66,25 @@ type Summary struct {
// could not dispatch, deduped, so the report can flag a spec exercising
// gestures this target does not support.
UnsupportedVerbs []string
// SkippedActions counts, by reason, the actions a step chose that never
// reached the app. Without it a run that dropped most of what it generated
// reads exactly like one that exercised it: the reasons reach the trace and
// a warn line, and nothing else.
SkippedActions map[string]int
// FailedObservations counts the steps whose device read produced no tree at
// all. Such a step verifies nothing, so a run that failed every observation
// finishes with no violations and reads as a clean one.
FailedObservations int
// DispatchedActions counts the steps whose chosen action reached the driver.
// A run at zero never touched the app, whatever its step count says, so its
// empty violation list is the reading of an instrument that measured nothing.
DispatchedActions int
// GeneratorActions counts the dispatched actions the generator chose. The
// spec's setup drives the app into its starting position before the
// generator is consulted, so a run at zero here explored nothing however
// many actions its login fired. Both generators separate the two the same
// way, by the producer each action names on the trace.
GeneratorActions int
}
type ViolationRecord struct {
@@ -84,7 +108,19 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// Gate on the app actually being on top before acting, so the first
// action never fires against a leftover screen or a system dialog. Done
// before the deadline is set so the settle time does not eat the run.
waitForForeground(ctx, options, logger)
//
// A run that cannot get its app on screen has no preconditions to explore
// from: every step after this would observe some other app and every
// property would judge it, so the run ends here and the trace records why.
if !waitForForeground(ctx, options, logger) {
if err := recordPreconditionFailure(options); err != nil {
return Summary{}, err
}
return Summary{}, ForegroundNotReachedError{
BundleID: options.BundleID,
Waited: foregroundReadyBudget,
}
}
// Pick the action and extractor sources once from the driver's
// capabilities so the step loop runs one uniform path with no per-step
@@ -94,6 +130,8 @@ func Run(ctx context.Context, options Options) (Summary, error) {
return Summary{}, err
}
_, pageExtractors := extractorSource.(webSource)
exceptionReporter, _ := options.Driver.(driver.ExceptionReporter)
navigationReporter, _ := options.Driver.(driver.NavigationReporter)
rereadHierarchy := driverIsAndroid(ctx, options, logger)
summary := Summary{StartTime: time.Now()}
@@ -122,7 +160,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// readings" and the runner has no business saying that: the action ran,
// and a property told otherwise convicts the app of an effect with no
// cause. See foreground_guard_last_action_test.go.
guard := ensureForeground(ctx, options, logger, stepIndex)
guard, inScope := ensureForeground(ctx, options, logger, stepIndex)
if lastAction != nil {
switch guard {
case foregroundRelaunched:
@@ -146,8 +184,10 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// gctx is bound to the errgroup so a returned error (or outer
// cancellation) propagates to every sibling read rather than leaving
// one blocked on a hung device.
g, gctx := errgroup.WithContext(ctx)
// one blocked on a hung device, and to observationTimeout so a read
// that never answers ends the step instead of the run.
observeCtx, observeCancel := context.WithTimeout(ctx, observationTimeout)
g, gctx := errgroup.WithContext(observeCtx)
si := stepIndex
// fetchSyncedState issues a single Snapshot RPC so hierarchy and
// screenshot describe the same frame, then re-fetches the pair
@@ -169,12 +209,18 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// All goroutines write to local variables and return nil, so the Wait
// error is always nil; ignored intentionally.
_ = g.Wait()
observeCancel()
navigations := collectNavigations(ctx, navigationReporter, logger, stepIndex)
observationError := ""
if hierarchyErr != nil {
if isWDADrop(hierarchyErr) {
return summary, fmt.Errorf("WDA connection permanently lost at step %d - re-run the test: %w", stepIndex, hierarchyErr)
}
logger.Warn("hierarchy fetch failed", "step", stepIndex, "err", hierarchyErr)
observationError = hierarchyErr.Error()
summary.FailedObservations++
}
treeSize := 0
if tree != nil {
@@ -208,6 +254,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
var violations []string
var extractorChanges map[string]trace.ExtractorChange
var witnesses map[string]trace.Witness
var exceptions []verifier.Exception
skippedVerification := false
if !transitional {
// The page-side extractors evaluate only on steps the verifier will
@@ -225,7 +272,9 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// lastAction and logs are the same values PushSnapshot hands the
// goja state below: the two engines evaluate this step against one
// action and one set of log entries.
v8Overrides, overridesErr := extractorSource.ExtractorOverrides(ctx, lastAction, logs)
overridesCtx, overridesCancel := context.WithTimeout(ctx, observationTimeout)
v8Overrides, overridesErr := extractorSource.ExtractorOverrides(overridesCtx, lastAction, logs)
overridesCancel()
if overridesErr != nil {
// Not a warning. Without the page's values this step's
// extractors keep goja's dump-derived readings while the
@@ -234,6 +283,12 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// wrong.
return summary, fmt.Errorf("step %d extractor overrides: %w", stepIndex, overridesErr)
}
exceptions = collectExceptions(
ctx,
exceptionReporter,
logger,
stepIndex,
)
if err := options.Verifier.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
ScreenshotPNG: screenshotPNG,
@@ -242,6 +297,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
StepIndex: stepIndex,
RunStart: summary.StartTime,
Logs: logs,
Exceptions: exceptions,
}); err != nil {
return summary, fmt.Errorf("step %d push: %w", stepIndex, err)
}
@@ -303,7 +359,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
nextErr := verifier.ErrNoAction
var traceAction *trace.Action
if !held {
nextAction, nextErr = actionSource.NextAction(ctx)
nextAction, nextErr = actionSource.NextAction(ctx, stepIndex)
if nextErr == nil {
traceAction = traceActionFor(nextAction, tree)
stampActionSource(traceAction, actionSource)
@@ -318,7 +374,15 @@ func Run(ctx context.Context, options Options) (Summary, error) {
}
applySkipped := held
if nextErr == nil && !appIsForeground(ctx, options) {
var actionSkipped actionSkipReason
if nextErr != nil && !held {
// The source was asked and handed nothing back. Recorded like every
// other non-action, because unrecorded it is the one that survives a
// whole run: a picker declining on all 200 steps leaves a trace, a
// summary and an exit status a run that exercised all 200 produces.
actionSkipped = actionSkippedNoActionProduced
lastAction = nil
} else if nextErr == nil && !appIsForeground(ctx, options) {
// The app left the foreground between observe and apply (a prior
// action's gesture settling late, or an async navigation). The
// chosen action's coordinates reference a tree that no longer
@@ -327,9 +391,36 @@ func Run(ctx context.Context, options Options) (Summary, error) {
logger.Warn("app not in foreground at action time; skipping (relaunch next step)",
"step", stepIndex, "action", nextAction.Kind)
applySkipped = true
actionSkipped = actionSkippedForeground
lastAction = nil
} else if nextErr == nil {
if err := applyAction(ctx, options.Driver, nextAction, tree); err != nil {
applyCtx, applyCancel := context.WithTimeout(ctx, applyBound(nextAction))
notDispatched, err := applyAction(applyCtx, options.Driver, nextAction, tree)
applyCancel()
if errors.Is(err, driver.ErrGestureUndelivered) {
// The gesture reached no element, so the app cannot have
// responded to it and the screen is still the one already
// verified. The device is healthy, so the step is neither
// transitional nor part of the apply-failure streak; it records
// that the action landed on nothing, which is what separates it
// from an action the app received and ignored.
logger.Warn("gesture reached no element",
"step", stepIndex, "action", nextAction.Kind, "err", err)
applySkipped = true
actionSkipped = actionSkippedGestureUndelivered
lastAction = nil
} else if errors.Is(err, driver.ErrSelectorMatchedNothing) {
// The selector named no element, so no point was resolved and
// nothing was dispatched. The screen is the one already
// verified and the device is healthy, so this is the same
// non-action the runner records when it cannot resolve a
// selector itself, not a device fault worth a failure streak.
logger.Warn("selector matched no element",
"step", stepIndex, "action", nextAction.Kind, "err", err)
applySkipped = true
actionSkipped = actionSkippedUnresolvedSelector
lastAction = nil
} else if err != nil {
if isWDADrop(err) {
return summary, fmt.Errorf("step %d: the iOS XCTest runner could not be restarted - re-run the test: %w", stepIndex, err)
}
@@ -346,7 +437,12 @@ func Run(ctx context.Context, options Options) (Summary, error) {
if consecutiveApplyFailures >= maxConsecutiveApplyFailures {
return summary, fmt.Errorf("step %d apply: %d consecutive failures; the device is not recovering: %w", stepIndex, consecutiveApplyFailures, err)
}
logger.Warn("apply error; marking step transitional", "step", stepIndex, "err", err)
actionSkipped = actionSkippedApplyError
if errors.Is(applyCtx.Err(), context.DeadlineExceeded) {
actionSkipped = actionSkippedApplyTimeout
}
logger.Warn("apply error; marking step transitional",
"step", stepIndex, "reason", actionSkipped, "err", err)
transitional = true
applySkipped = true
// The error says the call failed, not that the gesture never
@@ -354,16 +450,25 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// the effect committed. Reporting no action here would let a
// property convict the app for an effect with no cause, so the
// action is reported with its fate unknown instead.
unconfirmed := nextAction
unconfirmed := verifier.RecordedAction(nextAction, tree)
lastAction = &unconfirmed
} else if notDispatched != "" {
// The action was chosen but nothing reached the driver, so the
// screen is exactly the one already verified: the step stays
// non-transitional and only records why it acted on nothing.
// The apply-failure streak is left alone; a step that never
// reached the device says nothing about the device's health.
logger.Warn("action not dispatched",
"step", stepIndex, "action", nextAction.Kind, "reason", notDispatched)
applySkipped = true
actionSkipped = notDispatched
lastAction = nil
} else {
consecutiveApplyFailures = 0
applied := nextAction
applied := verifier.RecordedAction(nextAction, tree)
applied.Applied = true
lastAction = &applied
}
} else if !held {
lastAction = nil
}
// A held step leaves lastAction alone on purpose: nothing ran here, and
// the action it points at is still the one the next verified step has to
@@ -374,18 +479,36 @@ func Run(ctx context.Context, options Options) (Summary, error) {
Timestamp: stepStart,
Screen: screen,
NextAction: traceAction,
Logs: traceLogs(logs),
Exceptions: traceExceptions(exceptions),
Navigations: navigations,
Violations: violations,
Hierarchy: tree,
Residuals: residuals,
Metrics: metrics,
ExtractorChanges: extractorChanges,
Transitional: transitional,
ObservationError: observationError,
ActionSkipped: string(actionSkipped),
SkippedVerification: skippedVerification,
Witnesses: witnesses,
PreconditionFailure: preconditionFailure(inScope),
}
if err := options.TraceWriter.WriteStep(step); err != nil {
return summary, fmt.Errorf("step %d trace: %w", stepIndex, err)
}
if actionSkipped != "" {
if summary.SkippedActions == nil {
summary.SkippedActions = map[string]int{}
}
summary.SkippedActions[string(actionSkipped)]++
}
if nextErr == nil && !applySkipped {
summary.DispatchedActions++
if generatorChoseAction(traceAction) {
summary.GeneratorActions++
}
}
summary.Steps = stepIndex
if len(violations) > 0 {
summary.Violations = append(summary.Violations, violationRecords(violations, witnesses, stepIndex)...)
@@ -450,7 +573,8 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// record, and any unsupported verbs. The wall-clock duration is excluded so the
// output is deterministic and snapshot-testable; the CLI prints it separately.
func RenderSummary(w io.Writer, summary Summary, platform string) {
fmt.Fprintf(w, "\nrun complete: %d steps\n", summary.Steps)
fmt.Fprintf(w, "\nrun complete: %d steps, %d driven by the generator\n",
summary.Steps, summary.GeneratorActions)
if len(summary.Violations) == 0 {
fmt.Fprintln(w, "no violations.")
} else {
@@ -459,6 +583,21 @@ func RenderSummary(w io.Writer, summary Summary, platform string) {
fmt.Fprintf(w, " step %d: %v\n", violation.StepIndex, violation.Properties)
}
}
if len(summary.SkippedActions) > 0 {
total := 0
byReason := make([]string, 0, len(summary.SkippedActions))
for _, reason := range slices.Sorted(maps.Keys(summary.SkippedActions)) {
total += summary.SkippedActions[reason]
byReason = append(byReason,
fmt.Sprintf("%s %d", reason, summary.SkippedActions[reason]))
}
fmt.Fprintf(w, "%d action(s) never reached the app: %s\n",
total, strings.Join(byReason, ", "))
}
if summary.FailedObservations > 0 {
fmt.Fprintf(w, "%d step(s) observed nothing: the device state could not be read\n",
summary.FailedObservations)
}
if summary.SkippedVerification > 0 {
fmt.Fprintf(w, "%d step(s) judged by nothing: the screen was still moving when it was read\n",
summary.SkippedVerification)
@@ -528,22 +667,24 @@ const (
// ensureForeground keeps the app under test in the foreground. When the driver
// can report the foreground app and it no longer matches the bundle under test,
// the app is relaunched. Reports what it did so the caller can pass that on to
// the spec through the previous action. Drivers without ForegroundChecker (web,
// iOS) are a no-op.
// the spec through the previous action, and whether the app is in front at all:
// a false there is a step whose observation is not of the app under test, which
// the step records so a trace cannot pass it off as exploration. Drivers without
// ForegroundChecker (web, iOS) are a no-op.
func ensureForeground(
ctx context.Context,
options Options,
logger *slog.Logger,
stepIndex int,
) foregroundGuard {
) (foregroundGuard, bool) {
checker, ok := options.Driver.(driver.ForegroundChecker)
if !ok || options.BundleID == "" {
return foregroundIntact
return foregroundIntact, true
}
foreground, err := checker.ForegroundApp(ctx)
if err != nil {
logger.Warn("foreground check failed", "step", stepIndex, "err", err)
return foregroundIntact
return foregroundIntact, true
}
if foreground != "" && foreground != options.BundleID {
logger.Warn("app left foreground; relaunching",
@@ -555,8 +696,7 @@ func ensureForeground(
// InputText). awaitForeground re-checks the foreground and focused
// window, so it never acts outside the app no matter how slow the
// relaunch settles.
awaitForeground(ctx, options, logger, stepIndex)
return foregroundRelaunched
return foregroundRelaunched, awaitForeground(ctx, options, logger, stepIndex)
}
// The app is the resumed activity, but a system overlay can still own the
// focused window while the app stays resumed: a fuzzer swipe starting in the
@@ -566,15 +706,15 @@ func ensureForeground(
// the app again.
focusChecker, hasFocus := options.Driver.(driver.FocusedWindowChecker)
if !hasFocus {
return foregroundIntact
return foregroundIntact, true
}
focused, err := focusChecker.FocusedWindowApp(ctx)
if err != nil {
logger.Warn("focus check failed", "step", stepIndex, "err", err)
return foregroundIntact
return foregroundIntact, true
}
if focused == "" || focused == options.BundleID {
return foregroundIntact
return foregroundIntact, true
}
logger.Warn("system window obscuring app; dismissing",
"step", stepIndex, "focused", focused, "want", options.BundleID)
@@ -582,7 +722,7 @@ func ensureForeground(
logger.Warn("dismiss overlay failed", "step", stepIndex, "err", err)
}
settleForForeground(ctx, options)
return foregroundOverlayDismissed
return foregroundOverlayDismissed, true
}
// appIsForeground reports whether the app under test currently owns the
@@ -617,10 +757,44 @@ func appIsForeground(ctx context.Context, options Options) bool {
return focused == options.BundleID
}
// foregroundReadyAttempts bounds how many times waitForForeground tries to
// bring the app forward before the first step, so a stuck system dialog can
// never hang the run.
const foregroundReadyAttempts = 8
// foregroundReadyBudget bounds how long awaitForeground waits for the app to be
// on screen, so a stuck system dialog can never hang the run. It is wall-clock
// time and not a count of polls because a poll costs whatever the driver's idle
// wait happens to take, and that is a property of the device: on an API 34
// emulator settleForForeground returned in ~100ms, so eight polls gave up 1.2s
// into a launch whose window drew at ~1.9s, while on API 36 the same eight polls
// spanned 3s and cleared the same launch. Counted in polls, the gate's verdict
// describes the device it ran on rather than the app it was watching.
var foregroundReadyBudget = 15 * time.Second
// foregroundPollInterval floors how often the gate re-reads the device.
// settleForForeground returns the moment the device reports idle, which during a
// launch animation is immediately, so without a floor the gate would spin on adb
// for the whole budget.
var foregroundPollInterval = 250 * time.Millisecond
// preconditionAppNotForeground is the reason a trace step carries when the app
// under test was not in front of it: the startup gate's verdict on step 0, and
// the scope guard's on any later step it could not bring the app back for. It is
// a fixed token so a campaign counts these by decoding the trace rather than by
// grepping a log line.
const preconditionAppNotForeground = "app_not_in_foreground"
// ForegroundNotReachedError reports that the app under test never came to the
// foreground within foregroundReadyBudget, so the run never started. A run that
// ends this way has judged nothing, and reporting it as a clean run would count
// a harness failure as evidence about the app.
type ForegroundNotReachedError struct {
BundleID string
Waited time.Duration
}
func (e ForegroundNotReachedError) Error() string {
return fmt.Sprintf(
"%s never reached the foreground within %s: nothing was observed and no property "+
"judged anything, so this run holds no verdict about the app",
e.BundleID, e.Waited)
}
// focusTapSettle is the pause after tapping a field to focus it, before typing.
// Long enough for focus to land, short enough to avoid the ~500ms-1s full
@@ -638,13 +812,13 @@ var focusTapSettle = 250 * time.Millisecond
// alone lets the first observe read the outgoing app. When the driver can also
// report the focused window, the gate additionally waits for that window to
// name the app, which only happens once it is genuinely drawn.
func waitForForeground(ctx context.Context, options Options, logger *slog.Logger) {
awaitForeground(ctx, options, logger, 0)
func waitForForeground(ctx context.Context, options Options, logger *slog.Logger) bool {
return awaitForeground(ctx, options, logger, 0)
}
// awaitForeground brings the app under test forward when it is not already
// resumed and blocks until its window is actually drawn, bounded by
// foregroundReadyAttempts so a stuck system dialog can never hang the run. It
// foregroundReadyBudget so a stuck system dialog can never hang the run. It
// re-checks the foreground each iteration and only presses back + relaunches
// while the app is genuinely absent, so once the app is resumed it polls the
// focused-window signal instead of mashing back (which would re-exit the app
@@ -652,61 +826,116 @@ func waitForForeground(ctx context.Context, options Options, logger *slog.Logger
// the per-step scope guard so neither lets an observe or action land outside
// the app. Drivers without ForegroundChecker (web) and an unknown foreground
// both skip the gate.
func awaitForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) {
//
// It reports whether the app is on screen. False means the budget ran out with
// the app confirmed absent, which is a precondition the caller has to record:
// every signal that cannot answer the question (no capability, a read error, an
// unknown foreground, a cancelled run) reports true rather than manufacture a
// failure out of a gate that never got to judge.
func awaitForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) bool {
checker, ok := options.Driver.(driver.ForegroundChecker)
if !ok || options.BundleID == "" {
return
return true
}
focusChecker, hasFocus := options.Driver.(driver.FocusedWindowChecker)
for attempt := range foregroundReadyAttempts {
deadline := time.Now().Add(foregroundReadyBudget)
for attempt := 0; ; attempt++ {
if err := ctx.Err(); err != nil {
return
return true // the run is ending on its own; the gate holds no verdict
}
foreground, err := checker.ForegroundApp(ctx)
if err != nil {
logger.Warn("foreground check failed", "step", stepIndex, "err", err)
return
return true
}
if foreground == "" {
return // foreground unknowable (e.g. iOS); don't block the run
return true // foreground unknowable (e.g. iOS); don't block the run
}
if foreground != options.BundleID {
logger.Warn("app not in foreground; bringing it forward",
"step", stepIndex, "foreground", foreground, "want", options.BundleID, "attempt", attempt)
bringToForeground(ctx, options, logger, stepIndex)
if foreground == options.BundleID {
if !hasFocus {
return true // resumed is the app and no finer signal exists
}
focused, err := focusChecker.FocusedWindowApp(ctx)
if err != nil {
logger.Warn("focus check failed", "step", stepIndex, "err", err)
return true
}
if focused == options.BundleID {
return true // window is drawn; safe to observe
}
if !time.Now().Before(deadline) {
break
}
logger.Warn("app resumed but window not yet drawn; waiting",
"step", stepIndex, "focused", focused, "want", options.BundleID, "attempt", attempt)
awaitNextForegroundPoll(ctx, options)
continue
}
if !hasFocus {
return // resumed is the app and no finer signal exists
if !time.Now().Before(deadline) {
break
}
focused, err := focusChecker.FocusedWindowApp(ctx)
if err != nil {
logger.Warn("focus check failed", "step", stepIndex, "err", err)
return
}
if focused == options.BundleID {
return // window is drawn; safe to observe
}
logger.Warn("app resumed but window not yet drawn; waiting",
"step", stepIndex, "focused", focused, "want", options.BundleID, "attempt", attempt)
settleForForeground(ctx, options)
logger.Warn("app not in foreground; bringing it forward",
"step", stepIndex, "foreground", foreground, "want", options.BundleID, "attempt", attempt)
bringToForeground(ctx, options, logger, stepIndex)
awaitNextForegroundPoll(ctx, options)
}
logger.Warn("app never reached foreground", "step", stepIndex,
"want", options.BundleID, "waited", foregroundReadyBudget)
return false
}
// awaitNextForegroundPoll waits out one poll interval before the gate re-reads
// the device: one settle, then whatever is left of foregroundPollInterval, so a
// driver whose idle wait returns immediately cannot turn the gate into a spin.
func awaitNextForegroundPoll(ctx context.Context, options Options) {
start := time.Now()
settleForForeground(ctx, options)
remaining := foregroundPollInterval - time.Since(start)
if remaining <= 0 {
return
}
timer := time.NewTimer(remaining)
defer timer.Stop()
select {
case <-ctx.Done():
case <-timer.C:
}
logger.Warn("app never reached foreground; proceeding anyway",
"step", stepIndex, "want", options.BundleID)
}
// bringToForeground returns the app under test to the foreground. It first
// presses BACK to dismiss any modal system dialog (a relaunch alone does not
// close one), then relaunches and waits for the UI to settle.
// close one), then relaunches. The caller waits out the poll interval before
// looking again, so a relaunch that fails outright cannot spin the gate.
func bringToForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) {
if err := options.Driver.PressKey(ctx, "back"); err != nil {
logger.Warn("dismiss key before relaunch failed", "step", stepIndex, "err", err)
}
if err := options.Driver.Launch(ctx, options.BundleID, false, nil); err != nil {
logger.Warn("relaunch failed", "step", stepIndex, "err", err)
return
}
settleForForeground(ctx, options)
}
// preconditionFailure names what a step could not assume, empty when it could.
func preconditionFailure(inScope bool) string {
if inScope {
return ""
}
return preconditionAppNotForeground
}
// recordPreconditionFailure writes the startup gate's verdict to the trace as
// step 0, the step index no observation ever uses, so a run that never started
// is countable off the trace instead of off a warn line in a log.
func recordPreconditionFailure(options Options) error {
step := trace.Step{
Index: 0,
Timestamp: time.Now(),
PreconditionFailure: preconditionAppNotForeground,
}
if err := options.TraceWriter.WriteStep(step); err != nil {
return fmt.Errorf("write precondition failure to trace: %w", err)
}
return nil
}
// settleForForeground waits one idle window for the UI to settle, bounding the
@@ -717,34 +946,32 @@ func settleForForeground(ctx context.Context, options Options) {
cancel()
}
func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.Action, tree *hierarchy.Tree) error {
// applyAction dispatches one chosen action to the driver. The returned reason is
// empty exactly when the driver was called; a non-empty reason means nothing was
// dispatched and names why, so the step can record that it acted on nothing
// instead of showing a next_action that looks executed.
func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.Action, tree *hierarchy.Tree) (actionSkipReason, error) {
switch action.Kind {
case verifier.ActionKindTap:
x, y, ok := resolveCoordinates(action, tree)
if !ok {
if action.On == "" {
return nil
}
return drv.TapSelector(ctx, action.On)
return "", drv.TapSelector(ctx, action.On)
}
return drv.Tap(ctx, x, y)
return "", drv.Tap(ctx, x, y)
case verifier.ActionKindDoubleTap:
x, y, ok := resolveCoordinates(action, tree)
if !ok {
if action.On == "" {
return nil
}
return drv.DoubleTapSelector(ctx, action.On)
return "", drv.DoubleTapSelector(ctx, action.On)
}
return drv.DoubleTap(ctx, x, y)
return "", drv.DoubleTap(ctx, x, y)
case verifier.ActionKindLongPress:
x, y, ok := resolveCoordinates(action, tree)
if !ok {
// No long-press-by-selector RPC exists, so an unresolved target is
// nothing we can dispatch; skip rather than error.
return nil
// No long-press-by-selector RPC exists, so a selector that resolves
// to no coordinates is nothing we can dispatch.
return actionSkippedUnresolvedSelector, nil
}
return drv.LongPress(ctx, x, y)
return "", drv.LongPress(ctx, x, y)
case verifier.ActionKindScroll:
fromX, fromY, toX, toY := scrollEndpoints(action, tree)
fromX, fromY, toX, toY = clampGestureToSafeArea(fromX, fromY, toX, toY, screenBounds(tree))
@@ -752,17 +979,20 @@ func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.A
if duration <= 0 {
duration = 300 * time.Millisecond
}
return drv.Swipe(ctx, fromX, fromY, toX, toY, duration)
if scroller, ok := drv.(driver.Scroller); ok {
return "", scroller.Scroll(ctx, fromX, fromY, toX, toY, duration)
}
return "", drv.Swipe(ctx, fromX, fromY, toX, toY, duration)
case verifier.ActionKindInputText:
tapped := false
if x, y, ok := resolveCoordinates(action, tree); ok {
if err := drv.Tap(ctx, x, y); err != nil {
return err
return "", err
}
tapped = true
} else if action.On != "" {
if err := drv.TapSelector(ctx, action.On); err != nil {
return err
return "", err
}
tapped = true
}
@@ -776,9 +1006,12 @@ func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.A
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
return "", ctx.Err()
case <-timer.C:
}
if err := confirmFocus(ctx, drv, action.On, tree); err != nil {
return "", err
}
}
// InputText replaces the field's content: erase what the target
// holds before typing. Appending instead lets repeated draws grow
@@ -788,38 +1021,38 @@ func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.A
if !inputReplacesText(drv) {
if count := existingTextLength(action, tree); count > 0 {
if err := drv.EraseText(ctx, count); err != nil {
return err
return "", err
}
}
}
return drv.InputText(ctx, action.Text)
return "", drv.InputText(ctx, action.Text)
case verifier.ActionKindSwipe:
duration := time.Duration(action.DurationMillis) * time.Millisecond
if duration <= 0 {
duration = 250 * time.Millisecond
}
fromX, fromY, toX, toY := clampGestureToSafeArea(action.FromX, action.FromY, action.ToX, action.ToY, screenBounds(tree))
return drv.Swipe(ctx, fromX, fromY, toX, toY, duration)
return "", drv.Swipe(ctx, fromX, fromY, toX, toY, duration)
case verifier.ActionKindPressKey:
if action.Key == "" {
return nil
return actionSkippedMissingKey, nil
}
return drv.PressKey(ctx, action.Key)
return "", drv.PressKey(ctx, action.Key)
case verifier.ActionKindWait:
duration := time.Duration(action.DurationMillis) * time.Millisecond
if duration <= 0 {
return nil
return actionSkippedZeroDurationWait, nil
}
timer := time.NewTimer(duration)
defer timer.Stop()
select {
case <-ctx.Done():
return ctx.Err()
return "", ctx.Err()
case <-timer.C:
return nil
return "", nil
}
default:
return fmt.Errorf("unknown action kind %q", action.Kind)
return "", fmt.Errorf("unknown action kind %q", action.Kind)
}
}
@@ -855,6 +1088,81 @@ func collectLogs(
return result
}
// collectExceptions reads the app's captured uncaught errors. Like log
// capture it is best-effort: a failed read is warned on rather than ending
// the run, and a driver that cannot report them yields none.
func collectExceptions(
ctx context.Context,
reporter driver.ExceptionReporter,
logger *slog.Logger,
stepIndex int,
) []verifier.Exception {
if reporter == nil {
return nil
}
captured, err := reporter.Exceptions(ctx)
if err != nil {
logger.Warn("exception fetch failed", "step", stepIndex, "err", err)
return nil
}
result := make([]verifier.Exception, 0, len(captured))
for _, entry := range captured {
result = append(result, verifier.Exception{
Class: entry.Class,
Message: entry.Message,
StackTrace: entry.StackTrace,
UnixMillis: entry.UnixMillis,
})
}
return result
}
func collectNavigations(
ctx context.Context,
reporter driver.NavigationReporter,
logger *slog.Logger,
stepIndex int,
) []trace.Navigation {
if reporter == nil {
return nil
}
observed, err := reporter.Navigations(ctx)
if err != nil {
logger.Warn("navigation fetch failed", "step", stepIndex, "err", err)
return nil
}
records := make([]trace.Navigation, 0, len(observed))
for _, entry := range observed {
records = append(records, trace.Navigation{URL: entry.URL, UnixMillis: entry.UnixMillis})
}
if len(records) == 0 {
return nil
}
return records
}
func traceLogs(entries []verifier.LogEntry) []trace.LogEntry {
if len(entries) == 0 {
return nil
}
result := make([]trace.LogEntry, 0, len(entries))
for _, entry := range entries {
result = append(result, trace.LogEntry(entry))
}
return result
}
func traceExceptions(entries []verifier.Exception) []trace.Exception {
if len(entries) == 0 {
return nil
}
result := make([]trace.Exception, 0, len(entries))
for _, entry := range entries {
result = append(result, trace.Exception(entry))
}
return result
}
// inputReplacesText reports whether the driver's InputText replaces existing
// content, making the runner's pre-erase redundant.
func inputReplacesText(drv driver.DeviceDriver) bool {
@@ -862,6 +1170,93 @@ func inputReplacesText(drv driver.DeviceDriver) bool {
return ok && replacer.ReplacesTextOnInput()
}
// confirmFocus fails the action when the device reports focus on an element
// other than the one the focus tap aimed at. Typing is a blind write to
// whatever holds focus, so a tap the target never received (a keyboard overlay
// window covering it, a target that cannot take focus) would stream the
// characters into a different field, corrupting it and every property that
// reads it. Only a field that already holds focus can receive that text, so
// the confirming read is charged only when the pre-tap hierarchy shows focus
// somewhere other than the target: a target that already holds focus, a screen
// with nothing focused, and platforms that never report focus (iOS) all skip
// it and keep the round-trip.
func confirmFocus(
ctx context.Context,
drv driver.DeviceDriver,
selector string,
tree *hierarchy.Tree,
) error {
if selector == "" || !otherElementHoldsFocus(tree, selector) {
return nil
}
dump, err := drv.Hierarchy(ctx)
if err != nil {
return fmt.Errorf("focus check for %s: %w", selector, err)
}
current, err := hierarchy.Parse(dump)
if err != nil {
return fmt.Errorf("focus check for %s: %w", selector, err)
}
if !otherElementHoldsFocus(current, selector) {
return nil
}
return fmt.Errorf(
"focus tap on %s did not focus it: %s holds focus, so the text would land there",
selector, elementName(focusedElement(current)),
)
}
// otherElementHoldsFocus reports whether the hierarchy shows focus on
// something outside the selector's subtree, which is the state that sends
// typed text to the wrong field. A selector the hierarchy cannot resolve
// answers false: not knowing where the target is says nothing about where the
// text would land, and failing on it turns every step the dump has no node for
// into an apply error, which is an aborted run three steps later.
func otherElementHoldsFocus(tree *hierarchy.Tree, selector string) bool {
if tree == nil || focusedElement(tree) == nil {
return false
}
target := tree.FindNode(selector)
return target != nil && !holdsFocus(target)
}
func focusedElement(tree *hierarchy.Tree) *hierarchy.Element {
for _, element := range tree.Elements {
if element.Focused {
return element
}
}
return nil
}
// holdsFocus accepts focus anywhere in the target's subtree: a selector often
// names the field wrapper while the platform reports focus on the inner
// editable node.
func holdsFocus(node *hierarchy.Node) bool {
if node.Focused {
return true
}
for _, child := range node.Children {
if holdsFocus(child) {
return true
}
}
return false
}
func elementName(element *hierarchy.Element) string {
switch {
case element.ResourceID != "":
return element.ResourceID
case element.Description != "":
return element.Description
case element.Class != "":
return element.Class
default:
return "an unnamed element"
}
}
// existingTextLength returns the character count of the InputText target's
// current text, so the runner can erase it before typing. Zero when the
// target cannot be resolved or holds no text.
@@ -880,17 +1275,23 @@ func resolveCoordinates(action verifier.Action, tree *hierarchy.Tree) (int, int,
// When On is empty, X/Y are authoritative (web V8 path emits coordinates
// directly from getBoundingClientRect; the runtime nullifies unresolved
// actions upstream so a non-null InputText here always has real coords,
// even at (0,0)). When On is set, prefer the tree lookup so stale coords
// don't leak from earlier ticks.
// even at (0,0)). A point outside the viewport is off screen, not absent:
// only the driver knows whether it can scroll that point back into reach,
// so the judgement belongs there and not here. When On is set, prefer the
// tree lookup so stale coords don't leak from earlier ticks.
if action.On == "" {
if action.X >= 0 && action.Y >= 0 {
return action.X, action.Y, true
}
return 0, 0, false
return action.X, action.Y, true
}
if tree != nil {
if element := tree.Find(action.On); element != nil {
x, y := element.Bounds.Center()
// An ambiguous selector names several elements while the action's own
// coordinates name one, so the coordinates win. Attribute values match
// by substring, so a selector unique where the candidate was built can
// be ambiguous in the tree it resolves against. A bare-string target
// carries no coordinates, and there the name is all there is.
matches := tree.FindAll(action.On)
hasCoordinates := action.X > 0 && action.Y > 0
if len(matches) > 0 && (len(matches) == 1 || !hasCoordinates) {
x, y := matches[0].Bounds.Center()
if x > 0 && y > 0 {
return x, y, true
}
@@ -1178,13 +1579,18 @@ func driverIsAndroid(ctx context.Context, options Options, logger *slog.Logger)
}
func traceActionFor(action verifier.Action, tree *hierarchy.Tree) *trace.Action {
traceAction := &trace.Action{Kind: string(action.Kind), X: action.X, Y: action.Y}
traceAction := &trace.Action{
Kind: string(action.Kind),
X: action.X,
Y: action.Y,
Source: action.Source,
}
switch action.Kind {
case verifier.ActionKindTap, verifier.ActionKindDoubleTap, verifier.ActionKindLongPress:
traceAction.Selector = action.On
stampSelectorTarget(traceAction, action, tree)
case verifier.ActionKindInputText:
traceAction.Text = action.Text
traceAction.Text = verifier.RecordedActionText(action, tree)
traceAction.Selector = action.On
stampSelectorTarget(traceAction, action, tree)
case verifier.ActionKindSwipe:
@@ -1341,12 +1747,60 @@ func encodeResiduals(residuals map[string]ltl.Formula) (map[string]json.RawMessa
return encoded, firstErr
}
// actionSkipReason names why a chosen action was never dispatched. It is
// recorded on the step so a count of executed actions is not inflated by the
// next_action of a step that acted on nothing. Empty means the action ran.
type actionSkipReason string
const (
actionSkippedForeground actionSkipReason = "app_left_foreground"
actionSkippedApplyError actionSkipReason = "apply_error"
// The action was dispatched and the driver never came back inside the
// step's bound. Distinct from apply_error, which is a call that answered
// and said no, and from the reasons below, which are actions that never
// reached the device at all.
actionSkippedApplyTimeout actionSkipReason = "apply_timeout"
// The action named a selector that resolved to no on-screen coordinates,
// either because its verb has no by-selector dispatch to fall back to or
// because the driver's own lookup found nothing to tap.
actionSkippedUnresolvedSelector actionSkipReason = "unresolved_selector"
actionSkippedMissingKey actionSkipReason = "missing_key"
actionSkippedZeroDurationWait actionSkipReason = "zero_duration_wait"
// The driver resolved the action's point and found no element there, so
// the gesture was never dispatched. Recorded rather than counted as a
// device fault: a run that acts on nothing has to say so.
actionSkippedGestureUndelivered actionSkipReason = "gesture_undelivered"
// The step asked the action source for an action and was handed none: a
// generator with no candidate for this screen, or a model call that failed.
// The step then drives nothing, which is the one non-action a run can take
// on every one of its steps and still finish reporting a full step count.
actionSkippedNoActionProduced actionSkipReason = "no_action_produced"
)
// maxConsecutiveApplyFailures bounds how many transient apply failures in a
// row the run tolerates before aborting. One or two absorb a runner restart;
// an unbroken streak means the device is wedged and the rest of the budget
// would be spent doing nothing.
const maxConsecutiveApplyFailures = 3
// applyTimeout bounds one dispatched action and observationTimeout the device
// reads a step opens with. Options.Duration is a loop condition checked between
// steps and the drivers add no deadline of their own, so without these a call
// that never returns holds the run for as long as the process lives. Both are
// far above any healthy call and far below the timeout a campaign runner puts
// on a whole run. Variables so the timeout tests can shrink them.
var (
applyTimeout = 60 * time.Second
observationTimeout = 60 * time.Second
)
// applyBound is how long one dispatched action may take. An action that names
// its own duration carries it on top: the bound exists to end a call that
// stopped answering, not to cut a gesture the spec asked for.
func applyBound(action verifier.Action) time.Duration {
return applyTimeout + time.Duration(action.DurationMillis)*time.Millisecond
}
// isWDADrop reports that the sidecar could not restart the iOS XCTest
// runner: the channel is gone for good and the run must abort. Transient
// drops are classified by the sidecar itself (it reconnects and surfaces
File diff suppressed because it is too large. Load diff
+12 -4
View File
@@ -15,9 +15,11 @@ import (
// ActionSource resolves the next action for a step. Both runtimes (the goja
// picker and the web/V8 picker) implement it so the runner loop has one path
// and no per-step driver type assertion. NextAction returns verifier.ErrNoAction
// when the picker declined to act this tick.
// when the picker declined to act this tick. stepIndex is the trace line the
// decision belongs to, so a source that keeps its own records can key them to it
// rather than counting steps a second time.
type ActionSource interface {
NextAction(ctx context.Context) (verifier.Action, error)
NextAction(ctx context.Context, stepIndex int) (verifier.Action, error)
}
// ExtractorSource yields per-step extractor overrides the runner applies after
@@ -57,7 +59,7 @@ type gojaSource struct {
verifier *verifier.Verifier
}
func (s gojaSource) NextAction(context.Context) (verifier.Action, error) {
func (s gojaSource) NextAction(context.Context, int) (verifier.Action, error) {
return s.verifier.NextAction()
}
@@ -76,7 +78,7 @@ type webSource struct {
web driver.WebDriver
}
func (s webSource) NextAction(ctx context.Context) (verifier.Action, error) {
func (s webSource) NextAction(ctx context.Context, _ int) (verifier.Action, error) {
raw, err := s.web.NextActionFromV8(ctx)
if err != nil {
return verifier.Action{}, fmt.Errorf("v8 next action: %w", err)
@@ -160,8 +162,14 @@ func pickSources(options Options) (ActionSource, ExtractorSource, error) {
client: client,
model: config.Model,
instructions: config.Instructions,
labelSource: options.LabelSource,
logger: logger,
history: newActionHistory(llmHistorySize),
}
// Assigned only when present so the interface field stays nil rather than
// holding a typed nil pointer that would panic on the first record.
if options.TraceWriter != nil {
action.recorder = options.TraceWriter
}
return action, extractor, nil
}
+4 -4
View File
@@ -1,4 +1,4 @@
{"extractor_changes":{"extractor_0":{"prev":null,"curr":0}},"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120}},"residuals":{"balanceNonNegative":{"op":"true"}},"step":1,"timestamp":"0001-01-01T00:00:00Z"}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120}},"residuals":{"balanceNonNegative":{"op":"true"}},"step":2,"timestamp":"0001-01-01T00:00:00Z"}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120}},"residuals":{"balanceNonNegative":{"op":"true"}},"step":3,"timestamp":"0001-01-01T00:00:00Z"}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120}},"residuals":{"balanceNonNegative":{"op":"true"}},"step":4,"timestamp":"0001-01-01T00:00:00Z"}
{"extractor_changes":{"extractor_0":{"prev":null,"curr":0}},"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}],"depths":[0,1]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120},"source":"seeded"},"residuals":{"balanceNonNegative":{"op":"true"}},"step":1,"timestamp":"0001-01-01T00:00:00Z","trace_version":1}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}],"depths":[0,1]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120},"source":"seeded"},"residuals":{"balanceNonNegative":{"op":"true"}},"step":2,"timestamp":"0001-01-01T00:00:00Z","trace_version":1}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}],"depths":[0,1]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120},"source":"seeded"},"residuals":{"balanceNonNegative":{"op":"true"}},"step":3,"timestamp":"0001-01-01T00:00:00Z","trace_version":1}
{"hierarchy":{"elements":[{"resourceId":"HomeScreen","bounds":{"left":0,"top":0,"right":0,"bottom":0},"attrs":{"editable":"false","resource-id":"HomeScreen"}},{"resourceId":"next","clickable":true,"enabled":true,"bounds":{"left":40,"top":80,"right":240,"bottom":160},"attrs":{"bounds":"[40,80,240,160]","clickable":"true","editable":"false","enabled":"true","resource-id":"next"}}],"depths":[0,1]},"next_action":{"kind":"Tap","selector":"id:next","resolved_bounds":{"x":40,"y":80,"width":200,"height":80},"tap_point":{"x":140,"y":120},"source":"seeded"},"residuals":{"balanceNonNegative":{"op":"true"}},"step":4,"timestamp":"0001-01-01T00:00:00Z","trace_version":1}
+2 -1
View File
@@ -1,4 +1,5 @@
run complete: 4 steps
run complete: 4 steps, 0 driven by the generator
1 violation record(s):
step 1: [balanceNonNegative]
4 action(s) never reached the app: no_action_produced 4
+83
View File
@@ -0,0 +1,83 @@
// Package seedspec expands the seed specification a campaign is given into the
// explicit list of seeds it intends to run. The campaign tool and the tools
// that drive it have to read a specification the same way, or a sweep records
// an intent that differs from what ran.
package seedspec
import (
"fmt"
"strconv"
"strings"
)
// Parse expands a seed specification such as "1-10,20,30-32" into the
// explicit seed list a campaign intends to run.
func Parse(specification string) ([]int64, error) {
trimmed := strings.TrimSpace(specification)
if trimmed == "" {
return nil, fmt.Errorf("empty seed spec")
}
var seeds []int64
seen := map[int64]bool{}
for _, part := range strings.Split(trimmed, ",") {
part = strings.TrimSpace(part)
if part == "" {
return nil, fmt.Errorf("empty seed in %q", specification)
}
expanded, err := expandSeedPart(part)
if err != nil {
return nil, err
}
for _, seed := range expanded {
if seen[seed] {
return nil, fmt.Errorf("duplicate seed %d in %q", seed, specification)
}
seen[seed] = true
seeds = append(seeds, seed)
}
}
return seeds, nil
}
func expandSeedPart(part string) ([]int64, error) {
start, end, isRange := strings.Cut(part, "-")
if !isRange {
seed, err := parseSeed(part)
if err != nil {
return nil, err
}
return []int64{seed}, nil
}
first, err := parseSeed(strings.TrimSpace(start))
if err != nil {
return nil, fmt.Errorf("seed range %q: %w", part, err)
}
last, err := parseSeed(strings.TrimSpace(end))
if err != nil {
return nil, fmt.Errorf("seed range %q: %w", part, err)
}
if first > last {
return nil, fmt.Errorf("seed range %q: start %d is above end %d", part, first, last)
}
seeds := make([]int64, 0, last-first+1)
for seed := first; seed <= last; seed++ {
seeds = append(seeds, seed)
}
return seeds, nil
}
func parseSeed(text string) (int64, error) {
seed, err := strconv.ParseInt(text, 10, 64)
if err != nil {
return 0, fmt.Errorf("invalid seed %q: want a positive integer", text)
}
if seed == 0 {
// `sanderling test` reads --seed 0 as "derive a seed from the clock",
// so a campaign listing seed 0 records a run nobody can reproduce.
return 0, fmt.Errorf("seed 0 is not reproducible: sanderling test derives a random seed when --seed is 0, so list explicit non-zero seeds")
}
if seed < 0 {
return 0, fmt.Errorf("invalid seed %d: want a positive integer", seed)
}
return seed, nil
}
+49
View File
@@ -0,0 +1,49 @@
package seedspec
import (
"slices"
"strings"
"testing"
)
func TestParse_RangesAndLists(t *testing.T) {
cases := []struct {
specification string
want []int64
}{
{"1-5", []int64{1, 2, 3, 4, 5}},
{"1,5,9", []int64{1, 5, 9}},
{"1-3,20,30-32", []int64{1, 2, 3, 20, 30, 31, 32}},
{" 7 , 8 ", []int64{7, 8}},
{"4-4", []int64{4}},
}
for _, testCase := range cases {
got, err := Parse(testCase.specification)
if err != nil {
t.Fatalf("%q: %v", testCase.specification, err)
}
if !slices.Equal(got, testCase.want) {
t.Errorf("%q: got %v, want %v", testCase.specification, got, testCase.want)
}
}
}
func TestParse_RejectsSeedZero(t *testing.T) {
for _, specification := range []string{"0", "1,0,2", "0-3"} {
_, err := Parse(specification)
if err == nil {
t.Fatalf("%q: expected rejection of seed 0", specification)
}
if !strings.Contains(err.Error(), "not reproducible") {
t.Errorf("%q: error should explain why seed 0 is rejected: %v", specification, err)
}
}
}
func TestParse_RejectsMalformed(t *testing.T) {
for _, specification := range []string{"", " ", "abc", "1,,2", "5-1", "1-", "-5", "1-2-3", "1.5", "2,2"} {
if seeds, err := Parse(specification); err == nil {
t.Errorf("%q: expected error, got %v", specification, seeds)
}
}
}
+125
View File
@@ -5,9 +5,12 @@ import (
"context"
"fmt"
"io"
"maps"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"time"
"github.com/priyanshujain/sanderling/internal/android"
@@ -62,8 +65,21 @@ type Options struct {
// recorded violations as a ViolationsError, so a caller (CI) can tell
// "the run found the bug" from "the run finished clean".
ExitOnViolation bool
// AllowNoProperties lets a run proceed against a spec that registers no
// properties. The extraction sweeps pass it: they measure what a spec can
// read and report no detection count. Every other run without it is a false
// green.
AllowNoProperties bool
// AllowNoGeneratorActions lets a run finish having never been driven by its
// action generator. The exploration-reach sweeps pass it, because "the
// generator reached nothing on this build" is their measurement rather than
// a broken run.
AllowNoGeneratorActions bool
// Generator selects the action picker: "llm" or the default seeded picker.
Generator string
// LabelSource selects how candidates are named to the model picker, and is
// recorded in meta.json as part of the run's cell.
LabelSource string
// iosUDID, iosIsSimulator, and iosCoreDeviceID are filled by Execute after
// resolving the iOS target, then read by buildDriver to choose the simulator
@@ -79,6 +95,15 @@ type Options struct {
// recorded only when the LLM picker is the one that will actually run, so a
// spec that declares generator = llm() but is run under the seeded picker does
// not label its trace with a model it never called.
// runDevice reads whichever flag named the hardware for this platform. An ios
// run is selected with --ios-device and leaves --device empty.
func runDevice(options Options) string {
if options.Platform == "ios" {
return options.IosDevice
}
return options.Device
}
func buildRunMeta(options Options, bundleSHA256 string, seed int64, host string, llmConfig verifier.LLMConfig, hasLLMConfig bool) trace.Meta {
meta := trace.Meta{
Seed: seed,
@@ -90,9 +115,11 @@ func buildRunMeta(options Options, bundleSHA256 string, seed int64, host string,
SanderlingVersion: "0.0.1",
Arm: options.Arm,
Generator: options.Generator,
LabelSource: options.LabelSource,
MaxSteps: options.MaxSteps,
DurationMillis: options.Duration.Milliseconds(),
Host: host,
Device: runDevice(options),
}
if options.Generator == "llm" && hasLLMConfig {
meta.Model = llmConfig.Model
@@ -186,6 +213,9 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
if err := verifierInstance.Load(string(bundle.JavaScript)); err != nil {
return fmt.Errorf("load spec: %w", err)
}
if !options.AllowNoProperties && len(verifierInstance.PropertyNames()) == 0 {
return NoPropertiesError{Spec: options.Spec}
}
fmt.Fprintln(stdout, "spec loaded into verifier")
runDirectory := filepath.Join(options.Output, time.Now().UTC().Format("20060102-150405"))
@@ -222,6 +252,7 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
TraceWriter: traceWriter,
Logger: newProgressLogger(stdout),
Generator: options.Generator,
LabelSource: options.LabelSource,
StopOnViolation: options.ExitOnViolation,
})
@@ -247,6 +278,19 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
// because it holds no verdict to report. The threshold is every step and not a
// fraction of them: a screen that composes now and then costs a healthy android
// run a step or two, and a check that fired on those would be red on every run.
//
// A run whose generator dispatched no action fails on the same grounds, and the
// threshold is zero for the same reason: a generator with nothing to offer on
// some screens is ordinary, one with nothing to offer on every screen of a whole
// run drove nothing. The count is the generator's alone because a spec's setup
// drives the app before the generator is consulted, so a login that ran leaves
// dispatched actions behind whatever the generator then did.
//
// A recorded violation carries such a run through whether or not
// --exit-on-violation was passed: the refusal exists because a run with no
// verdict must not read as a clean one, and a run that recorded a violation
// holds a verdict. Campaigns pass no flags, so refusing them there would write
// exit_code 1 and lose a real detection to the analysis as missing data.
func runOutcome(options Options, summary runner.Summary) error {
if summary.Steps > 0 && summary.SkippedVerification == summary.Steps {
return VacuousRunError{Steps: summary.Steps}
@@ -254,6 +298,13 @@ func runOutcome(options Options, summary runner.Summary) error {
if options.ExitOnViolation && len(summary.Violations) > 0 {
return ViolationsError{Count: len(summary.Violations)}
}
if !options.AllowNoGeneratorActions && len(summary.Violations) == 0 &&
summary.Steps > 0 && summary.GeneratorActions == 0 {
return NoGeneratorActionsError{
Steps: summary.Steps,
SkippedActions: summary.SkippedActions,
}
}
return nil
}
@@ -269,6 +320,23 @@ func (e ViolationsError) Error() string {
return fmt.Sprintf("%d violation record(s)", e.Count)
}
// BundleSpec produces the goja bundle a run of this spec loaded, seeded as that
// run was. An offline replay of the run's trace has to load the same JavaScript
// the run did, and the seed is one of the bundle's defines, so it is part of
// the bundle's identity.
func BundleSpec(specPath string, seed int64) (bundler.Result, error) {
inputs, err := prepareBundleInputs(Options{Spec: specPath, Seed: seed})
if err != nil {
return bundler.Result{}, err
}
return bundler.Bundle(bundler.Options{
EntryFile: specPath,
RuntimeFile: inputs.gojaRuntimePath,
Defines: inputs.defines,
Aliases: inputs.aliases,
})
}
// VacuousRunError reports a run in which no step reached the verifier, so no
// property ever judged anything. It is not a clean run and it is not a found
// bug: it is a run that produced no evidence either way, and the absence of
@@ -285,6 +353,63 @@ func (e VacuousRunError) Error() string {
e.Steps)
}
// NoPropertiesError reports a spec that bundled and loaded cleanly and holds no
// properties. Nothing is broken: the run would drive the app, fill a trace and
// report no violations having judged nothing, and that green says as much about
// the app as an unplugged meter says about a wire. It stays untyped to the CLI's
// violation path like VacuousRunError, so it exits 1 as a run that cannot
// produce a verdict rather than 2.
type NoPropertiesError struct {
Spec string
}
func (e NoPropertiesError) Error() string {
return fmt.Sprintf(
"%s bundled and loaded into the verifier cleanly and registers no properties: "+
"nothing is wrong with the spec and nothing is wrong with the run, but this run "+
"would check nothing and report no violations. Pass --allow-no-properties for a "+
"run that measures what the spec extracts instead of judging the app",
e.Spec)
}
// NoGeneratorActionsError reports a run not one of whose steps was driven by the
// action generator. Setup can put the app in position, but only the generator
// explores it, so every screen this run judged was one setup left it on and its
// empty violation list says as much about the app as a spec with no properties
// would: the run observed, judged the same state over and over, and exercised
// nothing. It stays untyped to the CLI's violation path like VacuousRunError, so
// it exits 1 as a run that holds no verdict rather than 2.
type NoGeneratorActionsError struct {
Steps int
// SkippedActions is the runner's per-reason count of actions that never
// reached the app, which is where the cause is: a picker with no candidate
// reads differently from one whose every model call failed.
SkippedActions map[string]int
}
func (e NoGeneratorActionsError) Error() string {
return fmt.Sprintf(
"%d step(s) ran and the action generator drove the app in none of them: whatever "+
"the spec's setup did to get the app into position, nothing explored it from "+
"there, so the run judged one screen over and over and its violation count "+
"says nothing about the rest of the app%s. Pass --allow-no-generator-actions "+
"for a run that measures where a generator reaches instead of judging the app",
e.Steps, skipReasonSuffix(e.SkippedActions))
}
// skipReasonSuffix renders the skip-reason tally as a trailing clause, empty
// when the run recorded none.
func skipReasonSuffix(skipped map[string]int) string {
if len(skipped) == 0 {
return ""
}
reasons := make([]string, 0, len(skipped))
for _, reason := range slices.Sorted(maps.Keys(skipped)) {
reasons = append(reasons, fmt.Sprintf("%s %d", reason, skipped[reason]))
}
return ". Actions that never reached the app: " + strings.Join(reasons, ", ")
}
// bundleInputs holds the pre-driver assembly: alias map, seed, esbuild defines,
// and the resolved spec-API/goja-runtime paths the bundler consumes.
type bundleInputs struct {
+233 -11
View File
@@ -13,6 +13,7 @@ import (
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/runner"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
@@ -193,13 +194,14 @@ func TestResolveSpecAPIPath_ReturnsEmptyWhenMissing(t *testing.T) {
func TestBuildRunMeta_RecordsArmMembership(t *testing.T) {
options := Options{
Spec: "spec.ts",
BundleID: "com.example",
Platform: "android",
Duration: 3 * time.Minute,
MaxSteps: 300,
Arm: "llm-visible-text",
Generator: "llm",
Spec: "spec.ts",
BundleID: "com.example",
Platform: "android",
Duration: 3 * time.Minute,
MaxSteps: 300,
Arm: "llm-visible-text",
Generator: "llm",
LabelSource: verifier.LabelSourceVisibleText,
}
meta := buildRunMeta(options, "deadbeef", 7, "farm-01",
verifier.LLMConfig{Model: "claude-sonnet-5", Instructions: "exercise the outbox"}, true)
@@ -207,6 +209,9 @@ func TestBuildRunMeta_RecordsArmMembership(t *testing.T) {
if meta.Arm != "llm-visible-text" || meta.Generator != "llm" {
t.Errorf("arm membership: got arm=%q generator=%q", meta.Arm, meta.Generator)
}
if meta.LabelSource != verifier.LabelSourceVisibleText {
t.Errorf("label source: got %q, want %q", meta.LabelSource, verifier.LabelSourceVisibleText)
}
if meta.Model != "claude-sonnet-5" || meta.Instructions != "exercise the outbox" {
t.Errorf("llm config: got model=%q instructions=%q", meta.Model, meta.Instructions)
}
@@ -218,6 +223,35 @@ func TestBuildRunMeta_RecordsArmMembership(t *testing.T) {
}
}
// The device a run drove has to survive in the run's own artifact. Held only in
// the campaign's runs.jsonl, a trace read on its own cannot say which emulator,
// or which API level, produced it.
func TestBuildRunMeta_RecordsTheDeviceInMetaJSON(t *testing.T) {
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
defer writer.Close()
options := Options{Platform: "android", Generator: "seeded", Duration: time.Minute, Device: "emulator-5556"}
if err := writer.WriteMeta(buildRunMeta(options, "deadbeef", 3, "farm-01", verifier.LLMConfig{}, false)); err != nil {
t.Fatal(err)
}
body, err := os.ReadFile(filepath.Join(directory, "meta.json"))
if err != nil {
t.Fatal(err)
}
var stored trace.Meta
if err := json.Unmarshal(body, &stored); err != nil {
t.Fatalf("meta.json is not valid JSON: %v\n%s", err, body)
}
if stored.Device != "emulator-5556" {
t.Errorf("device in meta.json: got %q, want emulator-5556\n%s", stored.Device, body)
}
}
func TestBuildRunMeta_OmitsModelWhenSeededPickerRuns(t *testing.T) {
options := Options{Platform: "android", Generator: "seeded", Duration: time.Minute}
meta := buildRunMeta(options, "deadbeef", 1, "farm-01",
@@ -229,6 +263,55 @@ func TestBuildRunMeta_OmitsModelWhenSeededPickerRuns(t *testing.T) {
}
}
// TestBuildRunMeta_RecordsLabelSourceForASeededRun is the deliberate difference
// from Model and Instructions above. The seeded picker never reads a label, but
// the run still belongs to a labelling cell, and the pair of seeded runs across
// the two cells is the manipulation check. Omitting it here would leave those
// two runs indistinguishable in the artifact.
func TestBuildRunMeta_RecordsLabelSourceForASeededRun(t *testing.T) {
options := Options{
Platform: "android",
Generator: "seeded",
Duration: time.Minute,
LabelSource: verifier.LabelSourceResourceID,
}
meta := buildRunMeta(options, "deadbeef", 1, "farm-01", verifier.LLMConfig{}, false)
if meta.LabelSource != verifier.LabelSourceResourceID {
t.Errorf("label source: got %q, want %q", meta.LabelSource, verifier.LabelSourceResourceID)
}
}
// An ios run names its simulator with --ios-device, so reading the android
// --device flag left every ios trace unable to say what it executed on.
func TestBuildRunMeta_NamesTheIosSimulatorItRanOn(t *testing.T) {
options := Options{
Platform: "ios",
Generator: "seeded",
Duration: time.Minute,
IosDevice: "iPhone 17 Pro",
}
meta := buildRunMeta(options, "deadbeef", 1, "farm-01", verifier.LLMConfig{}, false)
if meta.Device != "iPhone 17 Pro" {
t.Errorf("device: got %q, want the ios simulator the run named", meta.Device)
}
}
func TestBuildRunMeta_NamesTheAndroidDeviceItRanOn(t *testing.T) {
options := Options{
Platform: "android",
Generator: "seeded",
Duration: time.Minute,
Device: "emulator-5554",
}
meta := buildRunMeta(options, "deadbeef", 1, "farm-01", verifier.LLMConfig{}, false)
if meta.Device != "emulator-5554" {
t.Errorf("device: got %q, want the android device the run named", meta.Device)
}
}
func TestBuildRunMeta_OmitsModelWhenSpecDeclaresNoLLMGenerator(t *testing.T) {
options := Options{Platform: "android", Generator: "llm", Duration: time.Minute}
meta := buildRunMeta(options, "deadbeef", 1, "farm-01", verifier.LLMConfig{}, false)
@@ -244,10 +327,12 @@ func TestBuildRunMeta_OmitsModelWhenSpecDeclaresNoLLMGenerator(t *testing.T) {
// stays a successful run, which is what every existing invocation expects.
func TestRunOutcome_ReportsViolationsOnlyUnderTheFlag(t *testing.T) {
violated := runner.Summary{
Steps: 7,
Violations: []runner.ViolationRecord{{StepIndex: 3, Properties: []string{"balanceMoves"}}},
Steps: 7,
DispatchedActions: 7,
GeneratorActions: 7,
Violations: []runner.ViolationRecord{{StepIndex: 3, Properties: []string{"balanceMoves"}}},
}
clean := runner.Summary{Steps: 7}
clean := runner.Summary{Steps: 7, DispatchedActions: 7, GeneratorActions: 7}
if err := runOutcome(Options{}, violated); err != nil {
t.Errorf("without --exit-on-violation a violated run must succeed, got %v", err)
@@ -285,12 +370,149 @@ func TestRunOutcome_ARunThatJudgedNothingIsNotASuccess(t *testing.T) {
// A screen that composes now and then costs a run steps, not its verdict. A
// check that fired here would turn every healthy android run red.
mostlyJudged := runner.Summary{Steps: 6, SkippedVerification: 5}
mostlyJudged := runner.Summary{
Steps: 6,
SkippedVerification: 5,
DispatchedActions: 1,
GeneratorActions: 1,
}
if err := runOutcome(Options{}, mostlyJudged); err != nil {
t.Errorf("a run that judged one of its 6 steps must succeed, got %v", err)
}
}
// A run that dispatched no action at all observed one screen for its whole life
// and never drove the app, so its empty violation list is what an unplugged
// instrument reports. The provider rate-limiting a model picker is the way this
// happens in practice: every step ends with the picker handing back nothing.
func TestRunOutcome_ARunThatDroveNothingIsNotASuccess(t *testing.T) {
droveNothing := runner.Summary{
Steps: 200,
SkippedActions: map[string]int{"no_action_produced": 200},
}
err := runOutcome(Options{}, droveNothing)
var dead NoGeneratorActionsError
if !errors.As(err, &dead) {
t.Fatalf("a run that dispatched none of its 200 steps' actions came back %v, "+
"want a NoGeneratorActionsError", err)
}
if dead.Steps != 200 {
t.Errorf("steps: got %d, want 200", dead.Steps)
}
if !strings.Contains(dead.Error(), "no_action_produced") {
t.Errorf("the error never names why nothing was dispatched: %v", dead)
}
// One generator action is exploration, however little. A check that fired
// here would be red on any run whose screen offers the generator nothing for
// a while.
droveOnce := runner.Summary{Steps: 200, DispatchedActions: 1, GeneratorActions: 1}
if err := runOutcome(Options{}, droveOnce); err != nil {
t.Errorf("a run whose generator dispatched one action must succeed, got %v", err)
}
// The sweeps that measure where a generator reaches ask for a run that
// explores nothing by name, and "the generator reached nothing here" is
// their measurement rather than their failure.
if err := runOutcome(Options{AllowNoGeneratorActions: true}, droveNothing); err != nil {
t.Errorf("the dead-run opt-out no longer carries a run through, got %v", err)
}
// A run cut short before it took a step never got going, which the deadline
// and the step count already say.
if err := runOutcome(Options{}, runner.Summary{}); err != nil {
t.Errorf("a run with no steps at all must not report as dead-on-arrival, got %v", err)
}
// CI reads exit 2 as "the run found the bug". A run whose first screen
// already violated must keep reporting that, or the found bug is downgraded
// to a broken harness.
foundOnTheFirstScreen := droveNothing
foundOnTheFirstScreen.Violations = []runner.ViolationRecord{
{StepIndex: 1, Properties: []string{"balanceNonNegative"}},
}
err = runOutcome(Options{ExitOnViolation: true}, foundOnTheFirstScreen)
var violations ViolationsError
if !errors.As(err, &violations) {
t.Errorf("a run that found a violation came back %v, want a ViolationsError", err)
}
}
// A spec whose setup logs in drives the app before the generator is ever
// consulted, so a run whose every model call failed still reached the driver a
// few times. Those actions are the harness getting into position: counting them
// as the run driving the app passed 86 steps of folio that explored nothing.
func TestRunOutcome_SetupActionsDoNotCarryARunWhoseGeneratorDroveNothing(t *testing.T) {
loginThenNothing := runner.Summary{
Steps: 86,
DispatchedActions: 3,
GeneratorActions: 0,
SkippedActions: map[string]int{"no_action_produced": 83},
}
err := runOutcome(Options{}, loginThenNothing)
var dead NoGeneratorActionsError
if !errors.As(err, &dead) {
t.Fatalf("a run whose 3 dispatched actions all came from setup came back %v, "+
"want a NoGeneratorActionsError", err)
}
if dead.Steps != 86 {
t.Errorf("steps: got %d, want 86", dead.Steps)
}
if !strings.Contains(dead.Error(), "no_action_produced") {
t.Errorf("the error never names why the generator dispatched nothing: %v", dead)
}
// The same run under --exit-on-violation still reports its evidence: CI
// reads exit 2 as "the run found the bug", and a setup-only run that found
// one must not be downgraded to a broken harness.
found := loginThenNothing
found.Violations = []runner.ViolationRecord{{StepIndex: 2, Properties: []string{"balanceMoves"}}}
var violations ViolationsError
if err := runOutcome(Options{ExitOnViolation: true}, found); !errors.As(err, &violations) {
t.Errorf("a setup-only run that found a violation came back %v, want a ViolationsError", err)
}
}
// Two refusals, two flags. A sweep that runs a property-free spec asked to
// judge nothing, not to explore nothing, and a run with properties that needs
// the dead-run exemption must be able to say so without claiming a waiver it
// does not want.
func TestRunOutcome_TheDeadRunRefusalHasItsOwnOptOut(t *testing.T) {
droveNothing := runner.Summary{
Steps: 200,
SkippedActions: map[string]int{"no_action_produced": 200},
}
var dead NoGeneratorActionsError
if err := runOutcome(Options{AllowNoProperties: true}, droveNothing); !errors.As(err, &dead) {
t.Errorf("--allow-no-properties waived the dead-run refusal, which it does not name: got %v", err)
}
if err := runOutcome(Options{AllowNoGeneratorActions: true}, droveNothing); err != nil {
t.Errorf("--allow-no-generator-actions did not carry a run that drove nothing through, got %v", err)
}
}
// A recorded violation is a verdict, and a run that reached one is not a dead
// run whatever drove the app to it. Campaigns are where this bites: they never
// pass --exit-on-violation, so refusing such a run writes exit_code 1 into the
// record and the analysis drops a real detection as missing data.
func TestRunOutcome_ARecordedViolationCarriesARunWhoseGeneratorDroveNothing(t *testing.T) {
setupReachedTheBug := runner.Summary{
Steps: 40,
DispatchedActions: 3,
GeneratorActions: 0,
SkippedActions: map[string]int{"no_action_produced": 40},
Violations: []runner.ViolationRecord{
{StepIndex: 3, Properties: []string{"noUncaughtExceptions"}},
},
}
if err := runOutcome(Options{}, setupReachedTheBug); err != nil {
t.Fatalf("a run that recorded a violation came back %v, want the run to succeed "+
"so the campaign records exit_code 0 and the detection survives", err)
}
}
// wedgedLaunchDriver never returns from Launch, standing in for a driver whose
// device-side session is stuck.
type wedgedLaunchDriver struct {
+132
View File
@@ -0,0 +1,132 @@
package trace
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
"time"
)
// LLMCallFileName is the run-directory file model-call records are appended to,
// one JSON object per line, in step order.
//
// These records live beside trace.jsonl rather than inside it because the two
// have different readers. Every trace line already carries a full accessibility
// hierarchy, and both the replay server and the campaign summarizer scan every
// line of it; folding prompts, candidate lists and raw responses in would grow
// the lines those readers parse by several KB each for data neither one reads.
// The join is by Step, which is the step field of the trace line whose
// next_action the call produced.
const LLMCallFileName = "llm-calls.jsonl"
// Outcomes of one step's action selection. Exactly one is recorded per step of
// a model-driven run, so a step the guard threw away is never confused with one
// where the picker legitimately had nothing to do.
const (
// LLMOutcomeSelected: the model picked a candidate and the action ran.
LLMOutcomeSelected = "selected"
// LLMOutcomeSetupAction: the spec's setup generator drove the step, so the
// model was not consulted.
LLMOutcomeSetupAction = "setup_action"
// LLMOutcomeSetupFailed: the setup generator itself errored, which aborts
// the run.
LLMOutcomeSetupFailed = "setup_failed"
// LLMOutcomeNoCandidates: the action tree yielded nothing on this screen, so
// no call was made.
LLMOutcomeNoCandidates = "no_candidates"
// LLMOutcomeCandidatesFailed: the action tree cannot be enumerated for this
// policy at all (an authored leaf samples), which aborts the run.
LLMOutcomeCandidatesFailed = "candidates_failed"
// LLMOutcomeRequestFailed: the provider call failed (transport, timeout,
// non-2xx).
LLMOutcomeRequestFailed = "request_failed"
// LLMOutcomeNoChoices: a 2xx response carrying an empty choices array.
LLMOutcomeNoChoices = "no_choices"
// LLMOutcomeUnparsableResponse: the content was empty, not JSON, or carried
// no choice.
LLMOutcomeUnparsableResponse = "unparsable_response"
// LLMOutcomeChoiceOutOfRange: the number picked is not in 1..len(candidates).
LLMOutcomeChoiceOutOfRange = "choice_out_of_range"
// LLMOutcomeEchoMismatch: the echoed chosen_action disagrees with the
// numbered entry, so the pick is discarded.
LLMOutcomeEchoMismatch = "echo_mismatch"
// LLMOutcomeActionBuildFailed: the chosen candidate could not be turned into
// an executable action (e.g. the input sampler was unavailable).
LLMOutcomeActionBuildFailed = "action_build_failed"
)
// LLMCall is one step's action-selection record: what was sent, what came back,
// what it cost, and how the step ended.
type LLMCall struct {
// Step joins this record to the trace line of the same step index.
Step int `json:"step"`
Timestamp time.Time `json:"timestamp"`
Outcome string `json:"outcome"`
// Model is the requested model id; ServedModel is the one the provider
// reported serving, which a router may vary per call.
Model string `json:"model,omitempty"`
ServedModel string `json:"served_model,omitempty"`
// SystemPrompt and UserPrompt are the text parts as sent, not the templates
// they were assembled from.
SystemPrompt string `json:"system_prompt,omitempty"`
UserPrompt string `json:"user_prompt,omitempty"`
// Candidates is the numbered list the prompt rendered, so an experiment that
// varies candidate labelling can recover the labels this call actually saw.
Candidates []LLMCandidate `json:"candidates,omitempty"`
// Screenshot is the run-relative path of the image sent with the call. It
// names the step the image was captured at, which lags Step when the runner
// skipped a transitional observation.
Screenshot string `json:"screenshot,omitempty"`
// RawResponse is the assistant content before parsing.
RawResponse string `json:"raw_response,omitempty"`
PromptTokens int `json:"prompt_tokens,omitempty"`
CompletionTokens int `json:"completion_tokens,omitempty"`
TotalTokens int `json:"total_tokens,omitempty"`
LatencyMillis int64 `json:"latency_millis,omitempty"`
// Choice, EchoedAction and Reasoning are the parsed output, recorded
// whenever parsing succeeded and so present on the outcomes that then
// discarded the pick. EchoedAction is verbatim, including the weight
// annotation models copy along with the line; on an echo_mismatch it is what
// disagreed with Candidates[Choice-1].
Choice int `json:"choice,omitempty"`
EchoedAction string `json:"echoed_action,omitempty"`
Reasoning string `json:"reasoning,omitempty"`
Error string `json:"error,omitempty"`
}
// LLMCandidate is one numbered line of the candidate list as the model saw it.
type LLMCandidate struct {
Index int `json:"index"`
Kind string `json:"kind,omitempty"`
// Description is the rendered line the model echoes back; Label is the
// target's visible text the description was built from.
Description string `json:"description"`
Label string `json:"label,omitempty"`
// Weight is the percentage annotation shown on the line, 0 when the spec's
// action tree declared no weights and none was shown.
Weight int `json:"weight,omitempty"`
}
// WriteLLMCall appends one record to llm-calls.jsonl, creating the file on the
// first call.
func (w *Writer) WriteLLMCall(call LLMCall) error {
w.mutex.Lock()
defer w.mutex.Unlock()
if w.file == nil {
return fmt.Errorf("trace: writer is closed")
}
if w.llmCallEncoder == nil {
file, err := os.OpenFile(
filepath.Join(w.directory, LLMCallFileName),
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
0o644,
)
if err != nil {
return fmt.Errorf("open %s: %w", LLMCallFileName, err)
}
w.llmCallFile = file
w.llmCallEncoder = json.NewEncoder(file)
}
return w.llmCallEncoder.Encode(call)
}
+131
View File
@@ -0,0 +1,131 @@
package trace
import (
"encoding/json"
"os"
"path/filepath"
"reflect"
"strings"
"testing"
"time"
)
func TestWriteLLMCall_RoundTrip(t *testing.T) {
directory := t.TempDir()
writer, err := NewWriter(directory)
if err != nil {
t.Fatal(err)
}
call := LLMCall{
Step: 4,
Timestamp: time.Date(2026, 8, 12, 9, 30, 0, 0, time.UTC),
Outcome: LLMOutcomeSelected,
Model: "vendor/model",
ServedModel: "vendor/model-2026-05",
SystemPrompt: "You are exercising a UI to find bugs.\n\n" +
"hunt for double submits",
UserPrompt: "Actions available on the current screen:\n1. Tap \"Submit\" (w60)\n",
Candidates: []LLMCandidate{
{Index: 1, Kind: "tap", Description: `Tap "Submit"`, Label: "Submit", Weight: 60},
{Index: 2, Kind: "inputText", Description: `Type into "Name"`, Label: "Name", Weight: 40},
},
Screenshot: ScreenshotReference(4),
RawResponse: `{"reasoning":"submit twice","choice":1,"chosen_action":"Tap \"Submit\"","text":""}`,
PromptTokens: 1200,
CompletionTokens: 34,
TotalTokens: 1234,
LatencyMillis: 812,
Choice: 1,
EchoedAction: `Tap "Submit"`,
Reasoning: "submit twice",
}
if err := writer.WriteLLMCall(call); err != nil {
t.Fatal(err)
}
if err := writer.WriteLLMCall(LLMCall{
Step: 5,
Timestamp: call.Timestamp.Add(time.Second),
Outcome: LLMOutcomeNoCandidates,
}); err != nil {
t.Fatal(err)
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
body, err := os.ReadFile(filepath.Join(directory, LLMCallFileName))
if err != nil {
t.Fatal(err)
}
lines := strings.Split(strings.TrimSpace(string(body)), "\n")
if len(lines) != 2 {
t.Fatalf("wrote %d lines, want 2", len(lines))
}
var got LLMCall
if err := json.Unmarshal([]byte(lines[0]), &got); err != nil {
t.Fatalf("%s line 1 is not valid JSON: %v", LLMCallFileName, err)
}
if !reflect.DeepEqual(got, call) {
t.Errorf("round-trip mismatch:\n got: %+v\nwant: %+v", got, call)
}
var declined LLMCall
if err := json.Unmarshal([]byte(lines[1]), &declined); err != nil {
t.Fatal(err)
}
if declined.Step != 5 || declined.Outcome != LLMOutcomeNoCandidates {
t.Errorf("second record = %+v, want step 5 with outcome %q", declined, LLMOutcomeNoCandidates)
}
}
// TestWriteLLMCall_FileAbsentWithoutCalls keeps a seeded run's directory free of
// model-call output, so its presence alone identifies a model-driven run.
func TestWriteLLMCall_FileAbsentWithoutCalls(t *testing.T) {
directory := t.TempDir()
writer, err := NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteStep(Step{Index: 1, Timestamp: time.Now()}); err != nil {
t.Fatal(err)
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
if _, err := os.Stat(filepath.Join(directory, LLMCallFileName)); !os.IsNotExist(err) {
t.Errorf("stat %s = %v, want the file never to be created", LLMCallFileName, err)
}
}
// TestScreenshotReferenceMatchesWrittenFile pins the reference records point at
// to the path WriteScreenshot actually writes.
func TestScreenshotReferenceMatchesWrittenFile(t *testing.T) {
directory := t.TempDir()
writer, err := NewWriter(directory)
if err != nil {
t.Fatal(err)
}
defer writer.Close()
if err := writer.WriteScreenshot(12, []byte("not really a png")); err != nil {
t.Fatal(err)
}
if _, err := os.Stat(filepath.Join(directory, ScreenshotReference(12))); err != nil {
t.Errorf("stat %s: %v", ScreenshotReference(12), err)
}
}
func TestWriteLLMCall_AfterCloseErrors(t *testing.T) {
directory := t.TempDir()
writer, err := NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
if err := writer.WriteLLMCall(LLMCall{Step: 1}); err == nil {
t.Error("expected an error writing to a closed writer")
}
}
+109 -12
View File
@@ -13,14 +13,31 @@ import (
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// TraceVersion is stamped on every step this build writes. A step decoding to
// version 0 predates the fields introduced with version 1 (the tree's stored
// depths, per-step logs, per-step exceptions), which is what separates "this
// trace cannot answer the question" from "this step had nothing to report".
const TraceVersion = 1
type Step struct {
Index int `json:"step"`
Timestamp time.Time `json:"timestamp"`
Screen string `json:"screen,omitempty"`
Snapshots map[string]json.RawMessage `json:"snapshots,omitempty"`
Index int `json:"step"`
TraceVersion int `json:"trace_version,omitempty"`
Timestamp time.Time `json:"timestamp"`
Screen string `json:"screen,omitempty"`
Snapshots map[string]json.RawMessage `json:"snapshots,omitempty"`
// NextAction is the action chosen for the next iteration based on observing this step.
NextAction *Action `json:"next_action,omitempty"`
Exceptions []Exception `json:"exceptions,omitempty"`
NextAction *Action `json:"next_action,omitempty"`
// Logs are the platform log lines collected for this step, Exceptions the
// uncaught errors read at verification time: the error surface behind
// state.logs and state.exceptions, which the default properties read and
// an offline oracle has no other source for.
Logs []LogEntry `json:"logs,omitempty"`
Exceptions []Exception `json:"exceptions,omitempty"`
// Navigations are the document-replacing navigations seen since the
// previous step: the app reloaded, submitted a form, or changed route.
// Each one restarts the app's own runtime, so without them a reload and a
// generator repeating itself read the same way in a trace.
Navigations []Navigation `json:"navigations,omitempty"`
Violations []string `json:"violations,omitempty"`
Hierarchy *hierarchy.Tree `json:"hierarchy,omitempty"`
Residuals map[string]json.RawMessage `json:"residuals,omitempty"`
@@ -31,15 +48,37 @@ type Step struct {
// retry budget. The verifier is skipped for these steps so transient
// state does not poison the previous/current extractor advance.
Transitional bool `json:"transitional,omitempty"`
// ActionSkipped names why this step dispatched nothing: an action chosen and
// then thrown away, or a source that was asked and produced none. A count of
// executed actions cannot be inflated by either. Empty when the action ran,
// and empty on a held step, which never asked for one.
ActionSkipped string `json:"action_skipped,omitempty"`
// ObservationError names why this step's device read produced no tree,
// empty when a tree was read. A screen with no elements on it is a tree,
// so without this a step that observed nothing at all is indistinguishable
// from a step that observed an app showing nothing.
ObservationError string `json:"observation_error,omitempty"`
// SkippedVerification is set true exactly when the verifier was skipped
// for this step, so downstream tooling can tell a deliberately-skipped
// step from one that was verified and came back clean.
SkippedVerification bool `json:"skipped_verification,omitempty"`
// PreconditionFailure names a precondition of the run that was not met, so
// a step that never had the app under test in front of it cannot be counted
// as one that explored the app. Index 0 carries the startup gate's verdict:
// a trace holding that record and nothing else is a run that never started,
// which is a different thing from a run that explored and found nothing.
PreconditionFailure string `json:"precondition_failure,omitempty"`
// Witnesses records the violation witness for each property that newly
// violated at this step: the cause and the extractor values at onset.
Witnesses map[string]Witness `json:"witnesses,omitempty"`
}
// Navigation is one document-replacing navigation the run observed.
type Navigation struct {
URL string `json:"url"`
UnixMillis int64 `json:"unix_millis,omitempty"`
}
// Witness is the trace-side record of a property violation: why it fired, the
// two steps a deferred obligation spans, and the extractor values behind it.
type Witness struct {
@@ -70,6 +109,16 @@ type Metrics struct {
TotalMemoryBytes int64 `json:"total_memory_bytes,omitempty"`
}
// The three producers an action can come from. The spec's setup drives the app
// into position and explores nothing, so a per-action rate divides by the other
// two; an action recorded before this distinction existed carries none of them
// and cannot be attributed after the fact.
const (
ActionSourceSetup = "setup"
ActionSourceSeeded = "seeded"
ActionSourceModel = "llm"
)
type Action struct {
Kind string `json:"kind"`
X int `json:"x,omitempty"`
@@ -84,8 +133,8 @@ type Action struct {
Selector string `json:"selector,omitempty"`
ResolvedBounds *BoundsRecord `json:"resolved_bounds,omitempty"`
TapPoint *PointRecord `json:"tap_point,omitempty"`
// Source names the backend that chose this action: "llm" when the LLM
// action backend selected it, empty for the seeded picker. LLMReasoning is
// Source names which of the three producers chose this action, and is empty
// only on a trace recorded before actions named themselves. LLMReasoning is
// the model's short rationale, shown by the replay UI to explain the pick.
Source string `json:"source,omitempty"`
LLMReasoning string `json:"llm_reasoning,omitempty"`
@@ -110,6 +159,15 @@ type PointRecord struct {
Y int `json:"y"`
}
// LogEntry mirrors one platform log line the runner collected for this step,
// in the shape state.logs exposes to a spec.
type LogEntry struct {
UnixMillis int64 `json:"unix_millis,omitempty"`
Level string `json:"level,omitempty"`
Tag string `json:"tag,omitempty"`
Message string `json:"message,omitempty"`
}
type Exception struct {
Class string `json:"class"`
Message string `json:"message,omitempty"`
@@ -136,6 +194,12 @@ type Meta struct {
Generator string `json:"generator,omitempty"`
Model string `json:"model,omitempty"`
Instructions string `json:"instructions,omitempty"`
// LabelSource records how candidates were named to the picker. It is written
// for a seeded run too, even though that picker selects by index and never
// reads a label: it is the cell the run was assigned to, and the pair of
// seeded runs across the two label modes is the manipulation check that says
// how much of any difference is just application nondeterminism.
LabelSource string `json:"label_source,omitempty"`
// MaxSteps and DurationMillis are the budget the run was given, which has
// to be identical across arms for a comparison to mean anything.
MaxSteps int `json:"max_steps,omitempty"`
@@ -144,6 +208,11 @@ type Meta struct {
// several hosts, so a per-host effect has to be detectable rather than
// invisible.
Host string `json:"host,omitempty"`
// Device is the target the run drove, from --device. One host drives
// several emulators at different API levels, so without it a trace on its
// own cannot say what produced it and a per-device split can only be
// recovered by joining against the campaign manifest.
Device string `json:"device,omitempty"`
}
type Writer struct {
@@ -151,6 +220,10 @@ type Writer struct {
mutex sync.Mutex
file io.WriteCloser
encoder *json.Encoder
// llmCallFile is opened on the first WriteLLMCall, so a run whose picker
// never called a model leaves no llm-calls.jsonl behind at all.
llmCallFile io.WriteCloser
llmCallEncoder *json.Encoder
}
// NewWriter ensures `directory` exists and opens trace.jsonl for append.
@@ -181,12 +254,15 @@ func (w *Writer) WriteMeta(meta Meta) error {
return os.WriteFile(filepath.Join(w.directory, "meta.json"), body, 0o644)
}
// WriteStep stamps the format version itself so no caller can write a step
// that cannot be told apart from one written before the format changed.
func (w *Writer) WriteStep(step Step) error {
w.mutex.Lock()
defer w.mutex.Unlock()
if w.file == nil {
return fmt.Errorf("trace: writer is closed")
}
step.TraceVersion = TraceVersion
return w.encoder.Encode(step)
}
@@ -194,14 +270,27 @@ func (w *Writer) WriteStep(step Step) error {
// file via os.WriteFile and touches no field of Writer, so concurrent calls
// never contend.
func (w *Writer) WriteScreenshot(stepIndex int, png []byte) error {
return w.writePNG(fmt.Sprintf("step-%05d.png", stepIndex), png)
return w.writePNG(screenshotName(stepIndex), png)
}
func screenshotName(stepIndex int) string {
return fmt.Sprintf("step-%05d.png", stepIndex)
}
// ScreenshotReference is the run-relative path WriteScreenshot puts a step's
// screenshot at. Records that describe an image point at it instead of copying
// the bytes.
func ScreenshotReference(stepIndex int) string {
return screenshotDirectory + "/" + screenshotName(stepIndex)
}
const screenshotDirectory = "screenshots"
func (w *Writer) writePNG(name string, png []byte) error {
if len(png) == 0 {
return nil
}
directory := filepath.Join(w.directory, "screenshots")
directory := filepath.Join(w.directory, screenshotDirectory)
if err := os.MkdirAll(directory, 0o755); err != nil {
return fmt.Errorf("mkdir screenshots: %w", err)
}
@@ -211,10 +300,18 @@ func (w *Writer) writePNG(name string, png []byte) error {
func (w *Writer) Close() error {
w.mutex.Lock()
defer w.mutex.Unlock()
var err error
if w.llmCallFile != nil {
err = w.llmCallFile.Close()
w.llmCallFile = nil
w.llmCallEncoder = nil
}
if w.file == nil {
return nil
return err
}
if closeErr := w.file.Close(); closeErr != nil {
err = closeErr
}
err := w.file.Close()
w.file = nil
return err
}
+64 -8
View File
@@ -143,11 +143,6 @@ func TestWriteStep_HierarchyAndResidualsRoundTrip(t *testing.T) {
t.Errorf("residuals round-trip wrong: %s", got.Residuals["prop1"])
}
// Intentionally-lossy contract: Tree marshals only Elements (Root and
// Node.Children are json:"-"). The flat element list survives; tree
// structure does not. Lock both halves so a regression that drops the
// element list, or one that silently starts persisting structure the
// replay UI would then depend on, is caught.
if got.Hierarchy == nil {
t.Fatal("hierarchy dropped from trace")
}
@@ -157,8 +152,22 @@ func TestWriteStep_HierarchyAndResidualsRoundTrip(t *testing.T) {
if got.Hierarchy.Elements[1].Text != "hi" {
t.Errorf("element field lost: %+v", got.Hierarchy.Elements[1])
}
if got.Hierarchy.Root != nil {
t.Errorf("Root is json:\"-\" and must decode nil, got %+v", got.Hierarchy.Root)
if got.Hierarchy.Root == nil {
t.Fatal("tree structure not reconstructed from the stored form")
}
if got.Hierarchy.Root.ResourceID != "root" ||
len(got.Hierarchy.Root.Children) != 1 {
t.Fatalf("root rebuilt wrong: %+v", got.Hierarchy.Root.Element)
}
if resolved := got.Hierarchy.Find("id:child"); resolved == nil ||
resolved.Text != "hi" {
t.Errorf(
"selector resolves online but not against the decoded tree: %+v",
resolved,
)
}
if resolved := got.Hierarchy.Find("id:child"); resolved != got.Hierarchy.Elements[1] {
t.Error("decoded elements and the rebuilt nodes are different pointers")
}
}
@@ -448,11 +457,13 @@ func TestWriteMeta_ArmMembershipRoundTrip(t *testing.T) {
SanderlingVersion: "0.0.1",
Arm: "llm-visible-text",
Generator: "llm",
LabelSource: "visible-text",
Model: "claude-sonnet-5",
Instructions: "exercise the outbox",
MaxSteps: 300,
DurationMillis: 180000,
Host: "emulator-farm-01",
Device: "emulator-5556",
}
if err := writer.WriteMeta(meta); err != nil {
t.Fatal(err)
@@ -486,9 +497,54 @@ func TestWriteMeta_OmitsArmMembershipWhenUnset(t *testing.T) {
if err != nil {
t.Fatal(err)
}
for _, key := range []string{"arm", "generator", "model", "instructions", "max_steps", "duration_millis", "host"} {
for _, key := range []string{"arm", "generator", "label_source", "model", "instructions", "max_steps", "duration_millis", "host", "device"} {
if strings.Contains(string(body), `"`+key+`"`) {
t.Errorf("meta.json carries %q when unset:\n%s", key, body)
}
}
}
// TestStepPredatingTheFormatIsDistinguishable is the backward-compatibility
// contract: the existing corpus must still load, and a step from it must be
// separable from one this build wrote with nothing to report, or "this trace
// predates the format" reads as "this step had no logs".
func TestStepPredatingTheFormatIsDistinguishable(t *testing.T) {
const stored = `{"step":3,"timestamp":"2026-06-10T21:22:07Z",` +
`"hierarchy":{"elements":[{"resourceId":"root"},{"resourceId":"child"}]}}`
var old Step
if err := json.Unmarshal([]byte(stored), &old); err != nil {
t.Fatalf(
"a trace written before the format change no longer loads: %v",
err,
)
}
if old.TraceVersion != 0 {
t.Errorf(
"trace_version = %d, want 0 for a step that predates the field",
old.TraceVersion,
)
}
if len(old.Hierarchy.Elements) != 2 || old.Hierarchy.Root != nil {
t.Errorf("old hierarchy reinterpreted: elements=%d root=%v",
len(old.Hierarchy.Elements), old.Hierarchy.Root)
}
directory := t.TempDir()
writer, _ := NewWriter(directory)
defer writer.Close()
if err := writer.WriteStep(Step{Index: 3}); err != nil {
t.Fatal(err)
}
body, _ := os.ReadFile(filepath.Join(directory, "trace.jsonl"))
var fresh Step
if err := json.Unmarshal(body, &fresh); err != nil {
t.Fatal(err)
}
if fresh.TraceVersion != TraceVersion {
t.Errorf("a step with nothing to report stamped version %d, want %d",
fresh.TraceVersion, TraceVersion)
}
if len(fresh.Logs) != 0 {
t.Errorf("logs = %v, want none", fresh.Logs)
}
}
+113
View File
@@ -0,0 +1,113 @@
// Package tracecorpus loads recorded runs for measures that read stored
// traces with no device attached.
package tracecorpus
import (
"bufio"
"encoding/json"
"fmt"
"os"
"path/filepath"
"sort"
"github.com/priyanshujain/sanderling/internal/trace"
)
const maxStepBytes = 64 * 1024 * 1024
// Run is one run directory: its meta and every step in file order, including
// the synthetic end-of-run record a finalised liveness obligation lands on.
type Run struct {
Directory string
Meta trace.Meta
Steps []trace.Step
}
// Load reads a run directory and refuses anything an offline measure cannot
// read. A step written before the format change stores no element depths, so
// its hierarchy decodes with a nil root: a structural hash over it is the
// empty string for every screen, and every run would look identical to every
// other. The refusal has to name the version rather than report that number.
func Load(directory string) (Run, error) {
metaBody, err := os.ReadFile(filepath.Join(directory, "meta.json"))
if err != nil {
return Run{}, fmt.Errorf("read meta: %w", err)
}
var meta trace.Meta
if err := json.Unmarshal(metaBody, &meta); err != nil {
return Run{}, fmt.Errorf("decode meta: %w", err)
}
file, err := os.Open(filepath.Join(directory, "trace.jsonl"))
if err != nil {
return Run{}, fmt.Errorf("open trace: %w", err)
}
defer file.Close()
run := Run{Directory: directory, Meta: meta}
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 1024*1024), maxStepBytes)
line := 0
for scanner.Scan() {
line++
if len(scanner.Bytes()) == 0 {
continue
}
var step trace.Step
if err := json.Unmarshal(scanner.Bytes(), &step); err != nil {
return Run{}, fmt.Errorf("decode step on line %d: %w", line, err)
}
if step.TraceVersion != trace.TraceVersion {
return Run{}, fmt.Errorf(
"step %d is trace_version %d and an offline measure reads version %d only: "+
"an older step stores no element depths, no logs and no exceptions, so "+
"its hierarchy decodes with a nil root and nothing retrofits it",
step.Index, step.TraceVersion, trace.TraceVersion)
}
if step.Hierarchy != nil && len(step.Hierarchy.Elements) > 0 &&
step.Hierarchy.Root == nil {
return Run{}, fmt.Errorf(
"step %d stores %d elements and no tree shape, so its hierarchy "+
"cannot be rebuilt",
step.Index, len(step.Hierarchy.Elements))
}
run.Steps = append(run.Steps, step)
}
if err := scanner.Err(); err != nil {
return Run{}, fmt.Errorf("read trace: %w", err)
}
if len(run.Steps) == 0 {
return Run{}, fmt.Errorf("trace has no steps")
}
return run, nil
}
// Discover finds every run directory at or below root, a run directory being
// one holding both meta.json and trace.jsonl.
func Discover(root string) ([]string, error) {
var directories []string
err := filepath.Walk(
root,
func(path string, info os.FileInfo, err error) error {
if err != nil {
return err
}
if !info.IsDir() {
return nil
}
if _, statErr := os.Stat(filepath.Join(path, "trace.jsonl")); statErr != nil {
return nil
}
if _, statErr := os.Stat(filepath.Join(path, "meta.json")); statErr != nil {
return nil
}
directories = append(directories, path)
return nil
},
)
if err != nil {
return nil, err
}
sort.Strings(directories)
return directories, nil
}
+122
View File
@@ -0,0 +1,122 @@
package tracecorpus
import (
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/trace"
)
const oneScreen = `{"attributes": {"text": "Hi", "bounds": "[0,0,10,10]"}, "children": [
{"attributes": {"text": "child"}, "children": []}
]}`
func TestLoadRebuildsTheTreeAStepStored(t *testing.T) {
directory := writeRun(t, oneScreen)
run, err := Load(directory)
if err != nil {
t.Fatalf("load: %v", err)
}
if len(run.Steps) != 1 {
t.Fatalf("steps = %d, want 1", len(run.Steps))
}
root := run.Steps[0].Hierarchy.Root
if root == nil {
t.Fatal("stored hierarchy came back with no root, so no selector or hash reads it")
}
if len(root.Children) != 1 || root.Children[0].Text != "child" {
t.Fatalf("rebuilt tree lost its child: %+v", root)
}
}
func TestLoadRefusesAVersionZeroStepByVersion(t *testing.T) {
directory := writeRun(t, oneScreen)
downgrade(t, filepath.Join(directory, "trace.jsonl"))
_, err := Load(directory)
if err == nil {
t.Fatal("a version 0 step must be refused, not measured")
}
if !strings.Contains(err.Error(), "trace_version 0") {
t.Fatalf("refusal must name the version, got %q", err)
}
}
// TestLoadRefusesElementsWithNoStoredShape covers the failure that would be
// silent: a tree that decodes with elements and no root hashes to the empty
// string, which reads as one state shared by every screen.
func TestLoadRefusesElementsWithNoStoredShape(t *testing.T) {
directory := writeRun(t, oneScreen)
path := filepath.Join(directory, "trace.jsonl")
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
var step map[string]any
if err := json.Unmarshal(body, &step); err != nil {
t.Fatal(err)
}
hierarchyField := step["hierarchy"].(map[string]any)
delete(hierarchyField, "depths")
rewrite(t, path, step)
if _, err := Load(directory); err == nil ||
!strings.Contains(err.Error(), "cannot be rebuilt") {
t.Fatalf("a shapeless tree must be refused, got %v", err)
}
}
func writeRun(t *testing.T, dumps ...string) string {
t.Helper()
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteMeta(trace.Meta{Seed: 7, Platform: "web"}); err != nil {
t.Fatal(err)
}
for index, dump := range dumps {
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteStep(trace.Step{Index: index + 1, Hierarchy: tree}); err != nil {
t.Fatal(err)
}
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
return directory
}
func downgrade(t *testing.T, path string) {
t.Helper()
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
var step map[string]any
if err := json.Unmarshal(body, &step); err != nil {
t.Fatal(err)
}
delete(step, "trace_version")
rewrite(t, path, step)
}
func rewrite(t *testing.T, path string, step map[string]any) {
t.Helper()
body, err := json.Marshal(step)
if err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, append(body, '\n'), 0o644); err != nil {
t.Fatal(err)
}
}
+78 -1
View File
@@ -3,6 +3,7 @@ package verifier
import (
"os"
"strconv"
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/hierarchy"
@@ -66,7 +67,7 @@ func TestStateAxObjectSelectorTestTagAlias(t *testing.T) {
func TestStateAxFindWorks(t *testing.T) {
jsonText, err := os.ReadFile("testdata/ax_find_tree.json")
if err != nil {
t.Skip("ax_find_tree.json fixture unreadable")
t.Fatalf("committed fixture unreadable, so the round trip never ran: %v", err)
}
tree, err := hierarchy.Parse(string(jsonText))
if err != nil {
@@ -101,6 +102,82 @@ func TestStateAxFindWorks(t *testing.T) {
}
}
// A selector key that can never match is a spec bug, and an empty result hides
// it: the generator yields no action, the runner waits out the step, and the
// run ends clean having explored nothing. The spec must fail instead.
func TestStateAxObjectSelectorRejectsAnUnknownKey(t *testing.T) {
tree, err := hierarchy.Parse(`{
"attributes": {"resource-id": "root"},
"children": [{"attributes": {"content-desc": "Supplier"}, "children": []}]
}`)
if err != nil {
t.Fatal(err)
}
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.probe = __sanderling__.extract(state => !!state.ax.find({ descripton: "Supplier" }));
`)
err = verifier.PushSnapshot(SnapshotInput{Snapshots: Snapshots{}, Tree: tree})
if err == nil {
t.Fatal("expected an unknown selector key to fail the spec")
}
if !strings.Contains(err.Error(), "descripton") {
t.Errorf("error does not name the offending key: %v", err)
}
if !strings.Contains(err.Error(), "accepted keys") {
t.Errorf("error does not list the accepted keys: %v", err)
}
}
// desc names the accessibility description in the element fields and in the
// string form, so the object form answers to it too rather than reporting it as
// a mistake.
func TestStateAxObjectSelectorAcceptsDesc(t *testing.T) {
tree, err := hierarchy.Parse(`{
"attributes": {"resource-id": "root"},
"children": [{"attributes": {"content-desc": "Supplier"}, "children": []}]
}`)
if err != nil {
t.Fatal(err)
}
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.probe = __sanderling__.extract(state => state.ax.find({ desc: "Supplier" }) ? "matched" : "miss");
`)
if err := verifier.PushSnapshot(SnapshotInput{Snapshots: Snapshots{}, Tree: tree}); err != nil {
t.Fatal(err)
}
got := verifier.runtime.GlobalObject().Get("probe").ToObject(verifier.runtime).Get("current").String()
if got != "matched" {
t.Fatalf("probe = %q, want matched", got)
}
}
// A key that belongs to another platform must stay silent: one spec runs on
// every platform, and iOS-only attributes are absent from an Android tree by
// design rather than by mistake.
func TestStateAxObjectSelectorKeepsCrossPlatformKeysSilent(t *testing.T) {
tree, err := hierarchy.Parse(`{
"attributes": {"resource-id": "root"},
"children": []
}`)
if err != nil {
t.Fatal(err)
}
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.probe = __sanderling__.extract(state => state.ax.find({ title: "Settings" }) ? "matched" : "miss");
`)
if err := verifier.PushSnapshot(SnapshotInput{Snapshots: Snapshots{}, Tree: tree}); err != nil {
t.Fatalf("a platform-specific key must not fail the run: %v", err)
}
got := verifier.runtime.GlobalObject().Get("probe").ToObject(verifier.runtime).Get("current").String()
if got != "miss" {
t.Fatalf("probe = %q, want miss", got)
}
}
// axSelectorFormsTree carries one node per id shape a dump produces: the bare
// tag Compose and the web driver emit, the package-qualified resource id
// Android emits, and the iOS accessibility identifier.
+47
View File
@@ -168,3 +168,50 @@ globalThis.properties = {
t.Errorf("residual lost the bound: %s", residual)
}
}
// TestTopLevelEventuallyWithinSteps_KeepsAuthoredWindow drives the shape the
// folio-web reachability properties now use, `eventually(p).within(n,
// "steps")`, through the goja runtime and pins what the trace records: one
// obligation, the window the spec authored, and the observation it closes at.
// Bug class: the residual reporting the part of the window that is left, so a
// reader cannot tell what the spec asked for or when the promise comes due.
func TestTopLevelEventuallyWithinSteps_KeepsAuthoredWindow(t *testing.T) {
const source = `
globalThis.seen = __sanderling__.extract(state => state.snapshots["seen"] ?? false, "seen");
globalThis.properties = {
reachable: __sanderling__.eventually(() => seen.current).within(1915, 'steps'),
};
`
verifier := newVerifier(t)
mustLoad(t, verifier, source)
base := time.Unix(1780000000, 0)
const steps = 40
for index := range steps {
if err := verifier.PushSnapshot(SnapshotInput{
Snapshots: Snapshots{"seen": json.RawMessage(`false`)},
StepIndex: index + 1,
StepTime: base.Add(time.Duration(index) * 1200 * time.Millisecond),
RunStart: base,
}); err != nil {
t.Fatal(err)
}
if got := verifier.EvaluateProperties()["reachable"]; got != ltl.VerdictPending {
t.Fatalf("step %d: got %v, want pending", index+1, got)
}
}
residual, err := json.Marshal(verifier.Residuals()["reachable"])
if err != nil {
t.Fatal(err)
}
if strings.Contains(string(residual), `"op":"and"`) {
t.Errorf("residual accumulated obligations: %s", residual)
}
if !strings.Contains(string(residual), `"amount":1915`) || !strings.Contains(string(residual), `"unit":"steps"`) {
t.Errorf("residual lost the authored window: %s", residual)
}
if !strings.Contains(string(residual), `"expiresAtObservation":1915`) {
t.Errorf("residual lost the closing observation: %s", residual)
}
}
@@ -43,6 +43,7 @@ const canonicalElement = `{
"enabled": true,
"focused": false,
"id": "TxnAmountField",
"secure": null,
"selected": false,
"text": "199",
"x": 200,
+350 -83
View File
@@ -65,6 +65,13 @@ func (v *Verifier) Screenshot() []byte {
return v.lastScreenshot
}
// SnapshotStep returns the step index of the most recent PushSnapshot. It lags
// the runner's current step whenever an observation was skipped (a transitional
// tree), which is exactly when Screenshot returns an older step's image.
func (v *Verifier) SnapshotStep() int {
return v.stepIndex
}
// CurrentScreen returns the screen id of the most recent snapshot's first
// element, matching the runner's own screen labeling. Empty when no tree is
// loaded.
@@ -75,6 +82,12 @@ func (v *Verifier) CurrentScreen() string {
return v.lastTree.Elements[0].Screen
}
// Tree returns the hierarchy of the most recent snapshot, nil when none was
// pushed. It is what a caller resolves an action's target against.
func (v *Verifier) Tree() *hierarchy.Tree {
return v.lastTree
}
// SampleInput draws one InputText value from the shared corpus via the bundled
// __sanderlingSampleInput__. It errors when the bundle did not install the
// callable (a raw-JS fixture) so the caller can skip typing rather than send an
@@ -106,9 +119,12 @@ type ActionCandidate struct {
// Kind is the resulting action kind.
Kind ActionKind
// Description is the rendered action shown to the model and echoed back as
// chosen_action, e.g. `Tap "Add credit"`. Dedup keys on it, so it is unique.
// chosen_action, e.g. `Tap "Add credit"`. Two entries may render the same
// when the screen holds two controls a user reads alike; Index is what tells
// them apart, and it is what the model picks by.
Description string
// Label is the visible-text target label (empty for gestures).
// Label is the target label the selected LabelSource named (empty for
// gestures).
Label string
// Weight is the effective selection weight as a percentage (1..100),
// meaningful only when Weighted is true (the tree used `weighted`).
@@ -129,72 +145,122 @@ type ActionCandidate struct {
// prob is the internal accumulated selection probability, summed across
// dedup, then rounded into Weight. Not exposed in the prompt directly.
prob float64
// secure is what the target reports about being a secure text entry, which
// is what Description is rendered under. It is not part of what the model
// sees.
secure secureFact
}
// maxLabelRunes caps a visible-text label so joined descendant text stays short
// enough to render on one numbered line.
const maxLabelRunes = 40
// gestureDurationMillis is how long a drag takes when the descriptor did not say.
// It mirrors DEFAULT_SWIPE_DURATION in runtime-entry.ts, which is what the
// seeded policy's action carries by the time it reaches the runner: the two
// policies must hand the driver the same gesture, not two speeds of it.
const gestureDurationMillis = 250
// The label sources a candidate's target can be named by. This is the
// observation channel the model reads, and nothing else: the seeded picker
// selects by index and never asks for a label, so the two seeded cells of a
// labelling factorial draw the identical stream.
const (
// LabelSourceVisibleText names a control by what a user would read. It is
// the default, and the channel every run so far was produced with.
LabelSourceVisibleText = "visible-text"
// LabelSourceResourceID names a control by the identifier the app assigned
// it, which no user ever sees.
LabelSourceResourceID = "resource-id"
)
// Candidates enumerates every action the spec's weighted actionsRoot yields at
// the current step, each tagged with a plainly-worded description and its
// effective weight, for the LLM generator to pick one number from. It walks the
// SAME tree the seeded picker draws: weighted branches recurse (accumulating the
// selection probability), authored actions()/whenRoute leaves are called once
// for their concrete actions, and builtin verbs come straight from the picker's
// own enumeration. Identical descriptions dedup, summing weight.
func (v *Verifier) Candidates() []ActionCandidate {
// own enumeration. Candidates that would execute the same action dedup, summing
// weight; two controls that merely read alike stay two entries.
//
// labelSource selects the channel each target is named by. An unrecognized
// value (including the zero value) names targets by visible text; the CLI
// rejects an unknown mode before a run starts, so only a test reaches that.
//
// The error is the spec refusing to be run by this policy at all: an authored
// leaf that samples one of several items reaches the seeded picker's rng but
// never this walk, so the model would be offered a fixed first item forever. It
// names the leaf and it is fatal, because degrading to that fixed item silently
// is what makes a policy comparison meaningless.
func (v *Verifier) Candidates(labelSource string) ([]ActionCandidate, error) {
if v.lastTree == nil {
return nil
return nil, nil
}
root := v.runtime.GlobalObject().Get("actions")
if root == nil || goja.IsUndefined(root) || goja.IsNull(root) {
return nil
return nil, nil
}
nodeIndex := buildNodeIndex(v.lastTree)
labels := labelContext{nodeIndex: buildNodeIndex(v.lastTree), source: labelSource}
var raw []ActionCandidate
v.collectNode(root, 1.0, false, nodeIndex, &raw)
return finalizeCandidates(raw)
v.setEnumeratingCandidates(true)
defer v.setEnumeratingCandidates(false)
if err := v.collectNode(root, 1.0, false, labels, &raw); err != nil {
return nil, err
}
return finalizeCandidates(raw), nil
}
// setEnumeratingCandidates tells the spec bundle that the authored leaves are
// being called by this policy rather than by the picker. A spec loaded without
// the runtime entry (a raw-JS unit fixture) has no such callable, and no
// sampler to refuse either.
func (v *Verifier) setEnumeratingCandidates(enumerating bool) {
if v.setEnumeratingCandidatesFn == nil {
return
}
_, _ = v.setEnumeratingCandidatesFn(goja.Undefined(), v.runtime.ToValue(enumerating))
}
// collectNode dispatches one GeneratorNode of the action tree. prob is the
// accumulated probability the seeded picker reaches this node; weighted records
// whether any weighted node lies on the path (so weights are shown only when the
// spec actually declared them).
func (v *Verifier) collectNode(node goja.Value, prob float64, weighted bool, nodeIndex map[*hierarchy.Element]*hierarchy.Node, out *[]ActionCandidate) {
func (v *Verifier) collectNode(node goja.Value, prob float64, weighted bool, labels labelContext, out *[]ActionCandidate) error {
object := node.ToObject(v.runtime)
if object == nil {
return
return nil
}
kind := object.Get("kind")
if kind == nil || goja.IsUndefined(kind) {
return
return nil
}
switch kind.String() {
case "weighted":
v.collectWeighted(object, prob, nodeIndex, out)
return v.collectWeighted(object, prob, labels, out)
case "actions":
v.collectActions(object, prob, weighted, nodeIndex, out)
return v.collectActions(object, prob, weighted, labels, out)
case "builtin":
verb := object.Get("verb")
if verb != nil && !goja.IsUndefined(verb) {
v.collectBuiltin(verb.String(), prob, weighted, nodeIndex, out)
v.collectBuiltin(verb.String(), prob, weighted, labels, out)
}
case "llm":
// The llm marker is the generator, not part of the candidate tree.
}
return nil
}
// collectWeighted recurses each branch, splitting the incoming probability by
// the branch weight over the sibling total (matching the seeded picker's single
// weighted draw).
func (v *Verifier) collectWeighted(object *goja.Object, prob float64, nodeIndex map[*hierarchy.Element]*hierarchy.Node, out *[]ActionCandidate) {
func (v *Verifier) collectWeighted(object *goja.Object, prob float64, labels labelContext, out *[]ActionCandidate) error {
branches := object.Get("branches")
if branches == nil {
return
return nil
}
array := branches.ToObject(v.runtime)
if array == nil {
return
return nil
}
length := int(array.Get("length").ToInteger())
weights := make([]float64, length)
@@ -215,36 +281,51 @@ func (v *Verifier) collectWeighted(object *goja.Object, prob float64, nodeIndex
total += weight
}
if total <= 0 {
return
return nil
}
for i := range length {
if children[i] == nil {
continue
}
v.collectNode(children[i], prob*weights[i]/total, true, nodeIndex, out)
// The branch number is the author's own path to a refused leaf, which
// its closure source alone does not give when the leaf is a whenRoute
// (whose closure belongs to the library, not the spec).
if err := v.collectNode(children[i], prob*weights[i]/total, true, labels, out); err != nil {
return fmt.Errorf("branch %d: %w", i+1, err)
}
}
return nil
}
// collectActions calls an authored leaf's generator once (safe: it reads state
// and, off-route, returns []), turning each concrete descriptor into a
// candidate. It runs OUTSIDE the picker's rng scope, so from(...).generate()
// draws nothing and no seed advances.
func (v *Verifier) collectActions(object *goja.Object, prob float64, weighted bool, nodeIndex map[*hierarchy.Element]*hierarchy.Node, out *[]ActionCandidate) {
generate, ok := goja.AssertFunction(object.Get("generate"))
//
// A generator that throws for its own reasons still contributes nothing and
// nothing more: this walk calls EVERY leaf every step, including leaves the
// seeded picker would have walked once in a hundred steps, so promoting those
// throws would kill runs the seeded arm survives.
func (v *Verifier) collectActions(object *goja.Object, prob float64, weighted bool, labels labelContext, out *[]ActionCandidate) error {
generatorValue := object.Get("generate")
generate, ok := goja.AssertFunction(generatorValue)
if !ok {
return
return nil
}
result, err := generate(goja.Undefined())
if err != nil {
return
if refusal, refused := v.samplerRefusal(err); refused {
return fmt.Errorf("authored action %s %s", authoredLeafIdentity(generatorValue), refusal)
}
return nil
}
array := result.ToObject(v.runtime)
if array == nil {
return
return nil
}
length := int(array.Get("length").ToInteger())
for i := range length {
candidate, ok := v.candidateFromDescriptor(array.Get(strconv.Itoa(i)), nodeIndex)
candidate, ok := v.candidateFromDescriptor(array.Get(strconv.Itoa(i)), labels)
if !ok {
continue
}
@@ -252,12 +333,55 @@ func (v *Verifier) collectActions(object *goja.Object, prob float64, weighted bo
candidate.Weighted = weighted
*out = append(*out, candidate)
}
return nil
}
// samplerRefusalName is the error name pkg/spec/src/sampler-rng.ts stamps on the
// refusal it throws, which is what tells that refusal apart from a spec's own
// runtime errors.
const samplerRefusalName = "SanderlingSamplerRefusal"
// samplerRefusal reports the refusal message when the authored leaf declined to
// sample for this policy.
func (v *Verifier) samplerRefusal(err error) (string, bool) {
var exception *goja.Exception
if !errors.As(err, &exception) {
return "", false
}
value := exception.Value()
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return "", false
}
thrown := value.ToObject(v.runtime)
if thrown == nil || stringField(thrown, "name") != samplerRefusalName {
return "", false
}
return stringField(thrown, "message"), true
}
// maxLeafSourceRunes caps the generator excerpt that names a leaf in an error.
const maxLeafSourceRunes = 160
// authoredLeafIdentity renders the leaf's generator source on one line. An
// authored leaf is an anonymous closure among identical-looking tree nodes, so
// its source is the handle an author can search the spec for.
func authoredLeafIdentity(generator goja.Value) string {
source := []rune(strings.Join(strings.Fields(generator.String()), " "))
if len(source) > maxLeafSourceRunes {
return strconv.Quote(string(source[:maxLeafSourceRunes]) + "...")
}
return strconv.Quote(string(source))
}
// candidateFromDescriptor lowers one authored ActionDescriptor (as a goja
// object) into a ready-to-run candidate, resolving the target's coordinates,
// selector, and visible-text label. Actions on a disabled control are dropped.
func (v *Verifier) candidateFromDescriptor(value goja.Value, nodeIndex map[*hierarchy.Element]*hierarchy.Node) (ActionCandidate, bool) {
// selector, and label.
//
// A disabled target is offered like any other. The seeded picker executes
// whatever the leaf authored, disabled or not, and attempting a disabled
// control is where boundary defects live: a control the app forgot to re-enable
// reads as disabled, and a policy that cannot attempt it cannot find that.
func (v *Verifier) candidateFromDescriptor(value goja.Value, labels labelContext) (ActionCandidate, bool) {
object := value.ToObject(v.runtime)
if object == nil {
return ActionCandidate{}, false
@@ -269,8 +393,8 @@ func (v *Verifier) candidateFromDescriptor(value goja.Value, nodeIndex map[*hier
kind := ActionKind(kindValue.String())
switch kind {
case ActionKindTap, ActionKindDoubleTap, ActionKindLongPress:
target := v.resolveTarget(object.Get("on"), nodeIndex)
if target.disabled {
target, ok := v.resolveTarget(object.Get("on"), labels)
if !ok {
return ActionCandidate{}, false
}
return ActionCandidate{
@@ -279,8 +403,8 @@ func (v *Verifier) candidateFromDescriptor(value goja.Value, nodeIndex map[*hier
Action: Action{Kind: kind, On: target.selector, X: target.x, Y: target.y},
}, true
case ActionKindInputText:
target := v.resolveTarget(object.Get("into"), nodeIndex)
if target.disabled {
target, ok := v.resolveTarget(object.Get("into"), labels)
if !ok {
return ActionCandidate{}, false
}
text := stringField(object, "text")
@@ -288,29 +412,45 @@ func (v *Verifier) candidateFromDescriptor(value goja.Value, nodeIndex map[*hier
Kind: kind,
Label: target.label,
InputType: target.inputType,
secure: target.secure,
Action: Action{Kind: kind, On: target.selector, X: target.x, Y: target.y, Text: text},
}, true
case ActionKindScroll:
target := v.resolveTarget(object.Get("in"), nodeIndex)
container, _ := v.resolveTarget(object.Get("in"), labels)
direction := stringField(object, "direction")
if direction == "" {
direction = "down"
}
return ActionCandidate{
Kind: kind,
Direction: direction,
Action: Action{Kind: kind, On: target.selector, Direction: direction},
}, true
action := Action{
Kind: kind,
On: container.selector,
Direction: direction,
DurationMillis: gestureDurationMillis,
}
// Endpoints only when the descriptor computed the whole gesture (the
// builtin generator does). Anchoring an authored scroll on the
// container's own point instead would hand the runner a drag from a
// point to itself, which it executes as written.
from, hasFrom := v.resolveTarget(object.Get("from"), labels)
to, hasTo := v.resolveTarget(object.Get("to"), labels)
if hasFrom && hasTo {
action.FromX, action.FromY = from.x, from.y
action.ToX, action.ToY = to.x, to.y
}
return ActionCandidate{Kind: kind, Direction: direction, Action: action}, true
case ActionKindSwipe:
from := v.resolveTarget(object.Get("from"), nodeIndex)
to := v.resolveTarget(object.Get("to"), nodeIndex)
from, hasFrom := v.resolveTarget(object.Get("from"), labels)
to, hasTo := v.resolveTarget(object.Get("to"), labels)
if !hasFrom || !hasTo {
return ActionCandidate{}, false
}
return ActionCandidate{
Kind: kind,
Action: Action{
Kind: kind,
FromX: from.x, FromY: from.y,
ToX: to.x, ToY: to.y,
DurationMillis: intField(object, "durationMillis"),
DurationMillis: intFieldOr(object, "durationMillis", gestureDurationMillis),
},
}, true
case ActionKindPressKey:
@@ -319,7 +459,10 @@ func (v *Verifier) candidateFromDescriptor(value goja.Value, nodeIndex map[*hier
Action: Action{Kind: kind, Key: stringField(object, "key")},
}, true
case ActionKindWait:
return ActionCandidate{Kind: kind, Action: Action{Kind: kind}}, true
return ActionCandidate{
Kind: kind,
Action: Action{Kind: kind, DurationMillis: intField(object, "durationMillis")},
}, true
default:
return ActionCandidate{}, false
}
@@ -332,54 +475,66 @@ type resolvedTarget struct {
selector string
label string
inputType string
disabled bool
secure secureFact
}
// resolveTarget reads an authored action's target. Ax element handles carry
// x/y/__sanderlingSelector plus their own text; a bare selector string resolves
// against the current tree; a point carries geometry only.
func (v *Verifier) resolveTarget(value goja.Value, nodeIndex map[*hierarchy.Element]*hierarchy.Node) resolvedTarget {
// x/y/__sanderlingSelector plus their own text and id; a bare selector string
// resolves against the current tree; a point carries geometry only.
//
// The second return is false when the value names no target the seeded policy
// could act on either: runtime-entry.ts pointOf accepts a non-empty selector
// string or an object with numeric coordinates, and drops the whole action
// otherwise. Lowering one of those to (0, 0) instead would offer the model an
// action the seeded policy never takes, aimed at the screen corner.
func (v *Verifier) resolveTarget(value goja.Value, labels labelContext) (resolvedTarget, bool) {
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return resolvedTarget{}
return resolvedTarget{}, false
}
if selector, ok := value.Export().(string); ok {
return v.targetFromSelector(selector, nodeIndex)
if selector == "" {
return resolvedTarget{}, false
}
return v.targetFromSelector(selector, labels), true
}
object := value.ToObject(v.runtime)
if object == nil {
return resolvedTarget{}
return resolvedTarget{}, false
}
x, hasX := numberField(object, "x")
y, hasY := numberField(object, "y")
if !hasX || !hasY {
return resolvedTarget{}, false
}
selector := stringField(object, tagSelector)
if selector == "" {
selector = stringField(object, "selector")
}
target := resolvedTarget{
x: int(object.Get("x").ToInteger()),
y: int(object.Get("y").ToInteger()),
selector: selector,
}
target := resolvedTarget{x: x, y: y, selector: selector, secure: secureFactFromHandle(object)}
if element := v.findBySelector(selector); element != nil {
target.label = visibleLabel(element, nodeIndex)
target.label = labels.label(element)
target.inputType = inputTypeHint(element)
target.disabled = !element.Enabled && hasEnabled(element)
if !target.secure.reported {
target.secure = secureFactOf(element)
}
}
if target.label == "" {
target.label = truncateLabel(stringField(object, "text"))
target.label = truncateLabel(v.handleLabel(object, labels))
}
return target
return target, true
}
// targetFromSelector resolves a bare selector-string target against the tree.
func (v *Verifier) targetFromSelector(selector string, nodeIndex map[*hierarchy.Element]*hierarchy.Node) resolvedTarget {
func (v *Verifier) targetFromSelector(selector string, labels labelContext) resolvedTarget {
target := resolvedTarget{selector: selector}
element := v.findBySelector(selector)
if element == nil {
return target
}
target.x, target.y = element.Bounds.Center()
target.label = visibleLabel(element, nodeIndex)
target.label = labels.label(element)
target.inputType = inputTypeHint(element)
target.disabled = !element.Enabled && hasEnabled(element)
target.secure = secureFactOf(element)
return target
}
@@ -396,7 +551,7 @@ func (v *Verifier) findBySelector(selector string) *hierarchy.Element {
// the two policies select over one action space and cannot drift apart. Each
// entry's action arrives on the wire contract DecodeAction already reads, so a
// chosen candidate executes the action the seeded draw would have executed.
func (v *Verifier) collectBuiltin(verb string, prob float64, weighted bool, nodeIndex map[*hierarchy.Element]*hierarchy.Node, out *[]ActionCandidate) {
func (v *Verifier) collectBuiltin(verb string, prob float64, weighted bool, labels labelContext, out *[]ActionCandidate) {
entries, err := v.enumerateBuiltin(verb)
if err != nil {
return
@@ -415,8 +570,9 @@ func (v *Verifier) collectBuiltin(verb string, prob float64, weighted bool, node
}
if entry.targetIndex >= 0 && entry.targetIndex < len(targets) {
element := targets[entry.targetIndex].element
candidate.Label = visibleLabel(element, nodeIndex)
candidate.Label = labels.label(element)
candidate.InputType = inputTypeHint(element)
candidate.secure = secureFactOf(element)
}
*out = append(*out, candidate)
}
@@ -466,20 +622,33 @@ func (v *Verifier) enumerateBuiltin(verb string) ([]builtinCandidate, error) {
return entries, nil
}
// finalizeCandidates renders each candidate's description, dedups identical
// descriptions (summing weight), numbers the survivors 1..N, and rounds the
// accumulated probability into a percentage Weight.
// candidateIdentity is what a candidate would DO. Two candidates sharing it are
// the same action reached through two paths of the action tree, so folding them
// into one numbered entry loses nothing; two that differ are different actions
// however alike they read, so folding them would put one of them out of reach.
// llmText is part of it because it decides where the typed text comes from: the
// model writes it for a builtin typing candidate, while an authored one replays
// the value already sitting in Action.Text.
type candidateIdentity struct {
action Action
llmText bool
}
// finalizeCandidates renders each candidate's description, dedups by what the
// candidate executes (summing weight), numbers the survivors 1..N, and rounds
// the accumulated probability into a percentage Weight.
func finalizeCandidates(raw []ActionCandidate) []ActionCandidate {
seen := make(map[string]int, len(raw))
seen := make(map[candidateIdentity]int, len(raw))
result := make([]ActionCandidate, 0, len(raw))
for _, candidate := range raw {
candidate.Description = describeCandidate(candidate)
if index, ok := seen[candidate.Description]; ok {
identity := candidateIdentity{action: candidate.Action, llmText: candidate.LLMText}
if index, ok := seen[identity]; ok {
result[index].prob += candidate.prob
result[index].Weighted = result[index].Weighted || candidate.Weighted
continue
}
seen[candidate.Description] = len(result)
seen[identity] = len(result)
result = append(result, candidate)
}
for i := range result {
@@ -492,8 +661,11 @@ func finalizeCandidates(raw []ActionCandidate) []ActionCandidate {
}
// describeCandidate renders the plain, echo-friendly description shown in the
// numbered list. It is the dedup key, so it must be stable and unique per
// distinct action.
// numbered list. It is display only: dedup keys on the action, so two entries
// may read alike without merging and Index is what separates them. Do not add an
// ordinal or a coordinate to pull those apart: this string IS the observation
// channel a labelling experiment varies, so a disambiguator here would name a
// target through a channel the label source deliberately withholds.
func describeCandidate(candidate ActionCandidate) string {
switch candidate.Kind {
case ActionKindTap:
@@ -509,14 +681,16 @@ func describeCandidate(candidate ActionCandidate) string {
}
return fmt.Sprintf("Type into %q", candidate.Label)
}
return fmt.Sprintf("Type %q into %q", candidate.Action.Text, candidate.Label)
return fmt.Sprintf("Type %q into %q",
recordedInputText(candidate.Action.Text, candidate.secure), candidate.Label)
case ActionKindScroll:
return "Scroll " + candidate.Direction
case ActionKindSwipe:
// A swipe carries endpoints and no selector, so the coordinates are what
// keep two swipes distinct. The label is prepended when the origin
// element has one, because "swipe that row" is the interaction a model
// reaches for and a bare pair of points does not say which row.
// A swipe is a drag across the screen, so where it runs is the whole of
// what it does and a reader needs the endpoints to picture it. The label
// is prepended when the origin element has one, because "swipe that row"
// is the interaction a model reaches for and a bare pair of points does
// not say which row.
where := fmt.Sprintf("from (%d,%d) to (%d,%d)",
candidate.Action.FromX, candidate.Action.FromY,
candidate.Action.ToX, candidate.Action.ToY)
@@ -551,6 +725,66 @@ func buildNodeIndex(tree *hierarchy.Tree) map[*hierarchy.Element]*hierarchy.Node
return index
}
// labelContext carries what naming a candidate's target takes: the node index
// descendant text is borrowed through, and the channel the name comes from.
type labelContext struct {
nodeIndex map[*hierarchy.Element]*hierarchy.Node
source string
}
func (l labelContext) label(element *hierarchy.Element) string {
if l.source == LabelSourceResourceID {
return resourceIdentifierLabel(element)
}
return visibleLabel(element, l.nodeIndex)
}
// handleLabel names a target from the ax handle alone, for the web tick path
// where the handle was built in V8 and carries no selector to resolve against
// the tree. It walks visibleLabel's rungs over the fields a handle has: an
// editable field's hint names its purpose, its own text is the transient typed
// value. The identifier arm reads the handle's id and nothing a user could
// read, which is the one thing that arm must not see.
func (v *Verifier) handleLabel(object *goja.Object, labels labelContext) string {
if labels.source == LabelSourceResourceID {
return stringField(object, "id")
}
hint := v.handleAttribute(object, "hintText")
if hint != "" && boolField(object, "editable") {
return hint
}
if text := stringField(object, "text"); text != "" {
return text
}
if desc := stringField(object, "desc"); desc != "" {
return desc
}
return hint
}
func (v *Verifier) handleAttribute(object *goja.Object, name string) string {
attrs := object.Get("attrs")
if attrs == nil || goja.IsUndefined(attrs) || goja.IsNull(attrs) {
return ""
}
return stringField(attrs.ToObject(v.runtime), name)
}
// resourceIdentifierLabel names a control by the identifier the app assigned it,
// then by its class, then by a bare word. Every rung a user could read (text,
// description, hint, descendant text) is deliberately absent: the point of this
// channel is that the model sees no visible text at all, so a fallback that
// reached for text would silently turn the arm back into the default one.
func resourceIdentifierLabel(element *hierarchy.Element) string {
if element.ResourceID != "" {
return truncateLabel(element.ResourceID)
}
if element.Class != "" {
return element.Class
}
return "control"
}
// visibleLabel names a control by what a user would read: its own text, then
// description, then a field hint, then text borrowed from its descendants (the
// case that fixes empty-text Compose buttons whose word lives on a child), then
@@ -635,13 +869,6 @@ func inputTypeHint(element *hierarchy.Element) string {
}
}
// hasEnabled reports whether the source tree carried an explicit enabled flag
// for the element, so a missing flag is not mistaken for "disabled".
func hasEnabled(element *hierarchy.Element) bool {
_, ok := element.Attributes["enabled"]
return ok
}
// stringField reads a string property off a goja object, returning "" when
// absent, null, or undefined.
func stringField(object *goja.Object, key string) string {
@@ -652,6 +879,16 @@ func stringField(object *goja.Object, key string) string {
return value.String()
}
// boolField reads a boolean property off a goja object, returning false when
// absent, null, or undefined.
func boolField(object *goja.Object, key string) bool {
value := object.Get(key)
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return false
}
return value.ToBoolean()
}
// intField reads a numeric property off a goja object, returning 0 when absent,
// null, or undefined.
func intField(object *goja.Object, key string) int {
@@ -661,3 +898,33 @@ func intField(object *goja.Object, key string) int {
}
return int(value.ToInteger())
}
// intFieldOr reads a numeric property, falling back when the descriptor left it
// out. It mirrors the serializer's `??`, so an explicit zero is kept.
func intFieldOr(object *goja.Object, key string, fallback int) int {
value := object.Get(key)
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return fallback
}
return int(value.ToInteger())
}
// numberField reads a property that must actually BE a number, which is what
// tells a target carrying no coordinates apart from one anchored at (0, 0).
func numberField(object *goja.Object, key string) (int, bool) {
value := object.Get(key)
if value == nil {
return 0, false
}
switch number := value.Export().(type) {
case int64:
return int(number), true
case float64:
if math.IsNaN(number) {
return 0, false
}
return int(number), true
default:
return 0, false
}
}
+474 -21
View File
@@ -1,6 +1,7 @@
package verifier
import (
"encoding/json"
"strings"
"testing"
@@ -24,6 +25,35 @@ const enumTreeJSON = `{
]
}`
// labelChannelTreeJSON pulls the two label channels apart: every identifier
// differs from the text a user reads, one control has a class but no identifier,
// one has neither, and one has an identifier but nothing readable at all.
const labelChannelTreeJSON = `{
"attributes": {"bounds": "[0,0,1080,2400]"},
"children": [
{"attributes": {"resource-id": "add_credit_button", "class": "android.widget.Button", "bounds": "[0,100,1080,200]"}, "clickable": true, "enabled": true, "children": [
{"attributes": {"text": "Add credit", "bounds": "[0,100,540,200]"}, "children": []}
]},
{"attributes": {"class": "android.widget.CheckBox", "text": "Remember me", "bounds": "[0,250,1080,300]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"class": "android.widget.CheckBox", "text": "Stay signed in", "bounds": "[0,300,1080,350]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"text": "Sign in", "bounds": "[0,400,1080,450]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"resource-id": "amount_field", "class": "EditText", "hintText": "Amount", "bounds": "[0,500,1080,600]"}, "enabled": true, "children": []},
{"attributes": {"resource-id": "silent_row", "bounds": "[0,650,1080,700]"}, "clickable": true, "enabled": true, "children": []}
]
}`
// sharedLabelTreeJSON is a list where two rows read exactly the same to a user
// ("Delete") while the app tells them apart by identifier. It is the shape a
// list of removable items has in any real app.
const sharedLabelTreeJSON = `{
"attributes": {"bounds": "[0,0,1080,2400]"},
"children": [
{"attributes": {"resource-id": "delete_alpha", "text": "Delete", "bounds": "[0,100,1080,200]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"resource-id": "delete_beta", "text": "Delete", "bounds": "[0,200,1080,300]"}, "clickable": true, "enabled": true, "children": []},
{"attributes": {"resource-id": "checkout", "text": "Checkout", "bounds": "[0,400,1080,500]"}, "clickable": true, "enabled": true, "children": []}
]
}`
// enumVerifier loads a spec whose actions root is the given plain-object graph
// and stages the given tree, so Candidates walks a controlled action tree. The
// spec is bundled with the goja runtime entry because the model arm reads the
@@ -40,6 +70,17 @@ func enumVerifier(t *testing.T, actionsJS, treeJSON string) *Verifier {
return v
}
// mustCandidates enumerates the model policy's list, failing the test on the
// refusal an authored multi-item sampler raises.
func mustCandidates(t *testing.T, v *Verifier, labelSource string) []ActionCandidate {
t.Helper()
candidates, err := v.Candidates(labelSource)
if err != nil {
t.Fatalf("Candidates: %v", err)
}
return candidates
}
func findCandidate(candidates []ActionCandidate, description string) (ActionCandidate, bool) {
for _, candidate := range candidates {
if candidate.Description == description {
@@ -56,7 +97,7 @@ func hasCandidate(candidates []ActionCandidate, description string) bool {
func TestCandidatesLabelsControlsByVisibleText(t *testing.T) {
v := enumVerifier(t, "{kind:'builtin', verb:'taps'}", enumTreeJSON)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
// The empty-text clickable wrapper is labeled by its child Text, NOT its
// resource-id.
@@ -79,9 +120,162 @@ func TestCandidatesLabelsControlsByVisibleText(t *testing.T) {
}
}
func TestCandidatesDropsDisabledControls(t *testing.T) {
func TestCandidatesLabelsControlsByResourceIdentifier(t *testing.T) {
candidates := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", labelChannelTreeJSON), LabelSourceResourceID)
if !hasCandidate(candidates, `Tap "add_credit_button"`) {
t.Errorf("want the control named by its identifier, got %v", descriptions(candidates))
}
// Nothing a user could read may reach this channel, including through a
// fallback rung: an arm that sees the text is the other arm.
for _, readable := range []string{`Tap "Add credit"`, `Tap "Remember me"`, `Tap "Sign in"`} {
if hasCandidate(candidates, readable) {
t.Errorf("visible text leaked in as %s: %v", readable, descriptions(candidates))
}
}
}
func TestCandidatesIdentifierChannelFallsBackToClassThenBareControl(t *testing.T) {
candidates := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", labelChannelTreeJSON), LabelSourceResourceID)
if !hasCandidate(candidates, `Tap "android.widget.CheckBox"`) {
t.Errorf("a control with no identifier falls back to its class, got %v", descriptions(candidates))
}
if !hasCandidate(candidates, `Tap "control"`) {
t.Errorf("a control with neither identifier nor class falls back to a bare word, got %v",
descriptions(candidates))
}
}
// TestCandidatesIdentifierChannelKeepsControlsItCannotNameApartReachable is the
// cost of the channel, bounded: two identifier-less controls of one class read
// the same in the numbered list, but they stay TWO entries, each carrying its
// own action, so the model can act on either by number. A channel that renames
// controls must never shrink the action space.
func TestCandidatesIdentifierChannelKeepsControlsItCannotNameApartReachable(t *testing.T) {
text := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", labelChannelTreeJSON), LabelSourceVisibleText)
identifier := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", labelChannelTreeJSON), LabelSourceResourceID)
if count(text, `Tap "Remember me"`) != 1 || count(text, `Tap "Stay signed in"`) != 1 {
t.Fatalf("the text channel should address both checkboxes, got %v", descriptions(text))
}
checkboxes := candidatesMatching(identifier, `Tap "android.widget.CheckBox"`)
if len(checkboxes) != 2 {
t.Fatalf("the two checkboxes should be two entries, got %d: %v",
len(checkboxes), descriptions(identifier))
}
if checkboxes[0].Action == checkboxes[1].Action {
t.Errorf("both entries execute the same action: %+v", checkboxes[0].Action)
}
if len(identifier) != len(text) {
t.Errorf("identifier list (%d) and text list (%d) must offer the same actions: %v vs %v",
len(identifier), len(text), descriptions(identifier), descriptions(text))
}
}
// TestCandidatesReachBothControlsSharingOneVisibleLabel is the reachability
// floor: two rows a user reads as the same word are two different controls, so
// both get a number and the second number taps the second row. Dedup that keyed
// on the rendered line dropped the second one, putting it out of reach of any
// prompt or policy.
func TestCandidatesReachBothControlsSharingOneVisibleLabel(t *testing.T) {
candidates := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", sharedLabelTreeJSON), LabelSourceVisibleText)
deletes := candidatesMatching(candidates, `Tap "Delete"`)
if len(deletes) != 2 {
t.Fatalf("want both Delete rows reachable, got %d: %v", len(deletes), descriptions(candidates))
}
if got := deletes[0].Action.On; got != "id:delete_alpha" {
t.Errorf("first entry targets %q, want id:delete_alpha", got)
}
if got := deletes[1].Action.On; got != "id:delete_beta" {
t.Errorf("second entry targets %q, want id:delete_beta", got)
}
if deletes[0].Index == deletes[1].Index {
t.Errorf("both entries share number %d, so the model cannot address them apart", deletes[0].Index)
}
}
// TestCandidatesVisibleTextFallsBackToTheIdentifier is where the two channels
// agree: a control carrying nothing readable is named by its identifier in both,
// so a screen built entirely from such controls is one cell, not two.
func TestCandidatesVisibleTextFallsBackToTheIdentifier(t *testing.T) {
candidates := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'taps'}", labelChannelTreeJSON), LabelSourceVisibleText)
if !hasCandidate(candidates, `Tap "silent_row"`) {
t.Errorf("a control with no readable text falls back to its identifier, got %v",
descriptions(candidates))
}
}
func TestCandidatesTypingLabelFollowsTheLabelSource(t *testing.T) {
text := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'typing'}", labelChannelTreeJSON), LabelSourceVisibleText)
if !hasCandidate(text, `Type into "Amount" (number)`) {
t.Errorf("want the field named by its hint, got %v", descriptions(text))
}
identifier := mustCandidates(t, enumVerifier(t, "{kind:'builtin', verb:'typing'}", labelChannelTreeJSON), LabelSourceResourceID)
if !hasCandidate(identifier, `Type into "amount_field" (number)`) {
t.Errorf("want the field named by its identifier, got %v", descriptions(identifier))
}
}
// TestLabelSourceChangesOnlyTheDescription is the claim the labelling factorial
// rests on: the channel renames every target and does nothing else. Both arms
// enumerate the same candidates, in the same order, at the same weights,
// carrying the same executable actions; the description and the label are the
// only things that move. Anything else and the two cells would be picking from
// different action spaces, so a difference in defect yield could not be
// attributed to how the controls were named.
func TestLabelSourceChangesOnlyTheDescription(t *testing.T) {
const everyLabelledVerb = `{kind:'weighted', branches:[
[1,{kind:'builtin',verb:'taps'}],
[1,{kind:'builtin',verb:'typing'}],
[1,{kind:'builtin',verb:'swipes'}]
]}`
fixtures := []struct {
name string
tree string
}{
{"identifiers collide", labelChannelTreeJSON},
{"visible text collides", sharedLabelTreeJSON},
}
withoutNames := func(candidate ActionCandidate) ActionCandidate {
candidate.Description = ""
candidate.Label = ""
return candidate
}
for _, fixture := range fixtures {
t.Run(fixture.name, func(t *testing.T) {
text := mustCandidates(t, enumVerifier(t, everyLabelledVerb, fixture.tree), LabelSourceVisibleText)
identifier := mustCandidates(t, enumVerifier(t, everyLabelledVerb, fixture.tree), LabelSourceResourceID)
if len(text) == 0 {
t.Fatal("fixture yielded no candidates")
}
if len(text) != len(identifier) {
t.Fatalf("different action spaces: %d text candidates vs %d identifier ones:\n%v\n%v",
len(text), len(identifier), descriptions(text), descriptions(identifier))
}
renamed := false
for i := range text {
if withoutNames(text[i]) != withoutNames(identifier[i]) {
t.Errorf("candidate %d differs beyond its name:\n text=%+v\n id=%+v",
i+1, text[i], identifier[i])
}
if text[i].Description != identifier[i].Description {
renamed = true
}
}
if !renamed {
t.Error("no candidate was renamed, so this fixture does not exercise the channel")
}
})
}
}
func TestCandidatesDropsDisabledControlsFromBuiltinVerbs(t *testing.T) {
v := enumVerifier(t, "{kind:'builtin', verb:'taps'}", enumTreeJSON)
for _, candidate := range v.Candidates() {
for _, candidate := range mustCandidates(t, v, LabelSourceVisibleText) {
if strings.Contains(candidate.Description, "Off") {
t.Errorf("disabled control surfaced as %q", candidate.Description)
}
@@ -90,7 +284,7 @@ func TestCandidatesDropsDisabledControls(t *testing.T) {
func TestCandidatesTypingExposesInputType(t *testing.T) {
v := enumVerifier(t, "{kind:'builtin', verb:'typing'}", enumTreeJSON)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
candidate, ok := findCandidate(candidates, `Type into "Amount" (number)`)
if !ok {
t.Fatalf("want typing candidate with input type, got %v", descriptions(candidates))
@@ -113,7 +307,7 @@ func TestCandidatesLabelsEditableFieldByHintNotTypedValue(t *testing.T) {
]
}`
v := enumVerifier(t, "{kind:'builtin', verb:'typing'}", tree)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
if hasCandidate(candidates, `Type into "99" (number)`) || hasCandidate(candidates, `Type into "99"`) {
t.Errorf("editable field labeled by its typed value: %v", descriptions(candidates))
}
@@ -126,7 +320,7 @@ func TestCandidatesKeepsGestureVerbsDistinct(t *testing.T) {
v := enumVerifier(t,
"{kind:'weighted', branches:[[1,{kind:'builtin',verb:'scrolls'}],[1,{kind:'builtin',verb:'swipes'}]]}",
enumTreeJSON)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
// `scrolls` folds to one directional pair over the single scrollable
// container, which is what keeps the list short.
@@ -168,7 +362,7 @@ func TestCandidatesWeightsCombineAcrossPaths(t *testing.T) {
v := enumVerifier(t,
"{kind:'weighted', branches:[[1,{kind:'builtin',verb:'taps'}],[1,{kind:'builtin',verb:'taps'}]]}",
oneClickable)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
if len(candidates) != 1 {
t.Fatalf("want one deduped candidate, got %v", descriptions(candidates))
}
@@ -185,7 +379,7 @@ func TestCandidatesWeightReflectsBranchShare(t *testing.T) {
v := enumVerifier(t,
"{kind:'weighted', branches:[[1,{kind:'builtin',verb:'taps'}],[3,{kind:'builtin',verb:'typing'}]]}",
enumTreeJSON)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
tap, ok := findCandidate(candidates, `Tap "Sign in"`)
if !ok {
t.Fatalf("missing tap candidate: %v", descriptions(candidates))
@@ -204,7 +398,7 @@ func TestCandidatesWeightReflectsBranchShare(t *testing.T) {
func TestCandidatesUnweightedTreeShowsNoWeight(t *testing.T) {
v := enumVerifier(t, "{kind:'builtin', verb:'taps'}", enumTreeJSON)
for _, candidate := range v.Candidates() {
for _, candidate := range mustCandidates(t, v, LabelSourceVisibleText) {
if candidate.Weighted || candidate.Weight != 0 {
t.Errorf("%q carries a weight despite no weighted node", candidate.Description)
}
@@ -218,20 +412,22 @@ func TestCandidatesCallsAuthoredLeafOnce(t *testing.T) {
{kind:'InputText', into:'id:Amount', text:'42'}
]}`
v := enumVerifier(t, actions, enumTreeJSON)
candidates := v.Candidates()
candidates := mustCandidates(t, v, LabelSourceVisibleText)
// Authored Tap resolves its selector to the visible-text label.
if !hasCandidate(candidates, `Tap "Sign in"`) {
t.Errorf("authored tap missing: %v", descriptions(candidates))
}
// A disabled authored target is dropped.
for _, candidate := range candidates {
if strings.Contains(candidate.Description, "Off") {
t.Errorf("authored action on disabled control surfaced: %q", candidate.Description)
}
// A disabled authored target is offered, not dropped: the seeded picker
// executes it, and attempting a disabled control is where boundary defects
// live, so a policy that cannot attempt it cannot find them.
if !hasCandidate(candidates, `Tap "Off"`) {
t.Errorf("authored action on a disabled control was dropped: %v", descriptions(candidates))
}
// Authored InputText replays its own sampled value (LLM does not supply it).
authored, ok := findCandidate(candidates, `Type "42" into "Amount"`)
// The fixture is an android tree, which reports no secure fact, so the
// rendered value is redacted while the action still carries it.
authored, ok := findCandidate(candidates, `Type "[redacted]" into "Amount"`)
if !ok {
t.Fatalf("authored typing missing: %v", descriptions(candidates))
}
@@ -252,7 +448,7 @@ func TestCandidatesSurfaceAuthoredUntargetedActions(t *testing.T) {
{kind:'PressKey', key:'back'},
{kind:'Wait'}
]}`
candidates := enumVerifier(t, actions, enumTreeJSON).Candidates()
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
for _, want := range []string{"Swipe from (10,600) to (10,100)", "Press back", "Wait"} {
if !hasCandidate(candidates, want) {
t.Errorf("authored %q missing: %v", want, descriptions(candidates))
@@ -260,9 +456,114 @@ func TestCandidatesSurfaceAuthoredUntargetedActions(t *testing.T) {
}
}
// TestCandidatesAuthoredWaitKeepsItsDuration: a Wait that loses its duration is
// a wait of zero, which the runner cannot dispatch at all, so the model would be
// idling on paper while the seeded arm really waits.
func TestCandidatesAuthoredWaitKeepsItsDuration(t *testing.T) {
actions := `{kind:'actions', generate: () => [{kind:'Wait', durationMillis: 500}]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
candidate, ok := findCandidate(candidates, "Wait")
if !ok {
t.Fatalf("authored wait missing: %v", descriptions(candidates))
}
if candidate.Action.DurationMillis != 500 {
t.Errorf("wait duration = %d, want the authored 500", candidate.Action.DurationMillis)
}
}
// TestCandidatesAuthoredScrollNamesItsContainer: the container is the whole
// point of an authored scroll. Dropped, the runner re-derives the gesture from
// the screen and the scroll lands on whatever else is scrollable.
func TestCandidatesAuthoredScrollNamesItsContainer(t *testing.T) {
actions := `{kind:'actions', generate: () => [{kind:'Scroll', direction:'down', in:'id:List'}]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
candidate, ok := findCandidate(candidates, "Scroll down")
if !ok {
t.Fatalf("authored scroll missing: %v", descriptions(candidates))
}
if candidate.Action.On != "id:List" {
t.Errorf("scroll container = %q, want id:List", candidate.Action.On)
}
if candidate.Action.DurationMillis != gestureDurationMillis {
t.Errorf("scroll duration = %d, want %d", candidate.Action.DurationMillis, gestureDurationMillis)
}
}
// TestCandidatesAuthoredScrollKeepsPrecomputedEndpoints: a descriptor that
// already carries the gesture is executed as written rather than re-derived.
func TestCandidatesAuthoredScrollKeepsPrecomputedEndpoints(t *testing.T) {
actions := `{kind:'actions', generate: () => [
{kind:'Scroll', direction:'down', in:'id:List', from:{x:540,y:1400}, to:{x:540,y:920}}
]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
candidate, ok := findCandidate(candidates, "Scroll down")
if !ok {
t.Fatalf("authored scroll missing: %v", descriptions(candidates))
}
action := candidate.Action
got := [4]int{action.FromX, action.FromY, action.ToX, action.ToY}
if got != [4]int{540, 1400, 540, 920} {
t.Errorf("scroll endpoints = %v, want the descriptor's (540,1400)->(540,920)", got)
}
}
// TestCandidatesAuthoredSwipeDefaultsItsDuration keeps the gesture default in
// one place: the serializer the seeded arm goes through fills an omitted
// duration, and a candidate that left it at zero would depend on the runner
// happening to pick the same fallback.
func TestCandidatesAuthoredSwipeDefaultsItsDuration(t *testing.T) {
actions := `{kind:'actions', generate: () => [
{kind:'Swipe', from:{x:10,y:600}, to:{x:10,y:100}},
{kind:'Swipe', from:{x:20,y:600}, to:{x:20,y:100}, durationMillis: 400}
]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
omitted, ok := findCandidate(candidates, "Swipe from (10,600) to (10,100)")
if !ok {
t.Fatalf("authored swipe missing: %v", descriptions(candidates))
}
if omitted.Action.DurationMillis != gestureDurationMillis {
t.Errorf("omitted duration = %d, want %d", omitted.Action.DurationMillis, gestureDurationMillis)
}
authored, ok := findCandidate(candidates, "Swipe from (20,600) to (20,100)")
if !ok {
t.Fatalf("authored swipe missing: %v", descriptions(candidates))
}
if authored.Action.DurationMillis != 400 {
t.Errorf("authored duration = %d, want 400", authored.Action.DurationMillis)
}
}
// TestCandidatesDropTargetsThatResolveToNothing: the seeded picker drops an
// action whose target resolves to neither coordinates nor a selector. Offering
// it to the model instead would put a tap on the screen origin within reach,
// which on Android is the corner that pulls the notification shade down.
func TestCandidatesDropTargetsThatResolveToNothing(t *testing.T) {
actions := `{kind:'actions', generate: () => [
{kind:'Tap', on: null},
{kind:'Tap', on: {}},
{kind:'InputText', into: {}, text:'x'},
{kind:'Swipe', from:{x:1,y:2}, to:{}}
]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
if len(candidates) != 0 {
t.Errorf("targetless actions reached the model: %v", descriptions(candidates))
}
}
// TestCandidatesKeepATargetOnTheScreenOrigin is the other side of that rule: a
// point at (0,0) IS a target the seeded picker executes, so the drop must key on
// a target with no coordinates rather than on coordinates that are zero.
func TestCandidatesKeepATargetOnTheScreenOrigin(t *testing.T) {
actions := `{kind:'actions', generate: () => [{kind:'Tap', on: {x: 0, y: 0}}]}`
candidates := mustCandidates(t, enumVerifier(t, actions, enumTreeJSON), LabelSourceVisibleText)
if len(candidates) != 1 {
t.Fatalf("want the origin tap kept, got %v", descriptions(candidates))
}
}
func TestCandidatesOffRouteLeafYieldsNothing(t *testing.T) {
v := enumVerifier(t, "{kind:'actions', generate: () => []}", enumTreeJSON)
if got := v.Candidates(); len(got) != 0 {
if got := mustCandidates(t, v, LabelSourceVisibleText); len(got) != 0 {
t.Errorf("off-route leaf should yield no candidates, got %v", descriptions(got))
}
}
@@ -280,7 +581,7 @@ func TestCandidatesSkipsCrossFadeFrames(t *testing.T) {
]
}`
v := enumVerifier(t, "{kind:'builtin', verb:'taps'}", crossFade)
if got := v.Candidates(); len(got) != 0 {
if got := mustCandidates(t, v, LabelSourceVisibleText); len(got) != 0 {
t.Errorf("cross-fade frame should yield no candidates, got %v", descriptions(got))
}
// The seeded policy is skipped by the SAME guard, in the shared producer,
@@ -292,13 +593,13 @@ func TestCandidatesSkipsCrossFadeFrames(t *testing.T) {
func TestCandidatesNilWithoutTreeOrActions(t *testing.T) {
withActions := newLoadedVerifier(t, "globalThis.actions = {kind:'builtin', verb:'taps'};")
if got := withActions.Candidates(); got != nil {
if got := mustCandidates(t, withActions, LabelSourceVisibleText); got != nil {
t.Errorf("Candidates with no tree = %v, want nil", got)
}
noActions := newLoadedVerifier(t, "globalThis.properties = {};")
tree, _ := hierarchy.Parse(enumTreeJSON)
noActions.lastTree = tree
if got := noActions.Candidates(); got != nil {
if got := mustCandidates(t, noActions, LabelSourceVisibleText); got != nil {
t.Errorf("Candidates with no actions root = %v, want nil", got)
}
}
@@ -311,6 +612,16 @@ func descriptions(candidates []ActionCandidate) []string {
return out
}
func candidatesMatching(candidates []ActionCandidate, description string) []ActionCandidate {
var matched []ActionCandidate
for _, candidate := range candidates {
if candidate.Description == description {
matched = append(matched, candidate)
}
}
return matched
}
func count(candidates []ActionCandidate, description string) int {
n := 0
for _, candidate := range candidates {
@@ -382,3 +693,145 @@ func newLoadedVerifier(t *testing.T, source string) *Verifier {
}
return v
}
// samplerSpec authors one leaf that taps a target drawn from the given list.
func samplerSpec(items string) string {
return `
import { actions, from, Tap } from "@sanderling/spec";
const targets = from(` + items + `);
globalThis.actions = actions(() => [Tap({ on: targets.generate() })]);
`
}
// TestCandidatesRefuseAMultiItemAuthoredSampler pins the refusal: a sampler
// reads the picker's rng, which this policy has no way to enter, so the draw
// would collapse to the first item on every step while the seeded picker keeps
// reaching all three. Offering that silently is what would make a comparison of
// the two policies meaningless, so the spec is refused instead.
func TestCandidatesRefuseAMultiItemAuthoredSampler(t *testing.T) {
v := newVerifier(t)
loadActionSpec(t, v, samplerSpec(`["id:SignIn", "id:Amount", "id:List"]`))
pushTree(t, v, enumTreeJSON)
_, err := v.Candidates(LabelSourceVisibleText)
if err == nil {
t.Fatal("a multi-item authored sampler must refuse to run under the model policy")
}
message := err.Error()
if !strings.Contains(message, "targets.generate()") {
t.Errorf("error does not name the offending leaf, so the author cannot find it: %s", message)
}
if !strings.Contains(message, "draws 1 of 3 sampled items") {
t.Errorf("error does not say what the leaf did: %s", message)
}
}
// TestCandidatesAcceptASingleItemAuthoredSampler: a one-item sampler short
// circuits before the rng, so both policies get that one value and there is no
// divergence to refuse.
func TestCandidatesAcceptASingleItemAuthoredSampler(t *testing.T) {
v := newVerifier(t)
loadActionSpec(t, v, samplerSpec(`["id:SignIn"]`))
pushTree(t, v, enumTreeJSON)
candidates, err := v.Candidates(LabelSourceVisibleText)
if err != nil {
t.Fatalf("a single-item sampler is not a divergence: %v", err)
}
if !hasCandidate(candidates, `Tap "Sign in"`) {
t.Errorf("sampled tap missing: %v", descriptions(candidates))
}
}
// webFieldTreeJSON is the add-transaction screen as the chrome driver dumps it:
// two inputs a user tells apart by the <label> bound to each, which the dump
// does not carry. Named from the tree alone they collapse onto their CSS class.
const webFieldTreeJSON = `{
"attributes": {"tag": "html", "bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "txn-amount", "tag": "input", "class": "input amount-input", "bounds": "[0,100,400,160]"}, "editable": true, "enabled": true, "secure": false, "children": []},
{"attributes": {"resource-id": "txn-note", "tag": "input", "class": "input", "bounds": "[0,200,400,260]"}, "editable": true, "enabled": true, "secure": false, "children": []}
]
}`
// The web tick evaluates extractors in V8 and injects the handles here, so an
// authored action's target is a plain object with no selector to resolve
// against the tree. Naming it by the handle's own text names every input "",
// and the model then cannot tell the amount field from the note field.
// pkg/spec/test/web-runtime.test.ts asserts the producing side builds these
// handles with the hint each assertion here reads.
func TestCandidatesNameWebAuthoredFieldsByTheirHint(t *testing.T) {
const amountHandle = `{
"id": "txn-amount", "text": "", "desc": "", "class": "input amount-input",
"clickable": true, "enabled": true, "editable": true, "focused": false, "secure": false,
"x": 200, "y": 130, "bounds": {"left": 0, "top": 100, "right": 400, "bottom": 160},
"attrs": {"tag": "input", "aria-label": "", "id": "txn-amount",
"class": "input amount-input", "placeholder": "0.00", "hintText": "Amount"}
}`
const noteHandle = `{
"id": "txn-note", "text": "", "desc": "", "class": "input",
"clickable": true, "enabled": true, "editable": true, "focused": false, "secure": false,
"x": 200, "y": 230, "bounds": {"left": 0, "top": 200, "right": 400, "bottom": 260},
"attrs": {"tag": "input", "aria-label": "", "id": "txn-note", "class": "input",
"placeholder": "What's this for?", "hintText": "Note (optional)"}
}`
v := newVerifier(t)
loadActionSpec(t, v, `
import { InputText, actions, extract } from "@sanderling/spec";
const amount = extract((s) => s.ax.find({ id: "txn-amount" })).named("amount");
const note = extract((s) => s.ax.find({ id: "txn-note" })).named("note");
globalThis.actions = actions(() => {
if (!amount.current || !note.current) return [];
return [
InputText({ into: amount.current, text: "12.34" }),
InputText({ into: note.current, text: "Coffee" }),
];
});
`)
pushTree(t, v, webFieldTreeJSON)
if _, err := v.OverrideExtractorValues(map[int]json.RawMessage{
0: json.RawMessage(amountHandle),
1: json.RawMessage(noteHandle),
}); err != nil {
t.Fatal(err)
}
candidates := mustCandidates(t, v, LabelSourceVisibleText)
if !hasCandidate(candidates, `Type "12.34" into "Amount"`) {
t.Errorf("amount field unnamed: %v", descriptions(candidates))
}
if !hasCandidate(candidates, `Type "Coffee" into "Note (optional)"`) {
t.Errorf("note field unnamed: %v", descriptions(candidates))
}
}
// The identifier arm must stay blind to anything a user reads, hint included.
func TestCandidatesNameWebAuthoredFieldsByIdentifierOnThatArm(t *testing.T) {
const amountHandle = `{
"id": "txn-amount", "text": "", "editable": true, "enabled": true, "secure": false,
"x": 200, "y": 130,
"attrs": {"tag": "input", "hintText": "Amount"}
}`
v := newVerifier(t)
loadActionSpec(t, v, `
import { InputText, actions, extract } from "@sanderling/spec";
const amount = extract((s) => s.ax.find({ id: "txn-amount" })).named("amount");
globalThis.actions = actions(() =>
amount.current ? [InputText({ into: amount.current, text: "12.34" })] : []);
`)
pushTree(t, v, webFieldTreeJSON)
if _, err := v.OverrideExtractorValues(map[int]json.RawMessage{
0: json.RawMessage(amountHandle),
}); err != nil {
t.Fatal(err)
}
candidates := mustCandidates(t, v, LabelSourceResourceID)
if !hasCandidate(candidates, `Type "12.34" into "txn-amount"`) {
t.Errorf("identifier arm lost the field name: %v", descriptions(candidates))
}
if hasCandidate(candidates, `Type "12.34" into "Amount"`) {
t.Error("identifier arm leaked the hint a user reads")
}
}
+84 -8
View File
@@ -4,6 +4,7 @@ import (
"bytes"
"encoding/json"
"fmt"
"slices"
"strings"
"time"
@@ -71,16 +72,17 @@ func accessibilityObject(runtime *goja.Runtime, tree *hierarchy.Tree) *goja.Obje
if node == nil {
return goja.Undefined()
}
return nodeObject(runtime, node, selectorStringFromJS(runtime, call.Argument(0)))
return nodeObject(runtime, tree, node, selectorStringFromJS(runtime, call.Argument(0)))
}
findAll := func(call goja.FunctionCall) goja.Value {
if tree == nil {
return goja.Undefined()
}
nodes := findAllNodesFromJS(runtime, tree, call.Argument(0))
selector := selectorStringFromJS(runtime, call.Argument(0))
array := runtime.NewArray()
for i, n := range nodes {
_ = array.Set(fmt.Sprintf("%d", i), nodeObject(runtime, n, selectorStringFromJS(runtime, call.Argument(0))))
_ = array.Set(fmt.Sprintf("%d", i), nodeObject(runtime, tree, n, selector))
}
return array
}
@@ -89,7 +91,26 @@ func accessibilityObject(runtime *goja.Runtime, tree *hierarchy.Tree) *goja.Obje
return accessibility
}
func nodeObject(runtime *goja.Runtime, node *hierarchy.Node, selector string) goja.Value {
// unambiguousSelector returns selector only when no node other than this one
// answers to it. The runner prefers tree.Find(action.On) over the coordinates
// the element reported (resolveCoordinates) and Find takes the first match, so
// naming an element by a selector its siblings share sends every one of their
// actions to the first sibling. An unnamed element keeps its own coordinates,
// which are already right, matching what selectorsFor does for the builtin
// target enumeration in pkg/spec/src/web-runtime.ts.
func unambiguousSelector(tree *hierarchy.Tree, node *hierarchy.Node, selector string) string {
if tree == nil || selector == "" {
return ""
}
for _, match := range tree.FindAllNodes(selector) {
if match != node {
return ""
}
}
return selector
}
func nodeObject(runtime *goja.Runtime, tree *hierarchy.Tree, node *hierarchy.Node, selector string) goja.Value {
element := &node.Element
object := runtime.NewObject()
centerX, centerY := element.Bounds.Center()
@@ -103,9 +124,17 @@ func nodeObject(runtime *goja.Runtime, node *hierarchy.Node, selector string) go
_ = object.Set("focused", element.Focused)
_ = object.Set("selected", element.Selected)
_ = object.Set("editable", element.Editable)
// Three-valued, unlike the other state flags: null where the platform
// reported nothing at all, which is what separates an ordinary field from a
// password field on a platform that cannot tell them apart.
var secure any
if element.SecureReported() {
secure = element.Secure
}
_ = object.Set("secure", secure)
_ = object.Set("x", centerX)
_ = object.Set("y", centerY)
_ = object.Set(tagSelector, selector)
_ = object.Set(tagSelector, unambiguousSelector(tree, node, selector))
bounds := runtime.NewObject()
_ = bounds.Set("left", element.Bounds.Left)
_ = bounds.Set("top", element.Bounds.Top)
@@ -123,14 +152,15 @@ func nodeObject(runtime *goja.Runtime, node *hierarchy.Node, selector string) go
if childNode == nil {
return goja.Undefined()
}
return nodeObject(runtime, childNode, selectorStringFromJS(runtime, arg))
return nodeObject(runtime, tree, childNode, selectorStringFromJS(runtime, arg))
}
childFindAll := func(call goja.FunctionCall) goja.Value {
arg := call.Argument(0)
childNodes := findAllNodesInSubtreeFromJS(runtime, node, arg)
childSelector := selectorStringFromJS(runtime, arg)
array := runtime.NewArray()
for i, n := range childNodes {
_ = array.Set(fmt.Sprintf("%d", i), nodeObject(runtime, n, selectorStringFromJS(runtime, arg)))
_ = array.Set(fmt.Sprintf("%d", i), nodeObject(runtime, tree, n, childSelector))
}
return array
}
@@ -152,13 +182,15 @@ func findNodeFromJS(runtime *goja.Runtime, tree *hierarchy.Tree, arg goja.Value)
return tree.FindNode(s)
}
if path, ok := selectorPathFromJS(runtime, arg); ok {
requireKnownSelectorKeys(runtime, tree, path...)
return tree.FindBySelectorPath(path)
}
sel := selectorFromJSObject(runtime, arg)
if len(sel.Filters) == 0 {
return nil
}
return tree.Root.FindBySelector(sel)
requireKnownSelectorKeys(runtime, tree, sel)
return tree.FindBySelector(sel)
}
// findAllNodesFromJS dispatches a JS value to Tree-level multi-node lookup.
@@ -170,13 +202,15 @@ func findAllNodesFromJS(runtime *goja.Runtime, tree *hierarchy.Tree, arg goja.Va
return tree.FindAllNodes(s)
}
if path, ok := selectorPathFromJS(runtime, arg); ok {
requireKnownSelectorKeys(runtime, tree, path...)
return tree.FindAllBySelectorPath(path)
}
sel := selectorFromJSObject(runtime, arg)
if len(sel.Filters) == 0 {
return nil
}
return tree.Root.FindAllBySelector(sel)
requireKnownSelectorKeys(runtime, tree, sel)
return tree.FindAllBySelector(sel)
}
// findNodeInSubtreeFromJS dispatches a JS value to Node-level scoped lookup.
@@ -188,12 +222,14 @@ func findNodeInSubtreeFromJS(runtime *goja.Runtime, node *hierarchy.Node, arg go
return node.Find(s)
}
if path, ok := selectorPathFromJS(runtime, arg); ok {
requireKnownSelectorKeys(runtime, node.Tree(), path...)
return node.FindBySelectorPath(path)
}
sel := selectorFromJSObject(runtime, arg)
if len(sel.Filters) == 0 {
return nil
}
requireKnownSelectorKeys(runtime, node.Tree(), sel)
return node.FindBySelector(sel)
}
@@ -206,15 +242,41 @@ func findAllNodesInSubtreeFromJS(runtime *goja.Runtime, node *hierarchy.Node, ar
return node.FindAll(s)
}
if path, ok := selectorPathFromJS(runtime, arg); ok {
requireKnownSelectorKeys(runtime, node.Tree(), path...)
return node.FindAllBySelectorPath(path)
}
sel := selectorFromJSObject(runtime, arg)
if len(sel.Filters) == 0 {
return nil
}
requireKnownSelectorKeys(runtime, node.Tree(), sel)
return node.FindAllBySelector(sel)
}
// requireKnownSelectorKeys throws a JS error when a selector names a key that
// can never match. Returning an empty result instead is indistinguishable from
// a screen that simply has no such element, so a spec built on a mistyped key
// generates no action, the runner waits out every step, and the campaign
// finishes clean having explored nothing.
func requireKnownSelectorKeys(
runtime *goja.Runtime,
tree *hierarchy.Tree,
selectors ...hierarchy.Selector,
) {
var unknown []string
for _, sel := range selectors {
for _, key := range tree.UnknownSelectorKeys(sel) {
if !slices.Contains(unknown, key) {
unknown = append(unknown, key)
}
}
}
if len(unknown) == 0 {
return
}
panic(runtime.NewTypeError(hierarchy.UnknownSelectorKeyMessage(unknown)))
}
// selectorFromJSObject converts a JS object {attr: value, ...} into a Selector.
func selectorFromJSObject(runtime *goja.Runtime, arg goja.Value) hierarchy.Selector {
obj := arg.ToObject(runtime)
@@ -549,6 +611,10 @@ type wireAction struct {
ToX int `json:"toX"`
ToY int `json:"toY"`
DurationMillis int `json:"durationMillis"`
// Source names the generator that produced the action: the spec's setup or
// the action root. Empty from the candidate enumeration, which serializes
// actions nothing has chosen yet.
Source string `json:"source"`
}
// DecodeAction turns one serialized action (the flat camelCase wire contract)
@@ -561,6 +627,15 @@ func DecodeAction(raw json.RawMessage) (Action, error) {
if err := json.Unmarshal(raw, &wire); err != nil {
return Action{}, fmt.Errorf("decode action: %w", err)
}
action, err := actionFromWire(wire)
if err != nil {
return Action{}, err
}
action.Source = wire.Source
return action, nil
}
func actionFromWire(wire wireAction) (Action, error) {
switch wire.Kind {
case "Tap":
return Action{Kind: ActionKindTap, On: wire.Selector, X: wire.X, Y: wire.Y}, nil
@@ -582,6 +657,7 @@ func DecodeAction(raw json.RawMessage) (Action, error) {
case "Scroll":
return Action{
Kind: ActionKindScroll,
On: wire.Selector,
Direction: wire.Direction,
FromX: wire.FromX,
FromY: wire.FromY,
+7
View File
@@ -47,6 +47,13 @@ func TestDecodeAction_AllKinds(t *testing.T) {
raw: `{"kind":"Scroll","direction":"down","fromX":1,"fromY":2,"toX":3,"toY":4,"durationMillis":100}`,
want: Action{Kind: ActionKindScroll, Direction: "down", FromX: 1, FromY: 2, ToX: 3, ToY: 4, DurationMillis: 100},
},
{
// An author names the container and leaves the drag to the runner,
// which needs the selector to size it against the right bounds.
name: "Scroll with a container and no endpoints",
raw: `{"kind":"Scroll","direction":"up","selector":"id:list","fromX":0,"fromY":0,"toX":0,"toY":0,"durationMillis":250}`,
want: Action{Kind: ActionKindScroll, On: "id:list", Direction: "up", DurationMillis: 250},
},
{
name: "PressKey",
raw: `{"kind":"PressKey","key":"back"}`,
+181 -8
View File
@@ -5,6 +5,7 @@ import (
"fmt"
"maps"
"slices"
"strings"
"testing"
)
@@ -62,7 +63,7 @@ func TestModelCandidateDescriptionsAreUniqueAndNamed(t *testing.T) {
t.Run(verb, func(t *testing.T) {
verifier := loadVerbSpec(t, verb)
seen := map[string]bool{}
for _, candidate := range verifier.Candidates() {
for _, candidate := range mustCandidates(t, verifier, LabelSourceVisibleText) {
if candidate.Description == "" {
t.Fatalf("%s produced a candidate with no description: %+v", verb, candidate.Action)
}
@@ -81,17 +82,64 @@ func TestModelCandidateDescriptionsAreUniqueAndNamed(t *testing.T) {
// dropped, so the model could never navigate back or let the app settle.
func TestModelIsOfferedTheUntargetedVerbs(t *testing.T) {
verifier := loadVerbSpec(t, "pressKeys")
if !hasCandidate(verifier.Candidates(), "Press back") {
if !hasCandidate(mustCandidates(t, verifier, LabelSourceVisibleText), "Press back") {
t.Errorf("pressKeys missing from the model's candidates: %v",
descriptions(verifier.Candidates()))
descriptions(mustCandidates(t, verifier, LabelSourceVisibleText)))
}
verifier = loadVerbSpec(t, "waitOnce")
if !hasCandidate(verifier.Candidates(), "Wait") {
if !hasCandidate(mustCandidates(t, verifier, LabelSourceVisibleText), "Wait") {
t.Errorf("waitOnce missing from the model's candidates: %v",
descriptions(verifier.Candidates()))
descriptions(mustCandidates(t, verifier, LabelSourceVisibleText)))
}
}
// TestSeededDrawStreamIgnoresLabelSource is the manipulation check the
// labelling factorial needs: the seeded picker selects by index and never asks
// for a label, so its draw stream must be bit-identical whichever channel the
// model policy would have been given, and identical again to a run where the
// candidate list was never enumerated at all. Any difference between two seeded
// cells is then the application and the harness, not the factor.
func TestSeededDrawStreamIgnoresLabelSource(t *testing.T) {
for _, verb := range policyVerbs {
t.Run(verb, func(t *testing.T) {
never := seededDrawStream(t, verb, "")
text := seededDrawStream(t, verb, LabelSourceVisibleText)
identifier := seededDrawStream(t, verb, LabelSourceResourceID)
if !slices.Equal(never, text) {
t.Errorf("enumerating visible-text candidates moved the seeded stream for %s", verb)
}
if !slices.Equal(never, identifier) {
t.Errorf("enumerating identifier candidates moved the seeded stream for %s", verb)
}
})
}
}
// seededDrawStream drives the seeded picker for the draw budget and returns
// every action in order. An empty labelSource enumerates nothing; otherwise the
// model's candidate list is built under that channel before each draw, which is
// the only way the two could ever touch.
func seededDrawStream(t *testing.T, verb, labelSource string) []string {
t.Helper()
verifier := loadVerbSpec(t, verb)
stream := make([]string, 0, seededDrawBudget)
for range seededDrawBudget {
if labelSource != "" {
mustCandidates(t, verifier, labelSource)
}
action, err := verifier.NextAction()
if errors.Is(err, ErrNoAction) {
stream = append(stream, "no action")
continue
}
if err != nil {
t.Fatalf("%s next action: %v", verb, err)
}
stream = append(stream, fmt.Sprintf("%+v", action))
}
return stream
}
// loadVerbSpec builds a verifier whose whole action tree is one builtin verb,
// with policyTreeJSON pushed as the current state.
func loadVerbSpec(t *testing.T, verb string) *Verifier {
@@ -128,7 +176,7 @@ func modelOfferedActions(t *testing.T, verb string) map[string]Action {
t.Helper()
verifier := loadVerbSpec(t, verb)
offered := map[string]Action{}
for _, candidate := range verifier.Candidates() {
for _, candidate := range mustCandidates(t, verifier, LabelSourceVisibleText) {
offered[actionIdentity(candidate.Action)] = candidate.Action
}
return offered
@@ -136,12 +184,15 @@ func modelOfferedActions(t *testing.T, verb string) map[string]Action {
// actionIdentity keys an action by everything except the values the policy owns
// rather than the candidate set: the typed text, which the seeded arm draws from
// the edge-case corpus and the model writes itself, and a swipe's drag distance,
// which the seeded arm draws and the enumeration lists at a nominal length.
// the edge-case corpus and the model writes itself, a swipe's drag distance,
// which the seeded arm draws and the enumeration lists at a nominal length, and
// the source, which names who produced an action rather than what it does (the
// enumeration names nobody, because the model has not chosen yet).
// Comparing those would compare policies instead of action spaces. A swipe's
// direction is NOT policy-owned, so it survives as the sign of the drag.
func actionIdentity(action Action) string {
action.Text = ""
action.Source = ""
if action.Kind == ActionKindSwipe {
action.ToX = sign(action.ToX - action.FromX)
action.ToY = sign(action.ToY - action.FromY)
@@ -159,3 +210,125 @@ func sign(value int) int {
return 0
}
}
// samplerParitySpec authors one leaf that taps a target drawn from three: the
// pattern the model policy refuses and the seeded picker exists to draw.
const samplerParitySpec = `
import { actions, from, Tap } from "@sanderling/spec";
const targets = from(["id:Save", "id:Cancel", "id:Amount"]);
globalThis.actions = actions(() => [Tap({ on: targets.generate() })]);
`
// TestSeededSamplingSurvivesTheModelPolicysRefusal keeps the refusal on the one
// policy it belongs to. Sampling inside the picker's rng scope is correct, so
// the seeded stream must be identical whether or not the model policy tried and
// failed to enumerate the same leaf first, and it must still reach every item.
func TestSeededSamplingSurvivesTheModelPolicysRefusal(t *testing.T) {
alone := seededSamplerStream(t, samplerParitySpec, false)
afterRefusal := seededSamplerStream(t, samplerParitySpec, true)
if !slices.Equal(alone, afterRefusal) {
t.Error("a refused enumeration moved the seeded draw stream")
}
drawn := map[string]bool{}
for _, action := range alone {
drawn[action] = true
}
if len(drawn) != 3 {
t.Errorf("seeded picker reached %d of the 3 sampled targets: %v", len(drawn), slices.Sorted(maps.Keys(drawn)))
}
}
// seededSamplerStream drives the seeded picker over the draw budget, optionally
// letting the model policy refuse the same spec before every draw.
func seededSamplerStream(t *testing.T, specSource string, enumerateFirst bool) []string {
t.Helper()
verifier := newVerifier(t, WithSeed(0x5eed))
loadActionSpec(t, verifier, specSource)
pushTree(t, verifier, policyTreeJSON)
stream := make([]string, 0, seededDrawBudget)
for range seededDrawBudget {
if enumerateFirst {
if _, err := verifier.Candidates(LabelSourceVisibleText); err == nil {
t.Fatal("the model policy must refuse a spec that samples")
}
}
action, err := verifier.NextAction()
if err != nil {
t.Fatalf("seeded picker declined the sampled tap: %v", err)
}
stream = append(stream, fmt.Sprintf("%+v", action))
}
return stream
}
// valueGeneratorSpec authors one leaf that types a drawn value, which is the
// same divergence from() has: the draw reaches the seeded picker's rng and never
// this enumeration, so the model would be handed one fixed value forever.
func valueGeneratorSpec(generator string) string {
return fmt.Sprintf(`
import { actions, InputText, integers, strings, emails, edgeCaseText } from "@sanderling/spec";
const authoredValues = %s;
globalThis.actions = actions(() => [InputText({ into: "id:Amount", text: String(authoredValues.generate()) })]);
`, generator)
}
// TestModelPolicyRefusesAnAuthoredValueGenerator covers every generator in
// values.ts whose span is wider than one value.
func TestModelPolicyRefusesAnAuthoredValueGenerator(t *testing.T) {
for _, generator := range []struct{ name, expression string }{
{"integers", "integers().between(1, 500)"},
{"strings", "strings().length(3, 6).alpha()"},
{"emails", `emails().domain("folio.app")`},
{"edgeCaseText", "edgeCaseText()"},
} {
t.Run(generator.name, func(t *testing.T) {
verifier := newVerifier(t, WithSeed(0x5eed))
loadActionSpec(t, verifier, valueGeneratorSpec(generator.expression))
pushTree(t, verifier, policyTreeJSON)
_, err := verifier.Candidates(LabelSourceVisibleText)
if err == nil {
t.Fatalf("%s was enumerated for the model policy, which cannot draw it", generator.name)
}
if !strings.Contains(err.Error(), "authoredValues.generate()") {
t.Errorf("error does not name the offending leaf: %v", err)
}
if !strings.Contains(err.Error(), generator.name+"()") {
t.Errorf("error does not name %s(): %v", generator.name, err)
}
})
}
}
// TestModelPolicyAcceptsASingleValuedGenerator is the boundary: a generator that
// spans one value hands both policies the same value, so refusing it would stop
// runs that have nothing wrong with them.
func TestModelPolicyAcceptsASingleValuedGenerator(t *testing.T) {
verifier := newVerifier(t, WithSeed(0x5eed))
loadActionSpec(t, verifier, valueGeneratorSpec("integers().between(7, 7)"))
pushTree(t, verifier, policyTreeJSON)
candidates := mustCandidates(t, verifier, LabelSourceVisibleText)
if len(candidates) != 1 || candidates[0].Action.Text != "7" {
t.Fatalf("model was offered %+v, want the one authored InputText typing 7", candidates)
}
}
// TestSeededValueDrawsSurviveTheModelPolicysRefusal is the values.ts half of the
// guard above: the seeded arm keeps its whole range, and its draw stream does not
// move because the model policy refused the same spec first.
func TestSeededValueDrawsSurviveTheModelPolicysRefusal(t *testing.T) {
spec := valueGeneratorSpec("integers().between(1, 500)")
alone := seededSamplerStream(t, spec, false)
afterRefusal := seededSamplerStream(t, spec, true)
if !slices.Equal(alone, afterRefusal) {
t.Error("a refused enumeration moved the seeded draw stream")
}
drawn := map[string]bool{}
for _, action := range alone {
drawn[action] = true
}
if len(drawn) < 100 {
t.Errorf("seeded picker typed %d distinct values over %d draws", len(drawn), seededDrawBudget)
}
}
+89
View File
@@ -0,0 +1,89 @@
package verifier
import (
"github.com/dop251/goja"
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// RedactedInputText stands in for a typed value in every record. It is fixed,
// so it gives away neither the value nor its length.
const RedactedInputText = "[redacted]"
// secureFact is what a target says about being a secure text entry. `reported`
// separates "the platform says this is not one" from "the platform says
// nothing", which are the two cases the redaction rule has to tell apart.
type secureFact struct {
reported bool
secure bool
}
// secureFactFromHandle reads an element handle's own report. The web host
// injects handles built in the page, which carry no selector to resolve against
// the goja-side tree, so the handle is the only thing that knows.
func secureFactFromHandle(object *goja.Object) secureFact {
value := object.Get("secure")
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return secureFact{}
}
return secureFact{reported: true, secure: value.ToBoolean()}
}
func secureFactOf(element *hierarchy.Element) secureFact {
if element == nil {
return secureFact{}
}
return secureFact{reported: element.SecureReported(), secure: element.Secure}
}
// RecordedActionText renders the typed value of an action for anything that is
// persisted or sent, resolving the action's target in the tree it was chosen
// against.
func RecordedActionText(action Action, tree *hierarchy.Tree) string {
if action.Kind != ActionKindInputText {
return action.Text
}
var target *hierarchy.Element
if tree != nil && action.On != "" {
target = tree.Find(action.On)
}
return recordedInputText(action.Text, secureFactOf(target))
}
// RecordedAction is the action as everything downstream of the dispatch sees
// it. The runner reports the previous step's action to the spec as
// state.lastAction, and a spec extracting it (examples/folio/sanderling/spec.ts)
// writes it to the trace, so the copy the runner keeps carries the recorded
// text rather than the typed one.
//
// The decision is made here rather than where the two hosts render
// state.lastAction because only the caller holds the tree that can answer it.
// The hosts hold the NEXT step's tree, where the selector may name a different
// element (a revealed password field reports secure:false) or none at all, and
// cmd/internal-tools/oracle-reduction replays state.lastAction from the trace's
// already-recorded text, which redacting here matches exactly.
func RecordedAction(action Action, tree *hierarchy.Tree) Action {
action.Text = RecordedActionText(action, tree)
return action
}
// recordedInputText is the one place a typed value is rendered for a record.
// The prompt's recent-action memory, the numbered candidate list, the trace and
// the action the runner reports back as state.lastAction all go through it, so
// a fifth record added later cannot publish a value the other four withhold.
// The driver dispatch reads Action.Text directly and is the only reader of the
// real value, which is what keeps the app receiving the keystrokes a user would
// have produced.
//
// A target the platform reports as a secure entry is redacted, and so is a
// target carrying no report at all. iOS and web state the fact on every
// editable element, so a missing one means Android, whose native tree mapper
// drops uiautomator's password attribute before the sidecar sees it. There a
// password field cannot be told from a search box, and the target that cannot
// be told apart is treated as the credential.
func recordedInputText(text string, target secureFact) string {
if target.reported && !target.secure {
return text
}
return RedactedInputText
}
+196
View File
@@ -0,0 +1,196 @@
package verifier
import (
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// typedCredential stands in for what a login setup types. Nothing rendered for
// a record may carry it.
const typedCredential = "fixture-passphrase-9f21"
// secureLoginTreeJSON is the shape iOS and web produce for a login form: both
// report `secure` on every editable field, so a field carrying `secure:false`
// is positively known not to be a credential entry.
const secureLoginTreeJSON = `{
"attributes": {"bounds": "[0,0,390,844]", "class": "Window"},
"children": [
{"attributes": {"identifier": "LoginEmail", "class": "TextField", "hintText": "Email", "bounds": "[20,200,370,244]"},
"editable": true, "enabled": true, "secure": false, "children": []},
{"attributes": {"identifier": "LoginPassword", "class": "SecureTextField", "hintText": "Password", "bounds": "[20,260,370,304]"},
"editable": true, "enabled": true, "secure": true, "children": []}
]
}`
// androidLoginTreeJSON is the same form as Android reports it. The native tree
// mapper drops uiautomator's password attribute, so neither field carries the
// fact and the two are indistinguishable.
const androidLoginTreeJSON = `{
"attributes": {"bounds": "[0,0,1080,2340]"},
"children": [
{"attributes": {"resource-id": "login_email", "class": "android.widget.EditText", "hintText": "Email", "bounds": "[20,200,1060,300]"},
"enabled": true, "children": []},
{"attributes": {"resource-id": "login_password", "class": "android.widget.EditText", "hintText": "Password", "bounds": "[20,320,1060,420]"},
"enabled": true, "children": []}
]
}`
func authoredTypingCandidates(t *testing.T, selector, text, treeJSON string) []ActionCandidate {
t.Helper()
actions := `{kind:'actions', generate: () => [{kind:'InputText', into:'` + selector +
`', text:'` + text + `'}]}`
return mustCandidates(t, enumVerifier(t, actions, treeJSON), LabelSourceVisibleText)
}
func TestCandidatesRedactTextTypedIntoASecureField(t *testing.T) {
candidates := authoredTypingCandidates(t, "id:LoginPassword", typedCredential, secureLoginTreeJSON)
if len(candidates) != 1 {
t.Fatalf("candidates = %v, want the one authored typing action", descriptions(candidates))
}
candidate := candidates[0]
if !strings.Contains(candidate.Description, RedactedInputText) {
t.Errorf("candidate description = %q, want the typed value redacted", candidate.Description)
}
if !strings.Contains(candidate.Description, "Password") {
t.Errorf("candidate description = %q, want it to still name the field typed into", candidate.Description)
}
if candidate.Action.Text != typedCredential {
t.Errorf("candidate action text = %q, want the real value so the driver still types it", candidate.Action.Text)
}
}
func TestCandidatesRedactEveryTypedValueWhereTheTargetCannotReportSecure(t *testing.T) {
for _, selector := range []string{"id:login_email", "id:login_password"} {
candidates := authoredTypingCandidates(t, selector, typedCredential, androidLoginTreeJSON)
if len(candidates) != 1 {
t.Fatalf("%s: candidates = %v, want the one authored typing action", selector, descriptions(candidates))
}
if strings.Contains(candidates[0].Description, typedCredential) {
t.Errorf("%s: candidate description carries the typed value: %q", selector, candidates[0].Description)
}
}
}
func TestCandidatesKeepTextTypedIntoAFieldReportedNotSecure(t *testing.T) {
const address = "[email protected]"
candidates := authoredTypingCandidates(t, "id:LoginEmail", address, secureLoginTreeJSON)
if len(candidates) != 1 {
t.Fatalf("candidates = %v, want the one authored typing action", descriptions(candidates))
}
if !strings.Contains(candidates[0].Description, address) {
t.Errorf("candidate description = %q, want the typed value on a field reported not secure", candidates[0].Description)
}
}
// A spec that types into an element HANDLE is the shape every login setup has
// (examples/folio/sanderling/spec.ts). The handle carries the goja host's own
// `secure` field, which is a plain boolean and reads false on a platform that
// reports nothing, so a rule trusting it would leave every Android value in the
// clear. The tree the handle names is the only carrier of the three-way fact.
func TestCandidatesRedactTextTypedIntoAnAndroidElementHandle(t *testing.T) {
v := newVerifier(t)
loadActionSpec(t, v, `
import { InputText, actions, extract } from "@sanderling/spec";
const field = extract((s) => s.ax.find({ "resource-id": "login_password" })).named("field");
globalThis.actions = actions(() =>
field.current ? [InputText({ into: field.current, text: "`+typedCredential+`" })] : []);
`)
pushTree(t, v, androidLoginTreeJSON)
candidates := mustCandidates(t, v, LabelSourceVisibleText)
if len(candidates) != 1 {
t.Fatalf("candidates = %v, want the one authored typing action", descriptions(candidates))
}
if strings.Contains(candidates[0].Description, typedCredential) {
t.Errorf("candidate description carries the typed value: %q", candidates[0].Description)
}
if candidates[0].Action.Text != typedCredential {
t.Errorf("candidate action text = %q, want the real value so the driver still types it",
candidates[0].Action.Text)
}
}
// lastActionOnBothHosts reports what a spec reading state.lastAction sees on
// each host for an action the runner has recorded: the goja host stringifies
// its own object, the web host receives EncodeLastAction's JSON and installs
// the parsed value in the page.
func lastActionOnBothHosts(t *testing.T, action Action, treeJSON string) (string, string) {
t.Helper()
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
recorded := RecordedAction(action, tree)
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.last = __sanderling__.extract(state => JSON.stringify(state.lastAction));
`)
if err := verifier.PushSnapshot(SnapshotInput{
Snapshots: Snapshots{},
Tree: tree,
LastAction: &recorded,
}); err != nil {
t.Fatalf("PushSnapshot: %v", err)
}
handle := verifier.runtime.GlobalObject().Get("last").ToObject(verifier.runtime)
return handle.Get("current").String(), string(EncodeLastAction(&recorded))
}
func typedInto(selector, text string) Action {
return Action{Kind: ActionKindInputText, On: selector, Text: text}
}
// A spec extracting state.lastAction (examples/folio/sanderling/spec.ts) writes
// what it reads into the trace, so state.lastAction is a record like the other
// three and carries the same rule on both hosts.
func TestStateLastActionRedactsATypedValueTheTargetCannotClear(t *testing.T) {
for _, testCase := range []struct {
name string
treeJSON string
selector string
}{
{"secure field", secureLoginTreeJSON, "id:LoginPassword"},
{"android field reported as neither", androidLoginTreeJSON, "id:login_email"},
{"android password field", androidLoginTreeJSON, "id:login_password"},
} {
t.Run(testCase.name, func(t *testing.T) {
goja, web := lastActionOnBothHosts(
t, typedInto(testCase.selector, typedCredential), testCase.treeJSON)
for _, host := range []struct{ name, reported string }{
{"goja", goja},
{"web", web},
} {
if strings.Contains(host.reported, typedCredential) {
t.Errorf("the %s host publishes the typed value in state.lastAction: %s",
host.name, host.reported)
}
if !strings.Contains(host.reported, RedactedInputText) {
t.Errorf("the %s host reports state.lastAction as %s, want the typed value "+
"redacted in place", host.name, host.reported)
}
}
if goja != web {
t.Errorf("the two hosts disagree on state.lastAction\n goja: %s\n web: %s", goja, web)
}
})
}
}
func TestStateLastActionKeepsATypedValueForAFieldReportedNotSecure(t *testing.T) {
const address = "[email protected]"
goja, web := lastActionOnBothHosts(t, typedInto("id:LoginEmail", address), secureLoginTreeJSON)
for _, host := range []struct{ name, reported string }{
{"goja", goja},
{"web", web},
} {
if !strings.Contains(host.reported, address) {
t.Errorf("the %s host reports state.lastAction as %s, want the typed value on a "+
"field reported not secure", host.name, host.reported)
}
}
if goja != web {
t.Errorf("the two hosts disagree on state.lastAction\n goja: %s\n web: %s", goja, web)
}
}
+82
View File
@@ -2,8 +2,10 @@ package verifier
import (
"errors"
"maps"
"os"
"path/filepath"
"slices"
"testing"
"github.com/priyanshujain/sanderling/internal/bundler"
@@ -93,3 +95,83 @@ export const actionsRoot = taps;
t.Fatalf("NextAction should draw from actionsRoot: %v", err)
}
}
// TestNextActionNamesTheGeneratorThatProducedIt is the native half of the
// cross-host marker contract: for the same spec shape, a setup-driven step and
// an action-root step must name different producers. The web half asserts the
// same two names in pkg/spec/test/web-runtime.test.ts, so a match on both sides
// proves the engines agree without either invoking the other. A per-action rate
// divides by the root's steps only, and an unnamed producer inflates it by
// however many steps the login consumed.
func TestNextActionNamesTheGeneratorThatProducedIt(t *testing.T) {
spec := `
import { Tap, actions, extract, taps } from "@sanderling/spec";
const signIn = extract("signIn", state => state.ax.find("id:SignIn"));
export const setup = actions(() => (signIn.current ? [Tap({ on: "id:SignIn" })] : []));
export const actionsRoot = taps;
`
v := loadBundled(t, spec, enumTreeJSON)
action, err := v.NextAction()
if err != nil {
t.Fatalf("NextAction: %v", err)
}
if action.On != "id:SignIn" || action.Source != "setup" {
t.Errorf("NextAction = %+v, want the Tap on id:SignIn named setup", action)
}
pushTree(t, v, policyTreeJSON)
action, err = v.NextAction()
if err != nil {
t.Fatalf("NextAction after setup went quiet: %v", err)
}
if action.Source != "seeded" {
t.Errorf("NextAction = %+v, want the action root's tap named seeded", action)
}
}
// TestSetupActionNamesSetup: the model policy reaches setup through its own
// entry, which must name the producer the same way the seeded entry does, or
// the two arms' per-action rates divide by different things.
func TestSetupActionNamesSetup(t *testing.T) {
spec := `
import { Tap, actions, taps } from "@sanderling/spec";
export const setup = actions(() => [Tap({ on: "id:SignIn" })]);
export const actionsRoot = taps;
`
v := loadBundled(t, spec, enumTreeJSON)
action, err := v.SetupAction()
if err != nil {
t.Fatalf("SetupAction: %v", err)
}
if action.Source != "setup" {
t.Errorf("SetupAction = %+v, want it named setup", action)
}
}
// TestSetupGeneratorDrawsUnderTheModelPolicy: setup walks through the picker
// with its rng under both policies, so a generator there is not the divergence
// the enumeration refuses and must keep drawing. Enumerating before every setup
// step is what a model-driven run does, and the refusal must not leak out of it.
func TestSetupGeneratorDrawsUnderTheModelPolicy(t *testing.T) {
spec := `
import { InputText, actions, integers, taps } from "@sanderling/spec";
const setupValues = integers().between(1, 500);
export const setup = actions(() => [InputText({ into: "id:Amount", text: String(setupValues.generate()) })]);
export const actionsRoot = taps;
`
v := loadBundled(t, spec, policyTreeJSON)
typed := map[string]bool{}
for range 16 {
if _, err := v.Candidates(LabelSourceVisibleText); err != nil {
t.Fatalf("Candidates: %v", err)
}
action, err := v.SetupAction()
if err != nil {
t.Fatalf("SetupAction: %v", err)
}
typed[action.Text] = true
}
if len(typed) < 2 {
t.Errorf("setup typed %v on every step; the picker's rng did not reach it", slices.Sorted(maps.Keys(typed)))
}
}
+4
View File
@@ -38,6 +38,10 @@ type Action struct {
// when the apply call failed and nothing can say whether the action
// reached the app. The spec is told which of the two it is.
Applied bool
// Source names the generator that produced this action, "setup" or
// "seeded", as the runtime entry tagged it. Empty on an action the runner
// built itself (the model policy's pick), which the runner names instead.
Source string
// Relaunched, like Applied, is meaningful only on state.lastAction: the
// runner brought the app back to the foreground after this action, so the
// two readings the spec compares straddle a restart. The action still
+125
View File
@@ -3,9 +3,11 @@ package verifier
import (
"encoding/json"
"errors"
"fmt"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"testing"
@@ -1217,6 +1219,52 @@ func TestChangedExtractors_DiffsBetweenSnapshots(t *testing.T) {
}
}
// TestPushSnapshot_ReportsAValueItCannotRecord covers the residue the
// projection cannot reach: a value with no JSON form at all. The author has to
// hear about it, because an extractor dropped from the diff is indistinguishable
// from one that never changed.
func TestPushSnapshot_ReportsAValueItCannotRecord(t *testing.T) {
verifier := newVerifier(t)
if err := verifier.runtime.GlobalObject().Set("hostChannel", make(chan int)); err != nil {
t.Fatal(err)
}
mustLoad(t, verifier, `__sanderling__.extract(() => globalThis.hostChannel, "wedged");`)
err := verifier.PushSnapshot(SnapshotInput{})
if err == nil {
t.Fatal("PushSnapshot accepted a value it cannot record; the extractor would vanish from the trace")
}
if !strings.Contains(err.Error(), "wedged") {
t.Errorf("error does not name the extractor: %v", err)
}
}
// TestChangedExtractors_RecordsValuesJSONCannotHold pins the projection's
// edges: a cycle and a NaN are recorded as null rather than costing the run,
// which is what JSON.stringify does with them on the web host.
func TestChangedExtractors_RecordsValuesJSONCannotHold(t *testing.T) {
verifier := newVerifier(t)
mustLoad(t, verifier, `
__sanderling__.extract(() => { const node = { n: 1 }; node.self = node; return node; }, "cyclic");
__sanderling__.extract(() => 0 / 0, "notANumber");
`)
if err := verifier.PushSnapshot(SnapshotInput{}); err != nil {
t.Fatalf("PushSnapshot: %v", err)
}
changes := verifier.ChangedExtractors()
cyclic, ok := changes["cyclic"]
if !ok {
t.Fatalf("cyclic extractor missing from the diff: %+v", changes)
}
if got := string(cyclic.Curr); got != `{"n":1,"self":null}` {
t.Errorf("cyclic recorded as %s, want the data with the cycle cut", got)
}
if got := string(verifier.extractorSnapshot()["notANumber"]); got != "null" {
t.Errorf("NaN recorded as %s, want null (a run must not die on it)", got)
}
}
// TestExtract_DefaultsAndNamedNames verifies bindExtract assigns a fallback
// `extractor_N` name when no name is supplied and respects an explicit one.
func TestExtract_DefaultsAndNamedNames(t *testing.T) {
@@ -1433,3 +1481,80 @@ func TestExtractorCount_ReportsEveryRegisteredExtractor(t *testing.T) {
t.Errorf("ExtractorCount() = %d, want 2 (helloSpec registers screen and balance)", got)
}
}
// TestAxSelectorTag_OnlyNamesTheElementItAloneResolvesTo pins the rule the
// runner depends on: resolveCoordinates prefers tree.Find(action.On) over the
// coordinates the element reported, and Find takes the first match, so a
// selector three sibling cards share would send all three taps to the first
// card. Siblings therefore carry no selector and keep their own coordinates;
// an element the selector alone resolves to still carries it.
func TestAxSelectorTag_OnlyNamesTheElementItAloneResolvesTo(t *testing.T) {
const treeJSON = `{
"attributes": {"resource-id": "root", "bounds": "[0,0,100,400]"},
"children": [
{"attributes": {"testTag": "AccountCard", "text": "Alpha", "bounds": "[0,0,100,100]"}, "clickable": true, "children": []},
{"attributes": {"testTag": "AccountCard", "text": "Beta", "bounds": "[0,100,100,200]"}, "clickable": true, "children": []},
{"attributes": {"testTag": "AccountCard", "text": "Gamma", "bounds": "[0,200,100,300]"}, "clickable": true, "children": []},
{"attributes": {"testTag": "AddAccount", "bounds": "[0,300,100,400]"}, "clickable": true, "children": []}
]
}`
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.cards = __sanderling__.extract(state =>
state.ax.findAll({ testTag: "AccountCard" })
);
globalThis.sole = __sanderling__.extract(state =>
state.ax.findAll({ testTag: "AddAccount" })
);
globalThis.soleFind = __sanderling__.extract(state =>
state.ax.find({ testTag: "AddAccount" })
);
`)
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatal(err)
}
if err := verifier.PushSnapshot(SnapshotInput{Snapshots: Snapshots{}, Tree: tree}); err != nil {
t.Fatal(err)
}
current := func(name string) *goja.Object {
handle := verifier.runtime.GlobalObject().Get(name).ToObject(verifier.runtime)
value := handle.Get("current")
if goja.IsUndefined(value) || goja.IsNull(value) {
t.Fatalf("%s: extractor produced no value", name)
}
return value.ToObject(verifier.runtime)
}
element := func(array *goja.Object, index int) *goja.Object {
return array.Get(strconv.Itoa(index)).ToObject(verifier.runtime)
}
cards := current("cards")
if got := cards.Get("length").ToInteger(); got != 3 {
t.Fatalf("findAll returned %d cards, want 3", got)
}
centers := map[string]bool{}
for index := range 3 {
card := element(cards, index)
if got := card.Get(tagSelector).String(); got != "" {
t.Errorf("card %d (%s) carries selector %q; three cards answer to it, so the runner would tap the first card three times",
index, card.Get("text"), got)
}
centers[fmt.Sprintf("%d,%d", card.Get("x").ToInteger(), card.Get("y").ToInteger())] = true
}
if len(centers) != 3 {
t.Errorf("the three cards report %d distinct centers, want 3: %v", len(centers), centers)
}
soleAll := current("sole")
if got := soleAll.Get("length").ToInteger(); got != 1 {
t.Fatalf("findAll returned %d AddAccount elements, want 1", got)
}
if got := element(soleAll, 0).Get(tagSelector).String(); got != "testTag:AddAccount" {
t.Errorf("findAll over a single match: selector = %q, want %q", got, "testTag:AddAccount")
}
if got := current("soleFind").Get(tagSelector).String(); got != "testTag:AddAccount" {
t.Errorf("find over a single match: selector = %q, want %q", got, "testTag:AddAccount")
}
}
+78 -11
View File
@@ -47,6 +47,13 @@ type Verifier struct {
// over the picker's action space rather than one of its own.
enumerateBuiltinFn goja.Callable
// setEnumeratingCandidatesFn is the bundle-installed
// __sanderlingSetEnumeratingCandidates__, which brackets the model policy's
// authored-leaf calls. Those run outside the picker's rng scope, where a
// sampler would quietly hand back its first item, so the bundle refuses to
// sample while it is set.
setEnumeratingCandidatesFn goja.Callable
evaluators map[string]*ltl.Evaluator
priorVerdicts map[string]ltl.Verdict
@@ -194,6 +201,12 @@ func (v *Verifier) Load(source string) error {
}
}
if fn := v.runtime.GlobalObject().Get("__sanderlingSetEnumeratingCandidates__"); fn != nil {
if callable, ok := goja.AssertFunction(fn); ok {
v.setEnumeratingCandidatesFn = callable
}
}
return nil
}
@@ -342,8 +355,14 @@ func (v *Verifier) PushSnapshot(input SnapshotInput) error {
return fmt.Errorf("extractor %d: %w", index, err)
}
extractor.currentValue = newValue
encoded, err := encodeExtractorValue(newValue)
if err != nil {
return fmt.Errorf(
"extractor %q: the value cannot be recorded in the trace: %w; return plain data instead",
extractor.name, err)
}
extractor.prev = extractor.curr
extractor.curr = encodeExtractorValue(newValue)
extractor.curr = encoded
}
return nil
}
@@ -358,17 +377,19 @@ func (v *Verifier) runExtractor(extractor *extractorState, state goja.Value) (go
}
// encodeExtractorValue produces a stable JSON encoding of an extractor's
// current value for diff comparison. Values that still don't survive encoding
// yield nil; callers treat nil as "unknown" and emit no diff entry.
func encodeExtractorValue(value goja.Value) []byte {
// current value for diff comparison. Host functions, cycles and non-finite
// numbers are projected away by recordableValue; anything still beyond JSON is
// an error, never a silently dropped value, because an extractor missing from
// the trace reads exactly like an extractor that never changed.
func encodeExtractorValue(value goja.Value) ([]byte, error) {
if value == nil || goja.IsUndefined(value) || goja.IsNull(value) {
return []byte("null")
return []byte("null"), nil
}
body, err := json.Marshal(recordableValue(value.Export(), 0, map[uintptr]bool{}))
if err != nil {
return nil
return nil, err
}
return body
return body, nil
}
// recordableMaxDepth mirrors SANITIZE_MAX_DEPTH in pkg/spec/src/web-runtime.ts.
@@ -437,6 +458,8 @@ func recordableValue(value any, depth int, seen map[uintptr]bool) any {
func (v *Verifier) ChangedExtractors() map[string]ExtractorChange {
changes := map[string]ExtractorChange{}
for _, extractor := range v.extractors {
// nil curr now means only that no snapshot has been pushed yet: an
// encoding that cannot be recorded fails PushSnapshot instead.
if extractor.curr == nil {
continue
}
@@ -464,6 +487,45 @@ func (v *Verifier) ExtractorCount() int {
return len(v.extractors)
}
// ExtractorNames returns every extractor's name in registration order, which is
// the order OverrideExtractorValues is keyed by. A trace records extractor
// values by name, so replaying one offline needs the name-to-index mapping the
// spec fixed at load.
func (v *Verifier) ExtractorNames() []string {
names := make([]string, 0, len(v.extractors))
for _, extractor := range v.extractors {
names = append(names, extractor.name)
}
return names
}
// PropertyNames returns the names of the properties the loaded spec registered,
// sorted.
func (v *Verifier) PropertyNames() []string {
names := make([]string, 0, len(v.properties))
for name := range v.properties {
names = append(names, name)
}
sort.Strings(names)
return names
}
// PropertyFormulas rebuilds each registered property's formula. The thunks are
// this verifier's own predicates, reading this verifier's extractor state, so
// an evaluator built over a rewritten formula observes exactly what the
// engine's evaluator does.
func (v *Verifier) PropertyFormulas() (map[string]ltl.Formula, error) {
formulas := make(map[string]ltl.Formula, len(v.properties))
for name, specIndex := range v.properties {
formula, err := v.buildFormula(specIndex)
if err != nil {
return nil, fmt.Errorf("property %q: %w", name, err)
}
formulas[name] = formula
}
return formulas, nil
}
// OverrideExtractorValues replaces each extractor's `current` slot with a
// caller-supplied value, keyed by registration index. Used by the web tick
// path so extractor bodies that ran in V8 (against the real DOM) drive the
@@ -497,7 +559,13 @@ func (v *Verifier) OverrideExtractorValues(overrides map[int]json.RawMessage) (s
return skipped, fmt.Errorf("extractor override %d: %w", index, conversionErr)
}
v.extractors[index].currentValue = value
v.extractors[index].curr = encodeExtractorValue(value)
encoded, encodeErr := encodeExtractorValue(value)
if encodeErr != nil {
return skipped, fmt.Errorf(
"extractor override %d (%q): the value cannot be recorded in the trace: %w",
index, v.extractors[index].name, encodeErr)
}
v.extractors[index].curr = encoded
}
return skipped, nil
}
@@ -603,9 +671,8 @@ func (v *Verifier) captureWitness(name string) {
}
}
// extractorSnapshot encodes every named extractor's current value as JSON. A
// nil value (extractor never advanced or its value did not survive Export)
// is recorded as JSON null.
// extractorSnapshot encodes every named extractor's current value as JSON. An
// extractor that never advanced is recorded as JSON null.
func (v *Verifier) extractorSnapshot() map[string]json.RawMessage {
if len(v.extractors) == 0 {
return nil