Files
pj 9b4ff5f247 record what the model picker did, and make both policies see the same actions (#74)
* feat(llmclient): parse usage and the served model

An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record one typed outcome per model-driven step

llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): expose the step a snapshot was observed at

It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a guard-skipped step is no longer a silent log line

The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): record when a chosen action was never dispatched

A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): document llm-calls.jsonl

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(analyze): survival analysis over campaign directories

Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): select the candidate label source

Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(runner): thread the label source to the model picker

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the label source as arm membership

Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --label-source

Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): dedup candidates by what they execute, not how they read

The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): report every action that was chosen and never dispatched

applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): the echo guard admits a repeated description

Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): count dispatched actions, not steps

A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): lower authored actions the way the seeded arm does

The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): decode a container-only scroll

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): carry the container on an authored scroll

serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): both policies must dispatch the same authored action

Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(hierarchy): match identifiers by role prefix

idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.

* feat(chrome): translate idPrefix to a starts-with id match

* feat(spec): match idPrefix in the web runtime

The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.

* feat(sidecar): match idPrefix in the tap-by-selector path

* docs(manual): document the idPrefix selector

* feat(replay-ui): render idPrefix targets as a prefix tag

* fix(spec): read the injected seed per call

Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.

* test(chrome): compare both selector matchers over one live page

Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.

* fix(hierarchy): give id and desc one meaning in both selector forms

The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.

* test(hierarchy): pin both selector forms and the unknown-key report

* feat(verifier): fail the spec on a selector key that cannot match

An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.

* feat(spec): reject an unknown object-selector key in the web runtime

Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.

* test(spec): pin the unknown-key diagnostic to one text

The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.

* fix(spec): match a merged label by its leading name on web too

The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.

* test(chrome): drive the live-page parity test through both selector forms

* docs(manual): document object-selector key rules

* feat(spec): refuse a multi-item authored sampler while enumerating

from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets

Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): abort on a candidate enumeration that refused

Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(spec): refuse a multi-value generator while enumerating

integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(verifier): setup still draws, and the seeded stream is unmoved

Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): value generators are refused under the model policy too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio): enumerate authored targets and values instead of sampling

Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): enumerate authored targets and values, declare the llm generator

The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: minimal changes, self-documenting code, tests as first-class

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): key web attrs by the names the markup writes

attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name a web handle by the same ladder as a tree element

The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): attrs carries raw attribute names on web too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): confirm focus moved before typing

InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): signal a timed-out run so it reaps its sidecar

CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* perf(runner): confirm focus only when another element holds it

Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): record both clocks a run was measured on

Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide per-hour rates by time actually worked

A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(campaign): wait for the trap instead of racing it

The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(ltl): keep the authored window on a step-bounded obligation

reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(ltl): pin that a slow policy does not fail on time alone

Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(spec): guard the step unit on the authoring surface

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): bound the reachability properties by steps

At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): a step bound counts observations, not runner steps

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): make the label source a cell dimension

A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): record the label source in the manifest

A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): name a web field by its hint, not its CSS class

visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): answer clickable for an element reached through ax

The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: every target runs on this machine, so start one rather than skip it

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a selector the tree cannot resolve is not a focus failure

otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit data-testid so both resolvers name the same element

The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name an element only when the selector names it alone

ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): hold the V8 host to the same naming gate

elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): sibling taps reach the driver at their own coordinates

Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): an ambiguous name loses to the coordinates it was built from

Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): record an element-valued extractor instead of dropping it

An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.

* test(verifier): an unrecordable extractor value is reported, not dropped

* test(runner): element-valued extractors reach the trace

* docs(spec-language): say what a trace records for an element-valued extractor

* docs(claude): add delegation and record-keeping sections

delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.

* feat(driver): declare undelivered-action errors and three optional capabilities

ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.

* feat(driver): add escape to the pressKey surface

escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.

* fix(ios): refuse a gesture the screen has no surface under

the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.

* feat(ios): derive scrollable from the snapshot's tree depth

the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.

* fix(sidecar): stop dropping gestures, selectors and keys in silence

a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.

* fix(sidecar): map the driver's refusals onto the gesture errors

OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.

* fix(selectors): resolve text to the innermost match and scan the root in both forms

an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.

* feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag

a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.

* fix(chrome): emit every markup attribute and read checked and selected off the property

the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.

* fix(chrome): scroll a gesture point into view and dispatch trusted input

getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.

* feat(chrome): read the page's exceptions and navigations, and hold the picker state across them

a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.

* feat(trace): version each step and record its logs, exceptions and navigations

a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.

* feat(runner): bound every device call and record the actions that never reached the app

observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.

* feat(verifier): expose extractor names and rebuilt property formulas

an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.

* feat(testrun): expose the seeded bundle a run loaded

BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity.

* feat(tracecorpus): load recorded runs for offline measures

reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.

* refactor(seedspec): move seed spec parsing out of the campaign command

the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.

* feat(analyze): time an event at the step it was detected and report the quartiles

an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.

* feat(analyze): add the seed-paired signed-rank comparison and record the holm family

--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.

* test(analyze): recover planted effects through the tool's own entry point

a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.

* feat(label-coverage): report the addressable share of an app's interactive surface

reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.

* feat(exploration-reach): count the distinct structural states a stored run visited

the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.

* feat(defect-identity): count distinct defects across stored runs

a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.

* feat(oracle-reduction): replay stored traces under four reduced oracles

re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.

* feat(implementation-sweep): run one campaign against every implementation of a requirement

installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.

* feat(corpus-sweep): run one specification against a served corpus of implementations

same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.

* docs(manual): document innermost text matching, escape and the web scroll verb

text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.

* test(browser): assert an uncaught page exception reaches the trace

the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.

* feat(trace): a step can name the precondition it could not meet

A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.

* fix(runner): budget the foreground gate in time, not in polls

Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.

* test(runner): the gate keeps looking until its budget runs out

Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.

* feat(campaign): count the runs that were never in the app

A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.

* docs(triage): name the trace field a run that never started leaves

* fix(selectors): tag names the whole tag, not a substring of it

matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.

* test(chrome): both resolvers agree on tag where a container's name contains its child's

* fix(make): build the binary instead of matching the build directory

build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.

* feat(verifier): expose the property names a loaded spec registered

* feat(testrun): refuse a run against a spec that registers no properties

A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.

* feat(cli): --allow-no-properties opts a run out of the refusal

* docs(cli): document --allow-no-properties

* feat(bundle-check): fail a spec that bundles but registers no properties

* test(bundle-check): cover the zero-property refusal and pin the reported bundle

* feat(folio-web): predicates for counting commits against submit actions

* feat(folio-web): judge one commit per submit over a home-card window

Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.

* fix(folio-web): keep submit live for 400ms after saving

Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.

* feat(confusion-matrix): score the checker against a blind reviewer

Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.

* test(confusion-matrix): reject malformed inputs and keep missing data out of the cells

* test(confusion-matrix): cover cell assignment, precision and recall

* fix(chrome): focus descends into the shadow root

document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.

* test(implementation-sweep): supply the binaries the missing-binary test does not test

resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.

* fix(replay-ui): read data-* attributes by their markup names

The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.

* fix(web-runtime): focus descends into the shadow root here too

The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.

* fix(implementation-sweep): name every missing binary, in flag order

Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.

* fix(chrome): focus follows the caret to the field it types into

Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.

* fix(corpus-sweep): name every missing binary, in flag order

Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.

* fix(web-runtime): a handle answers editable for itself, not its container

isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.

* test(chrome): a hinted field is not named by its css class

The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.

* test(chrome): the handle and the enumeration agree on editable too

The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.

* fix(web-runtime): focus follows the caret to the field it types into

Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.

* fix(campaign): name every missing required flag, in flag order

Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.

* fix(corpus-sweep): name every missing required flag, in flag order

* fix(implementation-sweep): name every missing required flag, in flag order

* fix(confusion-matrix): name every missing required flag, in flag order

* ci: pin the idb-companion tap to the formula the companion is staged from

The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.

* fix(campaign): refuse to start on a device that is not there

A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): quarantine a device that keeps failing fast

Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the device a run executed on

meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* ci: let a restored ios bundle survive make's mtime check

The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.

* fix(confusion-matrix): a campaign that died is missing data, not a true negative

The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.

* fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.

* fix(runner): a source that was asked and handed nothing says so

NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.

* feat(testrun): refuse a run that dispatched none of its actions

Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.

* docs(cli): document --label-source

* docs(spec-language): name the hintText selector's host divergence

The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.

* feat(bundle-check): --allow-no-properties opts out of the refusal

The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.

* test(verifier): an unreadable committed fixture fails, it does not skip

The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.

* fix(testrun): the refusal asks whether the generator drove, not whether anything did

A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.

* feat(runner): the summary says how many steps the generator drove

A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.

* fix(testrun): an ios run records the simulator it executed on

Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.

* fix(campaign): the action count leaves the setup's login out on a model run

Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.

* feat(hierarchy): an element reports whether it masks what is typed into it

ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".

* fix(verifier): a secure field's typed value never reaches the record

A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.

* fix(runner): a secure field's value does not reach state.lastAction either

folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.

* fix(trace): an action names the generator that produced it

The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.

* fix(defect-identity): degrade a redacted origin action to its selector

The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.

* fix(campaign): a record always says how many actions named no producer

An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.

* fix(analyze): read how much of a record's action count names no producer

A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.

* fix(analyze): refuse to compare attributed and unattributed denominators

One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.

* fix(analyze): mark an action count of unknown provenance in the report

* docs(manual): what an action count with no producer means for a rate

* fix(folio): install through adb so a remote adb server works

Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.

* docs(folio): say how to point just test at a remote adb server

* test(conformance): the g4 fixture holds what a redacted android trace holds

Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.

* fix(testrun): a recorded violation outranks the dead-run refusal

A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.

* fix(testrun): the dead-run refusal gets its own opt-out

--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.

* feat(cli): --allow-no-generator-actions

The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.

* refactor(analyze): open the log-rank up to a weight on the risk set

The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.

* feat(analyze): add the gehan generalized wilcoxon test

The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.

* fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.

* fix(conformance): g4 reads a doubling off the observed field value

The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.

* fix(spec): a secure selector names the password field on web

secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.

* test(chrome): resolve the secure selector on both matchers

the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.

* test(chrome): compare the secure fact across both producers

it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.

* docs(manual): state what a secure selector matches

* test(conformance): g4 keeps checking past an input typed at coordinates

An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.

* fix(conformance): g4 skips an input that names no field

jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.

* test(browser): the exit code a dead run and a violated one actually leave

Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.

* fix(spec): keep a secure selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.

* fix(analyze): score a seed pair by which run outlived the other

The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.

* docs(analyze): name the tests the tool actually runs

The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.

* docs(manual): exit 1 also means a run that holds no verdict

And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.

* docs(skills): quote the summary line the runner prints now

The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.

* test(conformance): g4 sees a doubling appended to what the field held

Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.

* refactor(analyze): hoist the sign test's loop bound

* fix(conformance): g4 reads a doubling out of what the field grew by

The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.

* fix(analyze): write an undefined paired p-value as null, not as NaN

A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.

* fix(spec): a boolean state selector names what the live element reports

clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.

* fix(spec): keep a tag selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.

* fix(chrome): state every boolean flag the dump can state

internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.

* test(chrome): resolve the five state selectors on both matchers

the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.

* test(chrome): compare checked, selected and focused across both producers

the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.

* fix(spec): keep a selector out of the head subtree

the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.

* docs(manual): state what the other boolean state selectors match

* fix(folio): refuse to install and fuzz a device nobody named

adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.

* docs(folio): state that android recipes need ANDROID_DEVICE

* fix(spec): and text with the keys written beside it

a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.

* test(spec): pin text against the key beside it in either order

object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.

* test(chrome): compare a compound text selector across both matchers

one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.

* docs(manual): state how text combines with the key beside it

the object selector section said every pair must match without saying
where the innermost rule then lands.

* fix(hierarchy): reach the class attribute through className

className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.

* test(chrome): compare className across both matchers

one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.

* test(spec): pin className and class on the same elements

this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.

* docs(manual): list className among the cross-platform aliases

the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.

* fix(hierarchy): reach the accessible label through every name for it

label and accessibilityLabel aliased onto accessibilityText alone, which
only the ios sidecar writes, and alias expansion is ONE level: the hop
from accessibilityText to content-desc was never taken, so both keys
matched nothing on android and on the chrome dump, which write the fact
under content-desc. ariaLabel and contentDescription aliased onto
nothing at all and matched nothing anywhere.

The web runtime resolves all four against the live DOM, so a selector
naming a field this way found it on one host and no element at all on
the other. The keys are accepted, so no unknown-key error fires, and a
property over the element that was never found passes having checked
nothing.

Each name lists both keys rather than chaining through accessibilityText:
transitive expansion would silently widen every existing key at once.

* fix(hierarchy): reach a web test tag through testTag and testID

Compose for Web writes a test tag as data-testid, which is what the web
runtime resolves both names against. testTag aliased onto the three
identifier keys and not that one, and testID aliased onto nothing at
all, so a tag the web runtime found on every row of a list named no
element here and every property over it passed vacuously.

* fix(spec): resolve the identifier, label and class aliases against the DOM

identifier, accessibilityIdentifier, accessibilityText and elementType
are the names ios writes four facts under, and internal/hierarchy
aliases each onto the key the other producers write. This table listed
none of them, so each fell through to a raw attribute lookup and built
[accessibilityIdentifier="summary_card"], which no element carries.

Every one of them resolved against the dump on the goja host and named
nothing here. The keys are accepted, so no unknown-key error fires, and
a property over the element that was never found passes having checked
nothing.

* fix(spec): an editable or scrollable selector names what this host derives

Both facts are derived from the live element rather than written by the
markup, and matching them as attributes built [editable="true"], which
no page carries. Both resolve against the dump on the goja host, so a
spec naming a field or a scroll container that way found it there and no
element at all here, with no unknown-key error to say so.

Each reads the same function the fact is derived with, so a selector
cannot name an element this host calls something else: the handle, the
picker's target list and the editable selector all go through
isEditable, and scrollable reads the overflow test collectTargets reads.

scrollable false names nothing rather than every element that does not
scroll: both producers state the fact only where it holds, the way an
element that is no field at all answers to neither value of secure.

* test(chrome): compare the alias keys and the two derived facts

one page, both resolvers, the ten names that resolved on one host only.
each alias is asked beside the key it resolves through, so the pair is
pinned to the same elements rather than each to itself.

the page grows a container that overflows its box and a neighbour that
does not, because scrollable is derived from the box: without one the
only scrolling element on the page is the document root, whose answer
moves with the window.

* fix(hierarchy): bounds is a raw attribute, not a cross-platform key

Every native dump writes the rectangle out as a string under bounds, and
no DOM element carries an attribute of that name, so the key resolved
against the dump and matched nothing on web on every page there is. It
is accepted, so no unknown-key error said so, and no mapping can be
invented for it: there is no DOM fact to map it to.

Off the accepted list the web runtime raises the unknown-key error
instead of matching nothing in silence, and the key still resolves
wherever a producer writes it, through the escape hatch every other raw
attribute already uses: a key some element carries is a key that can
match, on both sides.

* docs(manual): state which attribute each alias reads, and what bounds is

the table listed neither name for the accessible label that a web page
writes, nor the key a web test tag lands on, and said nothing about
elementType. editable and scrollable are boolean states like the rest,
and scrollable is the one of them the platforms state only where it
holds. bounds is a raw driver attribute rather than an accepted key.

* docs(hierarchy): the package doc names every key an alias reaches

it described the alias table as it stood before the label and test-tag
names reached the keys android and web write, and said nothing about
expansion being one level deep, which is why each name has to list every
key rather than hop through another alias.

* fix(spec): a hint selector names the ladder both producers derive

hintText and placeholderValue are the accessible-name ladder, derived
from the live element, and compiling them to [placeholder="..."] made
them name the wrong field or none at all. A field labelled by an
aria-label or a bound <label> carries no placeholder, so it resolved
against the dump on the goja host and reached nothing here; one carrying
both answered to its placeholder here where the dump answers to its
aria-label, which lands a find on an element nobody named.

Both keys read the same fieldHint elementHandle and the hierarchy dump
(internal/driver/chrome/driver.go) derive the fact with, so a selector
cannot name a field this host calls something else. An empty hint names
nothing rather than everything that is no field: both producers write
the fact only where the ladder answered.

placeholder stays the attribute the markup writes, which is what the
dump carries under that name too, so a field whose hint is something
else still answers to it on both hosts.

* fix(chrome): a hint target is not tapped by the placeholder attribute

TapSelector is a third resolver, and it built [placeholder="..."] for
hintText and placeholderValue too. Now that both matchers read the
accessible-name ladder, that CSS names a field whose hint is its
aria-label and whose placeholder happens to carry the value, which is an
element neither matcher named.

No CSS says what the ladder says, so both keys fall through to a match
that reaches nothing and the step fails naming the selector, the way
every other derived key in this file already does. A selector reaches
here only where the dump resolved it to no coordinates at all.

* test(chrome): compare the hint keys and placeholder across both matchers

The page gains four fields that differ one rung at a time: a bound
label, a placeholder, a placeholder an aria-label outranks, and the name
the form gives the field. Only the placeholder rung was reachable
before, so hintText and placeholderValue named a field on the goja host
and no element at all on web for the other three, and named the field
here and nothing there for the rung the ladder passed over.

placeholder was measured empty on both hosts because nothing on the page
carried the attribute, which said nothing about it. It now names the
field the markup wrote it on and not the field whose hint is its
aria-label.

The third resolver reads the same selectors: what TranslateStringSelector
builds for a hint key has to match nothing over CDP rather than the field
carrying the value as a placeholder.

* docs(manual): both hosts read the hint ladder, placeholder is the attribute

The web section said the hintText key does not read the ladder on both
hosts and told authors to select such a field by attrs.hintText instead.
Both hosts read it now, so that instruction is gone rather than left
standing beside a newer sentence.

placeholder is stated as the attribute the markup writes and nothing
more, the tap path is stated as failing by name where no CSS says what
the ladder says, and the alias table gains the row it was missing.
2026-08-19 11:07:49 +05:30

1378 lines
49 KiB
Go

// Package ioscompanion drives an iOS simulator through the native simulator
// companion. This file implements the DeviceDriver surface on top of the
// brand-free transport, supervises the companion child process, and recovers
// from a dropped connection with one in-place restart.
package ioscompanion
import (
"bytes"
"context"
"errors"
"fmt"
"image"
_ "image/png"
"io"
"net"
"os"
"os/exec"
"path/filepath"
"strings"
"sync"
"syscall"
"time"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/companionassets"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/transport"
"github.com/priyanshujain/sanderling/internal/hierarchy"
)
// startupTimeout bounds how long New waits for the spawned companion to accept
// a connection and answer a health probe.
const startupTimeout = 30 * time.Second
// runnerStartupTimeout bounds the in-simulator runner's startup. The runner is
// hosted by a test session whose cold start is far slower than the companion's.
const runnerStartupTimeout = 120 * time.Second
// shutdownGrace bounds how long the companion child gets to exit after SIGTERM
// before it is killed. A variable so the kill-escalation test can shrink it.
var shutdownGrace = 15 * time.Second
// launchTimeout bounds a single app lifecycle RPC. The runner serves lifecycle
// inside its XCTest session, and a launch the simulator rejects sends that
// session down a recovery chain (a 120s accessibility wait, a spindump, then an
// idle wait) that answers minutes late or never. Callers reach Launch with an
// undeadlined context, since it runs before the run's duration clock starts, so
// the bound has to come from here or a wedged session hangs the run with no
// trace, no error, and no end. Kept under runnerStartupTimeout: launching an
// app inside a live session must cost less than cold-starting that session.
// A variable so the timeout test can shrink it.
var launchTimeout = 90 * time.Second
// launchRecoveryTimeout bounds the whole recovery a blown launch bound
// triggers, the session restart and the second attempt together. It keeps the
// launch path inside the three minutes testrun allows it, so what a user sees
// when the app really cannot be launched stays the driver's error rather than
// that backstop firing over the top of it. A variable so the bound test can
// shrink it.
var launchRecoveryTimeout = 60 * time.Second
// longPressHoldMilliseconds is how long LongPress holds the finger down.
const longPressHoldMilliseconds = 600
// Options configures a Driver.
type Options struct {
// UniqueDeviceIdentifier selects the booted simulator the companion drives.
UniqueDeviceIdentifier string
// BundleID is the app under test. Launch and Terminate act on it.
BundleID string
// AppPath is the .app bundle directory. Required for clear-state reinstall;
// when empty, clear state falls back to resetting the data container.
AppPath string
// ClearState resets the app to first-launch state while New runs, before
// any automation session attaches. Clear state is a property of the driver
// rather than of a launch: see Launch.
ClearState bool
// Output receives companion stdout and stderr plus driver warnings.
Output io.Writer
// DoubleTapGapMilliseconds overrides the synthesized double-tap gap.
DoubleTapGapMilliseconds float64
// These are test seams. Production leaves them nil and New wires the real
// extraction, spawn, dial and simctl calls.
spawnChild func(ctx context.Context, address string) (*exec.Cmd, error)
dialCompanion func(address string) (transport.Companion, error)
pickAddress func() (string, error)
spawnRunner func(ctx context.Context, address string) (*exec.Cmd, error)
dialRunner func(address string) (transport.Companion, error)
reinstallApp func(ctx context.Context) error
resetContainer func(ctx context.Context) error
terminateApp func(ctx context.Context) error
}
// Driver implements driver.DeviceDriver against an iOS simulator companion.
type Driver struct {
companion transport.Companion
udid string
bundleID string
appPath string
output io.Writer
// clearedBundleID names the app New (or NewDevice) reset to first-launch
// state before attaching, which is the only point in a run where clearing
// is safe. Launch refuses a clear-state request for anything else rather
// than reinstalling under a live session or reporting a reset that only
// ever reached another bundle.
clearedBundleID string
screenWidth int
screenHeight int
doubleTapGapMilliseconds float64
// mu guards Snapshot's hierarchy+screenshot pairing and the lastTap record.
mu sync.Mutex
lastTap struct {
x, y float64
set bool
}
clearStateWarned bool
// restart rebuilds the transport in place after a connection-level failure.
// It is a seam so tests exercise the supervision logic without spawning a
// real companion. restarting guards against re-entrant restarts.
restart func(ctx context.Context) error
restarting bool
address string
// resetContainer wipes the app data container for the clear-state fallback.
// A seam so tests skip the xcrun shell-out.
resetContainer func(ctx context.Context) error
// reinstallApp uninstalls and reinstalls the app bundle for clear-state.
// A seam so tests skip the simctl shell-outs.
reinstallApp func(ctx context.Context) error
// terminateApp stops the app before its state is cleared. A seam so tests
// skip the simctl shell-out.
terminateApp func(ctx context.Context) error
// grantPaste pre-authorizes the app's pasteboard access. A seam so tests
// skip the sqlite shell-out.
grantPaste func(ctx context.Context) error
// idleClock drives WaitForIdle's settle poll. A seam so tests substitute a
// fake clock and avoid the real settle cap.
idleClock Clock
spawnChild func(ctx context.Context, address string) (*exec.Cmd, error)
dial func(address string) (transport.Companion, error)
child *exec.Cmd
// The hybrid simulator companion pairs the legacy child (HID gestures,
// lifecycle, screenshot) with an in-simulator runner that serves
// collapse-free accessibility snapshots and native unicode typing.
// runnerClient is nil on the legacy-only path.
runnerClient transport.Companion
runnerChild *exec.Cmd
runnerAddress string
spawnRunner func(ctx context.Context, address string) (*exec.Cmd, error)
dialRunner func(address string) (transport.Companion, error)
hybrid bool
// pickRunnerAddress hands every bring-up a free loopback port, on the
// simulator and the device alike. One field, so no path can be wired
// without it.
pickRunnerAddress func() (string, error)
// Device-mode fields. On the physical-device path d.companion is the runner
// dialed over a usbmux tunnel, hybrid is false, and runnerClient is nil.
// coreDeviceID feeds devicectl; tunnel is the in-process usbmux forwarder
// bridging the host loopback port to the runner's device-side port.
deviceMode bool
coreDeviceID string
tunnel io.Closer
startTunnel func(ctx context.Context, hardwareUDID, localAddress, devicePort string) (io.Closer, error)
// processContext owns the companion child's lifetime: it is derived from
// New's context (so a canceled run still reaps the child) and canceled by
// Close. Spawning under a startup-scoped context would SIGTERM the child
// the moment startup finishes.
processContext context.Context
processCancel context.CancelFunc
// deviceLock is the exclusive claim on the target, held for the driver's
// whole life and released by Close.
deviceLock io.Closer
}
// acquireDeviceLock takes an exclusive advisory lock on the target so only one
// run drives it at a time. Two runs on one device interleave app lifecycle: the
// second run's uninstall and reinstall land under the first's live automation
// session, leaving its app proxies bound to a bundle the simulator no longer
// knows, and every later snapshot and launch on that session stalls. Failing
// fast beats recovering silently, since the other run owns the device and would
// be corrupted either way. The lock lives on the file descriptor, so a crashed
// run's claim is released by the kernel and never strands the device.
func acquireDeviceLock(udid string) (io.Closer, error) {
path := filepath.Join(os.TempDir(), "sanderling-ios-"+udid+".lock")
file, err := os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0o644)
if err != nil {
return nil, fmt.Errorf("open device lock %s: %w", path, err)
}
if err := syscall.Flock(int(file.Fd()), syscall.LOCK_EX|syscall.LOCK_NB); err != nil {
file.Close()
return nil, fmt.Errorf(
"ios target %s is already driven by another sanderling run (lock %s); "+
"wait for that run to finish or point this one at a different device with --ios-device",
udid, path)
}
return file, nil
}
// New extracts the embedded companion, spawns it against the configured
// simulator, dials the transport, health-probes it, and caches the screen
// point dimensions. Call Close when done to stop the child.
func New(ctx context.Context, options Options) (*Driver, error) {
if options.UniqueDeviceIdentifier == "" {
return nil, errors.New("ios companion: UniqueDeviceIdentifier is required")
}
if options.ClearState && options.BundleID == "" {
return nil, errors.New("ios companion: clear-state needs BundleID: there is nothing to uninstall or wipe without it")
}
output := options.Output
if output == nil {
output = io.Discard
}
gap := options.DoubleTapGapMilliseconds
if gap <= 0 {
gap = DefaultDoubleTapGapMilliseconds
}
driverInstance := &Driver{
udid: options.UniqueDeviceIdentifier,
bundleID: options.BundleID,
appPath: options.AppPath,
output: output,
doubleTapGapMilliseconds: gap,
spawnChild: options.spawnChild,
dial: options.dialCompanion,
spawnRunner: options.spawnRunner,
dialRunner: options.dialRunner,
reinstallApp: options.reinstallApp,
resetContainer: options.resetContainer,
terminateApp: options.terminateApp,
hybrid: hybridCompanionEnabled(),
}
if driverInstance.spawnChild == nil {
driverInstance.spawnChild = driverInstance.realSpawnChild
}
if driverInstance.dial == nil {
driverInstance.dial = transport.Dial
}
if driverInstance.spawnRunner == nil {
driverInstance.spawnRunner = driverInstance.realSpawnRunner
}
if driverInstance.dialRunner == nil {
driverInstance.dialRunner = func(address string) (transport.Companion, error) {
return transport.DialRunner(address, driverInstance.udid, driverInstance.bundleID)
}
}
pickAddress := options.pickAddress
if pickAddress == nil {
pickAddress = pickLoopbackAddress
}
address, err := pickAddress()
if err != nil {
return nil, err
}
driverInstance.address = address
driverInstance.pickRunnerAddress = pickAddress
driverInstance.restart = driverInstance.respawnAndRedial
if driverInstance.resetContainer == nil {
driverInstance.resetContainer = driverInstance.resetDataContainer
}
if driverInstance.reinstallApp == nil {
driverInstance.reinstallApp = driverInstance.simctlReinstall
}
if driverInstance.terminateApp == nil {
driverInstance.terminateApp = driverInstance.simctlTerminate
}
driverInstance.grantPaste = driverInstance.grantPasteboardAccess
driverInstance.processContext, driverInstance.processCancel = context.WithCancel(ctx)
lock, err := acquireDeviceLock(driverInstance.udid)
if err != nil {
driverInstance.processCancel()
return nil, err
}
driverInstance.deviceLock = lock
if options.ClearState {
if err := driverInstance.clearAppState(ctx); err != nil {
driverInstance.Close()
return nil, err
}
driverInstance.clearedBundleID = options.BundleID
}
if err := driverInstance.bringUp(ctx); err != nil {
driverInstance.Close()
return nil, err
}
if driverInstance.hybrid {
if err := driverInstance.bringUpRunner(ctx); err != nil {
driverInstance.Close()
return nil, fmt.Errorf("simulator runner: %w (set SANDERLING_SIMULATOR_COMPANION=legacy to bypass)", err)
}
}
description, err := driverInstance.companion.Describe(ctx)
if err != nil {
driverInstance.Close()
return nil, fmt.Errorf("describe target: %w", err)
}
driverInstance.screenWidth = description.WidthPoints
driverInstance.screenHeight = description.HeightPoints
return driverInstance, nil
}
// bringUp spawns the companion child, waits for the listener, dials, and
// confirms health. It is used by New and by the in-place restart.
func (d *Driver) bringUp(ctx context.Context) error {
startupCtx, cancel := context.WithTimeout(ctx, startupTimeout)
defer cancel()
child, err := d.spawnChild(d.processContext, d.address)
if err != nil {
return fmt.Errorf("spawn companion: %w", err)
}
d.child = child
if err := waitForListener(startupCtx, d.address); err != nil {
d.stopChild()
return fmt.Errorf("companion listener: %w", err)
}
companion, err := d.dial(d.address)
if err != nil {
d.stopChild()
return fmt.Errorf("dial companion: %w", err)
}
d.companion = companion
if err := d.waitForHealth(startupCtx); err != nil {
_ = companion.Close()
d.stopChild()
return fmt.Errorf("companion health: %w", err)
}
return nil
}
// waitForHealth probes AccessibilityInfo until it succeeds or the context
// expires. A successful describe-all means the companion is attached to the
// simulator and ready to serve.
func (d *Driver) waitForHealth(ctx context.Context) error {
ticker := time.NewTicker(250 * time.Millisecond)
defer ticker.Stop()
for {
if _, err := d.companion.AccessibilityInfo(ctx); err == nil {
return nil
}
select {
case <-ctx.Done():
return ctx.Err()
case <-ticker.C:
}
}
}
// respawnAndRedial tears down the current transports and children, then brings
// fresh ones up at the same addresses. Used as the supervision restart. On the
// hybrid path both halves restart together: their failure modes overlap (a
// rebooted simulator drops both) and one orchestration keeps recovery simple.
func (d *Driver) respawnAndRedial(ctx context.Context) error {
// The dead transports are closed but kept in place until their fresh
// replacements land: if the restart fails, later calls error gracefully on
// the closed transport instead of dereferencing nil, and a later incident
// earns another restart attempt.
if d.companion != nil {
_ = d.companion.Close()
}
d.stopChild()
if d.runnerClient != nil {
_ = d.runnerClient.Close()
}
d.stopRunnerChild()
if err := d.bringUp(ctx); err != nil {
return err
}
if d.hybrid {
return d.bringUpRunner(ctx)
}
return nil
}
// hybridCompanionEnabled reports whether the simulator driver should pair the
// legacy companion with the in-simulator runner. The hybrid is the default;
// SANDERLING_SIMULATOR_COMPANION=legacy forces the legacy companion alone.
func hybridCompanionEnabled() bool {
return os.Getenv("SANDERLING_SIMULATOR_COMPANION") != "legacy"
}
// bringUpRunner spawns the in-simulator runner, waits for its listener, dials,
// and confirms it serves snapshots. Used by New and the in-place restart.
func (d *Driver) bringUpRunner(ctx context.Context) error {
startupCtx, cancel := context.WithTimeout(ctx, runnerStartupTimeout)
defer cancel()
// A fresh port every bring-up: after a restart the dying session's
// listener may still answer on the old port and would satisfy the wait
// below with a dead server.
address, err := d.pickRunnerAddress()
if err != nil {
return err
}
d.runnerAddress = address
child, err := d.spawnRunner(d.processContext, d.runnerAddress)
if err != nil {
return fmt.Errorf("spawn runner: %w", err)
}
d.runnerChild = child
if err := waitForListener(startupCtx, d.runnerAddress); err != nil {
d.stopRunnerChild()
return fmt.Errorf("runner listener: %w", err)
}
client, err := d.dialRunner(d.runnerAddress)
if err != nil {
d.stopRunnerChild()
return fmt.Errorf("dial runner: %w", err)
}
ticker := time.NewTicker(250 * time.Millisecond)
defer ticker.Stop()
for {
if _, healthErr := client.AccessibilityInfo(startupCtx); healthErr == nil {
break
}
select {
case <-startupCtx.Done():
_ = client.Close()
d.stopRunnerChild()
return fmt.Errorf("runner health: %w", startupCtx.Err())
case <-ticker.C:
}
}
d.runnerClient = client
return nil
}
// withRecovery runs call, and on a connection-level failure performs one
// in-place restart before retrying the call once. A non-connection error, or a
// second failure of any kind, surfaces to the caller. The restart budget is per
// failure incident: each healthy call resets restarting to false, so a later
// drop earns its own single restart.
func (d *Driver) withRecovery(ctx context.Context, call func() error) error {
err := call()
if err == nil || !isConnectionError(err) || d.restarting || d.restart == nil {
return err
}
d.restarting = true
defer func() { d.restarting = false }()
fmt.Fprintf(d.output, "companion connection lost (%v); restarting once\n", err)
// The restart runs under the driver's own lifetime context, not the
// failed call's: an action whose deadline already expired must not doom
// the recovery that later actions depend on.
restartCtx := d.processContext
if restartCtx == nil {
restartCtx = ctx
}
if restartErr := d.restart(restartCtx); restartErr != nil {
return fmt.Errorf("companion restart failed: %w (original: %v)", restartErr, err)
}
return call()
}
// isConnectionError reports whether err is a dropped-connection signal that a
// restart can recover from: the transport's unavailable sentinel, a gRPC
// Unavailable status, or an EOF.
func isConnectionError(err error) bool {
if err == nil {
return false
}
if errors.Is(err, io.EOF) || errors.Is(err, transport.ErrCompanionUnavailable) {
return true
}
if statusValue, ok := status.FromError(err); ok {
return statusValue.Code() == codes.Unavailable
}
return false
}
// isBudgetExpiry reports whether err is a call that outlived its bound rather
// than a failure the transport can name. Each transport says so differently:
// the runner wraps the context's error when cancellation has landed and the
// connection's i/o timeout when the deadline it armed from that context fires
// first, and the legacy companion returns a gRPC status. Only an expiry earns
// a session restart; an error the runner reports has already said what a fresh
// session would say.
func isBudgetExpiry(err error) bool {
if err == nil {
return false
}
if errors.Is(err, context.DeadlineExceeded) || errors.Is(err, os.ErrDeadlineExceeded) {
return true
}
if statusValue, ok := status.FromError(err); ok {
return statusValue.Code() == codes.DeadlineExceeded
}
return false
}
func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, env map[string]string) error {
if bundleID != "" {
d.bundleID = bundleID
}
if len(env) > 0 {
// The launch Start message carries an env map, but this backend does
// not pass it through: passing it would change the app's process
// environment in ways the rest of the run does not account for. Reject
// loudly rather than silently dropping the request.
return errors.New("ios companion: launch with environment variables is unsupported on this backend")
}
if clearState && (d.clearedBundleID == "" || d.clearedBundleID != d.bundleID) {
// Clearing here would uninstall and reinstall the app underneath a live
// automation session, which is what races FrontBoard's registration and
// leaves the session launching a bundle FrontBoard has not registered.
return fmt.Errorf("ios companion: clear-state must be requested when the driver is created (Options.ClearState) "+
"for the bundle being launched; this backend cleared %q before its automation session existed, not %q",
d.clearedBundleID, d.bundleID)
}
// Terminate first so the launch is a clean cold start regardless of the
// app's prior state. A not-running app is not an error here.
_ = d.lifecycleCall(ctx, func(callCtx context.Context, companion transport.Companion) error {
return companion.Terminate(callCtx, d.bundleID)
})
// Grant the app pasteboard access before it runs so unicode input (which
// must go through the pasteboard, since HID cannot express it) never trips
// the iOS paste-permission prompt. clearState reinstall resets the grant,
// so it is reapplied on every launch. Best effort: if it fails, the paste
// path still handles the prompt, just slower. The hybrid path types
// natively and never touches the pasteboard, so it skips the grant.
if d.bundleID != "" && d.runnerTyper() == nil {
if err := d.grantPaste(ctx); err != nil {
fmt.Fprintf(d.output, "grant pasteboard access failed (continuing): %v\n", err)
}
}
if err := d.launchWithSessionRecovery(ctx); err != nil {
return fmt.Errorf("launch %s: %w", d.bundleID, err)
}
return nil
}
// launchWithSessionRecovery runs the launch RPC and, when it blows its own
// bound, replaces the session and launches again.
//
// A launch the simulator refuses, which is what a clear-state reinstall racing
// FrontBoard's registration produces, never comes back as an error: XCTest
// records the refusal as a test failure the runner cannot observe, then holds
// the session's main thread for about four minutes walking a diagnostic chain
// (a 120s accessibility wait, a spindump, an idle wait). So there is no error
// text to key a retry on, only the expired bound, and every later call queues
// behind the same wedge. Only a session that never served the refused launch
// can serve the retry, which is why this restarts rather than calls again.
func (d *Driver) launchWithSessionRecovery(ctx context.Context) error {
launch := func(callCtx context.Context, companion transport.Companion) error {
return companion.Launch(callCtx, d.bundleID, true)
}
err := d.lifecycleCall(ctx, launch)
// A caller whose own budget ran out gets no restart: the bound that expired
// was the caller's to spend, and the second attempt would inherit it dead.
if !isBudgetExpiry(err) || ctx.Err() != nil || d.restart == nil {
return err
}
fmt.Fprintf(d.output, "launch %s blew its %v bound (%v); restarting the session and launching once more\n",
d.bundleID, launchTimeout, err)
// The restart runs under the driver's own lifetime context for the same
// reason withRecovery's does, while the second attempt stays on the
// caller's. Both end at one deadline, so a launch that already spent
// launchTimeout cannot then wait out a session cold start on top of it.
recoveryDeadline := time.Now().Add(launchRecoveryTimeout)
restartCtx := d.processContext
if restartCtx == nil {
restartCtx = ctx
}
restartCtx, cancelRestart := context.WithDeadline(restartCtx, recoveryDeadline)
defer cancelRestart()
if restartErr := d.restart(restartCtx); restartErr != nil {
return fmt.Errorf("session restart failed: %w (original: %v)", restartErr, err)
}
relaunchCtx, cancelRelaunch := context.WithDeadline(ctx, recoveryDeadline)
defer cancelRelaunch()
return d.lifecycleCall(relaunchCtx, launch)
}
// lifecycleCall runs an app lifecycle RPC against lifecycleCompanion under a
// launchTimeout-bounded context, with the usual one-restart recovery. The
// companion is resolved inside the retry so a restart's replacement client
// serves the second attempt.
func (d *Driver) lifecycleCall(ctx context.Context, call func(context.Context, transport.Companion) error) error {
boundedCtx, cancel := context.WithTimeout(ctx, launchTimeout)
defer cancel()
return d.withRecovery(boundedCtx, func() error {
return call(boundedCtx, d.lifecycleCompanion())
})
}
// lifecycleCompanion is the transport that owns app launch and terminate: the
// in-simulator runner when the hybrid is active, otherwise the legacy
// companion. Lifecycle performed outside the runner's automation session
// leaves the session's app proxies bound to dead processes, after which
// snapshots hang and typing asserts.
func (d *Driver) lifecycleCompanion() transport.Companion {
if d.runnerClient != nil {
return d.runnerClient
}
return d.companion
}
// clearAppState resets the app to a first-launch state. With an app path it
// uninstalls and reinstalls; without one it falls back to wiping the app's data
// container and warns once that a full reinstall needs the app path. Called
// only from construction, before any automation session is attached to the app.
func (d *Driver) clearAppState(ctx context.Context) error {
// Nothing may be writing to the state while it goes, which is the ordering
// Launch used to hold: uninstall copes with a running app, deleting the
// data container out from under one does not. Best effort, because an app
// that is not running reports a failure that means nothing here.
_ = d.terminateApp(ctx)
if d.appPath != "" {
if err := d.reinstallApp(ctx); err != nil {
return fmt.Errorf("reinstall %s: %w", d.appPath, err)
}
return nil
}
if !d.clearStateWarned {
fmt.Fprintln(d.output, "clear-state requested without an app path: resetting the data container only; pass the app path for a full reinstall")
d.clearStateWarned = true
}
return d.resetContainer(ctx)
}
// simctlReinstall uninstalls and reinstalls the app bundle via simctl. App
// lifecycle stays with simctl: the companion's install RPC misreads current
// simulator targets' architectures and rejects valid bundles.
// A failed uninstall ends the reinstall: `simctl install` over an installed app
// carries its data container across, so clear-state would be reported without
// happening. Uninstalling an app that is not installed exits 0, so there is no
// benign failure here to sort out from a real one.
func (d *Driver) simctlReinstall(ctx context.Context) error {
if output, err := exec.CommandContext(ctx, "xcrun", "simctl", "uninstall", d.udid, d.bundleID).CombinedOutput(); err != nil {
return fmt.Errorf("simctl uninstall %s: %w: %s", d.bundleID, err, strings.TrimSpace(string(output)))
}
output, err := exec.CommandContext(ctx, "xcrun", "simctl", "install", d.udid, d.appPath).CombinedOutput()
if err != nil {
return fmt.Errorf("simctl install: %w: %s", err, strings.TrimSpace(string(output)))
}
return nil
}
// simctlTerminate stops the app under test. Launch used to terminate through
// the automation session before clearing; the clear now runs before any session
// exists, so simctl is what is left to stop the app with. An app that is not
// running reports a failure that means nothing to the caller, which is why
// clearAppState treats this as best effort.
func (d *Driver) simctlTerminate(ctx context.Context) error {
if output, err := exec.CommandContext(ctx, "xcrun", "simctl", "terminate", d.udid, d.bundleID).CombinedOutput(); err != nil {
return fmt.Errorf("simctl terminate %s: %w: %s", d.bundleID, err, strings.TrimSpace(string(output)))
}
return nil
}
// grantPasteboardAccess authorizes the app to read the pasteboard without the
// iOS permission prompt, by writing an allow row into the simulator's privacy
// (TCC) database. This is the simulator counterpart to `simctl privacy grant`,
// which does not expose the pasteboard service. Without it, every unicode input
// (which must paste, since HID cannot express unicode) blocks on a modal that
// costs seconds; with it the paste lands in one frame.
func (d *Driver) grantPasteboardAccess(ctx context.Context) error {
databasePath := filepath.Join(
os.Getenv("HOME"), "Library", "Developer", "CoreSimulator", "Devices",
d.udid, "data", "Library", "TCC", "TCC.db",
)
if _, err := os.Stat(databasePath); err != nil {
return fmt.Errorf("locate privacy database: %w", err)
}
statement := fmt.Sprintf(
"INSERT OR REPLACE INTO access "+
"(service,client,client_type,auth_value,auth_reason,auth_version,indirect_object_identifier) "+
"VALUES ('kTCCServicePasteboard','%s',0,2,4,1,'UNUSED');",
d.bundleID,
)
output, err := exec.CommandContext(ctx, "sqlite3", databasePath, statement).CombinedOutput()
if err != nil {
return fmt.Errorf("write privacy grant: %w: %s", err, strings.TrimSpace(string(output)))
}
return nil
}
// resetDataContainer deletes the contents of the app's data container so the
// next launch starts with empty storage.
func (d *Driver) resetDataContainer(ctx context.Context) error {
output, err := exec.CommandContext(ctx, "xcrun", "simctl", "get_app_container", d.udid, d.bundleID, "data").Output()
if err != nil {
return fmt.Errorf("get app container: %w", err)
}
container := string(bytes.TrimSpace(output))
if container == "" {
return nil
}
entries, err := os.ReadDir(container)
if err != nil {
return fmt.Errorf("read app container: %w", err)
}
for _, entry := range entries {
if err := os.RemoveAll(filepath.Join(container, entry.Name())); err != nil {
return fmt.Errorf("clear app container: %w", err)
}
}
return nil
}
func (d *Driver) Terminate(ctx context.Context) error {
return d.lifecycleCall(ctx, func(callCtx context.Context, companion transport.Companion) error {
return companion.Terminate(callCtx, d.bundleID)
})
}
// offScreen reports a point the device has no surface under. The hierarchy
// reaches past the screen wherever a scroll container holds content beyond the
// fold, so an action derived from it can name a point no touch can land on.
// The far edge is exclusive: a touch at x == screenWidth arrives at
// screenWidth-1, which is a point the action never named.
func (d *Driver) offScreen(x, y int) bool {
return x < 0 || y < 0 || x >= d.screenWidth || y >= d.screenHeight
}
func (d *Driver) Tap(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
d.mu.Lock()
d.lastTap.x = float64(x)
d.lastTap.y = float64(y)
d.lastTap.set = true
d.mu.Unlock()
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, tapEvents(float64(x), float64(y))...)
})
}
func (d *Driver) DoubleTap(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, doubleTapEvents(float64(x), float64(y), d.doubleTapGapMilliseconds)...)
})
}
func (d *Driver) LongPress(ctx context.Context, x, y int) error {
if d.offScreen(x, y) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, x, y)
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, longPressEvents(float64(x), float64(y), longPressHoldMilliseconds)...)
})
}
func (d *Driver) Swipe(ctx context.Context, fromX, fromY, toX, toY int, duration time.Duration) error {
if d.offScreen(fromX, fromY) {
return fmt.Errorf("%w: (%d,%d)", driver.ErrGestureUndelivered, fromX, fromY)
}
seconds := duration.Seconds()
if seconds <= 0 {
seconds = 0.25
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, transport.SwipeEvent(
float64(fromX), float64(fromY), float64(toX), float64(toY), seconds))
})
}
func (d *Driver) PressKey(ctx context.Context, key string) error {
if d.textEditor() != nil {
return d.withRecovery(ctx, func() error {
return d.textEditor().PressKey(ctx, key)
})
}
usage, ok := pressKeyUsage(key)
if !ok {
return fmt.Errorf("ios companion: unsupported key %q", key)
}
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, transport.KeyDown(usage), transport.KeyUp(usage))
})
}
// textEditor returns the companion's native text-editing capability, or nil
// when the transport does not implement it. Resolved per call because a
// restart replaces d.companion.
func (d *Driver) textEditor() transport.TextEditor {
if editor, ok := d.companion.(transport.TextEditor); ok {
return editor
}
return nil
}
// runnerTyper returns the in-simulator runner's native typing capability, or
// nil outside the hybrid path. Resolved per call because a restart replaces
// d.runnerClient.
func (d *Driver) runnerTyper() transport.TextTyper {
if typer, ok := d.runnerClient.(transport.TextTyper); ok {
return typer
}
return nil
}
// pressKeyUsage maps the logical key names mobile runs emit to a HID usage.
// Return/Enter and Escape are the hardware-keyboard keys the simulator has;
// other names (notably "back" and "home") have no HID key and report
// unsupported.
func pressKeyUsage(key string) (uint32, bool) {
switch key {
case "enter", "return", "Enter", "Return":
return usageReturn, true
case "escape", "Escape":
return usageEscape, true
default:
return 0, false
}
}
func (d *Driver) TapSelector(ctx context.Context, selector string) error {
x, y, err := d.resolveSelectorCenter(ctx, selector)
if err != nil {
return err
}
return d.Tap(ctx, x, y)
}
func (d *Driver) DoubleTapSelector(ctx context.Context, selector string) error {
x, y, err := d.resolveSelectorCenter(ctx, selector)
if err != nil {
return err
}
return d.DoubleTap(ctx, x, y)
}
// resolveSelectorCenter fetches a fresh hierarchy and returns the center of the
// first element matching selector.
func (d *Driver) resolveSelectorCenter(ctx context.Context, selector string) (int, int, error) {
hierarchyJSON, err := d.Hierarchy(ctx)
if err != nil {
return 0, 0, err
}
tree, err := hierarchy.Parse(hierarchyJSON)
if err != nil {
return 0, 0, fmt.Errorf("parse hierarchy: %w", err)
}
element := tree.Find(selector)
if element == nil {
return 0, 0, fmt.Errorf("%w: %q", driver.ErrSelectorMatchedNothing, selector)
}
x, y := element.Bounds.Center()
return x, y, nil
}
func (d *Driver) InputText(ctx context.Context, text string) error {
// Hybrid path. Mappable text rides one HID stream: select-all chord plus
// keystrokes, atomic and strictly ordered on a single channel. Unicode
// (which HID cannot express) is typed natively by the runner after the
// chord; chord and typing ride different channels with no ordering
// guarantee between them, so the clear is verified through a snapshot
// before the first keystroke goes out.
if typer := d.runnerTyper(); typer != nil {
if !usesPasteboard(text) {
events := append(clearFieldEvents(), keyPressEvents(typeStringPresses(text))...)
return d.withRecovery(ctx, func() error {
return d.companion.SendHID(ctx, events...)
})
}
return d.withRecovery(ctx, func() error {
if err := d.companion.SendHID(ctx, clearFieldEvents()...); err != nil {
return fmt.Errorf("clear field: %w", err)
}
d.waitFieldCleared(ctx)
return d.runnerTyper().TypeText(ctx, text, false)
})
}
// A text-editing companion replaces the field's content natively, which
// covers unicode without the pasteboard and its permission dialog.
if d.textEditor() != nil {
return d.withRecovery(ctx, func() error {
return d.textEditor().InputText(ctx, text)
})
}
// The field target is only needed for the pasteboard path. Resolving it
// requires a describe-all, so the fast keyboard path skips that round-trip
// and lets inputText send the key presses directly.
var field fieldTarget
if usesPasteboard(text) {
field = d.resolveInputField(ctx)
}
return inputText(ctx, d.makeRunner(), text, field)
}
// fieldClearedWaitCap and fieldClearedPoll bound the verify-cleared loop
// between the HID clear chord and the runner's native typing.
const fieldClearedWaitCap = 1200 * time.Millisecond
const fieldClearedPoll = 150 * time.Millisecond
// waitFieldCleared polls the focused field (the editable element under the
// last tap) until its value reads empty, so the clear chord has demonstrably
// landed before typing starts on the other channel. Best effort: when the
// field cannot be resolved or the cap elapses, typing proceeds anyway.
func (d *Driver) waitFieldCleared(ctx context.Context) {
d.mu.Lock()
tap := d.lastTap
d.mu.Unlock()
if !tap.set {
return
}
deadline := time.Now().Add(fieldClearedWaitCap)
for time.Now().Before(deadline) {
dump, err := d.describeAllRaw(ctx)
if err != nil {
return
}
cleared := true
for _, element := range decodeDump(dump) {
if !isEditable(element.Type) {
continue
}
frame := element.Frame
if !finite(frame.X) || !finite(frame.Y) || !finite(frame.Width) || !finite(frame.Height) {
continue
}
if tap.x < frame.X || tap.x > frame.X+frame.Width ||
tap.y < frame.Y || tap.y > frame.Y+frame.Height {
continue
}
if value := stringValue(element.AXValue); value != "" && value != emptyFieldValueSentinel {
cleared = false
}
break
}
if cleared {
return
}
select {
case <-ctx.Done():
return
case <-time.After(fieldClearedPoll):
}
}
}
// resolveInputField finds the editable element under the last tap so the
// pasteboard fallback can confirm the paste landed and refocus after dismissing
// the permission dialog. The runner always taps a field before typing, so
// lastTap names the focus point. An empty fieldTarget is returned when no
// editable element contains the tap (the fast keyboard path ignores it).
func (d *Driver) resolveInputField(ctx context.Context) fieldTarget {
d.mu.Lock()
tap := d.lastTap
d.mu.Unlock()
if !tap.set {
return fieldTarget{}
}
dump, err := d.describeAll(ctx)
if err != nil {
return fieldTarget{}
}
for _, element := range decodeDump(dump) {
if !isEditable(element.Type) {
continue
}
frame := element.Frame
if !finite(frame.X) || !finite(frame.Y) || !finite(frame.Width) || !finite(frame.Height) {
continue
}
if tap.x < frame.X || tap.x > frame.X+frame.Width ||
tap.y < frame.Y || tap.y > frame.Y+frame.Height {
continue
}
return fieldTarget{
identifier: stringValue(element.AXUniqueID),
centerX: frame.X + frame.Width/2,
centerY: frame.Y + frame.Height/2,
}
}
return fieldTarget{}
}
func (d *Driver) EraseText(ctx context.Context, characterCount int) error {
if d.textEditor() != nil {
return d.withRecovery(ctx, func() error {
return d.textEditor().EraseText(ctx, characterCount)
})
}
return eraseText(ctx, d.makeRunner(), characterCount)
}
func (d *Driver) Hierarchy(ctx context.Context) (string, error) {
dump, err := d.describeAll(ctx)
if err != nil {
return "", err
}
mapped, err := MapHierarchy(dump, d.screenWidth, d.screenHeight)
if err != nil {
return "", err
}
return string(mapped), nil
}
func (d *Driver) Screenshot(ctx context.Context) (driver.Image, error) {
var data []byte
err := d.withRecovery(ctx, func() error {
var screenshotErr error
data, _, screenshotErr = d.companion.Screenshot(ctx)
return screenshotErr
})
if err != nil {
return driver.Image{}, fmt.Errorf("screenshot: %w", err)
}
return decodeScreenshot(data)
}
func (d *Driver) Snapshot(ctx context.Context) (string, driver.Image, error) {
d.mu.Lock()
defer d.mu.Unlock()
// The hierarchy and the screenshot ride different transports on the
// hybrid path, so they are captured concurrently. Only the hierarchy leg
// runs under withRecovery: two concurrent recoveries would race the
// restart bookkeeping, and a screenshot connection failure surfaces as a
// plain error that the next serialized call recovers from. The goroutine
// works through a captured local because a hierarchy-leg recovery
// reassigns d.companion mid-flight; a screenshot against the torn-down
// transport then fails as a plain error rather than racing the field.
var data []byte
screenshotDone := make(chan error, 1)
companion := d.companion
go func() {
imageData, _, callErr := companion.Screenshot(ctx)
data = imageData
screenshotDone <- callErr
}()
dump, err := d.describeAll(ctx)
screenshotErr := <-screenshotDone
if err != nil {
return "", driver.Image{}, err
}
if screenshotErr != nil {
return "", driver.Image{}, fmt.Errorf("screenshot: %w", screenshotErr)
}
mapped, err := MapHierarchy(dump, d.screenWidth, d.screenHeight)
if err != nil {
return "", driver.Image{}, err
}
image, err := decodeScreenshot(data)
if err != nil {
return string(mapped), driver.Image{}, err
}
return string(mapped), image, nil
}
// WaitForIdle polls the hierarchy until it settles. The duration argument is
// ignored: the ported settle constants (StabilityPollCap and friends) own the
// cap, matching the companion's own settle behavior.
func (d *Driver) WaitForIdle(ctx context.Context, _ time.Duration) error {
clock := d.idleClock
if clock == nil {
clock = SystemClock()
}
PollUntilStable(ctx, clock, func() *hierarchy.Tree {
dump, err := d.describeAllRaw(ctx)
if err != nil || dumpIsCollapsed(dump) {
// A collapsed dump is the bridge mid-transition; report it
// transitional so the streak resets and the poll waits for the
// real tree rather than settling on the empty shell.
return nil
}
mapped, err := MapHierarchy(dump, d.screenWidth, d.screenHeight)
if err != nil {
return nil
}
tree, err := hierarchy.Parse(string(mapped))
if err != nil {
return nil
}
return tree
})
return nil
}
// RecentLogs returns no entries: the companion log RPC is a follow-up, so v1
// reports an empty slice rather than failing. Every property reading state.logs
// therefore holds vacuously on iOS and nothing says so; closing it means
// tailing idb's streaming log RPC and mapping os_log levels onto the
// single-letter scale driver.LogEntry declares.
func (d *Driver) RecentLogs(_ context.Context, _ time.Time, _ string) ([]driver.LogEntry, error) {
return []driver.LogEntry{}, nil
}
func (d *Driver) Metrics(_ context.Context, _ string) (driver.Metrics, error) {
return driver.Metrics{}, nil
}
func (d *Driver) Health(_ context.Context) (driver.Health, error) {
return driver.Health{Ready: true, Platform: "ios"}, nil
}
// ForegroundApp reports the foreground app. It returns the app under test when
// it is running; otherwise it names another running user app, or "" when none
// is. The companion exposes process state but not a foreground flag, so "the
// app under test is running" stands in for "in the foreground".
func (d *Driver) ForegroundApp(ctx context.Context) (string, error) {
var apps []transport.InstalledApp
if err := d.withRecovery(ctx, func() error {
var listErr error
apps, listErr = d.companion.ListApps(ctx)
return listErr
}); err != nil {
return "", err
}
other := ""
for _, app := range apps {
if app.ProcessState != transport.ProcessStateRunning {
continue
}
if app.BundleID == d.bundleID {
return d.bundleID, nil
}
if app.InstallType == "user" && other == "" {
other = app.BundleID
}
}
return other, nil
}
// collapsedDumpRetries and collapsedDumpDelay bound how long describeAll waits
// out a collapsed accessibility dump. The bridge briefly reports only the app
// shell (no UI content) during cold start and screen transitions; it recovers
// within a few hundred milliseconds. Re-fetching past the collapse keeps the
// runner from acting on, and snapshotting, an empty tree.
const collapsedDumpRetries = 6
const collapsedDumpDelay = 150 * time.Millisecond
// snapshotCompanion is the transport that serves accessibility dumps: the
// in-simulator runner when the hybrid is active (its snapshots never collapse),
// otherwise the legacy companion.
func (d *Driver) snapshotCompanion() transport.Companion {
if d.runnerClient != nil {
return d.runnerClient
}
return d.companion
}
// describeAllRaw fetches the flat accessibility dump with one-restart recovery
// and no collapse handling. The settle loop uses it: it treats a collapsed dump
// as transitional itself, so an inner retry here would double the wait.
func (d *Driver) describeAllRaw(ctx context.Context) ([]byte, error) {
var dump []byte
err := d.withRecovery(ctx, func() error {
info, infoErr := d.snapshotCompanion().AccessibilityInfo(ctx)
if infoErr != nil {
return infoErr
}
dump = []byte(info)
return nil
})
return dump, err
}
// describeAll fetches the flat accessibility dump, retrying past a transient
// collapsed dump so one-shot reads (Snapshot, Hierarchy) see real UI content.
func (d *Driver) describeAll(ctx context.Context) ([]byte, error) {
dump, err := d.describeAllRaw(ctx)
if err != nil {
return dump, err
}
for attempt := 0; attempt < collapsedDumpRetries && dumpIsCollapsed(dump); attempt++ {
select {
case <-ctx.Done():
return dump, nil
case <-time.After(collapsedDumpDelay):
}
next, nextErr := d.describeAllRaw(ctx)
if nextErr != nil {
return dump, nil
}
dump = next
}
return dump, nil
}
// makeRunner builds the input runner backed by the current transport. The text
// runner does not route through withRecovery: it is invoked synchronously
// inside a single InputText call and a mid-paste connection drop surfaces as a
// normal error the runner retries.
func (d *Driver) makeRunner() runner {
return simctlRunner{companion: d.companion, udid: d.udid}
}
// Close stops the companion and runner children and releases the transports.
func (d *Driver) Close() {
if d.companion != nil {
_ = d.companion.Close()
d.companion = nil
}
d.stopChild()
if d.runnerClient != nil {
_ = d.runnerClient.Close()
d.runnerClient = nil
}
d.stopRunnerChild()
d.stopTunnel()
if d.processCancel != nil {
d.processCancel()
}
if d.deviceLock != nil {
_ = d.deviceLock.Close()
d.deviceLock = nil
}
}
// stopTunnel closes the in-process usbmux forwarder on the device path. Closing
// its listener ends the accept loop and lets the open bridges drain; a nil
// tunnel (the simulator path) is a no-op.
func (d *Driver) stopTunnel() {
tunnel := d.tunnel
d.tunnel = nil
if tunnel != nil {
_ = tunnel.Close()
}
}
// stopChild terminates the companion child gracefully (SIGTERM, grace window,
// then SIGKILL) so it leaves no orphan behind.
func (d *Driver) stopChild() {
child := d.child
d.child = nil
stopProcess(child)
}
// stopRunnerChild terminates the runner's hosting session the same way. The
// session tears down its in-simulator children on SIGTERM; killing it outright
// would orphan them.
func (d *Driver) stopRunnerChild() {
child := d.runnerChild
d.runnerChild = nil
stopProcess(child)
}
func stopProcess(child *exec.Cmd) {
if child == nil || child.Process == nil {
return
}
if err := child.Process.Signal(syscall.SIGTERM); err != nil {
_ = child.Process.Kill()
_ = child.Wait()
return
}
done := make(chan struct{})
go func() {
_ = child.Wait()
close(done)
}()
select {
case <-done:
case <-time.After(shutdownGrace):
_ = child.Process.Kill()
<-done
}
}
// decodeScreenshot sniffs the PNG magic and decodes the pixel dimensions. The
// companion leaves image_format empty in practice, so the magic bytes are the
// only reliable format signal. No scaling is applied: the dimensions are pixels.
func decodeScreenshot(data []byte) (driver.Image, error) {
if len(data) < 8 || !bytes.HasPrefix(data, []byte("\x89PNG\r\n\x1a\n")) {
return driver.Image{}, errors.New("screenshot: response is not a PNG")
}
config, _, err := image.DecodeConfig(bytes.NewReader(data))
if err != nil {
return driver.Image{}, fmt.Errorf("decode screenshot: %w", err)
}
return driver.Image{PNG: data, Width: config.Width, Height: config.Height}, nil
}
// realSpawnChild extracts the embedded companion and starts it on the given
// address. Cancel sends SIGTERM so the companion detaches cleanly from the
// simulator; WaitDelay bounds the grace before the runtime kills it.
func (d *Driver) realSpawnChild(ctx context.Context, address string) (*exec.Cmd, error) {
extractDirectory := filepath.Join(os.TempDir(), "sanderling-companion")
binaryPath, err := companionassets.Extract(extractDirectory)
if err != nil {
return nil, fmt.Errorf("extract companion: %w", err)
}
_, port, err := net.SplitHostPort(address)
if err != nil {
return nil, err
}
command := exec.CommandContext(ctx, binaryPath, "--udid", d.udid, "--grpc-port", port)
command.Stdout = d.output
command.Stderr = d.output
// The companion echoes its whole environment into the run log at startup,
// so it gets a minimal one: secrets in the parent environment must never
// reach run artifacts.
command.Env = []string{
"HOME=" + os.Getenv("HOME"),
"PATH=/usr/bin:/bin",
"TMPDIR=" + os.TempDir(),
}
command.Cancel = func() error { return command.Process.Signal(syscall.SIGTERM) }
command.WaitDelay = shutdownGrace
if err := command.Start(); err != nil {
return nil, fmt.Errorf("start companion: %w", err)
}
fmt.Fprintf(d.output, "companion pid=%d listening on %s\n", command.Process.Pid, address)
return command, nil
}
// pickLoopbackAddress reserves a free loopback port and returns its address.
func pickLoopbackAddress() (string, error) {
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
return "", err
}
defer listener.Close()
return listener.Addr().String(), nil
}
// waitForListener blocks until address accepts a TCP connection or ctx expires.
func waitForListener(ctx context.Context, address string) error {
ticker := time.NewTicker(100 * time.Millisecond)
defer ticker.Stop()
for {
dialer := net.Dialer{Timeout: time.Second}
conn, err := dialer.DialContext(ctx, "tcp", address)
if err == nil {
_ = conn.Close()
return nil
}
select {
case <-ctx.Done():
return ctx.Err()
case <-ticker.C:
}
}
}
// ReplacesTextOnInput reports that InputText replaces the field's content, so
// the runner skips its pre-erase. The driver clears the field inside InputText,
// which is robust even when a collapsed accessibility bridge would make the
// runner read the field length as zero and wrongly skip erasing.
func (d *Driver) ReplacesTextOnInput() bool { return true }
var (
_ driver.DeviceDriver = (*Driver)(nil)
_ driver.ForegroundChecker = (*Driver)(nil)
_ driver.TextReplacer = (*Driver)(nil)
)