OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.
a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.
the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.
the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.
escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.
ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.
delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.
An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.
Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.
A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.
The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.
Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.
The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.
Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.
Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Brings in #73, which landed web fact-parity work overlapping this branch:
selector-tagged ax handles, aria-disabled in `enabled`, shadow-DOM traversal
and a `scrollable`/`editable` dump, plus a third folio property and two new
testTags on HomeScreen.
Six files conflicted. The option-carrying ones took the union of both sides'
fields, so `--label-source` and `--exit-on-violation` both reach the pipeline.
In web-runtime.ts both sides changed how an element is described: master gave
`elementHandle` a selector to tag handles with, this branch gave it raw
attribute names and a field hint. Both survive, and `enabled` now answers
through master's isEnabled while `editable` stays.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(runner): stop the step loop at the first violation on request
* feat(testrun): report violations as a typed error under exit-on-violation
* feat(cli): add --exit-on-violation and exit 2 when it fires
* docs(cli): document --exit-on-violation, --max-steps, and exit codes
* fix(web): enumerate and query across shadow roots in both producers
* test(chrome): compare both producers on a shadow-dom parity page
* test(browser): drive a canvas-under-shadow-root fixture end to end
* fix(web): select the focused field inside a shadow root before typing
* fix(web): report the pathname as the screen when there is no hash route
* fix(web): settle on dom quiescence instead of returning at body ready
* feat(replay-ui): add data-testid hooks the dogfood spec drives
* feat(replay-ui): add the dogfood spec sanderling runs against the replay ui
* fix(replay-ui): scope the screenshot property to the named state panel
* chore(make): add per-platform sanderling build targets
* ci: add dispatch workflows for folio and the replay ui
* docs: describe the dispatch workflows and how to read a failure
* ci(folio): give the ios leg its jdk, android sdk, just, and a clean app start
* refactor(web): use max for the settle budget
* ci: pin calibrated seeds, skip the flaky ios reinstall, bound every job
* ci: authenticate and pin the buf setup step
the anonymous release download hit the shared runner ip rate limit and
failed the job with 'socket hang up' after three retries.
* docs: record that canvas apps need a dom proxy to be text-fuzzable
* fix(ios): bound lifecycle rpcs and claim the target device
a launch the simulator rejects sent the xctest session into a recovery
chain that answered minutes late or never, and the rpc had no deadline,
so the run hung with no trace and no error. also take a per-udid flock:
a second run's reinstall lands under the first's live automation session
and wedges it.
* docs(ci): correct the ios hang wording and note the device lock
* refactor(verifier): derive the lastAction shape from one field list
both hosts must show a spec the same lastAction. one ordered list now
feeds the goja object and the json the web host installs, so they
cannot drift.
* fix(web): install lastAction in the page before extractors read it
state.lastAction was hardcoded null on web, so every property reading it
was silently vacuous: a correct property passed without ever firing.
* fix(web): carry element identity on actions and fix findAll on paths
an action's target was coordinates only, so a property matching on which
element was acted upon could never fire. ax.findAll([a,b]) also returned
nothing on web.
* fix(chrome): wait out a route transition before sampling facts
the tree stays byte-identical and quiet across a cross-fade, so both the
quiet timer and the unchanged-tree escape called it settled mid-flight
and extractors read two screens at once.
* fix: bound the pre-run app launch
launch happens before the runner starts, so --duration never covered it
and a wedged driver hung with no trace and no error.
* fix(folio): read balances from merged cards and treat unreadable as unknown
compose for web merges the whole accountcard subtree, so the balance
child never exists there and every card parsed as 0. the property then
compared 0 to 0 and fired on any submit, which is a false positive
generator. unknown is now null and null is vacuously true.
* test(folio): cover merged-card parsing and unknown balances
* ci(folio): make web an expect-the-bug leg
the web runtime can observe the double submit now, so the health gate
understates it. seed 1 finds it at step 109, 3 runs out of 3.
* docs(ci): explain why a submit tap landing on home is the bug
* fix(ios): read a StaticText's label as its text
AXValue was the only source for text, but a StaticText carries its
string in AXLabel, so nothing on screen had .text on ios: a spec reading
it saw everything on android and nothing here.
* docs(ci): correct the calibrated step ranges
* fix(folio): stop convicting on arithmetic float64 cannot hold
past 2^53 cents the gap between representable values is 128, so a real
1600-cent move reads back as something else and the equality is false
for a healthy submit as readily as a double one. also match parseCents:
a sign or an oversized amount is rejected, not read as an amount.
* test(folio): pin the safe-integer guard and its boundary
* docs: stop teaching the zero-default that caused a false alarm
* docs: write down the silent-vacuity failure modes
* feat(folio): tag the home total and the card transaction count
the total was the only untagged node on the screen, so the spec had to
sum cards and a clipped card broke the sum.
* fix(folio): read the app's own total and refuse contaminated windows
summing cards went null when one was clipped, and the null poisoned the
carrier for the rest of the run. the balance window also spanned every
transaction since the last home visit, so the property convicted on
deltas it could not attribute: the old web witness was 3.16x the typed
amount, not 2x.
* test(folio): pin the window rules and the count invariant
* fix(folio): never read a frame that shows two screens
android dumps a cross-fade with both screens in the tree. the route said
add-transaction while an unscoped find said home, so the oracle took a
half-rendered total as fresh and convicted on a tap that committed
nothing. one function now decides the route and returns null when the
frame is ambiguous.
* test(folio): cover transition frames, card readings and creation
* fix(folio): only disambiguate counts that came from merged text
the equal-length digit rule exists because web merges the card and an
account named -1 makes '12' ambiguous. a dedicated count node has
nothing to disambiguate, so applying it there threw away real evidence.
* ci(folio): pin the recalibrated seeds and drop android to a health gate
web 3 and ios 7 convict 3 runs out of 3 with an exactly 2x witness.
android convicts 2 in 5 because the same seed does not walk the same
trajectory there, so it proves the app runs instead.
* docs(ci): describe the two properties and why android cannot convict
* fix(android): wait out a route cross-fade before snapshotting
the dump could hold two screens at once, and the runner refuses to act
on such a tree, so a quarter of android steps applied no action and the
count varied per run: the same seed never walked the same trajectory.
the ios companion and the chrome driver already do this.
* ci(folio): let the android leg run far enough to see its conviction
* docs: only the repo owner merges
* ci(folio): a thrown predicate is not a conviction
exit 2 means the run recorded a violation, and a predicate that throws
is recorded as one too. so was newAccountBalanceIsZero, an unrelated
property in the same spec. the gate read the exit code and went green
with detection dead.
* ci: install idb-companion from its tap and stop interpolating inputs
idb-companion is not in homebrew-core, so the ios leg died before it
built anything. replay-ui expanded dispatch inputs into the shell.
* docs: correct the snippets and numbers that drifted from the code
* test(sidecar): pin that a slow read counts toward the stability streak
* fix(web): read the page's extractors only on steps that count
the page advances the spec's carriers when it evaluates, but the runner
applied the result only on non-transitional steps. a discarded step
moved the window forward anyway, so the next accepted pair bracketed two
transactions while counting one submit, and convicted a healthy app.
extractor errors now fail the run instead of leaving goja's values in
current against v8's in previous.
* fix(chrome): anchor the transition deadline when the dom goes quiet
it was anchored at script start, so a page that churned past the window
reached the check already expired and returned mid cross-fade. the
driver now publishes the idle timeout it needs, since the caller's 1s
could never spend the 800ms window.
* fix(web): fail on a partial extractor override
same mixed-producer hazard as the install error: some extractors hold
the page's value and the rest hold goja's, and a property comparing
across that split fires on a healthy app.
* docs: six of seven, the seventh is the stock property
* fix(folio): drop a name two cards answer to
homeTxnCountsOf keyed on the account name and let the last card win, so
two accounts the fuzzer named the same collapsed into one entry. a
reading that saw one Travel card and a later one that saw both then
subtracted two different accounts' counts, and
submitCommitsOneTransactionPerAction convicted a healthy app of
double-submitting. it is a gated property in folio-run.sh, so that reads
as "found the submit bug" over a card scrolling into view.
same rule createdAccountHasNonZeroBalance already applies: a name
nothing can attribute is no evidence. counted over every card, since an
unreadable twin spoils the identity too.
* perf(folio): read each frame once
every extractor asked routeOf, and routeOf does five ax.find calls. on
web each find walks the document and every shadow root beneath it, so
the spec cost 110 tree walks a step; homeCards was parsed four times
over. now 5 and once.
keyed on the identity of the state object because both hosts build a new
one per step and hand that one object to every getter, so it cannot
outlive its frame. holding the reference is what keeps that true rather
than likely.
* fix(web): keep an undefined reading's index through JSON
json has no undefined, so an extractor whose getter returned one had its
whole index dropped by JSON.stringify. that index then kept goja's
dump-derived value while its neighbours held the page's, and a property
comparing previous to current across the split fires on a healthy app.
folio has nine on(route, tag) extractors, so this was most extractors on
most steps.
each reading is wrapped in a {value} envelope: the drop now happens
inside the entry, and an absent value means the getter returned
undefined, which is what the goja host records for the same getter. a
json null would instead claim it returned null and x.current ===
undefined would answer differently on the two hosts.
* feat(verifier): report the registered extractor count
the web path needs it to check the page sent one reading per extractor.
* fix(runner): fail when the page reports fewer readings than extractors
the comment here already claimed a partial override was fatal. it was
not: the skipped check only catches indices outside the extractor list,
so a page reporting values for some extractors and not others left the
rest holding goja's reading of the dump with nothing said.
* test(browser): drive an undefined reading through the whole web path
four layers carry it: the page's envelope, the driver's unwrap, the
runner's count check and the verifier's decode. each has a unit test and
only a run proves they compose. goes red both ways, decoding an absent
value as null and dropping the envelope.
* fix(web): offer the aria roles a user activates
only role=button was in the tappable set, so link, checkbox, radio,
switch, tab, option, the menuitems and treeitem were invisible to the
enumeration however plain the control looked. the replay ui builds its
step rows as <li role="option">, and the spec dogfooding it had to
hand-write an action to reach them because no default verb could see a
single row.
both producers build the set from the same role list, since the parity
test compares them element by element.
* test(browser): tap a role-based control end to end
every control on the page is an <li role="option">, the shape the
replay ui gives its step rows, and the spec carries no action of its
own: the property firing is the evidence the default enumeration offered
a tap on one.
* fix(web): read aria-disabled as disabled
the enabled fact came off the disabled property, which only real form
controls have. it reads undefined on the role-based controls the
tappable set now covers, so every one of them looked enabled however
plainly it was marked otherwise, and the fuzzer would spend actions on
inert ones.
both producers answer the same two ways, and the parity fixture carries
a disabled row so the comparison covers it: reverting one side alone
names the element and the fact.
* docs(replay-ui): the enumeration reaches step rows now
the comment said role="option" is not in the tappable selector set,
which stopped being true a few commits ago. selectAStep stays, for the
reason the tab weight below it stays: one row among the page's clickable
elements is a thin chance, and both step-facing properties go vacuous on
a run that never selects one.
* test(runner): bound the last-action test by steps, not wall clock
100ms of wall clock against an assertion that two steps ran fatals under
load with "the web path never installed it", which reads as a
regression. every sibling test in the package uses a long duration and
MaxSteps.
* ci: run the kotlin tests in make test
RouteTransitionTest and the stability poll cover the android settle and
nothing in ci ran them. :sidecar:test needs no android sdk, checked by
running it with ANDROID_HOME pointed at nothing.
* fix(sidecar): measure the stability streak as observed quiet
parameterising pollUntilStable also moved the clock to the start of the
read that opened a run of identical snapshots, so a read's own duration
counted as quiet. the pre-existing caller polls a real uiautomator dump:
at 400ms a read, 750ms of required quiet became 250ms of observed quiet
and the poll settled in two reads instead of four.
the parameters stay, the semantics go back.
* test(sidecar): pin the transition cap by driving it
it asserted 1500 >= 700 + 300, two constants, which can only fail if
someone edits a constant. it now drives awaitSettledTree against a fade
that lands after 700ms and asserts it hands back the settled tree before
the cap. cut the cap to 1000 and it goes red.
* ci: pin buf-setup-action to a commit
it takes a token now, so a floating tag is a token handed to whatever
that tag moves to. note v1 there is a branch, not a tag, so the ref
lookup that resolves it is matching-refs/heads/v1.
* ci: declare least-privilege permissions
none of the three declared any, so each got the repository default.
release.yml and docs.yml already do this. all three only check out,
build, test and upload artifacts.
* ci: fail fast when a server never comes up
the readiness loops fell through silently after 30 tries, so a server
that never started surfaced as an opaque driver failure minutes later.
each now says what did not answer and on which port.
* ci(folio): a missing trace is not a verdict
with no trace the android gate ran its grep against ./trace.jsonl and
reported "never reached AddTransactionScreen, so it never got past
login", which is not what happened. the web and ios branches had the
same misdiagnosis on exit 0.
same class, one line up: the classifier's own failure was swallowed, so
with the evidence reader dead the gate printed a healthy run and exited
0.
* ci(replay-ui): skip a run directory with no trace
the summarise step is if: always(), and under github's bash -eo pipefail
an unmatched glob stays literal, the redirect fails, pipefail carries it
into the assignment and -e kills the step. so a failed fuzz run went red
twice, once for the real reason.
At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.
The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.
The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.
Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.
The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.
The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.
Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.
Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.
Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.
Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.
A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.
Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.
An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.
The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.
Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.
Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.
Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.
The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.
idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.
Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.
Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.
No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.
An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.
Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.
A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.
The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.
This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.
The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(replay-ui): drop @font-face rules for fonts that were never shipped
* fix(replay-ui): drop unresolvable JetBrains Mono from --font-mono stack
* chore(replay-ui): remove vestigial empty public/fonts dir
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.
A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.
Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(cli): add --max-steps for step-bounded runs
runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(trace): record arm membership and host in meta.json
meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(cli): add --arm and populate run meta from it
Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): sweep seeds for one experiment cell
campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.
Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(runner): no silent generator fallback, and llm on web
--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.
pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): make the hierarchy dump agree with the web runtime
Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.
scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.
The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(conformance): give the gate reproducible seeds
SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.
The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): emit editable as a plain boolean
editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(spec): leave the head subtree out of the web target walk
collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* test(chrome): compare the facts both hosts derive from one DOM
The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.
This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* chore(make): run the browser packages one at a time
Both launch Chrome and launching two at once has failed with "Launch: context
canceled".
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style: remove every em-dash and en-dash
Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): honor the caller context in Launch
Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.
The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(sidecarassets): publish the extracted jar through a rename
Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): kill a run that outlives --run-timeout
A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style(test): gofmt browser_test.go
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(ltl): give every thunk a construction identity
Two distinct unnamed predicates both described as "Thunk(...)", so obligation
collapse merged their residuals and could drop a live violation. Identity is
assigned at construction and the fields are unexported, so a thunk cannot be
built without one.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): reduce a thrown-predicate residual instead of panicking
The verifier substitutes an ErrorFormula for the residual of a property whose
predicate threw, and that residual is fed back in on the next step. reduce had
no case for it, so the run crashed. It re-reports the same failure now.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): make a bounded always the dual of a bounded eventually
G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still
pending when the window closed, so nnf's negation normal form was not semantics
preserving. Both sides now range over the observations at which their inner can
definitely resolve: the eventually keeps a pending inner as a disjunct, and the
always discharges vacuously at window close.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): arm a one-shot root once per run
A root that carries its own horizon is one obligation for the whole run, not one
per observation. Re-instantiating a top-level eventually monitored G F<=n(p)
instead of F<=n(p) and left one live obligation per step behind; a bounded
always restarted its window every step and never closed.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): stop wrapping a top-level eventually in always
`eventually(p).within(300, "seconds")` as a property meant "within 300 seconds
of every step", which spawned an obligation per step with its own resolved
deadline. A 553-step run carried 553 of them and serialized a 75 KB residual.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): serialize the resolved deadline of a bounded window
Two obligations spawned at different steps from one duration-bounded formula
differ only in the deadline the evaluator resolved for them, so they serialized
identically and the trace erased a distinction the evaluator makes. The authored
window stays in amount/unit; the resolved deadline rides alongside.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): split a witness's origin step from its detection step
A deferred obligation spans two steps: the one that armed it and the one whose
reduction failed. They were conflated under one index, so the extractor snapshot
(which is the detecting step's state) was reported against the origin step.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(runner): record a witness's detection step in the trace
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* feat(replay-ui): show the step a violation was detected at
The witness evidence is the detecting step's state, so say which step that is
and let a reader jump to it.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): record the extractor state the predicates actually read
On the web path extractor bodies are evaluated in V8 and injected here, but only
the goja value was replaced. The trace diff and the violation witness therefore
described a state no property ever saw.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* refactor(spec): one candidate producer over one target-eligibility rule
Both hosts routed verbs themselves and both policies enumerated their own
actions, and all four drifted. Web sent `swipes` to scrollable containers only,
so swipe-to-dismiss on a list row was reachable on native and unreachable on
web; the model policy folded gestures its own way and could not reach what the
seeded picker drew.
A host now reports facts about every element and never decides which verb may
act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates
is the single enumeration, and the model policy reads it through
__sanderlingEnumerateBuiltin__ instead of reimplementing it in Go.
Gesture verbs change with it: scrolls stay vertical over scrollable containers,
swipes go free-form in all four directions from any element with real bounds.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(runner): name a builtin scroll by its drag origin
A builtin gesture carries endpoints and no selector, so every scroll rendered as
"Scroll down " in the prompt's recent-action memory and two scrollable regions
were indistinguishable.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(chrome): clear storage over cdp instead of scripting an opaque origin
Launch runs while the tab is still on about:blank, whose opaque origin denies
storage access, so localStorage.clear() threw SecurityError and every web run
died at launch. Storage.clearDataForOrigin needs no navigation. The exception
helper lands here because "Uncaught" is what hid this for so long.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(chrome): enable the swiftshader webgl fallback
Headless Chrome runs with --disable-gpu, and without this flag it refuses the
software WebGL backend: getContext returns null, so a canvas-rendered app paints
nothing and every screenshot is identical black.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(web): resolve testTag through data-testid or id
Compose Multiplatform emits its testTag into the element id, which the native
table already accepts via the resource-id alias. The two web selector tables
were the only place that rejected it.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* test(spec): type-check the spec api as part of make test
The fake runtime in api.test.ts did not return a chainable handle from extract,
so the file had not type-checked since named() was added. Wiring the check into
make test stops it drifting again.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* docs(manual): one-shot eventually and the gesture verbs
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* feat(spec): add llm() action-backend marker
* feat(spec): make llm marker inert on the JS picker
* feat(spec): expose __sanderlingSampleInput__ corpus draw
* feat(openrouter): minimal chat-completions client
* test(openrouter): cover request shape, parse, and errors
* feat(verifier): thread screenshot + capture corpus sampler
* feat(verifier): LLM accessors — candidates, config, sampler
* test(verifier): cover AllCandidates, LLMConfig, SampleInput
* feat(trace): record action Source and LLMReasoning
* feat(runner): thread step screenshot into PushSnapshot
* feat(runner): llmSource selects actions via OpenRouter
* feat(runner): wire llmSource selection and trace stamping
* test(runner): cover llmSource selection, mapping, downscale
* docs(folio): add llm action-backend example spec
* docs(folio): document the LLM action backend run
* feat(llmclient): support OPENAI_API_KEY, openrouter wins
* refactor(runner): rename openrouter package to llmclient
* docs: both api keys, example model gpt-5.4-nano
* docs: add pr style rules to claude.md
* fix(runner): explain action kinds in llm prompt to stop swipe loops
* feat(trace): record llm ranked list and chosen rank
* feat(runner): stamp llm ranked list and chosen rank on trace
* fix(runner): tap by selector to survive layout shift after observe
* revert(runner): drop selector-first tap; broke path/testTag selectors
* feat(spec): llm() accepts optional instructions
* feat(verifier): read llm instructions off config
* feat(runner): append spec instructions to llm system prompt
* docs(folio): describe app in llm spec instructions
* feat(bundler): map generator export to globalThis.generator
* feat(verifier): read llm config off globalThis.generator
* feat(runner): gate llm source on --generator flag
* feat(cmd): add --generator llm|seeded flag
* test: cover --generator flag parsing and pickSources gating
* feat(verifier): enumerate llm candidates by walking actionsRoot
collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.
* test(verifier): cover candidate enumeration walk
* feat(verifier): add SetupAction to walk setup without the seeded root
* test(verifier): cover SetupAction setup-only precedence
* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order
* feat(trace): record llm choice number and chosen_action echo
* feat(runner): llm picks one number from weighted candidates
drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).
* test(runner): cover choice schema, strict-skip, and setup precedence
* refactor(verifier): drop the superseded AllCandidates enumeration
* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts
* fix(verifier): label editable fields by hint, not the typed value
an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.
* test(runner): cover weight-suffixed echo and stripWeightSuffix
* fix(runner): accept chosen_action echo that carries the weight suffix
real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).
* fix(verifier): skip llm enumeration on cross-fade frames
a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.
* feat(folio): show current balance on the add-transaction screen
renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.
* fix(replay): derive device space from screen extent, not first node
the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.
* fix(folio): show balance as a compact one-line label
per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.
* fix(folio): move balance into the header, one compact line under the account name
* fix(replay): attribute deferred violations to the causing step, not detection
* fix(replay): show a step's own violations in both panels, no next-step bleed
* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check
* chore: ignore .playwright-mcp scratch output
* docs: document the llm generator and --generator flag
* docs(spec): correct the llm() comment; config reads off globalThis.generator
* docs: add pr description rules
* feat(sidecar): reach USB devices via the adb server by serial
* feat(test): add --device flag to target a specific Android device by serial
* feat(folio): select Android device via ANDROID_DEVICE in justfile
* feat(conformance): add android backend to the gate suite
* feat(android): keep device awake and unlocked so the app stays foreground
* feat(conformance): prep physical android device (autofill/verifier/stayon)
* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run
* fix(verifier): require positive bounds for swipe candidates
A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.
* fix(runner): harden app-scope guard against launcher and overlays
The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.
* feat(android): harden physical-device runs in device prep
Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.
* feat(driver): clear-state via APK reinstall when pm clear is blocked
When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.
* feat(cli): add --android-app-path for clear-state reinstall
Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.
* chore(folio): pass --android-app-path in just test
* fix(runner): clamp swipe/scroll origin out of edge gesture zones
A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.
* perf(sidecar): faster Android text input and drop redundant settle poll
inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.
* fix(verifier): exclude soft-keyboard region from action candidates
The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.
* perf(runner): replace focus-tap settle with a brief wait
The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.
* chore(conformance): platform-aware G5 p95 budget for android
The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.
* fix(sidecar): retry maestro android driver startup
The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.
* chore(conformance): widen android G5 budget to 5500ms
Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.
* web replay fix
* feat(android): force 3-button nav during runs to prevent app drift
On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.
* fix(android): target the selected device in adb reads; don't strand nav mode
Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
foreground/scope guard works when several devices are attached (the --device
path). Previously they ran bare `adb shell`, which errors with multiple
devices, silently disabling app-scope enforcement. The sidecar client passes
its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
the current mode is unknown or already 3-button it leaves nav untouched,
instead of switching and then stranding the device in 3-button. Logic split
into the pure navModeToRestore, now unit tested.
* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds
Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.
* test(verifier): cover keyboardRegionTop, including the decor-view guard
The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.
* style(cli): gofmt testOptions field alignment
* fix(sidecar): keep a leading dash off the fast input path
A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).
* refactor(verifier): scope action candidates by window ownership
Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.
This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.
* fix(runner): re-check foreground at apply time, skip stale actions
ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.
* fix(android): type long ASCII via fast guarded path to stop keystroke escape
A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.
* chore: ignore gate artifacts and local scratch files
* refactor(runner): narrow gesture clamp to the top shade strip
3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.
* chore(format): add .editorconfig enforcing 80-column limit
* chore(format): add prettier config with 80-char printWidth
* chore(deps): add prettier devDependency to replay-ui
* chore(deps): add prettier devDependency to folio-web
* chore(deps): add prettier devDependency to spec package
* chore(format): add swift-format config with 80-char lineLength
* feat(format): add make fmt targets for per-language 80-col formatting
* fix(runner): translate gesture to safe area so near-top scrolls keep direction
Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.
* fix(runner): apply-time guard consults focused window, not just resumed activity
ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.
* test(runner): cover apply-time foreground skip and appIsForeground table
Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.
* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker
- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.
* fix(conformance): pin self-test p95 budget and score install failures as run failures
self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.
* fix(android): require --device when several devices are connected
With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.
* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands
The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.
* perf(verifier): memoize scopedElements per tree
scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.
* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear
Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.
* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms
The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.
* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty
bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.
* perf(sidecar): reuse a single Jackson ObjectMapper
structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.
* refactor(android): remove unused AdbReverse/AdbReverseRemove
No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.
* style(runner): trim non-load-bearing comments from this PR's runner code and tests
* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
* docs(manual): add introduction page
* docs(manual): rewrite getting started as guided first run
* docs(manual): rewrite writing specs as a folio tutorial
* docs(manual): document missing spec API in reference
* docs(manual): plain-language rewrite of runs page
* docs: real introductions on index pages and README
* fix(docs): sibling links from directory-style pages need ../
* fix(docs): correct sampling and restart-cost claims to match implementation
* docs: nav lists Introduction and Case study; roadmap points to milestone
* docs(manual): make getting started target the reader's own app, not Folio
* docs(manual): add Folio case study page
* docs: point manual navigation at the case study
* docs(readme): lead with the case study, fix roadmap link
* docs: roadmap links to milestone, sync clear-data default and cross-links
* feat(companion): add appState, eraseText, pressKey runner handlers
The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.
* feat(ios): resolve physical devices from devicectl
ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.
* feat(ioscompanion): runner-only device driver mode
NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.
* test(ioscompanion): cover device wiring, routing, and shell-out argv
Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.
* feat(testrun): route physical-device iOS runs to the device driver
Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.
* feat(doctor): device prereqs replace java/sidecar for ios-device
iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.
* feat(conformance): device backend uses iphoneos app and tunnel orphan checks
The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).
* feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.
* docs(cli): document ios-device doctor checks and the device flags
The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.
* fix(ioscompanion): resolve signing key path to absolute
xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.
* fix(ioscompanion): re-enable signing for the device runner build
companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.
* fix(ioscompanion): key the device build cache on signing identity
The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.
* docs(getting-started): document physical iOS device setup
Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.
* feat(ios): native usbmux client and in-process tunnel forwarder
Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.
* refactor(ios): drive device tunnel via io.Closer seam
Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.
* refactor(ios): remove iproxy spawn from device runner
* test(ios): cover tunnel close via io.Closer not child process
* feat(doctor): check usbmuxd socket instead of iproxy on PATH
* chore(conformance): drop iproxy orphan check; tunnel is in-process
* docs(ios): device tunnel uses native usbmux, nothing to install
* chore: gitignore the signing keys directory
* feat(folio): add Android launcher icon (black bg, white dot)
* feat(folio): add iOS app icon (black bg, white dot)
* feat(folio): add web favicon (black bg, white dot)
* docs(ioscompanion): fix stale const comments
* refactor(ioscompanion): inline single-use devicectl argv builders
* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests
* refactor(ioscompanion): inline firstNonEmpty
* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam
* test(doctor): trim redundant signing-check test
* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env
* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
* refactor(sidecar): drop IosDriverBackend
* refactor(sidecar): route ios platform off the iOS backend
* chore(sidecar): remove maestro ios dependencies
* feat(testrun): reject physical iOS with a clear message
* refactor(sidecar): drop iOS hierarchy helpers and their test
* build(sidecar): strip iOS runner bundles and classes from the fat jar
* test(testrun): cover physical-iOS rejection
* perf(ios): use prebuilt XCTest runner to cut startup
* chore(ioscompanion): add companion asset prepare script
* feat(ioscompanion): embed and extract simulator companion bundle
* test(ioscompanion): cover companion stub and embedded extraction
* docs: add third party notices for vendored companion
* chore: ignore vendored companion bundle artifact
* build(proto): pin simulator companion proto v1.1.8
* build(proto): add dedicated buf module and gen template for pinned proto
* build(proto): exclude pinned companion proto from root buf workspace
* feat(ioscompanion): commit generated companion gRPC stubs
* feat(ioscompanion): map flat companion describe dump to TreeNode JSON
* test(ioscompanion): add hierarchy-map golden and unit tests
* feat(ioscompanion): port screen-settle stability polling to Go
* test(ioscompanion): cover settle transitional, hash, streak, and cap rules
* feat(ioscompanion): add USB HID keymap module
* test(ioscompanion): cover keymap branches and paste-chord constants
* build: embed companion assets via withcompanion tag
* feat(ioscompanion): add transport companion interface
* feat(ioscompanion): add HID event wrapper and builders
* feat(ioscompanion): wire gRPC companion client and Dial
* test(ioscompanion): cover HID builders and unit conversions
* test(ioscompanion): cover Dial, process-state mapping, and install archive
* test(ioscompanion): add gated simulator integration smoke test
* feat(ioscompanion): text input and gesture HID composition with pasteboard fallback
* test(ioscompanion): cover input composers, paste dialog loop, and pure helpers
* feat(ioscompanion): add Describe to companion transport
* feat(ioscompanion): implement DeviceDriver with companion supervision
* test(ioscompanion): unit tests with fake companion transport
* test(ioscompanion): gated companion smoke test
* feat(ios): add ResolveTarget for simulator vs physical-device routing
* feat(testrun): route iOS simulators through the native companion driver
* refactor(testrun): defer the java preflight check to the physical-device path
* feat(cli): add --ios-app-path flag
* feat(doctor): split iOS checks into simulator and physical-device paths
* test(folio): add gate-analyzer fixtures for G1-G5
* feat(folio): add iOS conformance gate script
* chore(folio): wire gates recipe, app path, and ignore gate output
* style: gofmt struct alignment drift
* fix(doctor): probe simctl via xcrun instead of PATH lookup
* fix(ioscompanion): spawn companion under driver-lifetime context
* test(ioscompanion): prove companion child outlives startup context
* fix(ioscompanion): chunk install payload under companion message cap
* test(ioscompanion): cover install payload chunking
* fix(ioscompanion): reinstall via simctl and sanitize companion env
* fix(ioscompanion): wait out unresolved accessibility values after launch
* perf(ioscompanion): paste long text for atomic landing
* test(ioscompanion): cover paste threshold, retry flow, and sentinel detection
* fix(ioscompanion): treat unresolved bridge values as transitional, never as content
* fix(ioscompanion): accept masked secure-field values as paste landing
* test(ioscompanion): cover sentinel mapping and masked-field landing
* fix(ioscompanion): atomic erase and single-send paste to prevent doubling
* test(ioscompanion): cover atomic erase, single chord, unverifiable field
* fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout
* test(ioscompanion): cover bridge-blackout paste verification
* fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle
* refactor(ioscompanion): name the empty-editable-field sentinel for what it is
* perf(ioscompanion): tighten settle streak for the fast companion transport
* feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt
* refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt
* test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests
* fix(ioscompanion): retry describe past transient collapsed accessibility dumps
* test(ioscompanion): cover collapsed-dump detection
* perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses
* perf(ioscompanion): tighten settle now that collapses are handled separately
* fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text
* test(ioscompanion): cover replace-on-input and TextReplacer capability
* refactor(ioscompanion): neutralize HID events behind the transport seam
* feat(companion): add simulator runner project skeleton
* feat(companion): serve accessibility snapshots over the wire protocol
* feat(companion): synthesize timestamped touch gestures
* feat(companion): type text with replace semantics
* feat(companion): serve the wire protocol from a parked runner
* feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam
* feat(ioscompanion): route text input through a text-editing companion when available
* fix(companion): bind listener by port and source screen size from snapshot
* feat(ioscompanion): add runner companion JSON transport
* test(ioscompanion): cover runner transport protocol mapping
* fix(companion): synthesize gestures synchronously to avoid the async completion crash
* fix(companion): type on the main thread and recover from focus assertions
* fix(companion): keep serving after an automation failure
* refactor(companion): tidy snapshot serialization
* fix(companion): honor sequential tap gaps and survive synthesis exceptions
* feat(ioscompanion): expose native typing with an explicit replace flag
* chore(companion): add runner asset prepare script
* feat(ioscompanion): embed and extract the runner test bundle
* test(ioscompanion): cover runner asset extraction
* build(ioscompanion): commit runner asset archive
* feat(ioscompanion): pair the legacy companion with the in-simulator runner
* test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding
* fix(ioscompanion): reconnect after interrupted runner calls instead of restarting
* fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts
* feat(companion): launch and terminate apps through the automation session
* build(ioscompanion): refresh runner asset with session lifecycle
* fix(ioscompanion): classify connection deadline expiry as caller budget
* fix(companion): capture snapshots on the main thread inside the catch bridge
* build(ioscompanion): refresh runner asset with main-thread snapshots
* perf(ioscompanion): count read spans toward settle and capture snapshots concurrently
* feat(ioscompanion): make the hybrid simulator companion the default
* test(folio): cover runner-session orphans in the gate harness
* test(ioscompanion): pin the child-lifetime test to the legacy path
* fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears
* fix(ioscompanion): pause the clear chord so selection applies before the delete
* fix(companion): prune the keyboard subtree from snapshots
* build(ioscompanion): refresh runner asset without keyboard elements
* fix(ioscompanion): capture the screenshot transport before a recovery can reassign it
* fix(companion): pin the runner listener to loopback
* fix(companion): size the replace delete prefix to cover any focused field
* build(ioscompanion): refresh runner asset with loopback bind and replace fix
* fix(cli): cancel the run context on SIGINT so spawned children are reaped
* fix(testrun): point the device java preflight hint at the ios-device doctor
* fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob
* test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision
* chore: add test-companion target for the withcompanion-tagged suite
* chore(ioscompanion): stop tracking the runner archive build artifact
* build: produce the runner archive from source like the companion bundle
* refactor(conformance): move the gate harness out of examples/folio
* chore(folio): drop the gate harness wiring from the example app
* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
* feat(ltl): attribute violations to the obligation origin step
* feat(verifier): label evaluator observations with the runner step index
* feat(trace): carry the causing step in violation witnesses and summary
* feat(replay): move the violation marker to the causing step
* feat(replay-ui): render witness evidence in the violations panel
* feat(replay-ui): wire witnesses and step jump into violation panels
* fix(ltl): treat next obligations as vacuous at run end
* fix(runner): give the finalize trace record its own step index
* fix(hierarchy): bounds-containment fallback for scoped and path queries
Compose on iOS surfaces a testTag node as an empty leaf sibling of the
content it labels instead of as an ancestor, so descendant search under
the tagged node finds nothing and every path or scoped query returns
null. When structural search yields no match, fall back to nodes whose
bounds lie inside the scope node's bounds.
* feat(sidecar): derive iOS clickable and editable from element type
The XCTest hierarchy mapping dropped the element type, leaving no
clickable or editable flags on iOS, so the fuzzer's tap and typing
verbs never found a candidate inside the app. Map the raw
accessibility tree directly and derive clickable, editable,
scrollable, and class from the XCUIElementType raw value.
* feat(proto): add EraseText RPC for InputText replace semantics
* feat(driver): add EraseText to the device driver surface
* fix(runner): erase existing field text before InputText
InputText appended on native platforms, so repeated draws grew fields
without bound. The folio fuzz run wedged on the add-account screen:
each draw concatenated another name until the 40-character validation
error became permanent. Replace semantics also makes retried typing
idempotent. The web driver already replaced via select-all; native now
matches.
* feat(sidecar): EraseText backend support on android and ios
* fix(folio): saturation-gate account creation in the spec
The 2-3 step add-account loop outcompeted the 5-step transaction chain
at every weighted re-draw, so runs filled with account creation and
rarely exercised the balance properties. Stop offering add-account once
three accounts exist; the renormalized weights then favor the
transaction flow at every step of its chain.
* fix(folio): author spec weights to match testing intent
Revert the account saturation gate: it starved newAccountBalanceIsZero
once it tripped, and a magic account count is app-state tuning, not
intent. Instead weight the generators by what the properties need:
the transaction chain leads, account creation stays exercised, and
doubleTaps gets explicit weight everywhere because double-submission
idempotency is what the spec is testing for.
* fix(folio): lower doubleTaps weight to 5
* fix(sidecar): surface visible text on iOS static elements
Static text and button strings live in the accessibility label on
iOS, so the text attribute came through empty and every balance
extractor parsed to zero, silently disarming both folio properties.
Non-editable elements now fall back title, value, then label;
editable fields keep value-only so an empty field's caption does not
read as content.
* feat(driver): native DoubleTap RPC for a tight inter-tap gap
Composing two Tap round trips from the Go client spread the taps by
hundreds of milliseconds on iOS, wide enough for the app to navigate
between them, so double-submission races could never reproduce. The
sidecar now lands both taps back-to-back next to the device transport.
* feat(sidecar): pipeline iOS double-tap requests
Queue the second tap at the XCTest runner while the first executes.
The runner serializes handlers, so this is the tightest gap the
transport allows (~350ms per tap round trip); recorded here with
measurements for the iOS double-tap limitation.
* refactor: rename inspect to replay across the codebase
Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/,
the CLI subcommand from `sanderling inspect` to `sanderling replay`, and
updates all references in docs, Makefile, README, and Go comments.
* feat(replay-ui): show spec filename with full path on hover
RunList and RunDetail now render the basename of spec_path (e.g.
login.spec.ts) with the full path available as a title tooltip.
* feat(ltl): bound fields on AlwaysFormula and named thunks
Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.
* feat(ltl): negation normal form pass
nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.
* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse
Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.
* test(ltl): property-based NNF laws
Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.
* test(ltl): Finalize, bounded eventually, latch, collapse
Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.
* feat(inspect): within clause on always residual node
A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.
* feat(ltl): witness violations and (bool,error) predicate thunks
* test(ltl): migrate thunk call sites to (bool,error)
* feat(ltl): flag thrown-predicate witnesses with IsError
* refactor(verifier): replace predicate err side-channel with violation witness
* test(verifier): witness API for thrown predicates
* feat(trace): witnesses map and skipped-verification marker on Step
* feat(runner): thread violation witnesses, finalize, skip marker into trace
* test(ltl): lock violation witness reason, IsError, and step
* test(verifier): finalize surfaces unmet eventually with witness
* fix(ltl): eliminate implies and bounded-always false-negatives
Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.
* test(ltl): lock implies and bounded-always false-negative regressions
* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick
* feat(testrun): inject seed into web bundle via SANDERLING_SEED define
* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring
* test(spec): add Go math/rand/v2 PCG oracle and golden fixture
* feat(spec): bit-exact PCG port of Go math/rand/v2
* test(spec): assert pcg.ts matches the PCG golden fixture
* feat(spec): shared input corpus and press-key pools
* feat(spec): action-tree types and Host interface
* feat(spec): verb support matrix and warn-once helper
* feat(spec): deterministic shared action picker
* test(spec): verb matrix and warn-once semantics
* test(spec): picker draw-order and determinism
* refactor(spec): actions.ts returns pure GeneratorNode data trees
* refactor(spec): wire from() sampling through the picker rng
* feat(spec): shared runtime-entry installs next-action over pick.ts
* feat(spec): export LongPress/Scroll/longPresses/scrolls factories
* test(spec): assert data-tree shapes for action factories
* test(spec): runtime-entry serializeAction wire-contract round-trip
* refactor(spec): bridge data-tree nodes to the legacy goja picker tags
* fix(spec): web runtime walks the spec's globalThis.actions data tree
* test(spec): tolerate legacy bridge fields on builtin nodes
* refactor(spec): installRuntime accepts a lazy root resolver
The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.
* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker
Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).
* test(spec): cover the WEB Host surface and seed precision
Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.
* refactor(spec): picker emits native selector + scroll endpoints, setup precedence
* feat(spec): goja runtime entry wires the shared picker over the Go host
* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin
* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker
* refactor(spec): drop the legacy goja bridge fields from action factories
* feat(spec): serialize selector-only string targets for the runner to re-resolve
* refactor(verifier): one DecodeAction reads the unified flat wire contract
* refactor(verifier): goja host + shared picker replace the duplicate Go picker
* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime
* test(verifier): author specs through the shared picker path
* test(runner): bundle authored specs with the goja runtime entry
* feat(verifier): collect unsupported verbs for the run report
* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource
* feat(testrun): surface unsupported verbs in run report
* test(verifier): cross-runtime goja/node parity gate on the shared picker
* test(verifier): unsupported verbs collected deduped in first-seen order
* test(runner): summary reports no unsupported verbs on a clean run
* test(spec): golden-fixture cross-runtime parity gate for the node picker
Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.
* test(verifier): assert goja picker against the same cross-runtime golden
Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.
* refactor(spec): rename pressKey generator export to pressKeys
* refactor(spec): update barrel re-exports for pressKeys
* test(spec): update pressKeys generator export name
* docs(spec): rename pressKey generator to pressKeys
* refactor(spec): extract samplerRng into shared sampler-rng module
* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)
* test(spec): cover fluent value generators determinism and chaining
* refactor(bundler): inject globalThis trailer from spec named exports
* refactor(bundler): reuse registration trailer in web bundler
* test(bundler): cover named-export globalThis registration
* feat(spec): add named() to Extracted handle type
* feat(web-runtime): named() and cross-extractor read guard
* feat(verifier): named() and cross-extractor read guard in goja
* test(verifier): cross-extractor read guard and named()
* test(web-runtime): export runtime and extractors for tests
* test(web-runtime): named() and cross-extractor read guard
* refactor(folio): drop manual globalThis trailer (bundler injects it)
* refactor(folio): seed txn amounts via integers().between(1,500)
* refactor(folio-web): drop manual globalThis trailer (bundler injects it)
* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs
* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts
* refactor(folio-web): name extractors so violation witnesses are readable
* fix(web-runtime): propagate extractor getter throws and unpoison locked global
Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.
* test(spec): install fake runtime via defineProperty to survive locked global
* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors
* feat(runner): add MaxSteps bound to Options
* test(runner): MaxSteps stops after exactly N steps
* test(driverpb): drop proto getter round-trip tautology
* test(sidecar): drop stub-mode placeholder tautology tests
* test(mock): drop default-field-value assertion test
* test(ltl): drop Verdict.String tautology tests
* refactor(runner): extract RenderSummary for snapshot testing
* test(runner): golden snapshots for trace stream and violation summary
* feat(web-runtime): capture uncaught errors into state.exceptions
* test(integration): add throwing and counter web fixtures
* test(integration): add specs for the web fixtures
* test(integration): drive web fixtures through the real pipeline in headless Chrome
* chore(make): add test-browser target for the Chrome-driven suite
* ci: run the Chrome-driven browser suite in a separate job
* refactor(test): relocate browser suite to test/browser
* refactor(permissions): delete dead internal/permissions package
* refactor(test): rename package to browser_test
* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets
* chore(make): point test-browser at test/browser
* docs(decisions): record internal/permissions deletion
* refactor(doctor): use sidecarassets package
* refactor(testrun): use sidecarassets package
* fix(test): resolve testdata relative to browser_test.go
* refactor(verifier): remove dead __sanderlingIndex compat alias
* refactor(bundler): use encoding/json for JS string literals
* docs(action-space): use vendor-neutral native driver wording
* refactor(hierarchy): scrub backend tool name from comments
* refactor(driver): scrub backend tool name from comments
* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver
* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap
* refactor(chrome): implement DoubleTap as two taps with the gap
* refactor(mock): record DoubleTap and DoubleTapSelector actions
* refactor(runner): delegate double-tap to driver, drop gesture timing
* test(runner): assert double-tap delegates to driver DoubleTap
* docs(cmd): add package docs to CLI and developer tools
* docs(driver): add package docs to driver interface and chrome backend
* docs(driver): add package docs to mock and sidecar backends
* docs(platform): add package docs to android and ios device prep
* docs: add package docs to bundler and inspect
* docs(ltl): add package doc to temporal logic evaluator
* docs: add package docs to runner and testrun pipeline
* docs: add package docs to trace and verifier
* docs(sidecarassets): add package doc for embedded JAR loader
* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI
* test(chrome): gate real-Chrome driver tests behind the browser tag
* chore(make): run chrome driver tests in the browser job
* fix(web-runtime): guard global error listeners for non-browser hosts
The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.
* ci(browser): re-enable unprivileged user namespaces for headless Chrome
ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.
* ci(browser): pin stable Chrome for the driver tests
setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.
* feat(defaults): add scroll and rebalance action weights
Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.
* feat(defaults): trim scroll action weight wiring
* fix(build): point sidecar jar ignore and embed paths at sidecarassets
* test(defaults): drop stale longPresses re-export assertion
longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.
* fix(chrome): raise DevTools websocket read timeout to 60s
Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.