correctness fixes from the first real folio dispatch, and spec authoring skills (#82)

* docs: add the apache 2.0 license text

package.json has declared Apache-2.0 since the first release and .goreleaser.yaml
globs LICENSE* into the archives, so that glob has been matching nothing. npm
only picks up a license from the package directory, hence the copy under
pkg/spec.

* fix(sidecar): close the soft keyboard after typing on android

* test(sidecar): pin the guarded ime dismissal

* fix(sidecar): treat a failed ime probe as no keyboard open

* ci(folio): let the ios leg clear state for itself

* ci(folio): drop the stale frontboard note from the ios job

* docs(ci): record what the ios calibration assumes and where it was measured

* fix(ios): replace the session when a launch blows its bound

a launch the simulator refuses is never reported: xctest records it as a test
failure the runner cannot see, then holds the session's main thread for about
four minutes on a diagnostic chain. so the only signal is the expired bound,
and every later call queues behind the same wedge. restart the session once and
launch again, bounded so the launch path stays inside testrun's backstop.

* test(ios): cover the session replacement a wedged launch needs

* fix(ios): share one deadline across the restart and the second launch

the recovery a blown bound triggers now costs at most launchRecoveryTimeout
whatever it spends it on, so the launch path tops out at 150s and testrun's
three minute backstop stays a backstop.

* test(ios): the restart a blown launch triggers has to be bounded

* docs(ci): the ios leg does not convict on the runner, and a seed cannot fix it

seed 28 reproduced its walk on macos-15 and reached the bug at the step it
convicts at locally. it still could not be judged: the run did not return
home between step 19 and step 136, so the counting invariant saw a rise of
15 against a window of 37 submits.

* docs(sidecar): record why the stale ime flag stays out of reach

* test(folio): add commonTest source sets to core and shared

* fix(folio): reject amounts parseCents cannot represent

* fix(folio): cap a transaction at one million dollars

* test(spec): give the fake dom a real tree and a walking querySelectorAll

* style(sidecar): make ktlint clean, formatting only

ktlint -F over every kotlin file except DriverBackend.kt, then hand
fixes where the reflow read worse and for the long lines ktlint cannot
break. No behaviour changes.

DriverBackend.kt is left untouched to avoid a conflict with concurrent
work; its three over-long lines still fail fmt-kotlin.

* fix(spec): deepQueryAll returns matches in document order

* ci(folio): say what the android gate found, not why

The gate proves only that AddTransactionScreen is absent from the trace.
Claiming the run never got past login was an inference it cannot make: the
run that produced it had logged in and was stuck on the new account screen.
Name the routes the trace does record instead.

* test(runner): a relaunch must not convict the submit counting property

* test(browser): compare ax.find across both hosts on one page

* test(spec): name the shadow match in the grammar both hosts parse

* ci: move every action off the node20 runtime

checkout v4->v7, setup-go v5->v7, setup-node v4->v7, setup-java v4->v5,
upload-artifact v4->v7, cache v4->v6, upload-pages-artifact v3->v5,
deploy-pages v4->v5, setup-chrome v1->v2, setup-android v3->v4,
goreleaser-action v6->v7. setup-bun and android-emulator-runner are
already node24; buf-setup-action stays on its deliberate SHA pin.

setup-chrome v2 resolves stable from Chrome for Testing rather than the
official installer, so the ci.yml comment about the action's default no
longer held.

* feat(verifier): report a relaunch on state.lastAction

* test(verifier): pin the relaunch field on both hosts

* fix(runner): keep the action the app was relaunched after

* test: pin that a nested undefined does not survive the wire

* fix(folio): uninstall before installing in just ios

folio's signed-in session lives in the data container, which an install
over the top keeps, so a local run started right after just ios opened on
the previous run's Home screen and diverged at step 1. On CI's fresh
simulator the uninstall is a no-op, so the ios leg is unchanged.

* style(sidecar): bring the last three lines under the line limit

* fix(ios): the runner must not answer ok for a launch that failed

XCTest records a refused launch as a test issue that never throws, so the
companion returned ok for an app that never started. Check the state the
app actually reached and report the refusal instead.

* test(ios): a refusal the runner names costs no session restart

The session restart is for a launch that never answers. A launch that
reports the app's state has already said what a fresh session would.

* fix(folio): attribute a created account by its whole key, not a suffix

createdAccountHasNonZeroBalance matched the created card with endsWith, so
an older account whose name ends with the typed one ("Emergency Fund" for a
typed "Fund") was judged instead whenever the new card was clipped out of
the reading. Build both keys the card can carry, the plain name and web's
initials + name, and compare them whole.

* fix(hierarchy): object selectors resolve by the same rule as string ones

* test(verifier): both ax.find selector forms resolve the same element

* test(browser): the cross-host fixture uses the object selector form

* fix(replay-ui): wrap the tab strip so its last tabs stay clickable

* test(replay-ui): drive the fuzzer onto a violating step with a panel

* feat(folio): judge a submit against the account's own balance

The counting invariant can only close its window on Home, and the iOS run
in #78 went 117 steps between two Home readings: 37 submits against a rise
of 15 transactions is no evidence about the double tap sitting inside it.
The ledger and the add-transaction screen both show the account's own
balance, and an accepted submit pops back to the ledger, so a window
bounded by those readings holds one action.

The bound is an upper one: a balance that has not moved is a commit still
in flight, a rejected submit or a tap that never landed, and none of those
is a violation. Moving by more than the one submit in the window typed is.

* test(runner): an overlay dismissal must not convict the counting property

* docs(runner): point the guard comment at the renamed test

* docs(replay-ui): name the viewport the tab overflow was measured at

* feat(folio): close the submit window on the account's own screens

submitCommitsOneTransactionPerAction now states its rule over two windows:
the Home counts it already compared, and the account balance the ledger and
the add-transaction screen redraw on nearly every frame of the transaction
flow. Same rule, and the second window is usually one action wide.

* fix(sidecar): reach the adb server the environment names

buildDadb hardcoded localhost:5037, so a serial-addressed device always
resolved through this machine's adb server and ADB_SERVER_SOCKET was
ignored. Read the endpoint the way the adb CLI does instead.

Fixes #79

* test(sidecar): pin the adb server endpoint parsing

* fix(sidecar): close a keyboard standing in the snapshot

A tap on a text field raises the keyboard and nothing closed it, so the
tree the picker chooses from was missing every app node underneath it,
the submit control included. Close it before the read rather than after
the tap: the picker only ever sees snapshots, and the keyboard is still
on its way up when the tap returns.

Fixes #78

* test(sidecar): pin the tree-guarded keyboard dismissal

* chore(make): a target that runs folio's unit tests

* chore(folio): a just recipe for the unit tests

* ci: run folio's unit tests on every pr

* ci: switch to jdk 21 only for the folio step

* docs(sidecar): put the measured read cost in the dismissal bound

* fix(folio): decline the two demanding properties across a relaunch

The runner now keeps lastAction and marks it relaunched: true where it used
to report nothing at all, so the two properties that demand an effect judge
a step whose process may have died before the write landed.
submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both
decline there. The counting bound does not: a relaunch cannot manufacture a
transaction, and the submit is counted, so declining would throw away the
detection the runner fix restored.

* docs(folio): say why the merged card key cannot be made injective

Folio rejects a duplicate account name, so the twin the drop rule guards
against is two names the web key cannot tell apart, not two accounts
sharing a name. State what closing the rest would cost and what the tree
would have to carry to close it properly.

* test(folio): pin that two accounts can render the same card text

The proof behind the comment: "Travel1" holding 25 transactions and
"Travel12" holding 5 merge to the same string, so no identity key read off
a web card can tell them apart.

* fix(sidecar): bound the diagnostic adb reads

adbOutput and readLogcat read to EOF and then waited with no timeout, so
a wedged adb held the step for as long as it liked; one stall over a
remote adb server measured ~100s. The bound has to sit on the read, not
on waitFor: a wedged adb never reaches EOF, so a bounded waitFor after
the read is a line that never runs.

* test(sidecar): pin the bound on a wedged adb read

* fix(sidecar): an unreadable animation count is not idle

Defaulting the count to zero made a dumpsys that said nothing mean
nothing is animating, so a degraded link broke out of the settle early
and handed the runner a frame caught mid-animation. Unknown now waits,
inside the deadline waitForIdle already holds.

* test(sidecar): unknown animation state must not read as idle

* fix(folio): stop spending the submit budget on taps the app refused

The window is an upper bound on the transactions an interval could hold, and
a bound inflated by taps that commit nothing is a bound the app can never
exceed: #78 read a rise of 15 transactions against 37 submits. TxnSubmit is
clickable(enabled = amount.isNotBlank()) and parseCents refuses anything its
regex misses, so a tap whose landing frame shows a refused amount cannot have
committed. Over four recorded android runs that is 19, 11, 25 and 25 of 35,
26, 42 and 42 submit taps.

A relaunch is excepted: a fresh process draws an empty field whatever was
submitted.

* feat(folio): read the amount field into every submit window

Each of the three windows asks whether the tap could have committed, off the
field as the landing frame shows it.

* fix(sidecar): a foreground read that fails degrades the typing guard

An unreadable dumpsys passed a null owner to typeChunks, which switches
the mid-type focus guard off outright and lets the rest of the string
spray into whatever holds the foreground. Fall back to the launched
bundle instead: the guard stays armed, typing still happens, and the
degradation is said out loud rather than assumed away.

* test(sidecar): pin the degraded typing guard both ways

* test(runner): answer Snapshot and Hierarchy off one tree in the fakes

* feat(runner): skip a step whose tree changed between two reads

* test(runner): cover the reread's cost to the existing snapshot rules

* fix(ios): clear app state before the automation session attaches

New performs the clear-state reset, so the uninstall and reinstall no
longer land underneath a live XCTest session that is already bound to
the app. Launch refuses a clear-state request the driver was not built
for rather than reinstalling under its own session.

* fix(ios): the device path clears before its runner session too

* fix(testrun): thread clear-data into the ios drivers

* test(runner): a skipped step must not swallow the action before it

* fix(runner): hold the action back on a step nothing verified

* refactor(runner): drop the empty branch from the hold path

* docs(runner): describe both modes of the composing test driver

* fix(sidecar): erase a field by selecting it, not one delete per character

maestro's eraseText sends one delete per character through its
instrumentation, measured 29.6 ms/char on the API 34 emulator. The
4096-character string the corpus types cost ~121s to clear, a fifth of a
20 minute run spent on one step, and it recurred every time that field
was typed into again.

Select the content and delete the selection instead: two key events at
any length, measured 0.15s to 1.16s for 4096 characters across API 34,
35 and 36. The result is read back off the tree, and a field that is not
empty, or that the tree cannot report on, is finished off per character
in batches rather than assumed clear.

Fixes #80

* test(sidecar): pin the constant-cost erase and its residue check

* fix(sidecar): find the erased field by class, past the keyboard's own focus

The check that decides whether the select-all worked looked for an
"editable" attribute maestro's tree does not carry, so it answered
"cannot tell" every time and every erase paid the per-character
fallback. Worse, an open keyboard puts a second focused node in the
tree, one of the IME's own keys, carrying no text: taking the first
focused node would read a field still holding 4096 characters as empty,
which is the one answer that stops the erase early.

Match the text field by class instead. Measured against the real
backend, 4096 characters now clear in 385ms on API 34, 409ms on API 35
and 870ms on API 36, verified empty, where the fallback took ~4s.

* test(sidecar): use the tree the device really returns

* docs(ci): the android step number describes a local emulator, not ci

the leg disables animations and the number was measured with them on. the
first real dispatch carries 4 transitional steps over 200, so the cross-fade
wait does still fire in ci, just far less often.

* fix(android): say what the sdk lookup checked, not just to set ANDROID_HOME

* fix(doctor): resolve adb and emulator the way a run does

* docs(cli): the android doctor checks are not path-only

* fix(testrun): preflight resolves adb through the sdk, not just PATH

* docs(skills): add a spec review skill and the skills index

* fix(testrun): report a sidecar that dies at startup as the exit it was

* test(testrun): cover the sidecar shutdown path after an early exit

* docs(skills): add a property patterns catalogue skill

* fix(folio): bound the total-balance move instead of demanding it exactly

The write finishes before AddTransactionViewModel navigates, but nothing
establishes that Home's total has re-rendered before the frame is read, and
an equality convicts a healthy app for a total one frame behind. A delta of
zero is exactly the shape nine of the eleven measured android false
convictions had. 2x still exceeds x, so all four recorded convictions
survive, checked against the traces.

The trade is real: a balance that moves by LESS than the amount typed is no
longer judged anywhere in this spec.

* docs(folio): say what property 2 demands now that it is a bound

* docs(skills): add a spec authoring skill

covers hooks, extractors, selectors, properties, actions and the order to write them in, with a complete sample spec that typechecks against the real export surface.

* docs(skills): ground the property patterns catalogue in the merged specs

* docs(skills): name the selector keys that still substring match

* docs(skills): add a setup skill for adopting sanderling

* docs(skills): add a run triage skill

* docs(skills): point the setup skill at its siblings

* docs(manual): correct the flags the cli reference gets wrong

--launcher-activity does not exist in cmd/sanderling/main.go. --device,
--android-app-path and --arm do and were undocumented. runs.md still listed
--max-steps and --exit-on-violation as unshipped, and described --clear-data
as opt-in when the default is already true, contradicting itself ten lines on.

* test(folio): pin that a commit stays in the window until Home reads it

The interaction that keeps a stale Home card list from ever banking counts
the budget has already forgotten: a submit lands on the ledger, so the
reading that resets the window is a whole action later and the submit is
still in it. Characterization, not a regression: no code changed and it
cannot go red first.

* docs(folio): record why a banked card reading can be trusted as current

The freshness rule rests on the app popping one entry back to the ledger,
not on anything the frame carries, so the assumption and the measurements
behind it belong next to it.

* refactor(folio): name the balance property for the bound it asserts

it stopped being an equality and became |delta| <= typed, so the old name
demanded more than the property does. renamed with the ci gate's
GATED_PROPERTIES in the same commit so the gate never sees a name it does
not know.

* fix(android): a refused uninstall must not pass for clear-state

adb uninstall answers Failure [DELETE_FAILED_INTERNAL_ERROR] both when the package was never installed and when it refuses to remove one, so the failure text cannot say which happened and the old code installed over the top either way, keeping the data clear-state was asked to drop. Ask pm path instead, and fall back to pm clear when the app is still there.

* fix(ios): a failed simctl uninstall must fail the reinstall

simctl install over an installed app carries its data container across, so discarding the uninstall error reported a clear-state that never happened. Uninstalling an app that is not installed exits 0 on a booted simulator, so every failure here is a real one.

* fix(ios): a failed devicectl uninstall must fail the reinstall

same hole as the simulator path: devicectl install over an app keeps its data, and the discarded uninstall error hid it. Uninstalling a bundle id that is not installed exits 0 with 'App uninstalled.' on a paired iPhone, so a failure here is always real.

* docs(android): say why the uninstall text cannot be read

* test(android): name the uninstall failure for what it says, not why

* docs(ci): the ios leg convicts on the runner now, and why it did not before

* docs(ci): the cross-fade wait does not fire on ci, say so

* fix(runner): a bounded hold puts the swallow back one step later

the hold carries one action; letting the runner act again while the verifier
is still skipped overwrites it, so the carried action reaches no spec. hold
for as long as the verifier is skipped, and settle on a held step so the
reread pair is not tighter than the window the detector was measured over.

* test(runner): pin what the two reads are compared on

structuralShape excluding text and bounds is the decision separating this
feature from a run that verifies nothing, and only prose held it. adding
either field back now turns a case red.

* fix(ci): close shell injection into the npm publish job

A refname is attacker-controlled and git permits backtick, $, (, ; and |
in it. Three sites substituted it into a run: block, and NODE_AUTH_TOKEN
sat at job level, so a pushed tag ran arbitrary commands with the publish
credential in reach.

The tag now goes through env:, is validated against an anchored version
pattern before anything consumes it, and reaches the other jobs as a job
output. The token is scoped to the publish step. release-npm declares
contents: read instead of inheriting the repo default.

* fix(ios): recognise every shape a blown launch bound arrives in

The runner transport reports a blown budget two ways, its own comment says
so: the context's error once cancellation has landed, and the connection's
i/o timeout when the deadline armed from that context fires first. The
legacy transport reports it as a gRPC status. errors.Is against
context.DeadlineExceeded only matches the first, so the session restart
never fired for the other two and a wedged session stayed wedged.

* test(ios): drive the launch recovery with what the transports return

The wedged-session fake answered with ctx.Err() raw, which is the one
shape the guard already matched. The recovery now runs against the error
each transport really produces for the same expiry, taken from a runner
and a legacy companion that never answer.

* test(runner): pin both guard writes to what the spec reads

deleting lastAction.Relaunched or lastAction.Applied left the whole suite
green, so the only producer of the two fields every spec-side guard reads
had nothing holding it. both now assert the value out of the trace.

* fix(testrun): a run that judged nothing is not a green run

every step skipped means no property ever evaluated, so no violations is the
absence of a verdict rather than a clean one. the hold makes that reachable
now, so the run says it instead of exiting 0.

* fix(ios): stop the app before clearing its state

Launch terminated and then cleared; the clear moved to construction and
left nothing stopping the app first. The container wipe deletes files a
live app still holds open, and the CI ios leg passes no app path so the
wipe is the path it takes. simctl stops it, since the clear now runs
before any automation session exists. On a device the uninstall that is
its only clear takes the running app with it.

* test(ios): pin the stop that has to precede a clear

The ordering probe now records the stop, and a scripted xcrun holds what
reaches the tool: terminate before get_app_container, with the previous
run's files gone after. A simctl terminate that finds nothing to stop
still leaves the clear a success.

* docs(spec): an unbounded eventually is violated at run end

* docs(skills): an unreached eventually convicts at run end

* docs(skills): noUncaughtExceptions only fires on web

* fix(ios): a device clear-state that cannot happen must fail

--clear-data on a physical device with no --ios-app-path warned and then
ran anyway, so the run started on the previous run's data while the flag
said it started clean. There is no data-container wipe on a device, so
there is nothing to fall back to.

* test(ios): a device clear-state without an app path ends the run

* docs(skills): the stock properties each cover one platform

* docs(ci): the balance property demands a bound, not an equality

* fix(ios): the clear-state guard checks the bundle that was cleared

A bool only said that something was cleared, so Launch(ctx, otherBundle,
clearState=true) passed the guard and reported a reset that had reached a
different app. Record what was cleared and compare against the bundle
being launched.

* test(ios): a clear-state launch for an uncleared bundle is refused

* docs(manual): the flagship property is a bound, and say what that costs

* fix(ios): one address picker for every bring-up

bringUpRunner reads the picker from a field, and NewDevice only ever set
the device one, so a device driver that reached bringUpRunner would call
nil. The two fields held the same function; keeping one leaves no path
that can be wired without it.

* test(ios): a device driver can bring a runner up

* docs(skills): both shipped balance forms are bounds now

* docs(skills): name the balance predicate that still exists

* docs(skills): quote the doctor the binary actually prints

* docs(skills): screen= is the chrome driver's url, web only

* docs(skills): substring selector matching is native only

* docs(skills): web selectors are exact, native ones are substrings

* fix(ci): a run that wrote no trace is not evidence about folio

run_dir is empty when the run produced no output directory, and the
fallback made trace ./trace.jsonl. A stray trace in the working directory
was then read as this run's, so a run that wrote nothing reported 'found
the submit bug' and exited 0, defeating the missing-trace check below it.

* fix(ci): fail folio when a gated property is not in the spec

Nothing tied GATED_PROPERTIES to the spec it gates. Renaming a property
left the classifier matching nothing: ios and web blamed the spec for
finding a different bug, and android silently reclassified a real
conviction as 'judging health only' and stayed green.

replay-ui-summary.sh already makes this check for its own list. The spec
path becomes SPEC-overridable the same way, so the check is testable.

* test(ci): cover the folio classifier's verdicts

21 cases through a stubbed sanderling: every exit path, the drift check,
a missing trace, a zero-byte trace, an empty glob and a truncated line.
Asserts the flags that reached the binary, not just the exit code.

Invoked as bash -eo pipefail -c, which is what a run: block does. Running
folio-run.sh itself under -e would kill it at the first non-zero
sanderling test, which is the exit code it exists to read.

* docs(skills): defaultActions bundles five of the eight generators

* test(ios): name the picker test for what it covers

* docs(skills): three of the replay-ui properties are cross-panel

* docs(manual): state.exceptions is web only and reportError does not exist

* fix(folio): the bound carries no unconfirmed-submit guard

deleting confirmedApplied here broke 0 of 355 tests: under a bound a submit
that may not have landed moves the balance by 0, which the bound already
permits, so the guard could only ever drop the double commit it exists to
catch. the relaunch guard stays for a reason the bound does not cover, and
both tests now assert a verdict that changes when their guard does.

* docs(ci): three of the replay-ui properties are cross-panel

* docs(manual): the starter property only fires on web

* test(folio): judge the conjunct on the landings a real run produces

three of the 18 frames the recorded ios run drove it down, each with the
second commit the bound is there to catch. neutering the comparison reddens
it: a second commit on 357900 went unjudged.

* test(folio): the walk drives the composition the spec runs

countSubmitsInWindow never saw an amountText here, so every walk test counted
submits the app must have refused. with the field passed, a refused submit no
longer buys a later double tap an alibi: without it the window reads 3, not 1.

* docs(folio): say which double submit the conjunct can see, and which it cannot

the home landing is the counting invariant's, three of three in the recorded
ios run; this one gets the interleaving whose second pop is cancelled. it is
still the only judge on the 18 ledger landings that run produced.

* docs(folio): the narrow window is not where the detection comes from

the double taps land on home, so the counting form convicts them; what turned
0 convictions into 4 on the recorded ios run is submitCouldCommit, which drops
the windows at those three steps from 5/4/7 to 2/1/2.

* docs(skills): folio drives three platforms from one spec

* fix(folio): an amount over the app's cap spends no window budget

the corpus reaches TxnSubmit with 999999999999999999999, AMOUNT_REGEX takes it
and AddTransactionViewModel refuses it against MAX_TRANSACTION_AMOUNT_CENTS, so
counting it was budget a double submit could hide behind.

* docs(folio): say which form judged one step, not which node was read once

* docs: a bound still needs the relaunch guard, and eventually does convict

* fix(sidecar): the hierarchy rpc serves the tree the snapshot reads

the runner compares the two per step, but snapshot settles and closes a
keyboard while hierarchy was a bare contentDescriptor. measured on emulator
-5556 (api 34) with an ime open: 489 nodes against the snapshot's 134. both
now come off snapshotTree under the same lock; the reread still costs ~75ms
when no keyboard is up.

* docs(runner): say what makes the two reads comparable

the reread's comment claimed the round trip was the only interval between
them; what it left out is that the two rpcs have to read the same way, which
the repo's own android backend did not do.

* refactor(runner): name the settle predicate for what it means

* ci: add a headless-chrome composite action

The setup-chrome / apparmor sysctl / launch-check trio is copied across
three jobs. The old comment described setup-chrome v1 semantics: under v2
stable is the default and the alternative is Chrome for Testing latest,
not a dev Chromium, so it is restated for what the pin actually does.

* ci(examples): add the folio setup actions

folio-app holds the per-platform toolchain and app build, so a caller
guards one step instead of eight. folio-simulator boots the simulator,
installs folio and leaves the app stopped.

* ci(examples): add the replay-ui fixture action

Records a trace and serves it with sanderling replay. The step page URL
is a composite output rather than GITHUB_ENV, so it is scoped to the one
step that drives it.

* ci(examples): one dispatch workflow for every example

folio.yml and replay-ui.yml ran the same operation: build sanderling for
a platform, bring a target up, run a spec against it, classify the trace,
upload the run. They are now one matrix over four examples, each naming
its own runner.

The job is named for what it fuzzes. 'dogfood' named why we run it, not
what runs, the same error as a diagnostic that reports a motivation
instead of an observation.

The matrix is computed by a plan job because jobs.<id>.if cannot read the
matrix context, so a static matrix has no way to leave a leg out. Seeds,
budgets, timeouts, runners and artifact names are unchanged.

* ci: reuse the headless-chrome action in the browser job

Same three steps the examples workflow needs, and the comment explaining
the AppArmor sysctl now lives in one place.

* ci: move the folio jdk step to setup-java v5

The only setup-java left on v4; every other one moved.

* ci: pin third-party actions to commit shas

buf-setup-action was already pinned with a comment saying why; the other
five rode mutable major tags, so a tag move is an unreviewed change to
what runs. Each major currently resolves to the release named in the
comment, so this freezes today's behaviour rather than changing it.

actions/* stay on major tags: they are first-party to the runner.

* ci(replay-ui): name the run directory for what it fuzzes

runs/dogfood and the '### replay-ui dogfood' heading carried the same
naming error as the job name: dogfooding is why the run exists, not what
it fuzzes.

* docs(driver): state the log level scale on LogEntry

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* docs(sidecar): name the device the node counts came off

* fix(chrome): keep a log entry the level scale cannot rank

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* fix(chrome): record console levels on the logcat scale

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* docs(ci): say why upload-pages-artifact needs no include-hidden-files

v3 to v5 crossed v4's change to exclude dot-files. build/site has none,
so nothing was dropped, and the underscore directory is not hidden.

* test(browser): drive a console error through to the spec

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* docs(ioscompanion): name the vacuity behind the empty log slice

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: run the folio classifier's test in make test-ci-scripts

* fix(chrome): keep the message of an object console argument

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(browser): cover console.error with an error object

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* feat(spec): let the runner install state.logs in the page

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* feat(verifier): encode state.logs for the web host

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* feat(chrome): install the step's logs in the page

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(ci): pin three real folio traces from run 31902501859

ios convicted on submitCommitsOneTransactionPerAction, web on both gated
properties, android ran its full 200 steps healthy. Every step is kept;
of each step only step, violations, witnesses and residuals survive.

hierarchy is replaced by the quoted "...Screen" resource ids it held, in
order. It cannot just be dropped: it is 95% of the bytes and also the
only place the android route gate's grep can match, so dropping it flips
that leg from healthy to 'never reached'. 8.4MB to 59KB.

* test(ci): drive the classifier over the real traces

Four cases on real data: each leg's real verdict, plus the android trace
cut before it reached the transaction screen, which is what proves the
route gate reads a real hierarchy dump.

Also corrects the hand-written fixtures. They set is_error to false on a
plain violation; internal/trace/writer.go tags that field omitempty, so a
real trace omits it entirely. Harmless to the classifier, but a fixture
that does not look like reality is the thing that hides drift.

* test(runner): teach the web fakes to take the step's logs

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* fix(runner): install the step's logs before the page extracts

On web every extractor reading is replaced by the one the page computed,
and the page answered logs: [], so noLogcatErrors counted an empty array
however full of errors the console was.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(runner): cover the logs reaching the page and failing to

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(spec): cover the host pushing state.logs into the page

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(browser): drive console.error through to a fired property

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* fix(runner): report a log fetch the driver could not make

The comment claimed the failure was warned about; nothing warned, so a
device whose log fetch failed every step held noLogcatErrors on evidence
nobody collected.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* test(runner): cover the silently dropped log fetch

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* fix(ci): read the route the spec reports, not the hierarchy dump

The android health gate grepped the trace for "AddTransactionScreen",
which occurs in exactly one place: the hierarchy dump, as a resource-id.
That is a debug artifact standing in for a fact the spec already reports,
and it was wrong in both directions. Against the real 8.4MB trace with
hierarchy stripped, the old gate failed a healthy 200-step run; against a
trace carrying the marker on a transition frame without the route ever
being reported, it passed and called it healthy.

It reads extractor_changes.route now, whose values come from SCREENS in
the spec, so the gate and the app agree on what being on a screen means.
routeOf answers null on a frame showing two screens, which is exactly the
frame the marker was matching.

The drift check grows to cover both new names: extract("route") and the
SCREENS key. The fixtures are re-derived keeping the route entry and no
hierarchy at all; the full artifacts and the fixtures give byte-identical
verdicts, which is what proves the coupling is gone.

Script and fixtures move together: either alone leaves the suite red.

* ci: check that the workflow references resolve

actionlint reads a local action's inputs but never checks its path
exists: uses: ./.github/actions/typo lints clean and fails only when the
job runs. Covers composite action paths, make targets including the ones
the examples matrix builds from $SANDERLING, and the scripts a run: step
invokes plus their executable bit.

Fails when it parses fewer references out of a file than that file
mentions, because a checker that matches nothing reports a safety it
never looked for.

* ci: lint the workflows on every pr

The workflows that fuzz the examples are dispatch-only, and GitHub will
not dispatch a workflow that is not on the default branch, so their first
real run is after merge. actionlint and the reference checker are the
only things that can fail before that.

actionlint is pinned by commit, and its tool version is pinned too so a
new release cannot change what CI enforces.

* ci: collapse the four workflows into one

Nine jobs written out one by one, each with its own steps and its own
calibrated seed and budget as literals. Triggers are pull requests, master
and v* tags, and a dispatch with no inputs.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: inline the two composite actions with one caller each

Both existed to give the matrix a per-target hook. folio-app and
headless-chrome stay: three and three callers.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: check that no run: block interpolates an expression

A ${{ }} lands in the script text before bash reads the line, and
actionlint only flags the contexts it already knows are attacker
controlled. Nothing enforced the rule the workflow follows. Also drops the
matrix table lookup, which has no table to read now.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci(folio): name the run, not the fuzzer, in the clean-run message

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* docs: point at the workflow that holds the release secrets now

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: name every job Category (variant), and gate the lot on one check

Follows the convention in antithesishq/bombadil: the display name is what
groups a run in the Actions UI, so Check (tests), Check (browser),
Check (workflows), Folio (android), Folio (ios), Folio (web), Replay UI,
Release and Docs. Every job carries a name, so none of them falls back to
its kebab-case id.

All checks passed needs all nine and runs with if: always(), so branch
protection has one check to point at and a skipped job cannot read as a
pass. Release and docs now gate on startsWith(github.ref, 'refs/tags/v')
alongside master, which is the form the trigger filter already uses.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: split the release job back in two

Collapsing them left the npm publish steps in a job holding contents:
write, because GoReleaser needs it, so npm ci ran its dependency lifecycle
scripts with a write-capable GITHUB_TOKEN in reach of the same job as a
live NPM_TOKEN. Release (npm) is back on contents: read and Release (cli)
keeps contents: write, which is what they each had before.

Each validates the tag from its own copy of the pattern rather than
waiting on a job that exists only to pass a string. Release (cli) is tags
only: there is no CLI to cut on a merge.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: run folio on pull requests

folio was skipped on pull requests, so ios, android and web only ever ran
after a merge. The three legs are 3 to 19 minutes and run in parallel, and
a superseded pull request run already cancels itself.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ

* ci: draw each group as its own box in the run graph

The run graph boxes jobs together when they share the same dependencies and
the same dependents. All ten jobs fed only all-checks-passed, so all ten drew
as one pile. A gate per group gives each group a dependent that is exactly
that group.

Release and docs now need the checks, which they should have all along: npm
publish and the pages deploy ran on a merge without waiting for the test job.
Folio stays unblocked so a 20 minute leg does not wait on a 3 minute one.

Claude-Session: https://claude.ai/code/session_01ShuAy8q8ZfPi8KHxwc8JpQ
This commit is contained in:
pj authored and GitHub committed 2026-08-16 13:23:53 +05:30
1 parent 11f72a722a
commit 9fb121e9d0
117 files changed
+11108 -1262

No files matched your search

+83 -13
View File
@@ -123,14 +123,17 @@ func antiFreezeCommands() [][]string {
// reinstalling it. This replaces `pm clear` for clear-state: ColorOS and other
// hardened OEM builds deny CLEAR_APP_USER_DATA even to the adb shell user, so a
// clear aborts the launch, whereas uninstall+install is always permitted.
// The uninstall is best effort so a not-installed app is not an error.
// A failed uninstall is not passed over: `install -r` keeps the app's data, so
// the reinstall would report a clear-state that never happened.
func ReinstallApp(ctx context.Context, serial, bundleID, apkPath string, stdout io.Writer) error {
adb, err := AdbBinary()
if err != nil {
return err
}
if output, err := exec.CommandContext(ctx, adb, adbArgs(serial, "uninstall", bundleID)...).CombinedOutput(); err != nil {
fmt.Fprintf(stdout, "clear-state: uninstall %s skipped (%v: %s)\n", bundleID, err, strings.TrimSpace(string(output)))
if err := clearDataUninstallLeftBehind(ctx, adb, serial, bundleID, strings.TrimSpace(string(output)), stdout); err != nil {
return err
}
}
if output, err := exec.CommandContext(ctx, adb, adbArgs(serial, "install", "-r", apkPath)...).CombinedOutput(); err != nil {
return fmt.Errorf("install %s: %w: %s", apkPath, err, strings.TrimSpace(string(output)))
@@ -138,6 +141,37 @@ func ReinstallApp(ctx context.Context, serial, bundleID, apkPath string, stdout
return nil
}
// clearDataUninstallLeftBehind reaches first-launch state after `adb uninstall`
// failed. The failure text cannot say why: an API 34 emulator answers
// "Failure [DELETE_FAILED_INTERNAL_ERROR]" both for a package that was never
// installed and for one it refuses to remove. So ask the package manager which
// happened. Nothing installed means nothing to clear. Still installed
// means the data survives the reinstall, and `pm clear` is the one remaining
// way to reach first-launch state; when that fails too, so does clear-state.
func clearDataUninstallLeftBehind(ctx context.Context, adb, serial, bundleID, uninstallOutput string, stdout io.Writer) error {
if !packageInstalled(ctx, adb, serial, bundleID) {
return nil
}
output, err := exec.CommandContext(ctx, adb, adbArgs(serial, "shell", "pm", "clear", bundleID)...).CombinedOutput()
cleared := strings.TrimSpace(string(output))
if err != nil || !strings.Contains(cleared, "Success") {
return fmt.Errorf(
"clear-state: %s is still installed after `adb uninstall` said %q, and `pm clear` said %q: its data was not cleared",
bundleID, uninstallOutput, cleared,
)
}
fmt.Fprintf(stdout, "clear-state: uninstall %s said %q and left it installed; cleared its data with `pm clear` instead\n", bundleID, uninstallOutput)
return nil
}
// packageInstalled reports whether the package manager resolves an APK path for
// bundleID. The printed path is the signal rather than the exit status, which
// `adb shell` does not forward from devices below API 24.
func packageInstalled(ctx context.Context, adb, serial, bundleID string) bool {
output, _ := exec.CommandContext(ctx, adb, adbArgs(serial, "shell", "pm", "path", bundleID)...).Output()
return strings.HasPrefix(strings.TrimSpace(string(output)), "package:")
}
const threeButtonNavOverlay = "com.android.internal.systemui.navbar.threebutton"
// navModeOverlays are the system navigation-mode overlays. Only one is active at
@@ -221,11 +255,7 @@ func navOverlayCommand(ctx context.Context, adb, serial, overlay string) *exec.C
// EnvWithAndroidPlatformTools returns env with the directory containing adb
// prepended to PATH, so child processes (the sidecar) can invoke adb even
// when the user hasn't set up their shell PATH.
func EnvWithAndroidPlatformTools(env []string) []string {
adb, err := AdbBinary()
if err != nil {
return env
}
func EnvWithAndroidPlatformTools(env []string, adb string) []string {
adbDir := filepath.Dir(adb)
result := make([]string, 0, len(env))
found := false
@@ -247,7 +277,9 @@ func EnvWithAndroidPlatformTools(env []string) []string {
// AdbBinary locates the adb binary via PATH or known Android SDK locations.
func AdbBinary() (string, error) { return findAndroidTool("adb", "platform-tools") }
func emulatorBinary() (string, error) { return findAndroidTool("emulator", "emulator") }
// EmulatorBinary locates the emulator binary via PATH or known Android SDK
// locations.
func EmulatorBinary() (string, error) { return findAndroidTool("emulator", "emulator") }
// findAndroidTool locates a binary from the Android SDK. It checks PATH,
// then $ANDROID_HOME/<subdir>/<name> and $ANDROID_SDK_ROOT/<subdir>/<name>,
@@ -264,7 +296,36 @@ func findAndroidTool(name, subdir string) (string, error) {
}
tried = append(tried, candidate)
}
return "", fmt.Errorf("could not locate %q: not on PATH and not under any known Android SDK root (set $ANDROID_HOME to point at your SDK; tried %v)", name, tried)
return "", fmt.Errorf(
"%s not found: not on PATH, and not at [%s]; %s\nput %s on PATH, or point $ANDROID_HOME at an Android SDK that has %s/%s",
name, strings.Join(tried, ", "), sdkRootStatus(), name, subdir, name,
)
}
// sdkRootStatus reports what the SDK root variables hold, so a lookup failure
// says whether they were unset or pointed somewhere that is not an SDK instead
// of leaving the reader to work out which from a list of paths.
func sdkRootStatus() string {
var reported []string
for _, variable := range []string{"ANDROID_HOME", "ANDROID_SDK_ROOT"} {
value := os.Getenv(variable)
if value == "" {
continue
}
info, err := os.Stat(value)
switch {
case err != nil:
reported = append(reported, fmt.Sprintf("$%s=%s does not exist", variable, value))
case !info.IsDir():
reported = append(reported, fmt.Sprintf("$%s=%s is not a directory", variable, value))
default:
reported = append(reported, fmt.Sprintf("$%s=%s", variable, value))
}
}
if len(reported) == 0 {
return "$ANDROID_HOME and $ANDROID_SDK_ROOT are unset"
}
return strings.Join(reported, ", ")
}
func androidSDKCandidates() []string {
@@ -283,11 +344,20 @@ func androidSDKCandidates() []string {
addRoot(filepath.Join(home, "Library", "Android", "sdk"))
addRoot(filepath.Join(home, "Android", "Sdk"))
}
addRoot("/opt/homebrew/share/android-commandlinetools")
addRoot("/usr/local/share/android-commandlinetools")
for _, root := range standardSDKRoots {
addRoot(root)
}
return roots
}
// standardSDKRoots are the install locations checked after the environment and
// the home directory. A var so a resolution test can point it at a fixture
// instead of whatever SDK the host running the test happens to have.
var standardSDKRoots = []string{
"/opt/homebrew/share/android-commandlinetools",
"/usr/local/share/android-commandlinetools",
}
func listAdbDevices(ctx context.Context) ([]string, error) {
adb, err := AdbBinary()
if err != nil {
@@ -317,7 +387,7 @@ func parseAdbDevices(output string) []string {
}
func listAVDs(ctx context.Context) ([]string, error) {
emulator, err := emulatorBinary()
emulator, err := EmulatorBinary()
if err != nil {
return nil, err
}
@@ -382,7 +452,7 @@ func pickAVD(requested string, available []string) (string, error) {
}
func bootAVD(_ context.Context, name string) error {
emulator, err := emulatorBinary()
emulator, err := EmulatorBinary()
if err != nil {
return err
}
+241
View File
@@ -1,6 +1,8 @@
package android
import (
"os"
"path/filepath"
"reflect"
"slices"
"strings"
@@ -136,6 +138,114 @@ func TestPickAVD_NoneAvailable(t *testing.T) {
}
}
// fakeSDK writes an SDK layout holding only the named tools ("emulator/emulator"),
// so a lookup test never resolves against the host's own SDK.
func fakeSDK(t *testing.T, tools ...string) string {
t.Helper()
root := t.TempDir()
for _, tool := range tools {
path := filepath.Join(root, filepath.FromSlash(tool))
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatalf("mkdir %s: %v", filepath.Dir(path), err)
}
if err := os.WriteFile(path, nil, 0o755); err != nil {
t.Fatalf("write %s: %v", path, err)
}
}
return root
}
// isolateSDKLookup closes every route to the host's own SDK: an empty PATH, no
// root variables, a home with nothing under it, and the standard install
// locations replaced by the given roots. It returns the fake home directory.
func isolateSDKLookup(t *testing.T, roots ...string) string {
t.Helper()
home := t.TempDir()
t.Setenv("PATH", t.TempDir())
t.Setenv("ANDROID_HOME", "")
t.Setenv("ANDROID_SDK_ROOT", "")
t.Setenv("HOME", home)
original := standardSDKRoots
t.Cleanup(func() { standardSDKRoots = original })
standardSDKRoots = roots
return home
}
func TestAdbBinary_ResolvesUnderStandardSDKRootWithAndroidHomeUnset(t *testing.T) {
sdk := fakeSDK(t, "platform-tools/adb")
isolateSDKLookup(t, sdk)
adb, err := AdbBinary()
if err != nil {
t.Fatalf("AdbBinary: %v", err)
}
if want := filepath.Join(sdk, "platform-tools", "adb"); adb != want {
t.Errorf("AdbBinary() = %q, want %q", adb, want)
}
}
func TestEmulatorBinary_ResolvesUnderStandardSDKRootWithAndroidHomeUnset(t *testing.T) {
sdk := fakeSDK(t, "emulator/emulator")
isolateSDKLookup(t, sdk)
emulator, err := EmulatorBinary()
if err != nil {
t.Fatalf("EmulatorBinary: %v", err)
}
if want := filepath.Join(sdk, "emulator", "emulator"); emulator != want {
t.Errorf("EmulatorBinary() = %q, want %q", emulator, want)
}
}
func TestAdbBinary_UsesAndroidHome(t *testing.T) {
sdk := fakeSDK(t, "platform-tools/adb")
isolateSDKLookup(t)
t.Setenv("ANDROID_HOME", sdk)
adb, err := AdbBinary()
if err != nil {
t.Fatalf("AdbBinary: %v", err)
}
if want := filepath.Join(sdk, "platform-tools", "adb"); adb != want {
t.Errorf("AdbBinary() = %q, want %q", adb, want)
}
}
func TestAdbBinary_NoSDKAnywhereReportsWhereItLooked(t *testing.T) {
empty := t.TempDir()
home := isolateSDKLookup(t, empty)
_, err := AdbBinary()
if err == nil {
t.Fatal("expected an error with no SDK anywhere")
}
for _, want := range []string{
"adb not found: not on PATH",
filepath.Join(home, "Library", "Android", "sdk", "platform-tools", "adb"),
filepath.Join(empty, "platform-tools", "adb"),
"$ANDROID_HOME and $ANDROID_SDK_ROOT are unset",
"put adb on PATH, or point $ANDROID_HOME at an Android SDK that has platform-tools/adb",
} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q missing %q", err, want)
}
}
}
func TestAdbBinary_AndroidHomePointingNowhereIsNamed(t *testing.T) {
isolateSDKLookup(t, t.TempDir())
missing := filepath.Join(t.TempDir(), "no-such-sdk")
t.Setenv("ANDROID_HOME", missing)
_, err := AdbBinary()
if err == nil {
t.Fatal("expected an error when ANDROID_HOME points nowhere")
}
if want := "$ANDROID_HOME=" + missing + " does not exist"; !strings.Contains(err.Error(), want) {
t.Errorf("error %q must say %q rather than leaving it in the tried list", err, want)
}
}
func TestPathContains(t *testing.T) {
path := "/usr/bin:/opt/tools:/usr/local/bin"
if !pathContains(path, "/opt/tools") {
@@ -331,3 +441,134 @@ func TestNavModeToRestore(t *testing.T) {
}
}
}
// scriptedAdb puts an adb under a fake SDK root that logs each invocation's
// arguments and answers from replies, a `case "$*" in` body. SDK lookup is
// isolated onto that root, so ReinstallApp runs its real command sequence
// against the script and the log holds what reached adb.
func scriptedAdb(t *testing.T, replies string) string {
t.Helper()
root := t.TempDir()
log := filepath.Join(root, "adb.log")
adb := filepath.Join(root, "platform-tools", "adb")
if err := os.MkdirAll(filepath.Dir(adb), 0o755); err != nil {
t.Fatalf("mkdir %s: %v", filepath.Dir(adb), err)
}
script := "#!/bin/sh\necho \"$*\" >> " + log + "\ncase \"$*\" in\n" + replies + "\nesac\n"
if err := os.WriteFile(adb, []byte(script), 0o755); err != nil {
t.Fatalf("write %s: %v", adb, err)
}
isolateSDKLookup(t, root)
return log
}
func adbCalls(t *testing.T, log string) []string {
t.Helper()
contents, err := os.ReadFile(log)
if err != nil {
if os.IsNotExist(err) {
return nil
}
t.Fatalf("read %s: %v", log, err)
}
return strings.Split(strings.TrimSpace(string(contents)), "\n")
}
const (
uninstallFails = `"uninstall "*) echo "Failure [DELETE_FAILED_INTERNAL_ERROR]"; exit 1;;`
stillInstalled = `"shell pm path "*) echo "package:/data/app/app.example-1/base.apk";;`
notInstalled = `"shell pm path "*) exit 1;;`
installSucceeds = `"install "*) echo "Success";;`
)
func TestReinstallApp_RefusedUninstallClearsTheDataItLeftBehind(t *testing.T) {
log := scriptedAdb(t, strings.Join([]string{
uninstallFails,
stillInstalled,
`"shell pm clear "*) echo "Success";;`,
installSucceeds,
}, "\n"))
output := &strings.Builder{}
if err := ReinstallApp(t.Context(), "", "app.example", "/tmp/app.apk", output); err != nil {
t.Fatalf("ReinstallApp: %v", err)
}
want := []string{
"uninstall app.example",
"shell pm path app.example",
"shell pm clear app.example",
"install -r /tmp/app.apk",
}
if got := adbCalls(t, log); !slices.Equal(got, want) {
t.Fatalf("adb calls = %v, want %v", got, want)
}
if !strings.Contains(output.String(), "pm clear") {
t.Errorf("output %q does not say the data was cleared some other way", output.String())
}
}
func TestReinstallApp_RefusedUninstallThatCannotBeClearedIsFatal(t *testing.T) {
log := scriptedAdb(t, strings.Join([]string{
uninstallFails,
stillInstalled,
`"shell pm clear "*) echo "Failed"; exit 1;;`,
installSucceeds,
}, "\n"))
err := ReinstallApp(t.Context(), "", "app.example", "/tmp/app.apk", &strings.Builder{})
if err == nil {
t.Fatal("ReinstallApp reported success while app.example kept the data clear-state was asked to remove")
}
for _, want := range []string{"app.example", "DELETE_FAILED_INTERNAL_ERROR", "Failed"} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q does not quote %q", err, want)
}
}
if slices.Contains(adbCalls(t, log), "install -r /tmp/app.apk") {
t.Error("installed over an app whose data survived, which is the reinstall reporting a clear it did not perform")
}
}
func TestReinstallApp_UninstallFailureWithNothingInstalledIsQuiet(t *testing.T) {
log := scriptedAdb(t, strings.Join([]string{
uninstallFails,
notInstalled,
installSucceeds,
}, "\n"))
output := &strings.Builder{}
if err := ReinstallApp(t.Context(), "", "app.example", "/tmp/app.apk", output); err != nil {
t.Fatalf("ReinstallApp: %v", err)
}
calls := adbCalls(t, log)
if !slices.Contains(calls, "install -r /tmp/app.apk") {
t.Fatalf("adb calls = %v, want the install to go ahead: a first run has no app to uninstall", calls)
}
if slices.Contains(calls, "shell pm clear app.example") {
t.Errorf("adb calls = %v, want no data clear: there was no app holding data", calls)
}
if output.String() != "" {
t.Errorf("output = %q, want nothing: a first run has no app to uninstall", output.String())
}
}
func TestReinstallApp_SuccessfulUninstallNeedsNoFallback(t *testing.T) {
log := scriptedAdb(t, strings.Join([]string{
`"uninstall "*) echo "Success";;`,
stillInstalled,
`"shell pm clear "*) echo "Success";;`,
installSucceeds,
}, "\n"))
if err := ReinstallApp(t.Context(), "", "app.example", "/tmp/app.apk", &strings.Builder{}); err != nil {
t.Fatalf("ReinstallApp: %v", err)
}
want := []string{"uninstall app.example", "install -r /tmp/app.apk"}
if got := adbCalls(t, log); !slices.Equal(got, want) {
t.Fatalf("adb calls = %v, want %v", got, want)
}
}
+64 -12
View File
@@ -70,23 +70,27 @@ func New() *Driver {
}
var parts []string
for _, arg := range e.Args {
if arg.Value != nil {
var s string
if err := json.Unmarshal(arg.Value, &s); err == nil {
parts = append(parts, s)
} else {
parts = append(parts, string(arg.Value))
// An object argument, which is what console.error(err) passes,
// carries no value at all: CDP sends a description instead. Reading
// only the value logged those calls with an empty message, so the
// entry named a level and nothing a reader could act on.
if arg.Value == nil {
if arg.Description != "" {
parts = append(parts, arg.Description)
}
continue
}
var s string
if err := json.Unmarshal(arg.Value, &s); err == nil {
parts = append(parts, s)
} else {
parts = append(parts, string(arg.Value))
}
}
level := strings.ToUpper(string(e.Type))
if level == "LOG" {
level = "I"
}
d.logsMu.Lock()
d.logs = append(d.logs, driver.LogEntry{
UnixMillis: int64(e.Timestamp.Time().UnixMilli()),
Level: level,
Level: consoleLevel(e.Type),
Tag: "console",
Message: strings.Join(parts, " "),
})
@@ -712,9 +716,33 @@ func (d *Driver) Metrics(ctx context.Context, _ string) (driver.Metrics, error)
}, nil
}
// consoleLevel places a console call on driver.LogEntry's logcat scale. The
// verbs a spec acts on are all named here; the rest are info rather than "E"
// because promoting them would convict an app of an error it never logged.
func consoleLevel(apiType runtime.APIType) string {
switch apiType {
case runtime.APITypeError, runtime.APITypeAssert:
return "E"
case runtime.APITypeWarning:
return "W"
case runtime.APITypeDebug:
return "D"
default:
return "I"
}
}
// meetsLevel keeps an entry whose level the scale cannot rank. Ranking an
// unknown level below every threshold drops it, and a dropped entry is
// indistinguishable from a quiet app: the caller sees silence and reports it as
// health.
func meetsLevel(level, minLevel string) bool {
order := map[string]int{"V": 0, "D": 1, "I": 2, "W": 3, "E": 4, "F": 5}
return order[level] >= order[minLevel]
rank, ranked := order[level]
if !ranked {
return true
}
return rank >= order[minLevel]
}
func pngDimensions(png []byte) (int, int) {
@@ -859,6 +887,30 @@ func (d *Driver) SetLastAction(ctx context.Context, encoded json.RawMessage) err
return nil
}
// SetLogs installs the entries this step's log fetch returned as state.logs
// inside the page runtime. The page cannot derive them: console output reaches
// the driver over CDP and nothing in the page reads it back. Without this call
// every web state.logs is empty, and since the page's reading of an extractor
// replaces the host's, the default noLogcatErrors then reports green on a run
// whose console was full of errors.
//
// Unguarded for the same reason as SetLastAction: on a page with no setter,
// "the page cannot accept logs" has to fail the run rather than be reported as
// a successful install.
func (d *Driver) SetLogs(ctx context.Context, encoded json.RawMessage) error {
payload := strings.TrimSpace(string(encoded))
if payload == "" {
payload = "[]"
}
script := fmt.Sprintf(`window.__sanderlingSetLogs__(%s)`, payload)
runCtx, cancel := d.runCtx(ctx)
defer cancel()
if err := chromedp.Run(runCtx, chromedp.Evaluate(script, nil)); err != nil {
return fmt.Errorf("set logs: %w", err)
}
return nil
}
// extractorScript resolves the extractor table once the page is not mid route
// transition, giving up on that wait after %d ms.
//
+48
View File
@@ -860,6 +860,54 @@ func TestSetLastAction_ReportsAPageThatCannotTakeIt(t *testing.T) {
}
}
// TestSetLogs_ReportsAPageThatCannotTakeThem is the same install on the channel
// the log properties hang off. The driver holding a console error changes
// nothing on web: the page's reading of every extractor replaces the host's, so
// unless the entries are put back into the page, noLogcatErrors counts an empty
// array and stays green through a run full of errors.
func TestSetLogs_ReportsAPageThatCannotTakeThem(t *testing.T) {
const withSetter = `<body><script>
window.__logsSeen = null;
window.__sanderlingSetLogs__ = function (value) { window.__logsSeen = value; };
</script></body>`
const withoutSetter = `<body><div id="app">no sanderling runtime here</div></body>`
pages := map[string]string{"/with": withSetter, "/without": withoutSetter}
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "text/html")
_, _ = w.Write([]byte(pages[r.URL.Path]))
}))
defer server.Close()
d := New()
defer d.Terminate(context.Background())
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := d.Launch(ctx, server.URL+"/with", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
logs := json.RawMessage(`[{"unixMillis":1,"level":"E","tag":"console","message":"boom"}]`)
if err := d.SetLogs(ctx, logs); err != nil {
t.Fatalf("SetLogs on a page that defines the setter: %v", err)
}
var seen []map[string]any
if err := chromedp.Run(d.tabCtx,
chromedp.Evaluate(`window.__logsSeen`, &seen)); err != nil {
t.Fatalf("read installed logs: %v", err)
}
if len(seen) != 1 || seen[0]["level"] != "E" || seen[0]["message"] != "boom" {
t.Errorf("the page received %v, want the error-level entry the driver captured", seen)
}
if err := d.Launch(ctx, server.URL+"/without", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
if err := d.SetLogs(ctx, logs); err == nil {
t.Error("SetLogs reported success on a page with no setter; " +
"a runtime that cannot take the step's logs is indistinguishable from one that did")
}
}
// TestEvaluateExtractors_ReportsAMissingTable is the same failure on the other
// sampler. An empty override map is what a spec with no extractors returns, so
// treating a missing table as {} makes "this page has no sanderling runtime"
+60
View File
@@ -0,0 +1,60 @@
package chrome
import (
"testing"
"github.com/chromedp/cdproto/runtime"
)
// Every console verb has to land on the logcat scale driver.LogEntry declares:
// the runner fetches at "E" and the default properties count entries whose
// level equals "E", so a level spelled any other way is an error the spec never
// sees. A verb with no mapping is info, which is honest about severity without
// fabricating an error the page never logged.
func TestConsoleLevel(t *testing.T) {
cases := map[runtime.APIType]string{
runtime.APITypeError: "E",
runtime.APITypeAssert: "E",
runtime.APITypeWarning: "W",
runtime.APITypeDebug: "D",
runtime.APITypeLog: "I",
runtime.APITypeInfo: "I",
runtime.APITypeTable: "I",
runtime.APIType("countReset"): "I",
}
for apiType, want := range cases {
if got := consoleLevel(apiType); got != want {
t.Errorf("consoleLevel(%q) = %q, want %q", apiType, got, want)
}
}
}
// A level the scale does not name is unknown, not verbose. Ranking it below
// every threshold is what silently emptied the web log channel: the entries
// existed, the filter dropped them, and the run reported nothing. Evidence the
// filter cannot rank has to reach the caller, who can at least see it.
func TestMeetsLevel(t *testing.T) {
cases := []struct {
level string
minLevel string
want bool
}{
{"E", "E", true},
{"F", "E", true},
{"W", "E", false},
{"I", "E", false},
{"D", "E", false},
{"V", "E", false},
{"W", "W", true},
{"I", "W", false},
{"D", "V", true},
{"ERROR", "E", true},
{"WARNING", "E", true},
{"", "E", true},
}
for _, tc := range cases {
if got := meetsLevel(tc.level, tc.minLevel); got != tc.want {
t.Errorf("meetsLevel(%q, %q) = %v, want %v", tc.level, tc.minLevel, got, tc.want)
}
}
}
+6
View File
@@ -82,6 +82,12 @@ type FocusedWindowChecker interface {
FocusedWindowApp(ctx context.Context) (string, error)
}
// LogEntry is one line of device log. Level is logcat's single-letter scale on
// every platform: "V", "D", "I", "W", "E", "F", ordered as written. The runner
// fetches at "E" and the default properties count entries whose level equals
// "E", so a driver that spells a level any other way empties the channel
// without failing anything: the entries never arrive and every property reading
// state.logs holds vacuously.
type LogEntry struct {
UnixMillis int64
Level string
+46 -19
View File
@@ -31,17 +31,22 @@ type DeviceOptions struct {
BundleID string
// AppPath is the .app bundle installed via devicectl for clear-state.
AppPath string
// ClearState reinstalls the app while NewDevice runs, before the runner's
// test session exists. Clear state is a property of the driver rather than
// of a launch: see Launch.
ClearState bool
// Output receives the runner session log path and driver warnings.
Output io.Writer
// DoubleTapGapMilliseconds overrides the synthesized double-tap gap.
DoubleTapGapMilliseconds float64
// Test seams. Production leaves them nil and NewDevice wires the real
// build/spawn/tunnel/dial.
spawnRunner func(ctx context.Context, address string) (*exec.Cmd, error)
startTunnel func(ctx context.Context, hardwareUDID, localAddress, devicePort string) (io.Closer, error)
dialRunner func(address string) (transport.Companion, error)
pickAddress func() (string, error)
// build/spawn/tunnel/dial/devicectl.
spawnRunner func(ctx context.Context, address string) (*exec.Cmd, error)
startTunnel func(ctx context.Context, hardwareUDID, localAddress, devicePort string) (io.Closer, error)
dialRunner func(address string) (transport.Companion, error)
pickAddress func() (string, error)
reinstallApp func(ctx context.Context) error
}
// deviceStartupTimeout bounds the runner's startup once its hosting test
@@ -61,6 +66,9 @@ func NewDevice(ctx context.Context, options DeviceOptions) (*Driver, error) {
if options.CoreDeviceID == "" {
return nil, errors.New("ios device: CoreDeviceID is required")
}
if options.ClearState && options.BundleID == "" {
return nil, errors.New("ios device: clear-state needs BundleID: there is nothing to uninstall without it")
}
output := options.Output
if output == nil {
output = io.Discard
@@ -95,17 +103,24 @@ func NewDevice(ctx context.Context, options DeviceOptions) (*Driver, error) {
}
}
if options.pickAddress != nil {
d.pickDeviceAddress = options.pickAddress
d.pickRunnerAddress = options.pickAddress
} else {
d.pickDeviceAddress = pickLoopbackAddress
d.pickRunnerAddress = pickLoopbackAddress
}
// Device seams: clear-state reinstalls via devicectl; the container reset and
// paste grant are simulator-only and become no-ops. The runner types
// natively, so no paste prompt is ever hit.
d.reinstallApp = d.devicectlReinstall
// natively, so no paste prompt is ever hit. Stopping the app before the
// clear is a no-op too: devicectl addresses processes by pid rather than by
// bundle, and the uninstall that is the device's only clear takes the
// running app with it, which is what a terminate here would be for.
d.reinstallApp = options.reinstallApp
if d.reinstallApp == nil {
d.reinstallApp = d.devicectlReinstall
}
d.resetContainer = d.deviceResetContainerUnsupported
d.grantPaste = func(context.Context) error { return nil }
d.terminateApp = func(context.Context) error { return nil }
d.restart = d.respawnDevice
d.processContext, d.processCancel = context.WithCancel(ctx)
@@ -116,6 +131,14 @@ func NewDevice(ctx context.Context, options DeviceOptions) (*Driver, error) {
}
d.deviceLock = lock
if options.ClearState {
if err := d.clearAppState(ctx); err != nil {
d.Close()
return nil, err
}
d.clearedBundleID = options.BundleID
}
if err := d.bringUpDevice(ctx); err != nil {
d.Close()
return nil, err
@@ -136,7 +159,7 @@ func NewDevice(ctx context.Context, options DeviceOptions) (*Driver, error) {
// health. The build runs inside spawnRunner under the process context, so the
// startup timeout only bounds the post-spawn wait, not the build.
func (d *Driver) bringUpDevice(ctx context.Context) error {
address, err := d.pickDeviceAddress()
address, err := d.pickRunnerAddress()
if err != nil {
return err
}
@@ -205,8 +228,14 @@ func (d *Driver) respawnDevice(ctx context.Context) error {
// devicectlReinstall uninstalls then installs the app bundle via devicectl,
// keyed on the CoreDevice id. App lifecycle stays with devicectl: the runner's
// own install path is simulator-specific.
// A failed uninstall ends the reinstall: installing over an app keeps its data,
// so clear-state would be reported without happening. Uninstalling an app that
// is not installed exits 0 ("App uninstalled." on a paired iPhone running iOS
// 26.5), so there is no benign failure here to sort out from a real one.
func (d *Driver) devicectlReinstall(ctx context.Context) error {
_ = exec.CommandContext(ctx, "xcrun", "devicectl", "device", "uninstall", "app", "--device", d.coreDeviceID, d.bundleID).Run()
if output, err := exec.CommandContext(ctx, "xcrun", "devicectl", "device", "uninstall", "app", "--device", d.coreDeviceID, d.bundleID).CombinedOutput(); err != nil {
return fmt.Errorf("devicectl uninstall %s: %w: %s", d.bundleID, err, strings.TrimSpace(string(output)))
}
output, err := exec.CommandContext(ctx, "xcrun", "devicectl", "device", "install", "app", "--device", d.coreDeviceID, d.appPath).CombinedOutput()
if err != nil {
return fmt.Errorf("devicectl install: %w: %s", err, strings.TrimSpace(string(output)))
@@ -214,13 +243,11 @@ func (d *Driver) devicectlReinstall(ctx context.Context) error {
return nil
}
// deviceResetContainerUnsupported warns once that device clear-state needs an
// app path for a devicectl reinstall: there is no simulator-style data-container
// wipe on a physical device.
// deviceResetContainerUnsupported ends the run: there is no simulator-style
// data-container wipe on a physical device, so a clear-state with no app path
// to reinstall from cannot happen. Warning and carrying on hands the run every
// previous run's data while the flag says it started clean.
func (d *Driver) deviceResetContainerUnsupported(context.Context) error {
if !d.clearStateWarned {
fmt.Fprintln(d.output, "clear-state on a physical device requires --ios-app-path for a reinstall; skipping (state not cleared)")
d.clearStateWarned = true
}
return nil
return errors.New("clear-state on a physical device requires --ios-app-path for a reinstall; " +
"there is no data-container wipe on a device, so the run would start on the previous run's state")
}
+113 -9
View File
@@ -6,6 +6,8 @@ import (
"io"
"net"
"os/exec"
"slices"
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/transport"
@@ -81,6 +83,25 @@ func TestNewDeviceWiresRunnerOnlyMode(t *testing.T) {
}
}
// TestNewDeviceWiresTheAddressPickerEveryBringUpUses keeps the device driver
// whole. bringUpRunner reads its picker from a field rather than calling the
// package function, and NewDevice left that field nil, so the only thing
// standing between a device run and a nil call was which restart path ran.
func TestNewDeviceWiresTheAddressPickerEveryBringUpUses(t *testing.T) {
address := startLoopbackListener(t)
options := testDeviceOptions(address, newDeviceCompanion())
options.HardwareUDID = "00008140-PICKER"
d, err := NewDevice(context.Background(), options)
if err != nil {
t.Fatalf("NewDevice: %v", err)
}
defer d.Close()
if err := d.bringUpRunner(context.Background()); err != nil {
t.Fatalf("bringUpRunner: %v", err)
}
}
func TestNewDeviceRequiresIdentifiers(t *testing.T) {
if _, err := NewDevice(context.Background(), DeviceOptions{CoreDeviceID: "x"}); err == nil {
t.Fatal("missing HardwareUDID must error")
@@ -152,17 +173,100 @@ func TestDeviceEraseAndPressKeyRouteThroughEditor(t *testing.T) {
}
}
func TestDeviceClearStateWithoutAppPathWarnsOnce(t *testing.T) {
output := &bytes.Buffer{}
d := &Driver{output: output, deviceMode: true}
d.resetContainer = d.deviceResetContainerUnsupported
for i := 0; i < 2; i++ {
if err := d.deviceResetContainerUnsupported(context.Background()); err != nil {
t.Fatal(err)
func TestNewDeviceReinstallsOnceBeforeTheRunnerSession(t *testing.T) {
address := startLoopbackListener(t)
probe := &clearStateProbe{}
options := testDeviceOptions(address, newDeviceCompanion())
options.HardwareUDID = "00008140-CLEAR"
options.AppPath = "/tmp/Sample.app"
options.ClearState = true
options.reinstallApp = func(context.Context) error { probe.record("reinstall"); return nil }
spawn := options.spawnRunner
options.spawnRunner = func(ctx context.Context, runnerAddress string) (*exec.Cmd, error) {
probe.record("runner session")
return spawn(ctx, runnerAddress)
}
d, err := NewDevice(context.Background(), options)
if err != nil {
t.Fatalf("NewDevice: %v", err)
}
defer d.Close()
if err := d.Launch(context.Background(), "", true, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
want := []string{"reinstall", "runner session"}
if got := probe.recorded(); !slices.Equal(got, want) {
t.Fatalf("calls = %v, want %v: devicectl must reinstall once, before the runner's test session attaches", got, want)
}
}
func TestDevicectlReinstallStopsWhenTheUninstallFails(t *testing.T) {
log := scriptedXcrun(t, `"devicectl device uninstall "*) echo "ERROR: Internal logic error: Connection was invalidated"; exit 1;;
"devicectl device install "*) :;;`)
d := &Driver{coreDeviceID: "CORE-DEVICE", bundleID: "app.example", appPath: "/tmp/Sample.app"}
err := d.devicectlReinstall(context.Background())
if err == nil {
t.Fatal("devicectlReinstall reported success while app.example kept the data clear-state was asked to remove")
}
for _, want := range []string{"app.example", "Connection was invalidated"} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q does not quote %q", err, want)
}
}
if got := bytes.Count(output.Bytes(), []byte("requires --ios-app-path")); got != 1 {
t.Fatalf("warning emitted %d times, want once", got)
calls := xcrunCalls(t, log)
if slices.ContainsFunc(calls, func(call string) bool { return strings.HasPrefix(call, "devicectl device install") }) {
t.Errorf("xcrun calls = %v: installing over the app carries its data into the run", calls)
}
}
func TestDevicectlReinstallProceedsWhenNothingIsInstalled(t *testing.T) {
log := scriptedXcrun(t, `"devicectl device uninstall "*) echo "App uninstalled.";;
"devicectl device install "*) :;;`)
d := &Driver{coreDeviceID: "CORE-DEVICE", bundleID: "app.example", appPath: "/tmp/Sample.app"}
if err := d.devicectlReinstall(context.Background()); err != nil {
t.Fatalf("devicectlReinstall: %v", err)
}
want := []string{
"devicectl device uninstall app --device CORE-DEVICE app.example",
"devicectl device install app --device CORE-DEVICE /tmp/Sample.app",
}
if got := xcrunCalls(t, log); !slices.Equal(got, want) {
t.Fatalf("xcrun calls = %v, want %v", got, want)
}
}
// TestNewDeviceRefusesClearStateWithoutAnAppPath keeps the device from starting
// a run whose clear-state cannot happen. There is no data-container wipe on a
// physical device, so without an app path to reinstall from, carrying on hands
// the run every previous run's data under a flag that says otherwise.
func TestNewDeviceRefusesClearStateWithoutAnAppPath(t *testing.T) {
address := startLoopbackListener(t)
options := testDeviceOptions(address, newDeviceCompanion())
options.HardwareUDID = "00008140-NO-APP-PATH"
options.ClearState = true
spawned := false
spawn := options.spawnRunner
options.spawnRunner = func(ctx context.Context, runnerAddress string) (*exec.Cmd, error) {
spawned = true
return spawn(ctx, runnerAddress)
}
d, err := NewDevice(context.Background(), options)
if err == nil {
d.Close()
t.Fatal("NewDevice returned a driver whose clear-state never happened")
}
if !strings.Contains(err.Error(), "--ios-app-path") {
t.Fatalf("err = %v, want it to name the flag that makes the clear possible", err)
}
if spawned {
t.Fatal("the run started anyway; a clear-state that cannot happen must end the run, not open it")
}
}
+171 -25
View File
@@ -53,6 +53,14 @@ var shutdownGrace = 15 * time.Second
// A variable so the timeout test can shrink it.
var launchTimeout = 90 * time.Second
// launchRecoveryTimeout bounds the whole recovery a blown launch bound
// triggers, the session restart and the second attempt together. It keeps the
// launch path inside the three minutes testrun allows it, so what a user sees
// when the app really cannot be launched stays the driver's error rather than
// that backstop firing over the top of it. A variable so the bound test can
// shrink it.
var launchRecoveryTimeout = 60 * time.Second
// longPressHoldMilliseconds is how long LongPress holds the finger down.
const longPressHoldMilliseconds = 600
@@ -65,16 +73,25 @@ type Options struct {
// AppPath is the .app bundle directory. Required for clear-state reinstall;
// when empty, clear state falls back to resetting the data container.
AppPath string
// ClearState resets the app to first-launch state while New runs, before
// any automation session attaches. Clear state is a property of the driver
// rather than of a launch: see Launch.
ClearState bool
// Output receives companion stdout and stderr plus driver warnings.
Output io.Writer
// DoubleTapGapMilliseconds overrides the synthesized double-tap gap.
DoubleTapGapMilliseconds float64
// spawnChild, dialCompanion, and pickAddress are test seams. Production
// leaves them nil and New wires the real extraction, spawn, and dial.
spawnChild func(ctx context.Context, address string) (*exec.Cmd, error)
dialCompanion func(address string) (transport.Companion, error)
pickAddress func() (string, error)
// These are test seams. Production leaves them nil and New wires the real
// extraction, spawn, dial and simctl calls.
spawnChild func(ctx context.Context, address string) (*exec.Cmd, error)
dialCompanion func(address string) (transport.Companion, error)
pickAddress func() (string, error)
spawnRunner func(ctx context.Context, address string) (*exec.Cmd, error)
dialRunner func(address string) (transport.Companion, error)
reinstallApp func(ctx context.Context) error
resetContainer func(ctx context.Context) error
terminateApp func(ctx context.Context) error
}
// Driver implements driver.DeviceDriver against an iOS simulator companion.
@@ -85,6 +102,13 @@ type Driver struct {
appPath string
output io.Writer
// clearedBundleID names the app New (or NewDevice) reset to first-launch
// state before attaching, which is the only point in a run where clearing
// is safe. Launch refuses a clear-state request for anything else rather
// than reinstalling under a live session or reporting a reset that only
// ever reached another bundle.
clearedBundleID string
screenWidth int
screenHeight int
@@ -114,6 +138,10 @@ type Driver struct {
// A seam so tests skip the simctl shell-outs.
reinstallApp func(ctx context.Context) error
// terminateApp stops the app before its state is cleared. A seam so tests
// skip the simctl shell-out.
terminateApp func(ctx context.Context) error
// grantPaste pre-authorizes the app's pasteboard access. A seam so tests
// skip the sqlite shell-out.
grantPaste func(ctx context.Context) error
@@ -136,15 +164,19 @@ type Driver struct {
dialRunner func(address string) (transport.Companion, error)
hybrid bool
// pickRunnerAddress hands every bring-up a free loopback port, on the
// simulator and the device alike. One field, so no path can be wired
// without it.
pickRunnerAddress func() (string, error)
// Device-mode fields. On the physical-device path d.companion is the runner
// dialed over a usbmux tunnel, hybrid is false, and runnerClient is nil.
// coreDeviceID feeds devicectl; tunnel is the in-process usbmux forwarder
// bridging the host loopback port to the runner's device-side port.
deviceMode bool
coreDeviceID string
tunnel io.Closer
startTunnel func(ctx context.Context, hardwareUDID, localAddress, devicePort string) (io.Closer, error)
pickDeviceAddress func() (string, error)
deviceMode bool
coreDeviceID string
tunnel io.Closer
startTunnel func(ctx context.Context, hardwareUDID, localAddress, devicePort string) (io.Closer, error)
// processContext owns the companion child's lifetime: it is derived from
// New's context (so a canceled run still reaps the child) and canceled by
@@ -189,6 +221,9 @@ func New(ctx context.Context, options Options) (*Driver, error) {
if options.UniqueDeviceIdentifier == "" {
return nil, errors.New("ios companion: UniqueDeviceIdentifier is required")
}
if options.ClearState && options.BundleID == "" {
return nil, errors.New("ios companion: clear-state needs BundleID: there is nothing to uninstall or wipe without it")
}
output := options.Output
if output == nil {
output = io.Discard
@@ -206,6 +241,11 @@ func New(ctx context.Context, options Options) (*Driver, error) {
doubleTapGapMilliseconds: gap,
spawnChild: options.spawnChild,
dial: options.dialCompanion,
spawnRunner: options.spawnRunner,
dialRunner: options.dialRunner,
reinstallApp: options.reinstallApp,
resetContainer: options.resetContainer,
terminateApp: options.terminateApp,
hybrid: hybridCompanionEnabled(),
}
if driverInstance.spawnChild == nil {
@@ -232,9 +272,17 @@ func New(ctx context.Context, options Options) (*Driver, error) {
return nil, err
}
driverInstance.address = address
driverInstance.pickRunnerAddress = pickAddress
driverInstance.restart = driverInstance.respawnAndRedial
driverInstance.resetContainer = driverInstance.resetDataContainer
driverInstance.reinstallApp = driverInstance.simctlReinstall
if driverInstance.resetContainer == nil {
driverInstance.resetContainer = driverInstance.resetDataContainer
}
if driverInstance.reinstallApp == nil {
driverInstance.reinstallApp = driverInstance.simctlReinstall
}
if driverInstance.terminateApp == nil {
driverInstance.terminateApp = driverInstance.simctlTerminate
}
driverInstance.grantPaste = driverInstance.grantPasteboardAccess
driverInstance.processContext, driverInstance.processCancel = context.WithCancel(ctx)
@@ -245,6 +293,14 @@ func New(ctx context.Context, options Options) (*Driver, error) {
}
driverInstance.deviceLock = lock
if options.ClearState {
if err := driverInstance.clearAppState(ctx); err != nil {
driverInstance.Close()
return nil, err
}
driverInstance.clearedBundleID = options.BundleID
}
if err := driverInstance.bringUp(ctx); err != nil {
driverInstance.Close()
return nil, err
@@ -358,7 +414,7 @@ func (d *Driver) bringUpRunner(ctx context.Context) error {
// A fresh port every bring-up: after a restart the dying session's
// listener may still answer on the old port and would satisfy the wait
// below with a dead server.
address, err := pickLoopbackAddress()
address, err := d.pickRunnerAddress()
if err != nil {
return err
}
@@ -441,6 +497,26 @@ func isConnectionError(err error) bool {
return false
}
// isBudgetExpiry reports whether err is a call that outlived its bound rather
// than a failure the transport can name. Each transport says so differently:
// the runner wraps the context's error when cancellation has landed and the
// connection's i/o timeout when the deadline it armed from that context fires
// first, and the legacy companion returns a gRPC status. Only an expiry earns
// a session restart; an error the runner reports has already said what a fresh
// session would say.
func isBudgetExpiry(err error) bool {
if err == nil {
return false
}
if errors.Is(err, context.DeadlineExceeded) || errors.Is(err, os.ErrDeadlineExceeded) {
return true
}
if statusValue, ok := status.FromError(err); ok {
return statusValue.Code() == codes.DeadlineExceeded
}
return false
}
func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, env map[string]string) error {
if bundleID != "" {
d.bundleID = bundleID
@@ -452,6 +528,14 @@ func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, e
// loudly rather than silently dropping the request.
return errors.New("ios companion: launch with environment variables is unsupported on this backend")
}
if clearState && (d.clearedBundleID == "" || d.clearedBundleID != d.bundleID) {
// Clearing here would uninstall and reinstall the app underneath a live
// automation session, which is what races FrontBoard's registration and
// leaves the session launching a bundle FrontBoard has not registered.
return fmt.Errorf("ios companion: clear-state must be requested when the driver is created (Options.ClearState) "+
"for the bundle being launched; this backend cleared %q before its automation session existed, not %q",
d.clearedBundleID, d.bundleID)
}
// Terminate first so the launch is a clean cold start regardless of the
// app's prior state. A not-running app is not an error here.
@@ -459,12 +543,6 @@ func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, e
return companion.Terminate(callCtx, d.bundleID)
})
if clearState {
if err := d.clearAppState(ctx); err != nil {
return err
}
}
// Grant the app pasteboard access before it runs so unicode input (which
// must go through the pasteboard, since HID cannot express it) never trips
// the iOS paste-permission prompt. clearState reinstall resets the grant,
@@ -477,14 +555,55 @@ func (d *Driver) Launch(ctx context.Context, bundleID string, clearState bool, e
}
}
if err := d.lifecycleCall(ctx, func(callCtx context.Context, companion transport.Companion) error {
return companion.Launch(callCtx, d.bundleID, true)
}); err != nil {
if err := d.launchWithSessionRecovery(ctx); err != nil {
return fmt.Errorf("launch %s: %w", d.bundleID, err)
}
return nil
}
// launchWithSessionRecovery runs the launch RPC and, when it blows its own
// bound, replaces the session and launches again.
//
// A launch the simulator refuses, which is what a clear-state reinstall racing
// FrontBoard's registration produces, never comes back as an error: XCTest
// records the refusal as a test failure the runner cannot observe, then holds
// the session's main thread for about four minutes walking a diagnostic chain
// (a 120s accessibility wait, a spindump, an idle wait). So there is no error
// text to key a retry on, only the expired bound, and every later call queues
// behind the same wedge. Only a session that never served the refused launch
// can serve the retry, which is why this restarts rather than calls again.
func (d *Driver) launchWithSessionRecovery(ctx context.Context) error {
launch := func(callCtx context.Context, companion transport.Companion) error {
return companion.Launch(callCtx, d.bundleID, true)
}
err := d.lifecycleCall(ctx, launch)
// A caller whose own budget ran out gets no restart: the bound that expired
// was the caller's to spend, and the second attempt would inherit it dead.
if !isBudgetExpiry(err) || ctx.Err() != nil || d.restart == nil {
return err
}
fmt.Fprintf(d.output, "launch %s blew its %v bound (%v); restarting the session and launching once more\n",
d.bundleID, launchTimeout, err)
// The restart runs under the driver's own lifetime context for the same
// reason withRecovery's does, while the second attempt stays on the
// caller's. Both end at one deadline, so a launch that already spent
// launchTimeout cannot then wait out a session cold start on top of it.
recoveryDeadline := time.Now().Add(launchRecoveryTimeout)
restartCtx := d.processContext
if restartCtx == nil {
restartCtx = ctx
}
restartCtx, cancelRestart := context.WithDeadline(restartCtx, recoveryDeadline)
defer cancelRestart()
if restartErr := d.restart(restartCtx); restartErr != nil {
return fmt.Errorf("session restart failed: %w (original: %v)", restartErr, err)
}
relaunchCtx, cancelRelaunch := context.WithDeadline(ctx, recoveryDeadline)
defer cancelRelaunch()
return d.lifecycleCall(relaunchCtx, launch)
}
// lifecycleCall runs an app lifecycle RPC against lifecycleCompanion under a
// launchTimeout-bounded context, with the usual one-restart recovery. The
// companion is resolved inside the retry so a restart's replacement client
@@ -511,8 +630,14 @@ func (d *Driver) lifecycleCompanion() transport.Companion {
// clearAppState resets the app to a first-launch state. With an app path it
// uninstalls and reinstalls; without one it falls back to wiping the app's data
// container and warns once that a full reinstall needs the app path.
// container and warns once that a full reinstall needs the app path. Called
// only from construction, before any automation session is attached to the app.
func (d *Driver) clearAppState(ctx context.Context) error {
// Nothing may be writing to the state while it goes, which is the ordering
// Launch used to hold: uninstall copes with a running app, deleting the
// data container out from under one does not. Best effort, because an app
// that is not running reports a failure that means nothing here.
_ = d.terminateApp(ctx)
if d.appPath != "" {
if err := d.reinstallApp(ctx); err != nil {
return fmt.Errorf("reinstall %s: %w", d.appPath, err)
@@ -529,8 +654,14 @@ func (d *Driver) clearAppState(ctx context.Context) error {
// simctlReinstall uninstalls and reinstalls the app bundle via simctl. App
// lifecycle stays with simctl: the companion's install RPC misreads current
// simulator targets' architectures and rejects valid bundles.
// A failed uninstall ends the reinstall: `simctl install` over an installed app
// carries its data container across, so clear-state would be reported without
// happening. Uninstalling an app that is not installed exits 0, so there is no
// benign failure here to sort out from a real one.
func (d *Driver) simctlReinstall(ctx context.Context) error {
_ = exec.CommandContext(ctx, "xcrun", "simctl", "uninstall", d.udid, d.bundleID).Run()
if output, err := exec.CommandContext(ctx, "xcrun", "simctl", "uninstall", d.udid, d.bundleID).CombinedOutput(); err != nil {
return fmt.Errorf("simctl uninstall %s: %w: %s", d.bundleID, err, strings.TrimSpace(string(output)))
}
output, err := exec.CommandContext(ctx, "xcrun", "simctl", "install", d.udid, d.appPath).CombinedOutput()
if err != nil {
return fmt.Errorf("simctl install: %w: %s", err, strings.TrimSpace(string(output)))
@@ -538,6 +669,18 @@ func (d *Driver) simctlReinstall(ctx context.Context) error {
return nil
}
// simctlTerminate stops the app under test. Launch used to terminate through
// the automation session before clearing; the clear now runs before any session
// exists, so simctl is what is left to stop the app with. An app that is not
// running reports a failure that means nothing to the caller, which is why
// clearAppState treats this as best effort.
func (d *Driver) simctlTerminate(ctx context.Context) error {
if output, err := exec.CommandContext(ctx, "xcrun", "simctl", "terminate", d.udid, d.bundleID).CombinedOutput(); err != nil {
return fmt.Errorf("simctl terminate %s: %w: %s", d.bundleID, err, strings.TrimSpace(string(output)))
}
return nil
}
// grantPasteboardAccess authorizes the app to read the pasteboard without the
// iOS permission prompt, by writing an allow row into the simulator's privacy
// (TCC) database. This is the simulator counterpart to `simctl privacy grant`,
@@ -941,7 +1084,10 @@ func (d *Driver) WaitForIdle(ctx context.Context, _ time.Duration) error {
}
// RecentLogs returns no entries: the companion log RPC is a follow-up, so v1
// reports an empty slice rather than failing.
// reports an empty slice rather than failing. Every property reading state.logs
// therefore holds vacuously on iOS and nothing says so; closing it means
// tailing idb's streaming log RPC and mapping os_log levels onto the
// single-letter scale driver.LogEntry declares.
func (d *Driver) RecentLogs(_ context.Context, _ time.Time, _ string) ([]driver.LogEntry, error) {
return []driver.LogEntry{}, nil
}
+555 -24
View File
@@ -12,15 +12,18 @@ import (
"os"
"os/exec"
"path/filepath"
"slices"
"strings"
"sync"
"testing"
"time"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/companionpb"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion/transport"
)
@@ -182,47 +185,280 @@ func TestLaunchContinuesWhenGrantFails(t *testing.T) {
}
}
func TestLaunchClearStateReinstallsWithAppPath(t *testing.T) {
// clearStateProbe records, in order, the calls a run makes to reset the app and
// to bring the runner's automation session up. A reinstall recorded after the
// session is the ordering that races FrontBoard.
type clearStateProbe struct {
mutex sync.Mutex
events []string
}
func (p *clearStateProbe) record(event string) {
p.mutex.Lock()
defer p.mutex.Unlock()
p.events = append(p.events, event)
}
func (p *clearStateProbe) recorded() []string {
p.mutex.Lock()
defer p.mutex.Unlock()
out := make([]string, len(p.events))
copy(out, p.events)
return out
}
// clearStateOptions wires every seam a hybrid bring-up needs, so New runs its
// real sequence against fakes: no simulator, no simctl, no XCTest session.
func clearStateOptions(t *testing.T, probe *clearStateProbe, udid string, clearState bool) Options {
t.Helper()
t.Setenv("SANDERLING_SIMULATOR_COMPANION", "")
address := startLoopbackListener(t)
return Options{
UniqueDeviceIdentifier: udid,
BundleID: "com.example.app",
ClearState: clearState,
Output: &bytes.Buffer{},
pickAddress: func() (string, error) { return address, nil },
spawnChild: func(context.Context, string) (*exec.Cmd, error) { return &exec.Cmd{}, nil },
dialCompanion: func(string) (transport.Companion, error) {
return &fakeCompanion{accessibilityJSON: "[]"}, nil
},
spawnRunner: func(context.Context, string) (*exec.Cmd, error) {
probe.record("runner session")
return &exec.Cmd{}, nil
},
dialRunner: func(string) (transport.Companion, error) {
return &fakeCompanion{accessibilityJSON: "[]"}, nil
},
reinstallApp: func(context.Context) error { probe.record("reinstall"); return nil },
resetContainer: func(context.Context) error { probe.record("reset container"); return nil },
terminateApp: func(context.Context) error { probe.record("stop app"); return nil },
}
}
func TestClearStateReinstallsOnceBeforeTheRunnerSession(t *testing.T) {
probe := &clearStateProbe{}
options := clearStateOptions(t, probe, "CLEAR-REINSTALL-UDID", true)
options.AppPath = "/tmp/Sample.app"
d, err := New(context.Background(), options)
if err != nil {
t.Fatalf("New: %v", err)
}
defer d.Close()
if err := d.Launch(context.Background(), "", true, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
want := []string{"stop app", "reinstall", "runner session"}
if got := probe.recorded(); !slices.Equal(got, want) {
t.Fatalf("calls = %v, want %v: the reinstall must run once, on a stopped app, before the automation session attaches", got, want)
}
}
func TestClearStateWithoutAppPathWipesContainerBeforeTheRunnerSession(t *testing.T) {
probe := &clearStateProbe{}
output := &bytes.Buffer{}
options := clearStateOptions(t, probe, "CLEAR-CONTAINER-UDID", true)
options.Output = output
d, err := New(context.Background(), options)
if err != nil {
t.Fatalf("New: %v", err)
}
defer d.Close()
if err := d.Launch(context.Background(), "", true, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
want := []string{"stop app", "reset container", "runner session"}
if got := probe.recorded(); !slices.Equal(got, want) {
t.Fatalf("calls = %v, want %v: the fallback must wipe a stopped app's container once, before the session, and never reinstall", got, want)
}
if warnings := strings.Count(output.String(), "resetting the data container only"); warnings != 1 {
t.Fatalf("warning emitted %d times, want once", warnings)
}
}
func TestWithoutClearStateTheAppIsLeftAlone(t *testing.T) {
probe := &clearStateProbe{}
options := clearStateOptions(t, probe, "NO-CLEAR-UDID", false)
options.AppPath = "/tmp/Sample.app"
d, err := New(context.Background(), options)
if err != nil {
t.Fatalf("New: %v", err)
}
defer d.Close()
if err := d.Launch(context.Background(), "", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
want := []string{"runner session"}
if got := probe.recorded(); !slices.Equal(got, want) {
t.Fatalf("calls = %v, want %v: a run that did not ask for clear state must not touch the install", got, want)
}
}
func TestLaunchRefusesClearStateTheDriverWasNotBuiltFor(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
d := newTestDriver(companion)
d.appPath = "/tmp/Sample.app"
reinstalls := 0
d.reinstallApp = func(context.Context) error { reinstalls++; return nil }
if err := d.Launch(context.Background(), "", true, nil); err != nil {
t.Fatalf("Launch: %v", err)
err := d.Launch(context.Background(), "", true, nil)
if err == nil || !strings.Contains(err.Error(), "clear-state") {
t.Fatalf("Launch err = %v, want a refusal naming clear-state", err)
}
if reinstalls != 1 {
t.Fatalf("clear-state with app path must reinstall exactly once; got %d", reinstalls)
if reinstalls != 0 {
t.Fatalf("reinstalls = %d, want 0: a live session must never have the app reinstalled under it", reinstalls)
}
if indexOf(companion.calls, "launch") < indexOf(companion.calls, "terminate") {
t.Fatalf("launch must still follow terminate; got %v", companion.calls)
if indexOf(companion.recorded(), "launch") >= 0 {
t.Fatalf("a refused launch must not reach the companion; got %v", companion.recorded())
}
}
func TestLaunchClearStateFallbackWarnsOnce(t *testing.T) {
// TestLaunchRefusesClearStateForABundleItDidNotClear holds the guard to the
// fact it is guarding. A driver built to clear one bundle has cleared nothing
// for another, so reporting that launch as a clear-state launch is a reset the
// caller was told happened and did not.
func TestLaunchRefusesClearStateForABundleItDidNotClear(t *testing.T) {
companion := &fakeCompanion{accessibilityJSON: "[]"}
output := &bytes.Buffer{}
d := newTestDriver(companion)
d.output = output
resets := 0
d.resetContainer = func(context.Context) error { resets++; return nil }
d.clearedBundleID = "com.example.app"
for i := 0; i < 2; i++ {
if err := d.Launch(context.Background(), "", true, nil); err != nil {
t.Fatalf("Launch %d: %v", i, err)
err := d.Launch(context.Background(), "com.other.app", true, nil)
if err == nil || !strings.Contains(err.Error(), "clear-state") {
t.Fatalf("Launch err = %v, want a refusal naming clear-state: com.other.app was never cleared", err)
}
if indexOf(companion.recorded(), "launch") >= 0 {
t.Fatalf("a launch reporting a clear that never happened must not reach the companion; got %v", companion.recorded())
}
}
func TestNewRejectsClearStateWithoutBundleID(t *testing.T) {
probe := &clearStateProbe{}
options := clearStateOptions(t, probe, "NO-BUNDLE-UDID", true)
options.BundleID = ""
if _, err := New(context.Background(), options); err == nil || !strings.Contains(err.Error(), "BundleID") {
t.Fatalf("New err = %v, want a refusal naming BundleID", err)
}
if got := probe.recorded(); len(got) != 0 {
t.Fatalf("calls = %v, want none: clearing an unnamed bundle would reinstall without resetting anything", got)
}
}
// scriptedXcrun puts an xcrun on PATH that logs each invocation's arguments and
// answers from replies, a `case "$*" in` body, so a reinstall runs its real
// command sequence and the log holds what reached the tool.
func scriptedXcrun(t *testing.T, replies string) string {
t.Helper()
directory := t.TempDir()
log := filepath.Join(directory, "xcrun.log")
script := "#!/bin/sh\necho \"$*\" >> " + log + "\ncase \"$*\" in\n" + replies + "\nesac\n"
if err := os.WriteFile(filepath.Join(directory, "xcrun"), []byte(script), 0o755); err != nil {
t.Fatalf("write xcrun: %v", err)
}
t.Setenv("PATH", directory)
return log
}
func xcrunCalls(t *testing.T, log string) []string {
t.Helper()
contents, err := os.ReadFile(log)
if err != nil {
if os.IsNotExist(err) {
return nil
}
t.Fatalf("read %s: %v", log, err)
}
return strings.Split(strings.TrimSpace(string(contents)), "\n")
}
func TestSimctlReinstallStopsWhenTheUninstallFails(t *testing.T) {
log := scriptedXcrun(t, `"simctl uninstall "*) echo "Simulator device failed to uninstall app.example."; echo "Uninstall prohibited."; exit 22;;
"simctl install "*) :;;`)
d := &Driver{udid: "SIM-UDID", bundleID: "app.example", appPath: "/tmp/Sample.app"}
err := d.simctlReinstall(context.Background())
if err == nil {
t.Fatal("simctlReinstall reported success while app.example kept the data clear-state was asked to remove")
}
for _, want := range []string{"app.example", "Uninstall prohibited."} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q does not quote %q", err, want)
}
}
if resets != 2 {
t.Fatalf("resetContainer called %d times, want 2", resets)
if calls := xcrunCalls(t, log); slices.Contains(calls, "simctl install SIM-UDID /tmp/Sample.app") {
t.Errorf("xcrun calls = %v: installing over the app carries its data into the run", calls)
}
warnings := strings.Count(output.String(), "resetting the data container only")
if warnings != 1 {
t.Fatalf("warning emitted %d times, want once", warnings)
}
func TestSimctlReinstallProceedsWhenNothingIsInstalled(t *testing.T) {
log := scriptedXcrun(t, `"simctl uninstall "*) :;;
"simctl install "*) :;;`)
d := &Driver{udid: "SIM-UDID", bundleID: "app.example", appPath: "/tmp/Sample.app"}
if err := d.simctlReinstall(context.Background()); err != nil {
t.Fatalf("simctlReinstall: %v", err)
}
for _, call := range companion.calls {
if call == "install" || call == "uninstall" {
t.Fatalf("fallback path must not install/uninstall; got %v", companion.calls)
}
want := []string{"simctl uninstall SIM-UDID app.example", "simctl install SIM-UDID /tmp/Sample.app"}
if got := xcrunCalls(t, log); !slices.Equal(got, want) {
t.Fatalf("xcrun calls = %v, want %v", got, want)
}
}
// TestClearStateStopsTheAppBeforeWipingItsContainer covers the ordering Launch
// used to hold. simctl uninstall copes with a running app; deleting the data
// container out from under one does not, and the CI iOS leg passes no app path
// so it is the wipe that runs. A run whose previous run was interrupted finds
// the app still up.
func TestClearStateStopsTheAppBeforeWipingItsContainer(t *testing.T) {
container := t.TempDir()
stale := filepath.Join(container, "Documents")
if err := os.Mkdir(stale, 0o755); err != nil {
t.Fatal(err)
}
log := scriptedXcrun(t, `"simctl terminate "*) :;;
"simctl get_app_container "*) echo `+container+`;;`)
d := &Driver{udid: "SIM-UDID", bundleID: "app.example", output: &bytes.Buffer{}}
d.terminateApp = d.simctlTerminate
d.resetContainer = d.resetDataContainer
if err := d.clearAppState(context.Background()); err != nil {
t.Fatalf("clearAppState: %v", err)
}
want := []string{
"simctl terminate SIM-UDID app.example",
"simctl get_app_container SIM-UDID app.example data",
}
if got := xcrunCalls(t, log); !slices.Equal(got, want) {
t.Fatalf("xcrun calls = %v, want %v: the app was still writing to the container being deleted", got, want)
}
if _, err := os.Stat(stale); !os.IsNotExist(err) {
t.Fatalf("stat %s = %v, want the previous run's state gone", stale, err)
}
}
// TestClearStateSurvivesAnAppThatIsNotRunning holds the terminate to best
// effort. simctl exits non-zero when there is nothing to stop, and a first run
// on a fresh simulator must not fail on it.
func TestClearStateSurvivesAnAppThatIsNotRunning(t *testing.T) {
container := t.TempDir()
scriptedXcrun(t, `"simctl terminate "*) echo "No matching processes belonging to bundle identifier app.example"; exit 3;;
"simctl get_app_container "*) echo `+container+`;;`)
d := &Driver{udid: "SIM-UDID", bundleID: "app.example", output: &bytes.Buffer{}}
d.terminateApp = d.simctlTerminate
d.resetContainer = d.resetDataContainer
if err := d.clearAppState(context.Background()); err != nil {
t.Fatalf("clearAppState: %v: an app that is not running is not a failure to clear", err)
}
}
@@ -951,6 +1187,301 @@ func TestLaunchLeavesATighterCallerDeadlineAlone(t *testing.T) {
}
}
// blownBudgetShape is one way a transport the driver launches through reports
// a lifecycle call outliving its budget.
type blownBudgetShape struct {
name string
err error
}
// blownBudgetShapes drives every such transport against a server that never
// answers and returns the error each one really produces. The runner transport
// has two: wrapTransport wraps the context's own error once cancellation has
// landed, and the connection's i/o timeout when the deadline it armed from
// that context fires first. The legacy transport reports the same expiry as a
// gRPC status. Only the first satisfies errors.Is(err, context.DeadlineExceeded),
// so a fake that returns ctx.Err() raw shows the driver a recovery that two
// thirds of production can never reach.
func blownBudgetShapes(t *testing.T) []blownBudgetShape {
t.Helper()
return []blownBudgetShape{
{"runner context deadline", runnerContextDeadlineError(t)},
{"runner connection deadline", silentRunnerLaunchError(t, connectionDeadlineOnly{time.Now().Add(100 * time.Millisecond)})},
{"legacy grpc deadline", silentGRPCLaunchError(t)},
}
}
// runnerContextDeadlineError is the shape the runner transport produces once
// the context's own cancellation has landed.
func runnerContextDeadlineError(t *testing.T) error {
t.Helper()
ctx, cancel := context.WithDeadline(context.Background(), time.Now().Add(-time.Second))
defer cancel()
return silentRunnerLaunchError(t, ctx)
}
// connectionDeadlineOnly carries a deadline the runner transport arms the
// connection with, while its own cancellation never lands. That is the race
// wrapTransport's second branch exists for: the connection's deadline fires
// while ctx.Err() is still nil.
type connectionDeadlineOnly struct{ deadline time.Time }
func (c connectionDeadlineOnly) Deadline() (time.Time, bool) { return c.deadline, true }
func (c connectionDeadlineOnly) Done() <-chan struct{} { return nil }
func (c connectionDeadlineOnly) Err() error { return nil }
func (c connectionDeadlineOnly) Value(any) any { return nil }
// silentRunnerLaunchError returns what the real runner transport produces for a
// launch nobody ever answers.
func silentRunnerLaunchError(t *testing.T, ctx context.Context) error {
t.Helper()
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
closed := make(chan struct{})
go func() {
conn, acceptErr := listener.Accept()
if acceptErr != nil {
return
}
defer conn.Close()
<-closed
}()
companion, err := transport.DialRunner(listener.Addr().String(), "SIM-UDID", "com.example.app")
if err != nil {
listener.Close()
t.Fatal(err)
}
t.Cleanup(func() {
close(closed)
_ = companion.Close()
listener.Close()
})
launchErr := companion.Launch(ctx, "com.example.app", true)
if launchErr == nil {
t.Fatal("the runner transport reported a launch no server ever answered")
}
return launchErr
}
// silentCompanionServer is the legacy companion with a launch that never
// answers, so the caller's own deadline is what ends the call.
type silentCompanionServer struct {
companionpb.UnimplementedCompanionServiceServer
}
func (silentCompanionServer) Launch(stream grpc.BidiStreamingServer[companionpb.LaunchRequest, companionpb.LaunchResponse]) error {
<-stream.Context().Done()
return stream.Context().Err()
}
// silentGRPCLaunchError returns what the legacy transport produces for the same
// launch, which SANDERLING_SIMULATOR_COMPANION=legacy still runs on.
func silentGRPCLaunchError(t *testing.T) error {
t.Helper()
listener, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
t.Fatal(err)
}
server := grpc.NewServer()
companionpb.RegisterCompanionServiceServer(server, silentCompanionServer{})
go server.Serve(listener)
t.Cleanup(server.Stop)
companion, err := transport.Dial(listener.Addr().String())
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { companion.Close() })
ctx, cancel := context.WithTimeout(context.Background(), 250*time.Millisecond)
defer cancel()
launchErr := companion.Launch(ctx, "com.example.app", true)
if launchErr == nil {
t.Fatal("the legacy transport reported a launch the companion never answered")
}
return launchErr
}
// wedgedUntilRestartCompanion models the session a refused launch leaves
// behind: the refusal is never reported, no later launch is answered until the
// session itself is replaced, and the expired bound reaches the driver in
// whatever shape its transport gives it.
type wedgedUntilRestartCompanion struct {
fakeCompanion
blownBudget error
mutex sync.Mutex
replaced bool
attempted int
}
func (w *wedgedUntilRestartCompanion) replaceSession() {
w.mutex.Lock()
defer w.mutex.Unlock()
w.replaced = true
}
func (w *wedgedUntilRestartCompanion) launchAttempts() int {
w.mutex.Lock()
defer w.mutex.Unlock()
return w.attempted
}
func (w *wedgedUntilRestartCompanion) Launch(ctx context.Context, _ string, _ bool) error {
w.mutex.Lock()
w.attempted++
replaced := w.replaced
w.mutex.Unlock()
if replaced {
return nil
}
<-ctx.Done()
return w.blownBudget
}
// TestLaunchReplacesTheSessionAfterALaunchBlowsItsBound covers the FrontBoard
// race: a clear-state reinstall the simulator has not finished registering
// makes the session refuse the launch, and XCTest answers that refusal with
// minutes of diagnostics instead of an error, so the bound expires and every
// later call queues behind the same wedge. Calling launch again on that session
// cannot work; the run only recovers if the session is replaced first. The
// recovery has to fire on every shape the driver's transports report that
// expiry in, because which one arrives is a race the driver does not control.
func TestLaunchReplacesTheSessionAfterALaunchBlowsItsBound(t *testing.T) {
for _, shape := range blownBudgetShapes(t) {
t.Run(shape.name, func(t *testing.T) {
previous := launchTimeout
launchTimeout = 100 * time.Millisecond
defer func() { launchTimeout = previous }()
companion := &wedgedUntilRestartCompanion{blownBudget: shape.err}
output := &bytes.Buffer{}
d := newTestDriver(companion)
d.output = output
restarts := 0
d.restart = func(context.Context) error {
restarts++
companion.replaceSession()
return nil
}
if err := d.Launch(context.Background(), "", false, nil); err != nil {
t.Fatalf("Launch: %v", err)
}
if restarts != 1 {
t.Fatalf("session restarts = %d, want exactly 1 (the session reported %v)", restarts, shape.err)
}
if attempts := companion.launchAttempts(); attempts != 2 {
t.Fatalf("launch attempts = %d, want 2: one that wedged and one on the replaced session", attempts)
}
if !strings.Contains(output.String(), "restarting the session") {
t.Fatalf("the recovery was silent, so a run that needed it never says so; output was %q", output.String())
}
})
}
}
// TestLaunchBoundsTheSessionRestartItTriggers keeps the recovery inside a
// budget of its own. The restart deliberately runs on the driver's lifetime
// context rather than the caller's, so without a deadline a session that never
// comes back would hang the launch path exactly the way #73 stopped it hanging.
func TestLaunchBoundsTheSessionRestartItTriggers(t *testing.T) {
previousLaunch, previousRecovery := launchTimeout, launchRecoveryTimeout
launchTimeout = 100 * time.Millisecond
launchRecoveryTimeout = 200 * time.Millisecond
defer func() { launchTimeout, launchRecoveryTimeout = previousLaunch, previousRecovery }()
d := newTestDriver(&wedgedUntilRestartCompanion{blownBudget: runnerContextDeadlineError(t)})
d.restart = func(restartCtx context.Context) error {
<-restartCtx.Done()
return restartCtx.Err()
}
done := make(chan error, 1)
go func() { done <- d.Launch(context.Background(), "", false, nil) }()
select {
case err := <-done:
if err == nil || !strings.Contains(err.Error(), "session restart failed") {
t.Fatalf("err = %v, want the failed restart named", err)
}
case <-time.After(10 * time.Second):
t.Fatal("Launch never returned: a session that never comes back hangs the launch path")
}
}
// TestLaunchKeepsTheSessionWhenTheCallersOwnDeadlineExpires holds the recovery
// to the driver's own bound. Spending a session restart on a caller that has
// already run out of budget cannot produce a launch, only a later failure.
func TestLaunchKeepsTheSessionWhenTheCallersOwnDeadlineExpires(t *testing.T) {
previous := launchTimeout
launchTimeout = 30 * time.Second
defer func() { launchTimeout = previous }()
companion := &wedgedUntilRestartCompanion{blownBudget: runnerContextDeadlineError(t)}
d := newTestDriver(companion)
restarts := 0
d.restart = func(context.Context) error {
restarts++
companion.replaceSession()
return nil
}
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
if err := d.Launch(ctx, "", false, nil); !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("err = %v, want a deadline-exceeded error", err)
}
if restarts != 0 {
t.Fatalf("session restarts = %d, want 0", restarts)
}
}
// refusedLaunchCompanion answers a launch the way the runner does once it
// checks the app's state after activating it: promptly, naming the app and the
// state it reached, over a session that is still serving.
type refusedLaunchCompanion struct {
fakeCompanion
attempts int
}
func (r *refusedLaunchCompanion) Launch(context.Context, string, bool) error {
r.attempts++
return errors.New(`runner launch: failed("com.example.app is not running after launch")`)
}
// TestLaunchKeepsTheSessionWhenTheRunnerNamesTheRefusal separates a launch that
// answers from a launch that never does. The session restart is the only
// recovery from a wedged session, and it costs a cold start; a runner that
// reports the app's state has already said what a fresh session would say, so
// restarting to hear it again only delays the error and hides the app under it.
func TestLaunchKeepsTheSessionWhenTheRunnerNamesTheRefusal(t *testing.T) {
companion := &refusedLaunchCompanion{}
output := &bytes.Buffer{}
d := newTestDriver(companion)
d.output = output
restarts := 0
d.restart = func(context.Context) error {
restarts++
return nil
}
err := d.Launch(context.Background(), "", false, nil)
if err == nil || !strings.Contains(err.Error(), "com.example.app is not running after launch") {
t.Fatalf("err = %v, want the runner's refusal reaching the caller intact", err)
}
if restarts != 0 {
t.Fatalf("session restarts = %d, want 0: a refusal the runner reported is not a wedged session", restarts)
}
if companion.attempts != 1 {
t.Fatalf("launch attempts = %d, want 1: relaunching an app the runner just refused cannot launch it", companion.attempts)
}
if strings.Contains(output.String(), "restarting the session") {
t.Fatalf("the driver announced a recovery it must not spend here; output was %q", output.String())
}
}
// newLockTestOptions builds New options that dial a seamed companion, so the
// device-lock tests exercise New without spawning anything.
func newLockTestOptions(t *testing.T, udid string) Options {
+11 -3
View File
@@ -11,7 +11,8 @@
// descPrefix:<prefix> - starts-with on content-desc / accessibilityText
//
// Object selectors (multi-attribute AND, element-scoped or global):
// { attr: value, ... } - all key/value pairs must match; substring / boolean semantics
// { attr: value, ... } - all key/value pairs must match, each key resolved by
// the same rule its string form above uses
//
// Path queries (global scan only, string form):
// <sel> > <sel> > ... - each segment matched within subtree of previous match
@@ -156,10 +157,17 @@ func matchAttr(element *Element, attr, value string) bool {
return false
}
// matchSelector returns true when all filters in sel match the element (AND semantics).
// matchSelector returns true when all filters in sel match the element (AND
// semantics). Each filter goes through match, the same rule the string form
// resolves a "kind:value" segment by, so {id: "Submit"} and "id:Submit" can
// never resolve to different elements. Applying matchAttr directly here made
// the object form skip the kind arms entirely: id, desc and descPrefix name no
// attribute any producer writes, so those keys matched NOTHING through an
// object selector while the string form matched, and every property over the
// missing element passed vacuously.
func matchSelector(element *Element, sel Selector) bool {
for _, f := range sel.Filters {
if !matchAttr(element, f.Attr, f.Value) {
if !match(element, f.Attr, f.Value) {
return false
}
}
+80
View File
@@ -1035,3 +1035,83 @@ func TestTreeTransitional(t *testing.T) {
t.Error("nil tree must not be flagged as transitional")
}
}
// selectorFormsDump carries one node per id shape a real dump produces, plus
// nodes carrying a description in the ", " form the desc rule knows about and a
// text the text rule matches on a substring.
const selectorFormsDump = `{
"attributes": {"resource-id": "root", "bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "BareThing", "bounds": "[0,0,100,50]"}, "children": []},
{"attributes": {"resource-id": "com.example.app:id/AndroidThing", "bounds": "[0,50,100,100]"},
"children": []},
{"attributes": {"accessibilityIdentifier": "IosThing", "bounds": "[0,100,100,150]"},
"children": []},
{"attributes": {"resource-id": "Described", "content-desc": "Save, button", "bounds": "[0,150,100,200]"},
"children": []},
{"attributes": {"resource-id": "Labelled", "text": "Total balance", "bounds": "[0,200,100,250]"},
"children": []}
]
}`
// TestSelectorFormsResolveTheSameElement holds the two selector forms a spec can
// write to ONE rule per key. A spec reaches these through state.ax.find: a
// string goes to FindNode, an object to FindBySelector, and the two ran
// different matchers. `id` has a kind arm that knows an Android resource id is
// package-qualified (com.example.app:id/Thing) and that a spec names the bare
// tail; the object form had no such arm and looked for a literal `id` attribute
// no producer writes, so {id: "Thing"} silently matched nothing on every
// platform while "id:Thing" matched. `desc` and `descPrefix` had the same
// split. A selector that resolves nothing makes every property over it
// vacuously true, which is the failure that reports a green run while checking
// nothing.
func TestSelectorFormsResolveTheSameElement(t *testing.T) {
tree, err := Parse(selectorFormsDump)
if err != nil {
t.Fatal(err)
}
for _, test := range []struct {
key string
value string
want string
}{
{"id", "BareThing", "BareThing"},
// A spec names the tail; an Android dump carries the package prefix.
{"id", "AndroidThing", "com.example.app:id/AndroidThing"},
{"id", "com.example.app:id/AndroidThing", "com.example.app:id/AndroidThing"},
{"id", "IosThing", "IosThing"},
{"desc", "Save, button", "Described"},
// The ", " form an accessibility label takes when a role is appended.
{"desc", "Save", "Described"},
{"descPrefix", "Sav", "Described"},
{"text", "Total", "Labelled"},
{"resource-id", "BareThing", "BareThing"},
{"testTag", "IosThing", "IosThing"},
} {
t.Run(test.key+":"+test.value, func(t *testing.T) {
stringForm := test.key + ":" + test.value
fromString := tree.FindNode(stringForm)
if fromString == nil {
t.Fatalf("the string form %q resolved nothing", stringForm)
}
if fromString.ResourceID != test.want {
t.Fatalf("the string form resolved %q, want %q", fromString.ResourceID, test.want)
}
fromObject := tree.Root.FindBySelector(
Selector{Filters: []AttrFilter{{Attr: test.key, Value: test.value}}},
)
if fromObject == nil {
t.Fatalf(
"the object form {%s: %q} resolved nothing while %q resolved %q",
test.key, test.value, stringForm, fromString.ResourceID,
)
}
if fromObject != fromString {
t.Errorf(
"one selector, two answers: {%s: %q} resolved %q and %q resolved %q",
test.key, test.value, fromObject.ResourceID, stringForm, fromString.ResourceID,
)
}
})
}
}
+3 -2
View File
@@ -84,8 +84,9 @@ type NextFormula struct {
}
// EventuallyFormula obliges its inner formula to hold at some step within the
// given bound. An unbounded eventually never triggers a violation within a
// finite run.
// given bound. An unbounded eventually that never fires is violated when the
// run ends, with the reason "eventually never satisfied", so an eventually over
// a state the run may not reach is red on every run that does not reach it.
//
// When Duration is non-zero and Deadline is the zero time, the evaluator
// resolves the absolute deadline on first reduction using the observation
+509
View File
@@ -0,0 +1,509 @@
package runner
import (
"bytes"
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"sync/atomic"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/trace"
)
// homeWithRows is one settled route whose list holds rows. A row arriving
// between two reads is what a Compose lazy list mounting over several frames
// looks like from the runner's side.
func homeWithRows(rows int) string {
var children strings.Builder
for row := range rows {
fmt.Fprintf(&children,
`,{"attributes":{"resource-id":"TxnRow%d","class":"android.view.View"},"children":[]}`, row)
}
return fmt.Sprintf(
`{"attributes":{"resource-id":"HomeScreen","class":"android.view.View"},"children":[
{"attributes":{"resource-id":"TxnList","class":"android.view.View"},"children":[]}%s
]}`, children.String())
}
// composesLateDriver answers the paired Snapshot with the frame the step
// records and the hierarchy read that follows with a tree that has grown a row,
// for the first composingReads reads of the run. After that both reads describe
// the same screen.
type composesLateDriver struct {
*mockdriver.Driver
composingReads int64
reads atomic.Int64
}
func (d *composesLateDriver) Snapshot(context.Context) (string, driver.Image, error) {
return homeWithRows(1), driver.Image{PNG: []byte("png"), Width: 1, Height: 1}, nil
}
func (d *composesLateDriver) Hierarchy(context.Context) (string, error) {
if d.reads.Add(1) <= d.composingReads {
return homeWithRows(2), nil
}
return homeWithRows(1), nil
}
// A route can settle before its content composes, so a tree read the moment the
// route arrives can describe a screen that is still filling in. Verifying that
// step compares a half-composed frame against a settled one and convicts an app
// that did nothing wrong. Two reads a read apart see it happening, and the step
// they disagree on is one the verifier must never be handed.
//
// The always-false property is the witness: it fires on the first step the
// verifier evaluates, so the step index of its violation says exactly which
// step reached the verifier.
func TestRunner_AStepWhoseTreeChangedBetweenReadsIsNotVerified(t *testing.T) {
run := func(t *testing.T, composingReads int64) (Summary, string) {
t.Helper()
state := newHarnessWithSpec(t, violationSpec)
device := &composesLateDriver{Driver: state.mock, composingReads: composingReads}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 3 {
t.Fatalf("steps = %d, want 3", summary.Steps)
}
return summary, state.writer.Directory()
}
t.Run("the step it changed on is skipped, the next one is judged", func(t *testing.T) {
summary, directory := run(t, 1)
if len(summary.Violations) != 1 {
t.Fatalf("violations = %v, want exactly one", summary.Violations)
}
violation := summary.Violations[0]
if violation.Properties[0] != "balanceNonNegative" {
t.Fatalf("violated %v, want balanceNonNegative", violation.Properties)
}
if violation.StepIndex != 2 {
t.Errorf("the property first judged step %d, want 2; the verifier was handed "+
"a screen that grew a row while the runner was reading it",
violation.StepIndex)
}
if summary.SkippedVerification != 1 {
t.Errorf("the run reports %d step(s) judged by nothing, want 1",
summary.SkippedVerification)
}
// Skipped is not lost: the step is still recorded, screenshot and all,
// so the run can be replayed over the frame nothing judged.
steps := traceSteps(t, directory)
if len(steps) != 3 {
t.Fatalf("trace holds %d step(s), want 3", len(steps))
}
if len(steps[0].Violations) != 0 {
t.Errorf("step 1 recorded violations %v; it was never verified", steps[0].Violations)
}
screenshot := filepath.Join(directory, "screenshots", "step-00001.png")
if _, err := os.Stat(screenshot); err != nil {
t.Errorf("expected the skipped step's screenshot at %s: %v", screenshot, err)
}
})
// The control. Two reads that agree must verify as they always did,
// otherwise the case above is just a runner that verifies nothing.
t.Run("two reads that agree verify the step", func(t *testing.T) {
summary, _ := run(t, 0)
if len(summary.Violations) != 1 {
t.Fatalf("violations = %v, want exactly one", summary.Violations)
}
if got := summary.Violations[0].StepIndex; got != 1 {
t.Errorf("the property first judged step %d, want 1; a settled screen must be "+
"verified on the step it was read", got)
}
})
}
// submitsOnTapDriver commits commitsPerTap transactions on every tap, shows the
// running total in the tree, and grows a row under the hierarchy read that
// follows the paired Snapshot: on one chosen step, on the run of steps from
// composingRead through composingThrough, or on every one of them.
type submitsOnTapDriver struct {
*mockdriver.Driver
commitsPerTap int64
composingRead int64
composingThrough int64
everyRead bool
reads atomic.Int64
committed atomic.Int64
}
func (d *submitsOnTapDriver) Tap(context.Context, int, int) error { return d.commit() }
func (d *submitsOnTapDriver) TapSelector(context.Context, string) error { return d.commit() }
func (d *submitsOnTapDriver) commit() error {
d.committed.Add(d.commitsPerTap)
return nil
}
func (d *submitsOnTapDriver) Snapshot(context.Context) (string, driver.Image, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), driver.Image{}, nil
}
func (d *submitsOnTapDriver) Hierarchy(context.Context) (string, error) {
read := d.reads.Add(1)
composing := d.everyRead || read == d.composingRead ||
(read >= d.composingRead && read <= d.composingThrough)
if composing {
return fmt.Sprintf(homeWithTxnCountComposing, d.committed.Load()), nil
}
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), nil
}
// The same tree with one more row in it, which is what the reread sees while
// the screen is still filling in.
const homeWithTxnCountComposing = `{"attributes":{"resource-id":"HomeScreen"},"children":[
{"attributes":{"resource-id":"TxnCount","text":"%d"},"children":[]},
{"attributes":{"resource-id":"TxnSubmit","bounds":"[40,80,240,160]"},"children":[],"clickable":true,"enabled":true},
{"attributes":{"resource-id":"TxnRowLate"},"children":[]}
]}`
// Skipping a step is only free if nothing the spec needs goes missing with it.
// The action a step applies is reported to the spec on the NEXT step the
// verifier accepts, so a skipped step in between swallows the action before it:
// the transaction it committed still turns up in the next reading, and
// submitCommitsOneTransactionPerAction sees a rise nothing in its window
// accounts for. That is the conviction #77 and #78 are about, arriving through
// the skip rather than through the runner's report.
//
// So a frame the verifier will not look at is not one to act on either, which
// is also what #75 asked for: the fuzzer must not tap into a screen that is
// still filling in.
func TestRunner_ASkippedStepDoesNotSwallowTheActionBeforeIt(t *testing.T) {
spec := specWithFolioPredicates(t)
run := func(t *testing.T, composingRead int64) (Summary, int64) {
t.Helper()
state := newHarnessWithSpec(t, spec)
device := &submitsOnTapDriver{
Driver: state.mock,
commitsPerTap: 1,
composingRead: composingRead,
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 3 {
t.Fatalf("steps = %d, want 3", summary.Steps)
}
return summary, device.committed.Load()
}
t.Run("a submit is not lost to the step that follows it", func(t *testing.T) {
summary, committed := run(t, 2)
if summary.SkippedVerification != 1 {
t.Fatalf("the run skipped %d step(s), want 1; the reread never fired, so this "+
"proves nothing", summary.SkippedVerification)
}
if committed == 0 {
t.Fatal("the device committed nothing; a runner that never acts passes this " +
"test without meaning anything")
}
if len(summary.Violations) != 0 {
t.Errorf("the counting property convicted a healthy app: %v\n"+
"one transaction per submit rose, and a submit went unreported because "+
"the step after it was skipped", summary.Violations)
}
})
// The control: with nothing composing, every step is verified and the same
// app is judged clean, so the case above is not just a runner that stopped
// judging.
t.Run("every step verified, same app, no violation", func(t *testing.T) {
summary, committed := run(t, 0)
if summary.SkippedVerification != 0 {
t.Fatalf("the run skipped %d step(s), want 0", summary.SkippedVerification)
}
if committed != 3 {
t.Fatalf("the device committed %d transaction(s), want 3", committed)
}
if len(summary.Violations) != 0 {
t.Errorf("the counting property convicted a healthy app: %v", summary.Violations)
}
})
}
// One skipped step is held; a run of them has to be held too. A bound that lets
// the runner act again while the verifier is still being skipped puts back the
// exact swallow the hold exists to prevent, only later: the action drawn on the
// step past the bound overwrites the one the hold was carrying, and the carried
// action is never reported to any spec.
//
// The screen composes on steps 3 through 5 of 6 and the device commits one
// transaction per tap throughout, so the property has a clean pair to judge
// (step 2 to step 6) and nothing in between it can be told about except the
// action step 2 applied.
func TestRunner_ARunOfSkippedStepsReportsEveryActionItApplied(t *testing.T) {
spec := specWithFolioPredicates(t)
run := func(t *testing.T, commitsPerTap int64) (Summary, int64) {
t.Helper()
state := newHarnessWithSpec(t, spec)
device := &submitsOnTapDriver{
Driver: state.mock,
commitsPerTap: commitsPerTap,
composingRead: 3,
composingThrough: 5,
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 6,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 6 {
t.Fatalf("steps = %d, want 6", summary.Steps)
}
if summary.SkippedVerification != 3 {
t.Fatalf("the run skipped %d step(s), want 3; the reread never fired across "+
"the run this test is about", summary.SkippedVerification)
}
return summary, device.committed.Load()
}
t.Run("no submit is lost to the run of skipped steps", func(t *testing.T) {
summary, committed := run(t, 1)
if committed == 0 {
t.Fatal("the device committed nothing; a runner that never acts passes this " +
"test without meaning anything")
}
if len(summary.Violations) != 0 {
t.Errorf("the counting property convicted a healthy app: %v\n"+
"one transaction per submit rose, and a submit went unreported because "+
"the runner acted on a step the verifier skipped", summary.Violations)
}
if committed != 3 {
t.Errorf("the device committed %d transaction(s), want 3: one per verified "+
"step (1, 2 and 6) and none from a step nothing would judge", committed)
}
})
// The control. Without it a green above proves nothing: a property handed
// no comparable pair is silently vacuous and reports the same empty list.
t.Run("two transactions per tap still convicts across the same run", func(t *testing.T) {
summary, _ := run(t, 2)
if len(summary.Violations) == 0 {
t.Fatal("the counting property missed a double submit; the skipped steps left " +
"it with nothing to judge, so the case above proves nothing")
}
if got := summary.Violations[0].Properties[0]; got != "submitCommitsOneTransactionPerAction" {
t.Errorf("violated %v, want submitCommitsOneTransactionPerAction",
summary.Violations[0].Properties)
}
})
}
// A screen that changes shape under every pair of reads (a live list, a spinner
// mounting and unmounting) costs the run its actions: an action applied onto it
// would be the one the next verified step never hears about, and there is no
// next verified step. What the run must not do is come back green off that,
// which is what the "judged by nothing" count and the run's outcome are for.
func TestRunner_AScreenThatNeverSettlesActsOnNothingAndSaysSo(t *testing.T) {
state := newHarnessWithSpec(t, specWithFolioPredicates(t))
device := &submitsOnTapDriver{Driver: state.mock, commitsPerTap: 1, everyRead: true}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 5,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 5 {
t.Fatalf("steps = %d, want 5; the run stalled instead of finishing its budget",
summary.Steps)
}
if summary.SkippedVerification != 5 {
t.Fatalf("the run verified some step of a screen that never settled: skipped %d of 5",
summary.SkippedVerification)
}
if got := device.committed.Load(); got != 0 {
t.Errorf("the fuzzer applied %d action(s) onto a screen no property would judge; "+
"each one is an action no spec will ever be told about", got)
}
}
// homeWithTicker is one settled route holding a total that ticks and a button
// whose measured bounds shift under it. The nodes, their ids and their classes
// are the same in every rendering of it.
const homeWithTicker = `{"attributes":{"resource-id":"HomeScreen","class":"android.view.View"},"children":[
{"attributes":{"resource-id":"Total","class":"android.widget.TextView","text":"%s"},"children":[]},
{"attributes":{"resource-id":"TxnSubmit","class":"android.widget.Button","bounds":"%s"},"children":[]}
]}`
// The same route with a node in it that was not there a read ago.
const homeWithTickerAndRow = `{"attributes":{"resource-id":"HomeScreen","class":"android.view.View"},"children":[
{"attributes":{"resource-id":"Total","class":"android.widget.TextView","text":"120.00"},"children":[]},
{"attributes":{"resource-id":"TxnSubmit","class":"android.widget.Button","bounds":"[40,80,240,160]"},"children":[]},
{"attributes":{"resource-id":"TxnRowLate","class":"android.view.View"},"children":[]}
]}`
// rereadsDriver answers the paired Snapshot with one fixed tree and the
// hierarchy read that follows with another, so a test can say exactly what
// moved between the two reads the detector compares.
type rereadsDriver struct {
*mockdriver.Driver
snapshotTree string
rereadTree string
}
func (d *rereadsDriver) Snapshot(context.Context) (string, driver.Image, error) {
return d.snapshotTree, driver.Image{}, nil
}
func (d *rereadsDriver) Hierarchy(context.Context) (string, error) {
return d.rereadTree, nil
}
// What the two reads are compared ON is the whole feature. Comparing the values
// in the tree instead of the nodes in it would fire on every step of a screen
// with a total on it or a measure pass in flight, and a runner that skips every
// step verifies nothing while reporting no violations: green and vacuous, which
// is a worse answer than the composition the comparison set out to catch.
//
// The always-false property is the witness: it fires on the first step that
// reaches the verifier, so its presence and its step index say whether the step
// was judged at all.
func TestRunner_OnlyAChangeOfShapeCostsAStepItsVerdict(t *testing.T) {
settled := fmt.Sprintf(homeWithTicker, "120.00", "[40,80,240,160]")
cases := []struct {
name string
reread string
verified bool
}{
{
name: "a total that ticked between the two reads",
reread: fmt.Sprintf(homeWithTicker, "121.00", "[40,80,240,160]"),
verified: true,
},
{
name: "a measure pass that moved the button",
reread: fmt.Sprintf(homeWithTicker, "120.00", "[40,84,240,164]"),
verified: true,
},
{
name: "a node that was not in the tree a read ago",
reread: homeWithTickerAndRow,
verified: false,
},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
device := &rereadsDriver{
Driver: state.mock,
snapshotTree: settled,
rereadTree: testCase.reread,
}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 3 {
t.Fatalf("steps = %d, want 3", summary.Steps)
}
if !testCase.verified {
if summary.SkippedVerification != 3 {
t.Errorf("the run judged %d of 3 steps whose tree grew a node between "+
"the two reads; a screen still composing must reach no property",
3-summary.SkippedVerification)
}
if len(summary.Violations) != 0 {
t.Errorf("a skipped step reached the verifier anyway: %v",
summary.Violations)
}
return
}
if summary.SkippedVerification != 0 {
t.Fatalf("the run judged nothing: %d of 3 steps were skipped over a tree "+
"whose nodes never changed", summary.SkippedVerification)
}
if len(summary.Violations) != 1 {
t.Fatalf("violations = %v, want exactly one", summary.Violations)
}
if got := summary.Violations[0].StepIndex; got != 1 {
t.Errorf("the property first judged step %d, want 1", got)
}
})
}
}
type traceLine struct {
Step int `json:"step"`
Violations []string `json:"violations"`
ExtractorChanges map[string]trace.ExtractorChange `json:"extractor_changes"`
}
func traceSteps(t *testing.T, directory string) []traceLine {
t.Helper()
body, err := os.ReadFile(filepath.Join(directory, "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
var steps []traceLine
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
steps = append(steps, line)
}
return steps
}
@@ -0,0 +1,365 @@
package runner
import (
"context"
"fmt"
"path/filepath"
"sync/atomic"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
)
const guardedBundleID = "app.folio"
// committingDevice is a device whose submit taps commit transactions the next
// hierarchy read shows, and which can say how many it has committed so a test
// can prove the taps landed before reading anything into a verdict.
type committingDevice interface {
driver.DeviceDriver
commits() int64
}
// leavesForegroundAfterSubmitDriver is the condition the app-scope guard exists
// for: the submit tap lands and commits, and the app is no longer the
// foreground app by the time the next step looks. Folio's transactions are in
// sqlite, so the commit survives the relaunch and the next reading shows it.
type leavesForegroundAfterSubmitDriver struct {
*mockdriver.Driver
commitsPerTap int64
committed atomic.Int64
away atomic.Bool
}
func (d *leavesForegroundAfterSubmitDriver) Tap(context.Context, int, int) error {
return d.commitThenLeave()
}
func (d *leavesForegroundAfterSubmitDriver) TapSelector(context.Context, string) error {
return d.commitThenLeave()
}
func (d *leavesForegroundAfterSubmitDriver) commitThenLeave() error {
d.committed.Add(d.commitsPerTap)
d.away.Store(true)
return nil
}
func (d *leavesForegroundAfterSubmitDriver) commits() int64 { return d.committed.Load() }
func (d *leavesForegroundAfterSubmitDriver) Launch(
ctx context.Context,
bundleID string,
clearState bool,
env map[string]string,
) error {
d.away.Store(false)
return d.Driver.Launch(ctx, bundleID, clearState, env)
}
func (d *leavesForegroundAfterSubmitDriver) ForegroundApp(context.Context) (string, error) {
if d.away.Load() {
return "com.android.launcher", nil
}
return guardedBundleID, nil
}
func (d *leavesForegroundAfterSubmitDriver) FocusedWindowApp(ctx context.Context) (string, error) {
return d.ForegroundApp(ctx)
}
func (d *leavesForegroundAfterSubmitDriver) Snapshot(context.Context) (string, driver.Image, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), driver.Image{}, nil
}
// The runner reads both per step and compares them, so a device that answered
// them off different trees would make every step of this test transitional and
// judged by nothing. The sidecar serves both off one read path (snapshotTree)
// for the same reason; this one answers them off the same commit count.
func (d *leavesForegroundAfterSubmitDriver) Hierarchy(context.Context) (string, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), nil
}
// obscuredAfterSubmitDriver is the other half of the same guard: the app stays
// the resumed activity, but a system window (the notification shade) owns the
// focused window when the next step looks, and the guard presses back to
// collapse it.
type obscuredAfterSubmitDriver struct {
*mockdriver.Driver
commitsPerTap int64
committed atomic.Int64
obscured atomic.Bool
}
func (d *obscuredAfterSubmitDriver) Tap(context.Context, int, int) error {
return d.commitThenObscure()
}
func (d *obscuredAfterSubmitDriver) TapSelector(context.Context, string) error {
return d.commitThenObscure()
}
func (d *obscuredAfterSubmitDriver) commitThenObscure() error {
d.committed.Add(d.commitsPerTap)
d.obscured.Store(true)
return nil
}
func (d *obscuredAfterSubmitDriver) commits() int64 { return d.committed.Load() }
func (d *obscuredAfterSubmitDriver) PressKey(ctx context.Context, key string) error {
if key == "back" {
d.obscured.Store(false)
}
return d.Driver.PressKey(ctx, key)
}
func (d *obscuredAfterSubmitDriver) ForegroundApp(context.Context) (string, error) {
return guardedBundleID, nil
}
func (d *obscuredAfterSubmitDriver) FocusedWindowApp(context.Context) (string, error) {
if d.obscured.Load() {
return "com.android.systemui", nil
}
return guardedBundleID, nil
}
func (d *obscuredAfterSubmitDriver) Snapshot(context.Context) (string, driver.Image, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), driver.Image{}, nil
}
func (d *obscuredAfterSubmitDriver) Hierarchy(context.Context) (string, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), nil
}
// runTwoSubmitSteps drives two steps of the shipped folio counting property
// against a device that commits on every tap, and hands back what the property
// decided. Both steps have to run: the first arms the comparison, the second is
// where the guard fires and the pair is judged.
func runTwoSubmitSteps(
t *testing.T,
state *harness,
device committingDevice,
commitsPerTap int64,
) []ViolationRecord {
t.Helper()
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
BundleID: guardedBundleID,
Driver: device,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps != 2 {
t.Fatalf("steps = %d, want 2; the run never reached the step that judges the pair",
summary.Steps)
}
if got := device.commits(); got != commitsPerTap*2 {
t.Fatalf("the device committed %d transaction(s), want %d; the taps never reached it",
got, commitsPerTap*2)
}
return summary.Violations
}
func countMockActions(state *harness, kind mockdriver.ActionKind, key string) int {
count := 0
for _, action := range state.mock.Actions() {
if action.Kind != kind {
continue
}
if key != "" && action.Key != key {
continue
}
count++
}
return count
}
func specWithFolioPredicates(t *testing.T) string {
t.Helper()
predicates, err := filepath.Abs("../../examples/folio/sanderling/predicates.ts")
if err != nil {
t.Fatal(err)
}
return fmt.Sprintf(submitCountingSpecTemplate, predicates)
}
// A relaunch is not proof that nothing ran before it. The submit was dispatched
// and confirmed; what the relaunch changed is that the app restarted between
// the two readings the property compares. Reporting "no action" for it hands
// submitCommitsOneTransactionPerAction a transaction rise of one against a
// window of zero submits, which is the conviction #77 fixed for the apply-error
// path, manufactured here out of the scope guard instead.
func TestRunner_ARelaunchDoesNotConvictTheSubmitCountingProperty(t *testing.T) {
spec := specWithFolioPredicates(t)
run := func(t *testing.T, commitsPerTap int64) []ViolationRecord {
t.Helper()
state := newHarnessWithSpec(t, spec)
device := &leavesForegroundAfterSubmitDriver{
Driver: state.mock,
commitsPerTap: commitsPerTap,
}
violations := runTwoSubmitSteps(t, state, device, commitsPerTap)
if countMockActions(state, mockdriver.ActionLaunch, "") == 0 {
t.Fatal("the app was never relaunched, so the guard this test is about never ran")
}
return violations
}
t.Run("one transaction per tap is not a double submit", func(t *testing.T) {
if violations := run(t, 1); len(violations) != 0 {
t.Errorf("the counting property convicted a healthy app: %v\n"+
"one transaction rose against a submit the runner confirmed, and the "+
"spec was told no action happened because the app was relaunched",
violations)
}
})
// The control. Without it a green above proves nothing: a property that
// never sees a comparable pair is silently vacuous and reports the same
// empty violation list.
t.Run("two transactions per tap still convicts", func(t *testing.T) {
violations := run(t, 2)
if len(violations) == 0 {
t.Fatal("the counting property missed a double submit; the harness never " +
"put the property in a position to fire, so the case above proves nothing")
}
if violations[0].Properties[0] != "submitCommitsOneTransactionPerAction" {
t.Errorf("violated %v, want submitCommitsOneTransactionPerAction", violations[0].Properties)
}
})
}
// The same hole through the guard's other branch. A system window holding the
// focus says nothing about whether the tap under it ran: it was dispatched, and
// what nobody can say afterwards is whether the app received it. That is the
// unknown `applied` already carries, and it counts toward the submits a window
// could hold. Reporting no action instead convicts the app of a transaction
// with no cause.
func TestRunner_AnOverlayDoesNotConvictTheSubmitCountingProperty(t *testing.T) {
spec := specWithFolioPredicates(t)
run := func(t *testing.T, commitsPerTap int64) []ViolationRecord {
t.Helper()
state := newHarnessWithSpec(t, spec)
device := &obscuredAfterSubmitDriver{
Driver: state.mock,
commitsPerTap: commitsPerTap,
}
violations := runTwoSubmitSteps(t, state, device, commitsPerTap)
if countMockActions(state, mockdriver.ActionPressKey, "back") == 0 {
t.Fatal("the overlay was never dismissed, so the guard this test is about never ran")
}
if countMockActions(state, mockdriver.ActionLaunch, "") != 0 {
t.Fatal("a resumed-but-obscured app must not be relaunched")
}
return violations
}
t.Run("one transaction per tap is not a double submit", func(t *testing.T) {
if violations := run(t, 1); len(violations) != 0 {
t.Errorf("the counting property convicted a healthy app: %v\n"+
"one transaction rose against a submit the runner dispatched, and the "+
"spec was told no action happened because a system window took the focus",
violations)
}
})
t.Run("two transactions per tap still convicts", func(t *testing.T) {
violations := run(t, 2)
if len(violations) == 0 {
t.Fatal("the counting property missed a double submit; the harness never " +
"put the property in a position to fire, so the case above proves nothing")
}
if violations[0].Properties[0] != "submitCommitsOneTransactionPerAction" {
t.Errorf("violated %v, want submitCommitsOneTransactionPerAction", violations[0].Properties)
}
})
}
// reportedActionSpec puts what the runner told the spec about the last action
// into an extractor, so a test can read it out of the trace. `applied` and
// `relaunched` have no other producer: the runner's two guard writes are the
// only thing that ever sets them, and every spec-side guard built on them (see
// acrossRelaunch and confirmedApplied in the folio predicates) reads nothing
// else. A regression in either write leaves those guards permanently off with
// no property anywhere able to notice.
const reportedActionSpec = `
import { actions, always, extract, Tap } from "@sanderling/spec";
const reportedAction = extract("reportedAction", state => {
const last = state.lastAction;
if (last == null) return "none";
const dispatch = last.applied === true ? "applied" : "unconfirmed";
const process = last.relaunched === true ? "relaunched" : "same-process";
return dispatch + "/" + process;
});
globalThis.properties = {
theGuardTheRunnerRanReachesTheSpec: always(
() => reportedAction.current !== "applied/same-process",
),
};
globalThis.actions = actions(() => [Tap({ on: "id:TxnSubmit" })]);
`
// runReportingTheGuard drives two steps against a device whose submit tap trips
// one of the foreground guards, and hands back what the spec read off
// state.lastAction on the step the guard fired.
func runReportingTheGuard(t *testing.T, device committingDevice, state *harness) string {
t.Helper()
if violations := runTwoSubmitSteps(t, state, device, 1); len(violations) != 0 {
t.Errorf("the spec was told the action ran untouched by any guard: %v", violations)
}
steps := traceSteps(t, state.writer.Directory())
if len(steps) != 2 {
t.Fatalf("trace holds %d step(s), want 2", len(steps))
}
change, ok := steps[1].ExtractorChanges["reportedAction"]
if !ok {
t.Fatalf("step 2 recorded no reading of the reported action: %+v", steps[1])
}
return string(change.Curr)
}
func TestRunner_TheSpecIsToldTheAppWasRelaunchedUnderTheAction(t *testing.T) {
state := newHarnessWithSpec(t, reportedActionSpec)
device := &leavesForegroundAfterSubmitDriver{Driver: state.mock, commitsPerTap: 1}
reported := runReportingTheGuard(t, device, state)
if countMockActions(state, mockdriver.ActionLaunch, "") == 0 {
t.Fatal("the app was never relaunched, so the write this test is about never ran")
}
if reported != `"applied/relaunched"` {
t.Errorf("the spec read %s off state.lastAction, want \"applied/relaunched\"; "+
"a property relaxed across a relaunch cannot fire on a run that never "+
"tells it one happened", reported)
}
}
func TestRunner_TheSpecIsToldAnObscuredActionWasNotConfirmed(t *testing.T) {
state := newHarnessWithSpec(t, reportedActionSpec)
device := &obscuredAfterSubmitDriver{Driver: state.mock, commitsPerTap: 1}
reported := runReportingTheGuard(t, device, state)
if countMockActions(state, mockdriver.ActionPressKey, "back") == 0 {
t.Fatal("the overlay was never dismissed, so the write this test is about never ran")
}
if reported != `"unconfirmed/same-process"` {
t.Errorf("the spec read %s off state.lastAction, want "+
"\"unconfirmed/same-process\"; a system window held the focused window, "+
"so whether the app received the tap is exactly what nobody can say",
reported)
}
}
+234 -35
View File
@@ -52,6 +52,11 @@ type Summary struct {
EndTime time.Time
Steps int
Violations []ViolationRecord
// SkippedVerification counts the steps whose tree was still moving when it
// was read, so no property judged them. A green run that skipped most of
// its steps checked almost nothing, and nothing else in the output would
// say so.
SkippedVerification int
// UnsupportedVerbs lists verbs the picker requested that the platform
// could not dispatch, deduped, so the report can flag a spec exercising
// gestures this target does not support.
@@ -89,6 +94,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
return Summary{}, err
}
_, pageExtractors := extractorSource.(webSource)
rereadHierarchy := driverIsAndroid(ctx, options, logger)
summary := Summary{StartTime: time.Now()}
deadline := summary.StartTime.Add(options.Duration)
@@ -110,8 +116,23 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// backed out of (or otherwise left) the app, relaunch it before we
// observe or act, so properties never evaluate against a foreign app
// and actions never land outside the app.
if ensureForeground(ctx, options, logger, stepIndex) {
lastAction = nil
//
// What the guard did is reported to the spec on the action it followed,
// because dropping that action says "nothing ran between these two
// readings" and the runner has no business saying that: the action ran,
// and a property told otherwise convicts the app of an effect with no
// cause. See foreground_guard_last_action_test.go.
guard := ensureForeground(ctx, options, logger, stepIndex)
if lastAction != nil {
switch guard {
case foregroundRelaunched:
lastAction.Relaunched = true
case foregroundOverlayDismissed:
// A system window owned the focused window, so whether the app
// itself ever received this action is exactly the unknown
// Applied already has a state for.
lastAction.Applied = false
}
}
// Hierarchy, metrics, and logs are independent device reads. Run
@@ -132,7 +153,8 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// screenshot describe the same frame, then re-fetches the pair
// while the tree still looks transitional.
g.Go(func() error {
tree, screenshotPNG, transitional, hierarchyErr = fetchSyncedState(gctx, options, logger, si)
tree, screenshotPNG, transitional, hierarchyErr = fetchSyncedState(
gctx, options, logger, si, rereadHierarchy)
return nil
})
g.Go(func() error {
@@ -141,7 +163,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
})
logSince := lastLogTime
g.Go(func() error {
logs = collectLogs(gctx, options.Driver, logSince)
logs = collectLogs(gctx, options.Driver, logger, si, logSince)
return nil
})
// All goroutines write to local variables and return nil, so the Wait
@@ -173,8 +195,10 @@ func Run(ctx context.Context, options Options) (Summary, error) {
screen = tree.Elements[0].Screen
}
// Transitional trees describe a NavHost mid cross-fade. Pushing
// one would poison the verifier's previous/current extractor
// A transitional tree is one nothing can vouch for: a NavHost mid
// cross-fade, a screen that changed shape between two reads, or a
// hierarchy that came back empty. Pushing one would poison the
// verifier's previous/current extractor
// advance, so the next clean step would compare against this
// transient state and emit false-positive violations. We still
// record the step (hierarchy + screenshot) for replay-side
@@ -198,9 +222,10 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// behind the hierarchy fetch; the fetch is what decides whether this
// step counts at all, so it has to go first.
//
// lastAction is the same value PushSnapshot hands the goja state
// below: the two engines evaluate this step against one action.
v8Overrides, overridesErr := extractorSource.ExtractorOverrides(ctx, lastAction)
// lastAction and logs are the same values PushSnapshot hands the
// goja state below: the two engines evaluate this step against one
// action and one set of log entries.
v8Overrides, overridesErr := extractorSource.ExtractorOverrides(ctx, lastAction, logs)
if overridesErr != nil {
// Not a warning. Without the page's values this step's
// extractors keep goja's dump-derived readings while the
@@ -247,18 +272,44 @@ func Run(ctx context.Context, options Options) (Summary, error) {
extractorChanges = encodeExtractorChanges(options.Verifier.ChangedExtractors())
} else {
skippedVerification = true
logger.Warn("transitional tree after retry budget; skipping verifier",
summary.SkippedVerification++
logger.Warn("unsettled tree; skipping verifier",
"step", stepIndex, "screen", screen, "nodes", treeSize)
}
logger.Info("step", "index", stepIndex, "screen", screen, "nodes", treeSize)
nextAction, nextErr := actionSource.NextAction(ctx)
// A frame the verifier would not look at is not one to act on either.
// #75 is the fuzzer tapping into a screen that is still filling in, and
// holding the action back is also what keeps the spec's view of the run
// continuous: the action a step applies is reported on the NEXT step the
// verifier accepts, so acting here would leave the action applied last
// step unreported for good, and a property counting actions against
// their effects would then see an effect whose cause the runner
// swallowed. See TestRunner_ASkippedStepDoesNotSwallowTheActionBeforeIt.
//
// Unbounded, because lastAction holds exactly one action: any bound that
// let the runner act again while the verifier was still being skipped
// would overwrite the action the hold was carrying, and that is the same
// swallow arriving one step later. A screen that keeps moving therefore
// costs the run its actions rather than its soundness, and a run that
// verified nothing says so in its outcome (internal/testrun).
held := skippedVerification
if held {
logger.Warn("screen still moving; holding this step's action back",
"step", stepIndex)
}
var nextAction verifier.Action
nextErr := verifier.ErrNoAction
var traceAction *trace.Action
if nextErr == nil {
traceAction = traceActionFor(nextAction, tree)
stampActionSource(traceAction, actionSource)
} else if !errors.Is(nextErr, verifier.ErrNoAction) {
return summary, fmt.Errorf("step %d next action: %w", stepIndex, nextErr)
if !held {
nextAction, nextErr = actionSource.NextAction(ctx)
if nextErr == nil {
traceAction = traceActionFor(nextAction, tree)
stampActionSource(traceAction, actionSource)
} else if !errors.Is(nextErr, verifier.ErrNoAction) {
return summary, fmt.Errorf("step %d next action: %w", stepIndex, nextErr)
}
}
residuals, residualErr := encodeResiduals(options.Verifier.Residuals())
@@ -266,7 +317,7 @@ func Run(ctx context.Context, options Options) (Summary, error) {
logger.Warn("residual encode failed", "step", stepIndex, "err", residualErr)
}
applySkipped := false
applySkipped := held
if nextErr == nil && !appIsForeground(ctx, options) {
// The app left the foreground between observe and apply (a prior
// action's gesture settling late, or an async navigation). The
@@ -311,9 +362,12 @@ func Run(ctx context.Context, options Options) (Summary, error) {
applied.Applied = true
lastAction = &applied
}
} else {
} else if !held {
lastAction = nil
}
// A held step leaves lastAction alone on purpose: nothing ran here, and
// the action it points at is still the one the next verified step has to
// be told about.
step := trace.Step{
Index: stepIndex,
@@ -347,7 +401,16 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// concurrent fetches observe a stable post-action state. A transient
// apply error means nothing landed, so the idle poll has nothing to
// settle and may itself hang on the same device condition.
if nextErr == nil && !applySkipped && nextAction.Kind != verifier.ActionKindWait {
//
// A held step settles too, and it is the only case here that waits with
// nothing applied. The reread that held it takes its two reads a round
// trip apart, which is a tighter window than the one the detector was
// measured over (an action and a settle); looping straight back into it
// would compare two reads of a composing screen closer together still,
// so the screen that most needs to settle is the one given least room.
mutated := nextErr == nil && !applySkipped &&
nextAction.Kind != verifier.ActionKindWait
if held || mutated {
idleCtx, idleCancel := context.WithTimeout(ctx, options.IdleTimeout)
idleErr := options.Driver.WaitForIdle(idleCtx, options.IdleTimeout)
if idleErr != nil && idleCtx.Err() == nil {
@@ -396,6 +459,10 @@ func RenderSummary(w io.Writer, summary Summary, platform string) {
fmt.Fprintf(w, " step %d: %v\n", violation.StepIndex, violation.Properties)
}
}
if summary.SkippedVerification > 0 {
fmt.Fprintf(w, "%d step(s) judged by nothing: the screen was still moving when it was read\n",
summary.SkippedVerification)
}
if len(summary.UnsupportedVerbs) > 0 {
fmt.Fprintf(w, "unsupported on %s: %s\n",
platform, strings.Join(summary.UnsupportedVerbs, ", "))
@@ -445,20 +512,38 @@ func resolveIdleTimeout(options Options) time.Duration {
return timeout
}
// foregroundGuard is what ensureForeground had to do to put the app back in
// front. The two interventions are separate values because they are separate
// facts about the action they follow: a relaunch leaves it confirmed but
// straddling a restart, while a system window holding the focus leaves it
// dispatched with no way to tell whether the app received it.
type foregroundGuard int
const (
foregroundIntact foregroundGuard = iota
foregroundOverlayDismissed
foregroundRelaunched
)
// ensureForeground keeps the app under test in the foreground. When the driver
// can report the foreground app and it no longer matches the bundle under test,
// the app is relaunched. Returns true when a relaunch happened so the caller
// can drop the now-stale lastAction. Drivers without ForegroundChecker (web,
// the app is relaunched. Reports what it did so the caller can pass that on to
// the spec through the previous action. Drivers without ForegroundChecker (web,
// iOS) are a no-op.
func ensureForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) bool {
func ensureForeground(
ctx context.Context,
options Options,
logger *slog.Logger,
stepIndex int,
) foregroundGuard {
checker, ok := options.Driver.(driver.ForegroundChecker)
if !ok || options.BundleID == "" {
return false
return foregroundIntact
}
foreground, err := checker.ForegroundApp(ctx)
if err != nil {
logger.Warn("foreground check failed", "step", stepIndex, "err", err)
return false
return foregroundIntact
}
if foreground != "" && foreground != options.BundleID {
logger.Warn("app left foreground; relaunching",
@@ -471,7 +556,7 @@ func ensureForeground(ctx context.Context, options Options, logger *slog.Logger,
// window, so it never acts outside the app no matter how slow the
// relaunch settles.
awaitForeground(ctx, options, logger, stepIndex)
return true
return foregroundRelaunched
}
// The app is the resumed activity, but a system overlay can still own the
// focused window while the app stays resumed: a fuzzer swipe starting in the
@@ -481,15 +566,15 @@ func ensureForeground(ctx context.Context, options Options, logger *slog.Logger,
// the app again.
focusChecker, hasFocus := options.Driver.(driver.FocusedWindowChecker)
if !hasFocus {
return false
return foregroundIntact
}
focused, err := focusChecker.FocusedWindowApp(ctx)
if err != nil {
logger.Warn("focus check failed", "step", stepIndex, "err", err)
return false
return foregroundIntact
}
if focused == "" || focused == options.BundleID {
return false
return foregroundIntact
}
logger.Warn("system window obscuring app; dismissing",
"step", stepIndex, "focused", focused, "want", options.BundleID)
@@ -497,7 +582,7 @@ func ensureForeground(ctx context.Context, options Options, logger *slog.Logger,
logger.Warn("dismiss overlay failed", "step", stepIndex, "err", err)
}
settleForForeground(ctx, options)
return true
return foregroundOverlayDismissed
}
// appIsForeground reports whether the app under test currently owns the
@@ -739,11 +824,23 @@ func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.A
}
// collectLogs pulls recent error-level log entries from the driver since the
// previous fetch. A failure is warned-on but not fatal: log capture is a
// best-effort observability channel, not a correctness dependency.
func collectLogs(ctx context.Context, drv driver.DeviceDriver, since time.Time) []verifier.LogEntry {
// previous fetch. A failure is warned-on but not fatal: one unreadable fetch on
// a flaky device should not end a run. It is not free either. This fetch is the
// whole evidence base for state.logs, so a step that could not make it leaves
// every log property (the default noLogcatErrors included) holding on an empty
// slice, and that has to be visible in the run's output rather than read as the
// app having logged nothing.
func collectLogs(
ctx context.Context,
drv driver.DeviceDriver,
logger *slog.Logger,
step int,
since time.Time,
) []verifier.LogEntry {
entries, err := drv.RecentLogs(ctx, since, "E")
if err != nil {
logger.Warn("log fetch failed; log properties hold vacuously this step",
"step", step, "err", err)
return nil
}
result := make([]verifier.LogEntry, 0, len(entries))
@@ -931,10 +1028,17 @@ const (
// orthogonal case where the frame itself is transitional.
//
// The transitional return reports whether the retry budget was exhausted
// on a still-transitional tree. Callers use it to skip the verifier for
// that step so the previous/current extractor advance does not absorb
// on a still-transitional tree, or (when reread is set) whether a second
// hierarchy read disagreed with the first. Callers use it to skip the verifier
// for that step so the previous/current extractor advance does not absorb
// transient state.
func fetchSyncedState(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) (tree *hierarchy.Tree, png []byte, transitional bool, err error) {
func fetchSyncedState(
ctx context.Context,
options Options,
logger *slog.Logger,
stepIndex int,
reread bool,
) (tree *hierarchy.Tree, png []byte, transitional bool, err error) {
var pngBytes []byte
var previousJSON string
retryLoop:
@@ -970,6 +1074,9 @@ retryLoop:
case <-timer.C:
}
}
if reread && err == nil && !transitional && changedOnReread(ctx, options, logger, stepIndex, tree) {
transitional = true
}
if len(pngBytes) > 0 {
if writeErr := options.TraceWriter.WriteScreenshot(stepIndex, pngBytes); writeErr != nil {
logger.Warn("screenshot write failed", "step", stepIndex, "err", writeErr)
@@ -978,6 +1085,98 @@ retryLoop:
return tree, pngBytes, transitional, err
}
// changedOnReread reads the hierarchy once more and reports whether the screen
// changed shape while we were looking at it. A Compose route can settle before
// its content composes (a lazy list mounts over several frames, a query lands a
// frame late), and a tree read in that window describes a screen that is still
// filling in. Two reads a read apart are the cheapest thing that can see it
// happening: the round trip IS the interval, so there is no sleep here.
//
// The comparison only means anything because the Hierarchy RPC serves the tree
// the snapshot's own read produces (see snapshotTree in the sidecar). Off the
// bare device read it does not: with an IME standing open, the snapshot answers
// with 134 nodes and the bare read with 489, and the pair then differs over
// whether the sidecar closed a keyboard between them rather than over anything
// the app did.
//
// Waiting for the change to stop was measured on an API 34 device and refused:
// a 750ms-quiet poll capped at 2s cost a median 1434ms against 76ms for one
// read, hit its cap on every frame it fired for, and still handed back a frame
// that might be filling. Detecting is what the runner can act on, because a
// step it declines to verify is at worst a missed conviction, never a false
// one.
//
// A read that fails reports no change. Nothing about a dropped RPC says the
// screen was moving, and skipping verification on it would quietly spend the
// run's evidence on a flaky link.
func changedOnReread(
ctx context.Context,
options Options,
logger *slog.Logger,
stepIndex int,
first *hierarchy.Tree,
) bool {
// An empty tree is skipped by the caller anyway, so the read buys nothing.
if first == nil || len(first.Elements) == 0 {
return false
}
hierarchyJSON, err := options.Driver.Hierarchy(ctx)
if err != nil {
logger.Warn("second hierarchy read failed", "step", stepIndex, "err", err)
return false
}
second, err := hierarchy.Parse(hierarchyJSON)
if err != nil || second == nil {
logger.Warn("second hierarchy parse failed", "step", stepIndex, "err", err)
return false
}
if structuralShape(first) == structuralShape(second) {
return false
}
logger.Warn("screen changed between two reads; skipping verifier",
"step", stepIndex, "nodes", len(first.Elements), "then", len(second.Elements))
return true
}
// structuralShape renders what is on screen as its nodes' identities in tree
// order: how many there are, and which ids and classes they carry.
//
// Text and bounds are deliberately absent. A measure pass that moves pixels is
// not a screen still composing, and neither is a value arriving into a node
// that already exists, which this cannot tell apart from a clock ticking. This
// decides whether a property gets to judge at all, so it reads only what a
// change in what is on screen can move: a detector that fires on every step of
// a screen with a timer on it would leave the run green and vacuous, which is
// worse than the composition it set out to catch. The trade is measured rather
// than assumed: over 100 folio steps on an API 35 emulator, text moved under
// an unchanged shape on 1 step, and the shape itself moved on 1 other.
//
// TestRunner_OnlyAChangeOfShapeCostsAStepItsVerdict is what holds the line:
// adding either field back to the shape turns one of its cases red.
func structuralShape(tree *hierarchy.Tree) string {
var shape strings.Builder
for _, element := range tree.Elements {
shape.WriteString(element.ResourceID)
shape.WriteByte(0x1f)
shape.WriteString(element.Class)
shape.WriteByte(0x1e)
}
return shape.String()
}
// driverIsAndroid asks the driver what it is, once per run, so the step loop
// never repeats the RPC. It gates the reread: #75 is about Compose composition,
// and web and iOS have their own settle paths and no measurement saying an
// extra hierarchy read there is cheap. An unreadable answer is not android.
func driverIsAndroid(ctx context.Context, options Options, logger *slog.Logger) bool {
health, err := options.Driver.Health(ctx)
if err != nil {
logger.Warn("health read failed; not rereading the hierarchy", "err", err)
return false
}
return health.Platform == "android"
}
func traceActionFor(action verifier.Action, tree *hierarchy.Tree) *trace.Action {
traceAction := &trace.Action{Kind: string(action.Kind), X: action.X, Y: action.Y}
switch action.Kind {
+30 -4
View File
@@ -196,6 +196,23 @@ func TestRenderSummary_OmitsUnsupportedLineWhenNone(t *testing.T) {
}
}
// A step nothing judged is not a step that passed. The run prints its count so
// a green summary cannot hide a run that skipped most of its steps, which is
// what a screen that keeps moving under the reads would produce.
func TestRenderSummary_CountsTheStepsNothingJudged(t *testing.T) {
var out bytes.Buffer
RenderSummary(&out, Summary{Steps: 10, SkippedVerification: 4}, "android")
if !strings.Contains(out.String(), "4 step(s) judged by nothing") {
t.Errorf("expected the unjudged-step count, got:\n%s", out.String())
}
out.Reset()
RenderSummary(&out, Summary{Steps: 10}, "android")
if strings.Contains(out.String(), "judged by nothing") {
t.Errorf("a run that judged every step must not print the line, got:\n%s", out.String())
}
}
func TestRunner_ViolationSurfacesInSummary(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
@@ -1073,8 +1090,15 @@ func TestRunner_UsesAtomicSnapshot(t *testing.T) {
if snapshotCalls == 0 {
t.Errorf("expected at least one Snapshot call, got %d", snapshotCalls)
}
if hierarchyCalls != 0 {
t.Errorf("expected zero standalone Hierarchy calls (runner must use Snapshot), got %d", hierarchyCalls)
// The recorded pair still comes from Snapshot. The standalone hierarchy
// reads are the composition detector (changedOnReread), one per step at
// most, and they are never the source of what the step records.
if hierarchyCalls > summary.Steps {
t.Errorf("expected at most one standalone Hierarchy call per step (runner must observe through Snapshot), got %d over %d steps",
hierarchyCalls, summary.Steps)
}
if snapshotCalls < summary.Steps {
t.Errorf("expected a Snapshot per step, got %d over %d steps", snapshotCalls, summary.Steps)
}
if screenshotCalls != 0 {
t.Errorf("expected zero standalone Screenshot calls (runner must use Snapshot), got %d", screenshotCalls)
@@ -1894,8 +1918,10 @@ func TestEnsureForeground_DismissesSystemOverlay(t *testing.T) {
logger := slog.New(slog.NewTextHandler(io.Discard, &slog.HandlerOptions{Level: slog.LevelWarn}))
options := Options{BundleID: "app.folio", Driver: m, IdleTimeout: 10 * time.Millisecond}
if !ensureForeground(context.Background(), options, logger, 5) {
t.Fatal("expected the guard to act on the focus-stealing overlay")
got := ensureForeground(context.Background(), options, logger, 5)
if got != foregroundOverlayDismissed {
t.Fatalf("the guard reported %v, want foregroundOverlayDismissed; "+
"an obscured app is not a relaunched one", got)
}
backs, relaunches := 0, 0
for _, a := range m.Actions() {
+31 -11
View File
@@ -24,14 +24,15 @@ type ActionSource interface {
// PushSnapshot. The mobile path has none (returns nil); the web path returns the
// values its extractors computed in V8 against the real DOM.
//
// lastAction is the action the previous step actually applied, the same value
// PushSnapshot hands the goja state. The web path has to install it in the page
// before its extractors run: a spec extractor reading state.lastAction runs in
// V8 there, and V8 has no way to know what the runner dispatched.
// lastAction and logs are what PushSnapshot hands the goja state. The web path
// has to install both in the page before its extractors run: a spec extractor
// reading state.lastAction or state.logs runs in V8 there, and V8 knows neither
// what the runner dispatched nor what the driver's log fetch returned.
type ExtractorSource interface {
ExtractorOverrides(
ctx context.Context,
lastAction *verifier.Action,
logs []verifier.LogEntry,
) (map[int]json.RawMessage, error)
}
@@ -42,6 +43,13 @@ type lastActionInstaller interface {
SetLastAction(ctx context.Context, encoded json.RawMessage) error
}
// logInstaller is the same channel for the entries this step's log fetch
// returned. Console output reaches the driver over CDP, so the page can only
// learn about it from the runner.
type logInstaller interface {
SetLogs(ctx context.Context, encoded json.RawMessage) error
}
// gojaSource drives both action selection and (trivially) extractor overrides
// for the mobile path, where the goja-bundled picker runs in-process and no V8
// extractor values exist.
@@ -56,6 +64,7 @@ func (s gojaSource) NextAction(context.Context) (verifier.Action, error) {
func (gojaSource) ExtractorOverrides(
context.Context,
*verifier.Action,
[]verifier.LogEntry,
) (map[int]json.RawMessage, error) {
return nil, nil
}
@@ -77,24 +86,35 @@ func (s webSource) NextAction(ctx context.Context) (verifier.Action, error) {
return verifier.DecodeAction(raw)
}
// ExtractorOverrides installs the previous step's action in the page, then
// reads back what the spec's extractors computed against the live DOM. The
// install is not best-effort: a web driver that cannot take it leaves
// state.lastAction null in V8, which silently turns every action-gated
// property vacuously true, so it is reported as an error instead.
// ExtractorOverrides installs the previous step's action and this step's log
// entries in the page, then reads back what the spec's extractors computed
// against the live DOM. Neither install is best-effort: a web driver that
// cannot take them leaves state.lastAction null and state.logs empty in V8,
// which silently turns every action-gated property and every log property
// vacuously true, so both are reported as errors instead.
func (s webSource) ExtractorOverrides(
ctx context.Context,
lastAction *verifier.Action,
logs []verifier.LogEntry,
) (map[int]json.RawMessage, error) {
installer, ok := s.web.(lastActionInstaller)
actions, ok := s.web.(lastActionInstaller)
if !ok {
return nil, fmt.Errorf(
"web driver %T cannot install state.lastAction; every property gated "+
"on the last action would be vacuously true", s.web)
}
if err := installer.SetLastAction(ctx, verifier.EncodeLastAction(lastAction)); err != nil {
if err := actions.SetLastAction(ctx, verifier.EncodeLastAction(lastAction)); err != nil {
return nil, fmt.Errorf("install last action: %w", err)
}
entries, ok := s.web.(logInstaller)
if !ok {
return nil, fmt.Errorf(
"web driver %T cannot install state.logs; every property reading the "+
"log stream would be vacuously true", s.web)
}
if err := entries.SetLogs(ctx, verifier.EncodeLogs(logs)); err != nil {
return nil, fmt.Errorf("install logs: %w", err)
}
return s.web.EvaluateExtractors(ctx)
}
@@ -22,7 +22,7 @@ import (
//
// The spec below is the real folio predicate pair, imported from the example,
// so what this asserts is the verdict the shipped property reaches.
const uncertainApplySpecTemplate = `
const submitCountingSpecTemplate = `
import { actions, always, extract, next, Tap } from "@sanderling/spec";
import {
committedTransactionsExceedSubmits,
@@ -90,12 +90,16 @@ func (d *dispatchThenFailDriver) Snapshot(context.Context) (string, driver.Image
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), driver.Image{}, nil
}
func (d *dispatchThenFailDriver) Hierarchy(context.Context) (string, error) {
return fmt.Sprintf(homeWithTxnCount, d.committed.Load()), nil
}
func TestRunner_ApplyErrorAfterDispatchDoesNotConvictTheSubmitCountingProperty(t *testing.T) {
predicates, err := filepath.Abs("../../examples/folio/sanderling/predicates.ts")
if err != nil {
t.Fatal(err)
}
spec := fmt.Sprintf(uncertainApplySpecTemplate, predicates)
spec := fmt.Sprintf(submitCountingSpecTemplate, predicates)
run := func(t *testing.T, commitsPerTap int64) []ViolationRecord {
t.Helper()
+56
View File
@@ -55,6 +55,12 @@ func (d *carrierWebDriver) Snapshot(ctx context.Context) (string, driver.Image,
func (d *carrierWebDriver) InstallBundle(context.Context, []byte) error { return nil }
// A web target says so. The runner's per-step hierarchy reread is android-only,
// and a fake claiming android would take a path no chrome run takes.
func (d *carrierWebDriver) Health(context.Context) (driver.Health, error) {
return driver.Health{Ready: true, Version: "fake", Platform: "web"}, nil
}
func (d *carrierWebDriver) EvaluateExtractors(context.Context) (map[int]json.RawMessage, error) {
d.reads++
return map[int]json.RawMessage{0: json.RawMessage(strconv.Itoa(d.reads))}, nil
@@ -69,6 +75,8 @@ func (d *carrierWebDriver) NextActionFromV8(context.Context) (json.RawMessage, e
func (d *carrierWebDriver) SetLastAction(context.Context, json.RawMessage) error { return nil }
func (d *carrierWebDriver) SetLogs(context.Context, json.RawMessage) error { return nil }
// TestRunner_TransitionalStepNeverAdvancesThePageCarrier pins the ordering the
// web path depends on. The page-side extractors must run only on steps the
// verifier accepts: their getters advance spec state every time they evaluate,
@@ -166,6 +174,8 @@ func (d *installFailsWebDriver) SetLastAction(context.Context, json.RawMessage)
return errors.New("__sanderlingSetLastAction__ is not a function")
}
func (d *installFailsWebDriver) SetLogs(context.Context, json.RawMessage) error { return nil }
// TestRunner_LastActionInstallFailureFailsTheRun covers the other half of the
// same trust boundary. A run that cannot install lastAction in the page cannot
// apply the page's extractor values either, so the step keeps goja's
@@ -194,3 +204,49 @@ func TestRunner_LastActionInstallFailureFailsTheRun(t *testing.T) {
t.Errorf("Run error = %v, want it to name the failed lastAction install", err)
}
}
// logInstallFailsWebDriver takes lastAction and refuses the logs, the shape a
// page carrying an older published @sanderling/spec runtime has: it knows the
// action setter and not the log one.
type logInstallFailsWebDriver struct {
*installFailsWebDriver
}
func (d *logInstallFailsWebDriver) SetLastAction(context.Context, json.RawMessage) error {
return nil
}
func (d *logInstallFailsWebDriver) SetLogs(context.Context, json.RawMessage) error {
return errors.New("__sanderlingSetLogs__ is not a function")
}
// TestRunner_LogInstallFailureFailsTheRun holds the log channel to the same
// standard as the action one. The driver having the console errors decides
// nothing on web: the page's reading of every extractor replaces the host's, so
// a run that cannot put the entries back into the page evaluates noLogcatErrors
// against an empty array and reports green on a console full of errors.
// Continuing past this is the vacuity the whole install exists to prevent.
func TestRunner_LogInstallFailureFailsTheRun(t *testing.T) {
state := newHarnessWithSpec(t, carrierSpec)
web := &logInstallFailsWebDriver{
installFailsWebDriver: &installFailsWebDriver{Driver: state.mock},
}
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 2 * time.Second,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 3,
Driver: web,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err == nil {
t.Fatal("Run succeeded with a page that cannot take the step's logs; " +
"every property reading the log stream ran against an empty array")
}
if !bytes.Contains([]byte(err.Error()), []byte("install logs")) {
t.Errorf("Run error = %v, want it to name the failed log install", err)
}
}
@@ -47,6 +47,8 @@ func (d *webMockDriver) NextActionFromV8(context.Context) (json.RawMessage, erro
func (d *webMockDriver) SetLastAction(context.Context, json.RawMessage) error { return nil }
func (d *webMockDriver) SetLogs(context.Context, json.RawMessage) error { return nil }
// TestRunner_TraceRecordsTheValueTheVerdictUsed fails if the trace and the
// verdict disagree about an extractor. A witness is only an explanation of a
// violation if it holds the state the violated property was evaluated against.
+80 -3
View File
@@ -1,12 +1,16 @@
package runner
import (
"bytes"
"context"
"encoding/json"
"errors"
"log/slog"
"strings"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
)
@@ -26,7 +30,8 @@ globalThis.properties = {};
// control, so the runner has a real applied action to report on the next step.
type tappingWebDriver struct {
*mockdriver.Driver
installed []string
installed []string
installedLogs []string
}
func (d *tappingWebDriver) InstallBundle(context.Context, []byte) error { return nil }
@@ -44,6 +49,11 @@ func (d *tappingWebDriver) SetLastAction(_ context.Context, encoded json.RawMess
return nil
}
func (d *tappingWebDriver) SetLogs(_ context.Context, encoded json.RawMessage) error {
d.installedLogs = append(d.installedLogs, string(encoded))
return nil
}
func TestRunner_WebInstallsLastActionInThePage(t *testing.T) {
state := newHarnessWithSpec(t, lastActionSpec)
web := &tappingWebDriver{Driver: state.mock}
@@ -72,12 +82,79 @@ func TestRunner_WebInstallsLastActionInThePage(t *testing.T) {
// Every later step carries what the runner actually applied. The shape is
// the goja host's (internal/verifier/marshal.go lastActionFields), pinned
// against it by TestLastAction_WebJSONMatchesTheGojaObject.
const want = `{"kind":"Tap","applied":true,"on":"id:TxnSubmit"}`
const want = `{"kind":"Tap","applied":true,"relaunched":null,"on":"id:TxnSubmit"}`
if web.installed[1] != want {
t.Errorf("step 2 installed %s, want %s", web.installed[1], want)
}
}
// The same hole on the other channel: state.logs was hardcoded [] in
// pkg/spec/src/web-runtime.ts, and because the page's reading of an extractor
// replaces the host's on web, the driver's error-level entries never reached a
// property. The default noLogcatErrors counted an empty array on every run.
func TestRunner_WebInstallsTheStepsLogsInThePage(t *testing.T) {
state := newHarnessWithSpec(t, lastActionSpec)
state.mock.LogEntries = []driver.LogEntry{
{UnixMillis: 1700000000123, Level: "E", Tag: "console", Message: "boom from the page"},
}
web := &tappingWebDriver{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
Driver: web,
Verifier: state.verifier,
TraceWriter: state.writer,
}); err != nil {
t.Fatalf("Run: %v", err)
}
if len(web.installedLogs) == 0 {
t.Fatal("the page was never handed the step's logs; every property reading " +
"state.logs evaluated against the empty array the page starts with")
}
// The shape is the goja host's (internal/verifier/marshal.go logFields),
// pinned against it by TestLogs_WebJSONMatchesTheGojaObject.
const want = `[{"unixMillis":1700000000123,"level":"E","tag":"console","message":"boom from the page"}]`
if web.installedLogs[0] != want {
t.Errorf("step 1 installed %s, want %s", web.installedLogs[0], want)
}
}
// A log fetch that fails decides the verdict of every log property: they all
// evaluate against an empty slice and hold. That is not a fact about the app,
// so the step it happened on has to be visible in the run's output. It used to
// be dropped in silence, under a comment claiming it was warned about.
func TestRunner_ReportsALogFetchItCouldNotMake(t *testing.T) {
state := newHarnessWithSpec(t, lastActionSpec)
state.mock.Failures[mockdriver.ActionRecentLogs] = errors.New("adb: device offline")
var buffer bytes.Buffer
logger := slog.New(slog.NewTextHandler(&buffer, &slog.HandlerOptions{Level: slog.LevelWarn}))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if _, err := Run(ctx, Options{
Duration: time.Hour,
IdleTimeout: 20 * time.Millisecond,
MaxSteps: 2,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: logger,
}); err != nil {
t.Fatalf("Run: %v", err)
}
if !strings.Contains(buffer.String(), "adb: device offline") {
t.Errorf("the run never reported the failed log fetch, so noLogcatErrors "+
"held on evidence nobody collected; log was %q", buffer.String())
}
}
// failingTapWebDriver dispatches the tap and then fails the call, the shape an
// RPC deadline takes: the page has the click, the runner has an error.
type failingTapWebDriver struct {
@@ -112,7 +189,7 @@ func TestRunner_WebInstallsAnUnconfirmedActionWithItsFateUnknown(t *testing.T) {
t.Fatalf("the page was handed lastAction %d time(s); the web path never installed it",
len(web.installed))
}
const want = `{"kind":"Tap","applied":null,"on":"id:TxnSubmit"}`
const want = `{"kind":"Tap","applied":null,"relaunched":null,"on":"id:TxnSubmit"}`
if web.installed[1] != want {
t.Errorf("step 2 installed %s, want %s", web.installed[1], want)
}
+96 -22
View File
@@ -84,6 +84,17 @@ var newDeviceDriver = func(ctx context.Context, options ioscompanion.DeviceOptio
return d, d.Close, nil
}
// newSimulatorDriver constructs the iOS simulator driver and its cleanup. A
// seam so routing tests assert the run's options reach ioscompanion.Options
// without spawning a companion.
var newSimulatorDriver = func(ctx context.Context, options ioscompanion.Options) (driver.DeviceDriver, func(), error) {
d, err := ioscompanion.New(ctx, options)
if err != nil {
return nil, nil, err
}
return d, d.Close, nil
}
// buildDriver creates the appropriate DeviceDriver for the platform and returns
// a cleanup function. For web, ChromeDriver is used directly. An iOS simulator
// is driven by the native simulator companion (no JVM). A physical iOS device
@@ -99,16 +110,17 @@ func buildDriver(ctx context.Context, options Options, stdout io.Writer) (driver
}
if options.Platform == "ios" && options.iosIsSimulator {
d, err := ioscompanion.New(ctx, ioscompanion.Options{
d, cleanup, err := newSimulatorDriver(ctx, ioscompanion.Options{
UniqueDeviceIdentifier: options.iosUDID,
BundleID: options.BundleID,
AppPath: options.IosAppPath,
ClearState: options.ClearData,
Output: stdout,
})
if err != nil {
return nil, nil, fmt.Errorf("ios simulator driver: %w", err)
}
return d, d.Close, nil
return d, cleanup, nil
}
if options.Platform == "ios" {
@@ -117,6 +129,7 @@ func buildDriver(ctx context.Context, options Options, stdout io.Writer) (driver
CoreDeviceID: options.iosCoreDeviceID,
BundleID: options.BundleID,
AppPath: options.IosAppPath,
ClearState: options.ClearData,
Output: stdout,
})
if err != nil {
@@ -148,10 +161,14 @@ func buildDriver(ctx context.Context, options Options, stdout io.Writer) (driver
if options.Device != "" {
sidecarArgs = append(sidecarArgs, "--serial", options.Device)
}
adbPath, err := android.AdbBinary()
if err != nil {
return nil, nil, preflightFailure("android", err)
}
sidecarCommand := exec.CommandContext(ctx, "java", sidecarArgs...)
sidecarCommand.Stdout = stdout
sidecarCommand.Stderr = stdout
sidecarCommand.Env = android.EnvWithAndroidPlatformTools(os.Environ())
sidecarCommand.Env = android.EnvWithAndroidPlatformTools(os.Environ(), adbPath)
// SIGTERM lets the sidecar's shutdown hook stop the iOS XCTest runner.
// SIGKILL skips the hook and orphans an xcodebuild session that later
// restarts its runner and hijacks the simulator mid-run.
@@ -162,11 +179,13 @@ func buildDriver(ctx context.Context, options Options, stdout io.Writer) (driver
if err := sidecarCommand.Start(); err != nil {
return nil, nil, fmt.Errorf("spawn sidecar: %w", err)
}
fmt.Fprintf(stdout, "sidecar pid=%d listening on 127.0.0.1:%d\n", sidecarCommand.Process.Pid, sidecarPort)
sidecarExited := watchSidecar(sidecarCommand)
address := fmt.Sprintf("127.0.0.1:%d", sidecarPort)
fmt.Fprintf(stdout, "sidecar pid=%d listening on %s (adb: %s)\n", sidecarCommand.Process.Pid, address, adbPath)
driverClient, err := driverSidecar.Dial(fmt.Sprintf("127.0.0.1:%d", sidecarPort))
driverClient, err := driverSidecar.Dial(address)
if err != nil {
stopSidecar(sidecarCommand)
stopSidecar(sidecarCommand, sidecarExited)
return nil, nil, fmt.Errorf("dial sidecar: %w", err)
}
driverClient.SetPlatform(options.Platform)
@@ -175,22 +194,82 @@ func buildDriver(ctx context.Context, options Options, stdout io.Writer) (driver
// (absorbing the XCUITest startup race) runs inside IosDriverBackend.init
// in the sidecar - no additional sleep needed here.
healthCtx, healthCancel := context.WithTimeout(ctx, sidecarStartupTimeout)
if err := driverClient.WaitForHealth(healthCtx, 250e6); err != nil {
healthCancel()
stopSidecar(sidecarCommand)
_ = driverClient.Close()
return nil, nil, fmt.Errorf("sidecar health check: %w", err)
}
healthErr := awaitSidecar(healthCtx, address, sidecarStartupTimeout, func(pollCtx context.Context) error {
return driverClient.WaitForHealth(pollCtx, 250e6)
}, sidecarExited)
healthCancel()
if healthErr != nil {
stopSidecar(sidecarCommand, sidecarExited)
_ = driverClient.Close()
return nil, nil, healthErr
}
fmt.Fprintln(stdout, "sidecar is healthy")
cleanup := func() {
_ = driverClient.Close()
stopSidecar(sidecarCommand)
stopSidecar(sidecarCommand, sidecarExited)
}
return driverClient, cleanup, nil
}
// watchSidecar reaps the sidecar and publishes its exit status. The channel is
// closed after the send so the shutdown path can still receive once the startup
// path has taken the status.
func watchSidecar(sidecarCommand *exec.Cmd) <-chan error {
exited := make(chan error, 1)
go func() {
exited <- sidecarCommand.Wait()
close(exited)
}()
return exited
}
// awaitSidecar waits for the sidecar to answer a health check, racing that
// against the process exiting so a sidecar that dies during startup is reported
// as the exit it was rather than as a deadline half a minute later. Neither
// failure knows why the sidecar was unhappy, so both name what to look at
// instead of picking a cause.
func awaitSidecar(
ctx context.Context,
address string,
timeout time.Duration,
health func(context.Context) error,
exited <-chan error,
) error {
healthy := make(chan error, 1)
go func() { healthy <- health(ctx) }()
select {
case exitErr := <-exited:
return sidecarExitedError(address, exitErr)
case err := <-healthy:
if err == nil {
return nil
}
select {
case exitErr := <-exited:
return sidecarExitedError(address, exitErr)
default:
return fmt.Errorf(
"sidecar did not answer a health check on %s within %s and is still running\n%s",
address, timeout, sidecarWhatToCheck,
)
}
}
}
const sidecarWhatToCheck = "check the sidecar output above, then `sanderling doctor --platform=android` (java 17+, adb, Android SDK)"
func sidecarExitedError(address string, exitErr error) error {
status := "exit status 0"
if exitErr != nil {
status = exitErr.Error()
}
return fmt.Errorf(
"sidecar exited before it answered a health check on %s: %s\n%s",
address, status, sidecarWhatToCheck,
)
}
// sidecarShutdownGrace bounds how long the sidecar gets to run its shutdown
// hook (terminate the app, stop the XCTest runner) before being killed.
const sidecarShutdownGrace = 15 * time.Second
@@ -198,25 +277,20 @@ const sidecarShutdownGrace = 15 * time.Second
// stopSidecar terminates the sidecar gracefully so its shutdown hook can stop
// the device-side runner processes, escalating to SIGKILL when it does not
// exit within the grace window.
func stopSidecar(sidecarCommand *exec.Cmd) {
func stopSidecar(sidecarCommand *exec.Cmd, exited <-chan error) {
if sidecarCommand.Process == nil {
return
}
if err := sidecarCommand.Process.Signal(syscall.SIGTERM); err != nil {
_ = sidecarCommand.Process.Kill()
_ = sidecarCommand.Wait()
<-exited
return
}
done := make(chan struct{})
go func() {
_ = sidecarCommand.Wait()
close(done)
}()
select {
case <-done:
case <-exited:
case <-time.After(sidecarShutdownGrace):
_ = sidecarCommand.Process.Kill()
<-done
<-exited
}
}
+157 -1
View File
@@ -4,7 +4,10 @@ import (
"context"
"errors"
"io"
"os/exec"
"strings"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion"
@@ -27,7 +30,7 @@ func TestBuildDriverRoutesPhysicalIOSToDeviceDriver(t *testing.T) {
return stubDeviceDriver{}, func() { closed = true }, nil
}
options := Options{Platform: "ios", BundleID: "app.folio", IosAppPath: "/tmp/iosApp.app"}
options := Options{Platform: "ios", BundleID: "app.folio", IosAppPath: "/tmp/iosApp.app", ClearData: true}
options.iosIsSimulator = false
options.iosUDID = "00008140-HW"
options.iosCoreDeviceID = "CORE-1"
@@ -45,12 +48,41 @@ func TestBuildDriverRoutesPhysicalIOSToDeviceDriver(t *testing.T) {
if got.BundleID != "app.folio" || got.AppPath != "/tmp/iosApp.app" {
t.Fatalf("DeviceOptions = %+v, want bundle and app path threaded through", got)
}
if !got.ClearState {
t.Fatalf("DeviceOptions = %+v, want clear-data threaded through: the driver clears before its session, so a launch cannot", got)
}
cleanup()
if !closed {
t.Fatal("cleanup must close the device driver")
}
}
func TestBuildDriverThreadsClearStateToTheSimulatorDriver(t *testing.T) {
stubPreflight(t)
original := newSimulatorDriver
t.Cleanup(func() { newSimulatorDriver = original })
var got ioscompanion.Options
newSimulatorDriver = func(_ context.Context, options ioscompanion.Options) (driver.DeviceDriver, func(), error) {
got = options
return stubDeviceDriver{}, func() {}, nil
}
options := Options{Platform: "ios", BundleID: "app.folio", IosAppPath: "/tmp/iosApp.app", ClearData: true}
options.iosIsSimulator = true
options.iosUDID = "SIM-UDID"
if _, _, err := buildDriver(context.Background(), options, io.Discard); err != nil {
t.Fatalf("buildDriver: %v", err)
}
if got.UniqueDeviceIdentifier != "SIM-UDID" || got.BundleID != "app.folio" || got.AppPath != "/tmp/iosApp.app" {
t.Fatalf("Options = %+v, want the resolved target, bundle and app path", got)
}
if !got.ClearState {
t.Fatalf("Options = %+v, want clear-data threaded through: the driver clears before its session, so a launch cannot", got)
}
}
func TestBuildDriverSurfacesDeviceConstructionError(t *testing.T) {
stubPreflight(t)
original := newDeviceDriver
@@ -65,6 +97,130 @@ func TestBuildDriverSurfacesDeviceConstructionError(t *testing.T) {
}
}
// A sidecar that dies during startup leaves the health poll with nothing to
// talk to, and reporting that as a deadline sends the reader after a gRPC
// timeout instead of the exit that already happened.
func TestAwaitSidecarReportsTheExitItSaw(t *testing.T) {
exited := make(chan error, 1)
exited <- errors.New("exit status 1")
close(exited)
err := awaitSidecar(
context.Background(),
"127.0.0.1:54321",
30*time.Second,
func(ctx context.Context) error { <-ctx.Done(); return ctx.Err() },
exited,
)
if err == nil {
t.Fatal("expected an error when the sidecar exits before it is healthy")
}
for _, want := range []string{
"sidecar exited before it answered a health check on 127.0.0.1:54321: exit status 1",
"check the sidecar output above",
"sanderling doctor --platform=android",
} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q missing %q", err, want)
}
}
}
func TestAwaitSidecarTimeoutSaysOnlyWhatItObserved(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Millisecond)
defer cancel()
err := awaitSidecar(
ctx,
"127.0.0.1:54321",
50*time.Millisecond,
func(ctx context.Context) error { <-ctx.Done(); return ctx.Err() },
make(chan error, 1),
)
if err == nil {
t.Fatal("expected an error when the sidecar never answers")
}
for _, want := range []string{
"sidecar did not answer a health check on 127.0.0.1:54321 within 50ms and is still running",
"check the sidecar output above",
"sanderling doctor --platform=android",
} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q missing %q", err, want)
}
}
if strings.Contains(err.Error(), "context deadline exceeded") {
t.Errorf("error %q must not hand the reader a bare gRPC deadline", err)
}
}
func TestAwaitSidecarHealthyReturnsNil(t *testing.T) {
err := awaitSidecar(
context.Background(),
"127.0.0.1:54321",
30*time.Second,
func(context.Context) error { return nil },
make(chan error, 1),
)
if err != nil {
t.Fatalf("expected a healthy sidecar to pass, got %v", err)
}
}
func TestStopSidecarTerminatesARunningSidecar(t *testing.T) {
command := exec.Command("sleep", "60")
if err := command.Start(); err != nil {
t.Fatalf("start: %v", err)
}
exited := watchSidecar(command)
stopped := make(chan struct{})
go func() {
stopSidecar(command, exited)
close(stopped)
}()
select {
case <-stopped:
case <-time.After(10 * time.Second):
t.Fatal("stopSidecar never returned for a running sidecar")
}
if got := command.ProcessState.String(); got != "signal: terminated" {
t.Errorf("sidecar ended as %q, want the SIGTERM its shutdown hook needs", got)
}
}
// The startup path takes the exit status to report it, so the shutdown path
// that follows must not sit waiting for a status nobody will send again.
func TestStopSidecarAfterTheStartupPathTookTheExitStatus(t *testing.T) {
command := exec.Command("sh", "-c", "exit 3")
if err := command.Start(); err != nil {
t.Fatalf("start: %v", err)
}
exited := watchSidecar(command)
err := awaitSidecar(
context.Background(),
"127.0.0.1:54321",
30*time.Second,
func(ctx context.Context) error { <-ctx.Done(); return ctx.Err() },
exited,
)
if err == nil || !strings.Contains(err.Error(), "exit status 3") {
t.Fatalf("expected the sidecar's real exit status, got %v", err)
}
stopped := make(chan struct{})
go func() {
stopSidecar(command, exited)
close(stopped)
}()
select {
case <-stopped:
case <-time.After(10 * time.Second):
t.Fatal("stopSidecar blocked on an exit status the startup path had already taken")
}
}
// stubPreflight bypasses the host-readiness checks so routing tests exercise
// driver construction on a Linux CI runner that lacks xcrun/java.
func stubPreflight(t *testing.T) {
+10
View File
@@ -4,6 +4,8 @@ import (
"context"
"fmt"
"os/exec"
"github.com/priyanshujain/sanderling/internal/android"
)
// Preflight runs platform-specific host checks before sidecar/driver setup.
@@ -18,7 +20,15 @@ func Preflight(ctx context.Context, platform string) error {
type preflightFunc func(name string) error
// preflightCheck resolves adb through the same helper every adb call in a run
// uses, so a host whose SDK is only reachable through $ANDROID_HOME or a
// standard install location is not turned away here and then driven fine by
// the rest of the pipeline.
func preflightCheck(name string) error {
if name == "adb" {
_, err := android.AdbBinary()
return err
}
if _, err := exec.LookPath(name); err != nil {
return fmt.Errorf("%s not found on PATH: %w", name, err)
}
+29
View File
@@ -3,6 +3,8 @@ package testrun
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
@@ -49,6 +51,33 @@ func TestPreflight_AndroidNeedsAdbAndJava(t *testing.T) {
}
}
// Every adb call in an android run resolves through $ANDROID_HOME and the
// standard SDK locations, so a preflight that only looks at PATH turns away a
// host the run itself would drive.
func TestPreflight_AndroidAcceptsAdbUnderAndroidHome(t *testing.T) {
sdk := t.TempDir()
writeExecutable(t, filepath.Join(sdk, "platform-tools", "adb"))
pathDirectory := t.TempDir()
writeExecutable(t, filepath.Join(pathDirectory, "java"))
t.Setenv("PATH", pathDirectory)
t.Setenv("ANDROID_HOME", sdk)
t.Setenv("ANDROID_SDK_ROOT", "")
if err := Preflight(context.Background(), "android"); err != nil {
t.Fatalf("Preflight with adb under $ANDROID_HOME: %v", err)
}
}
func writeExecutable(t *testing.T, path string) {
t.Helper()
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatalf("mkdir %s: %v", filepath.Dir(path), err)
}
if err := os.WriteFile(path, nil, 0o755); err != nil {
t.Fatalf("write %s: %v", path, err)
}
}
func TestPreflight_iOSNeedsXcrun(t *testing.T) {
check := func(name string) error {
if name == "xcrun" {
+24
View File
@@ -242,7 +242,15 @@ func Execute(ctx context.Context, options Options, stdout io.Writer) error {
// --exit-on-violation a run that found violations is still a successful run
// (the summary reports them), which is the behaviour every existing caller
// depends on.
//
// A run none of whose steps reached the verifier fails whatever the flags say,
// because it holds no verdict to report. The threshold is every step and not a
// fraction of them: a screen that composes now and then costs a healthy android
// run a step or two, and a check that fired on those would be red on every run.
func runOutcome(options Options, summary runner.Summary) error {
if summary.Steps > 0 && summary.SkippedVerification == summary.Steps {
return VacuousRunError{Steps: summary.Steps}
}
if options.ExitOnViolation && len(summary.Violations) > 0 {
return ViolationsError{Count: len(summary.Violations)}
}
@@ -261,6 +269,22 @@ func (e ViolationsError) Error() string {
return fmt.Sprintf("%d violation record(s)", e.Count)
}
// VacuousRunError reports a run in which no step reached the verifier, so no
// property ever judged anything. It is not a clean run and it is not a found
// bug: it is a run that produced no evidence either way, and the absence of
// violations in it says nothing about the app. It stays untyped to the CLI's
// violation path on purpose, so it exits 1 as a broken run rather than 2.
type VacuousRunError struct {
Steps int
}
func (e VacuousRunError) Error() string {
return fmt.Sprintf(
"%d step(s) ran and none of them reached the verifier: the screen was "+
"still moving every time it was read, so no property judged this run",
e.Steps)
}
// bundleInputs holds the pre-driver assembly: alias map, seed, esbuild defines,
// and the resolved spec-API/goja-runtime paths the bundler consumes.
type bundleInputs struct {
+25
View File
@@ -266,6 +266,31 @@ func TestRunOutcome_ReportsViolationsOnlyUnderTheFlag(t *testing.T) {
}
}
// A step the verifier skipped was judged by nothing, so a run whose every step
// was skipped holds no verdict at all: "no violations" there is the absence of
// an answer rather than a clean one. Reporting it as a successful run is the
// green and vacuous outcome structuralShape's own design notes call worse than
// the composition it catches, and the runner's hold is what makes a fully
// skipped run reachable.
func TestRunOutcome_ARunThatJudgedNothingIsNotASuccess(t *testing.T) {
nothingJudged := runner.Summary{Steps: 6, SkippedVerification: 6}
err := runOutcome(Options{}, nothingJudged)
var vacuous VacuousRunError
if !errors.As(err, &vacuous) {
t.Fatalf("a run that judged none of its 6 steps came back %v, want a VacuousRunError", err)
}
if vacuous.Steps != 6 {
t.Errorf("steps: got %d, want 6", vacuous.Steps)
}
// A screen that composes now and then costs a run steps, not its verdict. A
// check that fired here would turn every healthy android run red.
mostlyJudged := runner.Summary{Steps: 6, SkippedVerification: 5}
if err := runOutcome(Options{}, mostlyJudged); err != nil {
t.Errorf("a run that judged one of its 6 steps must succeed, got %v", err)
}
}
// wedgedLaunchDriver never returns from Launch, standing in for a driver whose
// device-side session is stuck.
type wedgedLaunchDriver struct {
+75
View File
@@ -2,6 +2,7 @@ package verifier
import (
"os"
"strconv"
"testing"
"github.com/priyanshujain/sanderling/internal/hierarchy"
@@ -99,3 +100,77 @@ func TestStateAxFindWorks(t *testing.T) {
t.Fatalf("findAll count = %d, want 1", count)
}
}
// axSelectorFormsTree carries one node per id shape a dump produces: the bare
// tag Compose and the web driver emit, the package-qualified resource id
// Android emits, and the iOS accessibility identifier.
const axSelectorFormsTree = `{
"attributes": {"resource-id": "root", "bounds": "[0,0,400,800]"},
"children": [
{"attributes": {"resource-id": "BareThing", "text": "bare", "bounds": "[0,0,100,50]"},
"children": []},
{"attributes": {"resource-id": "com.example.app:id/AndroidThing", "text": "android",
"bounds": "[0,50,100,100]"}, "children": []},
{"attributes": {"accessibilityIdentifier": "IosThing", "text": "ios",
"bounds": "[0,100,100,150]"}, "children": []}
]
}`
// TestStateAxSelectorFormsAgree drives both selector forms a spec can write
// through state.ax.find and holds them to the same element. The two forms
// dispatch to different lookups (findNodeFromJS sends a string to FindNode and
// an object to FindBySelector), and the object one used to skip the id rule
// that knows an Android resource id is package-qualified, so a spec that wrote
// ax.find({id: "AddAccountSubmit"}) got undefined on Android and every property
// reading it passed while checking nothing.
func TestStateAxSelectorFormsAgree(t *testing.T) {
tree, err := hierarchy.Parse(axSelectorFormsTree)
if err != nil {
t.Fatal(err)
}
for _, test := range []struct {
value string
want string
}{
{"BareThing", "bare"},
{"AndroidThing", "android"},
{"com.example.app:id/AndroidThing", "android"},
{"IosThing", "ios"},
} {
t.Run(test.value, func(t *testing.T) {
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.fromObject = __sanderling__.extract(
state => state.ax.find({ id: `+strconv.Quote(test.value)+` })?.text, "fromObject");
globalThis.fromString = __sanderling__.extract(
state => state.ax.find("id:" + `+strconv.Quote(test.value)+`)?.text, "fromString");
globalThis.properties = {};
`)
if err := verifier.PushSnapshot(SnapshotInput{Tree: tree}); err != nil {
t.Fatal(err)
}
fromObject := readCurrent(t, verifier, "fromObject")
fromString := readCurrent(t, verifier, "fromString")
if fromString != test.want {
t.Fatalf(`ax.find("id:%s") read %v, want %q`, test.value, fromString, test.want)
}
if fromObject != fromString {
t.Errorf(
`one selector, two answers: ax.find({id: %q}) read %v and ax.find("id:%s") read %v`,
test.value, fromObject, test.value, fromString,
)
}
})
}
}
// readCurrent returns a named extractor's current value, or nil when the getter
// returned undefined, which is what an unresolved selector produces.
func readCurrent(t *testing.T, verifier *Verifier, name string) any {
t.Helper()
handle := verifier.runtime.GlobalObject().Get(name)
if handle == nil {
t.Fatalf("%s is not defined", name)
}
return handle.ToObject(verifier.runtime).Get("current").Export()
}
@@ -177,3 +177,79 @@ func compactJSON(t *testing.T, source string) string {
}
return compact.String()
}
// TestExtractorEncoding_NestedUndefinedIsNotOnTheWire pins the one reading
// shape the two hosts do NOT encode alike, rather than hiding it.
//
// JSON has no undefined, so the page loses the whole key (asserted in
// pkg/spec/test/web-runtime.test.ts) while goja writes null. goja cannot mirror
// the drop: Export reports an undefined member and a null member identically as
// nil, so dropping those keys here would drop the genuine nulls the page keeps.
// Mirroring the other way, by writing null on the page, would break the one
// thing that does agree. Carrying the member across takes a wire format that
// can express undefined, which is a change to every layer that parses a reading
// and to the replay UI that renders one.
//
// So the guarantee is narrower than "the same object": both hosts answer
// undefined when a property READS the member. Key presence (`in`, Object.keys)
// is not part of it, and this test says so out loud, so closing the gap has to
// be a deliberate change to both hosts at once.
func TestExtractorEncoding_NestedUndefinedIsNotOnTheWire(t *testing.T) {
const reading = `({ absent: undefined, empty: null, present: 1 })`
const fromGoja = `{"absent":null,"empty":null,"present":1}`
// What the page sends for the same getter, with the key gone.
const fromWeb = `{"empty":null,"present":1}`
if got := encodeSpecValue(t, reading); got != fromGoja {
t.Errorf("goja encoded the reading as %s, want %s", got, fromGoja)
}
native := newVerifier(t)
mustLoad(t, native, "__sanderling__.extract(state => "+reading+", \"value\");\nglobalThis.properties = {};")
if err := native.PushSnapshot(SnapshotInput{}); err != nil {
t.Fatal(err)
}
web := newVerifier(t)
mustLoad(t, web, "__sanderling__.extract(state => null, \"value\");\nglobalThis.properties = {};")
if err := web.PushSnapshot(SnapshotInput{}); err != nil {
t.Fatal(err)
}
if _, err := web.OverrideExtractorValues(map[int]json.RawMessage{0: json.RawMessage(fromWeb)}); err != nil {
t.Fatal(err)
}
for _, probe := range []struct {
expression string
native bool
web bool
}{
{"reading.absent === undefined", true, true},
{"reading.empty === null", true, true},
{"reading.present === 1", true, true},
// The half that does not survive the wire.
{`"absent" in reading`, true, false},
} {
if got := evaluateAgainstReading(t, native, probe.expression); got != probe.native {
t.Errorf("goja host: %s is %v, want %v", probe.expression, got, probe.native)
}
if got := evaluateAgainstReading(t, web, probe.expression); got != probe.web {
t.Errorf("web host: %s is %v, want %v", probe.expression, got, probe.web)
}
}
}
// evaluateAgainstReading answers a boolean expression over the value a property
// would read out of the first extractor, which is where the two hosts have to
// agree.
func evaluateAgainstReading(t *testing.T, verifier *Verifier, expression string) bool {
t.Helper()
if err := verifier.runtime.GlobalObject().Set("reading", verifier.extractors[0].currentValue); err != nil {
t.Fatal(err)
}
value, err := verifier.runtime.RunString(expression)
if err != nil {
t.Fatalf("evaluate %s: %v", expression, err)
}
return value.ToBoolean()
}
+42 -6
View File
@@ -353,9 +353,20 @@ func lastActionFields(action *Action) []actionField {
if action.Applied {
applied = true
}
// A relaunch between two readings is not "no action happened", which is
// what dropping the action reported instead: the app restarted after an
// action that did run. Only the positive report is a fact the runner can
// vouch for, so the other side is null rather than false: a target whose
// foreground the runner cannot read (web, iOS) never relaunches the app and
// still cannot promise it never restarted.
var relaunched any
if action.Relaunched {
relaunched = true
}
fields := []actionField{
{key: "kind", value: string(action.Kind)},
{key: "applied", value: applied},
{key: "relaunched", value: relaunched},
}
if action.On != "" {
fields = append(fields, actionField{key: "on", value: action.On})
@@ -457,19 +468,44 @@ func runtimeMillis(stepTime, runStart time.Time) int64 {
return stepTime.Sub(runStart).Milliseconds()
}
// logFields is the ONE description of a state.logs entry, for the same reason
// lastActionFields is: the goja host turns it into a JS object (logsArray) and
// the web host receives the same fields as JSON (EncodeLogs), so a property
// counting error-level lines cannot read one shape on native and another on web.
func logFields(entry LogEntry) []actionField {
return []actionField{
{key: "unixMillis", value: entry.UnixMillis},
{key: "level", value: entry.Level},
{key: "tag", value: entry.Tag},
{key: "message", value: entry.Message},
}
}
func logsArray(runtime *goja.Runtime, logs []LogEntry) *goja.Object {
array := runtime.NewArray()
for index, entry := range logs {
item := runtime.NewObject()
_ = item.Set("unixMillis", entry.UnixMillis)
_ = item.Set("level", entry.Level)
_ = item.Set("tag", entry.Tag)
_ = item.Set("message", entry.Message)
_ = array.Set(fmt.Sprintf("%d", index), item)
_ = array.Set(fmt.Sprintf("%d", index), objectFromFields(runtime, logFields(entry)))
}
return array
}
// EncodeLogs renders this step's log entries for the web host, which has no
// Go-side state object to read: the runner pushes this JSON into the page
// before each extractor evaluation. No entries encodes as an empty array, the
// same value the goja host reports for a step whose log fetch found nothing.
func EncodeLogs(logs []LogEntry) json.RawMessage {
var buffer bytes.Buffer
buffer.WriteByte('[')
for index, entry := range logs {
if index > 0 {
buffer.WriteByte(',')
}
buffer.Write(encodeFields(logFields(entry)))
}
buffer.WriteByte(']')
return buffer.Bytes()
}
func exceptionsArray(runtime *goja.Runtime, exceptions []Exception) *goja.Object {
array := runtime.NewArray()
for index, exception := range exceptions {
+96
View File
@@ -169,6 +169,7 @@ func TestLastAction_WebJSONMatchesTheGojaObject(t *testing.T) {
{"nil", nil},
{"Tap", &Action{Kind: ActionKindTap, On: "id:TxnSubmit", X: 12, Y: 34}},
{"TapApplied", &Action{Kind: ActionKindTap, On: "id:TxnSubmit", Applied: true}},
{"TapRelaunched", &Action{Kind: ActionKindTap, On: "id:TxnSubmit", Applied: true, Relaunched: true}},
{"TapWithoutSelector", &Action{Kind: ActionKindTap, X: 12, Y: 34}},
{"DoubleTap", &Action{Kind: ActionKindDoubleTap, On: `desc:say "hi" <b>`}},
{"InputText", &Action{Kind: ActionKindInputText, On: "id:field", Text: "50"}},
@@ -233,3 +234,98 @@ func TestLastAction_SeparatesNoActionFromAnActionOfUnknownFate(t *testing.T) {
})
}
}
// The runner relaunches the app when it leaves the foreground, which used to be
// reported to the spec as "no action ran between these two readings". The
// action did run; what a property cannot assume across it is that app state was
// continuous, so the relaunch is its own fact on an action that keeps its
// confirmed dispatch.
func TestLastAction_ReportsARelaunchSeparatelyFromTheDispatch(t *testing.T) {
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.continuity = __sanderling__.extract(state =>
state.lastAction === null ? "no action"
: state.lastAction.applied !== true ? "unconfirmed"
: state.lastAction.relaunched === true ? "applied across a relaunch"
: state.lastAction.relaunched === null ? "applied, no relaunch reported"
: "unreadable");
`)
for _, testCase := range []struct {
name string
action *Action
want string
}{
{"nothing ran", nil, "no action"},
{
"confirmed, app stayed",
&Action{Kind: ActionKindTap, On: "id:TxnSubmit", Applied: true},
"applied, no relaunch reported",
},
{
"confirmed, app relaunched after it",
&Action{Kind: ActionKindTap, On: "id:TxnSubmit", Applied: true, Relaunched: true},
"applied across a relaunch",
},
} {
t.Run(testCase.name, func(t *testing.T) {
if err := verifier.PushSnapshot(SnapshotInput{
Snapshots: Snapshots{},
LastAction: testCase.action,
}); err != nil {
t.Fatal(err)
}
handle := verifier.runtime.GlobalObject().Get("continuity").ToObject(verifier.runtime)
if got := handle.Get("current").String(); got != testCase.want {
t.Errorf("the spec read %q, want %q", got, testCase.want)
}
})
}
}
// TestLogs_WebJSONMatchesTheGojaObject pins state.logs to ONE shape across the
// two hosts, for the same reason lastAction is pinned. On web the page's
// reading of every extractor replaces the host's, so state.logs is whatever
// EncodeLogs put in the page: a field this side renames or cases differently
// leaves the default noLogcatErrors counting nothing on web while it counts on
// native, with nothing reporting that it never saw an entry.
func TestLogs_WebJSONMatchesTheGojaObject(t *testing.T) {
verifier := newVerifier(t)
mustLoad(t, verifier, `
globalThis.lines = __sanderling__.extract(state => JSON.stringify(state.logs));
`)
for _, testCase := range []struct {
name string
logs []LogEntry
}{
{"none", nil},
{"empty", []LogEntry{}},
{
"one error",
[]LogEntry{{UnixMillis: 1700000000123, Level: "E", Tag: "console", Message: "boom from the page"}},
},
{
"mixed levels",
[]LogEntry{
{UnixMillis: 1, Level: "E", Tag: "console", Message: `say "hi" <b> & co`},
{UnixMillis: 2, Level: "W", Tag: "AndroidRuntime", Message: "a warning"},
},
},
} {
t.Run(testCase.name, func(t *testing.T) {
if err := verifier.PushSnapshot(SnapshotInput{
Snapshots: Snapshots{},
Logs: testCase.logs,
}); err != nil {
t.Fatal(err)
}
handle := verifier.runtime.GlobalObject().Get("lines").ToObject(verifier.runtime)
goja := handle.Get("current").String()
web := string(EncodeLogs(testCase.logs))
if goja != web {
t.Errorf("the two hosts disagree on state.logs\n goja: %s\n web: %s", goja, web)
}
})
}
}
+6
View File
@@ -38,6 +38,12 @@ type Action struct {
// when the apply call failed and nothing can say whether the action
// reached the app. The spec is told which of the two it is.
Applied bool
// Relaunched, like Applied, is meaningful only on state.lastAction: the
// runner brought the app back to the foreground after this action, so the
// two readings the spec compares straddle a restart. The action still
// happened; what a property cannot assume across it is that app state ran
// continuously from one reading to the next.
Relaunched bool
}
// LogEntry mirrors a logcat line captured between steps.