refactoring default action layer (#51)

* feat(hierarchy): add editable signal with native derivation

* feat(chrome): emit editable flag in hierarchy dump

* feat(verifier): expose editable on ax element objects

* feat(spec): add editable to selector and element types

* feat(verifier): register typing builtin generator

* feat(verifier): typing builtin types edge-case corpus into editable fields

* feat(spec): export typing builtin generator

* feat(spec): add defaultActions bundle

* feat(spec): export @sanderling/spec/defaults subpath

* feat(folio): layer defaultActions breadth over targeted flows

* test(verifier): typing builtin targets editable fields, declines otherwise

* test(hierarchy): editable derivation and selector matching

* test(spec): defaultActions, typing, and defaults barrel resolve

* fix(testrun): alias @sanderling/spec/defaults for the bundler

* test(chrome): editable flag for inputs, textarea, contenteditable

* feat(spec): typing builtin for the web (V8) action path

* chore(folio): auto-boot a bootable AVD in just test/install when none connected

* feat(driver): add ForegroundChecker optional capability

* feat(android): detect foreground package via adb dumpsys

* feat(sidecar): implement ForegroundApp via adb for android

* feat(runner): relaunch app when foreground escapes during exploration

* fix(spec): drop hardware back from defaultActions to stay in-app

* feat(spec): add DoubleTap action type and constructor

* feat(spec): wire DoubleTap through web-runtime serializer

* feat(verifier): bind doubleTap and decode DoubleTap actions

* feat(runner): dispatch DoubleTap as two taps inside one step

* test(doubleTap): cover constructor, verifier round-trip, and runner dispatch

* feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action

* fix(folio): track ledger row count across non-ledger steps; pin reproducer seed

* feat(spec): add doubleTaps random-target builtin to defaultActions

* feat(verifier): add doubleTaps random-target generator

* refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions

* fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives

* feat(verifier): track newly-violated property set per step

Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.

* refactor(runner): emit onset-only violations to trace and summary

Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.

* style(verifier): use maps.Copy for verdict snapshot

* fix(folio): make login spec content-driven (idempotent across re-entries)

* fix(verifier): canonicalize selector strings

Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.

* refactor(folio): replace txn invariants with balanceMatchesAddedTxn

Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.

* refactor(trace): drop WriteScreenshotAfter

Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.

* refactor(runner): one concurrent screenshot per step

Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.

* refactor(inspect-ui): use next step's screenshot for state after

Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.

* feat(sidecar): structural-hash settle poll

Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.

* test(sidecar): cover pollUntilStable and structuralHash

Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.

* feat(spec): accept optional name on extract()

Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.

* test(spec): cover extract name overload

Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.

* feat(verifier): name extractors for diff surfacing

bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.

* chore(folio): name every extract() call

Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.

* feat(verifier): track extractor value transitions

Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.

* test(verifier): cover ChangedExtractors diffs

Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.

* feat(trace): emit extractor_changes per step

Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.

* feat(inspect-ui): render extractor-change breadcrumbs at violations

Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.

* fix(sidecar): cap stability poll independently of settle budget

The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.

* feat(cli): default --clear-data on so runs start fresh

* feat(sidecar): streak-based settle with route-transition detection

Two changes layered into the stability poll:

1. stabilitySnapshot returns null while the tree carries more than one
   route-level Screen tag (resource-id / testTag / identifier ending
   in "Screen"), so the poll cannot declare a NavHost cross-fade
   stable. Apps following the Compose route convention get this
   detection for free; apps that don't fall through to the generic
   signal below.

2. pollUntilStable now requires an uninterrupted stable streak of at
   least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
   matches. A late transition that fires after a brief calm window
   breaks the streak instead of slipping past. Interval widened to
   250ms so UiAutomation isn't hammered under fuzz load.

* test(sidecar): cover streak reset and route-transition rejection

Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.

* feat(runner): re-fetch on transitional hierarchy capture

Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.

fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.

* feat(runner): gate first action on app reaching foreground

* test(runner): cover startup foreground gate and back-press

* feat(verifier): scope random-action targets to app package

Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.

* feat(testrun): pass app package into verifier scope filter

* test(verifier): cover package-scoped target selection

* feat(hierarchy): derive package from resource-id prefix

The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.

* test(hierarchy): cover package derivation from resource-id

* chore: stop tracking inspect-ui/dist build artifacts

* feat(android): detect focused-window package via dumpsys window

* feat(driver): add FocusedWindowChecker capability

* fix(runner): gate first observe on the app window being drawn, not just resumed

* test(mock): add FocusedWindowApp with foreground mirroring

* test(runner): cover startup gate waiting for app window to draw

* feat(proto): add Snapshot RPC for atomic hierarchy+screenshot

Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.

* feat(sidecar): add snapshot default on DriverBackend

Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.

* feat(sidecar): wire Snapshot handler with serialization lock

Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.

* test(sidecar): cover Snapshot wire path and serialization lock

SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.

* feat(driver): expose Snapshot on DeviceDriver and sidecar client

Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.

* feat(driver): add Snapshot to chrome and mock drivers

The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.

* refactor(runner): observe each step via the atomic Snapshot RPC

fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.

* test(runner): assert step uses Snapshot, not raw hierarchy/screenshot

TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.

* test(driver): cover Snapshot in proto descriptor and sidecar client

Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.

* feat(trace): add Transitional flag to Step

* fix(runner): skip verifier for transitional trees after retry budget

When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.

* test(runner): cover transitional step skips verifier and clean control

* refactor(trace): rename Step.Action to Step.NextAction

The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.

* refactor(runner): assign trace action to Step.NextAction field

Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.

* refactor(inspect): decode trace step's next_action JSON field

Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.

* test(inspect): update fixtures to use next_action trace field

Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.

* refactor(inspect-ui): rename Step.action to Step.next_action

Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.

* fix(folio): extract balanceMatchesAddedSum predicate as testable helper

Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.

* fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn

The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.

* test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases

Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.

* fix(build): rebuild sidecar JAR when Kotlin sources change

Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.

* fix(chrome): launch with no-sandbox so headless Chrome starts in CI

* fix(sidecar): type text at cursor instead of clearing the field

InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.

* test(sidecar): assert InputText types at cursor without clearing

Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.

* feat(proto): add LongPress RPC

* chore(proto): regenerate Go stubs for LongPress

* feat(driver): add LongPress to DeviceDriver interface

* feat(sidecar): add LongPress client method

* feat(mock): record LongPress action

* feat(chrome): implement LongPress as press-and-hold

* feat(sidecar): implement longPress across backends

* feat(sidecar): dispatch LongPress RPC to backend

* test(sidecar): cover LongPress dispatch

* test(sidecar): implement longPress in snapshot test backend

* feat(verifier): add LongPress and Scroll action kinds

* feat(folio-spec): predicate that gates balance check on TxnSubmit tap

Replaces the row-sum predicate (which always held by construction since
balance is derived from rows in Folio) with one that compares the typed
amount to the actual balance delta after a tap on TxnSubmit. Catches the
planted double-submit bug.

* feat(folio-spec): wire submitMovesBalanceByTypedAmount property

Adds lastAction and totalBalance extractors and uses them in the new
property. Drops ledgerRows/ledgerBalance extractors since nothing else
referenced them.

* feat(verifier): wire longPresses and scrolls generators

* test(verifier): cover longPresses and scrolls generators

* test(folio-spec): unit tests for submitChangesBalanceByTypedAmount

Covers single vs double submit, the DoubleTap variant, vacuous cases
(null action, wrong kind, wrong target, zero typed), and selector-as-
object coercion.

* feat(spec): add LongPress and Scroll authoring surface

* feat(spec): no-op LongPress and Scroll in web runtime

* feat(spec): re-export longPresses and scrolls as opt-in generators

* test(spec): cover LongPress and Scroll runtime members

* test(proto): expect LongPress in service descriptor

* feat(runner): dispatch LongPress and Scroll actions

* test(runner): cover LongPress and Scroll dispatch

* docs(action-space): move LongPress, Scroll, DoubleTap to current actions

* fix(runner): mark nil/empty hierarchy as transitional

A failed or empty sidecar hierarchy fetch was pushed straight to the
verifier, letting spec extractors crash with "Cannot read property 'map'
of undefined" when findAll returned null. Treat that case like a
transitional capture: skip the verifier push, still record the step, and
keep the loop progressing.

* fix(verifier): populate Action.On when tap chooser picks an element

Coordinate-targeted Taps/DoubleTaps left On empty, so action-gated
properties reading lastAction.on couldn't tell which target was hit and
were vacuously skipped. Resolve the picked element to a stable
key:value selector (resource-id, testTag, text, desc) and validate it
resolves back to the same element so we don't accidentally redirect the
tap to a sibling that shares the identifier.

* fix(folio): add parseTypedAmount helper matching app's parseCents

Raw user input like "50" must become 5000 cents, not 50. The existing
parseDollarCents helper strips non-digits and so reads "50" as 50 cents,
which is correct for formatted balance text but off by 100x for raw
input from the amount field.

* fix(folio): parse raw amount input as cents in submit predicate

txnAmountField holds raw user keystrokes, not formatted balance text.
Route it through parseTypedAmount so "50" reads as $50, matching how
the app commits the transaction.

* fix(folio): carry forward total balance across off-screen transitions

AddTransactionScreen shows neither AccountCard nor LedgerBalance, so the
extractor used to report 0 at the step before submit. That made every
non-zero current balance look like the full delta and tripped the typed
amount property on every honest submit. Remember the last-seen sum and
return it whenever the current snapshot has no balance signal.

* test(folio): cover submit predicate with raw typed-amount inputs

Pipes realistic raw keystrokes through parseTypedAmount + the predicate
so single submits clear and double submits fire as expected.

* feat(folio): add computeHomeTotalBalance helper

Pure helper that tracks Home multi-account total only and carries the last
Home sum across off-Home steps. Ledger's single-account balance is excluded
because mixing it would corrupt cross-screen scale comparisons.

* fix(folio): totalBalance carrier tracks only Home, not Ledger

Home cardSum is a multi-account total; Ledger's LedgerBalance is a single
account on a different scale. Blending them in the carrier produced bogus
cross-screen deltas (prev from Ledger, curr from Home), triggering false
positives in submitMovesBalanceByTypedAmount. Restrict the carrier to
Home AccountCard totals via the computeHomeTotalBalance helper.

* test(spec): cover computeHomeTotalBalance carrier behaviour

Tests Home sums, carrier passthrough on off-Home steps, the Ledger
scale-mismatch case, and a Home > off-Home > Home sequence.

* feat(runner): treat transient apply errors as transitional steps

Sidecar input RPCs occasionally hang with DEADLINE_EXCEEDED or
UNAVAILABLE on long fuzzing runs. The per-step loop previously
propagated any applyAction error and killed the run after a single
flake. Detect transient gRPC failures via status.FromError, mark the
step transitional, skip the post-action idle poll, and continue to the
next step. Fatal errors (outer ctx cancellation, non-transient codes,
verifier crashes) still propagate.

* test(runner): cover transient apply error resilience

TestRunner_TransientApplyErrorMarksTransitional drives the runner
through a wrapper that fails the first TapSelector with a gRPC
DeadlineExceeded then succeeds. Asserts the run does not exit, the
failed step is marked transitional with no violations, and the next
step runs cleanly. TestIsTransientApplyError_Classification covers the
helper's matching rules directly so future code changes don't quietly
drop a transient case.

* fix(folio): gate submit-balance property on Home route landing

totalBalance is only freshly computed when AccountCards are visible on
Home; off-Home landings return the carrier and would false-fire the
property, latching always(next(F)) to false and masking the real
double-submit bug. Skip vacuously when route is not "home".

* test(spec): cover route gate in submit-balance predicate

Adds route arg to existing cases (all use "home") and adds five new
cases: ledger landing with stale carrier, add-transaction with
double-insert delta, null route, plus home-landing positive and
double-insert negative cases anchoring the gate's allow path.
This commit is contained in:
pj authored and GitHub committed 2026-06-01 12:48:51 +05:30
1 parent f572c8ba66
commit 88db9653e5
65 files changed
+4748 -596

No files matched your search

+431 -101
View File
@@ -10,6 +10,8 @@ import (
"time"
"golang.org/x/sync/errgroup"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/driver"
"github.com/priyanshujain/sanderling/internal/hierarchy"
@@ -18,6 +20,11 @@ import (
"github.com/priyanshujain/sanderling/internal/verifier"
)
// doubleTapGap is the inter-tap delay for ActionKindDoubleTap: short enough to
// land both events inside a sub-100 ms race window, long enough for adb
// `input tap` to serialize two MotionEvent streams.
const doubleTapGap = 50 * time.Millisecond
type Options struct {
Duration time.Duration
IdleTimeout time.Duration
@@ -53,13 +60,16 @@ func Run(ctx context.Context, options Options) (Summary, error) {
logger = slog.Default()
}
// Gate on the app actually being on top before acting, so the first
// action never fires against a leftover screen or a system dialog. Done
// before the deadline is set so the settle time does not eat the run.
waitForForeground(ctx, options, logger)
summary := Summary{StartTime: time.Now()}
deadline := summary.StartTime.Add(options.Duration)
stepIndex := 0
var lastAction *verifier.Action
var lastLogTime time.Time
var pendingPostScreenshotStep int
pendingPostScreenshot := false
for time.Now().Before(deadline) {
if err := ctx.Err(); err != nil {
break
@@ -67,10 +77,19 @@ func Run(ctx context.Context, options Options) (Summary, error) {
stepIndex++
stepStart := time.Now()
// Hierarchy, metrics, and logs are independent device reads — run
// Keep exploration scoped to the app under test. If a prior action
// backed out of (or otherwise left) the app, relaunch it before we
// observe or act, so properties never evaluate against a foreign app
// and actions never land outside the app.
if ensureForeground(ctx, options, logger, stepIndex) {
lastAction = nil
}
// Hierarchy, metrics, and logs are independent device reads. Run
// them concurrently so metrics+logs hide behind the hierarchy fetch.
var tree *hierarchy.Tree
var hierarchyErr error
var transitional bool
var metrics *trace.Metrics
var logs []verifier.LogEntry
@@ -79,11 +98,14 @@ func Run(ctx context.Context, options Options) (Summary, error) {
// goroutine, whose CDP round-trip can otherwise outrun the step
// budget on a hung tab.
g, gctx := errgroup.WithContext(ctx)
si := stepIndex
// fetchSyncedState issues a single Snapshot RPC so hierarchy and
// screenshot describe the same frame, then re-fetches the pair
// while the tree still looks transitional.
g.Go(func() error {
tree, hierarchyErr = fetchHierarchy(gctx, options.Driver)
tree, transitional, hierarchyErr = fetchSyncedState(gctx, options, logger, si)
return nil
})
si := stepIndex
g.Go(func() error {
metrics = captureMetrics(gctx, options, logger, si)
return nil
@@ -105,14 +127,6 @@ func Run(ctx context.Context, options Options) (Summary, error) {
return nil
})
}
if pendingPostScreenshot {
postStep := pendingPostScreenshotStep
g.Go(func() error {
captureScreenshot(gctx, options, logger, postStep, true)
return nil
})
pendingPostScreenshot = false
}
// All goroutines write to local variables and return nil, so the Wait
// error is always nil; ignored intentionally.
_ = g.Wait()
@@ -127,38 +141,62 @@ func Run(ctx context.Context, options Options) (Summary, error) {
if tree != nil {
treeSize = len(tree.Elements)
}
// A nil or empty tree means the sidecar's hierarchy fetch failed or
// returned nothing (e.g. transient device-side timeout). Pushing it
// would let spec extractors call findAll() and chain .map() on a null
// result; treat it like a transitional capture so the verifier is
// skipped, the step is still recorded, and the loop progresses.
if treeSize == 0 {
transitional = true
}
lastLogTime = stepStart
if err := options.Verifier.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
LastAction: lastAction,
StepTime: stepStart,
RunStart: summary.StartTime,
Logs: logs,
}); err != nil {
return summary, fmt.Errorf("step %d push: %w", stepIndex, err)
}
skipped, overrideErr := options.Verifier.OverrideExtractorValues(v8Overrides)
if overrideErr != nil {
logger.Warn("v8 override apply failed", "step", stepIndex, "err", overrideErr)
}
if skipped > 0 {
logger.Warn("v8 override skipped out-of-range entries",
"step", stepIndex, "skipped", skipped, "have", len(v8Overrides))
}
screen := ""
if tree != nil && len(tree.Elements) > 0 {
screen = tree.Elements[0].Screen
}
logger.Info("step", "index", stepIndex, "screen", screen, "nodes", treeSize)
verdicts := options.Verifier.EvaluateProperties()
violations := violationNames(verdicts)
for _, name := range violations {
if predicateErr := options.Verifier.PredicateError(name); predicateErr != nil {
logger.Warn("predicate error", "step", stepIndex, "property", name, "err", predicateErr)
// Transitional trees describe a NavHost mid cross-fade. Pushing
// one would poison the verifier's previous/current extractor
// advance, so the next clean step would compare against this
// transient state and emit false-positive violations. We still
// record the step (hierarchy + screenshot) for inspect-side
// debugging, but skip the verifier entirely and pick the next
// action against the unchanged prior state to keep the loop
// progressing.
var violations []string
var extractorChanges map[string]trace.ExtractorChange
if !transitional {
if err := options.Verifier.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
LastAction: lastAction,
StepTime: stepStart,
RunStart: summary.StartTime,
Logs: logs,
}); err != nil {
return summary, fmt.Errorf("step %d push: %w", stepIndex, err)
}
skipped, overrideErr := options.Verifier.OverrideExtractorValues(v8Overrides)
if overrideErr != nil {
logger.Warn("v8 override apply failed", "step", stepIndex, "err", overrideErr)
}
if skipped > 0 {
logger.Warn("v8 override skipped out-of-range entries",
"step", stepIndex, "skipped", skipped, "have", len(v8Overrides))
}
options.Verifier.EvaluateProperties()
violations = options.Verifier.NewlyViolatedProperties()
for _, name := range violations {
if predicateErr := options.Verifier.PredicateError(name); predicateErr != nil {
logger.Warn("predicate error", "step", stepIndex, "property", name, "err", predicateErr)
}
}
extractorChanges = encodeExtractorChanges(options.Verifier.ChangedExtractors())
} else {
logger.Warn("transitional tree after retry budget; skipping verifier",
"step", stepIndex, "screen", screen, "nodes", treeSize)
}
logger.Info("step", "index", stepIndex, "screen", screen, "nodes", treeSize)
var nextAction verifier.Action
var nextErr error
@@ -179,20 +217,43 @@ func Run(ctx context.Context, options Options) (Summary, error) {
logger.Warn("residual encode failed", "step", stepIndex, "err", residualErr)
}
applySkipped := false
if nextErr == nil {
if err := applyAction(ctx, options.Driver, nextAction, tree); err != nil {
if isWDADrop(err) {
return summary, fmt.Errorf("step %d: iOS XCTest runner lost connection - known WDA startup flake, re-run the test: %w", stepIndex, err)
}
if isTransientApplyError(ctx, err) {
logger.Warn("transient apply error; marking step transitional", "step", stepIndex, "err", err)
transitional = true
applySkipped = true
lastAction = nil
} else {
return summary, fmt.Errorf("step %d apply: %w", stepIndex, err)
}
} else {
actionCopy := nextAction
lastAction = &actionCopy
}
} else {
lastAction = nil
}
step := trace.Step{
Index: stepIndex,
Timestamp: stepStart,
Screen: screen,
Action: traceAction,
Violations: violations,
Hierarchy: tree,
Residuals: residuals,
Metrics: metrics,
Index: stepIndex,
Timestamp: stepStart,
Screen: screen,
NextAction: traceAction,
Violations: violations,
Hierarchy: tree,
Residuals: residuals,
Metrics: metrics,
ExtractorChanges: extractorChanges,
Transitional: transitional,
}
if err := options.TraceWriter.WriteStep(step); err != nil {
return summary, fmt.Errorf("step %d trace: %w", stepIndex, err)
}
captureScreenshot(ctx, options, logger, stepIndex, false)
summary.Steps = stepIndex
if len(violations) > 0 {
summary.Violations = append(summary.Violations, ViolationRecord{
@@ -200,34 +261,19 @@ func Run(ctx context.Context, options Options) (Summary, error) {
Properties: violations,
})
}
if nextErr == nil {
if err := applyAction(ctx, options.Driver, nextAction, tree); err != nil {
if isWDADrop(err) {
return summary, fmt.Errorf("step %d: iOS XCTest runner lost connection - known WDA startup flake, re-run the test: %w", stepIndex, err)
}
return summary, fmt.Errorf("step %d apply: %w", stepIndex, err)
// Wait actions are themselves a settling: skip the idle poll. Actions
// that mutate the UI fall through to WaitForIdle so the next step's
// concurrent fetches observe a stable post-action state. A transient
// apply error means nothing landed, so the idle poll has nothing to
// settle and may itself hang on the same device condition.
if nextErr == nil && !applySkipped && nextAction.Kind != verifier.ActionKindWait {
idleCtx, idleCancel := context.WithTimeout(ctx, options.IdleTimeout)
idleErr := options.Driver.WaitForIdle(idleCtx, options.IdleTimeout)
if idleErr != nil && idleCtx.Err() == nil {
logger.Warn("wait_for_idle failed", "step", stepIndex, "err", idleErr)
}
actionCopy := nextAction
lastAction = &actionCopy
} else {
lastAction = nil
idleCancel()
}
idleCtx, idleCancel := context.WithTimeout(ctx, options.IdleTimeout)
idleErr := options.Driver.WaitForIdle(idleCtx, options.IdleTimeout)
if nextErr == nil {
pendingPostScreenshot = true
pendingPostScreenshotStep = stepIndex
}
if idleErr != nil && idleCtx.Err() == nil {
logger.Warn("wait_for_idle failed", "step", stepIndex, "err", idleErr)
}
idleCancel()
}
if pendingPostScreenshot {
captureScreenshot(ctx, options, logger, pendingPostScreenshotStep, true)
}
summary.EndTime = time.Now()
@@ -253,14 +299,109 @@ func validate(options Options) error {
return nil
}
func violationNames(verdicts map[string]ltl.Verdict) []string {
var names []string
for name, verdict := range verdicts {
if verdict == ltl.VerdictViolated {
names = append(names, name)
}
// ensureForeground keeps the app under test in the foreground. When the driver
// can report the foreground app and it no longer matches the bundle under test,
// the app is relaunched. Returns true when a relaunch happened so the caller
// can drop the now-stale lastAction. Drivers without ForegroundChecker (web,
// iOS) are a no-op.
func ensureForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) bool {
checker, ok := options.Driver.(driver.ForegroundChecker)
if !ok || options.BundleID == "" {
return false
}
return names
foreground, err := checker.ForegroundApp(ctx)
if err != nil {
logger.Warn("foreground check failed", "step", stepIndex, "err", err)
return false
}
if foreground == "" || foreground == options.BundleID {
return false
}
logger.Warn("app left foreground; relaunching",
"step", stepIndex, "foreground", foreground, "want", options.BundleID)
return bringToForeground(ctx, options, logger, stepIndex)
}
// foregroundReadyAttempts bounds how many times waitForForeground tries to
// bring the app forward before the first step, so a stuck system dialog can
// never hang the run.
const foregroundReadyAttempts = 8
// waitForForeground blocks until the app under test is actually on screen, so
// the first observe never captures a leftover screen or a freshly-booted
// device's system dialog (e.g. Android's "set a screen lock" prompt). Drivers
// without ForegroundChecker (web) and an unknown foreground both skip the gate.
//
// It is not enough that the app is the resumed activity: ResumedActivity flips
// to a freshly launched app ~before its first frame draws, so gating on it
// alone lets the first observe read the outgoing app. When the driver can also
// report the focused window, the gate additionally waits for that window to
// name the app, which only happens once it is genuinely drawn.
func waitForForeground(ctx context.Context, options Options, logger *slog.Logger) {
checker, ok := options.Driver.(driver.ForegroundChecker)
if !ok || options.BundleID == "" {
return
}
focusChecker, hasFocus := options.Driver.(driver.FocusedWindowChecker)
for attempt := range foregroundReadyAttempts {
if err := ctx.Err(); err != nil {
return
}
foreground, err := checker.ForegroundApp(ctx)
if err != nil {
logger.Warn("foreground check failed before first step", "err", err)
return
}
if foreground == "" {
return // foreground unknowable (e.g. iOS); don't block the run
}
if foreground != options.BundleID {
logger.Warn("app not in foreground at start; bringing it forward",
"foreground", foreground, "want", options.BundleID, "attempt", attempt)
bringToForeground(ctx, options, logger, 0)
continue
}
if !hasFocus {
return // resumed is the app and no finer signal exists
}
focused, err := focusChecker.FocusedWindowApp(ctx)
if err != nil {
logger.Warn("focus check failed before first step", "err", err)
return
}
if focused == options.BundleID {
return // window is drawn; safe to observe
}
logger.Warn("app resumed but window not yet drawn; waiting",
"focused", focused, "want", options.BundleID, "attempt", attempt)
settleForForeground(ctx, options)
}
logger.Warn("app never reached foreground before first step; proceeding anyway",
"want", options.BundleID)
}
// bringToForeground returns the app under test to the foreground. It first
// presses BACK to dismiss any modal system dialog (a relaunch alone does not
// close one), then relaunches and waits for the UI to settle. Returns true
// when the relaunch itself succeeded.
func bringToForeground(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) bool {
if err := options.Driver.PressKey(ctx, "back"); err != nil {
logger.Warn("dismiss key before relaunch failed", "step", stepIndex, "err", err)
}
if err := options.Driver.Launch(ctx, options.BundleID, false, nil); err != nil {
logger.Warn("relaunch failed", "step", stepIndex, "err", err)
return false
}
settleForForeground(ctx, options)
return true
}
// settleForForeground waits one idle window for the UI to settle, bounding the
// wait by the driver's idle timeout.
func settleForForeground(ctx context.Context, options Options) {
idleCtx, cancel := context.WithTimeout(ctx, options.IdleTimeout)
_ = options.Driver.WaitForIdle(idleCtx, options.IdleTimeout)
cancel()
}
func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.Action, tree *hierarchy.Tree) error {
@@ -274,6 +415,43 @@ func applyAction(ctx context.Context, drv driver.DeviceDriver, action verifier.A
return drv.TapSelector(ctx, action.On)
}
return drv.Tap(ctx, x, y)
case verifier.ActionKindDoubleTap:
x, y, ok := resolveCoordinates(action, tree)
tap := func() error {
if !ok {
if action.On == "" {
return nil
}
return drv.TapSelector(ctx, action.On)
}
return drv.Tap(ctx, x, y)
}
if err := tap(); err != nil {
return err
}
timer := time.NewTimer(doubleTapGap)
defer timer.Stop()
select {
case <-ctx.Done():
return ctx.Err()
case <-timer.C:
}
return tap()
case verifier.ActionKindLongPress:
x, y, ok := resolveCoordinates(action, tree)
if !ok {
// No long-press-by-selector RPC exists, so an unresolved target is
// nothing we can dispatch; skip rather than error.
return nil
}
return drv.LongPress(ctx, x, y)
case verifier.ActionKindScroll:
fromX, fromY, toX, toY := scrollEndpoints(action, tree)
duration := time.Duration(action.DurationMillis) * time.Millisecond
if duration <= 0 {
duration = 300 * time.Millisecond
}
return drv.Swipe(ctx, fromX, fromY, toX, toY, duration)
case verifier.ActionKindInputText:
if x, y, ok := resolveCoordinates(action, tree); ok {
if err := drv.Tap(ctx, x, y); err != nil {
@@ -360,18 +538,139 @@ func resolveCoordinates(action verifier.Action, tree *hierarchy.Tree) (int, int,
return 0, 0, false
}
func fetchHierarchy(ctx context.Context, drv driver.DeviceDriver) (*hierarchy.Tree, error) {
xmlText, err := drv.Hierarchy(ctx)
if err != nil {
return nil, err
// scrollEndpoints lowers a Scroll to a swipe's from/to points. Pre-computed
// endpoints (from the generator) win. Otherwise it derives them from the
// container bounds: the named node when On resolves, else the whole screen.
func scrollEndpoints(action verifier.Action, tree *hierarchy.Tree) (fromX, fromY, toX, toY int) {
if action.FromX != 0 || action.FromY != 0 || action.ToX != 0 || action.ToY != 0 {
return action.FromX, action.FromY, action.ToX, action.ToY
}
return hierarchy.Parse(xmlText)
bounds := scrollBounds(action, tree)
cx, cy := bounds.Center()
width := bounds.Width()
height := bounds.Height()
toX, toY = cx, cy
// Scroll direction names content motion; the gesture swipes the opposite
// way. Revealing lower content ("down") drags the finger up, so toY drops.
switch action.Direction {
case "down":
toY = cy - (4*height)/10
case "up":
toY = cy + (4*height)/10
case "left":
toX = cx + (4*width)/10
case "right":
toX = cx - (4*width)/10
}
if toX < 0 {
toX = 0
}
if toY < 0 {
toY = 0
}
return cx, cy, toX, toY
}
// scrollBounds returns the container bounds for an authored Scroll: the node
// named by On when it resolves, otherwise the root (whole-screen) bounds.
func scrollBounds(action verifier.Action, tree *hierarchy.Tree) hierarchy.Bounds {
if tree == nil {
return hierarchy.Bounds{}
}
if action.On != "" {
if element := tree.Find(action.On); element != nil {
return element.Bounds
}
}
if tree.Root != nil {
return tree.Root.Bounds
}
return hierarchy.Bounds{}
}
// transitionalRetryAttempts caps how many times we re-fetch hierarchy when a
// tree carries more than one route-level Screen tag (NavHost cross-fade in
// flight). Each retry pauses transitionalRetrySleep before the next fetch.
const (
transitionalRetryAttempts = 4
transitionalRetrySleep = 200 * time.Millisecond
)
// fetchSyncedState fetches hierarchy and screenshot together so the recorded
// pair shows the same UI moment. If the hierarchy looks like a NavHost
// cross-fade (multiple route-level *Screen tags), the function waits briefly
// and re-fetches the pair, up to transitionalRetryAttempts times. This
// handles transitions whose async work begins after the sidecar's settle
// poll has already exited.
//
// The driver's Snapshot RPC captures both reads under a backend-side mutex
// so they describe the same on-device frame; the retry exists for the
// orthogonal case where the frame itself is transitional.
//
// The transitional return reports whether the retry budget was exhausted
// on a still-transitional tree. Callers use it to skip the verifier for
// that step so the previous/current extractor advance does not absorb
// transient state.
func fetchSyncedState(ctx context.Context, options Options, logger *slog.Logger, stepIndex int) (tree *hierarchy.Tree, transitional bool, err error) {
var pngBytes []byte
retryLoop:
for attempt := range transitionalRetryAttempts {
hierarchyJSON, image, snapshotErr := options.Driver.Snapshot(ctx)
if snapshotErr != nil {
err = snapshotErr
tree = nil
} else {
tree, err = hierarchy.Parse(hierarchyJSON)
pngBytes = image.PNG
}
if err != nil || !isTransitionalHierarchy(tree) {
break
}
if attempt == transitionalRetryAttempts-1 {
transitional = true
break
}
timer := time.NewTimer(transitionalRetrySleep)
select {
case <-ctx.Done():
timer.Stop()
break retryLoop
case <-timer.C:
}
}
if len(pngBytes) > 0 {
if writeErr := options.TraceWriter.WriteScreenshot(stepIndex, pngBytes); writeErr != nil {
logger.Warn("screenshot write failed", "step", stepIndex, "err", writeErr)
}
}
return tree, transitional, err
}
// isTransitionalHierarchy returns true when the tree carries more than one
// resource-id ending in "Screen" - the marker of a Compose NavHost mid
// cross-fade where both source and destination route composables are alive.
// Mirrors the sidecar's stabilitySnapshot heuristic so runner-side rejection
// stays consistent with the settle poll.
func isTransitionalHierarchy(tree *hierarchy.Tree) bool {
if tree == nil {
return false
}
screens := 0
for _, element := range tree.Elements {
if strings.HasSuffix(element.ResourceID, "Screen") {
screens++
if screens > 1 {
return true
}
}
}
return false
}
func traceActionFor(action verifier.Action, tree *hierarchy.Tree) *trace.Action {
traceAction := &trace.Action{Kind: string(action.Kind), X: action.X, Y: action.Y}
switch action.Kind {
case verifier.ActionKindTap:
case verifier.ActionKindTap, verifier.ActionKindDoubleTap, verifier.ActionKindLongPress:
traceAction.Selector = action.On
stampSelectorTarget(traceAction, action, tree)
case verifier.ActionKindInputText:
@@ -386,6 +685,15 @@ func traceActionFor(action verifier.Action, tree *hierarchy.Tree) *trace.Action
traceAction.DurationMillis = action.DurationMillis
traceAction.X = 0
traceAction.Y = 0
case verifier.ActionKindScroll:
fromX, fromY, toX, toY := scrollEndpoints(action, tree)
traceAction.FromX = fromX
traceAction.FromY = fromY
traceAction.ToX = toX
traceAction.ToY = toY
traceAction.DurationMillis = action.DurationMillis
traceAction.X = 0
traceAction.Y = 0
case verifier.ActionKindPressKey:
traceAction.Key = action.Key
case verifier.ActionKindWait:
@@ -469,6 +777,8 @@ func nextActionFromV8(ctx context.Context, web driver.WebDriver) (verifier.Actio
switch decoded.Kind {
case "Tap":
return verifier.Action{Kind: verifier.ActionKindTap, X: decoded.X, Y: decoded.Y}, nil
case "DoubleTap":
return verifier.Action{Kind: verifier.ActionKindDoubleTap, X: decoded.X, Y: decoded.Y}, nil
case "InputText":
return verifier.Action{
Kind: verifier.ActionKindInputText,
@@ -493,24 +803,18 @@ func nextActionFromV8(ctx context.Context, web driver.WebDriver) (verifier.Actio
}
}
func captureScreenshot(ctx context.Context, options Options, logger *slog.Logger, stepIndex int, after bool) {
image, err := options.Driver.Screenshot(ctx)
if err != nil {
logger.Warn("screenshot capture failed", "step", stepIndex, "after", after, "err", err)
return
func encodeExtractorChanges(changes map[string]verifier.ExtractorChange) map[string]trace.ExtractorChange {
if len(changes) == 0 {
return nil
}
if len(image.PNG) == 0 {
return
}
var writeErr error
if after {
writeErr = options.TraceWriter.WriteScreenshotAfter(stepIndex, image.PNG)
} else {
writeErr = options.TraceWriter.WriteScreenshot(stepIndex, image.PNG)
}
if writeErr != nil {
logger.Warn("screenshot write failed", "step", stepIndex, "after", after, "err", writeErr)
out := make(map[string]trace.ExtractorChange, len(changes))
for name, change := range changes {
out[name] = trace.ExtractorChange{
Prev: json.RawMessage(change.Prev),
Curr: json.RawMessage(change.Curr),
}
}
return out
}
func encodeResiduals(residuals map[string]ltl.Formula) (map[string]json.RawMessage, error) {
@@ -537,3 +841,29 @@ func isWDADrop(err error) bool {
return strings.Contains(msg, "ConnectException") ||
(strings.Contains(msg, "code = Internal") && strings.Contains(msg, "SocketException"))
}
// isTransientApplyError reports whether an applyAction failure is a transient
// device-side hang (sidecar RPC deadline, momentary unavailability) rather than
// a fatal condition. Such steps are recorded as transitional and the loop
// continues. The run context being cancelled is never transient: it means the
// caller wants to stop.
func isTransientApplyError(runCtx context.Context, err error) bool {
if err == nil || runCtx.Err() != nil {
return false
}
if s, ok := status.FromError(err); ok {
switch s.Code() {
case codes.DeadlineExceeded, codes.Unavailable:
return true
case codes.Internal:
message := s.Message()
if strings.Contains(message, "DEADLINE_EXCEEDED") || strings.Contains(message, "UNAVAILABLE") {
return true
}
}
}
if errors.Is(err, context.DeadlineExceeded) {
return true
}
return false
}
+764 -17
View File
@@ -1,8 +1,10 @@
package runner
import (
"bufio"
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"log/slog"
@@ -13,8 +15,12 @@ import (
"testing"
"time"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
"github.com/priyanshujain/sanderling/internal/driver"
mockdriver "github.com/priyanshujain/sanderling/internal/driver/mock"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
@@ -118,6 +124,76 @@ func TestRunner_ViolationSurfacesInSummary(t *testing.T) {
}
}
func TestRunner_ViolationSurfacesOnlyOnOnsetStep(t *testing.T) {
// violationSpec uses always(() => false): onset fires on step 1 and the
// residual stays violated forever. The runner must record the violation
// exactly once (at the onset step) in both summary.Violations and trace
// lines, not on every subsequent step.
state := newHarnessWithSpec(t, violationSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to prove onset-only behavior, got %d", summary.Steps)
}
if len(summary.Violations) != 1 {
t.Fatalf("expected exactly one ViolationRecord (onset only), got %d: %v",
len(summary.Violations), summary.Violations)
}
if summary.Violations[0].StepIndex != 1 {
t.Errorf("onset step: got %d, want 1 (always(()=>false) violates immediately)",
summary.Violations[0].StepIndex)
}
if !slices.Equal(summary.Violations[0].Properties, []string{"balanceNonNegative"}) {
t.Errorf("onset properties: got %v, want [balanceNonNegative]",
summary.Violations[0].Properties)
}
file, err := os.Open(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
defer file.Close()
type traceLine struct {
Step int `json:"step"`
Violations []string `json:"violations"`
}
linesWithViolations := 0
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), 8*1024*1024)
for scanner.Scan() {
var line traceLine
if err := json.Unmarshal(scanner.Bytes(), &line); err != nil {
t.Fatalf("trace line decode: %v", err)
}
if len(line.Violations) == 0 {
continue
}
linesWithViolations++
if line.Step != 1 {
t.Errorf("step %d unexpectedly emitted violations %v (should be onset-only at step 1)",
line.Step, line.Violations)
}
}
if err := scanner.Err(); err != nil {
t.Fatalf("scan trace: %v", err)
}
if linesWithViolations != 1 {
t.Errorf("expected exactly 1 trace line with violations, got %d", linesWithViolations)
}
}
func TestRunner_ThrowingPredicateIsLoggedNotPanic(t *testing.T) {
const throwingSpec = `
globalThis.properties = {
@@ -294,6 +370,151 @@ func TestApplyAction_V8InputTextAtOriginStillTaps(t *testing.T) {
}
}
func TestApplyAction_DoubleTapDispatchesTwoTapsAtCoordinates(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindDoubleTap, X: 100, Y: 200}
start := time.Now()
if err := applyAction(context.Background(), driverMock, action, nil); err != nil {
t.Fatalf("apply action: %v", err)
}
elapsed := time.Since(start)
if elapsed < 40*time.Millisecond {
t.Errorf("expected >= 40ms gap between taps, elapsed %v", elapsed)
}
taps := 0
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionTap && a.X == 100 && a.Y == 200 {
taps++
}
}
if taps != 2 {
t.Errorf("expected 2 Tap calls at (100,200), got %d in %v", taps, driverMock.Actions())
}
}
func TestApplyAction_DoubleTapDispatchesTwoSelectorTaps(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindDoubleTap, On: "id:save"}
if err := applyAction(context.Background(), driverMock, action, nil); err != nil {
t.Fatalf("apply action: %v", err)
}
taps := 0
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionTapSelector && a.Selector == "id:save" {
taps++
}
}
if taps != 2 {
t.Errorf("expected 2 TapSelector calls with id:save, got %d in %v", taps, driverMock.Actions())
}
}
func TestApplyAction_LongPressDispatchesAtResolvedCoordinates(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{Kind: verifier.ActionKindLongPress, X: 120, Y: 240}
if err := applyAction(context.Background(), driverMock, action, nil); err != nil {
t.Fatalf("apply action: %v", err)
}
found := false
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionLongPress && a.X == 120 && a.Y == 240 {
found = true
}
}
if !found {
t.Errorf("expected LongPress at (120,240), got %v", driverMock.Actions())
}
}
func TestApplyAction_ScrollWithPrecomputedEndpointsSwipes(t *testing.T) {
driverMock := mockdriver.New()
action := verifier.Action{
Kind: verifier.ActionKindScroll,
Direction: "down",
FromX: 100,
FromY: 500,
ToX: 100,
ToY: 300,
DurationMillis: 300,
}
if err := applyAction(context.Background(), driverMock, action, nil); err != nil {
t.Fatalf("apply action: %v", err)
}
found := false
for _, a := range driverMock.Actions() {
if a.Kind == mockdriver.ActionSwipe && a.FromX == 100 && a.FromY == 500 && a.ToX == 100 && a.ToY == 300 {
found = true
}
}
if !found {
t.Errorf("expected Swipe with precomputed endpoints, got %v", driverMock.Actions())
}
}
func TestApplyAction_ScrollDirectionUsesInversion(t *testing.T) {
driverMock := mockdriver.New()
treeJSON := `{"attributes":{"resource-id":"com.fixture:id/list","bounds":"[0,0,400,800]"},"children":[],"enabled":true}`
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
action := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "down", On: "id:list"}
if err := applyAction(context.Background(), driverMock, action, tree); err != nil {
t.Fatalf("apply action: %v", err)
}
var swipe *mockdriver.Action
for i := range driverMock.Actions() {
if driverMock.Actions()[i].Kind == mockdriver.ActionSwipe {
a := driverMock.Actions()[i]
swipe = &a
}
}
if swipe == nil {
t.Fatalf("expected a Swipe, got %v", driverMock.Actions())
}
// "down" reveals lower content by dragging the finger up, so toY < fromY.
if swipe.ToY >= swipe.FromY {
t.Errorf("expected toY < fromY for scroll down, got from=%d to=%d", swipe.FromY, swipe.ToY)
}
}
func TestApplyAction_ScrollScreenFallback(t *testing.T) {
driverMock := mockdriver.New()
treeJSON := `{"attributes":{"bounds":"[0,0,400,800]"},"children":[],"enabled":true}`
tree, err := hierarchy.Parse(treeJSON)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
// On unset: container falls back to whole-screen (root) bounds.
action := verifier.Action{Kind: verifier.ActionKindScroll, Direction: "up"}
if err := applyAction(context.Background(), driverMock, action, tree); err != nil {
t.Fatalf("apply action: %v", err)
}
var swipe *mockdriver.Action
for i := range driverMock.Actions() {
if driverMock.Actions()[i].Kind == mockdriver.ActionSwipe {
a := driverMock.Actions()[i]
swipe = &a
}
}
if swipe == nil {
t.Fatalf("expected a Swipe, got %v", driverMock.Actions())
}
if swipe.FromX != 200 || swipe.FromY != 400 {
t.Errorf("expected swipe to start at screen center (200,400), got (%d,%d)", swipe.FromX, swipe.FromY)
}
// "up" reveals upper content by dragging the finger down, so toY > fromY.
if swipe.ToY <= swipe.FromY {
t.Errorf("expected toY > fromY for scroll up, got from=%d to=%d", swipe.FromY, swipe.ToY)
}
}
func TestRunner_ParallelFetchCallsAllDriverMethods(t *testing.T) {
state := newHarness(t)
state.mock.MetricsData = driver.Metrics{CPUPercent: 5.0, HeapBytes: 1024, TotalMemoryBytes: 4096}
@@ -316,19 +537,19 @@ func TestRunner_ParallelFetchCallsAllDriverMethods(t *testing.T) {
}
actions := state.mock.Actions()
var hasHierarchy, hasMetrics, hasLogs bool
var hasSnapshot, hasMetrics, hasLogs bool
for _, a := range actions {
switch a.Kind {
case mockdriver.ActionHierarchy:
hasHierarchy = true
case mockdriver.ActionSnapshot:
hasSnapshot = true
case mockdriver.ActionMetrics:
hasMetrics = true
case mockdriver.ActionRecentLogs:
hasLogs = true
}
}
if !hasHierarchy {
t.Error("expected Hierarchy call in mock actions")
if !hasSnapshot {
t.Error("expected Snapshot call in mock actions")
}
if !hasMetrics {
t.Error("expected Metrics call in mock actions")
@@ -338,7 +559,56 @@ func TestRunner_ParallelFetchCallsAllDriverMethods(t *testing.T) {
}
}
func TestRunner_PipelinedPostScreenshotWritten(t *testing.T) {
// TestRunner_UsesAtomicSnapshot ensures the runner observes a step's UI
// through the paired Snapshot RPC instead of racing two independent
// hierarchy + screenshot reads. The pair must come from one on-device
// frame; a regression to separate calls is what this test catches.
func TestRunner_UsesAtomicSnapshot(t *testing.T) {
state := newHarness(t)
state.mock.ImageData = driver.Image{PNG: []byte("png"), Width: 1, Height: 1}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected at least one step")
}
var snapshotCalls, hierarchyCalls, screenshotCalls int
for _, action := range state.mock.Actions() {
switch action.Kind {
case mockdriver.ActionSnapshot:
snapshotCalls++
case mockdriver.ActionHierarchy:
hierarchyCalls++
case mockdriver.ActionScreenshot:
screenshotCalls++
}
}
if snapshotCalls == 0 {
t.Errorf("expected at least one Snapshot call, got %d", snapshotCalls)
}
if hierarchyCalls != 0 {
t.Errorf("expected zero standalone Hierarchy calls (runner must use Snapshot), got %d", hierarchyCalls)
}
if screenshotCalls != 0 {
t.Errorf("expected zero standalone Screenshot calls (runner must use Snapshot), got %d", screenshotCalls)
}
}
// TestRunner_OneScreenshotPerStep verifies the runner writes a single
// screenshot per step, captured concurrently with hierarchy so the two
// observations describe the same UI moment.
func TestRunner_OneScreenshotPerStep(t *testing.T) {
state := newHarness(t)
state.mock.ImageData = driver.Image{PNG: []byte("fakepng"), Width: 100, Height: 200}
@@ -355,24 +625,362 @@ func TestRunner_PipelinedPostScreenshotWritten(t *testing.T) {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps for pipelining test, got %d", summary.Steps)
t.Fatalf("need at least 2 steps for screenshot test, got %d", summary.Steps)
}
screenshotDir := filepath.Join(state.writer.Directory(), "screenshots")
preFile := filepath.Join(screenshotDir, "step-00001.png")
if _, err := os.Stat(preFile); os.IsNotExist(err) {
t.Errorf("expected pre-screenshot for step 1: %s", preFile)
for step := 1; step <= summary.Steps; step++ {
path := filepath.Join(screenshotDir, fmt.Sprintf("step-%05d.png", step))
if _, err := os.Stat(path); os.IsNotExist(err) {
t.Errorf("expected screenshot for step %d at %s", step, path)
}
}
postFile := filepath.Join(screenshotDir, "step-00001-after.png")
if _, err := os.Stat(postFile); os.IsNotExist(err) {
t.Errorf("expected pipelined post-screenshot for step 1: %s", postFile)
entries, err := os.ReadDir(screenshotDir)
if err != nil {
t.Fatal(err)
}
for _, entry := range entries {
if strings.Contains(entry.Name(), "-after") {
t.Errorf("unexpected -after screenshot remains: %s", entry.Name())
}
}
}
// TestIsTransitionalHierarchy_DetectsMultipleScreens covers the runner-side
// guard that re-fetches when the hierarchy still carries two route-level
// *Screen ids - the NavHost cross-fade signature.
func TestIsTransitionalHierarchy_DetectsMultipleScreens(t *testing.T) {
multi, err := hierarchy.Parse(`{"attributes":{"resource-id":"root"},"children":[
{"attributes":{"resource-id":"AddAccountScreen"},"children":[]},
{"attributes":{"resource-id":"HomeScreen"},"children":[]}
]}`)
if err != nil {
t.Fatal(err)
}
if !isTransitionalHierarchy(multi) {
t.Error("expected multi-screen tree to be flagged as transitional")
}
lastAfter := filepath.Join(screenshotDir, fmt.Sprintf("step-%05d-after.png", summary.Steps))
if _, err := os.Stat(lastAfter); os.IsNotExist(err) {
t.Errorf("expected flushed post-screenshot for last step %d: %s", summary.Steps, lastAfter)
single, err := hierarchy.Parse(`{"attributes":{"resource-id":"HomeScreen"},"children":[]}`)
if err != nil {
t.Fatal(err)
}
if isTransitionalHierarchy(single) {
t.Error("single-screen tree must not be flagged as transitional")
}
if isTransitionalHierarchy(nil) {
t.Error("nil tree must not be flagged as transitional")
}
}
// TestRunner_TransitionalSkipsVerifier feeds a driver whose hierarchy stays
// transitional (multiple route-level *Screen ids) on every Snapshot call.
// Every step must be marked transitional in the trace, no violations may be
// emitted (the verifier never ran), and the summary must stay clean even
// though the spec is a guaranteed always-false predicate.
func TestRunner_TransitionalSkipsVerifier(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"root"},"children":[
{"attributes":{"resource-id":"AddAccountScreen"},"children":[]},
{"attributes":{"resource-id":"HomeScreen"},"children":[]}
]}`
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps == 0 {
t.Fatal("expected at least one step")
}
if len(summary.Violations) != 0 {
t.Fatalf("verifier must be skipped on transitional steps; got %v", summary.Violations)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
lines := 0
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
lines++
if !line.Transitional {
t.Errorf("step %d: expected transitional=true on every step, got false", line.Step)
}
if len(line.Violations) != 0 {
t.Errorf("step %d: verifier must be skipped, got violations %v", line.Step, line.Violations)
}
}
if lines == 0 {
t.Fatal("expected trace lines, got none")
}
}
// TestRunner_CleanTreeStillVerified is the control: a single-screen hierarchy
// must not be marked transitional and the verifier must still run, surfacing
// the always-false predicate's violation on the onset step.
func TestRunner_CleanTreeStillVerified(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"HomeScreen"},"children":[]}`
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if !containsProperty(summary.Violations, "balanceNonNegative") {
t.Fatalf("expected verifier to surface balanceNonNegative on a clean tree, got %v", summary.Violations)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
for _, raw := range bytes.Split(bytes.TrimSpace(body), []byte("\n")) {
var line traceLine
if err := json.Unmarshal(raw, &line); err != nil {
t.Fatalf("decode trace line: %v", err)
}
if line.Transitional {
t.Errorf("step %d: clean tree must not be marked transitional", line.Step)
}
}
}
// snapshotFailFirst wraps a mock driver so the first Snapshot call returns an
// error (mimicking a sidecar timeout while fetching view hierarchy), then
// delegates every subsequent call back to the mock.
type snapshotFailFirst struct {
*mockdriver.Driver
calls int
}
func (d *snapshotFailFirst) Snapshot(ctx context.Context) (string, driver.Image, error) {
d.calls++
if d.calls == 1 {
return "", driver.Image{}, errors.New("Timeout while fetching view hierarchy")
}
return d.Driver.Snapshot(ctx)
}
// TestRunner_NilHierarchyMarksTransitional verifies that when the sidecar's
// hierarchy fetch fails (nil tree), the runner marks the step transitional and
// skips the verifier instead of pushing a nil tree that would crash the spec.
// Subsequent steps with a clean tree still drive the verifier normally.
func TestRunner_NilHierarchyMarksTransitional(t *testing.T) {
state := newHarnessWithSpec(t, violationSpec)
state.mock.HierarchyJSON = `{"attributes":{"resource-id":"HomeScreen"},"children":[]}`
wrapped := &snapshotFailFirst{Driver: state.mock}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 200 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to verify the first is skipped and the second runs, got %d", summary.Steps)
}
// violationSpec always() => false fires on the first verifier push. With
// step 1's verifier skipped, onset moves to step 2.
if len(summary.Violations) != 1 {
t.Fatalf("expected exactly one onset record, got %d: %v", len(summary.Violations), summary.Violations)
}
if summary.Violations[0].StepIndex != 2 {
t.Errorf("onset step: got %d, want 2 (step 1 verifier skipped due to nil tree)", summary.Violations[0].StepIndex)
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
var first traceLine
if err := json.Unmarshal(bytes.SplitN(bytes.TrimSpace(body), []byte("\n"), 2)[0], &first); err != nil {
t.Fatalf("decode first trace line: %v", err)
}
if first.Step != 1 {
t.Fatalf("first trace line step: got %d, want 1", first.Step)
}
if !first.Transitional {
t.Error("first step must be marked transitional when the hierarchy fetch failed")
}
if len(first.Violations) != 0 {
t.Errorf("step 1 must skip the verifier; got violations %v", first.Violations)
}
}
// tapSelectorFailFirst wraps a mock driver so the first TapSelector call
// returns a gRPC DeadlineExceeded error (mimicking a sidecar-side RPC hang),
// then delegates every subsequent call back to the mock.
type tapSelectorFailFirst struct {
*mockdriver.Driver
calls int
}
func (d *tapSelectorFailFirst) TapSelector(ctx context.Context, selector string) error {
d.calls++
if d.calls == 1 {
return status.Error(codes.DeadlineExceeded, "boom")
}
return d.Driver.TapSelector(ctx, selector)
}
// TestRunner_TransientApplyErrorMarksTransitional verifies that a transient
// gRPC error from applyAction (e.g. sidecar RPC deadline) does not kill the
// run: the step is marked transitional, the verifier is skipped for it, and
// the loop continues with the next step running cleanly.
func TestRunner_TransientApplyErrorMarksTransitional(t *testing.T) {
state := newHarness(t)
wrapped := &tapSelectorFailFirst{Driver: state.mock}
var logBuf bytes.Buffer
logger := slog.New(slog.NewTextHandler(&logBuf, &slog.HandlerOptions{Level: slog.LevelWarn}))
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
summary, err := Run(ctx, Options{
Duration: 300 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
Driver: wrapped,
Verifier: state.verifier,
TraceWriter: state.writer,
Logger: logger,
})
if err != nil {
t.Fatalf("Run must not return on transient apply error, got %v", err)
}
if summary.Steps < 2 {
t.Fatalf("need at least 2 steps to prove the loop continued past the failed apply, got %d", summary.Steps)
}
if len(summary.Violations) != 0 {
t.Errorf("transient apply error must not surface as a violation, got %v", summary.Violations)
}
if !strings.Contains(logBuf.String(), "transient apply error") {
t.Errorf("expected transient-apply WARN log, got %q", logBuf.String())
}
type traceLine struct {
Step int `json:"step"`
Transitional bool `json:"transitional"`
Violations []string `json:"violations"`
}
body, err := os.ReadFile(filepath.Join(state.writer.Directory(), "trace.jsonl"))
if err != nil {
t.Fatal(err)
}
lines := bytes.Split(bytes.TrimSpace(body), []byte("\n"))
var first, second traceLine
if err := json.Unmarshal(lines[0], &first); err != nil {
t.Fatalf("decode first trace line: %v", err)
}
if err := json.Unmarshal(lines[1], &second); err != nil {
t.Fatalf("decode second trace line: %v", err)
}
if first.Step != 1 || !first.Transitional {
t.Errorf("step 1 must be transitional after transient apply error, got step=%d transitional=%v", first.Step, first.Transitional)
}
if len(first.Violations) != 0 {
t.Errorf("transient apply step must have no violations, got %v", first.Violations)
}
if second.Step != 2 || second.Transitional {
t.Errorf("step 2 must run cleanly after the transient step, got step=%d transitional=%v", second.Step, second.Transitional)
}
}
// TestIsTransientApplyError_Classification covers the helper's matching rules
// directly so future code changes don't quietly drop a transient case.
func TestIsTransientApplyError_Classification(t *testing.T) {
cleanCtx := context.Background()
cancelledCtx, cancel := context.WithCancel(context.Background())
cancel()
cases := []struct {
name string
ctx context.Context
err error
want bool
}{
{"nil error", cleanCtx, nil, false},
{"deadline exceeded", cleanCtx, status.Error(codes.DeadlineExceeded, "boom"), true},
{"unavailable", cleanCtx, status.Error(codes.Unavailable, "boom"), true},
{"internal wrapping deadline", cleanCtx, status.Error(codes.Internal, "io.grpc.StatusRuntimeException: DEADLINE_EXCEEDED: ..."), true},
{"internal wrapping unavailable", cleanCtx, status.Error(codes.Internal, "io.grpc.StatusRuntimeException: UNAVAILABLE: ..."), true},
{"internal generic", cleanCtx, status.Error(codes.Internal, "boom"), false},
{"raw context deadline", cleanCtx, context.DeadlineExceeded, true},
{"run context cancelled overrides", cancelledCtx, status.Error(codes.DeadlineExceeded, "boom"), false},
}
for _, testCase := range cases {
t.Run(testCase.name, func(t *testing.T) {
if got := isTransientApplyError(testCase.ctx, testCase.err); got != testCase.want {
t.Errorf("got %v, want %v", got, testCase.want)
}
})
}
}
// TestRunner_WaitActionSkipsIdle ensures the runner does not call WaitForIdle
// after a Wait action - the action already provides settling time.
func TestRunner_WaitActionSkipsIdle(t *testing.T) {
const waitSpec = `
globalThis.actions = __sanderling__.actions(() => [__sanderling__.wait({ durationMillis: 5 })]);
`
state := newHarnessWithSpec(t, waitSpec)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 150 * time.Millisecond,
IdleTimeout: 50 * time.Millisecond,
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
for _, action := range state.mock.Actions() {
if action.Kind == mockdriver.ActionWaitForIdle {
t.Fatalf("Wait action must skip WaitForIdle, got: %v", action)
}
}
}
@@ -426,3 +1034,142 @@ func containsProperty(records []ViolationRecord, property string) bool {
}
return false
}
func TestRunner_RelaunchesWhenAppLeavesForeground(t *testing.T) {
state := newHarness(t)
// Always report a foreign app, so every step's guard must relaunch.
state.mock.ForegroundResults = []string{"com.android.chrome"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
relaunches := 0
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionLaunch && a.BundleID == "app.folio" && !a.ClearState {
relaunches++
}
}
if relaunches == 0 {
t.Fatal("expected runner to relaunch app.folio when foreground escaped, got none")
}
}
func TestRunner_NoRelaunchWhenAppInForeground(t *testing.T) {
state := newHarness(t)
state.mock.ForegroundResults = []string{"app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionLaunch {
t.Fatalf("expected no relaunch while app in foreground, got %v", a)
}
}
}
// TestRunner_WaitsForForegroundBeforeFirstAction verifies the startup gate
// brings the app forward (back-press + relaunch) before any tap fires when the
// device boots showing a system dialog.
func TestRunner_WaitsForForegroundBeforeFirstAction(t *testing.T) {
state := newHarness(t)
// First the device shows a system setup screen, then the app is on top.
state.mock.ForegroundResults = []string{"com.google.android.setupwizard", "app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
actions := state.mock.Actions()
firstLaunch, firstTap := -1, -1
backPressed := false
for i, a := range actions {
switch {
case a.Kind == mockdriver.ActionLaunch && firstLaunch < 0:
firstLaunch = i
case a.Kind == mockdriver.ActionTap && firstTap < 0:
firstTap = i
case a.Kind == mockdriver.ActionPressKey && a.Key == "back":
backPressed = true
}
}
if firstLaunch < 0 {
t.Fatal("expected a relaunch to bring the app forward, got none")
}
if !backPressed {
t.Fatal("expected a back-press to dismiss the system dialog, got none")
}
if firstTap >= 0 && firstLaunch > firstTap {
t.Fatalf("expected the foreground gate (launch at %d) before the first tap (at %d)", firstLaunch, firstTap)
}
}
// TestRunner_WaitsForWindowDrawnBeforeFirstAction verifies the startup gate
// keeps waiting while the app is the resumed activity but its window has not
// drawn yet (a leftover screen still focused). It must poll the focused-window
// signal rather than relaunching, and only proceed once the window names the
// app.
func TestRunner_WaitsForWindowDrawnBeforeFirstAction(t *testing.T) {
state := newHarness(t)
// The app is resumed immediately, but its window lags: the outgoing
// settings screen stays focused for two checks before the app draws.
state.mock.ForegroundResults = []string{"app.folio"}
state.mock.FocusedWindowResults = []string{"com.android.settings", "com.android.settings", "app.folio"}
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
_, err := Run(ctx, Options{
Duration: 100 * time.Millisecond,
IdleTimeout: 20 * time.Millisecond,
BundleID: "app.folio",
Driver: state.mock,
Verifier: state.verifier,
TraceWriter: state.writer,
})
if err != nil {
t.Fatalf("Run: %v", err)
}
// The gate must have polled the focused window until it named the app,
// i.e. at least the three queued results were consumed.
if calls := state.mock.FocusedWindowCalls(); calls < 3 {
t.Fatalf("expected the gate to poll the focused window until drawn (>=3 calls), got %d", calls)
}
// The resumed app was never a foreign app, so the gate must not relaunch
// or back-press to "fix" a window that simply had not drawn yet.
for _, a := range state.mock.Actions() {
if a.Kind == mockdriver.ActionPressKey && a.Key == "back" {
t.Fatal("expected no back-press while waiting for the window to draw")
}
}
}