Files
sanderling/conformance/gates.sh
T
pj 6b0d6cb971 WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial

* feat(test): add --device flag to target a specific Android device by serial

* feat(folio): select Android device via ANDROID_DEVICE in justfile

* feat(conformance): add android backend to the gate suite

* feat(android): keep device awake and unlocked so the app stays foreground

* feat(conformance): prep physical android device (autofill/verifier/stayon)

* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run

* fix(verifier): require positive bounds for swipe candidates

A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.

* fix(runner): harden app-scope guard against launcher and overlays

The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.

* feat(android): harden physical-device runs in device prep

Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.

* feat(driver): clear-state via APK reinstall when pm clear is blocked

When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.

* feat(cli): add --android-app-path for clear-state reinstall

Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.

* chore(folio): pass --android-app-path in just test

* fix(runner): clamp swipe/scroll origin out of edge gesture zones

A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.

* perf(sidecar): faster Android text input and drop redundant settle poll

inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.

* fix(verifier): exclude soft-keyboard region from action candidates

The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.

* perf(runner): replace focus-tap settle with a brief wait

The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.

* chore(conformance): platform-aware G5 p95 budget for android

The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.

* fix(sidecar): retry maestro android driver startup

The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.

* chore(conformance): widen android G5 budget to 5500ms

Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.

* web replay fix

* feat(android): force 3-button nav during runs to prevent app drift

On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.

* fix(android): target the selected device in adb reads; don't strand nav mode

Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
  foreground/scope guard works when several devices are attached (the --device
  path). Previously they ran bare `adb shell`, which errors with multiple
  devices, silently disabling app-scope enforcement. The sidecar client passes
  its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
  duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
  the current mode is unknown or already 3-button it leaves nav untouched,
  instead of switching and then stranding the device in 3-button. Logic split
  into the pure navModeToRestore, now unit tested.

* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds

Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.

* test(verifier): cover keyboardRegionTop, including the decor-view guard

The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.

* style(cli): gofmt testOptions field alignment

* fix(sidecar): keep a leading dash off the fast input path

A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).

* refactor(verifier): scope action candidates by window ownership

Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.

This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.

* fix(runner): re-check foreground at apply time, skip stale actions

ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.

* fix(android): type long ASCII via fast guarded path to stop keystroke escape

A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.

* chore: ignore gate artifacts and local scratch files

* refactor(runner): narrow gesture clamp to the top shade strip

3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.

* chore(format): add .editorconfig enforcing 80-column limit

* chore(format): add prettier config with 80-char printWidth

* chore(deps): add prettier devDependency to replay-ui

* chore(deps): add prettier devDependency to folio-web

* chore(deps): add prettier devDependency to spec package

* chore(format): add swift-format config with 80-char lineLength

* feat(format): add make fmt targets for per-language 80-col formatting

* fix(runner): translate gesture to safe area so near-top scrolls keep direction

Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.

* fix(runner): apply-time guard consults focused window, not just resumed activity

ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.

* test(runner): cover apply-time foreground skip and appIsForeground table

Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.

* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker

- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.

* fix(conformance): pin self-test p95 budget and score install failures as run failures

self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.

* fix(android): require --device when several devices are connected

With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.

* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands

The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.

* perf(verifier): memoize scopedElements per tree

scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.

* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear

Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.

* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms

The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.

* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty

bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.

* perf(sidecar): reuse a single Jackson ObjectMapper

structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.

* refactor(android): remove unused AdbReverse/AdbReverseRemove

No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.

* style(runner): trim non-load-bearing comments from this PR's runner code and tests

* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
2026-06-11 10:10:05 +05:30

476 lines
19 KiB
Bash
Executable File

#!/usr/bin/env bash
# Scripted conformance gate for the mobile drivers. Runs five serial,
# non-overlapping 3-minute fuzz runs against the folio example app
# (examples/folio) and scores five gates (G1..G5) over the captured traces and
# output. Exits non-zero if any gate fails.
#
# Backends:
# BACKEND=simulator (default) drive the booted iOS simulator
# BACKEND=device drive an attached physical iPhone via the
# driver's runner-only device path; select it
# with IOS_DEVICE="<name>" (passed as --ios-device)
# BACKEND=android drive an Android device/emulator over the JVM
# sidecar; select a specific device with
# ANDROID_DEVICE="<adb serial>" (passed as --device)
#
# Usage:
# ./gates.sh run the simulator gates
# BACKEND=device IOS_DEVICE="iPhone" ./gates.sh
# BACKEND=android ANDROID_DEVICE="663c91b1" ./gates.sh
# ./gates.sh --self-test run the offline analyzer tests only
#
# Tunables (environment):
# RUNS=5 number of serial runs
# DURATION=3m per-run fuzz duration
# SEED=0 fuzz seed
# P95_LIMIT_MS G5 p95 step-latency ceiling in ms (default 2500 for iOS,
# 5500 for the android backend's higher and more variable
# per-step USB cost, especially cold right after a reboot)
# SANDERLING=sanderling binary to invoke
set -euo pipefail
script_directory="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
folio_directory="$(cd "${script_directory}/../examples/folio" && pwd)"
BACKEND="${BACKEND:-simulator}"
RUNS="${RUNS:-5}"
DURATION="${DURATION:-3m}"
SEED="${SEED:-0}"
# The p95 ceiling is backend-specific: a physical Android device drives every
# step over USB (snapshot + settle + adb round-trips), so its per-step floor is
# several times the iOS simulator's in-process cost. 2500ms was calibrated on
# the simulator; holding a physical device to it would force ripping out the
# settle/retry logic the correctness gates depend on. Override with P95_LIMIT_MS.
if [[ "$BACKEND" == "android" ]]; then
P95_LIMIT_MS="${P95_LIMIT_MS:-5500}"
else
P95_LIMIT_MS="${P95_LIMIT_MS:-2500}"
fi
SANDERLING="${SANDERLING:-sanderling}"
IOS_DEVICE="${IOS_DEVICE:-iPhone 17 Pro}"
ANDROID_DEVICE="${ANDROID_DEVICE:-}"
bundle_id="app.folio"
spec_path="${folio_directory}/sanderling/spec.ts"
android_apk="${folio_directory}/app/androidApp/build/outputs/apk/debug/androidApp-debug.apk"
# The built app bundle differs by SDK: the simulator build lands under
# Debug-iphonesimulator, the device build under Debug-iphoneos.
if [[ "$BACKEND" == "device" ]]; then
ios_app="${folio_directory}/app/iosApp/build/Build/Products/Debug-iphoneos/iosApp.app"
else
ios_app="${folio_directory}/app/iosApp/build/Build/Products/Debug-iphonesimulator/iosApp.app"
fi
# Run adb against the selected Android device, or the only one if unset.
adb_target() {
if [[ -n "$ANDROID_DEVICE" ]]; then adb -s "$ANDROID_DEVICE" "$@"; else adb "$@"; fi
}
# The companion binary, embedded for simulator runs. Referenced by file name
# only for the orphan-process check; prose elsewhere says "the companion".
companion_process_name="idb_companion"
# Known-benign stderr lines the companion always prints. These are matched as
# fixed substrings and excluded from the G2 ERROR scan. Keep this list tight:
# only lines that are provably harmless and emitted on every healthy run.
benign_stderr_substrings=(
# The dynamic linker reports the same Objective-C class registered by two
# loaded images. The companion runs fine; this is cosmetic. The real line
# reads "objc[<pid>]: Class ...", so the fixed substring is "]: Class".
"]: Class"
"is implemented in both"
"One of the two will be used. Which one is undefined."
# gRPC and absl emit informational banner lines on startup.
"WARNING: All log messages before absl::InitializeLog()"
)
# ---- gate analyzers (pure, operate on a single run directory) --------------
# G2 helper: strip benign companion noise, then report sanderling ERROR lines.
# sanderling's progress logger renders error-level records as lines beginning
# "error:" (see internal/testrun/progress.go). We also catch the upper-case
# ERROR token (word-bounded, so ERRORS/ERRORLESS and path fragments do not
# trip the gate) for safety against future handlers.
error_lines_in() {
local output_file="$1"
local filtered
filtered="$(cat "$output_file")"
local pattern
for pattern in "${benign_stderr_substrings[@]}"; do
filtered="$(printf '%s\n' "$filtered" | grep -vF "$pattern" || true)"
done
printf '%s\n' "$filtered" | grep -E '(^error:|\bERROR\b)' || true
}
# G1: process exit status recorded by the runner loop.
gate_exit_zero() {
local run_directory="$1"
[[ "$(cat "${run_directory}/exit_status")" == "0" ]]
}
# G2: no sanderling ERROR lines after filtering benign companion noise.
gate_no_error_lines() {
local run_directory="$1"
local found
found="$(error_lines_in "${run_directory}/output.log")"
[[ -z "$found" ]]
}
# G3: the first hierarchy snapshot shows the login screen with empty email and
# password fields (clear-state proof). Reads the first trace line carrying a
# hierarchy and asserts LoginEmail/LoginPassword carry no text.
gate_clear_state() {
local run_directory="$1"
local trace_file="${run_directory}/trace.jsonl"
[[ -f "$trace_file" ]] || return 1
local verdict
verdict="$(jq -s -r '
[ .[] | select(.hierarchy != null) ] as $withHierarchy
| if ($withHierarchy | length) == 0 then "fail:no-hierarchy"
else ($withHierarchy[0].hierarchy.elements // []) as $elements
| ($elements | map(select(.resourceId == "LoginEmail")) | first) as $email
| ($elements | map(select(.resourceId == "LoginPassword")) | first) as $password
| if $email == null or $password == null then "fail:no-login-fields"
elif (($email.attrs.text // $email.text // "") != "") then "fail:email-not-empty"
elif (($password.attrs.text // $password.text // "") != "") then "fail:password-not-empty"
else "pass" end
end
' "$trace_file")"
[[ "$verdict" == "pass" ]]
}
# G4: no doubled text after InputText. For each InputText action targeting a
# field, the field's value in the NEXT hierarchy must not contain the input
# concatenated with itself (catches append-vs-replace and double-paste bugs).
# The action chosen at step N is applied before step N+1 is observed, so the
# effect lands in the following snapshot.
gate_no_doubled_text() {
local run_directory="$1"
local trace_file="${run_directory}/trace.jsonl"
[[ -f "$trace_file" ]] || return 1
local verdict
verdict="$(jq -s -r '
# Map a selector like "testTag:LoginScreen > testTag:LoginEmail" to its
# target field id: the token after the final ":" of the last segment.
def target_field(selector):
(selector | split(">") | last | gsub("^\\s+|\\s+$";"")) as $last
| ($last | split(":") | last);
[ .[] | select(.hierarchy != null) ] as $steps
| reduce range(0; ($steps | length)) as $i ([];
($steps[$i]) as $current
| (if $i + 1 < ($steps | length) then $steps[$i + 1] else null end) as $next
| if ($current.next_action.kind == "InputText") and ($next != null)
then
(target_field($current.next_action.selector // "")) as $field
| ($current.next_action.text // "") as $typed
| (($next.hierarchy.elements // [])
| map(select(.resourceId == $field)) | first) as $element
| if $element != null and $typed != ""
then
(($element.attrs.text // $element.text // "")) as $value
| if ($value | contains($typed + $typed))
then . + [{field: $field, typed: $typed, value: $value}]
else . end
else . end
else . end)
| if length == 0 then "pass"
else "fail:" + (.[0].field) + ":" + (.[0].value) end
' "$trace_file")"
[[ "$verdict" == "pass" ]]
}
# Emit one step-latency sample per consecutive trace-step pair, in
# milliseconds, on stdout. The trace records a wall-clock timestamp per step;
# the latency of a step is the gap to the next step's observation. Used by G5,
# which aggregates samples across all runs before computing the p95. Timestamps
# are RFC3339 with fractional seconds and a numeric offset, so they are parsed
# with python3 rather than jq's UTC-only fromdateiso8601.
emit_step_latencies() {
local trace_file="$1"
[[ -f "$trace_file" ]] || return 0
jq -r 'select(.timestamp != null) | .timestamp' "$trace_file" \
| python3 -c '
import sys
from datetime import datetime
stamps = [datetime.fromisoformat(line.strip()) for line in sys.stdin if line.strip()]
for earlier, later in zip(stamps, stamps[1:]):
print(int((later - earlier).total_seconds() * 1000))
'
}
# Compute the p95 (nearest-rank) of the latency samples on stdin, in ms.
p95_of() {
python3 -c '
import sys, math
samples = sorted(int(float(line)) for line in sys.stdin if line.strip())
if not samples:
print(0)
sys.exit(0)
rank = max(1, math.ceil(0.95 * len(samples)))
print(samples[rank - 1])
'
}
# G5 orphan check: report any lingering companion, runner session (the hybrid
# simulator driver hosts an in-simulator runner), and, on the device backend,
# the device runner session. The usbmux tunnel is an in-process forwarder that
# dies with sanderling, so it leaves no process to check. Empty output is clean.
orphan_processes() {
local found=""
if [[ "$BACKEND" == "android" ]]; then
# sanderling SIGTERMs the JVM sidecar on shutdown; a survivor is an orphan.
if pgrep -f "sanderling-sidecar.*\.jar" >/dev/null 2>&1; then found+="sidecar "; fi
printf '%s' "$found"
return
fi
if pgrep -f "$companion_process_name" >/dev/null 2>&1; then
found+="companion "
fi
if pgrep -f "sanderling-runner.*xctestrun" >/dev/null 2>&1; then
found+="runner-session "
fi
if pgrep -f "CompanionRunnerUITests-Runner" >/dev/null 2>&1; then
found+="runner-app "
fi
if [[ "$BACKEND" == "device" ]]; then
# The device test session that hosts the runner. Its destination carries
# platform=iOS,id=<udid>.
if pgrep -f "xctestrun.*platform=iOS,id=" >/dev/null 2>&1; then
found+="device-session "
fi
fi
printf '%s' "$found"
}
# ---- run orchestration -----------------------------------------------------
invoke_sanderling() {
local output_directory="$1"
local output_log="$2"
local exit_status_file="$3"
local platform target_flags=()
if [[ "$BACKEND" == "android" ]]; then
# pm clear is blocked on some OEM ROMs, so a fresh install (which wipes
# /data/data) provides the clean clear-state start; --clear-data=false
# then skips the sidecar's pm clear. Mirrors the iOS per-run reinstall.
platform=android
adb_target uninstall "$bundle_id" >/dev/null 2>&1 || true
# A transient install hiccup must score this run as a failure, not abort the
# whole harness under `set -e` and discard the other runs' data.
if ! adb_target install "$android_apk" >"$output_log" 2>&1; then
echo "adb install failed for ${android_apk}; recording run as a failure" >>"$output_log"
printf '1' >"$exit_status_file"
return
fi
target_flags=(--clear-data=false)
[[ -n "$ANDROID_DEVICE" ]] && target_flags+=(--device "$ANDROID_DEVICE")
else
# The iOS backends pass --ios-app-path so each run reinstalls the current
# build for a clean start (device via devicectl, simulator via simctl).
platform=ios
target_flags=(--ios-device "$IOS_DEVICE" --ios-app-path "$ios_app")
fi
local status=0
"$SANDERLING" test \
--platform "$platform" \
--spec "$spec_path" \
--bundle-id "$bundle_id" \
"${target_flags[@]}" \
--duration "$DURATION" \
--seed "$SEED" \
--output "$output_directory" \
>"$output_log" 2>&1 || status=$?
printf '%s' "$status" >"$exit_status_file"
}
# Locate the run directory sanderling created under output_directory (it nests
# a timestamped subdirectory) and normalise its artifacts up one level.
collect_run_artifacts() {
local output_directory="$1"
[[ -f "${output_directory}/trace.jsonl" ]] && return 0
local produced
produced="$(find "$output_directory" -mindepth 2 -name 'trace.jsonl' 2>/dev/null | head -1 || true)"
if [[ -n "$produced" ]]; then
cp "$produced" "${output_directory}/trace.jsonl" 2>/dev/null || true
fi
}
run_gates() {
if [[ "$BACKEND" == "android" ]]; then
# Build only; invoke_sanderling reinstalls per run via adb (gradle's ddmlib
# install is flaky on some physical devices).
echo "preparing folio android build"
( cd "$folio_directory" && just build >/dev/null )
# A physical device, unlike an emulator, lets system UI steal the foreground
# from the app the fuzzer is exploring. Keep the screen on so it never
# re-locks, silence the autofill save-password prompt that pops over the
# login form, and stop Play Protect from intercepting the per-run reinstall.
# The device must already be unlocked (a secure lock cannot be opened here).
adb_target shell svc power stayon true >/dev/null 2>&1 || true
adb_target shell settings put secure autofill_service null >/dev/null 2>&1 || true
adb_target shell settings put global verifier_verify_adb_installs 0 >/dev/null 2>&1 || true
elif [[ "$BACKEND" == "simulator" ]]; then
echo "preparing folio build for the simulator backend"
( cd "$folio_directory" && just ios >/dev/null )
else
echo "preparing folio device build"
( cd "$folio_directory" && just ios-device >/dev/null )
fi
local timestamp
timestamp="$(date +%Y%m%d-%H%M%S)"
local gate_root="${script_directory}/runs/${timestamp}"
mkdir -p "$gate_root"
echo "gate root: ${gate_root}"
echo "backend=${BACKEND} runs=${RUNS} duration=${DURATION} seed=${SEED} p95_limit_ms=${P95_LIMIT_MS}"
local all_latencies="${gate_root}/all-latencies.txt"
: >"$all_latencies"
local -a g1 g2 g3 g4 g5
local run_index
for ((run_index = 1; run_index <= RUNS; run_index++)); do
local run_directory="${gate_root}/run-${run_index}"
mkdir -p "$run_directory"
echo "run ${run_index}/${RUNS} -> ${run_directory}"
invoke_sanderling "$run_directory" "${run_directory}/output.log" "${run_directory}/exit_status"
collect_run_artifacts "$run_directory"
g1[run_index]=$(gate_exit_zero "$run_directory" && echo PASS || echo FAIL)
g2[run_index]=$(gate_no_error_lines "$run_directory" && echo PASS || echo FAIL)
g3[run_index]=$(gate_clear_state "$run_directory" && echo PASS || echo FAIL)
g4[run_index]=$(gate_no_doubled_text "$run_directory" && echo PASS || echo FAIL)
emit_step_latencies "${run_directory}/trace.jsonl" >>"$all_latencies"
local orphans
orphans="$(orphan_processes)"
if [[ -n "$orphans" ]]; then
g5[run_index]="FAIL"
echo " orphaned processes after run ${run_index}: ${orphans}"
else
g5[run_index]="PENDING"
fi
done
local p95
p95="$(p95_of <"$all_latencies")"
local final_orphans
final_orphans="$(orphan_processes)"
# G5 is global: p95 is computed over every run's samples and an orphan after
# any single run fails the whole gate. Decide once, then stamp every row.
local g5_global="PASS"
if [[ "$p95" -ge "$P95_LIMIT_MS" ]]; then
g5_global="FAIL"
fi
if [[ -n "$final_orphans" ]]; then
g5_global="FAIL"
echo "orphaned processes at end: ${final_orphans}"
fi
for ((run_index = 1; run_index <= RUNS; run_index++)); do
if [[ "${g5[run_index]}" == "FAIL" ]]; then
g5_global="FAIL"
fi
done
for ((run_index = 1; run_index <= RUNS; run_index++)); do
g5[run_index]="$g5_global"
done
echo
printf 'run G1 G2 G3 G4 G5\n'
local verdict="PASS"
for ((run_index = 1; run_index <= RUNS; run_index++)); do
printf '%-4s %-5s %-5s %-5s %-5s %-5s\n' \
"$run_index" "${g1[run_index]}" "${g2[run_index]}" "${g3[run_index]}" "${g4[run_index]}" "${g5[run_index]}"
for cell in "${g1[run_index]}" "${g2[run_index]}" "${g3[run_index]}" "${g4[run_index]}" "${g5[run_index]}"; do
[[ "$cell" == "PASS" ]] || verdict="FAIL"
done
done
echo
echo "p95 step latency: ${p95}ms (limit ${P95_LIMIT_MS}ms)"
if [[ "$verdict" == "PASS" ]]; then
echo "GATES PASS"
return 0
fi
echo "GATES FAIL"
return 1
}
# ---- offline self-test -----------------------------------------------------
# Exercises every analyzer against canned passing/failing run directories under
# testdata/ so the parsing logic can be checked without a device.
self_test() {
local testdata="${script_directory}/testdata"
local failures=0
# The self-test fixtures (g5-slow-p95 = 4000ms) were calibrated against the
# 2500ms ceiling, so pin it here. Without this the backend-dependent default
# (5500ms under BACKEND=android) would rate the slow fixture as a PASS and the
# offline, device-free analyzer check would fail purely from an env var.
local P95_LIMIT_MS=2500
assert() {
local label="$1" expected="$2" actual="$3"
if [[ "$expected" == "$actual" ]]; then
printf 'ok %s\n' "$label"
else
printf 'FAIL %s (expected %s, got %s)\n' "$label" "$expected" "$actual"
failures=$((failures + 1))
fi
}
assert "G1 pass run exits zero" PASS \
"$(gate_exit_zero "${testdata}/pass" && echo PASS || echo FAIL)"
assert "G1 fail run nonzero exit" FAIL \
"$(gate_exit_zero "${testdata}/g1-nonzero-exit" && echo PASS || echo FAIL)"
assert "G2 pass run no errors" PASS \
"$(gate_no_error_lines "${testdata}/pass" && echo PASS || echo FAIL)"
assert "G2 benign noise tolerated" PASS \
"$(gate_no_error_lines "${testdata}/g2-benign-only" && echo PASS || echo FAIL)"
assert "G2 real error caught" FAIL \
"$(gate_no_error_lines "${testdata}/g2-real-error" && echo PASS || echo FAIL)"
assert "G3 pass clear state" PASS \
"$(gate_clear_state "${testdata}/pass" && echo PASS || echo FAIL)"
assert "G3 dirty fields caught" FAIL \
"$(gate_clear_state "${testdata}/g3-dirty-field" && echo PASS || echo FAIL)"
assert "G4 pass no doubling" PASS \
"$(gate_no_doubled_text "${testdata}/pass" && echo PASS || echo FAIL)"
assert "G4 doubled text caught" FAIL \
"$(gate_no_doubled_text "${testdata}/g4-doubled-text" && echo PASS || echo FAIL)"
local pass_p95 slow_p95
pass_p95="$(emit_step_latencies "${testdata}/pass/trace.jsonl" | p95_of)"
slow_p95="$(emit_step_latencies "${testdata}/g5-slow-p95/trace.jsonl" | p95_of)"
assert "G5 pass p95 under limit" PASS \
"$([[ "$pass_p95" -lt "$P95_LIMIT_MS" ]] && echo PASS || echo FAIL)"
assert "G5 slow p95 over limit" FAIL \
"$([[ "$slow_p95" -lt "$P95_LIMIT_MS" ]] && echo PASS || echo FAIL)"
echo
if [[ "$failures" -eq 0 ]]; then
echo "SELF-TEST PASS"
return 0
fi
echo "SELF-TEST FAIL (${failures} failed)"
return 1
}
main() {
case "${1:-}" in
--self-test) self_test ;;
"") run_gates ;;
*) echo "usage: $0 [--self-test]" >&2; exit 2 ;;
esac
}
main "$@"