mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 11:07:10 +00:00
* feat(sidecar): reach USB devices via the adb server by serial * feat(test): add --device flag to target a specific Android device by serial * feat(folio): select Android device via ANDROID_DEVICE in justfile * feat(conformance): add android backend to the gate suite * feat(android): keep device awake and unlocked so the app stays foreground * feat(conformance): prep physical android device (autofill/verifier/stayon) * fix(android): make device prep best-effort so OEM-blocked commands don't abort the run * fix(verifier): require positive bounds for swipe candidates A zero-bounds element centers at (0,0); a downward swipe from the top-left corner is the system gesture that pulls down the notification shade, dragging the fuzzer out of the app. Swipes now require positive bounds like every other verb. * fix(runner): harden app-scope guard against launcher and overlays The per-step guard now relaunches and waits until the app window is actually drawn before proceeding, so a slow physical-device relaunch no longer lets an observe or action land on the launcher. It also detects a system overlay (notification shade) stealing window focus while the app stays resumed, and dismisses it with back. * feat(android): harden physical-device runs in device prep Device prep now disables the AOSP cached-app freezer, phantom-process killer, and Doze (and exempts the driver) so OEM background management stops suspending the driver mid-run. Adds ReinstallApp for clear-state on ROMs that deny pm clear, and teaches focus detection to report the notification shade as systemui so the scope guard can dismiss it. * feat(driver): clear-state via APK reinstall when pm clear is blocked When an APK path is set, Android clear-state resets the app by uninstalling and reinstalling instead of asking the sidecar to pm clear, which hardened OEM builds (ColorOS) deny even to the adb shell user. Falls back to the sidecar clear path when no APK path is provided. * feat(cli): add --android-app-path for clear-state reinstall Wires the APK path from the test command through to the sidecar client so Android clear-state can reset apps on OEM builds that deny pm clear. * chore(folio): pass --android-app-path in just test * fix(runner): clamp swipe/scroll origin out of edge gesture zones A gesture starting in the top status-bar strip pulls down the notification shade; the bottom and side strips are the home and back gestures. Any of them drags the fuzzer out of the app. Swipe and scroll origins are now clamped into a safe inner area sized from the maximum element extent (the Android hierarchy root reports zero bounds, so the extent is the reliable screen size). Calibrated on device: origins below ~7% of height no longer open the shade. * perf(sidecar): faster Android text input and drop redundant settle poll inputText now uses adb `input text` for short shell-safe ASCII (~5x faster than the driver's per-character path) and falls back to the driver for unicode, injection payloads, and overflow-length strings. waitForIdle drops the structural-hash poll that followed waitForAppToSettle: each hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per mutating step for marginal benefit, and the runner already re-fetches transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4 still pass. * fix(verifier): exclude soft-keyboard region from action candidates The fuzzer was tapping Gboard's "Settings" key, navigating out of the app. That key is a bare FrameLayout with a content-desc and no package or resource-id, so the package-based scope filter missed it. Candidates whose center falls in the keyboard region (derived from the IME elements' bounds) are now dropped, so no tap or long-press lands on a key. Opt-in with app scoping; unscoped runs keep every node. * perf(runner): replace focus-tap settle with a brief wait The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText step on a physical device while the keyboard animated in. The tap registers focus immediately and text is injected into the focused view, so a short fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay green. * chore(conformance): platform-aware G5 p95 budget for android The 2500ms ceiling was calibrated on the iOS simulator. A physical Android device drives every step over USB (snapshot + settle + adb round-trips), so its per-step floor is several times higher; holding it to 2500ms would force removing the settle/retry logic the correctness gates depend on. The android backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500. * fix(sidecar): retry maestro android driver startup The maestro Android driver's dadb.open() occasionally misses its startup deadline (its instrumentation host is slow to come up right after a reboot or per-run reinstall), which aborted the whole run. Retry the open a few times with a short backoff so a transient timeout recovers. * chore(conformance): widen android G5 budget to 5500ms Physical-device p95 swung 3209-4612ms across sessions (cold runs right after a reboot are slower). 4500ms was too tight for that jitter; 5500ms covers the observed ceiling with headroom. * web replay fix * feat(android): force 3-button nav during runs to prevent app drift On gesture navigation a fuzzer swipe can trigger swipe-up-home or edge-back and fling the app off screen. Device-prep now switches to 3-button navigation for the run (no edge gestures; the nav bar's buttons are systemui-owned and already excluded from action candidates) and restores the original navigation mode when the run ends. Best effort: leaves nav untouched if the overlay command is unavailable. * fix(android): target the selected device in adb reads; don't strand nav mode Review fixes: - ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the foreground/scope guard works when several devices are attached (the --device path). Previously they ran bare `adb shell`, which errors with multiple devices, silently disabling app-scope enforcement. The sidecar client passes its serial through. - Extract an adbArgs helper and route every adb call through it, removing four duplicated serial-arg builders. - ForceThreeButtonNav now decides what to restore before changing anything: if the current mode is unknown or already 3-button it leaves nav untouched, instead of switching and then stranding the device in 3-button. Logic split into the pure navModeToRestore, now unit tested. * fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was stranded above screenBounds by an insertion). Extend the clamp test to assert an off-screen destination is clamped onto the screen and that the origin lands exactly on the margin. * test(verifier): cover keyboardRegionTop, including the decor-view guard The full-screen IME decor view rejection had no test; removing it left the suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in favor of the real keyboard line, and decor-only -> sentinel. * style(cli): gofmt testOptions field alignment * fix(sidecar): keep a leading dash off the fast input path A value starting with '-' could be read as an option by `adb input text`, so the fast-path regex now requires a non-dash first character; such values fall back to the driver. Also cover the dadb-target branch where a colon precedes a non-numeric port (a USB serial, not host:port). * refactor(verifier): scope action candidates by window ownership Replaces the leaky per-element package check and the keyboard-region Y heuristic with one rule: walk the window tree propagating each node's owning package (empty and the neutral android framework package are transparent); a node is in scope only when no concrete foreign package owns it (the app's own window carries no package on Compose apps) or the owner is the app package. This drops whole foreign windows (soft keyboard, system UI, launcher) AND their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which the old empty-package-is-in-scope rule admitted and which navigated out of the app. Deletes keyboardRegionTop/isInputMethodElement. * fix(runner): re-check foreground at apply time, skip stale actions ensureForeground runs before observe, but the app can leave between observe and apply (a prior gesture settling late); swipes/keys then fire stale coordinates onto whatever screen is now up. Re-check foreground immediately before applying and, when the app is gone, skip the action and log it (making the escape visible) so the next step's guard relaunches instead. * fix(android): type long ASCII via fast guarded path to stop keystroke escape A 4096-char corpus string exceeded the fast input cap and fell to the per-character driver path, which takes ~120s. During that uninterruptible window focus could leave the app and the remaining keystrokes sprayed into the launcher search box. Route shell-safe ASCII of any length through adb input text, chunked, re-checking the foreground app between chunks and stopping if it changed. * chore: ignore gate artifacts and local scratch files * refactor(runner): narrow gesture clamp to the top shade strip 3-button nav (forced for every run) disables the side back and bottom home gestures at the OS level. On-device probing confirmed side and bottom swipe origins no longer drift, leaving the notification shade as the only edge gesture a swipe can trigger. Clamp only the top strip; keep origin and destination on screen otherwise. * chore(format): add .editorconfig enforcing 80-column limit * chore(format): add prettier config with 80-char printWidth * chore(deps): add prettier devDependency to replay-ui * chore(deps): add prettier devDependency to folio-web * chore(deps): add prettier devDependency to spec package * chore(format): add swift-format config with 80-char lineLength * feat(format): add make fmt targets for per-language 80-col formatting * fix(runner): translate gesture to safe area so near-top scrolls keep direction Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp. * fix(runner): apply-time guard consults focused window, not just resumed activity ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay. * test(runner): cover apply-time foreground skip and appIsForeground table Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised. * fix(sidecar): harden android driver open, input guard, pressKey, foreground marker - openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF). - pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract. - the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner. - foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard. * fix(conformance): pin self-test p95 budget and score install failures as run failures self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue. * fix(android): require --device when several devices are connected With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous. * refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test. * perf(verifier): memoize scopedElements per tree scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot. * fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail. * test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup. * refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true. * perf(sidecar): reuse a single Jackson ObjectMapper structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val. * refactor(android): remove unused AdbReverse/AdbReverseRemove No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed. * style(runner): trim non-load-bearing comments from this PR's runner code and tests * style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
476 lines
19 KiB
Bash
Executable File
476 lines
19 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Scripted conformance gate for the mobile drivers. Runs five serial,
|
|
# non-overlapping 3-minute fuzz runs against the folio example app
|
|
# (examples/folio) and scores five gates (G1..G5) over the captured traces and
|
|
# output. Exits non-zero if any gate fails.
|
|
#
|
|
# Backends:
|
|
# BACKEND=simulator (default) drive the booted iOS simulator
|
|
# BACKEND=device drive an attached physical iPhone via the
|
|
# driver's runner-only device path; select it
|
|
# with IOS_DEVICE="<name>" (passed as --ios-device)
|
|
# BACKEND=android drive an Android device/emulator over the JVM
|
|
# sidecar; select a specific device with
|
|
# ANDROID_DEVICE="<adb serial>" (passed as --device)
|
|
#
|
|
# Usage:
|
|
# ./gates.sh run the simulator gates
|
|
# BACKEND=device IOS_DEVICE="iPhone" ./gates.sh
|
|
# BACKEND=android ANDROID_DEVICE="663c91b1" ./gates.sh
|
|
# ./gates.sh --self-test run the offline analyzer tests only
|
|
#
|
|
# Tunables (environment):
|
|
# RUNS=5 number of serial runs
|
|
# DURATION=3m per-run fuzz duration
|
|
# SEED=0 fuzz seed
|
|
# P95_LIMIT_MS G5 p95 step-latency ceiling in ms (default 2500 for iOS,
|
|
# 5500 for the android backend's higher and more variable
|
|
# per-step USB cost, especially cold right after a reboot)
|
|
# SANDERLING=sanderling binary to invoke
|
|
|
|
set -euo pipefail
|
|
|
|
script_directory="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
folio_directory="$(cd "${script_directory}/../examples/folio" && pwd)"
|
|
|
|
BACKEND="${BACKEND:-simulator}"
|
|
RUNS="${RUNS:-5}"
|
|
DURATION="${DURATION:-3m}"
|
|
SEED="${SEED:-0}"
|
|
# The p95 ceiling is backend-specific: a physical Android device drives every
|
|
# step over USB (snapshot + settle + adb round-trips), so its per-step floor is
|
|
# several times the iOS simulator's in-process cost. 2500ms was calibrated on
|
|
# the simulator; holding a physical device to it would force ripping out the
|
|
# settle/retry logic the correctness gates depend on. Override with P95_LIMIT_MS.
|
|
if [[ "$BACKEND" == "android" ]]; then
|
|
P95_LIMIT_MS="${P95_LIMIT_MS:-5500}"
|
|
else
|
|
P95_LIMIT_MS="${P95_LIMIT_MS:-2500}"
|
|
fi
|
|
SANDERLING="${SANDERLING:-sanderling}"
|
|
IOS_DEVICE="${IOS_DEVICE:-iPhone 17 Pro}"
|
|
ANDROID_DEVICE="${ANDROID_DEVICE:-}"
|
|
|
|
bundle_id="app.folio"
|
|
spec_path="${folio_directory}/sanderling/spec.ts"
|
|
android_apk="${folio_directory}/app/androidApp/build/outputs/apk/debug/androidApp-debug.apk"
|
|
# The built app bundle differs by SDK: the simulator build lands under
|
|
# Debug-iphonesimulator, the device build under Debug-iphoneos.
|
|
if [[ "$BACKEND" == "device" ]]; then
|
|
ios_app="${folio_directory}/app/iosApp/build/Build/Products/Debug-iphoneos/iosApp.app"
|
|
else
|
|
ios_app="${folio_directory}/app/iosApp/build/Build/Products/Debug-iphonesimulator/iosApp.app"
|
|
fi
|
|
|
|
# Run adb against the selected Android device, or the only one if unset.
|
|
adb_target() {
|
|
if [[ -n "$ANDROID_DEVICE" ]]; then adb -s "$ANDROID_DEVICE" "$@"; else adb "$@"; fi
|
|
}
|
|
|
|
# The companion binary, embedded for simulator runs. Referenced by file name
|
|
# only for the orphan-process check; prose elsewhere says "the companion".
|
|
companion_process_name="idb_companion"
|
|
|
|
# Known-benign stderr lines the companion always prints. These are matched as
|
|
# fixed substrings and excluded from the G2 ERROR scan. Keep this list tight:
|
|
# only lines that are provably harmless and emitted on every healthy run.
|
|
benign_stderr_substrings=(
|
|
# The dynamic linker reports the same Objective-C class registered by two
|
|
# loaded images. The companion runs fine; this is cosmetic. The real line
|
|
# reads "objc[<pid>]: Class ...", so the fixed substring is "]: Class".
|
|
"]: Class"
|
|
"is implemented in both"
|
|
"One of the two will be used. Which one is undefined."
|
|
# gRPC and absl emit informational banner lines on startup.
|
|
"WARNING: All log messages before absl::InitializeLog()"
|
|
)
|
|
|
|
# ---- gate analyzers (pure, operate on a single run directory) --------------
|
|
|
|
# G2 helper: strip benign companion noise, then report sanderling ERROR lines.
|
|
# sanderling's progress logger renders error-level records as lines beginning
|
|
# "error:" (see internal/testrun/progress.go). We also catch the upper-case
|
|
# ERROR token (word-bounded, so ERRORS/ERRORLESS and path fragments do not
|
|
# trip the gate) for safety against future handlers.
|
|
error_lines_in() {
|
|
local output_file="$1"
|
|
local filtered
|
|
filtered="$(cat "$output_file")"
|
|
local pattern
|
|
for pattern in "${benign_stderr_substrings[@]}"; do
|
|
filtered="$(printf '%s\n' "$filtered" | grep -vF "$pattern" || true)"
|
|
done
|
|
printf '%s\n' "$filtered" | grep -E '(^error:|\bERROR\b)' || true
|
|
}
|
|
|
|
# G1: process exit status recorded by the runner loop.
|
|
gate_exit_zero() {
|
|
local run_directory="$1"
|
|
[[ "$(cat "${run_directory}/exit_status")" == "0" ]]
|
|
}
|
|
|
|
# G2: no sanderling ERROR lines after filtering benign companion noise.
|
|
gate_no_error_lines() {
|
|
local run_directory="$1"
|
|
local found
|
|
found="$(error_lines_in "${run_directory}/output.log")"
|
|
[[ -z "$found" ]]
|
|
}
|
|
|
|
# G3: the first hierarchy snapshot shows the login screen with empty email and
|
|
# password fields (clear-state proof). Reads the first trace line carrying a
|
|
# hierarchy and asserts LoginEmail/LoginPassword carry no text.
|
|
gate_clear_state() {
|
|
local run_directory="$1"
|
|
local trace_file="${run_directory}/trace.jsonl"
|
|
[[ -f "$trace_file" ]] || return 1
|
|
local verdict
|
|
verdict="$(jq -s -r '
|
|
[ .[] | select(.hierarchy != null) ] as $withHierarchy
|
|
| if ($withHierarchy | length) == 0 then "fail:no-hierarchy"
|
|
else ($withHierarchy[0].hierarchy.elements // []) as $elements
|
|
| ($elements | map(select(.resourceId == "LoginEmail")) | first) as $email
|
|
| ($elements | map(select(.resourceId == "LoginPassword")) | first) as $password
|
|
| if $email == null or $password == null then "fail:no-login-fields"
|
|
elif (($email.attrs.text // $email.text // "") != "") then "fail:email-not-empty"
|
|
elif (($password.attrs.text // $password.text // "") != "") then "fail:password-not-empty"
|
|
else "pass" end
|
|
end
|
|
' "$trace_file")"
|
|
[[ "$verdict" == "pass" ]]
|
|
}
|
|
|
|
# G4: no doubled text after InputText. For each InputText action targeting a
|
|
# field, the field's value in the NEXT hierarchy must not contain the input
|
|
# concatenated with itself (catches append-vs-replace and double-paste bugs).
|
|
# The action chosen at step N is applied before step N+1 is observed, so the
|
|
# effect lands in the following snapshot.
|
|
gate_no_doubled_text() {
|
|
local run_directory="$1"
|
|
local trace_file="${run_directory}/trace.jsonl"
|
|
[[ -f "$trace_file" ]] || return 1
|
|
local verdict
|
|
verdict="$(jq -s -r '
|
|
# Map a selector like "testTag:LoginScreen > testTag:LoginEmail" to its
|
|
# target field id: the token after the final ":" of the last segment.
|
|
def target_field(selector):
|
|
(selector | split(">") | last | gsub("^\\s+|\\s+$";"")) as $last
|
|
| ($last | split(":") | last);
|
|
|
|
[ .[] | select(.hierarchy != null) ] as $steps
|
|
| reduce range(0; ($steps | length)) as $i ([];
|
|
($steps[$i]) as $current
|
|
| (if $i + 1 < ($steps | length) then $steps[$i + 1] else null end) as $next
|
|
| if ($current.next_action.kind == "InputText") and ($next != null)
|
|
then
|
|
(target_field($current.next_action.selector // "")) as $field
|
|
| ($current.next_action.text // "") as $typed
|
|
| (($next.hierarchy.elements // [])
|
|
| map(select(.resourceId == $field)) | first) as $element
|
|
| if $element != null and $typed != ""
|
|
then
|
|
(($element.attrs.text // $element.text // "")) as $value
|
|
| if ($value | contains($typed + $typed))
|
|
then . + [{field: $field, typed: $typed, value: $value}]
|
|
else . end
|
|
else . end
|
|
else . end)
|
|
| if length == 0 then "pass"
|
|
else "fail:" + (.[0].field) + ":" + (.[0].value) end
|
|
' "$trace_file")"
|
|
[[ "$verdict" == "pass" ]]
|
|
}
|
|
|
|
# Emit one step-latency sample per consecutive trace-step pair, in
|
|
# milliseconds, on stdout. The trace records a wall-clock timestamp per step;
|
|
# the latency of a step is the gap to the next step's observation. Used by G5,
|
|
# which aggregates samples across all runs before computing the p95. Timestamps
|
|
# are RFC3339 with fractional seconds and a numeric offset, so they are parsed
|
|
# with python3 rather than jq's UTC-only fromdateiso8601.
|
|
emit_step_latencies() {
|
|
local trace_file="$1"
|
|
[[ -f "$trace_file" ]] || return 0
|
|
jq -r 'select(.timestamp != null) | .timestamp' "$trace_file" \
|
|
| python3 -c '
|
|
import sys
|
|
from datetime import datetime
|
|
stamps = [datetime.fromisoformat(line.strip()) for line in sys.stdin if line.strip()]
|
|
for earlier, later in zip(stamps, stamps[1:]):
|
|
print(int((later - earlier).total_seconds() * 1000))
|
|
'
|
|
}
|
|
|
|
# Compute the p95 (nearest-rank) of the latency samples on stdin, in ms.
|
|
p95_of() {
|
|
python3 -c '
|
|
import sys, math
|
|
samples = sorted(int(float(line)) for line in sys.stdin if line.strip())
|
|
if not samples:
|
|
print(0)
|
|
sys.exit(0)
|
|
rank = max(1, math.ceil(0.95 * len(samples)))
|
|
print(samples[rank - 1])
|
|
'
|
|
}
|
|
|
|
# G5 orphan check: report any lingering companion, runner session (the hybrid
|
|
# simulator driver hosts an in-simulator runner), and, on the device backend,
|
|
# the device runner session. The usbmux tunnel is an in-process forwarder that
|
|
# dies with sanderling, so it leaves no process to check. Empty output is clean.
|
|
orphan_processes() {
|
|
local found=""
|
|
if [[ "$BACKEND" == "android" ]]; then
|
|
# sanderling SIGTERMs the JVM sidecar on shutdown; a survivor is an orphan.
|
|
if pgrep -f "sanderling-sidecar.*\.jar" >/dev/null 2>&1; then found+="sidecar "; fi
|
|
printf '%s' "$found"
|
|
return
|
|
fi
|
|
if pgrep -f "$companion_process_name" >/dev/null 2>&1; then
|
|
found+="companion "
|
|
fi
|
|
if pgrep -f "sanderling-runner.*xctestrun" >/dev/null 2>&1; then
|
|
found+="runner-session "
|
|
fi
|
|
if pgrep -f "CompanionRunnerUITests-Runner" >/dev/null 2>&1; then
|
|
found+="runner-app "
|
|
fi
|
|
if [[ "$BACKEND" == "device" ]]; then
|
|
# The device test session that hosts the runner. Its destination carries
|
|
# platform=iOS,id=<udid>.
|
|
if pgrep -f "xctestrun.*platform=iOS,id=" >/dev/null 2>&1; then
|
|
found+="device-session "
|
|
fi
|
|
fi
|
|
printf '%s' "$found"
|
|
}
|
|
|
|
# ---- run orchestration -----------------------------------------------------
|
|
|
|
invoke_sanderling() {
|
|
local output_directory="$1"
|
|
local output_log="$2"
|
|
local exit_status_file="$3"
|
|
|
|
local platform target_flags=()
|
|
if [[ "$BACKEND" == "android" ]]; then
|
|
# pm clear is blocked on some OEM ROMs, so a fresh install (which wipes
|
|
# /data/data) provides the clean clear-state start; --clear-data=false
|
|
# then skips the sidecar's pm clear. Mirrors the iOS per-run reinstall.
|
|
platform=android
|
|
adb_target uninstall "$bundle_id" >/dev/null 2>&1 || true
|
|
# A transient install hiccup must score this run as a failure, not abort the
|
|
# whole harness under `set -e` and discard the other runs' data.
|
|
if ! adb_target install "$android_apk" >"$output_log" 2>&1; then
|
|
echo "adb install failed for ${android_apk}; recording run as a failure" >>"$output_log"
|
|
printf '1' >"$exit_status_file"
|
|
return
|
|
fi
|
|
target_flags=(--clear-data=false)
|
|
[[ -n "$ANDROID_DEVICE" ]] && target_flags+=(--device "$ANDROID_DEVICE")
|
|
else
|
|
# The iOS backends pass --ios-app-path so each run reinstalls the current
|
|
# build for a clean start (device via devicectl, simulator via simctl).
|
|
platform=ios
|
|
target_flags=(--ios-device "$IOS_DEVICE" --ios-app-path "$ios_app")
|
|
fi
|
|
|
|
local status=0
|
|
"$SANDERLING" test \
|
|
--platform "$platform" \
|
|
--spec "$spec_path" \
|
|
--bundle-id "$bundle_id" \
|
|
"${target_flags[@]}" \
|
|
--duration "$DURATION" \
|
|
--seed "$SEED" \
|
|
--output "$output_directory" \
|
|
>"$output_log" 2>&1 || status=$?
|
|
printf '%s' "$status" >"$exit_status_file"
|
|
}
|
|
|
|
# Locate the run directory sanderling created under output_directory (it nests
|
|
# a timestamped subdirectory) and normalise its artifacts up one level.
|
|
collect_run_artifacts() {
|
|
local output_directory="$1"
|
|
[[ -f "${output_directory}/trace.jsonl" ]] && return 0
|
|
local produced
|
|
produced="$(find "$output_directory" -mindepth 2 -name 'trace.jsonl' 2>/dev/null | head -1 || true)"
|
|
if [[ -n "$produced" ]]; then
|
|
cp "$produced" "${output_directory}/trace.jsonl" 2>/dev/null || true
|
|
fi
|
|
}
|
|
|
|
run_gates() {
|
|
if [[ "$BACKEND" == "android" ]]; then
|
|
# Build only; invoke_sanderling reinstalls per run via adb (gradle's ddmlib
|
|
# install is flaky on some physical devices).
|
|
echo "preparing folio android build"
|
|
( cd "$folio_directory" && just build >/dev/null )
|
|
# A physical device, unlike an emulator, lets system UI steal the foreground
|
|
# from the app the fuzzer is exploring. Keep the screen on so it never
|
|
# re-locks, silence the autofill save-password prompt that pops over the
|
|
# login form, and stop Play Protect from intercepting the per-run reinstall.
|
|
# The device must already be unlocked (a secure lock cannot be opened here).
|
|
adb_target shell svc power stayon true >/dev/null 2>&1 || true
|
|
adb_target shell settings put secure autofill_service null >/dev/null 2>&1 || true
|
|
adb_target shell settings put global verifier_verify_adb_installs 0 >/dev/null 2>&1 || true
|
|
elif [[ "$BACKEND" == "simulator" ]]; then
|
|
echo "preparing folio build for the simulator backend"
|
|
( cd "$folio_directory" && just ios >/dev/null )
|
|
else
|
|
echo "preparing folio device build"
|
|
( cd "$folio_directory" && just ios-device >/dev/null )
|
|
fi
|
|
|
|
local timestamp
|
|
timestamp="$(date +%Y%m%d-%H%M%S)"
|
|
local gate_root="${script_directory}/runs/${timestamp}"
|
|
mkdir -p "$gate_root"
|
|
echo "gate root: ${gate_root}"
|
|
echo "backend=${BACKEND} runs=${RUNS} duration=${DURATION} seed=${SEED} p95_limit_ms=${P95_LIMIT_MS}"
|
|
|
|
local all_latencies="${gate_root}/all-latencies.txt"
|
|
: >"$all_latencies"
|
|
|
|
local -a g1 g2 g3 g4 g5
|
|
local run_index
|
|
for ((run_index = 1; run_index <= RUNS; run_index++)); do
|
|
local run_directory="${gate_root}/run-${run_index}"
|
|
mkdir -p "$run_directory"
|
|
echo "run ${run_index}/${RUNS} -> ${run_directory}"
|
|
|
|
invoke_sanderling "$run_directory" "${run_directory}/output.log" "${run_directory}/exit_status"
|
|
collect_run_artifacts "$run_directory"
|
|
|
|
g1[run_index]=$(gate_exit_zero "$run_directory" && echo PASS || echo FAIL)
|
|
g2[run_index]=$(gate_no_error_lines "$run_directory" && echo PASS || echo FAIL)
|
|
g3[run_index]=$(gate_clear_state "$run_directory" && echo PASS || echo FAIL)
|
|
g4[run_index]=$(gate_no_doubled_text "$run_directory" && echo PASS || echo FAIL)
|
|
|
|
emit_step_latencies "${run_directory}/trace.jsonl" >>"$all_latencies"
|
|
|
|
local orphans
|
|
orphans="$(orphan_processes)"
|
|
if [[ -n "$orphans" ]]; then
|
|
g5[run_index]="FAIL"
|
|
echo " orphaned processes after run ${run_index}: ${orphans}"
|
|
else
|
|
g5[run_index]="PENDING"
|
|
fi
|
|
done
|
|
|
|
local p95
|
|
p95="$(p95_of <"$all_latencies")"
|
|
local final_orphans
|
|
final_orphans="$(orphan_processes)"
|
|
|
|
# G5 is global: p95 is computed over every run's samples and an orphan after
|
|
# any single run fails the whole gate. Decide once, then stamp every row.
|
|
local g5_global="PASS"
|
|
if [[ "$p95" -ge "$P95_LIMIT_MS" ]]; then
|
|
g5_global="FAIL"
|
|
fi
|
|
if [[ -n "$final_orphans" ]]; then
|
|
g5_global="FAIL"
|
|
echo "orphaned processes at end: ${final_orphans}"
|
|
fi
|
|
for ((run_index = 1; run_index <= RUNS; run_index++)); do
|
|
if [[ "${g5[run_index]}" == "FAIL" ]]; then
|
|
g5_global="FAIL"
|
|
fi
|
|
done
|
|
for ((run_index = 1; run_index <= RUNS; run_index++)); do
|
|
g5[run_index]="$g5_global"
|
|
done
|
|
|
|
echo
|
|
printf 'run G1 G2 G3 G4 G5\n'
|
|
local verdict="PASS"
|
|
for ((run_index = 1; run_index <= RUNS; run_index++)); do
|
|
printf '%-4s %-5s %-5s %-5s %-5s %-5s\n' \
|
|
"$run_index" "${g1[run_index]}" "${g2[run_index]}" "${g3[run_index]}" "${g4[run_index]}" "${g5[run_index]}"
|
|
for cell in "${g1[run_index]}" "${g2[run_index]}" "${g3[run_index]}" "${g4[run_index]}" "${g5[run_index]}"; do
|
|
[[ "$cell" == "PASS" ]] || verdict="FAIL"
|
|
done
|
|
done
|
|
echo
|
|
echo "p95 step latency: ${p95}ms (limit ${P95_LIMIT_MS}ms)"
|
|
if [[ "$verdict" == "PASS" ]]; then
|
|
echo "GATES PASS"
|
|
return 0
|
|
fi
|
|
echo "GATES FAIL"
|
|
return 1
|
|
}
|
|
|
|
# ---- offline self-test -----------------------------------------------------
|
|
# Exercises every analyzer against canned passing/failing run directories under
|
|
# testdata/ so the parsing logic can be checked without a device.
|
|
|
|
self_test() {
|
|
local testdata="${script_directory}/testdata"
|
|
local failures=0
|
|
# The self-test fixtures (g5-slow-p95 = 4000ms) were calibrated against the
|
|
# 2500ms ceiling, so pin it here. Without this the backend-dependent default
|
|
# (5500ms under BACKEND=android) would rate the slow fixture as a PASS and the
|
|
# offline, device-free analyzer check would fail purely from an env var.
|
|
local P95_LIMIT_MS=2500
|
|
|
|
assert() {
|
|
local label="$1" expected="$2" actual="$3"
|
|
if [[ "$expected" == "$actual" ]]; then
|
|
printf 'ok %s\n' "$label"
|
|
else
|
|
printf 'FAIL %s (expected %s, got %s)\n' "$label" "$expected" "$actual"
|
|
failures=$((failures + 1))
|
|
fi
|
|
}
|
|
|
|
assert "G1 pass run exits zero" PASS \
|
|
"$(gate_exit_zero "${testdata}/pass" && echo PASS || echo FAIL)"
|
|
assert "G1 fail run nonzero exit" FAIL \
|
|
"$(gate_exit_zero "${testdata}/g1-nonzero-exit" && echo PASS || echo FAIL)"
|
|
|
|
assert "G2 pass run no errors" PASS \
|
|
"$(gate_no_error_lines "${testdata}/pass" && echo PASS || echo FAIL)"
|
|
assert "G2 benign noise tolerated" PASS \
|
|
"$(gate_no_error_lines "${testdata}/g2-benign-only" && echo PASS || echo FAIL)"
|
|
assert "G2 real error caught" FAIL \
|
|
"$(gate_no_error_lines "${testdata}/g2-real-error" && echo PASS || echo FAIL)"
|
|
|
|
assert "G3 pass clear state" PASS \
|
|
"$(gate_clear_state "${testdata}/pass" && echo PASS || echo FAIL)"
|
|
assert "G3 dirty fields caught" FAIL \
|
|
"$(gate_clear_state "${testdata}/g3-dirty-field" && echo PASS || echo FAIL)"
|
|
|
|
assert "G4 pass no doubling" PASS \
|
|
"$(gate_no_doubled_text "${testdata}/pass" && echo PASS || echo FAIL)"
|
|
assert "G4 doubled text caught" FAIL \
|
|
"$(gate_no_doubled_text "${testdata}/g4-doubled-text" && echo PASS || echo FAIL)"
|
|
|
|
local pass_p95 slow_p95
|
|
pass_p95="$(emit_step_latencies "${testdata}/pass/trace.jsonl" | p95_of)"
|
|
slow_p95="$(emit_step_latencies "${testdata}/g5-slow-p95/trace.jsonl" | p95_of)"
|
|
assert "G5 pass p95 under limit" PASS \
|
|
"$([[ "$pass_p95" -lt "$P95_LIMIT_MS" ]] && echo PASS || echo FAIL)"
|
|
assert "G5 slow p95 over limit" FAIL \
|
|
"$([[ "$slow_p95" -lt "$P95_LIMIT_MS" ]] && echo PASS || echo FAIL)"
|
|
|
|
echo
|
|
if [[ "$failures" -eq 0 ]]; then
|
|
echo "SELF-TEST PASS"
|
|
return 0
|
|
fi
|
|
echo "SELF-TEST FAIL (${failures} failed)"
|
|
return 1
|
|
}
|
|
|
|
main() {
|
|
case "${1:-}" in
|
|
--self-test) self_test ;;
|
|
"") run_gates ;;
|
|
*) echo "usage: $0 [--self-test]" >&2; exit 2 ;;
|
|
esac
|
|
}
|
|
|
|
main "$@"
|