Commit Graph
8 Commits
Author SHA1 Message Date
pj f4b9ff1468 fix(conformance): g4 skips an input that names no field
jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.
2026-08-18 20:14:28 +05:30
pj faa1a33a68 test(conformance): g4 keeps checking past an input typed at coordinates
An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.
2026-08-18 20:14:14 +05:30
pj 8414953135 fix(conformance): g4 reads a doubling off the observed field value
The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.
2026-08-18 20:13:52 +05:30
pj f631415cb6 test(conformance): the g4 fixture holds what a redacted android trace holds
Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.
2026-08-18 20:07:39 +05:30
pj 76dce1a75e experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs

runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record arm membership and host in meta.json

meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --arm and populate run meta from it

Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): sweep seeds for one experiment cell

campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.

Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): no silent generator fallback, and llm on web

--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.

pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): make the hierarchy dump agree with the web runtime

Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.

scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.

The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(conformance): give the gate reproducible seeds

SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.

The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit editable as a plain boolean

editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): leave the head subtree out of the web target walk

collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(chrome): compare the facts both hosts derive from one DOM

The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.

This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* chore(make): run the browser packages one at a time

Both launch Chrome and launching two at once has failed with "Launch: context
canceled".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style: remove every em-dash and en-dash

Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): honor the caller context in Launch

Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.

The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(sidecarassets): publish the extracted jar through a rename

Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): kill a run that outlives --run-timeout

A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style(test): gofmt browser_test.go

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:20:31 +05:30
pj 6b0d6cb971 WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial

* feat(test): add --device flag to target a specific Android device by serial

* feat(folio): select Android device via ANDROID_DEVICE in justfile

* feat(conformance): add android backend to the gate suite

* feat(android): keep device awake and unlocked so the app stays foreground

* feat(conformance): prep physical android device (autofill/verifier/stayon)

* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run

* fix(verifier): require positive bounds for swipe candidates

A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.

* fix(runner): harden app-scope guard against launcher and overlays

The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.

* feat(android): harden physical-device runs in device prep

Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.

* feat(driver): clear-state via APK reinstall when pm clear is blocked

When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.

* feat(cli): add --android-app-path for clear-state reinstall

Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.

* chore(folio): pass --android-app-path in just test

* fix(runner): clamp swipe/scroll origin out of edge gesture zones

A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.

* perf(sidecar): faster Android text input and drop redundant settle poll

inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.

* fix(verifier): exclude soft-keyboard region from action candidates

The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.

* perf(runner): replace focus-tap settle with a brief wait

The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.

* chore(conformance): platform-aware G5 p95 budget for android

The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.

* fix(sidecar): retry maestro android driver startup

The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.

* chore(conformance): widen android G5 budget to 5500ms

Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.

* web replay fix

* feat(android): force 3-button nav during runs to prevent app drift

On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.

* fix(android): target the selected device in adb reads; don't strand nav mode

Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
  foreground/scope guard works when several devices are attached (the --device
  path). Previously they ran bare `adb shell`, which errors with multiple
  devices, silently disabling app-scope enforcement. The sidecar client passes
  its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
  duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
  the current mode is unknown or already 3-button it leaves nav untouched,
  instead of switching and then stranding the device in 3-button. Logic split
  into the pure navModeToRestore, now unit tested.

* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds

Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.

* test(verifier): cover keyboardRegionTop, including the decor-view guard

The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.

* style(cli): gofmt testOptions field alignment

* fix(sidecar): keep a leading dash off the fast input path

A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).

* refactor(verifier): scope action candidates by window ownership

Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.

This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.

* fix(runner): re-check foreground at apply time, skip stale actions

ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.

* fix(android): type long ASCII via fast guarded path to stop keystroke escape

A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.

* chore: ignore gate artifacts and local scratch files

* refactor(runner): narrow gesture clamp to the top shade strip

3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.

* chore(format): add .editorconfig enforcing 80-column limit

* chore(format): add prettier config with 80-char printWidth

* chore(deps): add prettier devDependency to replay-ui

* chore(deps): add prettier devDependency to folio-web

* chore(deps): add prettier devDependency to spec package

* chore(format): add swift-format config with 80-char lineLength

* feat(format): add make fmt targets for per-language 80-col formatting

* fix(runner): translate gesture to safe area so near-top scrolls keep direction

Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.

* fix(runner): apply-time guard consults focused window, not just resumed activity

ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.

* test(runner): cover apply-time foreground skip and appIsForeground table

Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.

* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker

- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.

* fix(conformance): pin self-test p95 budget and score install failures as run failures

self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.

* fix(android): require --device when several devices are connected

With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.

* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands

The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.

* perf(verifier): memoize scopedElements per tree

scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.

* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear

Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.

* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms

The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.

* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty

bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.

* perf(sidecar): reuse a single Jackson ObjectMapper

structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.

* refactor(android): remove unused AdbReverse/AdbReverseRemove

No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.

* style(runner): trim non-load-bearing comments from this PR's runner code and tests

* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
2026-06-11 10:10:05 +05:30
pj 90224dfd06 Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers

The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.

* feat(ios): resolve physical devices from devicectl

ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.

* feat(ioscompanion): runner-only device driver mode

NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.

* test(ioscompanion): cover device wiring, routing, and shell-out argv

Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.

* feat(testrun): route physical-device iOS runs to the device driver

Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.

* feat(doctor): device prereqs replace java/sidecar for ios-device

iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.

* feat(conformance): device backend uses iphoneos app and tunnel orphan checks

The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).

* feat(folio): device build linking the iosArm64 framework

project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.

* docs(cli): document ios-device doctor checks and the device flags

The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.

* fix(ioscompanion): resolve signing key path to absolute

xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.

* fix(ioscompanion): re-enable signing for the device runner build

companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.

* fix(ioscompanion): key the device build cache on signing identity

The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.

* docs(getting-started): document physical iOS device setup

Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.

* feat(ios): native usbmux client and in-process tunnel forwarder

Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.

* refactor(ios): drive device tunnel via io.Closer seam

Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.

* refactor(ios): remove iproxy spawn from device runner

* test(ios): cover tunnel close via io.Closer not child process

* feat(doctor): check usbmuxd socket instead of iproxy on PATH

* chore(conformance): drop iproxy orphan check; tunnel is in-process

* docs(ios): device tunnel uses native usbmux, nothing to install

* chore: gitignore the signing keys directory

* feat(folio): add Android launcher icon (black bg, white dot)

* feat(folio): add iOS app icon (black bg, white dot)

* feat(folio): add web favicon (black bg, white dot)

* docs(ioscompanion): fix stale const comments

* refactor(ioscompanion): inline single-use devicectl argv builders

* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests

* refactor(ioscompanion): inline firstNonEmpty

* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam

* test(doctor): trim redundant signing-check test

* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env

* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
2026-06-09 18:38:52 +05:30
pj 406b7516b3 iOS simulator driver: Go-native companion-backed backend (#62)
* perf(ios): use prebuilt XCTest runner to cut startup

* chore(ioscompanion): add companion asset prepare script

* feat(ioscompanion): embed and extract simulator companion bundle

* test(ioscompanion): cover companion stub and embedded extraction

* docs: add third party notices for vendored companion

* chore: ignore vendored companion bundle artifact

* build(proto): pin simulator companion proto v1.1.8

* build(proto): add dedicated buf module and gen template for pinned proto

* build(proto): exclude pinned companion proto from root buf workspace

* feat(ioscompanion): commit generated companion gRPC stubs

* feat(ioscompanion): map flat companion describe dump to TreeNode JSON

* test(ioscompanion): add hierarchy-map golden and unit tests

* feat(ioscompanion): port screen-settle stability polling to Go

* test(ioscompanion): cover settle transitional, hash, streak, and cap rules

* feat(ioscompanion): add USB HID keymap module

* test(ioscompanion): cover keymap branches and paste-chord constants

* build: embed companion assets via withcompanion tag

* feat(ioscompanion): add transport companion interface

* feat(ioscompanion): add HID event wrapper and builders

* feat(ioscompanion): wire gRPC companion client and Dial

* test(ioscompanion): cover HID builders and unit conversions

* test(ioscompanion): cover Dial, process-state mapping, and install archive

* test(ioscompanion): add gated simulator integration smoke test

* feat(ioscompanion): text input and gesture HID composition with pasteboard fallback

* test(ioscompanion): cover input composers, paste dialog loop, and pure helpers

* feat(ioscompanion): add Describe to companion transport

* feat(ioscompanion): implement DeviceDriver with companion supervision

* test(ioscompanion): unit tests with fake companion transport

* test(ioscompanion): gated companion smoke test

* feat(ios): add ResolveTarget for simulator vs physical-device routing

* feat(testrun): route iOS simulators through the native companion driver

* refactor(testrun): defer the java preflight check to the physical-device path

* feat(cli): add --ios-app-path flag

* feat(doctor): split iOS checks into simulator and physical-device paths

* test(folio): add gate-analyzer fixtures for G1-G5

* feat(folio): add iOS conformance gate script

* chore(folio): wire gates recipe, app path, and ignore gate output

* style: gofmt struct alignment drift

* fix(doctor): probe simctl via xcrun instead of PATH lookup

* fix(ioscompanion): spawn companion under driver-lifetime context

* test(ioscompanion): prove companion child outlives startup context

* fix(ioscompanion): chunk install payload under companion message cap

* test(ioscompanion): cover install payload chunking

* fix(ioscompanion): reinstall via simctl and sanitize companion env

* fix(ioscompanion): wait out unresolved accessibility values after launch

* perf(ioscompanion): paste long text for atomic landing

* test(ioscompanion): cover paste threshold, retry flow, and sentinel detection

* fix(ioscompanion): treat unresolved bridge values as transitional, never as content

* fix(ioscompanion): accept masked secure-field values as paste landing

* test(ioscompanion): cover sentinel mapping and masked-field landing

* fix(ioscompanion): atomic erase and single-send paste to prevent doubling

* test(ioscompanion): cover atomic erase, single chord, unverifiable field

* fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout

* test(ioscompanion): cover bridge-blackout paste verification

* fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle

* refactor(ioscompanion): name the empty-editable-field sentinel for what it is

* perf(ioscompanion): tighten settle streak for the fast companion transport

* feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt

* refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt

* test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests

* fix(ioscompanion): retry describe past transient collapsed accessibility dumps

* test(ioscompanion): cover collapsed-dump detection

* perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses

* perf(ioscompanion): tighten settle now that collapses are handled separately

* fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text

* test(ioscompanion): cover replace-on-input and TextReplacer capability

* refactor(ioscompanion): neutralize HID events behind the transport seam

* feat(companion): add simulator runner project skeleton

* feat(companion): serve accessibility snapshots over the wire protocol

* feat(companion): synthesize timestamped touch gestures

* feat(companion): type text with replace semantics

* feat(companion): serve the wire protocol from a parked runner

* feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam

* feat(ioscompanion): route text input through a text-editing companion when available

* fix(companion): bind listener by port and source screen size from snapshot

* feat(ioscompanion): add runner companion JSON transport

* test(ioscompanion): cover runner transport protocol mapping

* fix(companion): synthesize gestures synchronously to avoid the async completion crash

* fix(companion): type on the main thread and recover from focus assertions

* fix(companion): keep serving after an automation failure

* refactor(companion): tidy snapshot serialization

* fix(companion): honor sequential tap gaps and survive synthesis exceptions

* feat(ioscompanion): expose native typing with an explicit replace flag

* chore(companion): add runner asset prepare script

* feat(ioscompanion): embed and extract the runner test bundle

* test(ioscompanion): cover runner asset extraction

* build(ioscompanion): commit runner asset archive

* feat(ioscompanion): pair the legacy companion with the in-simulator runner

* test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding

* fix(ioscompanion): reconnect after interrupted runner calls instead of restarting

* fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts

* feat(companion): launch and terminate apps through the automation session

* build(ioscompanion): refresh runner asset with session lifecycle

* fix(ioscompanion): classify connection deadline expiry as caller budget

* fix(companion): capture snapshots on the main thread inside the catch bridge

* build(ioscompanion): refresh runner asset with main-thread snapshots

* perf(ioscompanion): count read spans toward settle and capture snapshots concurrently

* feat(ioscompanion): make the hybrid simulator companion the default

* test(folio): cover runner-session orphans in the gate harness

* test(ioscompanion): pin the child-lifetime test to the legacy path

* fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears

* fix(ioscompanion): pause the clear chord so selection applies before the delete

* fix(companion): prune the keyboard subtree from snapshots

* build(ioscompanion): refresh runner asset without keyboard elements

* fix(ioscompanion): capture the screenshot transport before a recovery can reassign it

* fix(companion): pin the runner listener to loopback

* fix(companion): size the replace delete prefix to cover any focused field

* build(ioscompanion): refresh runner asset with loopback bind and replace fix

* fix(cli): cancel the run context on SIGINT so spawned children are reaped

* fix(testrun): point the device java preflight hint at the ios-device doctor

* fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob

* test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision

* chore: add test-companion target for the withcompanion-tagged suite

* chore(ioscompanion): stop tracking the runner archive build artifact

* build: produce the runner archive from source like the companion bundle

* refactor(conformance): move the gate harness out of examples/folio

* chore(folio): drop the gate harness wiring from the example app
2026-06-08 19:10:54 +05:30