android dumps a cross-fade with both screens in the tree. the route said
add-transaction while an unscoped find said home, so the oracle took a
half-rendered total as fresh and convicted on a tap that committed
nothing. one function now decides the route and returns null when the
frame is ambiguous.
summing cards went null when one was clipped, and the null poisoned the
carrier for the rest of the run. the balance window also spanned every
transaction since the last home visit, so the property convicted on
deltas it could not attribute: the old web witness was 3.16x the typed
amount, not 2x.
past 2^53 cents the gap between representable values is 128, so a real
1600-cent move reads back as something else and the equality is false
for a healthy submit as readily as a double one. also match parseCents:
a sign or an oversized amount is rejected, not read as an amount.
AXValue was the only source for text, but a StaticText carries its
string in AXLabel, so nothing on screen had .text on ios: a spec reading
it saw everything on android and nothing here.
compose for web merges the whole accountcard subtree, so the balance
child never exists there and every card parsed as 0. the property then
compared 0 to 0 and fired on any submit, which is a false positive
generator. unknown is now null and null is vacuously true.
the tree stays byte-identical and quiet across a cross-fade, so both the
quiet timer and the unchanged-tree escape called it settled mid-flight
and extractors read two screens at once.
an action's target was coordinates only, so a property matching on which
element was acted upon could never fire. ax.findAll([a,b]) also returned
nothing on web.
a launch the simulator rejects sent the xctest session into a recovery
chain that answered minutes late or never, and the rpc had no deadline,
so the run hung with no trace and no error. also take a per-udid flock:
a second run's reinstall lands under the first's live automation session
and wedges it.
* fix(replay-ui): drop @font-face rules for fonts that were never shipped
* fix(replay-ui): drop unresolvable JetBrains Mono from --font-mono stack
* chore(replay-ui): remove vestigial empty public/fonts dir
* feat(cli): add --max-steps for step-bounded runs
runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(trace): record arm membership and host in meta.json
meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(cli): add --arm and populate run meta from it
Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): sweep seeds for one experiment cell
campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.
Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(runner): no silent generator fallback, and llm on web
--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.
pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): make the hierarchy dump agree with the web runtime
Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.
scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.
The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(conformance): give the gate reproducible seeds
SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.
The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): emit editable as a plain boolean
editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(spec): leave the head subtree out of the web target walk
collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* test(chrome): compare the facts both hosts derive from one DOM
The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.
This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* chore(make): run the browser packages one at a time
Both launch Chrome and launching two at once has failed with "Launch: context
canceled".
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style: remove every em-dash and en-dash
Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): honor the caller context in Launch
Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.
The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(sidecarassets): publish the extracted jar through a rename
Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): kill a run that outlives --run-timeout
A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style(test): gofmt browser_test.go
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(ltl): give every thunk a construction identity
Two distinct unnamed predicates both described as "Thunk(...)", so obligation
collapse merged their residuals and could drop a live violation. Identity is
assigned at construction and the fields are unexported, so a thunk cannot be
built without one.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): reduce a thrown-predicate residual instead of panicking
The verifier substitutes an ErrorFormula for the residual of a property whose
predicate threw, and that residual is fed back in on the next step. reduce had
no case for it, so the run crashed. It re-reports the same failure now.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): make a bounded always the dual of a bounded eventually
G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still
pending when the window closed, so nnf's negation normal form was not semantics
preserving. Both sides now range over the observations at which their inner can
definitely resolve: the eventually keeps a pending inner as a disjunct, and the
always discharges vacuously at window close.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): arm a one-shot root once per run
A root that carries its own horizon is one obligation for the whole run, not one
per observation. Re-instantiating a top-level eventually monitored G F<=n(p)
instead of F<=n(p) and left one live obligation per step behind; a bounded
always restarted its window every step and never closed.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): stop wrapping a top-level eventually in always
`eventually(p).within(300, "seconds")` as a property meant "within 300 seconds
of every step", which spawned an obligation per step with its own resolved
deadline. A 553-step run carried 553 of them and serialized a 75 KB residual.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(ltl): serialize the resolved deadline of a bounded window
Two obligations spawned at different steps from one duration-bounded formula
differ only in the deadline the evaluator resolved for them, so they serialized
identically and the trace erased a distinction the evaluator makes. The authored
window stays in amount/unit; the resolved deadline rides alongside.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): split a witness's origin step from its detection step
A deferred obligation spans two steps: the one that armed it and the one whose
reduction failed. They were conflated under one index, so the extractor snapshot
(which is the detecting step's state) was reported against the origin step.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(runner): record a witness's detection step in the trace
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* feat(replay-ui): show the step a violation was detected at
The witness evidence is the detecting step's state, so say which step that is
and let a reader jump to it.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(verifier): record the extractor state the predicates actually read
On the web path extractor bodies are evaluated in V8 and injected here, but only
the goja value was replaced. The trace diff and the violation witness therefore
described a state no property ever saw.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* refactor(spec): one candidate producer over one target-eligibility rule
Both hosts routed verbs themselves and both policies enumerated their own
actions, and all four drifted. Web sent `swipes` to scrollable containers only,
so swipe-to-dismiss on a list row was reachable on native and unreachable on
web; the model policy folded gestures its own way and could not reach what the
seeded picker drew.
A host now reports facts about every element and never decides which verb may
act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates
is the single enumeration, and the model policy reads it through
__sanderlingEnumerateBuiltin__ instead of reimplementing it in Go.
Gesture verbs change with it: scrolls stay vertical over scrollable containers,
swipes go free-form in all four directions from any element with real bounds.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(runner): name a builtin scroll by its drag origin
A builtin gesture carries endpoints and no selector, so every scroll rendered as
"Scroll down " in the prompt's recent-action memory and two scrollable regions
were indistinguishable.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(chrome): clear storage over cdp instead of scripting an opaque origin
Launch runs while the tab is still on about:blank, whose opaque origin denies
storage access, so localStorage.clear() threw SecurityError and every web run
died at launch. Storage.clearDataForOrigin needs no navigation. The exception
helper lands here because "Uncaught" is what hid this for so long.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(chrome): enable the swiftshader webgl fallback
Headless Chrome runs with --disable-gpu, and without this flag it refuses the
software WebGL backend: getContext returns null, so a canvas-rendered app paints
nothing and every screenshot is identical black.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* fix(web): resolve testTag through data-testid or id
Compose Multiplatform emits its testTag into the element id, which the native
table already accepts via the resource-id alias. The two web selector tables
were the only place that rejected it.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* test(spec): type-check the spec api as part of make test
The fake runtime in api.test.ts did not return a chainable handle from extract,
so the file had not type-checked since named() was added. Wiring the check into
make test stops it drifting again.
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* docs(manual): one-shot eventually and the gesture verbs
Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
* feat(spec): add llm() action-backend marker
* feat(spec): make llm marker inert on the JS picker
* feat(spec): expose __sanderlingSampleInput__ corpus draw
* feat(openrouter): minimal chat-completions client
* test(openrouter): cover request shape, parse, and errors
* feat(verifier): thread screenshot + capture corpus sampler
* feat(verifier): LLM accessors — candidates, config, sampler
* test(verifier): cover AllCandidates, LLMConfig, SampleInput
* feat(trace): record action Source and LLMReasoning
* feat(runner): thread step screenshot into PushSnapshot
* feat(runner): llmSource selects actions via OpenRouter
* feat(runner): wire llmSource selection and trace stamping
* test(runner): cover llmSource selection, mapping, downscale
* docs(folio): add llm action-backend example spec
* docs(folio): document the LLM action backend run
* feat(llmclient): support OPENAI_API_KEY, openrouter wins
* refactor(runner): rename openrouter package to llmclient
* docs: both api keys, example model gpt-5.4-nano
* docs: add pr style rules to claude.md
* fix(runner): explain action kinds in llm prompt to stop swipe loops
* feat(trace): record llm ranked list and chosen rank
* feat(runner): stamp llm ranked list and chosen rank on trace
* fix(runner): tap by selector to survive layout shift after observe
* revert(runner): drop selector-first tap; broke path/testTag selectors
* feat(spec): llm() accepts optional instructions
* feat(verifier): read llm instructions off config
* feat(runner): append spec instructions to llm system prompt
* docs(folio): describe app in llm spec instructions
* feat(bundler): map generator export to globalThis.generator
* feat(verifier): read llm config off globalThis.generator
* feat(runner): gate llm source on --generator flag
* feat(cmd): add --generator llm|seeded flag
* test: cover --generator flag parsing and pickSources gating
* feat(verifier): enumerate llm candidates by walking actionsRoot
collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.
* test(verifier): cover candidate enumeration walk
* feat(verifier): add SetupAction to walk setup without the seeded root
* test(verifier): cover SetupAction setup-only precedence
* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order
* feat(trace): record llm choice number and chosen_action echo
* feat(runner): llm picks one number from weighted candidates
drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).
* test(runner): cover choice schema, strict-skip, and setup precedence
* refactor(verifier): drop the superseded AllCandidates enumeration
* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts
* fix(verifier): label editable fields by hint, not the typed value
an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.
* test(runner): cover weight-suffixed echo and stripWeightSuffix
* fix(runner): accept chosen_action echo that carries the weight suffix
real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).
* fix(verifier): skip llm enumeration on cross-fade frames
a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.
* feat(folio): show current balance on the add-transaction screen
renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.
* fix(replay): derive device space from screen extent, not first node
the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.
* fix(folio): show balance as a compact one-line label
per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.
* fix(folio): move balance into the header, one compact line under the account name
* fix(replay): attribute deferred violations to the causing step, not detection
* fix(replay): show a step's own violations in both panels, no next-step bleed
* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check
* chore: ignore .playwright-mcp scratch output
* docs: document the llm generator and --generator flag
* docs(spec): correct the llm() comment; config reads off globalThis.generator
* docs: add pr description rules
* feat(sidecar): reach USB devices via the adb server by serial
* feat(test): add --device flag to target a specific Android device by serial
* feat(folio): select Android device via ANDROID_DEVICE in justfile
* feat(conformance): add android backend to the gate suite
* feat(android): keep device awake and unlocked so the app stays foreground
* feat(conformance): prep physical android device (autofill/verifier/stayon)
* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run
* fix(verifier): require positive bounds for swipe candidates
A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.
* fix(runner): harden app-scope guard against launcher and overlays
The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.
* feat(android): harden physical-device runs in device prep
Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.
* feat(driver): clear-state via APK reinstall when pm clear is blocked
When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.
* feat(cli): add --android-app-path for clear-state reinstall
Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.
* chore(folio): pass --android-app-path in just test
* fix(runner): clamp swipe/scroll origin out of edge gesture zones
A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.
* perf(sidecar): faster Android text input and drop redundant settle poll
inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.
* fix(verifier): exclude soft-keyboard region from action candidates
The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.
* perf(runner): replace focus-tap settle with a brief wait
The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.
* chore(conformance): platform-aware G5 p95 budget for android
The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.
* fix(sidecar): retry maestro android driver startup
The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.
* chore(conformance): widen android G5 budget to 5500ms
Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.
* web replay fix
* feat(android): force 3-button nav during runs to prevent app drift
On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.
* fix(android): target the selected device in adb reads; don't strand nav mode
Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
foreground/scope guard works when several devices are attached (the --device
path). Previously they ran bare `adb shell`, which errors with multiple
devices, silently disabling app-scope enforcement. The sidecar client passes
its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
the current mode is unknown or already 3-button it leaves nav untouched,
instead of switching and then stranding the device in 3-button. Logic split
into the pure navModeToRestore, now unit tested.
* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds
Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.
* test(verifier): cover keyboardRegionTop, including the decor-view guard
The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.
* style(cli): gofmt testOptions field alignment
* fix(sidecar): keep a leading dash off the fast input path
A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).
* refactor(verifier): scope action candidates by window ownership
Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.
This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.
* fix(runner): re-check foreground at apply time, skip stale actions
ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.
* fix(android): type long ASCII via fast guarded path to stop keystroke escape
A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.
* chore: ignore gate artifacts and local scratch files
* refactor(runner): narrow gesture clamp to the top shade strip
3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.
* chore(format): add .editorconfig enforcing 80-column limit
* chore(format): add prettier config with 80-char printWidth
* chore(deps): add prettier devDependency to replay-ui
* chore(deps): add prettier devDependency to folio-web
* chore(deps): add prettier devDependency to spec package
* chore(format): add swift-format config with 80-char lineLength
* feat(format): add make fmt targets for per-language 80-col formatting
* fix(runner): translate gesture to safe area so near-top scrolls keep direction
Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.
* fix(runner): apply-time guard consults focused window, not just resumed activity
ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.
* test(runner): cover apply-time foreground skip and appIsForeground table
Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.
* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker
- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.
* fix(conformance): pin self-test p95 budget and score install failures as run failures
self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.
* fix(android): require --device when several devices are connected
With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.
* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands
The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.
* perf(verifier): memoize scopedElements per tree
scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.
* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear
Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.
* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms
The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.
* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty
bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.
* perf(sidecar): reuse a single Jackson ObjectMapper
structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.
* refactor(android): remove unused AdbReverse/AdbReverseRemove
No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.
* style(runner): trim non-load-bearing comments from this PR's runner code and tests
* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
* docs(manual): add introduction page
* docs(manual): rewrite getting started as guided first run
* docs(manual): rewrite writing specs as a folio tutorial
* docs(manual): document missing spec API in reference
* docs(manual): plain-language rewrite of runs page
* docs: real introductions on index pages and README
* fix(docs): sibling links from directory-style pages need ../
* fix(docs): correct sampling and restart-cost claims to match implementation
* docs: nav lists Introduction and Case study; roadmap points to milestone
* docs(manual): make getting started target the reader's own app, not Folio
* docs(manual): add Folio case study page
* docs: point manual navigation at the case study
* docs(readme): lead with the case study, fix roadmap link
* docs: roadmap links to milestone, sync clear-data default and cross-links
* feat(companion): add appState, eraseText, pressKey runner handlers
The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.
* feat(ios): resolve physical devices from devicectl
ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.
* feat(ioscompanion): runner-only device driver mode
NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.
* test(ioscompanion): cover device wiring, routing, and shell-out argv
Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.
* feat(testrun): route physical-device iOS runs to the device driver
Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.
* feat(doctor): device prereqs replace java/sidecar for ios-device
iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.
* feat(conformance): device backend uses iphoneos app and tunnel orphan checks
The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).
* feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.
* docs(cli): document ios-device doctor checks and the device flags
The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.
* fix(ioscompanion): resolve signing key path to absolute
xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.
* fix(ioscompanion): re-enable signing for the device runner build
companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.
* fix(ioscompanion): key the device build cache on signing identity
The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.
* docs(getting-started): document physical iOS device setup
Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.
* feat(ios): native usbmux client and in-process tunnel forwarder
Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.
* refactor(ios): drive device tunnel via io.Closer seam
Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.
* refactor(ios): remove iproxy spawn from device runner
* test(ios): cover tunnel close via io.Closer not child process
* feat(doctor): check usbmuxd socket instead of iproxy on PATH
* chore(conformance): drop iproxy orphan check; tunnel is in-process
* docs(ios): device tunnel uses native usbmux, nothing to install
* chore: gitignore the signing keys directory
* feat(folio): add Android launcher icon (black bg, white dot)
* feat(folio): add iOS app icon (black bg, white dot)
* feat(folio): add web favicon (black bg, white dot)
* docs(ioscompanion): fix stale const comments
* refactor(ioscompanion): inline single-use devicectl argv builders
* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests
* refactor(ioscompanion): inline firstNonEmpty
* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam
* test(doctor): trim redundant signing-check test
* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env
* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun