mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
85358f007e691f9305e86498d6809e6b7ae54dbf
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7343085614 |
llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker * feat(spec): make llm marker inert on the JS picker * feat(spec): expose __sanderlingSampleInput__ corpus draw * feat(openrouter): minimal chat-completions client * test(openrouter): cover request shape, parse, and errors * feat(verifier): thread screenshot + capture corpus sampler * feat(verifier): LLM accessors — candidates, config, sampler * test(verifier): cover AllCandidates, LLMConfig, SampleInput * feat(trace): record action Source and LLMReasoning * feat(runner): thread step screenshot into PushSnapshot * feat(runner): llmSource selects actions via OpenRouter * feat(runner): wire llmSource selection and trace stamping * test(runner): cover llmSource selection, mapping, downscale * docs(folio): add llm action-backend example spec * docs(folio): document the LLM action backend run * feat(llmclient): support OPENAI_API_KEY, openrouter wins * refactor(runner): rename openrouter package to llmclient * docs: both api keys, example model gpt-5.4-nano * docs: add pr style rules to claude.md * fix(runner): explain action kinds in llm prompt to stop swipe loops * feat(trace): record llm ranked list and chosen rank * feat(runner): stamp llm ranked list and chosen rank on trace * fix(runner): tap by selector to survive layout shift after observe * revert(runner): drop selector-first tap; broke path/testTag selectors * feat(spec): llm() accepts optional instructions * feat(verifier): read llm instructions off config * feat(runner): append spec instructions to llm system prompt * docs(folio): describe app in llm spec instructions * feat(bundler): map generator export to globalThis.generator * feat(verifier): read llm config off globalThis.generator * feat(runner): gate llm source on --generator flag * feat(cmd): add --generator llm|seeded flag * test: cover --generator flag parsing and pickSources gating * feat(verifier): enumerate llm candidates by walking actionsRoot collect-walk the weighted action tree: recurse weighted branches accumulating selection probability, call authored leaves once for concrete actions, enumerate builtins per element. label controls by visible text (borrowing descendant text), fold gestures into directional scrolls over scrollable containers, drop disabled, dedup descriptions. * test(verifier): cover candidate enumeration walk * feat(verifier): add SetupAction to walk setup without the seeded root * test(verifier): cover SetupAction setup-only precedence * refactor(llmclient): make JSONSchema.Schema raw json for pinned field order * feat(trace): record llm choice number and chosen_action echo * feat(runner): llm picks one number from weighted candidates drop the seeded-root call for a setup-only precedence path, render a numbered weighted candidate list, pin a reasoning-first choice schema, strict-skip when chosen_action does not echo the numbered entry, and let the model supply typed values (corpus fallback when empty). * test(runner): cover choice schema, strict-skip, and setup precedence * refactor(verifier): drop the superseded AllCandidates enumeration * feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts * fix(verifier): label editable fields by hint, not the typed value an editable field's own text is its transient content; prefer the hint so the field is named by purpose and the label stays stable. * test(runner): cover weight-suffixed echo and stripWeightSuffix * fix(runner): accept chosen_action echo that carries the weight suffix real runs showed the model copies the whole numbered line including the trailing (w34) weight annotation, so strict-skip rejected ~91% of picks and the llm was paralyzed. strip the weight suffix before comparing. also nudge the prompt to stress-test repeated submissions (idempotency). * fix(verifier): skip llm enumeration on cross-fade frames a navhost mid-transition carries >1 route *Screen in a collapsed coordinate space; acting on it taps garbage (soft keyboard). real runs showed the llm acting on 44% of steps being such frames. skip them so the llm re-observes a settled frame next step. * feat(folio): show current balance on the add-transaction screen renders the account's balance (testTag TxnCurrentBalance) below the account name, above the credit/debit toggle, so before/after screenshots carry comparison data. * fix(replay): derive device space from screen extent, not first node the first positive-bounds element is often a short status-bar node (320x24 on android); using it gave a 320/24 aspect ratio that squashed the screenshot overlay into a grey horizontal band. use the max extent across elements (like the runner's screenBounds) instead. * fix(folio): show balance as a compact one-line label per review: one line, account-name-sized, e.g. "Balance: $0.00" instead of a large balance card. * fix(folio): move balance into the header, one compact line under the account name * fix(replay): attribute deferred violations to the causing step, not detection * fix(replay): show a step's own violations in both panels, no next-step bleed * refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check * chore: ignore .playwright-mcp scratch output * docs: document the llm generator and --generator flag * docs(spec): correct the llm() comment; config reads off globalThis.generator * docs: add pr description rules |
||
|
|
6b0d6cb971 |
WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial * feat(test): add --device flag to target a specific Android device by serial * feat(folio): select Android device via ANDROID_DEVICE in justfile * feat(conformance): add android backend to the gate suite * feat(android): keep device awake and unlocked so the app stays foreground * feat(conformance): prep physical android device (autofill/verifier/stayon) * fix(android): make device prep best-effort so OEM-blocked commands don't abort the run * fix(verifier): require positive bounds for swipe candidates A zero-bounds element centers at (0,0); a downward swipe from the top-left corner is the system gesture that pulls down the notification shade, dragging the fuzzer out of the app. Swipes now require positive bounds like every other verb. * fix(runner): harden app-scope guard against launcher and overlays The per-step guard now relaunches and waits until the app window is actually drawn before proceeding, so a slow physical-device relaunch no longer lets an observe or action land on the launcher. It also detects a system overlay (notification shade) stealing window focus while the app stays resumed, and dismisses it with back. * feat(android): harden physical-device runs in device prep Device prep now disables the AOSP cached-app freezer, phantom-process killer, and Doze (and exempts the driver) so OEM background management stops suspending the driver mid-run. Adds ReinstallApp for clear-state on ROMs that deny pm clear, and teaches focus detection to report the notification shade as systemui so the scope guard can dismiss it. * feat(driver): clear-state via APK reinstall when pm clear is blocked When an APK path is set, Android clear-state resets the app by uninstalling and reinstalling instead of asking the sidecar to pm clear, which hardened OEM builds (ColorOS) deny even to the adb shell user. Falls back to the sidecar clear path when no APK path is provided. * feat(cli): add --android-app-path for clear-state reinstall Wires the APK path from the test command through to the sidecar client so Android clear-state can reset apps on OEM builds that deny pm clear. * chore(folio): pass --android-app-path in just test * fix(runner): clamp swipe/scroll origin out of edge gesture zones A gesture starting in the top status-bar strip pulls down the notification shade; the bottom and side strips are the home and back gestures. Any of them drags the fuzzer out of the app. Swipe and scroll origins are now clamped into a safe inner area sized from the maximum element extent (the Android hierarchy root reports zero bounds, so the extent is the reliable screen size). Calibrated on device: origins below ~7% of height no longer open the shade. * perf(sidecar): faster Android text input and drop redundant settle poll inputText now uses adb `input text` for short shell-safe ASCII (~5x faster than the driver's per-character path) and falls back to the driver for unicode, injection payloads, and overflow-length strings. waitForIdle drops the structural-hash poll that followed waitForAppToSettle: each hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per mutating step for marginal benefit, and the runner already re-fetches transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4 still pass. * fix(verifier): exclude soft-keyboard region from action candidates The fuzzer was tapping Gboard's "Settings" key, navigating out of the app. That key is a bare FrameLayout with a content-desc and no package or resource-id, so the package-based scope filter missed it. Candidates whose center falls in the keyboard region (derived from the IME elements' bounds) are now dropped, so no tap or long-press lands on a key. Opt-in with app scoping; unscoped runs keep every node. * perf(runner): replace focus-tap settle with a brief wait The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText step on a physical device while the keyboard animated in. The tap registers focus immediately and text is injected into the focused view, so a short fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay green. * chore(conformance): platform-aware G5 p95 budget for android The 2500ms ceiling was calibrated on the iOS simulator. A physical Android device drives every step over USB (snapshot + settle + adb round-trips), so its per-step floor is several times higher; holding it to 2500ms would force removing the settle/retry logic the correctness gates depend on. The android backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500. * fix(sidecar): retry maestro android driver startup The maestro Android driver's dadb.open() occasionally misses its startup deadline (its instrumentation host is slow to come up right after a reboot or per-run reinstall), which aborted the whole run. Retry the open a few times with a short backoff so a transient timeout recovers. * chore(conformance): widen android G5 budget to 5500ms Physical-device p95 swung 3209-4612ms across sessions (cold runs right after a reboot are slower). 4500ms was too tight for that jitter; 5500ms covers the observed ceiling with headroom. * web replay fix * feat(android): force 3-button nav during runs to prevent app drift On gesture navigation a fuzzer swipe can trigger swipe-up-home or edge-back and fling the app off screen. Device-prep now switches to 3-button navigation for the run (no edge gestures; the nav bar's buttons are systemui-owned and already excluded from action candidates) and restores the original navigation mode when the run ends. Best effort: leaves nav untouched if the overlay command is unavailable. * fix(android): target the selected device in adb reads; don't strand nav mode Review fixes: - ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the foreground/scope guard works when several devices are attached (the --device path). Previously they ran bare `adb shell`, which errors with multiple devices, silently disabling app-scope enforcement. The sidecar client passes its serial through. - Extract an adbArgs helper and route every adb call through it, removing four duplicated serial-arg builders. - ForceThreeButtonNav now decides what to restore before changing anything: if the current mode is unknown or already 3-button it leaves nav untouched, instead of switching and then stranding the device in 3-button. Logic split into the pure navModeToRestore, now unit tested. * fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was stranded above screenBounds by an insertion). Extend the clamp test to assert an off-screen destination is clamped onto the screen and that the origin lands exactly on the margin. * test(verifier): cover keyboardRegionTop, including the decor-view guard The full-screen IME decor view rejection had no test; removing it left the suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in favor of the real keyboard line, and decor-only -> sentinel. * style(cli): gofmt testOptions field alignment * fix(sidecar): keep a leading dash off the fast input path A value starting with '-' could be read as an option by `adb input text`, so the fast-path regex now requires a non-dash first character; such values fall back to the driver. Also cover the dadb-target branch where a colon precedes a non-numeric port (a USB serial, not host:port). * refactor(verifier): scope action candidates by window ownership Replaces the leaky per-element package check and the keyboard-region Y heuristic with one rule: walk the window tree propagating each node's owning package (empty and the neutral android framework package are transparent); a node is in scope only when no concrete foreign package owns it (the app's own window carries no package on Compose apps) or the owner is the app package. This drops whole foreign windows (soft keyboard, system UI, launcher) AND their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which the old empty-package-is-in-scope rule admitted and which navigated out of the app. Deletes keyboardRegionTop/isInputMethodElement. * fix(runner): re-check foreground at apply time, skip stale actions ensureForeground runs before observe, but the app can leave between observe and apply (a prior gesture settling late); swipes/keys then fire stale coordinates onto whatever screen is now up. Re-check foreground immediately before applying and, when the app is gone, skip the action and log it (making the escape visible) so the next step's guard relaunches instead. * fix(android): type long ASCII via fast guarded path to stop keystroke escape A 4096-char corpus string exceeded the fast input cap and fell to the per-character driver path, which takes ~120s. During that uninterruptible window focus could leave the app and the remaining keystrokes sprayed into the launcher search box. Route shell-safe ASCII of any length through adb input text, chunked, re-checking the foreground app between chunks and stopping if it changed. * chore: ignore gate artifacts and local scratch files * refactor(runner): narrow gesture clamp to the top shade strip 3-button nav (forced for every run) disables the side back and bottom home gestures at the OS level. On-device probing confirmed side and bottom swipe origins no longer drift, leaving the notification shade as the only edge gesture a swipe can trigger. Clamp only the top strip; keep origin and destination on screen otherwise. * chore(format): add .editorconfig enforcing 80-column limit * chore(format): add prettier config with 80-char printWidth * chore(deps): add prettier devDependency to replay-ui * chore(deps): add prettier devDependency to folio-web * chore(deps): add prettier devDependency to spec package * chore(format): add swift-format config with 80-char lineLength * feat(format): add make fmt targets for per-language 80-col formatting * fix(runner): translate gesture to safe area so near-top scrolls keep direction Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp. * fix(runner): apply-time guard consults focused window, not just resumed activity ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay. * test(runner): cover apply-time foreground skip and appIsForeground table Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised. * fix(sidecar): harden android driver open, input guard, pressKey, foreground marker - openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF). - pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract. - the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner. - foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard. * fix(conformance): pin self-test p95 budget and score install failures as run failures self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue. * fix(android): require --device when several devices are connected With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous. * refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test. * perf(verifier): memoize scopedElements per tree scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot. * fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail. * test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup. * refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true. * perf(sidecar): reuse a single Jackson ObjectMapper structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val. * refactor(android): remove unused AdbReverse/AdbReverseRemove No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed. * style(runner): trim non-load-bearing comments from this PR's runner code and tests * style(sidecar): trim non-load-bearing comments from this PR's driver code and tests |
||
|
|
90224dfd06 |
Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers
The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.
* feat(ios): resolve physical devices from devicectl
ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.
* feat(ioscompanion): runner-only device driver mode
NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.
* test(ioscompanion): cover device wiring, routing, and shell-out argv
Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.
* feat(testrun): route physical-device iOS runs to the device driver
Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.
* feat(doctor): device prereqs replace java/sidecar for ios-device
iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.
* feat(conformance): device backend uses iphoneos app and tunnel orphan checks
The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).
* feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.
* docs(cli): document ios-device doctor checks and the device flags
The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.
* fix(ioscompanion): resolve signing key path to absolute
xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.
* fix(ioscompanion): re-enable signing for the device runner build
companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.
* fix(ioscompanion): key the device build cache on signing identity
The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.
* docs(getting-started): document physical iOS device setup
Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.
* feat(ios): native usbmux client and in-process tunnel forwarder
Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.
* refactor(ios): drive device tunnel via io.Closer seam
Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.
* refactor(ios): remove iproxy spawn from device runner
* test(ios): cover tunnel close via io.Closer not child process
* feat(doctor): check usbmuxd socket instead of iproxy on PATH
* chore(conformance): drop iproxy orphan check; tunnel is in-process
* docs(ios): device tunnel uses native usbmux, nothing to install
* chore: gitignore the signing keys directory
* feat(folio): add Android launcher icon (black bg, white dot)
* feat(folio): add iOS app icon (black bg, white dot)
* feat(folio): add web favicon (black bg, white dot)
* docs(ioscompanion): fix stale const comments
* refactor(ioscompanion): inline single-use devicectl argv builders
* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests
* refactor(ioscompanion): inline firstNonEmpty
* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam
* test(doctor): trim redundant signing-check test
* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env
* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
|
||
|
|
406b7516b3 |
iOS simulator driver: Go-native companion-backed backend (#62)
* perf(ios): use prebuilt XCTest runner to cut startup * chore(ioscompanion): add companion asset prepare script * feat(ioscompanion): embed and extract simulator companion bundle * test(ioscompanion): cover companion stub and embedded extraction * docs: add third party notices for vendored companion * chore: ignore vendored companion bundle artifact * build(proto): pin simulator companion proto v1.1.8 * build(proto): add dedicated buf module and gen template for pinned proto * build(proto): exclude pinned companion proto from root buf workspace * feat(ioscompanion): commit generated companion gRPC stubs * feat(ioscompanion): map flat companion describe dump to TreeNode JSON * test(ioscompanion): add hierarchy-map golden and unit tests * feat(ioscompanion): port screen-settle stability polling to Go * test(ioscompanion): cover settle transitional, hash, streak, and cap rules * feat(ioscompanion): add USB HID keymap module * test(ioscompanion): cover keymap branches and paste-chord constants * build: embed companion assets via withcompanion tag * feat(ioscompanion): add transport companion interface * feat(ioscompanion): add HID event wrapper and builders * feat(ioscompanion): wire gRPC companion client and Dial * test(ioscompanion): cover HID builders and unit conversions * test(ioscompanion): cover Dial, process-state mapping, and install archive * test(ioscompanion): add gated simulator integration smoke test * feat(ioscompanion): text input and gesture HID composition with pasteboard fallback * test(ioscompanion): cover input composers, paste dialog loop, and pure helpers * feat(ioscompanion): add Describe to companion transport * feat(ioscompanion): implement DeviceDriver with companion supervision * test(ioscompanion): unit tests with fake companion transport * test(ioscompanion): gated companion smoke test * feat(ios): add ResolveTarget for simulator vs physical-device routing * feat(testrun): route iOS simulators through the native companion driver * refactor(testrun): defer the java preflight check to the physical-device path * feat(cli): add --ios-app-path flag * feat(doctor): split iOS checks into simulator and physical-device paths * test(folio): add gate-analyzer fixtures for G1-G5 * feat(folio): add iOS conformance gate script * chore(folio): wire gates recipe, app path, and ignore gate output * style: gofmt struct alignment drift * fix(doctor): probe simctl via xcrun instead of PATH lookup * fix(ioscompanion): spawn companion under driver-lifetime context * test(ioscompanion): prove companion child outlives startup context * fix(ioscompanion): chunk install payload under companion message cap * test(ioscompanion): cover install payload chunking * fix(ioscompanion): reinstall via simctl and sanitize companion env * fix(ioscompanion): wait out unresolved accessibility values after launch * perf(ioscompanion): paste long text for atomic landing * test(ioscompanion): cover paste threshold, retry flow, and sentinel detection * fix(ioscompanion): treat unresolved bridge values as transitional, never as content * fix(ioscompanion): accept masked secure-field values as paste landing * test(ioscompanion): cover sentinel mapping and masked-field landing * fix(ioscompanion): atomic erase and single-send paste to prevent doubling * test(ioscompanion): cover atomic erase, single chord, unverifiable field * fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout * test(ioscompanion): cover bridge-blackout paste verification * fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle * refactor(ioscompanion): name the empty-editable-field sentinel for what it is * perf(ioscompanion): tighten settle streak for the fast companion transport * feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt * refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt * test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests * fix(ioscompanion): retry describe past transient collapsed accessibility dumps * test(ioscompanion): cover collapsed-dump detection * perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses * perf(ioscompanion): tighten settle now that collapses are handled separately * fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text * test(ioscompanion): cover replace-on-input and TextReplacer capability * refactor(ioscompanion): neutralize HID events behind the transport seam * feat(companion): add simulator runner project skeleton * feat(companion): serve accessibility snapshots over the wire protocol * feat(companion): synthesize timestamped touch gestures * feat(companion): type text with replace semantics * feat(companion): serve the wire protocol from a parked runner * feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam * feat(ioscompanion): route text input through a text-editing companion when available * fix(companion): bind listener by port and source screen size from snapshot * feat(ioscompanion): add runner companion JSON transport * test(ioscompanion): cover runner transport protocol mapping * fix(companion): synthesize gestures synchronously to avoid the async completion crash * fix(companion): type on the main thread and recover from focus assertions * fix(companion): keep serving after an automation failure * refactor(companion): tidy snapshot serialization * fix(companion): honor sequential tap gaps and survive synthesis exceptions * feat(ioscompanion): expose native typing with an explicit replace flag * chore(companion): add runner asset prepare script * feat(ioscompanion): embed and extract the runner test bundle * test(ioscompanion): cover runner asset extraction * build(ioscompanion): commit runner asset archive * feat(ioscompanion): pair the legacy companion with the in-simulator runner * test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding * fix(ioscompanion): reconnect after interrupted runner calls instead of restarting * fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts * feat(companion): launch and terminate apps through the automation session * build(ioscompanion): refresh runner asset with session lifecycle * fix(ioscompanion): classify connection deadline expiry as caller budget * fix(companion): capture snapshots on the main thread inside the catch bridge * build(ioscompanion): refresh runner asset with main-thread snapshots * perf(ioscompanion): count read spans toward settle and capture snapshots concurrently * feat(ioscompanion): make the hybrid simulator companion the default * test(folio): cover runner-session orphans in the gate harness * test(ioscompanion): pin the child-lifetime test to the legacy path * fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears * fix(ioscompanion): pause the clear chord so selection applies before the delete * fix(companion): prune the keyboard subtree from snapshots * build(ioscompanion): refresh runner asset without keyboard elements * fix(ioscompanion): capture the screenshot transport before a recovery can reassign it * fix(companion): pin the runner listener to loopback * fix(companion): size the replace delete prefix to cover any focused field * build(ioscompanion): refresh runner asset with loopback bind and replace fix * fix(cli): cancel the run context on SIGINT so spawned children are reaped * fix(testrun): point the device java preflight hint at the ios-device doctor * fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob * test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision * chore: add test-companion target for the withcompanion-tagged suite * chore(ioscompanion): stop tracking the runner archive build artifact * build: produce the runner archive from source like the companion bundle * refactor(conformance): move the gate harness out of examples/folio * chore(folio): drop the gate harness wiring from the example app |
||
|
|
b44077afde |
replay ui fix (#56)
* refactor: rename inspect to replay across the codebase Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/, the CLI subcommand from `sanderling inspect` to `sanderling replay`, and updates all references in docs, Makefile, README, and Go comments. * feat(replay-ui): show spec filename with full path on hover RunList and RunDetail now render the basename of spec_path (e.g. login.spec.ts) with the full path available as a title tooltip. |
||
|
|
c5bb176be8 |
UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed, and surface both in describe() and MarshalJSON. * feat(ltl): negation normal form pass nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error leaf, dualizing Always<->Eventually and preserving bounds. * feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse Apply nnf on construction, reduce bounded Always symmetric to bounded Eventually (vacuous holds once the window closes), add Finalize to resolve undischarged liveness obligations to Violated at run end, and collapse structurally-identical pending obligations. * test(ltl): property-based NNF laws Lock double-negation identity, Always/Eventually duality with bound preservation, leaf pushdown, and not(always true) reaching Violated. * test(ltl): Finalize, bounded eventually, latch, collapse Property tests for monotonic violation latch and eventually-within violating iff n consecutive false, plus Finalize and collapse cases. * feat(inspect): within clause on always residual node A negated bounded eventually serializes as a bounded always; render its bound instead of dropping it. * feat(ltl): witness violations and (bool,error) predicate thunks * test(ltl): migrate thunk call sites to (bool,error) * feat(ltl): flag thrown-predicate witnesses with IsError * refactor(verifier): replace predicate err side-channel with violation witness * test(verifier): witness API for thrown predicates * feat(trace): witnesses map and skipped-verification marker on Step * feat(runner): thread violation witnesses, finalize, skip marker into trace * test(ltl): lock violation witness reason, IsError, and step * test(verifier): finalize surfaces unmet eventually with witness * fix(ltl): eliminate implies and bounded-always false-negatives Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent can no longer defer the whole implication and drop a consequent that was false at the current step. Carry a pending inner past a bounded-Always window close instead of dropping it to holds, so a deferred obligation is resolved by a later step or Finalize. * test(ltl): lock implies and bounded-always false-negative regressions * fix(web-runtime): seed PRNG for reproducible runs and align weighted pick * feat(testrun): inject seed into web bundle via SANDERLING_SEED define * test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring * test(spec): add Go math/rand/v2 PCG oracle and golden fixture * feat(spec): bit-exact PCG port of Go math/rand/v2 * test(spec): assert pcg.ts matches the PCG golden fixture * feat(spec): shared input corpus and press-key pools * feat(spec): action-tree types and Host interface * feat(spec): verb support matrix and warn-once helper * feat(spec): deterministic shared action picker * test(spec): verb matrix and warn-once semantics * test(spec): picker draw-order and determinism * refactor(spec): actions.ts returns pure GeneratorNode data trees * refactor(spec): wire from() sampling through the picker rng * feat(spec): shared runtime-entry installs next-action over pick.ts * feat(spec): export LongPress/Scroll/longPresses/scrolls factories * test(spec): assert data-tree shapes for action factories * test(spec): runtime-entry serializeAction wire-contract round-trip * refactor(spec): bridge data-tree nodes to the legacy goja picker tags * fix(spec): web runtime walks the spec's globalThis.actions data tree * test(spec): tolerate legacy bridge fields on builtin nodes * refactor(spec): installRuntime accepts a lazy root resolver The web bundle imports the runtime before the spec, so the action root on globalThis.actions only exists after the spec evaluates. Accept a function form so the goja and web hosts resolve the root per tick. * refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/ randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG, and the snake_case serializeAction) plus the __sanderling__ action factory binds. web-runtime now implements Host (platform/seedHi/seedLo from the injected 64-bit seed via BigInt, queryCandidates over the live DOM with a per-tick cache, reportUnsupported) and calls installRuntime so both engines run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts matrix instead of silently returning null. Keeps the DOM helpers (selector translation, queryElement, elementHandle, buildState, sanitize, extractors) and the global locking. Net -214 lines (741 -> 527). * test(spec): cover the WEB Host surface and seed precision Replace the deleted-picker tests with Host coverage: platform()==web, seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0, reportUnsupported warning, the installed next-action/extractor globals, and queryCandidates verb routing + per-tick caching over a querySelectorAll stub. * refactor(spec): picker emits native selector + scroll endpoints, setup precedence * feat(spec): goja runtime entry wires the shared picker over the Go host * feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin * feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker * refactor(spec): drop the legacy goja bridge fields from action factories * feat(spec): serialize selector-only string targets for the runner to re-resolve * refactor(verifier): one DecodeAction reads the unified flat wire contract * refactor(verifier): goja host + shared picker replace the duplicate Go picker * refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime * test(verifier): author specs through the shared picker path * test(runner): bundle authored specs with the goja runtime entry * feat(verifier): collect unsupported verbs for the run report * refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource * feat(testrun): surface unsupported verbs in run report * test(verifier): cross-runtime goja/node parity gate on the shared picker * test(verifier): unsupported verbs collected deduped in first-seen order * test(runner): summary reports no unsupported verbs on a clean run * test(spec): golden-fixture cross-runtime parity gate for the node picker Replace the env-driven parity harness with a shared scenario module and a committed golden the node picker asserts independently. The goja side asserts the same golden, so neither runtime invokes the other at test time. * test(verifier): assert goja picker against the same cross-runtime golden Drop the node-subprocess coupling: the goja side now installs a stub __sanderlingHost__ with the fixed candidate list and asserts the committed golden, matching pkg/spec/test/parity.test.ts. * refactor(spec): rename pressKey generator export to pressKeys * refactor(spec): update barrel re-exports for pressKeys * test(spec): update pressKeys generator export name * docs(spec): rename pressKey generator to pressKeys * refactor(spec): extract samplerRng into shared sampler-rng module * feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText) * test(spec): cover fluent value generators determinism and chaining * refactor(bundler): inject globalThis trailer from spec named exports * refactor(bundler): reuse registration trailer in web bundler * test(bundler): cover named-export globalThis registration * feat(spec): add named() to Extracted handle type * feat(web-runtime): named() and cross-extractor read guard * feat(verifier): named() and cross-extractor read guard in goja * test(verifier): cross-extractor read guard and named() * test(web-runtime): export runtime and extractors for tests * test(web-runtime): named() and cross-extractor read guard * refactor(folio): drop manual globalThis trailer (bundler injects it) * refactor(folio): seed txn amounts via integers().between(1,500) * refactor(folio-web): drop manual globalThis trailer (bundler injects it) * fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs * refactor(folio-web): weight valid generators against edgeCaseText for names/amounts * refactor(folio-web): name extractors so violation witnesses are readable * fix(web-runtime): propagate extractor getter throws and unpoison locked global Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake. * test(spec): install fake runtime via defineProperty to survive locked global * test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors * feat(runner): add MaxSteps bound to Options * test(runner): MaxSteps stops after exactly N steps * test(driverpb): drop proto getter round-trip tautology * test(sidecar): drop stub-mode placeholder tautology tests * test(mock): drop default-field-value assertion test * test(ltl): drop Verdict.String tautology tests * refactor(runner): extract RenderSummary for snapshot testing * test(runner): golden snapshots for trace stream and violation summary * feat(web-runtime): capture uncaught errors into state.exceptions * test(integration): add throwing and counter web fixtures * test(integration): add specs for the web fixtures * test(integration): drive web fixtures through the real pipeline in headless Chrome * chore(make): add test-browser target for the Chrome-driven suite * ci: run the Chrome-driven browser suite in a separate job * refactor(test): relocate browser suite to test/browser * refactor(permissions): delete dead internal/permissions package * refactor(test): rename package to browser_test * refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets * chore(make): point test-browser at test/browser * docs(decisions): record internal/permissions deletion * refactor(doctor): use sidecarassets package * refactor(testrun): use sidecarassets package * fix(test): resolve testdata relative to browser_test.go * refactor(verifier): remove dead __sanderlingIndex compat alias * refactor(bundler): use encoding/json for JS string literals * docs(action-space): use vendor-neutral native driver wording * refactor(hierarchy): scrub backend tool name from comments * refactor(driver): scrub backend tool name from comments * refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver * refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap * refactor(chrome): implement DoubleTap as two taps with the gap * refactor(mock): record DoubleTap and DoubleTapSelector actions * refactor(runner): delegate double-tap to driver, drop gesture timing * test(runner): assert double-tap delegates to driver DoubleTap * docs(cmd): add package docs to CLI and developer tools * docs(driver): add package docs to driver interface and chrome backend * docs(driver): add package docs to mock and sidecar backends * docs(platform): add package docs to android and ios device prep * docs: add package docs to bundler and inspect * docs(ltl): add package doc to temporal logic evaluator * docs: add package docs to runner and testrun pipeline * docs: add package docs to trace and verifier * docs(sidecarassets): add package doc for embedded JAR loader * fix(chrome): add disable-dev-shm-usage so Chrome starts in CI * test(chrome): gate real-Chrome driver tests behind the browser tag * chore(make): run chrome driver tests in the browser job * fix(web-runtime): guard global error listeners for non-browser hosts The module registered window error/unhandledrejection listeners at top level, which threw under Node (the spec-api test runner) where globalThis.addEventListener is absent. Register only when the API exists; the real browser run is unaffected. * ci(browser): re-enable unprivileged user namespaces for headless Chrome ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged user namespaces stops headless Chrome from opening its DevTools socket even with --no-sandbox, surfacing as the driver's 'websocket url timeout'. Relax the sysctl for the job and add a direct launch check so a future breakage shows Chrome's own stderr rather than an opaque driver timeout. * ci(browser): pin stable Chrome for the driver tests setup-chrome's default latest pulled a dev Chromium (150) whose remote debugging socket never came up under chromedp, while plain --dump-dom worked. Pin the stable channel, which the driver is tested against. * feat(defaults): add scroll and rebalance action weights Use relative-integer weights (taps/typing co-primary 100, scrolls 50, swipes 25, doubleTaps 10); the picker normalizes by their total. Adds scrolls to defaultActions as a first-class reveal behavior. * feat(defaults): trim scroll action weight wiring * fix(build): point sidecar jar ignore and embed paths at sidecarassets * test(defaults): drop stale longPresses re-export assertion longPresses is opt-in vocabulary, no longer re-exported from defaults/actions.ts since e0d3b20; its builtin resolution is already covered by api.test.ts. Trim the defaults test to scrolls, which is an actual default export. * fix(chrome): raise DevTools websocket read timeout to 60s Chrome cold-start on a loaded CI runner can exceed chromedp's 20s default for reading the DevTools websocket URL, flaking the browser tests with "websocket url timeout reached". Give launch more headroom. |
||
|
|
2a1b263b8c |
fix: WDA startup flakiness - warmup + connection drop message (#40)
* feat(ios): add simulator management package * feat(testrun): add iOS platform path (simctl launch + direct TCP) * feat(cli): add --ios-device flag and IosDevice option * feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser) * feat(folio-ios): wire SanderlingIos.start() in MainViewController * feat(folio-ios): add test-ios justfile recipe * fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings * fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects * feat(proto): add env map to LaunchRequest * feat(driver): add env param to Launch interface + all implementations * feat(testrun): launch iOS app via XCTest with env vars instead of simctl * feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through * feat(sidecar): wire env map in DriverService + IosDriverBackend in Main * fix(sidecar): use LocalIOSDevice (WDA+simctl) + stop before relaunch * fix(sidecar): include exception type in gRPC error description * fix(sidecar): pick free WDA port instead of hardcoded 9100 Use SocketUtils.nextFreePort to pick a free port in the 22000-23000 range rather than hardcoding 9100, which only worked if a previous WDA session left a listener there. * fix(sdk-ios): check semaphore wait result and throw on snapshot timeout dispatch_semaphore_wait returns nonzero on timeout; ignoring the return value caused pauseAndSnapshot to silently return an empty map, sending a garbage empty STATE frame to the host. Now throws so the agent loop reconnects instead. * fix(folio-ios): register snapshot extractors before starting agent SanderlingIos.start() was called before the snapshot objects were initialized, so a PAUSE message arriving early produced an empty snapshot. Move start() to after all extractors are registered. * chore(ios): remove dead LaunchApp function LaunchApp had no callers since |
||
|
|
88db0cbea8 |
docs: web platform + clean URLs + dark/light mode (#37)
* chore(docs): replace d2 diagram pipeline with mermaid Remove docs/_diagrams/ and d2 build step from Makefile. The HTML template already initialises mermaid.js; diagrams are now inline code fences rendered client-side. * docs(architecture): add mermaid diagram + web/CDP platform docs Replace SVG img tag with inline mermaid flowchart showing both native (Maestro sidecar + in-app SDK) and web (Chrome CDP) paths. Update Processes, Transports table, and per-step cycle sections. * docs(design-principles): update principles 1-4 for web platform Principles 1, 2, 3, and 4 referenced Maestro and native-only concepts. Add web/CDP context and update driver-is-an-interface to name both sidecar and chrome implementations. * docs(manual): add web prerequisites and folio-web example Update --platform flag to list android, ios, web. Add web prerequisites section (Chrome, no SDK needed) and folio-web quick-start to getting-started. * chore(gitignore): untrack inspect dist/index.html build artifact index.html is regenerated by vite on every build with a new content hash, making it permanently dirty. Only .gitkeep is needed for //go:embed to compile on a fresh checkout. Also remove duplicate dist/* line and stale d2 diagram ignore entries. * feat(docs): click-to-zoom for mermaid diagrams * docs(architecture): change diagram layout from LR to TB * docs(getting-started): link npm and Maven Central package headers * update docs root * docs(spec): rewrite npm package README Update usage example to current API, drop stale version-compatibility and license sections. * build(docs): output pages as pagename/index.html for clean URLs Split DOCS_OUT into INDEX_OUT (index.md files stay as index.html) and PAGE_OUT (all other pages become pagename/index.html). The __ROOT__ depth computation already handles the extra directory level correctly. * chore(docs): update sidebar links to directory-style URLs * docs: update cross-links from .html to directory-style paths * ci(docs): remove d2 install step * feat(docs): dark/light mode toggle Add theme toggle button (top-right, fixed). Persists preference in localStorage; falls back to prefers-color-scheme. Flash-free via inline script in <head> that sets data-theme before first paint. * fix(docs): fix inspect image path broken by directory URL restructure * feat(docs): click-to-fullscreen for all article images * fix(inspect): allow AssetsFS override in ServerOptions; drop unused request param from serveIndex * fix(inspect): use in-memory FS in tests so TestAssets_FallbackToIndexHTML passes without web build |
||
|
|
af7b7e27b0 |
feat(folio-web): web sample app + CDP spec tests (#36)
* fix(runner): allow nil connection for web platform * fix(testrun): skip SDK handshake for web platform * feat(folio-web): add React/Vite web sample app * feat(folio-web): add sanderling spec * fix(chrome): use InsertText for multi-char text input * feat(hierarchy): add Screen field populated from sanderling-screen attr * fix(chrome): auto-detect viewport from CSS vars, fix InputText accumulation, expose route as screen * fix(runner): fall back to hierarchy root screen when snapshot screen is empty * fix(folio-web): broaden loggedIn extractor to all authenticated pages |
||
|
|
eed99e58aa |
refactor: code organization cleanup (#35)
* chore: fix gitignore + decisions doc after web->inspect-ui rename Update web/ references to inspect-ui/ in .gitignore and Makefile. Add decisions.md tracking architectural decisions from code-org discussion. * refactor: rename pkg/spec-api to pkg/spec Aligns the directory name with the npm package name @sanderling/spec. Updates Makefile, package.json directory field, and resolveSpecAPIPath. * refactor(verifier): split bindings.go into types.go + bindings.go Move shared public types (Action, ActionKind, LogEntry, Exception) to types.go. bindings.go retains internal JS runtime wiring only. * refactor(inspect): split runs.go into runs.go, runs_cache.go, runs_decode.go runs.go: types (RunSummary, StepSummary, RunDetail, Run) and Scan. runs_cache.go: Cache type, Open/Step/Detail methods, parseRun, scanSteps. runs_decode.go: readMeta, tallyTrace, decodeStepSummary, validRunID. * refactor: move android_env.go to internal/android/ Extracts Android device/AVD/adb logic into internal/android package. Exports EnsureDevice, AdbReverse, AdbReverseRemove, EnvWithAndroidPlatformTools, AdbBinary. Moves tests to internal/android/android_test.go. cmd/sanderling becomes a thin caller. * refactor: extract test pipeline to internal/testrun/ runTestPipeline logic moves to testrun.Execute. buildDriver, resolveSpecAPIPath, pickFreePort, and the progress logger move to internal/testrun/. cmd/sanderling/test_run.go becomes a thin adapter. Tests follow their code. * ci: update workflow paths after pkg/spec-api -> pkg/spec rename |
||
|
|
08288202cb |
docs+install: post-rename docs polish, install script, d2 architecture diagram (#26)
* docs: add sanderling bird artwork to README and docs index * chore: add one-line install script for macOS and Linux Detects os/arch, resolves latest (or pre-)release, verifies sha256, and installs the binary into $HOME/.sanderling/bin. * docs: use the one-line installer in getting-started Replaces the broken `<version>` placeholder snippets with the install.sh one-liner. * docs: drop filler line under the install one-liner * docs: rename index heading to Sanderling Manual * docs(style): adopt JetBrains Mono and uppercase brand mark * docs(getting-started): use justfile flow for folio sample * docs(writing-specs): drop 'coming soon' notes for eventually and implies * docs(inspect): rewrite layout for tabbed state panels and metrics chart * docs(inspect): add UI screenshot * docs: add Inspect to sidebar and index * docs: restore mermaid bootstrap script in page template * docs(style): constrain article images to content width * docs(inspect): drop layout prose, keep what the screenshot doesn't show * docs(architecture): add d2 source for architecture diagram Replaces the in-page mermaid block with a d2-rendered SVG. Generated outputs (svg/png) stay out of git; only the .d2 source is checked in. * build(docs): render d2 diagrams into build/site/_assets/diagrams * docs(architecture): swap mermaid block for rendered d2 svg * ci(docs): install d2 before building the site * docs(architecture): tighten layout and reroute label-crossing edges Flip device/sidecar order so trace writer drops cleanly to runs/ without cutting through the JVM cell, right-align the inspect row via a pad column, and tune grid gaps to keep gRPC and Unix socket labels off the SANDERLING boundary. |
||
|
|
8ccf95c1cf |
refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling
Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.
* chore(proto): regenerate driverpb after module path rename
The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.
* refactor: rename CLI binary uatu -> sanderling
Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.
* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk
Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.
* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar
Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.
* refactor: rename Uatu API surface -> Sanderling
- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
__uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.
* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling
Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.
* chore(build): rename gradle property + rootProject.name uatu -> sanderling
- Renames the uatu.version gradle property and all its -P references in
Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.
* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1
Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.
* refactor: rename npm package @uatu/spec -> @sanderling/spec
Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.
* docs: rename uatu -> sanderling in README, docs, and URLs
- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
internal/inspect/server.go.
* refactor: rename remaining internal uatu strings -> sanderling
- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
|
||
|
|
13bb2feb82 |
feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI
Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.
* test(trace): cover EndedAt + new step fields round-trip
* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()
Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.
* feat(runner): stamp residuals, hierarchy, selector targets, ended_at
Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.
* feat(inspect): scaffold embed dist for SPA assets
Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).
* chore(web): ignore web/ build output in root .gitignore
* chore(web): add bun + vite + vitest scaffold config
* feat(web): monochrome design tokens, typography, app shell CSS
* chore(web): placeholder for self-hosted JetBrains Mono fonts
* feat(web): index.html entry with style links and root mount
* feat(web): typescript types mirroring run/step trace schema
* feat(web): typed fetchers for runs/steps/screenshots
* feat(web): App shell with router and run/step routes
* feat(inspect): runs scan, lazy step parse, mtime-aware cache
* feat(web): RunList route with table, loading, and error states
* feat(web): RunDetail route shell with three placeholder panels
* feat(inspect): fsnotify-backed runs watcher with debounce
* fix(web): use jest-dom/vitest entry so matchers register
* test(web): cover listRuns happy path and error response
* test(web): render RunList with mocked fetch and assert row
* chore(web): commit bun lockfile
* feat(inspect): http handlers for runs/steps/screenshots/SSE
* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy
* feat(cmd): add 'uatu inspect' subcommand
* fix(web): align TS types with snake_case wire format
Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.
* feat(web): add ActionList panel for run-detail step navigation
* feat(web): add SnapshotTable panel with diff highlighting
Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.
* test(web): cover SnapshotTable rendering and diff behavior
Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.
* feat(web): add Screenshot panel with bounds and tap overlays
Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.
* fix(web): guard scrollIntoView call for jsdom compatibility
* test(web): cover Screenshot panel rendering and overlays
* test(web): cover ActionList rendering, selection, keyboard, and markers
* feat(web): add ExceptionsPanel component
* test(web): add ExceptionsPanel tests
* feat(web): add Timeline panel with property swimlanes
Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.
* test(web): cover Timeline empty state, cells, status, click, highlight
* feat(web): add ResidualNode recursive AST renderer
* test(web): cover ResidualNode operators, predicate, and error chip
* feat(web): add ViolationsPanel with status badges and jump button
* test(web): cover ViolationsPanel rows, status grouping, and jump button
* test(web): register testing-library cleanup globally
All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.
* feat(web): hooks for url/keyboard/theme/sse
* feat(web): wire all panels into run-detail with phone-dominant grid
ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.
* test(web): add three reference run fixtures (clean, violation, exception)
* build: web targets in Makefile + bun in CI; docs(inspect)
- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.
* feat(runner): capture a screenshot per step
The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.
* feat(inspect): include action_label in StepSummary
Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.
* test(inspect): accept either #app or #root in SPA shell fallback
* feat(web): render action_label and screen in ActionList rows
Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.
* feat(sidecar): implement screencap for android driver backend
Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.
* feat(proto): add Metrics RPC for per-step CPU and memory capture
* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes
* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status
* feat(runner): capture metrics + before/after screenshots per step
Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.
* fix(runner,sidecar): measure CPU across step via /proc stat delta
'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.
* feat(web): add Metrics type for per-step cpu and memory
* refactor(web): replace --accent-change with --accent-positive token
* refactor(web): recolor chip-progress as neutral outlined chip
* refactor(web): use neutral border for changed snapshot rows
* feat(web): add MetricsChart panel with HEAP and CPU lanes
SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.
* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows
Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.
* fix(runner): stop copying Tap selector into action.text
The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.
* fix(web): use text-muted for swipe arrow after accent-change removal
* fix(web): snapshot values truncate with ellipsis + title tooltip
Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.
* feat(web): state-before/after columns + metrics chart at bottom
RunDetail now renders a four-column grid:
actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.
* fix(web): skip zero-value ticks + add exception markers to metrics
HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.
* fix(web): let action body column shrink below its content
Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.
* feat(web): bigger state screenshots + single properties row
Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.
* feat(web): add minimal Tabs component
Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.
* feat(web): fold timeline into MetricsChart as STEPS lane
Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.
* refactor(web): tabbed state cards, drop standalone Timeline panel
State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.
* feat(web): lock app shell to 100vh with no page scroll
html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.
* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)
Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).
* feat(web): ViolationsPanel supports violationsOnly filter
* feat(web): ActionList arrow-key nav + listbox semantics + smaller font
Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.
* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets
Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.
* feat(web): fourth 'Violations' tab + wider actions + shorter metrics
Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.
* feat(web): compact RunDetail layout using 1px borders instead of panel padding
* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis
Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.
* fix(web): RunList rows no longer stretch to fill viewport height
Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.
* misc changes
* fix(web): hoist useState above early return in MetricsChart
Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.
* fix(web): subscribe to named SSE event instead of 'message'
Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.
* fix(inspect): unsubscribe SSE clients on disconnect
Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.
Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.
* fix(trace): rename resolvedBounds/tapPoint to snake_case
Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.
* chore(web): drop vitest and remove UI tests from CI
No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).
Fixes CI failure where `vitest run` exits 1 with no test files.
* chore(make): dedupe sidecar embed and drop recursive make
Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
|
||
|
|
e62319e916 |
docs: pandoc-based site and v0.1.0 groundwork (#5)
* chore(prose): remove em-dashes from config files * chore(prose): remove em-dashes from android sdk config * docs(spec-api): remove em-dash from README * fix(doctor): reword sidecar-jar error without em-dash * test(sidecar): reword assertion message without em-dash * docs: add CLAUDE.md with project conventions * build: add docs target for pandoc site * docs(site): add pandoc template and stylesheet * docs(site): add pandoc build script * docs(site): add landing pages * docs(manual): add getting-started * docs(manual): add writing-specs * docs(manual): add runs * docs(manual): add cli reference * docs(dev): add design principles * docs(dev): add architecture * ci: deploy docs site to github pages * docs: rewrite README as entry point to docs site |
||
|
|
0570719e6f |
Publish pipeline: goreleaser + Maven Central + npm (#1)
* feat(cli): add Version var and version subcommand * build(gradle): introduce uatu.version property for lockstep releases * build(sdk-android): swap GitHub Packages for vanniktech Maven Central plugin * build(spec-api): make package publish-ready for npm * ci(release): add goreleaser config for cross-platform uatu CLI builds * ci: add ci and release GitHub Actions workflows * ci: restrict ci.yml to PR + workflow_dispatch (no direct push to master) * docs(release): add local release targets, env example, and install docs * build(sdk-android): make signAllPublications conditional on signing key * ci(release): stage sidecar JAR at embed path before go build * chore(spec-api): regenerate package-lock for updated package.json * ci: install protoc-gen-go plugins before buf generate * ci: bump Node to 22 (required for --experimental-strip-types) |
||
|
|
6e4c832678 |
chore: stop tracking sidecar JAR (build artifact)
make uatu copies the real fat JAR into assets/ before go build -tags withsidecar. Keeping that path tracked was the root cause of the 130 MB push rejection. |
||
|
|
0c68e72341 |
feat(sdk-android): publish to GitHub Packages at maven.pkg.github.com
Coordinates: dev.uatu:sdk-android:0.0.1. Credentials read from GH_TOKEN/GH_USERNAME (or GITHUB_TOKEN/GITHUB_ACTOR for CI). .env is gitignored so local tokens stay out of git. Consumers add the maven repo + debugImplementation in their build.gradle and we're done. |
||
|
|
5988e93057 |
chore(spec-api): scaffold @uatu/spec package metadata
TypeScript 5.6 with strict mode, bundler module resolution. Tests run via node --test --experimental-strip-types so we don't need vitest or a build step. Adds node_modules to .gitignore. |
||
|
|
32cd729f9d | chore: add .gitignore for gradle, IDE, and run outputs |