Commit Graph
22 Commits
Author SHA1 Message Date
pj 8d29b2906f test(runner): cover transitional step skips verifier and clean control 2026-05-31 15:36:04 +05:30
pj 5c58610181 test(runner): assert step uses Snapshot, not raw hierarchy/screenshot
TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.
2026-05-31 15:27:45 +05:30
pj a1390aecdd test(runner): cover startup gate waiting for app window to draw 2026-05-31 14:35:04 +05:30
pj 16a9136b28 test(runner): cover startup foreground gate and back-press 2026-05-31 12:12:58 +05:30
pj ddec95a2c5 feat(runner): re-fetch on transitional hierarchy capture
Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.

fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.
2026-05-31 12:11:29 +05:30
pj fb69c1b774 refactor(runner): one concurrent screenshot per step
Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.
2026-05-30 21:48:21 +05:30
pj 60c4ef7458 refactor(runner): emit onset-only violations to trace and summary
Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.
2026-05-30 16:32:17 +05:30
pj 87834caf96 test(doubleTap): cover constructor, verifier round-trip, and runner dispatch 2026-05-30 11:05:58 +05:30
pj dff91095b8 feat(runner): relaunch app when foreground escapes during exploration 2026-05-28 11:23:50 +05:30
pj b23fb0c723 feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag

Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.

* feat(testrun): add Preflight() before sidecar/driver setup

Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.

* refactor(chrome): split tag (HTML name) from class (CSS classList)

Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.

* feat(chrome): translate legacy string selectors to CSS/XPath

TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.

* feat(trace): add WriteHTML + Step.HTMLAvailable

Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.

* feat(driver): add WebDriver capability + chrome implementation

WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.

* feat(verifier): OverrideExtractorValues for V8-driven extractors

Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.

* feat(spec): add WebState + camelCase attribute aliases

WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.

* feat(runner): per-tick HTML capture for WebDriver-capable drivers

Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.

* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>

Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.

* feat(bundler): BundleWeb + V8-side runtime shim

web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.

* feat(runner): V8 extractor overrides + V8 action source for WebDriver

When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.

* feat(testrun): bundle + install web runtime when platform=web

BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.

* feat(inspect-ui): hierarchy + html panels in run detail

HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.

* fix(folio-web): drop aria-label data-carrier abuse

Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.

* chore: rebuild inspect-ui dist + folio-web .gitignore

Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.

* revert(trace): drop WriteHTML + Step.HTMLAvailable

Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.

* revert(runner): drop per-tick HTML capture

Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.

* revert(driver): drop WebDriver.Document

Document was only consumed by the runner's HTML capture which is gone.

* revert(inspect): drop /html route

Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.

* revert(inspect-ui): drop htmlUrl + html_available type

API surface no longer needs the HTML route; Step.html_available has no
producer.

* revert(inspect-ui): drop HtmlPanel + html tab

Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.

* test(inspect-ui): drop htmlUrl test, add @types/bun

Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.

* chore: rebuild inspect-ui dist without HtmlPanel

Embedded SPA bundle no longer ships the iframe HTML viewer.

* fix(web-runtime): retry action resolution + implement taps/swipes

V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.

* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys

Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.

* chore(folio-web): drop swipes from action root

Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.

* fix(inspect-ui): correct HierarchyPanel CSS variable names

Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.

* fix(chrome): correct PressKey mappings to chromedp/kb constants

Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.

Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.

* fix(cli): -h/--help exits 0 instead of error code

parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.

Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.

* fix(chrome): harden cssEscape for control chars + use [class~=]

Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.

Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.

* fix(web-runtime): use CSS.escape and validate tag-name selectors

The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.

The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.

Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.

* fix(chrome): validate attribute name in unknown-prefix branch

A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.

* fix(selectors): emit valid XPath 1.0 string literals via concat()

Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.

Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.

* fix(runtime): surface unresolved action targets instead of dropping silently

serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.

Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.

* fix(runner): use errgroup-bound ctx so siblings cancel on failure

The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.

* fix(chrome): propagate caller ctx cancellation to CDP calls

InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.

Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.

* fix(verifier): tolerate out-of-range override indices

A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.

Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.

* test(verifier): cover object-shaped extractor overrides

Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.

* fix(web-runtime): lock global runtime hooks against page shadowing

AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.

* perf(web-runtime): cache randomTap candidate DOM scan per tick

The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.

Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.

* fix(web-runtime): cap sanitize recursion to prevent stack overflow

State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.

* fix(web-runtime): enforce pressKey allowlist in factory

The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.

* chore(chrome): drop dead bundleSource/bundleMu

bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.

* fix(chrome): use strconv.Atoi for extractor key parsing

fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.

* fix(doctor): raise per-check timeout to 15s for chromium launch

5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.

* fix(runner): trust V8 coordinates for InputText, even at origin

resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.

Add applyAction tests covering both the typical web case and the (0,0)
edge case.

* test(bundler): lock down deterministic output across builds

The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
2026-05-03 11:21:21 +07:00
pj 776becdf4b Remove in-app SDK (#43)
* chore: delete internal/agent package

* chore(build): remove sdk-android from gradle settings

* chore(makefile): remove sdk-android targets

* chore(ci): remove release-android job from release workflow

* chore(folio): remove sdk-android dependency

* chore(folio): remove SDK initialization from FolioApplication

* chore(folio): delete snapshot extractor files

* feat(folio): add balance to account card content description

* feat(folio): add hierarchy content descriptions to LedgerScreen

* refactor(folio): rewrite spec.ts to use ax extractors

* docs: remove in-app SDK from README

* feat(folio): add focused_input indicator to App

* docs: remove in-app SDK from index

* refactor(runner): remove agent SDK connection and snapshot step

* test(runner): update tests for SDK removal

* docs: remove Android SDK section from getting-started

* refactor(testrun): remove agent SDK connection setup

* docs: remove snapshots from writing-specs

* docs: remove in-app SDK from architecture doc

* docs(folio): update README for SDK removal

* docs: update per-step cycle diagram in architecture doc

* fix(folio): detect screens from unique element presence, not id: selectors

testTag() in Compose is not exposed as resource-id without testTagsAsResourceId.
Use desc: selectors for elements unique to each screen instead of id: path queries.

* feat(folio): add screen root contentDescription for scoped ax selection

Each screen root gets semantics { contentDescription = "ScreenName" } so
sanderling specs can scope element lookups through the screen: desc:LoginScreen > desc:login_submit.

* fix(folio): scope all ax selectors through screen root nodes

Use desc:ScreenName > desc:element path queries so every selector is
rooted at the screen level. focusedInput stays unscoped since it lives
in the app root, outside any screen.

* fix(folio): guard newAccountBalanceIsZero against navigation false positives

Scoped selectors return [] when not on HomeScreen so accounts vanish and
reappear as apparently-new on each visit. Skip the check when prev was empty.

* chore(folio): link @sanderling/spec to local pkg/spec for IDE type checking

* feat(spec): add desc, class, clickable, enabled, checked, focused, selected to AccessibilityElement

Runtime fields set by the verifier were missing from the TypeScript type,
causing linting errors on el.desc and related accesses in specs.

* chore(folio): switch to bun, add tsconfig.json for IDE type checking

- Remove package-lock.json, add bun.lock
- Add tsconfig.json so VSCode resolves @sanderling/spec types
- Fix parseAccount/parseLedgerRow to accept string | undefined
2026-04-25 20:04:29 +07:00
pj af7b7e27b0 feat(folio-web): web sample app + CDP spec tests (#36)
* fix(runner): allow nil connection for web platform

* fix(testrun): skip SDK handshake for web platform

* feat(folio-web): add React/Vite web sample app

* feat(folio-web): add sanderling spec

* fix(chrome): use InsertText for multi-char text input

* feat(hierarchy): add Screen field populated from sanderling-screen attr

* fix(chrome): auto-detect viewport from CSS vars, fix InputText accumulation, expose route as screen

* fix(runner): fall back to hierarchy root screen when snapshot screen is empty

* fix(folio-web): broaden loggedIn extractor to all authenticated pages
2026-04-22 23:59:30 +07:00
pj d5b6f6ccaf feat: Maestro driver integration + DeviceDriver architecture (#32)
* proto(driver): drop launcher_activity from LaunchRequest

* refactor(driver): rename Driver to DeviceDriver, drop launcherActivity from Launch

* refactor(driver): rename package maestro to sidecar

* feat(driver): add ChromeDriver backed by chromedp

* feat(hierarchy): replace XML parser with TreeNode JSON parser

* refactor(driver): update mock and runner to DeviceDriver, drop launcherActivity

* feat(runner): add platform routing for web vs sidecar

* test(verifier): update hierarchy fixtures from XML to TreeNode JSON

* feat(sidecar): extract readLogcat/readProcMetrics; add MaestroDriverBackend

- Extract readLogcat() and readProcMetrics() as internal package-level
  helpers parameterised by serial
- Add MaestroDriverBackend wiring maestro-client AndroidDriver
- Drop launcherActivity from DriverBackend interface and StubDriverBackend
- Add maestro-utils and micrometer-core as explicit compile deps

* refactor(sidecar): inject DriverBackend into DriverService; drop launcherActivity

- Remove serial and default backend from DriverService constructor
- Make backend a required parameter
- Drop launcherActivity from launch RPC handler
- Create MaestroDriverBackend in Main.kt when platform is android

* test(sidecar): update tests for dropped launcherActivity and required backend

* chore(sidecar): remove unnecessary micrometer-core direct dependency
2026-04-22 19:02:39 +07:00
pj b667abbbca perf: reduce per-step iteration time (~3.9s to ~2.3s) (#29)
* perf(sidecar): use exec-out + tmpfs for hierarchy dump

Avoids FUSE overhead on /sdcard and shell startup cost by using
exec-out with /data/local/tmp. Saves ~100ms per hierarchy fetch.

* perf(sidecar): replace Thread.sleep waitForIdle with real idle detection

Poll `dumpsys window -a` for mAnimating=true every 50ms instead of
blindly sleeping. Breaks early when device is idle, saving 500-800ms
per step since most settle in <200ms after an action.

* test(sidecar): add idle detection parsing tests

* perf(runner): parallelize hierarchy, metrics, and logs fetch

Run fetchHierarchy, captureMetrics, and collectLogs concurrently via
errgroup so metrics+logs (~150ms) hide behind the hierarchy fetch
(~2s) instead of running serially.

* perf(runner): pipeline post-action screenshot with next step

Defer the post-action screenshot from step N and run it concurrently
with step N+1's hierarchy/metrics/logs fetch. Saves ~335ms per step
by hiding screenshot latency behind the hierarchy fetch.

* test(runner): add tests for parallel fetch and pipelined screenshots

Verify that hierarchy, metrics, and logs are all called per step.
Verify post-action screenshots are written with correct step indices
when pipelined, including the final flush after the loop.

* perf(sidecar): grep mAnimating on-device instead of pulling full dump

The full `dumpsys window -a` output is ~88KB per poll. Running
grep on-device transfers only a count byte, cutting per-poll
overhead from ~63ms to ~56ms and eliminating 88KB of ADB transfer.
2026-04-22 17:43:27 +07:00
pj 8ccf95c1cf refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling

Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.

* chore(proto): regenerate driverpb after module path rename

The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.

* refactor: rename CLI binary uatu -> sanderling

Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.

* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk

Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.

* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar

Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.

* refactor: rename Uatu API surface -> Sanderling

- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
  spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
  __uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.

* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling

Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.

* chore(build): rename gradle property + rootProject.name uatu -> sanderling

- Renames the uatu.version gradle property and all its -P references in
  Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.

* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1

Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.

* refactor: rename npm package @uatu/spec -> @sanderling/spec

Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.

* docs: rename uatu -> sanderling in README, docs, and URLs

- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
  'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
  internal/inspect/server.go.

* refactor: rename remaining internal uatu strings -> sanderling

- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
  localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
2026-04-21 11:57:49 +07:00
pj 13bb2feb82 feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI

Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.

* test(trace): cover EndedAt + new step fields round-trip

* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()

Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.

* feat(runner): stamp residuals, hierarchy, selector targets, ended_at

Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.

* feat(inspect): scaffold embed dist for SPA assets

Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).

* chore(web): ignore web/ build output in root .gitignore

* chore(web): add bun + vite + vitest scaffold config

* feat(web): monochrome design tokens, typography, app shell CSS

* chore(web): placeholder for self-hosted JetBrains Mono fonts

* feat(web): index.html entry with style links and root mount

* feat(web): typescript types mirroring run/step trace schema

* feat(web): typed fetchers for runs/steps/screenshots

* feat(web): App shell with router and run/step routes

* feat(inspect): runs scan, lazy step parse, mtime-aware cache

* feat(web): RunList route with table, loading, and error states

* feat(web): RunDetail route shell with three placeholder panels

* feat(inspect): fsnotify-backed runs watcher with debounce

* fix(web): use jest-dom/vitest entry so matchers register

* test(web): cover listRuns happy path and error response

* test(web): render RunList with mocked fetch and assert row

* chore(web): commit bun lockfile

* feat(inspect): http handlers for runs/steps/screenshots/SSE

* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy

* feat(cmd): add 'uatu inspect' subcommand

* fix(web): align TS types with snake_case wire format

Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.

* feat(web): add ActionList panel for run-detail step navigation

* feat(web): add SnapshotTable panel with diff highlighting

Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.

* test(web): cover SnapshotTable rendering and diff behavior

Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.

* feat(web): add Screenshot panel with bounds and tap overlays

Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.

* fix(web): guard scrollIntoView call for jsdom compatibility

* test(web): cover Screenshot panel rendering and overlays

* test(web): cover ActionList rendering, selection, keyboard, and markers

* feat(web): add ExceptionsPanel component

* test(web): add ExceptionsPanel tests

* feat(web): add Timeline panel with property swimlanes

Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.

* test(web): cover Timeline empty state, cells, status, click, highlight

* feat(web): add ResidualNode recursive AST renderer

* test(web): cover ResidualNode operators, predicate, and error chip

* feat(web): add ViolationsPanel with status badges and jump button

* test(web): cover ViolationsPanel rows, status grouping, and jump button

* test(web): register testing-library cleanup globally

All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.

* feat(web): hooks for url/keyboard/theme/sse

* feat(web): wire all panels into run-detail with phone-dominant grid

ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.

* test(web): add three reference run fixtures (clean, violation, exception)

* build: web targets in Makefile + bun in CI; docs(inspect)

- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
  install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
  vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.

* feat(runner): capture a screenshot per step

The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.

* feat(inspect): include action_label in StepSummary

Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.

* test(inspect): accept either #app or #root in SPA shell fallback

* feat(web): render action_label and screen in ActionList rows

Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.

* feat(sidecar): implement screencap for android driver backend

Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.

* feat(proto): add Metrics RPC for per-step CPU and memory capture

* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes

* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status

* feat(runner): capture metrics + before/after screenshots per step

Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.

* fix(runner,sidecar): measure CPU across step via /proc stat delta

'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.

* feat(web): add Metrics type for per-step cpu and memory

* refactor(web): replace --accent-change with --accent-positive token

* refactor(web): recolor chip-progress as neutral outlined chip

* refactor(web): use neutral border for changed snapshot rows

* feat(web): add MetricsChart panel with HEAP and CPU lanes

SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.

* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows

Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.

* fix(runner): stop copying Tap selector into action.text

The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.

* fix(web): use text-muted for swipe arrow after accent-change removal

* fix(web): snapshot values truncate with ellipsis + title tooltip

Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.

* feat(web): state-before/after columns + metrics chart at bottom

RunDetail now renders a four-column grid:
  actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.

* fix(web): skip zero-value ticks + add exception markers to metrics

HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.

* fix(web): let action body column shrink below its content

Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.

* feat(web): bigger state screenshots + single properties row

Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.

* feat(web): add minimal Tabs component

Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.

* feat(web): fold timeline into MetricsChart as STEPS lane

Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.

* refactor(web): tabbed state cards, drop standalone Timeline panel

State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.

* feat(web): lock app shell to 100vh with no page scroll

html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.

* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)

Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).

* feat(web): ViolationsPanel supports violationsOnly filter

* feat(web): ActionList arrow-key nav + listbox semantics + smaller font

Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.

* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets

Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.

* feat(web): fourth 'Violations' tab + wider actions + shorter metrics

Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.

* feat(web): compact RunDetail layout using 1px borders instead of panel padding

* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis

Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.

* fix(web): RunList rows no longer stretch to fill viewport height

Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.

* misc changes

* fix(web): hoist useState above early return in MetricsChart

Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.

* fix(web): subscribe to named SSE event instead of 'message'

Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.

* fix(inspect): unsubscribe SSE clients on disconnect

Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.

Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.

* fix(trace): rename resolvedBounds/tapPoint to snake_case

Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.

* chore(web): drop vitest and remove UI tests from CI

No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).

Fixes CI failure where `vitest run` exits 1 with no test files.

* chore(make): dedupe sidecar embed and drop recursive make

Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
2026-04-21 11:32:17 +07:00
pj 323878c34a fix(verifier): don't crash on throwing JS predicates (#21)
* fix(verifier): don't crash on throwing JS predicates

formulaThunk used to panic whenever goja returned an error from a
predicate callable, and nothing on the LTL -> runner path recovered, so
a malformed spec (e.g. a property whose body throws or touches an
undefined field) would kill the verifier process.

Latch the first error on formulaState, return false so LTL marks the
property violated, and expose PredicateError(name) that walks the
property's formula-spec tree and surfaces the latched cause.

* fix(runner): log predicate errors alongside violations

For each violated property, surface the verifier's latched predicate
error via logger.Warn so operators can distinguish a genuine false
verdict from a malformed spec. Add a runner-level test asserting that a
throwing predicate no longer crashes the run and that the error message
appears in the log.
2026-04-20 16:55:10 +07:00
pj 69002b07b1 fix(runner,sample-app): surface silent errors and demo a failing property (#15)
* fix(runner): surface non-deadline WaitForIdle errors

Previously the WaitForIdle return value was discarded entirely, hiding
real driver failures (gRPC transport errors, sidecar crashes) behind
the expected deadline-exceeded case. Log non-deadline errors so they
are visible without changing control flow.

* chore(sample-app): drop unused uptime_millis extractor

Registered in SampleApplication but never consumed by spec.ts.

* fix(sample-app): drop trivial appIsRunning property

app_state was hardcoded to 'running' so the property was a tautology
that could never fail. Removing both the extractor and the property
is the simplest fix; demo-grade properties that can fail land next.

* feat(sample-app): add Reset button that zeroes clickCount

Pairs with the next commit's tap-reset action so the fuzzer can
violate clickCountNeverDecreases and demonstrate uatu actually
finding a property violation.

* feat(sample-app): add tap-reset action to exercise Reset button

Weighted at 10/122, fuzzer reaches it within a short run. Pairs with
the Reset button to demonstrate uatu detecting the
clickCountNeverDecreases violation.

* fix(runner): filter WaitForIdle errors via context state, not errors.Is

errors.Is(err, context.DeadlineExceeded) misses gRPC's wrapped
status.DeadlineExceeded, so every step under the maestro driver
logged a spurious warning. Check idleCtx.Err() instead — captures
both deadline-fired and parent-canceled cases regardless of how the
driver wraps them.

* chore(sample-app): tune action weights so demo violates in ~30s

Prior weights left tap-reset rare enough that short demo runs missed
the violation by chance. Bumped to 30/107, with typeUsername reduced
since username noise doesn't help exercise clickCount.

* refactor(runner): route warnings through slog

Adds Options.Logger (defaults to slog.Default()) and converts the
three warning sites that were using fmt.Printf. Progress line stays
on Printf since it's user-facing UI, not a log. Makes the warnings
testable via a capturing handler.

* test(runner): assert WaitForIdle driver errors are logged

Captures slog output via TextHandler into a buffer and asserts the
warning message + injected error text appear when the mock driver
returns a non-context error from WaitForIdle. Guards against a
regression of the silent-error swallow.
2026-04-18 18:48:25 +07:00
pj d0578dbaaa fix(runner): warn on malformed screen snapshot (#13)
* fix(runner): warn on malformed screen snapshot

screenFromSnapshot swallowed json.Unmarshal errors, so a non-string
screen value silently became "" in the step log and trace while the
verifier still saw the raw JSON. Return the error and warn at the
call site, matching the hierarchy warning pattern.

* docs: clarify --avd is optional for uatu test

The CLI accepts --avd as an empty-string default (cmd/uatu/main.go:49)
and only requires it when no device is connected and multiple AVDs
exist (cmd/uatu/android_env.go:63). Docs and examples that showed it
as required or always-passed were misleading.
2026-04-18 17:14:31 +07:00
pj 16e55086d8 fix(runner): surface focus-tap errors in InputText (#12)
* fix(runner): surface focus-tap errors in InputText action

A failed Tap/TapSelector before InputText was swallowed, so text typed
into the wrong field (or no field) still reported success. Return the
error so the step fails explicitly.

* feat(sample-app): add username EditText and snapshot

Gives the spec a real EditText target (content-desc: username_field)
so the InputText action path can be exercised end-to-end. The typed
value is mirrored into MainActivity.username and surfaced as the
"username" snapshot for spec assertions.

* feat(sample-app): exercise InputText action against username field

Adds typeUsername action and usernameNeverShrinks property to the
sample spec, and extends the integration test to assert the bundled
spec emits an InputText(desc:username_field, "alice") action and that
the property correctly violates when a snapshot reports a shorter
string.
2026-04-18 16:51:35 +07:00
pj e7b3e2ba9c refactor(runner): caller manages app launch/terminate
Removes Launch + Terminate from runner.Run so the CLI can launch
the app first, wait for the SDK to connect, then start the loop.
The previous shape forced runner to launch internally which fought
with the SDK-must-be-connected-first ordering.

BundleID/ClearState fields go away too since runner no longer
launches; the CLI keeps them on its testOptions struct.
2026-04-18 00:51:39 +07:00
pj d4a6e33aa6 feat(runner): pause-snapshot-evaluate-resume loop
Wires agent.Conn + driver.Driver + verifier.Verifier + trace.Writer
into the v0.1 step cycle: snapshot the SDK, push to verifier,
evaluate properties, write the trace step (with violations), release
the SDK pause, apply the next action via the driver, wait for idle.

Driver.Launch happens once before the loop and Terminate runs in
defer so even an early error tears down the app cleanly. Summary
returns step count and per-step violation records for the caller
to print or persist.
2026-04-17 23:54:05 +07:00