mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
113282ffb2c940e8e68867211d24d41775526a92
24
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0abbcd7d0c | feat(cli): default --clear-data on so runs start fresh | ||
|
|
b23fb0c723 |
feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag
Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.
* feat(testrun): add Preflight() before sidecar/driver setup
Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.
* refactor(chrome): split tag (HTML name) from class (CSS classList)
Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.
* feat(chrome): translate legacy string selectors to CSS/XPath
TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.
* feat(trace): add WriteHTML + Step.HTMLAvailable
Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.
* feat(driver): add WebDriver capability + chrome implementation
WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.
* feat(verifier): OverrideExtractorValues for V8-driven extractors
Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.
* feat(spec): add WebState + camelCase attribute aliases
WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.
* feat(runner): per-tick HTML capture for WebDriver-capable drivers
Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.
* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>
Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.
* feat(bundler): BundleWeb + V8-side runtime shim
web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.
* feat(runner): V8 extractor overrides + V8 action source for WebDriver
When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.
* feat(testrun): bundle + install web runtime when platform=web
BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.
* feat(inspect-ui): hierarchy + html panels in run detail
HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.
* fix(folio-web): drop aria-label data-carrier abuse
Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.
* chore: rebuild inspect-ui dist + folio-web .gitignore
Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.
* revert(trace): drop WriteHTML + Step.HTMLAvailable
Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.
* revert(runner): drop per-tick HTML capture
Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.
* revert(driver): drop WebDriver.Document
Document was only consumed by the runner's HTML capture which is gone.
* revert(inspect): drop /html route
Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.
* revert(inspect-ui): drop htmlUrl + html_available type
API surface no longer needs the HTML route; Step.html_available has no
producer.
* revert(inspect-ui): drop HtmlPanel + html tab
Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.
* test(inspect-ui): drop htmlUrl test, add @types/bun
Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.
* chore: rebuild inspect-ui dist without HtmlPanel
Embedded SPA bundle no longer ships the iframe HTML viewer.
* fix(web-runtime): retry action resolution + implement taps/swipes
V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.
* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys
Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.
* chore(folio-web): drop swipes from action root
Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.
* fix(inspect-ui): correct HierarchyPanel CSS variable names
Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.
* fix(chrome): correct PressKey mappings to chromedp/kb constants
Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.
Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.
* fix(cli): -h/--help exits 0 instead of error code
parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.
Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.
* fix(chrome): harden cssEscape for control chars + use [class~=]
Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.
Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.
* fix(web-runtime): use CSS.escape and validate tag-name selectors
The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.
The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.
Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.
* fix(chrome): validate attribute name in unknown-prefix branch
A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.
* fix(selectors): emit valid XPath 1.0 string literals via concat()
Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.
Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.
* fix(runtime): surface unresolved action targets instead of dropping silently
serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.
Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.
* fix(runner): use errgroup-bound ctx so siblings cancel on failure
The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.
* fix(chrome): propagate caller ctx cancellation to CDP calls
InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.
Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.
* fix(verifier): tolerate out-of-range override indices
A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.
Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.
* test(verifier): cover object-shaped extractor overrides
Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.
* fix(web-runtime): lock global runtime hooks against page shadowing
AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.
* perf(web-runtime): cache randomTap candidate DOM scan per tick
The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.
Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.
* fix(web-runtime): cap sanitize recursion to prevent stack overflow
State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.
* fix(web-runtime): enforce pressKey allowlist in factory
The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.
* chore(chrome): drop dead bundleSource/bundleMu
bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.
* fix(chrome): use strconv.Atoi for extractor key parsing
fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.
* fix(doctor): raise per-check timeout to 15s for chromium launch
5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.
* fix(runner): trust V8 coordinates for InputText, even at origin
resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.
Add applyAction tests covering both the typical web case and the (0,0)
edge case.
* test(bundler): lock down deterministic output across builds
The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
|
||
|
|
dd54c24c4e |
feat: --clear-data flag + typed attribute selectors (#48)
* feat(test): add --clear-data flag to clear app data on launch * test+docs: cover --clear-data flag in CLI parser test and reference * feat(spec): type AttrSelector with known attribute names Replace AttrSelector = Record<string, string> with KnownAttrSelectors plus a string|boolean index signature, so authors get autocomplete and type-checking on testTag / focused / clickable / etc. while raw driver attributes still type-check via the fallback. Boolean state attributes accept native booleans; goja stringifies them at the marshal boundary. AccessibilityElement.attrs becomes RawAttrs (typed string-valued shape of the same canonical names) so element.attrs.testTag autocompletes. * test(verifier): native boolean selector value matches focused=true * docs+folio: use native boolean for focused selector and document typed attrs |
||
|
|
97154cf580 |
feat(ios): launch via XCTest with env vars for hierarchy/tap access (#39)
* feat(proto): add env map to LaunchRequest * feat(driver): add env param to Launch interface and all implementations * feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through * feat(testrun): launch iOS app via XCTest with env vars instead of simctl * fix(ios): replace LaunchApp with BootedUDID; simctl launch moved to XCTest path * test(cli): add tests for ios platform flag and ios-device flag parsing * fix(sdk-ios): check semaphore wait result; resolve port from args and env; register extractors before start |
||
|
|
007dcddd69 |
feat(ios): iOS e2e support — Kotlin Native SDK + simulator driver + test-ios (#38)
* feat(ios): add simulator management package * feat(testrun): add iOS platform path (simctl launch + direct TCP) * feat(cli): add --ios-device flag and IosDevice option * feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser) * feat(folio-ios): wire SanderlingIos.start() in MainViewController * feat(folio-ios): add test-ios justfile recipe * fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings * fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects |
||
|
|
eed99e58aa |
refactor: code organization cleanup (#35)
* chore: fix gitignore + decisions doc after web->inspect-ui rename Update web/ references to inspect-ui/ in .gitignore and Makefile. Add decisions.md tracking architectural decisions from code-org discussion. * refactor: rename pkg/spec-api to pkg/spec Aligns the directory name with the npm package name @sanderling/spec. Updates Makefile, package.json directory field, and resolveSpecAPIPath. * refactor(verifier): split bindings.go into types.go + bindings.go Move shared public types (Action, ActionKind, LogEntry, Exception) to types.go. bindings.go retains internal JS runtime wiring only. * refactor(inspect): split runs.go into runs.go, runs_cache.go, runs_decode.go runs.go: types (RunSummary, StepSummary, RunDetail, Run) and Scan. runs_cache.go: Cache type, Open/Step/Detail methods, parseRun, scanSteps. runs_decode.go: readMeta, tallyTrace, decodeStepSummary, validRunID. * refactor: move android_env.go to internal/android/ Extracts Android device/AVD/adb logic into internal/android package. Exports EnsureDevice, AdbReverse, AdbReverseRemove, EnvWithAndroidPlatformTools, AdbBinary. Moves tests to internal/android/android_test.go. cmd/sanderling becomes a thin caller. * refactor: extract test pipeline to internal/testrun/ runTestPipeline logic moves to testrun.Execute. buildDriver, resolveSpecAPIPath, pickFreePort, and the progress logger move to internal/testrun/. cmd/sanderling/test_run.go becomes a thin adapter. Tests follow their code. * ci: update workflow paths after pkg/spec-api -> pkg/spec rename |
||
|
|
d5b6f6ccaf |
feat: Maestro driver integration + DeviceDriver architecture (#32)
* proto(driver): drop launcher_activity from LaunchRequest * refactor(driver): rename Driver to DeviceDriver, drop launcherActivity from Launch * refactor(driver): rename package maestro to sidecar * feat(driver): add ChromeDriver backed by chromedp * feat(hierarchy): replace XML parser with TreeNode JSON parser * refactor(driver): update mock and runner to DeviceDriver, drop launcherActivity * feat(runner): add platform routing for web vs sidecar * test(verifier): update hierarchy fixtures from XML to TreeNode JSON * feat(sidecar): extract readLogcat/readProcMetrics; add MaestroDriverBackend - Extract readLogcat() and readProcMetrics() as internal package-level helpers parameterised by serial - Add MaestroDriverBackend wiring maestro-client AndroidDriver - Drop launcherActivity from DriverBackend interface and StubDriverBackend - Add maestro-utils and micrometer-core as explicit compile deps * refactor(sidecar): inject DriverBackend into DriverService; drop launcherActivity - Remove serial and default backend from DriverService constructor - Make backend a required parameter - Drop launcherActivity from launch RPC handler - Create MaestroDriverBackend in Main.kt when platform is android * test(sidecar): update tests for dropped launcherActivity and required backend * chore(sidecar): remove unnecessary micrometer-core direct dependency |
||
|
|
8ccf95c1cf |
refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling
Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.
* chore(proto): regenerate driverpb after module path rename
The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.
* refactor: rename CLI binary uatu -> sanderling
Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.
* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk
Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.
* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar
Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.
* refactor: rename Uatu API surface -> Sanderling
- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
__uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.
* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling
Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.
* chore(build): rename gradle property + rootProject.name uatu -> sanderling
- Renames the uatu.version gradle property and all its -P references in
Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.
* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1
Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.
* refactor: rename npm package @uatu/spec -> @sanderling/spec
Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.
* docs: rename uatu -> sanderling in README, docs, and URLs
- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
internal/inspect/server.go.
* refactor: rename remaining internal uatu strings -> sanderling
- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
|
||
|
|
13bb2feb82 |
feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI
Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.
* test(trace): cover EndedAt + new step fields round-trip
* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()
Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.
* feat(runner): stamp residuals, hierarchy, selector targets, ended_at
Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.
* feat(inspect): scaffold embed dist for SPA assets
Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).
* chore(web): ignore web/ build output in root .gitignore
* chore(web): add bun + vite + vitest scaffold config
* feat(web): monochrome design tokens, typography, app shell CSS
* chore(web): placeholder for self-hosted JetBrains Mono fonts
* feat(web): index.html entry with style links and root mount
* feat(web): typescript types mirroring run/step trace schema
* feat(web): typed fetchers for runs/steps/screenshots
* feat(web): App shell with router and run/step routes
* feat(inspect): runs scan, lazy step parse, mtime-aware cache
* feat(web): RunList route with table, loading, and error states
* feat(web): RunDetail route shell with three placeholder panels
* feat(inspect): fsnotify-backed runs watcher with debounce
* fix(web): use jest-dom/vitest entry so matchers register
* test(web): cover listRuns happy path and error response
* test(web): render RunList with mocked fetch and assert row
* chore(web): commit bun lockfile
* feat(inspect): http handlers for runs/steps/screenshots/SSE
* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy
* feat(cmd): add 'uatu inspect' subcommand
* fix(web): align TS types with snake_case wire format
Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.
* feat(web): add ActionList panel for run-detail step navigation
* feat(web): add SnapshotTable panel with diff highlighting
Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.
* test(web): cover SnapshotTable rendering and diff behavior
Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.
* feat(web): add Screenshot panel with bounds and tap overlays
Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.
* fix(web): guard scrollIntoView call for jsdom compatibility
* test(web): cover Screenshot panel rendering and overlays
* test(web): cover ActionList rendering, selection, keyboard, and markers
* feat(web): add ExceptionsPanel component
* test(web): add ExceptionsPanel tests
* feat(web): add Timeline panel with property swimlanes
Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.
* test(web): cover Timeline empty state, cells, status, click, highlight
* feat(web): add ResidualNode recursive AST renderer
* test(web): cover ResidualNode operators, predicate, and error chip
* feat(web): add ViolationsPanel with status badges and jump button
* test(web): cover ViolationsPanel rows, status grouping, and jump button
* test(web): register testing-library cleanup globally
All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.
* feat(web): hooks for url/keyboard/theme/sse
* feat(web): wire all panels into run-detail with phone-dominant grid
ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.
* test(web): add three reference run fixtures (clean, violation, exception)
* build: web targets in Makefile + bun in CI; docs(inspect)
- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.
* feat(runner): capture a screenshot per step
The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.
* feat(inspect): include action_label in StepSummary
Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.
* test(inspect): accept either #app or #root in SPA shell fallback
* feat(web): render action_label and screen in ActionList rows
Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.
* feat(sidecar): implement screencap for android driver backend
Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.
* feat(proto): add Metrics RPC for per-step CPU and memory capture
* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes
* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status
* feat(runner): capture metrics + before/after screenshots per step
Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.
* fix(runner,sidecar): measure CPU across step via /proc stat delta
'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.
* feat(web): add Metrics type for per-step cpu and memory
* refactor(web): replace --accent-change with --accent-positive token
* refactor(web): recolor chip-progress as neutral outlined chip
* refactor(web): use neutral border for changed snapshot rows
* feat(web): add MetricsChart panel with HEAP and CPU lanes
SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.
* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows
Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.
* fix(runner): stop copying Tap selector into action.text
The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.
* fix(web): use text-muted for swipe arrow after accent-change removal
* fix(web): snapshot values truncate with ellipsis + title tooltip
Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.
* feat(web): state-before/after columns + metrics chart at bottom
RunDetail now renders a four-column grid:
actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.
* fix(web): skip zero-value ticks + add exception markers to metrics
HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.
* fix(web): let action body column shrink below its content
Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.
* feat(web): bigger state screenshots + single properties row
Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.
* feat(web): add minimal Tabs component
Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.
* feat(web): fold timeline into MetricsChart as STEPS lane
Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.
* refactor(web): tabbed state cards, drop standalone Timeline panel
State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.
* feat(web): lock app shell to 100vh with no page scroll
html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.
* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)
Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).
* feat(web): ViolationsPanel supports violationsOnly filter
* feat(web): ActionList arrow-key nav + listbox semantics + smaller font
Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.
* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets
Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.
* feat(web): fourth 'Violations' tab + wider actions + shorter metrics
Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.
* feat(web): compact RunDetail layout using 1px borders instead of panel padding
* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis
Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.
* fix(web): RunList rows no longer stretch to fill viewport height
Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.
* misc changes
* fix(web): hoist useState above early return in MetricsChart
Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.
* fix(web): subscribe to named SSE event instead of 'message'
Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.
* fix(inspect): unsubscribe SSE clients on disconnect
Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.
Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.
* fix(trace): rename resolvedBounds/tapPoint to snake_case
Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.
* chore(web): drop vitest and remove UI tests from CI
No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).
Fixes CI failure where `vitest run` exits 1 with no test files.
* chore(make): dedupe sidecar embed and drop recursive make
Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
|
||
|
|
36188ca906 |
test+refactor: real sidecar test, deterministic test sleeps, slog step line (#22)
* test(sidecar): replace assertTrue(true) placeholder with real server test
MainTest.mainExists() always passed and inflated the green-check count.
DriverServiceTest covers RPCs, but SidecarServer start/stop had no
coverage. Drop the placeholder and add SidecarServerTest that binds to
port 0, asserts a real ephemeral port, and stops cleanly.
* test(agent): drop 50ms sleep before cancel in TestServer_AcceptCancelsOnContext
Accept's closeListenerOnCancel watcher closes the listener as soon as
ctx fires, regardless of whether the outer Accept has reached
listener.Accept() yet. The sleep was a CI-flake surface (50ms is not
enough on a slow runner), and dropping it still exercises the same
outcome — Accept returns with ctx.Err() after cancellation.
Stable across 50x -count runs.
* test(agent): replace 2s sleep with done-chan in TestConn_SnapshotTimesOutIfSDKSilent
The silent-SDK fake held the connection open via time.Sleep(2s), which
coupled the test's wall clock to the server's 200ms snapshot-timeout
assertion. Swap for a done channel closed by t.Cleanup — the goroutine
exits when the test ends, independent of timing.
* test(maestro): make WaitForHealth_PollsUntilReady deterministic
Replace the 50ms wall-clock sleep that flipped healthReady with a
healthReadyAfterCall counter in the fake server. The handler returns
ready=true once healthCalls reaches the threshold, so the test's
"at least 2 polls before ready" assertion is satisfied by call
count rather than a race between the flip goroutine and the 25ms
poll loop.
* refactor(runner): route per-step progress through slog instead of fmt.Printf
The runner already carries a *slog.Logger for warnings (logger.Warn on
decode failures, predicate errors). The per-step status line was the
outlier — a bare fmt.Printf that wrote to os.Stdout unconditionally,
bypassing both the injected logger and any caller-configured writer.
Switch it to logger.Info("step", "index", ..., "screen", ..., "nodes", ...).
The caller (cmd/uatu) is responsible for wiring a logger whose handler
renders to the right stream; the next commit adds that wiring.
* feat(cli): render runner progress via a thin slog handler on stdout
progressHandler writes Info records as "msg key=value ..." and prefixes
warnings/errors with their level, matching the prose style of the
surrounding CLI status prints. Wired into the runner via
runner.Options.Logger so the per-step status line still lands on stdout
without slog's default time= / level= framing.
|
||
|
|
a2e96af1af |
WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio Directory-level rename and path references in Go tests, bundle-check, top-level README, and getting-started docs. Package declarations, Gradle config, iOS bundle IDs, and class names follow in later commits. * refactor(folio): rename Kotlin package dev.uatu.sample to app.folio Moves source dirs and sqldelight schema from dev/uatu/sample to app/folio, updates package declarations and imports, and switches Android namespace/applicationId, iOS binaryOption bundleId, and sqldelight database packageName to the new identifier. * refactor(folio): rename SampleApplication to FolioApplication Android manifest now points at .FolioApplication with label 'Folio' instead of 'Uatu Sample'. * refactor(folio): set iOS bundle id and display name to Folio bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio. CFBundleName + CFBundleDisplayName -> 'Folio'. * refactor(folio): update demo email to [email protected] * refactor(folio): point justfile at app.folio bundle id Updates xcrun simctl launch target, uatu test --bundle-id, and the build/uninstall comments to reference folio instead of sample. * test: update fixture package ids to app.folio Sidecar activity-resolver test and verifier spec-integration XML fixtures referenced the old dev.uatu.sample Android package. Updates them to match the folio app's real package id so the tests stay representative of what the CLI sees on-device. * test(verifier): rename SampleApp identifiers to Folio Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and sampleAppHierarchyXML const (now loginHierarchyXML for consistency with the other per-screen fixtures). Updates trailing sample-app mentions in comments and assertion messages. * refactor(folio): rename Gradle/npm/wasm project identifiers to folio settings.gradle.kts rootProject.name, package.json + package-lock.json name, and the WasmJS index.html <title> all still read 'uatu-sample' / 'Uatu Sample'. Realigns them with the Folio brand. * docs(folio): rewrite README title + getting-started bundle id examples/folio/README.md is now titled 'Folio' with the Kotlin source paths corrected to app/folio. Getting-started example uses --bundle-id app.folio. Harness launch message is now generic ('app under test') since uatu-sample-harness is not specific to folio. * chore(folio): drop trailing 'sample' reference in gradle.properties * refactor(folio): rename LoginPage composable to LoginScreen Align with KMP/Android industry convention (NowInAndroid, Cash App, JetBrains samples use Screen, not Page). * refactor(folio): rename HomePage composable to HomeScreen * refactor(folio): rename AddAccountPage composable to AddAccountScreen * refactor(folio): rename LedgerPage composable to LedgerScreen * refactor(folio): rename AddTransactionPage composable to AddTransactionScreen * refactor(folio): split Models.kt into app.folio.data package Account, Transaction (with TxnType), and Session move into their own files under app.folio.data, matching NowInAndroid-style per-type organization. * refactor(folio): move data layer into app.folio.data package Repository, LedgerStore (expect + interface), SqlLedgerStore, WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext, and Snapshot move into app.folio.data. Update all consumer imports. * refactor(folio): move Navigation into app.folio.navigation package Split the former Navigation.kt into Route.kt (sealed interface) and Navigator.kt (singleton). Update consumer imports across screens, App.kt, and FolioApplication. * refactor(folio): move Platform and Format into app.folio.platform Both files carry expect declarations (Platform object, formatDate); grouping them into a dedicated platform package makes the KMP seam obvious and mirrors the structure used by JetBrains samples. * refactor(folio): move login into feature/auth package Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into LoginScreen since it is the sole caller. * refactor(folio): move HomeScreen into feature/home package * refactor(folio): move account creation into feature/account package AddAccountScreen gets its own AddAccountUiState colocated with the screen, replacing the shared UiState.addAccountError. * refactor(folio): move ledger screens into feature/ledger package LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger with AddTransactionUiState (txnError, txnFormType) colocated. The former catch-all UiState.kt is removed now that each screen owns its state alongside its UI. * refactor(folio): split Theme.kt; move theme and icons to subpackages Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and Type.kt (typography) under app.folio.ui.theme. Icons moves to app.folio.ui.icon. Update every consumer's imports to match. * refactor(folio): split ui components into per-file under ui/component Former Widgets.kt and Components.kt become 10 focused files: AppButton, Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/ BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style of one composable per file in a designsystem/component package. * chore(folio): consolidate uatu testing files under uatu/ folder Move spec.ts, package.json, package-lock.json into examples/folio/uatu so all uatu-specific testing artifacts live in one place. runs/ and node_modules/ follow the same convention (both remain gitignored). Update justfile, README, and the two Go consumers (bundle-check tool + verifier/trace tests) that referenced the old path. * refactor(trace): drop folio path in writer test Round-trip only needs a non-empty string; neutralize to keep the library free of folio references. * refactor(sidecar): neutralize ResolveActivity test fixtures Swap app.folio for com.example.app in the fixture strings so the sidecar tests don't reference the example app by name. * refactor(bundle-check): take spec path as argument Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a positional spec path instead so the tool works for any example and leaves no folio reference in the library surface. * test(verifier): add neutral integration spec and hierarchy fixtures Adds testdata/integration_spec.ts with two routes ("list", "form"), an InputText on text_field, a Tap on primary/secondary_action, a safety property (itemCountNonNegative), and a liveness property (submitEventually). Adds hierarchies_test.go with matching XML fixtures. Constants intentionally go in a _test.go at package root rather than testdata/hierarchies.go because go skips .go files under testdata/. * refactor(verifier): replace folio integration tests with neutral ones Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test* entry points to TestIntegrationSpec*. Uses the synthetic spec and hierarchies added in the previous commit so the library's test suite no longer references examples/folio at all. Folio-specific coverage remains covered by examples/folio/justfile's 'just test' which exercises the real spec on device/emulator. * chore: remove cmd/uatu-sample-harness Not referenced by Makefile, docs, CI, or any script. Duplicates the adb reverse helpers already in cmd/uatu/test_run.go, and its name implies ownership by the sample app which violates the library/ example decoupling. If a bare-protocol debugging tool is later needed it belongs inside cmd/uatu/. * docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax The directory-tree Layout section rots faster than the code and duplicates what ls shows for free. The expect/actual paragraph in Stack was the same kind of filler. The README also claimed 'just AVD=Pixel_7 test' but the justfile reads AVD as an env var via env_var_or_default, so the correct invocation is 'AVD=Pixel_7 just test'. |
||
|
|
7493945251 |
feat: LTL operators, sampling, and default generators (#17)
* feat(ltl): add Now/Next/Eventually/Implies/Or/And/Not formulas Replace the fold-with-latch evaluator with a residual-formula reducer. Each Observe() instantiates a fresh obligation from the root (stripping an outer Always), reduces each pending obligation against current state, latches Violated on first failure, and surfaces Pending verdicts for deferred obligations. Existing Always/Pure/Thunk tests continue to pass. * feat(ltl): support relative duration for eventually().within() * feat(proto): add Swipe, PressKey, RecentLogs RPCs * feat(verifier,runner): formula handles, new action kinds, rich state - verifier: add formula-spec registry; bindNow/bindNext/bindEventually with chainable .implies/.or/.and/.not and .within(n,unit) on eventually; bindFrom for uniform sampling. bindAlways keeps accepting plain predicates. - verifier: store lastTree, lastAction, step time, logs, exceptions on the Verifier; SnapshotInput replaces the (snapshots, tree) pair. stateObject now produces state.lastAction/time/logs/exceptions matching the TS State type. - verifier: make taps/swipes/waitOnce/pressKey built-in generators actually fire; taps picks a clickable, enabled element from the last hierarchy. - agent: add exceptions field to Message wire format. - driver: add Swipe/PressKey/RecentLogs to Driver interface; wire maestro client and mock driver. LogEntry exposed for runner consumption. - runner: apply Swipe/PressKey/Wait actions; collect logcat and exceptions; pass lastAction and step time into PushSnapshot. * feat(spec-api): LTL operators, new actions, richer State - ltl.ts exports now/next/eventually; always overload accepts a Formula - types.ts: Formula gains implies/or/and/not; EventuallyFormula adds .within; State gains lastAction/time/logs/exceptions; Swipe/PressKey/Wait action types - actions.ts: Swipe/PressKey/Wait/from constructors; waitOnce + pressKey default generators - tests exercise the chaining, sampling, and new actions through a recorded fake runtime * feat(sidecar): add swipe, pressKey, recentLogs RPC handlers * feat(sdk-android): capture uncaught exceptions Install a default uncaught handler on Uatu.start, chained with any existing handler so Android's crash reporter still runs. Expose Uatu.reportError for callers to forward caught throwables. A bounded circular buffer (default 50) drains into each STATE message's new exceptions field. Protocol.kt serializes/deserializes the field, matching the Go wire format added to internal/agent/protocol.go. * feat(spec-api): add @uatu/spec/defaults/properties bundle * feat(sample-app): exercise new LTL operators + defaults spec.ts now imports eventually/next/now/from from @uatu/spec and noUncaughtExceptions from @uatu/spec/defaults/properties. It declares three properties that exercise the new surface: - accountCountNonNegative: plain always() safety - addAccountAdvances: always(now(x).implies(next(y))) - eventuallyLoggedIn: eventually(p).within(30, "seconds") - noUncaughtExceptions: imported default The weighted actions root uses from() for random phone/name sampling and entries for taps/swipes/waitOnce/pressKey built-ins. SampleApplication gains a debug hook gated on the system property uatu.inject_error so the e2e run can synthesize an Uatu.reportError and verify noUncaughtExceptions violates. cmd/uatu/test_run.go adds a subpath alias so specs importing "@uatu/spec/defaults/properties" resolve against the in-tree source when running from the uatu checkout. The spec-integration tests swap the old click-counter fixtures for the new login hierarchy. * feat(trace): record swipe/key/wait details + exceptions trace.Step gains an Exceptions array so the trace captures the class/message/stackTrace for each SDK-reported throwable in a step. trace.Action gains FromX/FromY/ToX/ToY/Key/DurationMillis so the full payload of Swipe/PressKey/Wait actions is visible in trace.jsonl. sample-app's debug error hook now gates on ApplicationInfo.DEBUGGABLE instead of a system property (adb setprop fails on non-rooted emulators). |
||
|
|
6b1c17328c |
fix(sample-app): make it standalone, no repo_root assumptions (#8)
* fix(cli): fall back to node_modules for @uatu/spec resolution
Drop the hard failure when the uatu source tree is not reachable from
the spec file. Users integrating uatu in their own app have @uatu/spec
installed via npm; esbuild now resolves it from node_modules.
* build(gradle): drop :sample-app include from root settings
The sample now has its own Gradle project in examples/sample-app/android.
* build(sample-app): vendor gradle wrapper
Users running the sample build the APK via ./gradlew from inside the
sample-app's own android/ directory, no repo-root wrapper required.
* build(sample-app): make gradle project standalone
Drop the project(':sdk-android') dependency in favor of the Maven
Central coordinate io.github.priyanshujain:sdk-android. The sample now
owns its settings.gradle.kts and gradle.properties, so it builds
without any pieces of the uatu source tree.
* chore(sample-app): declare @uatu/spec npm dependency
Mirrors what a downstream user would put in their own package.json.
Uses file: for pre-release development; becomes a normal semver pin
once @uatu/spec ships to npm.
* docs(sample-app): rewrite justfile and add README
Justfile drops repo_root; all recipes run against the local gradle
wrapper and uatu from PATH. README is scoped to what a user needs to
run the sample against their own device.
* docs(manual): update sample install steps to standalone layout
./gradlew :sample-app:installDebug no longer exists; the sample owns
its own wrapper under android/.
* feat(sdk): log when Uatu.start succeeds
Silent SDK start makes the "SDK didn't connect" failure mode
impossible to debug. One INFO line at start time is enough.
* fix(sidecar): launch via am start -W instead of monkey
monkey -p <pkg> -c LAUNCHER 1 is unreliable on API 36+: it reports no
error but silently fails to start the activity, so the SDK never runs
and the CLI times out on the SDK-accept handshake.
Resolve the launcher activity via `cmd package resolve-activity
--brief` and launch it with `am start -W -n`. -W makes the call block
until the activity is up, which also makes the subsequent SDK
accept timing deterministic.
* feat(cli): auto-resolve Android device; boot AVD if none connected
--avd becomes optional. Resolution order:
- use any already-connected adb device;
- else if --avd names an existing AVD, boot it and wait for boot;
- else if --avd is missing or names no AVD, error with a clear message.
Falls back to $ANDROID_HOME/emulator/emulator when the binary is not on
PATH, so a standard Android SDK install works without extra shell setup.
* docs(sample-app): AVD is optional; document both paths
just test runs against any connected device. If none, pass AVD=<name>
to have uatu boot the emulator for you.
* feat(cli): auto-discover Android SDK; auto-pick the lone AVD
adb and emulator are looked up via PATH, then $ANDROID_HOME,
$ANDROID_SDK_ROOT, ~/Library/Android/sdk, ~/Android/Sdk, and the
Homebrew cask path. The discovered platform-tools directory is
prepended to the sidecar's PATH so its adb subprocess calls work too.
When --avd isn't passed and no device is connected, the CLI picks the
sole local AVD and boots it. Multiple AVDs → error listing them.
* build(make): add `make install` that go-installs uatu onto PATH
Puts `uatu` into $GOBIN (or $GOPATH/bin) so the sample and any local
dev flow can call it without PATH= prefixes.
* docs(sample-app): zero-config just test; dotenv-load for persistence
Drop the expectation that users prefix commands with PATH=, ANDROID_HOME=,
or AVD=. `just test` now works as-is; optional knobs can be pinned in a
.env file alongside the justfile.
* fix(sample-app): auto-detect ANDROID_HOME for Gradle tasks
The Go CLI finds the SDK itself, but AGP still needs ANDROID_HOME to
resolve `sdk.dir`. The justfile now resolves it from env or canonical
install paths before invoking ./gradlew, so `just install` works out
of the box on a standard Android SDK setup.
* gitignore runs directory for sample app
|
||
|
|
e62319e916 |
docs: pandoc-based site and v0.1.0 groundwork (#5)
* chore(prose): remove em-dashes from config files * chore(prose): remove em-dashes from android sdk config * docs(spec-api): remove em-dash from README * fix(doctor): reword sidecar-jar error without em-dash * test(sidecar): reword assertion message without em-dash * docs: add CLAUDE.md with project conventions * build: add docs target for pandoc site * docs(site): add pandoc template and stylesheet * docs(site): add pandoc build script * docs(site): add landing pages * docs(manual): add getting-started * docs(manual): add writing-specs * docs(manual): add runs * docs(manual): add cli reference * docs(dev): add design principles * docs(dev): add architecture * ci: deploy docs site to github pages * docs: rewrite README as entry point to docs site |
||
|
|
0c775c199d |
feat(examples): add sample-app spec
Introduce examples/sample-app/spec.ts — a minimal property-based spec that taps the sample app's "Click me" button and asserts click_count is monotonic. Bundle-check and the trace writer test now reference the new path. |
||
|
|
0570719e6f |
Publish pipeline: goreleaser + Maven Central + npm (#1)
* feat(cli): add Version var and version subcommand * build(gradle): introduce uatu.version property for lockstep releases * build(sdk-android): swap GitHub Packages for vanniktech Maven Central plugin * build(spec-api): make package publish-ready for npm * ci(release): add goreleaser config for cross-platform uatu CLI builds * ci: add ci and release GitHub Actions workflows * ci: restrict ci.yml to PR + workflow_dispatch (no direct push to master) * docs(release): add local release targets, env example, and install docs * build(sdk-android): make signAllPublications conditional on signing key * ci(release): stage sidecar JAR at embed path before go build * chore(spec-api): regenerate package-lock for updated package.json * ci: install protoc-gen-go plugins before buf generate * ci: bump Node to 22 (required for --experimental-strip-types) |
||
|
|
6ce4eabe86 | feat(spec): drive merchant onboarding (English, phone, OTP, multi-device, home) | ||
|
|
c4cf0ff530 | feat(driver): optional launcher_activity on Launch RPC for multi-alias apps | ||
|
|
e99c85c363 |
feat(cli): orchestrate full uatu test pipeline
runTestPipeline assembles the v0.1 stack end to end:
1. Bundle the spec (with @uatu/spec alias resolution)
2. Extract the embedded sidecar JAR
3. Spawn java -jar sidecar --port <free>
4. Wait for sidecar Health
5. Listen on a host TCP port + adb reverse to the device's
localabstract:uatu-agent socket
6. Launch the app via the maestro driver
7. Accept the SDK HELLO
8. Load the bundle into the verifier
9. Open the trace writer + write meta.json
10. runner.Run for the requested duration
11. Terminate the app + clean up adb reverse
Test subcommand now prints a bundle error for a missing spec,
verifying the flag surface reaches the pipeline.
|
||
|
|
4dbad9073e |
feat(spec): merchant-ledger property + login/navigation generators
Property ledgerBalanceMatchesTxns asserts displayed balance equals sum(Given) - sum(Received) on either ledger screen. Catches the class of bug where the server-fed CustomerModel.balance diverges from the local-DB-fed CoreDatabaseDao sums (stale cache, partial sync, deleted-txn handling glitch, etc.). Generators are gated by screen state and weighted to push the run through login -> home -> ledger quickly: enterPhone (100), enterOtp (100), openCustomerOrSupplier (80), taps (10), swipes (2). Phone/OTP read from process.env via esbuild defines so credentials never land in source. cmd/internal-tools/bundle-check is a quick sanity tool to confirm the spec bundles before running uatu test. |
||
|
|
8c6a3171d7 |
feat(cli): doctor check that surfaces placeholder sidecar JAR
Doctor flags a shipped binary that's still carrying the build-time placeholder so `uatu test` won't silently fail trying to launch a nonexistent sidecar. Makefile uatu target now copies the freshly built fat JAR into the embed directory before `go build`. |
||
|
|
b6b18c3681 |
feat(cli): real doctor subcommand with adb/emulator/java checks
Each check is a value so tests can swap in fakes. Java check parses both legacy (1.8) and modern (17+) version strings. Emulator check falls back to ANDROID_HOME/emulator/emulator since the brew cask ships it without putting it on PATH. |
||
|
|
d991d99e0f |
feat: uatu-sample-harness for SDK round-trip validation
Listens on TCP, runs adb reverse to expose the host port as the device's localabstract:uatu-agent socket, then drives N PAUSE/STATE /RESUME cycles against the connected SDK. Cleanly removes the adb forward on exit. |
||
|
|
1c4b0dd00f |
feat(cli): scaffold uatu CLI with test and doctor subcommands
Wires the flag surface from plan §3 (--spec, --bundle-id, --platform, --avd, --duration, --seed, --output). Subcommand bodies are stubs; actual work lands with runner/doctor tasks. |