Commit Graph
46 Commits
Author SHA1 Message Date
pj 8ccf95c1cf refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling

Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.

* chore(proto): regenerate driverpb after module path rename

The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.

* refactor: rename CLI binary uatu -> sanderling

Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.

* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk

Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.

* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar

Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.

* refactor: rename Uatu API surface -> Sanderling

- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
  spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
  __uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.

* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling

Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.

* chore(build): rename gradle property + rootProject.name uatu -> sanderling

- Renames the uatu.version gradle property and all its -P references in
  Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.

* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1

Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.

* refactor: rename npm package @uatu/spec -> @sanderling/spec

Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.

* docs: rename uatu -> sanderling in README, docs, and URLs

- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
  'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
  internal/inspect/server.go.

* refactor: rename remaining internal uatu strings -> sanderling

- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
  localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
2026-04-21 11:57:49 +07:00
pj 13bb2feb82 feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI

Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.

* test(trace): cover EndedAt + new step fields round-trip

* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()

Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.

* feat(runner): stamp residuals, hierarchy, selector targets, ended_at

Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.

* feat(inspect): scaffold embed dist for SPA assets

Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).

* chore(web): ignore web/ build output in root .gitignore

* chore(web): add bun + vite + vitest scaffold config

* feat(web): monochrome design tokens, typography, app shell CSS

* chore(web): placeholder for self-hosted JetBrains Mono fonts

* feat(web): index.html entry with style links and root mount

* feat(web): typescript types mirroring run/step trace schema

* feat(web): typed fetchers for runs/steps/screenshots

* feat(web): App shell with router and run/step routes

* feat(inspect): runs scan, lazy step parse, mtime-aware cache

* feat(web): RunList route with table, loading, and error states

* feat(web): RunDetail route shell with three placeholder panels

* feat(inspect): fsnotify-backed runs watcher with debounce

* fix(web): use jest-dom/vitest entry so matchers register

* test(web): cover listRuns happy path and error response

* test(web): render RunList with mocked fetch and assert row

* chore(web): commit bun lockfile

* feat(inspect): http handlers for runs/steps/screenshots/SSE

* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy

* feat(cmd): add 'uatu inspect' subcommand

* fix(web): align TS types with snake_case wire format

Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.

* feat(web): add ActionList panel for run-detail step navigation

* feat(web): add SnapshotTable panel with diff highlighting

Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.

* test(web): cover SnapshotTable rendering and diff behavior

Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.

* feat(web): add Screenshot panel with bounds and tap overlays

Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.

* fix(web): guard scrollIntoView call for jsdom compatibility

* test(web): cover Screenshot panel rendering and overlays

* test(web): cover ActionList rendering, selection, keyboard, and markers

* feat(web): add ExceptionsPanel component

* test(web): add ExceptionsPanel tests

* feat(web): add Timeline panel with property swimlanes

Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.

* test(web): cover Timeline empty state, cells, status, click, highlight

* feat(web): add ResidualNode recursive AST renderer

* test(web): cover ResidualNode operators, predicate, and error chip

* feat(web): add ViolationsPanel with status badges and jump button

* test(web): cover ViolationsPanel rows, status grouping, and jump button

* test(web): register testing-library cleanup globally

All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.

* feat(web): hooks for url/keyboard/theme/sse

* feat(web): wire all panels into run-detail with phone-dominant grid

ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.

* test(web): add three reference run fixtures (clean, violation, exception)

* build: web targets in Makefile + bun in CI; docs(inspect)

- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
  install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
  vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.

* feat(runner): capture a screenshot per step

The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.

* feat(inspect): include action_label in StepSummary

Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.

* test(inspect): accept either #app or #root in SPA shell fallback

* feat(web): render action_label and screen in ActionList rows

Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.

* feat(sidecar): implement screencap for android driver backend

Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.

* feat(proto): add Metrics RPC for per-step CPU and memory capture

* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes

* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status

* feat(runner): capture metrics + before/after screenshots per step

Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.

* fix(runner,sidecar): measure CPU across step via /proc stat delta

'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.

* feat(web): add Metrics type for per-step cpu and memory

* refactor(web): replace --accent-change with --accent-positive token

* refactor(web): recolor chip-progress as neutral outlined chip

* refactor(web): use neutral border for changed snapshot rows

* feat(web): add MetricsChart panel with HEAP and CPU lanes

SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.

* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows

Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.

* fix(runner): stop copying Tap selector into action.text

The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.

* fix(web): use text-muted for swipe arrow after accent-change removal

* fix(web): snapshot values truncate with ellipsis + title tooltip

Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.

* feat(web): state-before/after columns + metrics chart at bottom

RunDetail now renders a four-column grid:
  actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.

* fix(web): skip zero-value ticks + add exception markers to metrics

HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.

* fix(web): let action body column shrink below its content

Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.

* feat(web): bigger state screenshots + single properties row

Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.

* feat(web): add minimal Tabs component

Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.

* feat(web): fold timeline into MetricsChart as STEPS lane

Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.

* refactor(web): tabbed state cards, drop standalone Timeline panel

State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.

* feat(web): lock app shell to 100vh with no page scroll

html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.

* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)

Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).

* feat(web): ViolationsPanel supports violationsOnly filter

* feat(web): ActionList arrow-key nav + listbox semantics + smaller font

Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.

* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets

Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.

* feat(web): fourth 'Violations' tab + wider actions + shorter metrics

Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.

* feat(web): compact RunDetail layout using 1px borders instead of panel padding

* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis

Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.

* fix(web): RunList rows no longer stretch to fill viewport height

Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.

* misc changes

* fix(web): hoist useState above early return in MetricsChart

Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.

* fix(web): subscribe to named SSE event instead of 'message'

Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.

* fix(inspect): unsubscribe SSE clients on disconnect

Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.

Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.

* fix(trace): rename resolvedBounds/tapPoint to snake_case

Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.

* chore(web): drop vitest and remove UI tests from CI

No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).

Fixes CI failure where `vitest run` exits 1 with no test files.

* chore(make): dedupe sidecar embed and drop recursive make

Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
2026-04-21 11:32:17 +07:00
pj 36188ca906 test+refactor: real sidecar test, deterministic test sleeps, slog step line (#22)
* test(sidecar): replace assertTrue(true) placeholder with real server test

MainTest.mainExists() always passed and inflated the green-check count.
DriverServiceTest covers RPCs, but SidecarServer start/stop had no
coverage. Drop the placeholder and add SidecarServerTest that binds to
port 0, asserts a real ephemeral port, and stops cleanly.

* test(agent): drop 50ms sleep before cancel in TestServer_AcceptCancelsOnContext

Accept's closeListenerOnCancel watcher closes the listener as soon as
ctx fires, regardless of whether the outer Accept has reached
listener.Accept() yet. The sleep was a CI-flake surface (50ms is not
enough on a slow runner), and dropping it still exercises the same
outcome — Accept returns with ctx.Err() after cancellation.

Stable across 50x -count runs.

* test(agent): replace 2s sleep with done-chan in TestConn_SnapshotTimesOutIfSDKSilent

The silent-SDK fake held the connection open via time.Sleep(2s), which
coupled the test's wall clock to the server's 200ms snapshot-timeout
assertion. Swap for a done channel closed by t.Cleanup — the goroutine
exits when the test ends, independent of timing.

* test(maestro): make WaitForHealth_PollsUntilReady deterministic

Replace the 50ms wall-clock sleep that flipped healthReady with a
healthReadyAfterCall counter in the fake server. The handler returns
ready=true once healthCalls reaches the threshold, so the test's
"at least 2 polls before ready" assertion is satisfied by call
count rather than a race between the flip goroutine and the 25ms
poll loop.

* refactor(runner): route per-step progress through slog instead of fmt.Printf

The runner already carries a *slog.Logger for warnings (logger.Warn on
decode failures, predicate errors). The per-step status line was the
outlier — a bare fmt.Printf that wrote to os.Stdout unconditionally,
bypassing both the injected logger and any caller-configured writer.

Switch it to logger.Info("step", "index", ..., "screen", ..., "nodes", ...).
The caller (cmd/uatu) is responsible for wiring a logger whose handler
renders to the right stream; the next commit adds that wiring.

* feat(cli): render runner progress via a thin slog handler on stdout

progressHandler writes Info records as "msg key=value ..." and prefixes
warnings/errors with their level, matching the prose style of the
surrounding CLI status prints. Wired into the runner via
runner.Options.Logger so the per-step status line still lands on stdout
without slog's default time= / level= framing.
2026-04-20 17:58:41 +07:00
pj 323878c34a fix(verifier): don't crash on throwing JS predicates (#21)
* fix(verifier): don't crash on throwing JS predicates

formulaThunk used to panic whenever goja returned an error from a
predicate callable, and nothing on the LTL -> runner path recovered, so
a malformed spec (e.g. a property whose body throws or touches an
undefined field) would kill the verifier process.

Latch the first error on formulaState, return false so LTL marks the
property violated, and expose PredicateError(name) that walks the
property's formula-spec tree and surfaces the latched cause.

* fix(runner): log predicate errors alongside violations

For each violated property, surface the verifier's latched predicate
error via logger.Warn so operators can distinguish a genuine false
verdict from a malformed spec. Add a runner-level test asserting that a
throwing predicate no longer crashes the run and that the error message
appears in the log.
2026-04-20 16:55:10 +07:00
pj a2e96af1af WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio

Directory-level rename and path references in Go tests, bundle-check,
top-level README, and getting-started docs. Package declarations,
Gradle config, iOS bundle IDs, and class names follow in later commits.

* refactor(folio): rename Kotlin package dev.uatu.sample to app.folio

Moves source dirs and sqldelight schema from dev/uatu/sample to
app/folio, updates package declarations and imports, and switches
Android namespace/applicationId, iOS binaryOption bundleId, and
sqldelight database packageName to the new identifier.

* refactor(folio): rename SampleApplication to FolioApplication

Android manifest now points at .FolioApplication with label 'Folio'
instead of 'Uatu Sample'.

* refactor(folio): set iOS bundle id and display name to Folio

bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio.
CFBundleName + CFBundleDisplayName -> 'Folio'.

* refactor(folio): update demo email to [email protected]

* refactor(folio): point justfile at app.folio bundle id

Updates xcrun simctl launch target, uatu test --bundle-id, and the
build/uninstall comments to reference folio instead of sample.

* test: update fixture package ids to app.folio

Sidecar activity-resolver test and verifier spec-integration XML
fixtures referenced the old dev.uatu.sample Android package. Updates
them to match the folio app's real package id so the tests stay
representative of what the CLI sees on-device.

* test(verifier): rename SampleApp identifiers to Folio

Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and
sampleAppHierarchyXML const (now loginHierarchyXML for consistency with
the other per-screen fixtures). Updates trailing sample-app mentions in
comments and assertion messages.

* refactor(folio): rename Gradle/npm/wasm project identifiers to folio

settings.gradle.kts rootProject.name, package.json + package-lock.json
name, and the WasmJS index.html <title> all still read 'uatu-sample' /
'Uatu Sample'. Realigns them with the Folio brand.

* docs(folio): rewrite README title + getting-started bundle id

examples/folio/README.md is now titled 'Folio' with the Kotlin source
paths corrected to app/folio. Getting-started example uses --bundle-id
app.folio. Harness launch message is now generic ('app under test')
since uatu-sample-harness is not specific to folio.

* chore(folio): drop trailing 'sample' reference in gradle.properties

* refactor(folio): rename LoginPage composable to LoginScreen

Align with KMP/Android industry convention (NowInAndroid, Cash App,
JetBrains samples use Screen, not Page).

* refactor(folio): rename HomePage composable to HomeScreen

* refactor(folio): rename AddAccountPage composable to AddAccountScreen

* refactor(folio): rename LedgerPage composable to LedgerScreen

* refactor(folio): rename AddTransactionPage composable to AddTransactionScreen

* refactor(folio): split Models.kt into app.folio.data package

Account, Transaction (with TxnType), and Session move into their own
files under app.folio.data, matching NowInAndroid-style per-type
organization.

* refactor(folio): move data layer into app.folio.data package

Repository, LedgerStore (expect + interface), SqlLedgerStore,
WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext,
and Snapshot move into app.folio.data. Update all consumer imports.

* refactor(folio): move Navigation into app.folio.navigation package

Split the former Navigation.kt into Route.kt (sealed interface) and
Navigator.kt (singleton). Update consumer imports across screens,
App.kt, and FolioApplication.

* refactor(folio): move Platform and Format into app.folio.platform

Both files carry expect declarations (Platform object, formatDate);
grouping them into a dedicated platform package makes the KMP seam
obvious and mirrors the structure used by JetBrains samples.

* refactor(folio): move login into feature/auth package

Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline
the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into
LoginScreen since it is the sole caller.

* refactor(folio): move HomeScreen into feature/home package

* refactor(folio): move account creation into feature/account package

AddAccountScreen gets its own AddAccountUiState colocated with the
screen, replacing the shared UiState.addAccountError.

* refactor(folio): move ledger screens into feature/ledger package

LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger
with AddTransactionUiState (txnError, txnFormType) colocated. The
former catch-all UiState.kt is removed now that each screen owns its
state alongside its UI.

* refactor(folio): split Theme.kt; move theme and icons to subpackages

Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and
Type.kt (typography) under app.folio.ui.theme. Icons moves to
app.folio.ui.icon. Update every consumer's imports to match.

* refactor(folio): split ui components into per-file under ui/component

Former Widgets.kt and Components.kt become 10 focused files: AppButton,
Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/
BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style
of one composable per file in a designsystem/component package.

* chore(folio): consolidate uatu testing files under uatu/ folder

Move spec.ts, package.json, package-lock.json into examples/folio/uatu
so all uatu-specific testing artifacts live in one place. runs/ and
node_modules/ follow the same convention (both remain gitignored).
Update justfile, README, and the two Go consumers (bundle-check tool +
verifier/trace tests) that referenced the old path.

* refactor(trace): drop folio path in writer test

Round-trip only needs a non-empty string; neutralize to keep the
library free of folio references.

* refactor(sidecar): neutralize ResolveActivity test fixtures

Swap app.folio for com.example.app in the fixture strings so the
sidecar tests don't reference the example app by name.

* refactor(bundle-check): take spec path as argument

Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a
positional spec path instead so the tool works for any example and
leaves no folio reference in the library surface.

* test(verifier): add neutral integration spec and hierarchy fixtures

Adds testdata/integration_spec.ts with two routes ("list", "form"),
an InputText on text_field, a Tap on primary/secondary_action, a
safety property (itemCountNonNegative), and a liveness property
(submitEventually). Adds hierarchies_test.go with matching XML
fixtures. Constants intentionally go in a _test.go at package root
rather than testdata/hierarchies.go because go skips .go files
under testdata/.

* refactor(verifier): replace folio integration tests with neutral ones

Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test*
entry points to TestIntegrationSpec*. Uses the synthetic spec and
hierarchies added in the previous commit so the library's test suite
no longer references examples/folio at all.

Folio-specific coverage remains covered by examples/folio/justfile's
'just test' which exercises the real spec on device/emulator.

* chore: remove cmd/uatu-sample-harness

Not referenced by Makefile, docs, CI, or any script. Duplicates the
adb reverse helpers already in cmd/uatu/test_run.go, and its name
implies ownership by the sample app which violates the library/
example decoupling. If a bare-protocol debugging tool is later
needed it belongs inside cmd/uatu/.

* docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax

The directory-tree Layout section rots faster than the code and
duplicates what ls shows for free. The expect/actual paragraph in
Stack was the same kind of filler. The README also claimed 'just
AVD=Pixel_7 test' but the justfile reads AVD as an env var via
env_var_or_default, so the correct invocation is 'AVD=Pixel_7
just test'.
2026-04-20 16:04:20 +07:00
pj c1de4bf57a fix(sample-app): address PR #18 review findings (#19)
- WebLedgerStore.accountExistsByName case-insensitive (matches SQL COLLATE NOCASE)
- drop SampleApplication.maybeInjectDebugError; sample spec now passes clean
- drop noLogcatErrors from sample spec (default matches system-wide E logs)
- openRandomAccount picks uniformly from findAll instead of first match
- clear txnError on any add-transaction interaction, not only valid amount input
- escapeForAdbInputText quotes shell metacharacters (quotes, backslash, etc.)
- extract buildClearKeyevents helper with empty/normal/cap tests
- broaden spec_integration_test.go to cover home, add-account, ledger,
  add-transaction action generators
- document FocusTracker single-focus invariant
2026-04-20 13:31:50 +07:00
pj 8381a98aaf feat(sample-app): ledger spec parity (#18)
* feat(sample-app): add FocusTracker

* feat(sample-app): support description on TextInput

* feat(sample-app): support description on AppButton and Segmented

* feat(sample-app): stable ids on login screen

* feat(sample-app): stable ids on add-account screen

* feat(sample-app): stable ids on add-transaction screen

* feat(sample-app): stable ids on home and ledger screens

* feat(sample-app): hoist error + form state into UiState

* feat(sample-app): register auth_status + accounts snapshots

* feat(sample-app): register ledger_rows + ledger_balance snapshots

* feat(sample-app): register error + focused_input snapshots

* refactor(sample-app): scaffold spec.ts extractors + safety

* feat(sample-app): spec accounting invariants

* feat(sample-app): spec state-machine monotonicity

* feat(sample-app): spec liveness properties

* feat(sample-app): spec auth + account action generators

* feat(sample-app): spec transaction action generators

* feat(sample-app): spec weighted workflow

* feat(sample-app): publish login form input snapshots

* feat(sample-app): publish account + txn input snapshots

* feat(sample-app): register form input snapshots

* fix(sample-app): sequence login + idempotent text inputs in spec

* fix(sample-app): clear UiState on page dispose

* fix(sample-app): tighten adversarial login, allow retype on error

* test(verifier): update sample-app spec integration for new selectors

* feat(sidecar): clear focused field before InputText types

* refactor(sample-app): simplify loginHelper to focus-driven sequencing

* refactor(sample-app): drop input-value UiState mirrors

* refactor(sample-app): drop input-value snapshot registrations

* test(verifier): adjust sample-app integration for replace-on-type
2026-04-20 13:28:41 +07:00
pj 7493945251 feat: LTL operators, sampling, and default generators (#17)
* feat(ltl): add Now/Next/Eventually/Implies/Or/And/Not formulas

Replace the fold-with-latch evaluator with a residual-formula reducer.
Each Observe() instantiates a fresh obligation from the root (stripping
an outer Always), reduces each pending obligation against current state,
latches Violated on first failure, and surfaces Pending verdicts for
deferred obligations. Existing Always/Pure/Thunk tests continue to pass.

* feat(ltl): support relative duration for eventually().within()

* feat(proto): add Swipe, PressKey, RecentLogs RPCs

* feat(verifier,runner): formula handles, new action kinds, rich state

- verifier: add formula-spec registry; bindNow/bindNext/bindEventually with
  chainable .implies/.or/.and/.not and .within(n,unit) on eventually; bindFrom
  for uniform sampling. bindAlways keeps accepting plain predicates.
- verifier: store lastTree, lastAction, step time, logs, exceptions on the
  Verifier; SnapshotInput replaces the (snapshots, tree) pair. stateObject now
  produces state.lastAction/time/logs/exceptions matching the TS State type.
- verifier: make taps/swipes/waitOnce/pressKey built-in generators actually
  fire; taps picks a clickable, enabled element from the last hierarchy.
- agent: add exceptions field to Message wire format.
- driver: add Swipe/PressKey/RecentLogs to Driver interface; wire maestro
  client and mock driver. LogEntry exposed for runner consumption.
- runner: apply Swipe/PressKey/Wait actions; collect logcat and exceptions;
  pass lastAction and step time into PushSnapshot.

* feat(spec-api): LTL operators, new actions, richer State

- ltl.ts exports now/next/eventually; always overload accepts a Formula
- types.ts: Formula gains implies/or/and/not; EventuallyFormula adds .within;
  State gains lastAction/time/logs/exceptions; Swipe/PressKey/Wait action types
- actions.ts: Swipe/PressKey/Wait/from constructors; waitOnce + pressKey
  default generators
- tests exercise the chaining, sampling, and new actions through a recorded
  fake runtime

* feat(sidecar): add swipe, pressKey, recentLogs RPC handlers

* feat(sdk-android): capture uncaught exceptions

Install a default uncaught handler on Uatu.start, chained with any
existing handler so Android's crash reporter still runs. Expose
Uatu.reportError for callers to forward caught throwables. A bounded
circular buffer (default 50) drains into each STATE message's new
exceptions field. Protocol.kt serializes/deserializes the field,
matching the Go wire format added to internal/agent/protocol.go.

* feat(spec-api): add @uatu/spec/defaults/properties bundle

* feat(sample-app): exercise new LTL operators + defaults

spec.ts now imports eventually/next/now/from from @uatu/spec and
noUncaughtExceptions from @uatu/spec/defaults/properties. It declares
three properties that exercise the new surface:

- accountCountNonNegative: plain always() safety
- addAccountAdvances: always(now(x).implies(next(y)))
- eventuallyLoggedIn: eventually(p).within(30, "seconds")
- noUncaughtExceptions: imported default

The weighted actions root uses from() for random phone/name sampling
and entries for taps/swipes/waitOnce/pressKey built-ins.

SampleApplication gains a debug hook gated on the system property
uatu.inject_error so the e2e run can synthesize an Uatu.reportError and
verify noUncaughtExceptions violates.

cmd/uatu/test_run.go adds a subpath alias so specs importing
"@uatu/spec/defaults/properties" resolve against the in-tree source
when running from the uatu checkout. The spec-integration tests swap
the old click-counter fixtures for the new login hierarchy.

* feat(trace): record swipe/key/wait details + exceptions

trace.Step gains an Exceptions array so the trace captures the
class/message/stackTrace for each SDK-reported throwable in a step.
trace.Action gains FromX/FromY/ToX/ToY/Key/DurationMillis so the full
payload of Swipe/PressKey/Wait actions is visible in trace.jsonl.

sample-app's debug error hook now gates on ApplicationInfo.DEBUGGABLE
instead of a system property (adb setprop fails on non-rooted
emulators).
2026-04-20 02:19:39 +07:00
pj 69002b07b1 fix(runner,sample-app): surface silent errors and demo a failing property (#15)
* fix(runner): surface non-deadline WaitForIdle errors

Previously the WaitForIdle return value was discarded entirely, hiding
real driver failures (gRPC transport errors, sidecar crashes) behind
the expected deadline-exceeded case. Log non-deadline errors so they
are visible without changing control flow.

* chore(sample-app): drop unused uptime_millis extractor

Registered in SampleApplication but never consumed by spec.ts.

* fix(sample-app): drop trivial appIsRunning property

app_state was hardcoded to 'running' so the property was a tautology
that could never fail. Removing both the extractor and the property
is the simplest fix; demo-grade properties that can fail land next.

* feat(sample-app): add Reset button that zeroes clickCount

Pairs with the next commit's tap-reset action so the fuzzer can
violate clickCountNeverDecreases and demonstrate uatu actually
finding a property violation.

* feat(sample-app): add tap-reset action to exercise Reset button

Weighted at 10/122, fuzzer reaches it within a short run. Pairs with
the Reset button to demonstrate uatu detecting the
clickCountNeverDecreases violation.

* fix(runner): filter WaitForIdle errors via context state, not errors.Is

errors.Is(err, context.DeadlineExceeded) misses gRPC's wrapped
status.DeadlineExceeded, so every step under the maestro driver
logged a spurious warning. Check idleCtx.Err() instead — captures
both deadline-fired and parent-canceled cases regardless of how the
driver wraps them.

* chore(sample-app): tune action weights so demo violates in ~30s

Prior weights left tap-reset rare enough that short demo runs missed
the violation by chance. Bumped to 30/107, with typeUsername reduced
since username noise doesn't help exercise clickCount.

* refactor(runner): route warnings through slog

Adds Options.Logger (defaults to slog.Default()) and converts the
three warning sites that were using fmt.Printf. Progress line stays
on Printf since it's user-facing UI, not a log. Makes the warnings
testable via a capturing handler.

* test(runner): assert WaitForIdle driver errors are logged

Captures slog output via TextHandler into a buffer and asserts the
warning message + injected error text appear when the mock driver
returns a non-context error from WaitForIdle. Guards against a
regression of the silent-error swallow.
2026-04-18 18:48:25 +07:00
pj 00f66ba48d fix: pre-v0.1 review feedback (#14)
* fix(sdk-android): add @JvmOverloads to Uatu.start

Java callers can now invoke start(application) without supplying a
Configuration, matching the Kotlin default-arg ergonomics.

* feat(agent): add protocol_version to HELLO handshake

ProtocolVersion=1 lives on Message and is set by Hello(). Server.Accept
rejects mismatches with a clear error. SDK upgrades that don't change
the wire format keep the same protocol_version; bump on breaking changes.

* test(agent): assert protocol_version in Hello round-trip

* feat(sdk-android): send protocol_version=1 in HELLO

Mirrors agent.ProtocolVersion on the Go side. Bump in lockstep with
the Go constant when the wire format breaks.

* chore(sample-app): pull @uatu/spec from npm next tag

Replaces the file: dep. Copy-paste users can now npm install against
the registry. The release workflow publishes pre-release tags to
npm dist-tag 'next', so the sample tracks the latest rc without
manual version bumps. Lockfile currently resolves to 0.0.1-rc3.
2026-04-18 18:00:12 +07:00
pj d0578dbaaa fix(runner): warn on malformed screen snapshot (#13)
* fix(runner): warn on malformed screen snapshot

screenFromSnapshot swallowed json.Unmarshal errors, so a non-string
screen value silently became "" in the step log and trace while the
verifier still saw the raw JSON. Return the error and warn at the
call site, matching the hierarchy warning pattern.

* docs: clarify --avd is optional for uatu test

The CLI accepts --avd as an empty-string default (cmd/uatu/main.go:49)
and only requires it when no device is connected and multiple AVDs
exist (cmd/uatu/android_env.go:63). Docs and examples that showed it
as required or always-passed were misleading.
2026-04-18 17:14:31 +07:00
pj 16e55086d8 fix(runner): surface focus-tap errors in InputText (#12)
* fix(runner): surface focus-tap errors in InputText action

A failed Tap/TapSelector before InputText was swallowed, so text typed
into the wrong field (or no field) still reported success. Return the
error so the step fails explicitly.

* feat(sample-app): add username EditText and snapshot

Gives the spec a real EditText target (content-desc: username_field)
so the InputText action path can be exercised end-to-end. The typed
value is mirrored into MainActivity.username and surfaced as the
"username" snapshot for spec assertions.

* feat(sample-app): exercise InputText action against username field

Adds typeUsername action and usernameNeverShrinks property to the
sample spec, and extends the integration test to assert the bundled
spec emits an InputText(desc:username_field, "alice") action and that
the property correctly violates when a snapshot reports a shorter
string.
2026-04-18 16:51:35 +07:00
pj b457e22569 refactor(runner): drop hardcoded per-app selectors from step log (#10)
interestingTags hardcoded selectors from a specific app (etMobileNumber,
customer_row_, supplier_row_, etc.) inside the generic runner. None of
these selectors exist in the checked-in sample spec. Debug log now just
reports screen + hierarchy size; specs that want richer visibility can
log from state.ax.find themselves.
2026-04-18 16:41:12 +07:00
pj ae595526da fix(agent): race in readWithDeadline clobbers conn deadline (#7)
* refactor(docs): inline pandoc build into Makefile, drop scripts dir

* fix(agent): wait for deadline watcher before returning

readWithDeadline's watcher goroutine could clobber the conn's read
deadline with time.Unix(1, 0) after the main function reset it to zero.
When the Accept ctx was canceled shortly after Accept returned, the
watcher raced with close(done) in select and sometimes picked ctx.Done()
even though we were already done reading, leaving the conn unusable for
the next read (instant i/o timeout on step 1 snapshot).

Synchronize on the watcher's exit before resetting the deadline so it
can never override the reset.

* test(agent): cover readWithDeadline race on Accept ctx cancel

Drives Accept with a short-timeout ctx, cancels it right after Accept
returns, then does a Snapshot. Reliably fails without the readWithDeadline
synchronization fix (watcher goroutine overwrites the deadline to past).
2026-04-18 14:16:15 +07:00
pj 40c92d4569 test(verifier): rewrite integration tests for sample-app spec
The old tests loaded merchant-android uiautomator dumps from /tmp and
skipped when absent. Replace with two hermetic tests that bundle
examples/sample-app/spec.ts against a synthetic hierarchy: one checks
tapClickMe fires, the other drives the three properties through a
holds/holds/violated snapshot sequence.
2026-04-18 10:33:43 +07:00
pj 0c775c199d feat(examples): add sample-app spec
Introduce examples/sample-app/spec.ts — a minimal property-based spec that
taps the sample app's "Click me" button and asserts click_count is
monotonic. Bundle-check and the trace writer test now reference the new
path.
2026-04-18 10:33:35 +07:00
pj 6e4c832678 chore: stop tracking sidecar JAR (build artifact)
make uatu copies the real fat JAR into assets/ before
go build -tags withsidecar. Keeping that path tracked was
the root cause of the 130 MB push rejection.
2026-04-18 06:46:32 +07:00
pj ae10354ff0 test(sidecar): split tests across stub and withsidecar builds
Default !withsidecar build: assert IsPlaceholder, empty JAR,
Extract errors. withsidecar build: existing extract/checksum
coverage.
2026-04-18 06:46:24 +07:00
pj 4744736dc5 refactor(sidecar): gate JAR embed behind withsidecar build tag
Splits embed.go so the go:embed directive only fires under
-tags withsidecar. Default builds get a stub with a nil JAR
and IsPlaceholder()=true. This removes the landmine where
make uatu overwrote a tracked placeholder file, making any
git add silently stage 130 MB.
2026-04-18 06:46:20 +07:00
pj f2fe54debc test(verifier): verify dismissMultiDevice fires against real confirm dialog 2026-04-18 06:39:00 +07:00
pj 4b43b9a22a chore(runner): log screen + tag hits per step for debuggability 2026-04-18 06:39:00 +07:00
pj cefb398437 feat(verifier): retry NextAction up to 16x so gated generators eventually fire 2026-04-18 06:39:00 +07:00
pj 6c7ea35a7d feat(hierarchy): expose checked/focused/selected on element objects 2026-04-18 06:39:00 +07:00
pj b248ad2e5a test(verifier): integration tests against live uiautomator dumps 2026-04-18 01:59:54 +07:00
pj eea9a0067a feat(hierarchy): descPrefix selector for Compose testTag + UUID suffixes 2026-04-18 01:59:46 +07:00
pj fb6c3ca517 fix(runner): fetch hierarchy before SDK pause to avoid stale uiautomator dumps 2026-04-18 01:59:41 +07:00
pj c4cf0ff530 feat(driver): optional launcher_activity on Launch RPC for multi-alias apps 2026-04-18 01:59:36 +07:00
pj 447fd39363 feat(runner): fetch hierarchy per step and resolve Tap coords 2026-04-18 01:24:00 +07:00
pj 5f2a2d52ab feat(verifier): wire hierarchy into state.ax and carry element coords in Action 2026-04-18 01:23:55 +07:00
pj e539e89438 feat(hierarchy): XML parser and selector resolver for uiautomator dumps 2026-04-18 01:23:51 +07:00
pj e7b3e2ba9c refactor(runner): caller manages app launch/terminate
Removes Launch + Terminate from runner.Run so the CLI can launch
the app first, wait for the SDK to connect, then start the loop.
The previous shape forced runner to launch internally which fought
with the SDK-must-be-connected-first ordering.

BundleID/ClearState fields go away too since runner no longer
launches; the CLI keeps them on its testOptions struct.
2026-04-18 00:51:39 +07:00
pj 7aa75f3d9f feat(bundler): add Aliases option for spec import resolution
Specs import from "@uatu/spec"; the bundler maps that to the
vendored pkg/spec-api/src/index.ts via esbuild's Alias map.
2026-04-18 00:08:31 +07:00
pj 6a9fac7ec0 feat(sidecar): embed sidecar JAR via go:embed
Ships with a 24-byte placeholder so fresh clones build without
requiring a sidecar build first. `make uatu` copies the real
fat JAR into internal/sidecar/assets before `go build`, so
shipping binaries carry the full sidecar (~130 MB).

Extract writes the JAR to a temp dir alongside a SHA-256 file
and skips rewrite when the checksum already matches.
2026-04-18 00:06:27 +07:00
pj 69d0800ea3 feat(driver/maestro): gRPC client implementing driver.Driver
Wraps each v0.1 RPC, plus a WaitForHealth helper that polls until
the sidecar reports Ready=true (used at startup before any other
calls happen). Tests stub the gRPC server in-process so they don't
need a real sidecar JAR.
2026-04-18 00:04:24 +07:00
pj d856c40d7c feat(permissions): grant dangerous permissions via aapt + adb
Inspector parses uses-permission entries from `aapt dump permissions`
and the granter shells out to `adb shell pm grant`. Both are pluggable
so tests can drive logic without aapt or a device. Granter failures
become warnings rather than errors — Android refuses non-runtime
permissions and we'd rather keep going than abort the run.
2026-04-17 23:57:17 +07:00
pj d4a6e33aa6 feat(runner): pause-snapshot-evaluate-resume loop
Wires agent.Conn + driver.Driver + verifier.Verifier + trace.Writer
into the v0.1 step cycle: snapshot the SDK, push to verifier,
evaluate properties, write the trace step (with violations), release
the SDK pause, apply the next action via the driver, wait for idle.

Driver.Launch happens once before the loop and Terminate runs in
defer so even an early error tears down the app cleanly. Summary
returns step count and per-step violation records for the caller
to print or persist.
2026-04-17 23:54:05 +07:00
pj d2fae427f9 feat(driver): add TapSelector and make mock WaitForIdle honor duration
Real maestro driver will resolve selectors itself rather than
forcing the runner to look up coordinates from the hierarchy. Mock
now waits the requested duration so tests don't busy-spin and
starve the SDK fixture goroutine.
2026-04-17 23:54:00 +07:00
pj abeded0b19 test(verifier): cover spec lifecycle and weighted action selection 2026-04-17 23:48:59 +07:00
pj 72175bb13c feat(verifier): goja runtime hosting the spec API
Installs globalThis.__uatu__ with extract, always, actions,
weighted, tap, inputText, and stub taps/swipes. Load runs the
bundled spec, then pulls properties + actions out of globalThis.
PushSnapshot rebuilds state.snapshots and refreshes every
extractor handle's current/previous in registration order so
chained extractors observe up-to-date values.

Properties are wired through internal/ltl as Always(Thunk(...)),
so verdicts latch to violated as soon as a predicate returns
false. NextAction resolves actions/weighted recursively with a
seedable rand source for reproducible runs.
2026-04-17 23:48:58 +07:00
pj 72c4e4c4d3 feat(trace): JSONL writer for steps + meta + screenshots
Each WriteStep appends one JSON object per line so the trace can
be filtered with jq directly. Screenshots land under screenshots/
with zero-padded indexes. Writer is concurrency-safe; Close is
idempotent and a write after Close errors loudly rather than
silently dropping.
2026-04-17 23:43:45 +07:00
pj 0c277b2833 feat(ltl): formula AST and step evaluator for v0.1
Supports Always over Pure/Thunk leaves; eventually/next/bounds
deferred. Once a thunk returns false under an Always, the verdict
latches to violated so the runner can surface the offending step
without later observations masking it.
2026-04-17 23:42:44 +07:00
pj cee90694da feat(bundler): esbuild wrapper for spec compilation
Bundles a TypeScript entry into an IIFE ES2020 blob with optional
inline sourcemaps. process.env defines are JSON-quoted before being
fed to esbuild so values with quotes/backslashes survive correctly.
Returns the bundle's SHA-256 for trace meta.json.
2026-04-17 23:41:43 +07:00
pj 6353afdf83 feat(driver/mock): in-memory Driver for unit tests
Records every call as an Action and lets tests program hierarchy,
screenshot, health, and per-method failures. Actions() returns a
copy so test code can't mutate the driver's history.
2026-04-17 23:40:28 +07:00
pj d533aef04b feat(driver): Driver interface for v0.1 RPC surface
Mirrors proto/driverpb/driver.proto: Launch, Terminate, Tap,
InputText, Hierarchy, Screenshot, WaitForIdle, Health. Keeps
Image/Health as dedicated structs so callers don't depend on
generated proto types.
2026-04-17 23:40:28 +07:00
pj b9511c0718 feat(agent): socket server with PAUSE/STATE/RESUME flow
Accept waits for an SDK HELLO then hands back a Conn. Conn.Snapshot
sends a PAUSE, blocks on the matching STATE (id-correlated), and
leaves the SDK paused until Release sends RESUME. Conn.Close sends
GOODBYE best-effort.

v0.1 supports one client at a time; transport is left to the caller
so tests can use TCP loopback while production wires via adb reverse
to localabstract:uatu-agent.
2026-04-17 22:50:54 +07:00
pj ca92bb1492 feat(agent): Go wire protocol with length-prefixed JSON framing
Six message types: HELLO, PAUSE, RESUME, STATE, EXTRACT_RESULT,
GOODBYE. 4-byte big-endian length + JSON payload. 16MB frame cap.
Typed constructors per message; snapshots carry json.RawMessage so
downstream goja evaluation sees exact types.
2026-04-17 22:48:07 +07:00