* feat(ltl): add Now/Next/Eventually/Implies/Or/And/Not formulas
Replace the fold-with-latch evaluator with a residual-formula reducer.
Each Observe() instantiates a fresh obligation from the root (stripping
an outer Always), reduces each pending obligation against current state,
latches Violated on first failure, and surfaces Pending verdicts for
deferred obligations. Existing Always/Pure/Thunk tests continue to pass.
* feat(ltl): support relative duration for eventually().within()
* feat(proto): add Swipe, PressKey, RecentLogs RPCs
* feat(verifier,runner): formula handles, new action kinds, rich state
- verifier: add formula-spec registry; bindNow/bindNext/bindEventually with
chainable .implies/.or/.and/.not and .within(n,unit) on eventually; bindFrom
for uniform sampling. bindAlways keeps accepting plain predicates.
- verifier: store lastTree, lastAction, step time, logs, exceptions on the
Verifier; SnapshotInput replaces the (snapshots, tree) pair. stateObject now
produces state.lastAction/time/logs/exceptions matching the TS State type.
- verifier: make taps/swipes/waitOnce/pressKey built-in generators actually
fire; taps picks a clickable, enabled element from the last hierarchy.
- agent: add exceptions field to Message wire format.
- driver: add Swipe/PressKey/RecentLogs to Driver interface; wire maestro
client and mock driver. LogEntry exposed for runner consumption.
- runner: apply Swipe/PressKey/Wait actions; collect logcat and exceptions;
pass lastAction and step time into PushSnapshot.
* feat(spec-api): LTL operators, new actions, richer State
- ltl.ts exports now/next/eventually; always overload accepts a Formula
- types.ts: Formula gains implies/or/and/not; EventuallyFormula adds .within;
State gains lastAction/time/logs/exceptions; Swipe/PressKey/Wait action types
- actions.ts: Swipe/PressKey/Wait/from constructors; waitOnce + pressKey
default generators
- tests exercise the chaining, sampling, and new actions through a recorded
fake runtime
* feat(sidecar): add swipe, pressKey, recentLogs RPC handlers
* feat(sdk-android): capture uncaught exceptions
Install a default uncaught handler on Uatu.start, chained with any
existing handler so Android's crash reporter still runs. Expose
Uatu.reportError for callers to forward caught throwables. A bounded
circular buffer (default 50) drains into each STATE message's new
exceptions field. Protocol.kt serializes/deserializes the field,
matching the Go wire format added to internal/agent/protocol.go.
* feat(spec-api): add @uatu/spec/defaults/properties bundle
* feat(sample-app): exercise new LTL operators + defaults
spec.ts now imports eventually/next/now/from from @uatu/spec and
noUncaughtExceptions from @uatu/spec/defaults/properties. It declares
three properties that exercise the new surface:
- accountCountNonNegative: plain always() safety
- addAccountAdvances: always(now(x).implies(next(y)))
- eventuallyLoggedIn: eventually(p).within(30, "seconds")
- noUncaughtExceptions: imported default
The weighted actions root uses from() for random phone/name sampling
and entries for taps/swipes/waitOnce/pressKey built-ins.
SampleApplication gains a debug hook gated on the system property
uatu.inject_error so the e2e run can synthesize an Uatu.reportError and
verify noUncaughtExceptions violates.
cmd/uatu/test_run.go adds a subpath alias so specs importing
"@uatu/spec/defaults/properties" resolve against the in-tree source
when running from the uatu checkout. The spec-integration tests swap
the old click-counter fixtures for the new login hierarchy.
* feat(trace): record swipe/key/wait details + exceptions
trace.Step gains an Exceptions array so the trace captures the
class/message/stackTrace for each SDK-reported throwable in a step.
trace.Action gains FromX/FromY/ToX/ToY/Key/DurationMillis so the full
payload of Swipe/PressKey/Wait actions is visible in trace.jsonl.
sample-app's debug error hook now gates on ApplicationInfo.DEBUGGABLE
instead of a system property (adb setprop fails on non-rooted
emulators).
* fix(runner): surface non-deadline WaitForIdle errors
Previously the WaitForIdle return value was discarded entirely, hiding
real driver failures (gRPC transport errors, sidecar crashes) behind
the expected deadline-exceeded case. Log non-deadline errors so they
are visible without changing control flow.
* chore(sample-app): drop unused uptime_millis extractor
Registered in SampleApplication but never consumed by spec.ts.
* fix(sample-app): drop trivial appIsRunning property
app_state was hardcoded to 'running' so the property was a tautology
that could never fail. Removing both the extractor and the property
is the simplest fix; demo-grade properties that can fail land next.
* feat(sample-app): add Reset button that zeroes clickCount
Pairs with the next commit's tap-reset action so the fuzzer can
violate clickCountNeverDecreases and demonstrate uatu actually
finding a property violation.
* feat(sample-app): add tap-reset action to exercise Reset button
Weighted at 10/122, fuzzer reaches it within a short run. Pairs with
the Reset button to demonstrate uatu detecting the
clickCountNeverDecreases violation.
* fix(runner): filter WaitForIdle errors via context state, not errors.Is
errors.Is(err, context.DeadlineExceeded) misses gRPC's wrapped
status.DeadlineExceeded, so every step under the maestro driver
logged a spurious warning. Check idleCtx.Err() instead — captures
both deadline-fired and parent-canceled cases regardless of how the
driver wraps them.
* chore(sample-app): tune action weights so demo violates in ~30s
Prior weights left tap-reset rare enough that short demo runs missed
the violation by chance. Bumped to 30/107, with typeUsername reduced
since username noise doesn't help exercise clickCount.
* refactor(runner): route warnings through slog
Adds Options.Logger (defaults to slog.Default()) and converts the
three warning sites that were using fmt.Printf. Progress line stays
on Printf since it's user-facing UI, not a log. Makes the warnings
testable via a capturing handler.
* test(runner): assert WaitForIdle driver errors are logged
Captures slog output via TextHandler into a buffer and asserts the
warning message + injected error text appear when the mock driver
returns a non-context error from WaitForIdle. Guards against a
regression of the silent-error swallow.
* fix(runner): warn on malformed screen snapshot
screenFromSnapshot swallowed json.Unmarshal errors, so a non-string
screen value silently became "" in the step log and trace while the
verifier still saw the raw JSON. Return the error and warn at the
call site, matching the hierarchy warning pattern.
* docs: clarify --avd is optional for uatu test
The CLI accepts --avd as an empty-string default (cmd/uatu/main.go:49)
and only requires it when no device is connected and multiple AVDs
exist (cmd/uatu/android_env.go:63). Docs and examples that showed it
as required or always-passed were misleading.
* fix(runner): surface focus-tap errors in InputText action
A failed Tap/TapSelector before InputText was swallowed, so text typed
into the wrong field (or no field) still reported success. Return the
error so the step fails explicitly.
* feat(sample-app): add username EditText and snapshot
Gives the spec a real EditText target (content-desc: username_field)
so the InputText action path can be exercised end-to-end. The typed
value is mirrored into MainActivity.username and surfaced as the
"username" snapshot for spec assertions.
* feat(sample-app): exercise InputText action against username field
Adds typeUsername action and usernameNeverShrinks property to the
sample spec, and extends the integration test to assert the bundled
spec emits an InputText(desc:username_field, "alice") action and that
the property correctly violates when a snapshot reports a shorter
string.
interestingTags hardcoded selectors from a specific app (etMobileNumber,
customer_row_, supplier_row_, etc.) inside the generic runner. None of
these selectors exist in the checked-in sample spec. Debug log now just
reports screen + hierarchy size; specs that want richer visibility can
log from state.ax.find themselves.
Removes Launch + Terminate from runner.Run so the CLI can launch
the app first, wait for the SDK to connect, then start the loop.
The previous shape forced runner to launch internally which fought
with the SDK-must-be-connected-first ordering.
BundleID/ClearState fields go away too since runner no longer
launches; the CLI keeps them on its testOptions struct.
Wires agent.Conn + driver.Driver + verifier.Verifier + trace.Writer
into the v0.1 step cycle: snapshot the SDK, push to verifier,
evaluate properties, write the trace step (with violations), release
the SDK pause, apply the next action via the driver, wait for idle.
Driver.Launch happens once before the loop and Terminate runs in
defer so even an early error tears down the app cleanly. Summary
returns step count and per-step violation records for the caller
to print or persist.