Commit Graph
140 Commits
Author SHA1 Message Date
pj abdc8b45a3 test(verifier): pin element encoding to one rule on both hosts 2026-08-15 13:38:51 +05:30
pj b1b6f8558c fix(verifier): encode element-valued extractors into the trace
An ax element exports with its find/findAll host functions attached, and
json.Marshal refuses the whole value over them: json: unsupported type:
func(goja.FunctionCall) goja.Value. The encoding failed, curr stayed nil,
and the goja hosts (ios, android) recorded null for every element-valued
extractor in both the per-step diff and the violation witness.

Apply the web host's sanitize rule before marshaling, so one rule encodes
an element on both hosts.
2026-08-15 13:38:47 +05:30
pj b02e86b2e3 ci: dispatch workflows for folio and the replay ui (#73)
* feat(runner): stop the step loop at the first violation on request

* feat(testrun): report violations as a typed error under exit-on-violation

* feat(cli): add --exit-on-violation and exit 2 when it fires

* docs(cli): document --exit-on-violation, --max-steps, and exit codes

* fix(web): enumerate and query across shadow roots in both producers

* test(chrome): compare both producers on a shadow-dom parity page

* test(browser): drive a canvas-under-shadow-root fixture end to end

* fix(web): select the focused field inside a shadow root before typing

* fix(web): report the pathname as the screen when there is no hash route

* fix(web): settle on dom quiescence instead of returning at body ready

* feat(replay-ui): add data-testid hooks the dogfood spec drives

* feat(replay-ui): add the dogfood spec sanderling runs against the replay ui

* fix(replay-ui): scope the screenshot property to the named state panel

* chore(make): add per-platform sanderling build targets

* ci: add dispatch workflows for folio and the replay ui

* docs: describe the dispatch workflows and how to read a failure

* ci(folio): give the ios leg its jdk, android sdk, just, and a clean app start

* refactor(web): use max for the settle budget

* ci: pin calibrated seeds, skip the flaky ios reinstall, bound every job

* ci: authenticate and pin the buf setup step

the anonymous release download hit the shared runner ip rate limit and
failed the job with 'socket hang up' after three retries.

* docs: record that canvas apps need a dom proxy to be text-fuzzable

* fix(ios): bound lifecycle rpcs and claim the target device

a launch the simulator rejects sent the xctest session into a recovery
chain that answered minutes late or never, and the rpc had no deadline,
so the run hung with no trace and no error. also take a per-udid flock:
a second run's reinstall lands under the first's live automation session
and wedges it.

* docs(ci): correct the ios hang wording and note the device lock

* refactor(verifier): derive the lastAction shape from one field list

both hosts must show a spec the same lastAction. one ordered list now
feeds the goja object and the json the web host installs, so they
cannot drift.

* fix(web): install lastAction in the page before extractors read it

state.lastAction was hardcoded null on web, so every property reading it
was silently vacuous: a correct property passed without ever firing.

* fix(web): carry element identity on actions and fix findAll on paths

an action's target was coordinates only, so a property matching on which
element was acted upon could never fire. ax.findAll([a,b]) also returned
nothing on web.

* fix(chrome): wait out a route transition before sampling facts

the tree stays byte-identical and quiet across a cross-fade, so both the
quiet timer and the unchanged-tree escape called it settled mid-flight
and extractors read two screens at once.

* fix: bound the pre-run app launch

launch happens before the runner starts, so --duration never covered it
and a wedged driver hung with no trace and no error.

* fix(folio): read balances from merged cards and treat unreadable as unknown

compose for web merges the whole accountcard subtree, so the balance
child never exists there and every card parsed as 0. the property then
compared 0 to 0 and fired on any submit, which is a false positive
generator. unknown is now null and null is vacuously true.

* test(folio): cover merged-card parsing and unknown balances

* ci(folio): make web an expect-the-bug leg

the web runtime can observe the double submit now, so the health gate
understates it. seed 1 finds it at step 109, 3 runs out of 3.

* docs(ci): explain why a submit tap landing on home is the bug

* fix(ios): read a StaticText's label as its text

AXValue was the only source for text, but a StaticText carries its
string in AXLabel, so nothing on screen had .text on ios: a spec reading
it saw everything on android and nothing here.

* docs(ci): correct the calibrated step ranges

* fix(folio): stop convicting on arithmetic float64 cannot hold

past 2^53 cents the gap between representable values is 128, so a real
1600-cent move reads back as something else and the equality is false
for a healthy submit as readily as a double one. also match parseCents:
a sign or an oversized amount is rejected, not read as an amount.

* test(folio): pin the safe-integer guard and its boundary

* docs: stop teaching the zero-default that caused a false alarm

* docs: write down the silent-vacuity failure modes

* feat(folio): tag the home total and the card transaction count

the total was the only untagged node on the screen, so the spec had to
sum cards and a clipped card broke the sum.

* fix(folio): read the app's own total and refuse contaminated windows

summing cards went null when one was clipped, and the null poisoned the
carrier for the rest of the run. the balance window also spanned every
transaction since the last home visit, so the property convicted on
deltas it could not attribute: the old web witness was 3.16x the typed
amount, not 2x.

* test(folio): pin the window rules and the count invariant

* fix(folio): never read a frame that shows two screens

android dumps a cross-fade with both screens in the tree. the route said
add-transaction while an unscoped find said home, so the oracle took a
half-rendered total as fresh and convicted on a tap that committed
nothing. one function now decides the route and returns null when the
frame is ambiguous.

* test(folio): cover transition frames, card readings and creation

* fix(folio): only disambiguate counts that came from merged text

the equal-length digit rule exists because web merges the card and an
account named -1 makes '12' ambiguous. a dedicated count node has
nothing to disambiguate, so applying it there threw away real evidence.

* ci(folio): pin the recalibrated seeds and drop android to a health gate

web 3 and ios 7 convict 3 runs out of 3 with an exactly 2x witness.
android convicts 2 in 5 because the same seed does not walk the same
trajectory there, so it proves the app runs instead.

* docs(ci): describe the two properties and why android cannot convict

* fix(android): wait out a route cross-fade before snapshotting

the dump could hold two screens at once, and the runner refuses to act
on such a tree, so a quarter of android steps applied no action and the
count varied per run: the same seed never walked the same trajectory.
the ios companion and the chrome driver already do this.

* ci(folio): let the android leg run far enough to see its conviction

* docs: only the repo owner merges

* ci(folio): a thrown predicate is not a conviction

exit 2 means the run recorded a violation, and a predicate that throws
is recorded as one too. so was newAccountBalanceIsZero, an unrelated
property in the same spec. the gate read the exit code and went green
with detection dead.

* ci: install idb-companion from its tap and stop interpolating inputs

idb-companion is not in homebrew-core, so the ios leg died before it
built anything. replay-ui expanded dispatch inputs into the shell.

* docs: correct the snippets and numbers that drifted from the code

* test(sidecar): pin that a slow read counts toward the stability streak

* fix(web): read the page's extractors only on steps that count

the page advances the spec's carriers when it evaluates, but the runner
applied the result only on non-transitional steps. a discarded step
moved the window forward anyway, so the next accepted pair bracketed two
transactions while counting one submit, and convicted a healthy app.
extractor errors now fail the run instead of leaving goja's values in
current against v8's in previous.

* fix(chrome): anchor the transition deadline when the dom goes quiet

it was anchored at script start, so a page that churned past the window
reached the check already expired and returned mid cross-fade. the
driver now publishes the idle timeout it needs, since the caller's 1s
could never spend the 800ms window.

* fix(web): fail on a partial extractor override

same mixed-producer hazard as the install error: some extractors hold
the page's value and the rest hold goja's, and a property comparing
across that split fires on a healthy app.

* docs: six of seven, the seventh is the stock property

* fix(folio): drop a name two cards answer to

homeTxnCountsOf keyed on the account name and let the last card win, so
two accounts the fuzzer named the same collapsed into one entry. a
reading that saw one Travel card and a later one that saw both then
subtracted two different accounts' counts, and
submitCommitsOneTransactionPerAction convicted a healthy app of
double-submitting. it is a gated property in folio-run.sh, so that reads
as "found the submit bug" over a card scrolling into view.

same rule createdAccountHasNonZeroBalance already applies: a name
nothing can attribute is no evidence. counted over every card, since an
unreadable twin spoils the identity too.

* perf(folio): read each frame once

every extractor asked routeOf, and routeOf does five ax.find calls. on
web each find walks the document and every shadow root beneath it, so
the spec cost 110 tree walks a step; homeCards was parsed four times
over. now 5 and once.

keyed on the identity of the state object because both hosts build a new
one per step and hand that one object to every getter, so it cannot
outlive its frame. holding the reference is what keeps that true rather
than likely.

* fix(web): keep an undefined reading's index through JSON

json has no undefined, so an extractor whose getter returned one had its
whole index dropped by JSON.stringify. that index then kept goja's
dump-derived value while its neighbours held the page's, and a property
comparing previous to current across the split fires on a healthy app.
folio has nine on(route, tag) extractors, so this was most extractors on
most steps.

each reading is wrapped in a {value} envelope: the drop now happens
inside the entry, and an absent value means the getter returned
undefined, which is what the goja host records for the same getter. a
json null would instead claim it returned null and x.current ===
undefined would answer differently on the two hosts.

* feat(verifier): report the registered extractor count

the web path needs it to check the page sent one reading per extractor.

* fix(runner): fail when the page reports fewer readings than extractors

the comment here already claimed a partial override was fatal. it was
not: the skipped check only catches indices outside the extractor list,
so a page reporting values for some extractors and not others left the
rest holding goja's reading of the dump with nothing said.

* test(browser): drive an undefined reading through the whole web path

four layers carry it: the page's envelope, the driver's unwrap, the
runner's count check and the verifier's decode. each has a unit test and
only a run proves they compose. goes red both ways, decoding an absent
value as null and dropping the envelope.

* fix(web): offer the aria roles a user activates

only role=button was in the tappable set, so link, checkbox, radio,
switch, tab, option, the menuitems and treeitem were invisible to the
enumeration however plain the control looked. the replay ui builds its
step rows as <li role="option">, and the spec dogfooding it had to
hand-write an action to reach them because no default verb could see a
single row.

both producers build the set from the same role list, since the parity
test compares them element by element.

* test(browser): tap a role-based control end to end

every control on the page is an <li role="option">, the shape the
replay ui gives its step rows, and the spec carries no action of its
own: the property firing is the evidence the default enumeration offered
a tap on one.

* fix(web): read aria-disabled as disabled

the enabled fact came off the disabled property, which only real form
controls have. it reads undefined on the role-based controls the
tappable set now covers, so every one of them looked enabled however
plainly it was marked otherwise, and the fuzzer would spend actions on
inert ones.

both producers answer the same two ways, and the parity fixture carries
a disabled row so the comparison covers it: reverting one side alone
names the element and the fact.

* docs(replay-ui): the enumeration reaches step rows now

the comment said role="option" is not in the tappable selector set,
which stopped being true a few commits ago. selectAStep stays, for the
reason the tab weight below it stays: one row among the page's clickable
elements is a thin chance, and both step-facing properties go vacuous on
a run that never selects one.

* test(runner): bound the last-action test by steps, not wall clock

100ms of wall clock against an assertion that two steps ran fatals under
load with "the web path never installed it", which reads as a
regression. every sibling test in the package uses a long duration and
MaxSteps.

* ci: run the kotlin tests in make test

RouteTransitionTest and the stability poll cover the android settle and
nothing in ci ran them. :sidecar:test needs no android sdk, checked by
running it with ANDROID_HOME pointed at nothing.

* fix(sidecar): measure the stability streak as observed quiet

parameterising pollUntilStable also moved the clock to the start of the
read that opened a run of identical snapshots, so a read's own duration
counted as quiet. the pre-existing caller polls a real uiautomator dump:
at 400ms a read, 750ms of required quiet became 250ms of observed quiet
and the poll settled in two reads instead of four.

the parameters stay, the semantics go back.

* test(sidecar): pin the transition cap by driving it

it asserted 1500 >= 700 + 300, two constants, which can only fail if
someone edits a constant. it now drives awaitSettledTree against a fade
that lands after 700ms and asserts it hands back the settled tree before
the cap. cut the cap to 1000 and it goes red.

* ci: pin buf-setup-action to a commit

it takes a token now, so a floating tag is a token handed to whatever
that tag moves to. note v1 there is a branch, not a tag, so the ref
lookup that resolves it is matching-refs/heads/v1.

* ci: declare least-privilege permissions

none of the three declared any, so each got the repository default.
release.yml and docs.yml already do this. all three only check out,
build, test and upload artifacts.

* ci: fail fast when a server never comes up

the readiness loops fell through silently after 30 tries, so a server
that never started surfaced as an opaque driver failure minutes later.
each now says what did not answer and on which port.

* ci(folio): a missing trace is not a verdict

with no trace the android gate ran its grep against ./trace.jsonl and
reported "never reached AddTransactionScreen, so it never got past
login", which is not what happened. the web and ios branches had the
same misdiagnosis on exit 0.

same class, one line up: the classifier's own failure was swallowed, so
with the evidence reader dead the gate printed a healthy run and exited
0.

* ci(replay-ui): skip a run directory with no trace

the summarise step is if: always(), and under github's bash -eo pipefail
an unmatched glob stays literal, the redirect fails, pipefail carries it
into the assignment and -e kills the step. so a failed fuzz run went red
twice, once for the real reason.
2026-08-15 13:01:27 +05:30
pj 1f71e052d7 clean up dead jetbrains mono wiring in replay-ui (#70)
* fix(replay-ui): drop @font-face rules for fonts that were never shipped

* fix(replay-ui): drop unresolvable JetBrains Mono from --font-mono stack

* chore(replay-ui): remove vestigial empty public/fonts dir
2026-08-12 23:18:03 +05:30
pj 76dce1a75e experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs

runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record arm membership and host in meta.json

meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --arm and populate run meta from it

Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): sweep seeds for one experiment cell

campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.

Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): no silent generator fallback, and llm on web

--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.

pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): make the hierarchy dump agree with the web runtime

Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.

scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.

The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(conformance): give the gate reproducible seeds

SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.

The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit editable as a plain boolean

editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): leave the head subtree out of the web target walk

collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(chrome): compare the facts both hosts derive from one DOM

The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.

This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* chore(make): run the browser packages one at a time

Both launch Chrome and launching two at once has failed with "Launch: context
canceled".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style: remove every em-dash and en-dash

Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): honor the caller context in Launch

Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.

The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(sidecarassets): publish the extracted jar through a rename

Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): kill a run that outlives --run-timeout

A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style(test): gofmt browser_test.go

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:20:31 +05:30
pj 26b49b379a fix ltl semantics and unify action enumeration (#71)
* fix(ltl): give every thunk a construction identity

Two distinct unnamed predicates both described as "Thunk(...)", so obligation
collapse merged their residuals and could drop a live violation. Identity is
assigned at construction and the fields are unexported, so a thunk cannot be
built without one.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): reduce a thrown-predicate residual instead of panicking

The verifier substitutes an ErrorFormula for the residual of a property whose
predicate threw, and that residual is fed back in on the next step. reduce had
no case for it, so the run crashed. It re-reports the same failure now.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): make a bounded always the dual of a bounded eventually

G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still
pending when the window closed, so nnf's negation normal form was not semantics
preserving. Both sides now range over the observations at which their inner can
definitely resolve: the eventually keeps a pending inner as a disjunct, and the
always discharges vacuously at window close.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): arm a one-shot root once per run

A root that carries its own horizon is one obligation for the whole run, not one
per observation. Re-instantiating a top-level eventually monitored G F<=n(p)
instead of F<=n(p) and left one live obligation per step behind; a bounded
always restarted its window every step and never closed.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): stop wrapping a top-level eventually in always

`eventually(p).within(300, "seconds")` as a property meant "within 300 seconds
of every step", which spawned an obligation per step with its own resolved
deadline. A 553-step run carried 553 of them and serialized a 75 KB residual.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): serialize the resolved deadline of a bounded window

Two obligations spawned at different steps from one duration-bounded formula
differ only in the deadline the evaluator resolved for them, so they serialized
identically and the trace erased a distinction the evaluator makes. The authored
window stays in amount/unit; the resolved deadline rides alongside.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): split a witness's origin step from its detection step

A deferred obligation spans two steps: the one that armed it and the one whose
reduction failed. They were conflated under one index, so the extractor snapshot
(which is the detecting step's state) was reported against the origin step.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): record a witness's detection step in the trace

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* feat(replay-ui): show the step a violation was detected at

The witness evidence is the detecting step's state, so say which step that is
and let a reader jump to it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): record the extractor state the predicates actually read

On the web path extractor bodies are evaluated in V8 and injected here, but only
the goja value was replaced. The trace diff and the violation witness therefore
described a state no property ever saw.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* refactor(spec): one candidate producer over one target-eligibility rule

Both hosts routed verbs themselves and both policies enumerated their own
actions, and all four drifted. Web sent `swipes` to scrollable containers only,
so swipe-to-dismiss on a list row was reachable on native and unreachable on
web; the model policy folded gestures its own way and could not reach what the
seeded picker drew.

A host now reports facts about every element and never decides which verb may
act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates
is the single enumeration, and the model policy reads it through
__sanderlingEnumerateBuiltin__ instead of reimplementing it in Go.

Gesture verbs change with it: scrolls stay vertical over scrollable containers,
swipes go free-form in all four directions from any element with real bounds.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): name a builtin scroll by its drag origin

A builtin gesture carries endpoints and no selector, so every scroll rendered as
"Scroll down " in the prompt's recent-action memory and two scrollable regions
were indistinguishable.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): clear storage over cdp instead of scripting an opaque origin

Launch runs while the tab is still on about:blank, whose opaque origin denies
storage access, so localStorage.clear() threw SecurityError and every web run
died at launch. Storage.clearDataForOrigin needs no navigation. The exception
helper lands here because "Uncaught" is what hid this for so long.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): enable the swiftshader webgl fallback

Headless Chrome runs with --disable-gpu, and without this flag it refuses the
software WebGL backend: getContext returns null, so a canvas-rendered app paints
nothing and every screenshot is identical black.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(web): resolve testTag through data-testid or id

Compose Multiplatform emits its testTag into the element id, which the native
table already accepts via the resource-id alias. The two web selector tables
were the only place that rejected it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* test(spec): type-check the spec api as part of make test

The fake runtime in api.test.ts did not return a chainable handle from extract,
so the file had not type-checked since named() was added. Wiring the check into
make test stops it drifting again.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* docs(manual): one-shot eventually and the gesture verbs

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
2026-08-12 18:06:04 +05:30
pj 7343085614 llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker

* feat(spec): make llm marker inert on the JS picker

* feat(spec): expose __sanderlingSampleInput__ corpus draw

* feat(openrouter): minimal chat-completions client

* test(openrouter): cover request shape, parse, and errors

* feat(verifier): thread screenshot + capture corpus sampler

* feat(verifier): LLM accessors — candidates, config, sampler

* test(verifier): cover AllCandidates, LLMConfig, SampleInput

* feat(trace): record action Source and LLMReasoning

* feat(runner): thread step screenshot into PushSnapshot

* feat(runner): llmSource selects actions via OpenRouter

* feat(runner): wire llmSource selection and trace stamping

* test(runner): cover llmSource selection, mapping, downscale

* docs(folio): add llm action-backend example spec

* docs(folio): document the LLM action backend run

* feat(llmclient): support OPENAI_API_KEY, openrouter wins

* refactor(runner): rename openrouter package to llmclient

* docs: both api keys, example model gpt-5.4-nano

* docs: add pr style rules to claude.md

* fix(runner): explain action kinds in llm prompt to stop swipe loops

* feat(trace): record llm ranked list and chosen rank

* feat(runner): stamp llm ranked list and chosen rank on trace

* fix(runner): tap by selector to survive layout shift after observe

* revert(runner): drop selector-first tap; broke path/testTag selectors

* feat(spec): llm() accepts optional instructions

* feat(verifier): read llm instructions off config

* feat(runner): append spec instructions to llm system prompt

* docs(folio): describe app in llm spec instructions

* feat(bundler): map generator export to globalThis.generator

* feat(verifier): read llm config off globalThis.generator

* feat(runner): gate llm source on --generator flag

* feat(cmd): add --generator llm|seeded flag

* test: cover --generator flag parsing and pickSources gating

* feat(verifier): enumerate llm candidates by walking actionsRoot

collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.

* test(verifier): cover candidate enumeration walk

* feat(verifier): add SetupAction to walk setup without the seeded root

* test(verifier): cover SetupAction setup-only precedence

* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order

* feat(trace): record llm choice number and chosen_action echo

* feat(runner): llm picks one number from weighted candidates

drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).

* test(runner): cover choice schema, strict-skip, and setup precedence

* refactor(verifier): drop the superseded AllCandidates enumeration

* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts

* fix(verifier): label editable fields by hint, not the typed value

an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.

* test(runner): cover weight-suffixed echo and stripWeightSuffix

* fix(runner): accept chosen_action echo that carries the weight suffix

real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).

* fix(verifier): skip llm enumeration on cross-fade frames

a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.

* feat(folio): show current balance on the add-transaction screen

renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.

* fix(replay): derive device space from screen extent, not first node

the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.

* fix(folio): show balance as a compact one-line label

per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.

* fix(folio): move balance into the header, one compact line under the account name

* fix(replay): attribute deferred violations to the causing step, not detection

* fix(replay): show a step's own violations in both panels, no next-step bleed

* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check

* chore: ignore .playwright-mcp scratch output

* docs: document the llm generator and --generator flag

* docs(spec): correct the llm() comment; config reads off globalThis.generator

* docs: add pr description rules
2026-07-31 21:12:00 +05:30
pj 6b0d6cb971 WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial

* feat(test): add --device flag to target a specific Android device by serial

* feat(folio): select Android device via ANDROID_DEVICE in justfile

* feat(conformance): add android backend to the gate suite

* feat(android): keep device awake and unlocked so the app stays foreground

* feat(conformance): prep physical android device (autofill/verifier/stayon)

* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run

* fix(verifier): require positive bounds for swipe candidates

A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.

* fix(runner): harden app-scope guard against launcher and overlays

The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.

* feat(android): harden physical-device runs in device prep

Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.

* feat(driver): clear-state via APK reinstall when pm clear is blocked

When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.

* feat(cli): add --android-app-path for clear-state reinstall

Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.

* chore(folio): pass --android-app-path in just test

* fix(runner): clamp swipe/scroll origin out of edge gesture zones

A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.

* perf(sidecar): faster Android text input and drop redundant settle poll

inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.

* fix(verifier): exclude soft-keyboard region from action candidates

The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.

* perf(runner): replace focus-tap settle with a brief wait

The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.

* chore(conformance): platform-aware G5 p95 budget for android

The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.

* fix(sidecar): retry maestro android driver startup

The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.

* chore(conformance): widen android G5 budget to 5500ms

Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.

* web replay fix

* feat(android): force 3-button nav during runs to prevent app drift

On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.

* fix(android): target the selected device in adb reads; don't strand nav mode

Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
  foreground/scope guard works when several devices are attached (the --device
  path). Previously they ran bare `adb shell`, which errors with multiple
  devices, silently disabling app-scope enforcement. The sidecar client passes
  its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
  duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
  the current mode is unknown or already 3-button it leaves nav untouched,
  instead of switching and then stranding the device in 3-button. Logic split
  into the pure navModeToRestore, now unit tested.

* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds

Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.

* test(verifier): cover keyboardRegionTop, including the decor-view guard

The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.

* style(cli): gofmt testOptions field alignment

* fix(sidecar): keep a leading dash off the fast input path

A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).

* refactor(verifier): scope action candidates by window ownership

Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.

This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.

* fix(runner): re-check foreground at apply time, skip stale actions

ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.

* fix(android): type long ASCII via fast guarded path to stop keystroke escape

A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.

* chore: ignore gate artifacts and local scratch files

* refactor(runner): narrow gesture clamp to the top shade strip

3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.

* chore(format): add .editorconfig enforcing 80-column limit

* chore(format): add prettier config with 80-char printWidth

* chore(deps): add prettier devDependency to replay-ui

* chore(deps): add prettier devDependency to folio-web

* chore(deps): add prettier devDependency to spec package

* chore(format): add swift-format config with 80-char lineLength

* feat(format): add make fmt targets for per-language 80-col formatting

* fix(runner): translate gesture to safe area so near-top scrolls keep direction

Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.

* fix(runner): apply-time guard consults focused window, not just resumed activity

ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.

* test(runner): cover apply-time foreground skip and appIsForeground table

Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.

* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker

- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.

* fix(conformance): pin self-test p95 budget and score install failures as run failures

self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.

* fix(android): require --device when several devices are connected

With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.

* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands

The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.

* perf(verifier): memoize scopedElements per tree

scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.

* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear

Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.

* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms

The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.

* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty

bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.

* perf(sidecar): reuse a single Jackson ObjectMapper

structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.

* refactor(android): remove unused AdbReverse/AdbReverseRemove

No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.

* style(runner): trim non-load-bearing comments from this PR's runner code and tests

* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
2026-06-11 10:10:05 +05:30
pj 991c583eb9 docs update with case study (#63)
* docs(manual): add introduction page

* docs(manual): rewrite getting started as guided first run

* docs(manual): rewrite writing specs as a folio tutorial

* docs(manual): document missing spec API in reference

* docs(manual): plain-language rewrite of runs page

* docs: real introductions on index pages and README

* fix(docs): sibling links from directory-style pages need ../

* fix(docs): correct sampling and restart-cost claims to match implementation

* docs: nav lists Introduction and Case study; roadmap points to milestone

* docs(manual): make getting started target the reader's own app, not Folio

* docs(manual): add Folio case study page

* docs: point manual navigation at the case study

* docs(readme): lead with the case study, fix roadmap link

* docs: roadmap links to milestone, sync clear-data default and cross-links
2026-06-09 20:13:54 +05:30
pj 90224dfd06 Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers

The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.

* feat(ios): resolve physical devices from devicectl

ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.

* feat(ioscompanion): runner-only device driver mode

NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.

* test(ioscompanion): cover device wiring, routing, and shell-out argv

Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.

* feat(testrun): route physical-device iOS runs to the device driver

Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.

* feat(doctor): device prereqs replace java/sidecar for ios-device

iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.

* feat(conformance): device backend uses iphoneos app and tunnel orphan checks

The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).

* feat(folio): device build linking the iosArm64 framework

project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.

* docs(cli): document ios-device doctor checks and the device flags

The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.

* fix(ioscompanion): resolve signing key path to absolute

xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.

* fix(ioscompanion): re-enable signing for the device runner build

companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.

* fix(ioscompanion): key the device build cache on signing identity

The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.

* docs(getting-started): document physical iOS device setup

Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.

* feat(ios): native usbmux client and in-process tunnel forwarder

Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.

* refactor(ios): drive device tunnel via io.Closer seam

Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.

* refactor(ios): remove iproxy spawn from device runner

* test(ios): cover tunnel close via io.Closer not child process

* feat(doctor): check usbmuxd socket instead of iproxy on PATH

* chore(conformance): drop iproxy orphan check; tunnel is in-process

* docs(ios): device tunnel uses native usbmux, nothing to install

* chore: gitignore the signing keys directory

* feat(folio): add Android launcher icon (black bg, white dot)

* feat(folio): add iOS app icon (black bg, white dot)

* feat(folio): add web favicon (black bg, white dot)

* docs(ioscompanion): fix stale const comments

* refactor(ioscompanion): inline single-use devicectl argv builders

* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests

* refactor(ioscompanion): inline firstNonEmpty

* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam

* test(doctor): trim redundant signing-check test

* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env

* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
2026-06-09 18:38:52 +05:30
pj e04631d4c9 Remove the JVM sidecar's iOS backend (#65)
* refactor(sidecar): drop IosDriverBackend

* refactor(sidecar): route ios platform off the iOS backend

* chore(sidecar): remove maestro ios dependencies

* feat(testrun): reject physical iOS with a clear message

* refactor(sidecar): drop iOS hierarchy helpers and their test

* build(sidecar): strip iOS runner bundles and classes from the fat jar

* test(testrun): cover physical-iOS rejection
2026-06-08 20:29:35 +05:30
pj 406b7516b3 iOS simulator driver: Go-native companion-backed backend (#62)
* perf(ios): use prebuilt XCTest runner to cut startup

* chore(ioscompanion): add companion asset prepare script

* feat(ioscompanion): embed and extract simulator companion bundle

* test(ioscompanion): cover companion stub and embedded extraction

* docs: add third party notices for vendored companion

* chore: ignore vendored companion bundle artifact

* build(proto): pin simulator companion proto v1.1.8

* build(proto): add dedicated buf module and gen template for pinned proto

* build(proto): exclude pinned companion proto from root buf workspace

* feat(ioscompanion): commit generated companion gRPC stubs

* feat(ioscompanion): map flat companion describe dump to TreeNode JSON

* test(ioscompanion): add hierarchy-map golden and unit tests

* feat(ioscompanion): port screen-settle stability polling to Go

* test(ioscompanion): cover settle transitional, hash, streak, and cap rules

* feat(ioscompanion): add USB HID keymap module

* test(ioscompanion): cover keymap branches and paste-chord constants

* build: embed companion assets via withcompanion tag

* feat(ioscompanion): add transport companion interface

* feat(ioscompanion): add HID event wrapper and builders

* feat(ioscompanion): wire gRPC companion client and Dial

* test(ioscompanion): cover HID builders and unit conversions

* test(ioscompanion): cover Dial, process-state mapping, and install archive

* test(ioscompanion): add gated simulator integration smoke test

* feat(ioscompanion): text input and gesture HID composition with pasteboard fallback

* test(ioscompanion): cover input composers, paste dialog loop, and pure helpers

* feat(ioscompanion): add Describe to companion transport

* feat(ioscompanion): implement DeviceDriver with companion supervision

* test(ioscompanion): unit tests with fake companion transport

* test(ioscompanion): gated companion smoke test

* feat(ios): add ResolveTarget for simulator vs physical-device routing

* feat(testrun): route iOS simulators through the native companion driver

* refactor(testrun): defer the java preflight check to the physical-device path

* feat(cli): add --ios-app-path flag

* feat(doctor): split iOS checks into simulator and physical-device paths

* test(folio): add gate-analyzer fixtures for G1-G5

* feat(folio): add iOS conformance gate script

* chore(folio): wire gates recipe, app path, and ignore gate output

* style: gofmt struct alignment drift

* fix(doctor): probe simctl via xcrun instead of PATH lookup

* fix(ioscompanion): spawn companion under driver-lifetime context

* test(ioscompanion): prove companion child outlives startup context

* fix(ioscompanion): chunk install payload under companion message cap

* test(ioscompanion): cover install payload chunking

* fix(ioscompanion): reinstall via simctl and sanitize companion env

* fix(ioscompanion): wait out unresolved accessibility values after launch

* perf(ioscompanion): paste long text for atomic landing

* test(ioscompanion): cover paste threshold, retry flow, and sentinel detection

* fix(ioscompanion): treat unresolved bridge values as transitional, never as content

* fix(ioscompanion): accept masked secure-field values as paste landing

* test(ioscompanion): cover sentinel mapping and masked-field landing

* fix(ioscompanion): atomic erase and single-send paste to prevent doubling

* test(ioscompanion): cover atomic erase, single chord, unverifiable field

* fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout

* test(ioscompanion): cover bridge-blackout paste verification

* fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle

* refactor(ioscompanion): name the empty-editable-field sentinel for what it is

* perf(ioscompanion): tighten settle streak for the fast companion transport

* feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt

* refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt

* test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests

* fix(ioscompanion): retry describe past transient collapsed accessibility dumps

* test(ioscompanion): cover collapsed-dump detection

* perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses

* perf(ioscompanion): tighten settle now that collapses are handled separately

* fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text

* test(ioscompanion): cover replace-on-input and TextReplacer capability

* refactor(ioscompanion): neutralize HID events behind the transport seam

* feat(companion): add simulator runner project skeleton

* feat(companion): serve accessibility snapshots over the wire protocol

* feat(companion): synthesize timestamped touch gestures

* feat(companion): type text with replace semantics

* feat(companion): serve the wire protocol from a parked runner

* feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam

* feat(ioscompanion): route text input through a text-editing companion when available

* fix(companion): bind listener by port and source screen size from snapshot

* feat(ioscompanion): add runner companion JSON transport

* test(ioscompanion): cover runner transport protocol mapping

* fix(companion): synthesize gestures synchronously to avoid the async completion crash

* fix(companion): type on the main thread and recover from focus assertions

* fix(companion): keep serving after an automation failure

* refactor(companion): tidy snapshot serialization

* fix(companion): honor sequential tap gaps and survive synthesis exceptions

* feat(ioscompanion): expose native typing with an explicit replace flag

* chore(companion): add runner asset prepare script

* feat(ioscompanion): embed and extract the runner test bundle

* test(ioscompanion): cover runner asset extraction

* build(ioscompanion): commit runner asset archive

* feat(ioscompanion): pair the legacy companion with the in-simulator runner

* test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding

* fix(ioscompanion): reconnect after interrupted runner calls instead of restarting

* fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts

* feat(companion): launch and terminate apps through the automation session

* build(ioscompanion): refresh runner asset with session lifecycle

* fix(ioscompanion): classify connection deadline expiry as caller budget

* fix(companion): capture snapshots on the main thread inside the catch bridge

* build(ioscompanion): refresh runner asset with main-thread snapshots

* perf(ioscompanion): count read spans toward settle and capture snapshots concurrently

* feat(ioscompanion): make the hybrid simulator companion the default

* test(folio): cover runner-session orphans in the gate harness

* test(ioscompanion): pin the child-lifetime test to the legacy path

* fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears

* fix(ioscompanion): pause the clear chord so selection applies before the delete

* fix(companion): prune the keyboard subtree from snapshots

* build(ioscompanion): refresh runner asset without keyboard elements

* fix(ioscompanion): capture the screenshot transport before a recovery can reassign it

* fix(companion): pin the runner listener to loopback

* fix(companion): size the replace delete prefix to cover any focused field

* build(ioscompanion): refresh runner asset with loopback bind and replace fix

* fix(cli): cancel the run context on SIGINT so spawned children are reaped

* fix(testrun): point the device java preflight hint at the ios-device doctor

* fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob

* test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision

* chore: add test-companion target for the withcompanion-tagged suite

* chore(ioscompanion): stop tracking the runner archive build artifact

* build: produce the runner archive from source like the companion bundle

* refactor(conformance): move the gate harness out of examples/folio

* chore(folio): drop the gate harness wiring from the example app
2026-06-08 19:10:54 +05:30
pj 94d9511312 test: full test-suite refactor sweep (#61)
* chore(test): start test-suite refactor sweep

* test(ltl): pin exact multi-obligation residual AST

* test(ltl): table-test finalize Kleene connective combinations

* test(ltl): pin reduce over pending inner for bound, Or, Not

* test(ltl): marshal bounded Always steps/duration/deadline

* test(verifier): cover LTL combinator verdict transitions and within unit panic

* test(verifier): table-test DecodeAction kinds and lastAction field exposure

* test(verifier): assert WithPlatform(ios) reaches the picker host and key pool

* test(verifier): widen weighted-selection assertion to a 5x skew margin

* test(verifier): un-skip ax-find round trip with a committed tree fixture

* test(runner): pin isWDADrop to sidecar reconnect-failed message origin

* test(runner): assert PressKey/Wait trace encoding records kind-specific fields

* test(runner): cover RenderSummary unsupported-verbs surfacing branch

* test(trace): set Hierarchy in round-trip and lock lossy Tree contract

Also add a -race concurrent WriteStep test that asserts N well-formed JSONL lines, catching torn lines if the writer mutex is dropped.

* test(trace): round-trip witnesses/changes/metrics/exceptions, pin step-0 witness

* test(trace): document ViolationsAreGreppable grep contract and lock-free WriteScreenshot

* test(hierarchy): cover invalid-JSON and malformed-bounds parser paths

* test(trace): guard writer mutex via WriteStep/Close race on w.file

* test(replay): drop unfailable assets and devproxy assertions

* test(replay): cache reuses on equal mtime, reparses after append

* test(replay): violation marker falls back to detection step when attributed missing

* test(replay): corrupt meta/trace dirs return 500 with error body

* test(replay): SSE client receives runs.changed after a broadcast

* test(replay): Run coalesces creates, ignores write/chmod, closes subs on cancel

* fix(sidecar): synchronize health fixture writes and exercise healthError

* test(sidecar): cover swipe/longpress/doubletap/erase/presskey/metrics/logs translations

* test(sidecar): cover DoubleTapSelector composition and mid-gesture cancel

* test(sidecar): assert gRPC error status surfaces from action RPC

* fix(chrome): route action methods through runCtx so caller cancellation aborts CDP

* fix(chrome): route hierarchy/screenshot/waitidle/metrics through runCtx

* refactor(ios): extract pure simctl JSON parsers

* refactor(ios): add command-runner seams for EnsureSimulator

* test(ios): table-test simctl parsers and EnsureSimulator seams

* test(sidecarassets): cover placeholder build path

* test(sidecarassets): assert reuse via sentinel bytes not mtime

* test(bundler): cover properties-only spec registration

* refactor(testrun): extract prepareBundleInputs from Execute

* test(testrun): cover prepareBundleInputs aliases and missing-runtime error

* test(testrun): table-test resolveRuntimeSibling search edges

* test(testrun): exact-output tests for progressHandler line format

* fix(cmd): point bundle-check aliases at pkg/spec/src

* test(cmd): smoke-test bundle-check resolves spec aliases

* test(cmd): table-test hier-check parse and FindAll on fixture

* test(cmd): unit-test buildBrowseURL deep-link vs root

* test(cmd): drop flaky TestRun_Doctor that launched real Chromium

* test(cmd): pin pipeline error to bundle resolution on web platform

* test(replay-ui): add bun test script

* ci(replay-ui): run bun test via make web-test target

* ci(replay-ui): point bun cache key at replay-ui/bun.lock

* test(replay-ui): exercise real URL encoding and non-ok throw in getJson

* refactor(replay-ui): extract snapshot flatten/getAtPath into lib module

* test(replay-ui): pin snapshot flatten/getAtPath path round-trip

* refactor(replay-ui): extract action selector/format into lib module

* test(replay-ui): pin action selector parse and row formatting

* refactor(replay-ui): share one statusFor between panels

* refactor(replay-ui): extract run-history derivation into lib module

* test(replay-ui): pin shared statusFor precedence and ordering

* test(replay-ui): pin run-history derivation alignment

* refactor(replay-ui): export clampIndex for testing

* refactor(replay-ui): extract keyboard-nav dispatch into pure module

* refactor(replay-ui): extract metrics formatters into lib module

* test(replay-ui): pin clampIndex step boundaries

* test(replay-ui): pin keyboard-nav ownership and key routing

* test(replay-ui): pin metrics formatters and path gap handling

* refactor(sidecar): expose device-output parsers as internal for testing

* test(sidecar): table-test device-output parsers against malformed input

* test(sidecar): cover logcat parsing year inference and line skipping

* test(sidecar): pin pressKey keycode mapping and unknown-key rejection

* test(sidecar): metrics bundleId falls back to launched app and honors override

* test(sidecar): loosen deadline upper bound to tolerate slow CI scheduling

* test(web-runtime): export selector builders for unit tests

* test(web-runtime): guard sanitize cycle, function, and depth limits

* test(web-runtime): table-test selector builder quoting and escaping

* test(sidecar): collapse scalar-forwarding RPC tests into a table

* test(replay-ui): dedup step/summary fixtures into shared module

* test(ios): collapse pickSimulator point-tests into a table
2026-06-06 13:59:08 +05:30
pj 410602d2e1 Fix action-system audit findings (#60)
* fix(replay-ui): size overlay viewBox from hierarchy root bounds

Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.

* fix(runner): derive trace tap point from resolveCoordinates

stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.

* fix(runner): settle after InputText focus tap before key events

The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.

* fix(driver): skip pre-erase for replace-on-input drivers

The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.

* fix(hierarchy): rank spatial-fallback matches by specificity

The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.

* fix(runner): treat an unchanging transitional tree as settled

A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.

* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf

The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.

* fix(sidecar): never replay non-idempotent actions after reconnect

A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.

* fix(sidecar): land the second double-tap sequentially on gesture collision

The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.

* fix(sidecar): map non-Exception throwables to INTERNAL status

The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.

* feat(sidecar): close the driver and app under test on shutdown

* test(sidecar): cover service shutdown paths

* fix(testrun): stop the sidecar with SIGTERM before killing

* fix(sidecar): reap orphaned XCTest runner sessions at iOS init

* fix(sidecar): probe channel liveness before restarting the XCTest runner

* test(sidecar): cover WdaRecovery restart and retry policy

* fix(sidecar): absorb first-leg double-tap collision sequentially

* fix(runner): scope WDA-drop detection and cap consecutive transient failures

* chore(sidecar): silence vendored loggers on expected failure paths

* fix(runner): absorb one-off apply errors; only an unbroken streak aborts

* fix(folio): install the current build before the Android fuzz run

* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs

The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
2026-06-06 12:21:11 +05:30
pj 9958b0ddc8 Attribute violations to the causing step and render witness evidence (#59)
* feat(ltl): attribute violations to the obligation origin step

* feat(verifier): label evaluator observations with the runner step index

* feat(trace): carry the causing step in violation witnesses and summary

* feat(replay): move the violation marker to the causing step

* feat(replay-ui): render witness evidence in the violations panel

* feat(replay-ui): wire witnesses and step jump into violation panels

* fix(ltl): treat next obligations as vacuous at run end

* fix(runner): give the finalize trace record its own step index
2026-06-05 23:46:36 +05:30
pj 19a470121f fix: iOS spec driving, InputText replace semantics, native DoubleTap (#58)
* fix(hierarchy): bounds-containment fallback for scoped and path queries

Compose on iOS surfaces a testTag node as an empty leaf sibling of the
content it labels instead of as an ancestor, so descendant search under
the tagged node finds nothing and every path or scoped query returns
null. When structural search yields no match, fall back to nodes whose
bounds lie inside the scope node's bounds.

* feat(sidecar): derive iOS clickable and editable from element type

The XCTest hierarchy mapping dropped the element type, leaving no
clickable or editable flags on iOS, so the fuzzer's tap and typing
verbs never found a candidate inside the app. Map the raw
accessibility tree directly and derive clickable, editable,
scrollable, and class from the XCUIElementType raw value.

* feat(proto): add EraseText RPC for InputText replace semantics

* feat(driver): add EraseText to the device driver surface

* fix(runner): erase existing field text before InputText

InputText appended on native platforms, so repeated draws grew fields
without bound. The folio fuzz run wedged on the add-account screen:
each draw concatenated another name until the 40-character validation
error became permanent. Replace semantics also makes retried typing
idempotent. The web driver already replaced via select-all; native now
matches.

* feat(sidecar): EraseText backend support on android and ios

* fix(folio): saturation-gate account creation in the spec

The 2-3 step add-account loop outcompeted the 5-step transaction chain
at every weighted re-draw, so runs filled with account creation and
rarely exercised the balance properties. Stop offering add-account once
three accounts exist; the renormalized weights then favor the
transaction flow at every step of its chain.

* fix(folio): author spec weights to match testing intent

Revert the account saturation gate: it starved newAccountBalanceIsZero
once it tripped, and a magic account count is app-state tuning, not
intent. Instead weight the generators by what the properties need:
the transaction chain leads, account creation stays exercised, and
doubleTaps gets explicit weight everywhere because double-submission
idempotency is what the spec is testing for.

* fix(folio): lower doubleTaps weight to 5

* fix(sidecar): surface visible text on iOS static elements

Static text and button strings live in the accessibility label on
iOS, so the text attribute came through empty and every balance
extractor parsed to zero, silently disarming both folio properties.
Non-editable elements now fall back title, value, then label;
editable fields keep value-only so an empty field's caption does not
read as content.

* feat(driver): native DoubleTap RPC for a tight inter-tap gap

Composing two Tap round trips from the Go client spread the taps by
hundreds of milliseconds on iOS, wide enough for the app to navigate
between them, so double-submission races could never reproduce. The
sidecar now lands both taps back-to-back next to the device transport.

* feat(sidecar): pipeline iOS double-tap requests

Queue the second tap at the XCTest runner while the first executes.
The runner serializes handlers, so this is the tightest gap the
transport allows (~350ms per tap round trip); recorded here with
measurements for the iOS double-tap limitation.
2026-06-05 23:42:39 +05:30
pj ce65cc32a4 Fix gradle deprecation warnings (#57)
* chore(build): add foojay toolchain resolver to silence Gradle 9 toolchain warning

* chore(sidecar): upgrade protobuf gradle plugin 0.9.5 -> 0.10.0

0.10.0 uses single-string notation internally when resolving platform-specific
protoc and grpc-java binaries, eliminating two Gradle 9 deprecation warnings.
2026-06-03 17:35:24 +05:30
pj b44077afde replay ui fix (#56)
* refactor: rename inspect to replay across the codebase

Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/,
the CLI subcommand from `sanderling inspect` to `sanderling replay`, and
updates all references in docs, Makefile, README, and Go comments.

* feat(replay-ui): show spec filename with full path on hover

RunList and RunDetail now render the basename of spec_path (e.g.
login.spec.ts) with the full path available as a title tooltip.
2026-06-03 16:17:26 +05:30
pj a6f43e15b2 Update README for clarity and roadmap link 2026-06-02 13:28:46 +05:30
pj c5bb176be8 UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks

Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.

* feat(ltl): negation normal form pass

nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.

* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse

Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.

* test(ltl): property-based NNF laws

Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.

* test(ltl): Finalize, bounded eventually, latch, collapse

Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.

* feat(inspect): within clause on always residual node

A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.

* feat(ltl): witness violations and (bool,error) predicate thunks

* test(ltl): migrate thunk call sites to (bool,error)

* feat(ltl): flag thrown-predicate witnesses with IsError

* refactor(verifier): replace predicate err side-channel with violation witness

* test(verifier): witness API for thrown predicates

* feat(trace): witnesses map and skipped-verification marker on Step

* feat(runner): thread violation witnesses, finalize, skip marker into trace

* test(ltl): lock violation witness reason, IsError, and step

* test(verifier): finalize surfaces unmet eventually with witness

* fix(ltl): eliminate implies and bounded-always false-negatives

Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.

* test(ltl): lock implies and bounded-always false-negative regressions

* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick

* feat(testrun): inject seed into web bundle via SANDERLING_SEED define

* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring

* test(spec): add Go math/rand/v2 PCG oracle and golden fixture

* feat(spec): bit-exact PCG port of Go math/rand/v2

* test(spec): assert pcg.ts matches the PCG golden fixture

* feat(spec): shared input corpus and press-key pools

* feat(spec): action-tree types and Host interface

* feat(spec): verb support matrix and warn-once helper

* feat(spec): deterministic shared action picker

* test(spec): verb matrix and warn-once semantics

* test(spec): picker draw-order and determinism

* refactor(spec): actions.ts returns pure GeneratorNode data trees

* refactor(spec): wire from() sampling through the picker rng

* feat(spec): shared runtime-entry installs next-action over pick.ts

* feat(spec): export LongPress/Scroll/longPresses/scrolls factories

* test(spec): assert data-tree shapes for action factories

* test(spec): runtime-entry serializeAction wire-contract round-trip

* refactor(spec): bridge data-tree nodes to the legacy goja picker tags

* fix(spec): web runtime walks the spec's globalThis.actions data tree

* test(spec): tolerate legacy bridge fields on builtin nodes

* refactor(spec): installRuntime accepts a lazy root resolver

The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.

* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker

Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).

* test(spec): cover the WEB Host surface and seed precision

Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.

* refactor(spec): picker emits native selector + scroll endpoints, setup precedence

* feat(spec): goja runtime entry wires the shared picker over the Go host

* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin

* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker

* refactor(spec): drop the legacy goja bridge fields from action factories

* feat(spec): serialize selector-only string targets for the runner to re-resolve

* refactor(verifier): one DecodeAction reads the unified flat wire contract

* refactor(verifier): goja host + shared picker replace the duplicate Go picker

* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime

* test(verifier): author specs through the shared picker path

* test(runner): bundle authored specs with the goja runtime entry

* feat(verifier): collect unsupported verbs for the run report

* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource

* feat(testrun): surface unsupported verbs in run report

* test(verifier): cross-runtime goja/node parity gate on the shared picker

* test(verifier): unsupported verbs collected deduped in first-seen order

* test(runner): summary reports no unsupported verbs on a clean run

* test(spec): golden-fixture cross-runtime parity gate for the node picker

Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.

* test(verifier): assert goja picker against the same cross-runtime golden

Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.

* refactor(spec): rename pressKey generator export to pressKeys

* refactor(spec): update barrel re-exports for pressKeys

* test(spec): update pressKeys generator export name

* docs(spec): rename pressKey generator to pressKeys

* refactor(spec): extract samplerRng into shared sampler-rng module

* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)

* test(spec): cover fluent value generators determinism and chaining

* refactor(bundler): inject globalThis trailer from spec named exports

* refactor(bundler): reuse registration trailer in web bundler

* test(bundler): cover named-export globalThis registration

* feat(spec): add named() to Extracted handle type

* feat(web-runtime): named() and cross-extractor read guard

* feat(verifier): named() and cross-extractor read guard in goja

* test(verifier): cross-extractor read guard and named()

* test(web-runtime): export runtime and extractors for tests

* test(web-runtime): named() and cross-extractor read guard

* refactor(folio): drop manual globalThis trailer (bundler injects it)

* refactor(folio): seed txn amounts via integers().between(1,500)

* refactor(folio-web): drop manual globalThis trailer (bundler injects it)

* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs

* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts

* refactor(folio-web): name extractors so violation witnesses are readable

* fix(web-runtime): propagate extractor getter throws and unpoison locked global

Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.

* test(spec): install fake runtime via defineProperty to survive locked global

* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors

* feat(runner): add MaxSteps bound to Options

* test(runner): MaxSteps stops after exactly N steps

* test(driverpb): drop proto getter round-trip tautology

* test(sidecar): drop stub-mode placeholder tautology tests

* test(mock): drop default-field-value assertion test

* test(ltl): drop Verdict.String tautology tests

* refactor(runner): extract RenderSummary for snapshot testing

* test(runner): golden snapshots for trace stream and violation summary

* feat(web-runtime): capture uncaught errors into state.exceptions

* test(integration): add throwing and counter web fixtures

* test(integration): add specs for the web fixtures

* test(integration): drive web fixtures through the real pipeline in headless Chrome

* chore(make): add test-browser target for the Chrome-driven suite

* ci: run the Chrome-driven browser suite in a separate job

* refactor(test): relocate browser suite to test/browser

* refactor(permissions): delete dead internal/permissions package

* refactor(test): rename package to browser_test

* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets

* chore(make): point test-browser at test/browser

* docs(decisions): record internal/permissions deletion

* refactor(doctor): use sidecarassets package

* refactor(testrun): use sidecarassets package

* fix(test): resolve testdata relative to browser_test.go

* refactor(verifier): remove dead __sanderlingIndex compat alias

* refactor(bundler): use encoding/json for JS string literals

* docs(action-space): use vendor-neutral native driver wording

* refactor(hierarchy): scrub backend tool name from comments

* refactor(driver): scrub backend tool name from comments

* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver

* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap

* refactor(chrome): implement DoubleTap as two taps with the gap

* refactor(mock): record DoubleTap and DoubleTapSelector actions

* refactor(runner): delegate double-tap to driver, drop gesture timing

* test(runner): assert double-tap delegates to driver DoubleTap

* docs(cmd): add package docs to CLI and developer tools

* docs(driver): add package docs to driver interface and chrome backend

* docs(driver): add package docs to mock and sidecar backends

* docs(platform): add package docs to android and ios device prep

* docs: add package docs to bundler and inspect

* docs(ltl): add package doc to temporal logic evaluator

* docs: add package docs to runner and testrun pipeline

* docs: add package docs to trace and verifier

* docs(sidecarassets): add package doc for embedded JAR loader

* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI

* test(chrome): gate real-Chrome driver tests behind the browser tag

* chore(make): run chrome driver tests in the browser job

* fix(web-runtime): guard global error listeners for non-browser hosts

The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.

* ci(browser): re-enable unprivileged user namespaces for headless Chrome

ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.

* ci(browser): pin stable Chrome for the driver tests

setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.

* feat(defaults): add scroll and rebalance action weights

Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.

* feat(defaults): trim scroll action weight wiring

* fix(build): point sidecar jar ignore and embed paths at sidecarassets

* test(defaults): drop stale longPresses re-export assertion

longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.

* fix(chrome): raise DevTools websocket read timeout to 60s

Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.
2026-06-02 09:52:53 +05:30
pj 88db9653e5 refactoring default action layer (#51)
* feat(hierarchy): add editable signal with native derivation

* feat(chrome): emit editable flag in hierarchy dump

* feat(verifier): expose editable on ax element objects

* feat(spec): add editable to selector and element types

* feat(verifier): register typing builtin generator

* feat(verifier): typing builtin types edge-case corpus into editable fields

* feat(spec): export typing builtin generator

* feat(spec): add defaultActions bundle

* feat(spec): export @sanderling/spec/defaults subpath

* feat(folio): layer defaultActions breadth over targeted flows

* test(verifier): typing builtin targets editable fields, declines otherwise

* test(hierarchy): editable derivation and selector matching

* test(spec): defaultActions, typing, and defaults barrel resolve

* fix(testrun): alias @sanderling/spec/defaults for the bundler

* test(chrome): editable flag for inputs, textarea, contenteditable

* feat(spec): typing builtin for the web (V8) action path

* chore(folio): auto-boot a bootable AVD in just test/install when none connected

* feat(driver): add ForegroundChecker optional capability

* feat(android): detect foreground package via adb dumpsys

* feat(sidecar): implement ForegroundApp via adb for android

* feat(runner): relaunch app when foreground escapes during exploration

* fix(spec): drop hardware back from defaultActions to stay in-app

* feat(spec): add DoubleTap action type and constructor

* feat(spec): wire DoubleTap through web-runtime serializer

* feat(verifier): bind doubleTap and decode DoubleTap actions

* feat(runner): dispatch DoubleTap as two taps inside one step

* test(doubleTap): cover constructor, verifier round-trip, and runner dispatch

* feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action

* fix(folio): track ledger row count across non-ledger steps; pin reproducer seed

* feat(spec): add doubleTaps random-target builtin to defaultActions

* feat(verifier): add doubleTaps random-target generator

* refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions

* fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives

* feat(verifier): track newly-violated property set per step

Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.

* refactor(runner): emit onset-only violations to trace and summary

Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.

* style(verifier): use maps.Copy for verdict snapshot

* fix(folio): make login spec content-driven (idempotent across re-entries)

* fix(verifier): canonicalize selector strings

Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.

* refactor(folio): replace txn invariants with balanceMatchesAddedTxn

Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.

* refactor(trace): drop WriteScreenshotAfter

Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.

* refactor(runner): one concurrent screenshot per step

Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.

* refactor(inspect-ui): use next step's screenshot for state after

Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.

* feat(sidecar): structural-hash settle poll

Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.

* test(sidecar): cover pollUntilStable and structuralHash

Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.

* feat(spec): accept optional name on extract()

Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.

* test(spec): cover extract name overload

Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.

* feat(verifier): name extractors for diff surfacing

bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.

* chore(folio): name every extract() call

Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.

* feat(verifier): track extractor value transitions

Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.

* test(verifier): cover ChangedExtractors diffs

Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.

* feat(trace): emit extractor_changes per step

Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.

* feat(inspect-ui): render extractor-change breadcrumbs at violations

Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.

* fix(sidecar): cap stability poll independently of settle budget

The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.

* feat(cli): default --clear-data on so runs start fresh

* feat(sidecar): streak-based settle with route-transition detection

Two changes layered into the stability poll:

1. stabilitySnapshot returns null while the tree carries more than one
   route-level Screen tag (resource-id / testTag / identifier ending
   in "Screen"), so the poll cannot declare a NavHost cross-fade
   stable. Apps following the Compose route convention get this
   detection for free; apps that don't fall through to the generic
   signal below.

2. pollUntilStable now requires an uninterrupted stable streak of at
   least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
   matches. A late transition that fires after a brief calm window
   breaks the streak instead of slipping past. Interval widened to
   250ms so UiAutomation isn't hammered under fuzz load.

* test(sidecar): cover streak reset and route-transition rejection

Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.

* feat(runner): re-fetch on transitional hierarchy capture

Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.

fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.

* feat(runner): gate first action on app reaching foreground

* test(runner): cover startup foreground gate and back-press

* feat(verifier): scope random-action targets to app package

Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.

* feat(testrun): pass app package into verifier scope filter

* test(verifier): cover package-scoped target selection

* feat(hierarchy): derive package from resource-id prefix

The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.

* test(hierarchy): cover package derivation from resource-id

* chore: stop tracking inspect-ui/dist build artifacts

* feat(android): detect focused-window package via dumpsys window

* feat(driver): add FocusedWindowChecker capability

* fix(runner): gate first observe on the app window being drawn, not just resumed

* test(mock): add FocusedWindowApp with foreground mirroring

* test(runner): cover startup gate waiting for app window to draw

* feat(proto): add Snapshot RPC for atomic hierarchy+screenshot

Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.

* feat(sidecar): add snapshot default on DriverBackend

Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.

* feat(sidecar): wire Snapshot handler with serialization lock

Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.

* test(sidecar): cover Snapshot wire path and serialization lock

SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.

* feat(driver): expose Snapshot on DeviceDriver and sidecar client

Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.

* feat(driver): add Snapshot to chrome and mock drivers

The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.

* refactor(runner): observe each step via the atomic Snapshot RPC

fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.

* test(runner): assert step uses Snapshot, not raw hierarchy/screenshot

TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.

* test(driver): cover Snapshot in proto descriptor and sidecar client

Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.

* feat(trace): add Transitional flag to Step

* fix(runner): skip verifier for transitional trees after retry budget

When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.

* test(runner): cover transitional step skips verifier and clean control

* refactor(trace): rename Step.Action to Step.NextAction

The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.

* refactor(runner): assign trace action to Step.NextAction field

Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.

* refactor(inspect): decode trace step's next_action JSON field

Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.

* test(inspect): update fixtures to use next_action trace field

Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.

* refactor(inspect-ui): rename Step.action to Step.next_action

Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.

* fix(folio): extract balanceMatchesAddedSum predicate as testable helper

Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.

* fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn

The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.

* test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases

Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.

* fix(build): rebuild sidecar JAR when Kotlin sources change

Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.

* fix(chrome): launch with no-sandbox so headless Chrome starts in CI

* fix(sidecar): type text at cursor instead of clearing the field

InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.

* test(sidecar): assert InputText types at cursor without clearing

Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.

* feat(proto): add LongPress RPC

* chore(proto): regenerate Go stubs for LongPress

* feat(driver): add LongPress to DeviceDriver interface

* feat(sidecar): add LongPress client method

* feat(mock): record LongPress action

* feat(chrome): implement LongPress as press-and-hold

* feat(sidecar): implement longPress across backends

* feat(sidecar): dispatch LongPress RPC to backend

* test(sidecar): cover LongPress dispatch

* test(sidecar): implement longPress in snapshot test backend

* feat(verifier): add LongPress and Scroll action kinds

* feat(folio-spec): predicate that gates balance check on TxnSubmit tap

Replaces the row-sum predicate (which always held by construction since
balance is derived from rows in Folio) with one that compares the typed
amount to the actual balance delta after a tap on TxnSubmit. Catches the
planted double-submit bug.

* feat(folio-spec): wire submitMovesBalanceByTypedAmount property

Adds lastAction and totalBalance extractors and uses them in the new
property. Drops ledgerRows/ledgerBalance extractors since nothing else
referenced them.

* feat(verifier): wire longPresses and scrolls generators

* test(verifier): cover longPresses and scrolls generators

* test(folio-spec): unit tests for submitChangesBalanceByTypedAmount

Covers single vs double submit, the DoubleTap variant, vacuous cases
(null action, wrong kind, wrong target, zero typed), and selector-as-
object coercion.

* feat(spec): add LongPress and Scroll authoring surface

* feat(spec): no-op LongPress and Scroll in web runtime

* feat(spec): re-export longPresses and scrolls as opt-in generators

* test(spec): cover LongPress and Scroll runtime members

* test(proto): expect LongPress in service descriptor

* feat(runner): dispatch LongPress and Scroll actions

* test(runner): cover LongPress and Scroll dispatch

* docs(action-space): move LongPress, Scroll, DoubleTap to current actions

* fix(runner): mark nil/empty hierarchy as transitional

A failed or empty sidecar hierarchy fetch was pushed straight to the
verifier, letting spec extractors crash with "Cannot read property 'map'
of undefined" when findAll returned null. Treat that case like a
transitional capture: skip the verifier push, still record the step, and
keep the loop progressing.

* fix(verifier): populate Action.On when tap chooser picks an element

Coordinate-targeted Taps/DoubleTaps left On empty, so action-gated
properties reading lastAction.on couldn't tell which target was hit and
were vacuously skipped. Resolve the picked element to a stable
key:value selector (resource-id, testTag, text, desc) and validate it
resolves back to the same element so we don't accidentally redirect the
tap to a sibling that shares the identifier.

* fix(folio): add parseTypedAmount helper matching app's parseCents

Raw user input like "50" must become 5000 cents, not 50. The existing
parseDollarCents helper strips non-digits and so reads "50" as 50 cents,
which is correct for formatted balance text but off by 100x for raw
input from the amount field.

* fix(folio): parse raw amount input as cents in submit predicate

txnAmountField holds raw user keystrokes, not formatted balance text.
Route it through parseTypedAmount so "50" reads as $50, matching how
the app commits the transaction.

* fix(folio): carry forward total balance across off-screen transitions

AddTransactionScreen shows neither AccountCard nor LedgerBalance, so the
extractor used to report 0 at the step before submit. That made every
non-zero current balance look like the full delta and tripped the typed
amount property on every honest submit. Remember the last-seen sum and
return it whenever the current snapshot has no balance signal.

* test(folio): cover submit predicate with raw typed-amount inputs

Pipes realistic raw keystrokes through parseTypedAmount + the predicate
so single submits clear and double submits fire as expected.

* feat(folio): add computeHomeTotalBalance helper

Pure helper that tracks Home multi-account total only and carries the last
Home sum across off-Home steps. Ledger's single-account balance is excluded
because mixing it would corrupt cross-screen scale comparisons.

* fix(folio): totalBalance carrier tracks only Home, not Ledger

Home cardSum is a multi-account total; Ledger's LedgerBalance is a single
account on a different scale. Blending them in the carrier produced bogus
cross-screen deltas (prev from Ledger, curr from Home), triggering false
positives in submitMovesBalanceByTypedAmount. Restrict the carrier to
Home AccountCard totals via the computeHomeTotalBalance helper.

* test(spec): cover computeHomeTotalBalance carrier behaviour

Tests Home sums, carrier passthrough on off-Home steps, the Ledger
scale-mismatch case, and a Home > off-Home > Home sequence.

* feat(runner): treat transient apply errors as transitional steps

Sidecar input RPCs occasionally hang with DEADLINE_EXCEEDED or
UNAVAILABLE on long fuzzing runs. The per-step loop previously
propagated any applyAction error and killed the run after a single
flake. Detect transient gRPC failures via status.FromError, mark the
step transitional, skip the post-action idle poll, and continue to the
next step. Fatal errors (outer ctx cancellation, non-transient codes,
verifier crashes) still propagate.

* test(runner): cover transient apply error resilience

TestRunner_TransientApplyErrorMarksTransitional drives the runner
through a wrapper that fails the first TapSelector with a gRPC
DeadlineExceeded then succeeds. Asserts the run does not exit, the
failed step is marked transitional with no violations, and the next
step runs cleanly. TestIsTransientApplyError_Classification covers the
helper's matching rules directly so future code changes don't quietly
drop a transient case.

* fix(folio): gate submit-balance property on Home route landing

totalBalance is only freshly computed when AccountCards are visible on
Home; off-Home landings return the carrier and would false-fire the
property, latching always(next(F)) to false and masking the real
double-submit bug. Skip vacuously when route is not "home".

* test(spec): cover route gate in submit-balance predicate

Adds route arg to existing cases (all use "home") and adds five new
cases: ledger landing with stale carrier, add-transaction with
double-insert delta, null route, plus home-landing positive and
double-insert negative cases anchoring the gate's allow path.
2026-06-01 12:48:51 +05:30
pj f572c8ba66 WIP: docs: refresh after iOS + web support (#50)
* docs: README covers iOS + web, surface both example apps

* docs(cli): document --ios-device and per-platform doctor

* docs: tighten README, fold examples into Docs list

* docs(runs): correct --clear-data lifecycle wording

Default behavior no longer wipes app data between runs; --clear-data is now opt-in.

* docs(getting-started): add iOS path, separate folio and folio-web

Document just test-ios under examples/folio, and distinguish the KMP
sample from the React + Vite folio-web sample.

* docs(inspect): document the eight panels

Lists Screenshot, ActionList, Timeline, ViolationsPanel, HierarchyPanel,
SnapshotTable, MetricsChart, ExceptionsPanel. Cross-links HierarchyPanel
to the spec language reference.

* docs(writing-specs): document setup export, flag noLogcatErrors as android-only

Mirrors pkg/spec/README.md so the manual covers the runner's setup-first
fall-through. Marks noLogcatErrors as Android-only so iOS/web spec
authors know it silently no-ops.

* docs(folio): document web target and iOS sanderling test recipe

After the KMP refactor folio also runs on wasmJs and the justfile exposes
just web, just web-build, and just test-ios. Surface all three.

* docs(folio-web): add README

Covers prerequisites, demo credentials, just test recipe, and how the
React + Vite host exposes state to the sanderling spec via stable ids
and data-* attributes.

* docs: scrub driver-implementation name from user docs

Drop the implementation tool name from README, cli.md doctor table, and
spec-language.md. These docs should describe behaviour, not the specific
underlying tool the native sidecar wraps.

* docs(development): scrub driver-implementation name from dev docs

architecture, design-principles, decisions now describe the native
sidecar by role (gRPC surface over OS UI-test pipeline) rather than by
the specific tool it wraps.
2026-05-25 16:17:21 +05:30
pj b23fb0c723 feat: web-native specs + per-platform doctor (#49)
* feat(doctor): per-platform check sets + --platform flag

Replaces unconditional defaultDoctorChecks with doctorChecksFor(platform);
web-only users no longer see scary FAILs for adb/emulator/java/sidecar.

* feat(testrun): add Preflight() before sidecar/driver setup

Mobile platforms get a friendly install hint pointing at
`sanderling doctor --platform=<p>` instead of `fork/exec java: not found`.
Web is a no-op.

* refactor(chrome): split tag (HTML name) from class (CSS classList)

Hierarchy attributes now expose HTML tag under 'tag' and CSS classes
under 'class', stopping the conflation of the two.

* feat(chrome): translate legacy string selectors to CSS/XPath

TapSelector now maps id:/desc:/descPrefix:/testTag:/etc. through
TranslateStringSelector. Unknown prefixes pass through to a CSS
attribute selector so a future Maestro key works without a release.

* feat(trace): add WriteHTML + Step.HTMLAvailable

Per-step HTML lives in <run>/html/step-NNNNN.html so trace.jsonl stays
line-greppable on apps with hundreds-of-KB DOMs.

* feat(driver): add WebDriver capability + chrome implementation

WebDriver exposes InstallBundle/EvaluateExtractors/NextActionFromV8/Document
for the V8-native web tick path. Mobile drivers stay binary-compatible.

* feat(verifier): OverrideExtractorValues for V8-driven extractors

Web tick path runs extractor bodies in V8 against the real DOM, then
overrides goja-side .current slots so LTL predicates evaluate against
those values. Mobile callers can pass nil for a no-op.

* feat(spec): add WebState + camelCase attribute aliases

WebState extends State with live `document`/`window` for V8-side web
extractors. KnownAttrSelectors gains camelCase aliases (contentDescription,
ariaLabel, testID, etc.) so cross-framework specs autocomplete.

* feat(runner): per-tick HTML capture for WebDriver-capable drivers

Type-asserts driver.WebDriver and writes <run>/html/step-NNNNN.html in
parallel with screenshot/hierarchy/metrics. Step.HTMLAvailable flips so
the inspect UI can hide the html tab on mobile runs.

* feat(inspect): serveHTML route under /api/runs/<id>/html/<name>

Mirrors serveScreenshot path validation; rejects traversal segments and
unknown extensions. text/html content-type so the iframe renders cleanly.

* feat(bundler): BundleWeb + V8-side runtime shim

web-runtime.ts installs globalThis.__sanderling__ with extractor / action
registries, plus __sanderlingExtractors__ + __sanderlingNextAction__
globals. BundleWeb composes user spec + runtime under esbuild's
PlatformBrowser into one IIFE.

* feat(runner): V8 extractor overrides + V8 action source for WebDriver

When the driver implements WebDriver, the runner sources extractor values
from V8 (real DOM) and the next action from the V8-side action generator.
LTL property predicates still run host-side in goja.

* feat(testrun): bundle + install web runtime when platform=web

BundleWeb composes the user spec with web-runtime.ts; the chrome driver
installs the resulting IIFE via Page.AddScriptToEvaluateOnNewDocument
post-Launch so the per-tick V8 extractor + action evaluation can begin
on step 1.

* feat(inspect-ui): hierarchy + html panels in run detail

HierarchyPanel renders the captured DOM/AX tree with a filter input.
HtmlPanel renders the per-step HTML in an iframe (sandboxed) with a
toggle to view source. HTML tab only shows when the step actually has
HTML captured.

* fix(folio-web): drop aria-label data-carrier abuse

Account cards now expose data-account-id + data-balance attrs and use a
human-readable aria-label. total-balance / ledger / ledger-balance carry
data-cents and data-txn-count instead of stuffing values into title.
Spec rewritten to read structured attrs via object-form selectors.

* chore: rebuild inspect-ui dist + folio-web .gitignore

Embeds the new HierarchyPanel + HtmlPanel into the inspect-ui dist that
ships with sanderling. Adds folio-web/.gitignore so generated runs/
don't leak into commits.

* revert(trace): drop WriteHTML + Step.HTMLAvailable

Screenshots already cover inspection; HTML capture bloats disk by
50-200MB per run with no payoff.

* revert(runner): drop per-tick HTML capture

Removes captureHTML helper and its three call sites; HTMLAvailable
flag no longer set on Step.

* revert(driver): drop WebDriver.Document

Document was only consumed by the runner's HTML capture which is gone.

* revert(inspect): drop /html route

Removes htmlPathPattern, serveHTML, and the dispatch block that called
it; HTML capture no longer exists on disk.

* revert(inspect-ui): drop htmlUrl + html_available type

API surface no longer needs the HTML route; Step.html_available has no
producer.

* revert(inspect-ui): drop HtmlPanel + html tab

Removes the iframe-based HTML viewer and its before/after tab wiring
from RunDetail.

* test(inspect-ui): drop htmlUrl test, add @types/bun

Pulls bun-types into tsconfig so api.test.ts (which uses bun:test)
typechecks; this was broken from the original feature commit.

* chore: rebuild inspect-ui dist without HtmlPanel

Embedded SPA bundle no longer ships the iframe HTML viewer.

* fix(web-runtime): retry action resolution + implement taps/swipes

V8-side runtime previously returned null when weighted picked a
generator that returned [] (page-gated), causing 80%+ of ticks on
narrow routes to emit no action and no post-screenshot. Now retries
up to 16x like goja, and the taps/swipes builtins query the live DOM
for clickable elements / dispatch random swipes instead of returning
null.

* fix(web-runtime): drop swipe, restrict pressKey to browser-meaningful keys

Web has no swipe gesture, so swipes dispatched pointer events into empty
divs. Make swipe() and the swipes builtin no-op. For PressKey, replace
the always-"back" choice with a random pick from {enter, tab, escape,
up, down, left, right} - keys that have real semantics in a browser.

* chore(folio-web): drop swipes from action root

Web runtime no-ops Swipe; remove the import and weighted entry so the
spec doesn't request actions that won't fire.

* fix(inspect-ui): correct HierarchyPanel CSS variable names

Tokens --surface-1/--surface-2/--text-secondary/--border-subtle don't
exist in tokens.css, so sticky thead had no background and tag/bounds
text fell back to inherited color. Map to the canonical --surface,
--surface-elevated, --text-muted, --border that other panels use.

* fix(chrome): correct PressKey mappings to chromedp/kb constants

Old keyMap had "home":"\x00" (NUL byte) and arrow keys mapped to
random punctuation runes (\x25-\x28 = % & ' () instead of arrow
keys. "escape" was missing entirely while the V8 runtime emits it.

Drop back/home (no browser navigation semantics) and route the
remaining keys through chromedp/kb constants so they actually
dispatch as the named keys.

* fix(cli): -h/--help exits 0 instead of error code

parseDoctorArgs hand-rolled its own flag loop and surfaced help text
as an error; parseTestArgs used flag.ContinueOnError but propagated
flag.ErrHelp to main() which printed "error: flag: help requested"
and exited 1.

Switch parseDoctorArgs to flag.NewFlagSet matching parseTestArgs, then
recognise flag.ErrHelp in main() so all subcommands exit 0 on -h.

* fix(chrome): harden cssEscape for control chars + use [class~=]

Previous cssEscape only handled " and \, leaving NUL/newlines/control
chars to break out of the CSS string literal. Port the CSSOM string
serialization rules: NUL becomes U+FFFD, control chars become \HEX,
quotes/backslashes get escaped.

Class selector switched from `.x` (which would need separate identifier
escaping) to `[class~="x"]`, which is also semantically correct for
multi-class elements.

* fix(web-runtime): use CSS.escape and validate tag-name selectors

The previous cssEscape only handled " and \, leaving newlines/control
chars to break out of attribute string literals. Delegate to the
platform CSS.escape per CSSOM spec.

The `tag` selector branch returned the bare value through cssEscape,
which doesn't prevent pseudo-classes (`*:hover`) from injecting into
the surrounding selector. Add a positive whitelist; values that don't
match a tag-name pattern collapse to a never-matching `:not(*)`.

Also switch class selectors to `[class~="..."]` to remove the only
identifier-context use of cssEscape.

* fix(chrome): validate attribute name in unknown-prefix branch

A selector like `foo]:has(*),body[x:value` previously produced
[foo]:has(*),body[x="..."], a syntactically valid CSS selector that
escaped the attribute match and selected `body`. Reject anything that
isn't a plain HTML attribute name.

* fix(selectors): emit valid XPath 1.0 string literals via concat()

Both the Go translator and the V8 runtime escaped " by prepending \,
which XPath 1.0 doesn't accept (its string literals have no escape
syntax). A `text:` value containing a quote produced malformed XPath
that chromedp/document.evaluate rejected.

Use the standard concat() composition: when the value contains both
' and ", split on " and join with `, '"', ` so each fragment is
wrapped in single or double quotes individually.

* fix(runtime): surface unresolved action targets instead of dropping silently

serializeAction emitted {x:0,y:0} via `?? 0` whenever a Tap/InputText/Swipe
target failed to resolve to coordinates. The runner then collapsed those
to ErrNoAction, so every selector typo became a silent no-op tick.

Have the runtime return null on unresolved targets and log a console
warning (visible via chromedp's runtime listener). Drop the now-redundant
{0,0} -> ErrNoAction guard so a deliberate Tap at the origin actually
fires.

* fix(runner): use errgroup-bound ctx so siblings cancel on failure

The errgroup's bound ctx was discarded; goroutines closed over the
outer ctx, so neither a sibling failure nor the future ability to
propagate per-step cancellation reached the V8 extractor's CDP
round-trip. Switch closures to gctx and document why Wait()'s error
is intentionally discarded.

* fix(chrome): propagate caller ctx cancellation to CDP calls

InstallBundle, EvaluateExtractors, NextActionFromV8 ignored the caller
ctx and ran chromedp.Run on d.tabCtx alone, so step deadlines and
Ctrl-C couldn't interrupt an in-flight CDP round-trip on a hung tab.

Add a runCtx helper that derives a chromedp-bound context which also
cancels when the caller's ctx cancels, and route the three V8 entry
points through it.

* fix(verifier): tolerate out-of-range override indices

A single stale index from V8 aborted the entire override map, so any
valid entries alongside it were dropped and verification ran on stale
extractor values. V8 and goja register from the same bundle so a
mismatch is unusual but recoverable.

Skip out-of-range entries instead of erroring, and return the skipped
count so the runner logs the mismatch without losing valid overrides.

* test(verifier): cover object-shaped extractor overrides

Existing tests only override scalars (777, 200), so a future jsonToJSValue
regression around nested object propagation would slip through. Lock down
the contract: a JSON object override should make {attrs.testTag, balance}
readable from goja predicates.

* fix(web-runtime): lock global runtime hooks against page shadowing

AddScriptToEvaluateOnNewDocument runs first, but a page script can still
delete or replace window.__sanderling{,Extractors__,NextAction__} between
install and host invocation. Define them as non-writable, non-configurable
properties so any attempt to shadow them throws in strict mode rather than
silently breaking the run.

* perf(web-runtime): cache randomTap candidate DOM scan per tick

The 16-attempt retry loop in __sanderlingNextAction__ called
randomTap repeatedly; each call ran querySelectorAll over a-button-
input-... and re-flushed layout per match via getBoundingClientRect.
On heavy SPA routes that's the per-tick budget gone.

Cache the scan in a module-level slot, reset at the top of each
__sanderlingNextAction__ invocation so the cache doesn't outlive a tick.

* fix(web-runtime): cap sanitize recursion to prevent stack overflow

State exposes document and window (per WebState in types.ts). A user
extractor returning either crashes the runtime via stack overflow on
the circular DOM/Window references. Track seen objects in a WeakSet
and bail at depth 32 so the worst case becomes a truncated value, not
a process kill.

* fix(web-runtime): enforce pressKey allowlist in factory

The factory accepted any string while randomPressKey only emitted
enter/tab/escape/arrows. A spec emitting pressKey({key:"home"}) would
flow through to the chrome driver, which rejects unsupported keys with
a runtime error mid-step. Reject at the factory so the spec author
sees the failure where it originates.

* chore(chrome): drop dead bundleSource/bundleMu

bundleSource was written under bundleMu but never read. Either remove it
or wire a re-install path; remove until the second is actually needed.

* fix(chrome): use strconv.Atoi for extractor key parsing

fmt.Sscanf("%d", ...) silently accepts trailing garbage like "3abc"
as 3. strconv.Atoi rejects the same input outright, so a malformed
key surfaces as an error instead of a wrong-bucket override.

* fix(doctor): raise per-check timeout to 15s for chromium launch

5s could time out the headless chromium check on cold CI. Most checks
finish in milliseconds, so a longer ceiling doesn't slow real
failures.

* fix(runner): trust V8 coordinates for InputText, even at origin

resolveCoordinates required strict positive X/Y, so a V8-emitted
InputText for an element at viewport (0, *) or (*, 0) skipped the
focus tap and typed into whatever was focused. Distinguish the
selector-driven path (mobile) from the coords-only path (web V8) so
edge coordinates are honored without breaking the existing tree-lookup
fallback.

Add applyAction tests covering both the typical web case and the (0,0)
edge case.

* test(bundler): lock down deterministic output across builds

The review flagged map-iteration nondeterminism as a possible cause of
unstable bundle SHAs. Empirically esbuild's Define handling is order-
independent (parallel substitution rules), so output is already stable.
Add a regression test that builds 10x with multiple Defines and asserts
SHA equality so any future change that introduces ordering surfaces.
2026-05-03 11:21:21 +07:00
pj dd54c24c4e feat: --clear-data flag + typed attribute selectors (#48)
* feat(test): add --clear-data flag to clear app data on launch

* test+docs: cover --clear-data flag in CLI parser test and reference

* feat(spec): type AttrSelector with known attribute names

Replace AttrSelector = Record<string, string> with KnownAttrSelectors
plus a string|boolean index signature, so authors get autocomplete and
type-checking on testTag / focused / clickable / etc. while raw driver
attributes still type-check via the fallback. Boolean state attributes
accept native booleans; goja stringifies them at the marshal boundary.

AccessibilityElement.attrs becomes RawAttrs (typed string-valued shape
of the same canonical names) so element.attrs.testTag autocompletes.

* test(verifier): native boolean selector value matches focused=true

* docs+folio: use native boolean for focused selector and document typed attrs
2026-04-27 01:02:21 +07:00
pj c76745e5f1 WIP: folio refactor - KotlinConf-style production-app shape (#47)
* feat(hierarchy): testTag alias resolves to resource-id and accessibilityIdentifier

Compose's testTag surfaces as resource-id on Android and as
accessibilityIdentifier on iOS. Selectors written as
{ testTag: "Foo" } now match either, so Sanderling specs can use the
same tag on both platforms.

Also rounds out the iOS identifier aliases so resource-id /
identifier / accessibilityIdentifier all resolve to one another.

* chore(folio): add gradle/libs.versions.toml

Centralises versions for all folio modules ahead of the module split.
Adds new entries for kotlinx-serialization, navigation3, Metro, KSP,
and the JetBrains lifecycle-viewmodel-compose multiplatform artifact.

* refactor(folio): introduce nested KotlinConf-style modules

Split the monolithic :composeApp into :core, :app:shared,
:app:ui-components, and :app:androidApp. The old module is still
present and remains the source of truth until the next commits remove
it; both compile in parallel to keep iOS/Android builds green during
the cut-over.

Highlights:
- :core - SQLDelight schema + LedgerStore + Repository (now an
  injectable class, not a singleton object). Methods are suspend to
  match generateAsync = true.
- :app:ui-components - design system primitives. IconButton/AppButton
  APIs revised: label = real contentDescription, testTag = stable
  selector. Drops the data-carrier description argument.
- :app:shared - per-screen ViewModels colocated with screens; pure
  composables on (state, onEvent); LocalAppComponent CompositionLocal
  for hand-rolled DI; @Serializable Route. Hosts the iOS framework
  (baseName Shared).
- :app:androidApp - thin Android entry that constructs the
  DriverFactory and hands it to App().
- gradle/libs.versions.toml centralises versions; settings.gradle.kts
  enables type-safe project accessors.

Deferred to follow-up PRs (per the design discussion):
- Metro DI: hand-rolled AppComponent for now; Metro graphs are mostly
  ceremony for an app this size and add KSP/version risk.
- Navigation3: kept the existing Navigator-as-backstack class,
  injected rather than singleton; nav3 isn't shipping a stable
  multiplatform artifact for commonMain consumption yet.
- :app:webApp + OPFS sqlite worker: web persistence is real new
  wiring (custom worker on @sqlite.org/sqlite-wasm). Master's
  WebLedgerStore + Snapshot is being removed by this PR; web stays
  buildable as a klib but no app-level wasm binary lands here.

* refactor(folio): delete :composeApp and retarget tooling

Removes the old monolithic module now that :core / :app:shared /
:app:ui-components / :app:androidApp own the source. Updates:

- justfile install/uninstall recipes -> :app:androidApp
- iosApp/project.yml framework path -> ../app/shared/...,
  baseName Shared (was ComposeApp); pre-build script invokes
  :app:shared:linkDebugFrameworkIosSimulatorArm64
- iosApp/iosApp/iOSApp.swift -> import Shared
- README -> mentions SQLDelight unification, drops the
  data-carrier contentDescription notes (now stale), no Layout
  section per repo convention

* refactor(folio-spec): query testTag and identify items by visible text

Replaces every accessibilityText / descPrefix data-carrier read with
testTag selectors that resolve to resource-id (Android) or
accessibilityIdentifier (iOS) via the SDK's alias table.

- Routes detected via testTag (LoginScreen, HomeScreen, etc.)
- Account identity = visible account name (no synthetic id encoded
  in semantics).
- Ledger row identity = joined text content of the row.
- Active account derived from route alone (not parsed from
  contentDescription).
- Focused input read from native focused="true" attribute, not from
  a custom focused_input data carrier.

* fix(folio): build green on Android assemble + iOS framework link

- Drop ksp/metro/navigation3 plugin aliases - not actually applied
  by any module in this PR (deferred follow-up).
- import awaitAsOne from app.cash.sqldelight.async.coroutines for
  the suspend single-row reads enabled by generateAsync = true.
- Drop kotlin.js.ExperimentalWasmJsInterop opt-in from common
  compilerOptions (it isn't valid for android/jvm targets).
- :app:shared androidMain pulls in androidx.activity:activity-compose
  for the BackHandler actual.

* fix(folio): testTagsAsResourceId at App root + JS-bridge regression test

App.kt sets testTagsAsResourceId=true on the root Box semantics so
Compose's testTag surfaces as Android resource-id (and equivalent on
iOS via accessibilityIdentifier). Without this, testTag stays in the
Compose semantics tree but never reaches the runtime hierarchy that
UIAutomator and Sanderling read.

Also adds TestStateAxObjectSelectorTestTagAlias as a regression
test for the {testTag: ...} object selector resolving through the
SDK alias to resource-id at the JS bridge layer.

* test(verifier): expose PredicateError latching across steps

The runner logs PredicateError once per step. The current implementation
latches the first error per thunk, so the log freezes on step 1 forever
even when later steps would observe different errors. This test fails
today and locks in the contract: PredicateError must reflect the most
recent step.

* fix(verifier): refresh predicate errors per step

EvaluateProperties short-circuits once an Always-property latches to
violated, so the underlying goja predicate stops being called and
formula.err keeps whatever it threw at step 1. The runner logs
PredicateError every step a property is violated, which made every
subsequent log line repeat the step-1 throw. That looks like the spec
runtime is seeing stale state, but it is just stale error reporting.

EvaluateProperties now invokes every registered predicate once per step
purely to refresh formula.err. Verdicts are unaffected. The thunk
itself stops latching so the new value wins on whichever path runs first.

* chore(folio): add Metro DI plugin (1.0.0-RC4) to versions catalog

Adds dev.zacsweers.metro plugin alias and applies it to :core
as a smoke test. Compiler-plugin only, no KSP required.

* chore(folio): apply Metro plugin to :app:shared and :app:androidApp

* feat(folio-core): annotate Repository and SqlLedgerStore with @Inject

* feat(folio-core): scope Repository and SqlLedgerStore as @SingleIn(AppScope)

Both are app-wide singletons so the SqlDelight-backed flows remain
shared across the graph.

* feat(folio): annotate ViewModels with Metro @Inject / @AssistedInject

LedgerViewModel and AddTransactionViewModel use @AssistedInject for
their accountId param plus a nested @AssistedFactory; the rest are
plain @Inject constructor classes.

* feat(folio): introduce Metro AppGraph in commonMain

Single shared @DependencyGraph(AppScope::class) that exposes
Repository, Navigator, and ViewModels. LedgerDatabase enters the
graph via @DependencyGraph.Factory.create(database) so the suspend
DriverFactory.create() can stay outside the DI surface.

@Binds wires SqlLedgerStore to LedgerStore; Navigator is provided
explicitly so its Route.Home start state stays in DI rather than
relying on a default-parameter being honored by the graph.

* fix(folio): expect/actual testTagsAsResourceId so iOS link succeeds

Compose's androidx.compose.ui.semantics.testTagsAsResourceId is
Android-only. Calling it directly from commonMain broke
linkDebugFrameworkIosSimulatorArm64. Replace with an expect Modifier
extension that wires the semantics on Android and is a no-op on
iOS / wasmJs.

* refactor(folio): replace AppComponent with Metro AppGraph in App.kt

App now takes a suspend graph builder; the platform constructs
LedgerDatabase off the suspend DriverFactory.create() before invoking
the Metro graph factory. Routes resolve VMs through LocalAppGraph
instead of the hand-rolled LocalAppComponent.

Drops the loading-state placeholder comment (the empty Box is enough).

* refactor(folio): resolve ViewModels through LocalAppGraph in routes

Each *Route composable now reads the AppGraph from CompositionLocal
and pulls its VM via the appropriate accessor or AssistedFactory.

* refactor(folio): build AppGraph from platform entry points

MainActivity (Android) and MainViewController (iOS) now own the
suspend DriverFactory.create() and feed the resulting LedgerDatabase
into Metro's createGraphFactory<AppGraph.Factory>().

* chore(folio): add navigation-compose 2.9.2 dependency

Adds the JetBrains KMP navigation-compose library to the shared
module. Used in subsequent commits to replace the hand-rolled
Navigator with a typed-route NavHost.

* refactor(folio): replace custom Navigator with NavHost backstack

Wraps androidx.navigation.NavHostController behind the existing
push/replace/back surface so call sites in ViewModels stay unchanged.
App.kt now wires a typed NavHost with @Serializable Route entries
and observes the controller's currentBackStackEntry to drive the
session-based Login/Home redirect.

* fix(folio-core): wire kotlinx-browser so wasmJs DriverFactory compiles

org.w3c.dom.Worker on wasmJs lives in kotlinx-browser, not the stdlib.
Pin 0.5.0 alongside the @sqlite.org/sqlite-wasm 3.53.0-build1 version
that the upcoming web app will depend on, and switch the worker
constructor to the module-worker form that webpack expects.

* feat(folio): scaffold :app:webApp wasmJs module

Compose Multiplatform target that depends on :app:shared and pulls
@sqlite.org/sqlite-wasm 3.53.0-build1 as the npm runtime for the
SQLDelight web worker.

* feat(folio-webApp): add main entrypoint and index.html

main.kt mirrors the iOS entry point: builds DriverFactory + AppGraph
factory, hooks browser back-gesture into WebBackGesture, then mounts
the shared App composable into ComposeViewport.

* feat(folio-webApp): OPFS-backed sqlite worker + webpack config

sqlite.worker.js implements the SQLDelight web-worker protocol
(exec/begin_transaction/end_transaction/rollback_transaction) on top
of @sqlite.org/sqlite-wasm. Prefers the OPFS SAH pool VFS for
persistent storage and falls back to in-memory when OPFS is
unavailable.

webpack.config.d/coopcoep.js sends COOP/COEP headers on the dev
server so cross-origin isolation is available, even though the SAH
pool itself does not require it. webpack.config.d/sqlite-wasm.js
enables asyncWebAssembly so webpack can bundle sqlite3.wasm via the
'new URL("sqlite3.wasm", import.meta.url)' reference inside the
sqlite-wasm package.

* chore(folio): add web/web-build just recipes and refresh yarn lock

Yarn lock picks up @sqlite.org/sqlite-wasm 3.53.0-build1.

* fix(folio-core): probe schema before create on wasmJs

Wasm SqlDriver doesn't auto-track user_version like the Android
driver, so awaitCreate() ran on every page load and tripped over
already-created tables. Read PRAGMA user_version, run
awaitCreate/awaitMigrate based on it, and self-heal pre-existing
tables with version 0 by stamping the current schema version.

* chore(folio-webApp): pin dev-server port and trim worker logging

webpack-dev-server now binds 8088 (or WEBAPP_PORT) so it doesn't
collide with the docs server on 8080. Drop the per-message reply
log; keep only the OPFS init line and error logging.

* chore(folio): nest iosApp under app/ for KotlinConf parity

Match KotlinConf-app's filesystem layout where every entry point (android,
ios, web, shared, ui-components) lives under app/. iosApp is still an Xcode
project, not a Gradle module, so settings.gradle.kts is unchanged.

* feat(folio): testTag identity for AccountName and ledger row cells

Replaces string-heuristic identity in the spec extractors with stable
testTags. AccountCard exposes AccountName; LedgerRow exposes TxnNote
and TxnDate. Spec extractors read those directly instead of filtering
visible text by "starts with $" / "matches digit".

* fix(folio-app): branch start destination on initial session

Read repository.session.value at first composition and pick
Route.Home or Route.Login as the NavHost startDestination. Avoids
the one-frame Home flash on cold start with no persisted session.

* refactor(folio-webApp): hard-fail when OPFS unavailable

Drops the silent in-memory fallback. The README claims OPFS
persistence; falling back without surfacing the degrade made data
loss invisible across reloads. Now the worker errors out and the
Kotlin DriverFactory rejects the create() call instead.

* docs(verifier): document extractor advancement and refresh invariants

Extractor previous/current advance only on PushSnapshot, never per
thunk-call. refreshPredicateErrors depends on this for safe re-entry.
Also flags that re-invoked predicates run outside their LTL gate, so
they must be side-effect-free reads.

* fix(hierarchy): populate ResourceID from accessibilityIdentifier

iOS Compose surfaces testTag as accessibilityIdentifier. Previously
only resource-id and identifier seeded element.ResourceID, leaving
element.id empty for iOS Compose nodes and forcing specs to walk
attrs to recover stable identifiers.

* refactor(folio-spec): use element.id for focused field tag

Now that ResourceID populates uniformly across Android/iOS Compose,
the spec can read element.id directly instead of probing attrs for
each platform's underlying field name.

* fix(folio-spec): pick account card via seeded from(), not Math.random

Math.random() breaks --seed reproducibility. The verifier's seeded
RNG flows through from(), so re-running a seed now produces the
same card pick sequence.

* docs(spec): fix README example to use scoped extractors

The previous snippet referenced `state.ax.find` inside an actions()
body where state is not in scope, and shadowed the imported actions
helper with an export of the same name.

* feat(hierarchy): add FindBySelectorPath for chained object selectors

Each selector in the chain is matched within the descendants of the
previous match. Returns the deepest match (or nil) for FindBySelectorPath
and every deepest match for FindAllBySelectorPath.

* feat(verifier): dispatch JS array selectors to FindBySelectorPath

`state.ax.find([{...}, {...}])` now walks each segment scoped under
the previous match. Strings and single objects keep their existing
single-shot lookup.

* feat(spec): expose SelectorPath in find/findAll signatures

* fix(spec): satisfy AccessibilityElement interface in Tap test fixture

* refactor(folio-spec): collapse chained finds into selector paths

* feat(spec): add keyedBy(element, tags) identity helper

Joins element.find({testTag: tag})?.text per tag with U+001F as the
delimiter so user-visible text can never collide with the separator.
Returns empty string for an undefined element.

* refactor(folio-spec): use keyedBy for ledger row identity

* feat(spec): add whenRoute action gating helper

whenRoute(route, allowedRoutes, body) wraps an actions() generator
that returns [] unless route.current matches one of the allowed
values. Accepts a single route or an array.

* test(spec): cover whenRoute matching, gating, and array routes

* refactor(folio-spec): gate addAccount and addTxn with whenRoute

* refactor(hierarchy): use maps.Copy for attribute merge

Linter flagged the manual loop after recent edits surfaced the hint.

* feat(verifier): dispatch setup generator before actions root

Setup is consulted every step; when it returns ErrNoAction the call falls
through to the existing actionGenerator retry loop. This lets specs split
deterministic preconditions (login, onboarding) out of the weighted action
pool while auto-reengaging if state regresses (e.g. logout under fuzz).

* docs(spec): document setup precondition action generator

* refactor(folio-spec): export login as setup, remove from action pool

Login is deterministic and yields no actions once the app is past the
login screen; sitting at weight 50 in the action pool wasted half of step
picks on a no-op. Promote it to setup so the runner only consults it
while it has work to do, and rebalance remaining weights to round numbers
(addAccount 50, addTxn 40, back 10).
2026-04-27 00:33:57 +07:00
pj 28909d954b docs: spec language reference page + writing-specs rewrite (#46)
* docs(architecture): add device/emulator node and XCTest edge to diagram

* docs(manual): rewrite writing-specs with accurate API and updated examples

* docs(manual): add spec-language reference page

* docs(sidebar): add spec-language entry to sidebar nav
2026-04-26 13:07:31 +07:00
pj 34fb73d6cb fix(folio): stop abusing contentDescription as data carrier (#45)
* feat(hierarchy): full-attribute selector system

- Add Attributes map to Element (raw platform attrs + serialized booleans)
- Add Selector / AttrFilter types for multi-filter AND matching
- Add matchAttr with alias expansion and substring/boolean semantics
- Add matchSelector (AND of all filters)
- id: and desc: keep exact/suffix/prefix semantics for backward compat
- text: widens to substring via matchAttr
- default: case routes unknown kinds to matchAttr (NEW)
- Add Tree.FindNode / Tree.FindAllNodes returning *Node
- Add Node.Find / Node.FindAll for scoped subtree string search
- Add Node.FindBySelector / Node.FindAllBySelector for object AND search
- Add attributeAliases for cross-platform name expansion

* test(hierarchy): full-attribute selector coverage

- raw resource-id: substring match
- label:/content-desc: alias expansion to accessibilityText on iOS
- scrollable:true/false boolean exact match
- title: iOS-only attribute, graceful nil on Android
- text: substring widening
- Selector AND: both filters must match; single miss returns nil
- Node.Find scoped search: descendants only, not siblings

* feat(verifier): object-form selectors + attrs + element-level find

- ax.find/findAll accept string or {attr:value} JS objects
- Object form builds Selector with AND semantics
- Returned element objects expose attrs sub-object (raw platform attrs)
- Returned element objects expose .find() and .findAll() scoped to subtree
- Element-level .find/.findAll accept string or object selectors

* feat(spec): extend AccessibilityElement and AccessibilityTree types

- AccessibilityElement gains attrs, find(), findAll()
- find/findAll on both Tree and Element accept string | AttrSelector
- AttrSelector = Record<string, string> for object-form AND matching

* feat(folio): migrate to chained object-form selectors

- Replace path queries (desc:X > desc:Y) with chained API
- Screen root lookups use { accessibilityText: "ScreenName" }
- Element-scoped searches use find/findAll with string or object
- Keep string selectors for descPrefix: and desc:Back (shows both forms)

* fix(folio): use account_card:id desc, expose balance via text semantics

* fix(folio): embed accountId in LedgerScreen desc, expose values via text semantics

* fix(spec): replace desc-parsing with text-based parseDollarCents extraction
2026-04-25 23:43:19 +07:00
pj 776becdf4b Remove in-app SDK (#43)
* chore: delete internal/agent package

* chore(build): remove sdk-android from gradle settings

* chore(makefile): remove sdk-android targets

* chore(ci): remove release-android job from release workflow

* chore(folio): remove sdk-android dependency

* chore(folio): remove SDK initialization from FolioApplication

* chore(folio): delete snapshot extractor files

* feat(folio): add balance to account card content description

* feat(folio): add hierarchy content descriptions to LedgerScreen

* refactor(folio): rewrite spec.ts to use ax extractors

* docs: remove in-app SDK from README

* feat(folio): add focused_input indicator to App

* docs: remove in-app SDK from index

* refactor(runner): remove agent SDK connection and snapshot step

* test(runner): update tests for SDK removal

* docs: remove Android SDK section from getting-started

* refactor(testrun): remove agent SDK connection setup

* docs: remove snapshots from writing-specs

* docs: remove in-app SDK from architecture doc

* docs(folio): update README for SDK removal

* docs: update per-step cycle diagram in architecture doc

* fix(folio): detect screens from unique element presence, not id: selectors

testTag() in Compose is not exposed as resource-id without testTagsAsResourceId.
Use desc: selectors for elements unique to each screen instead of id: path queries.

* feat(folio): add screen root contentDescription for scoped ax selection

Each screen root gets semantics { contentDescription = "ScreenName" } so
sanderling specs can scope element lookups through the screen: desc:LoginScreen > desc:login_submit.

* fix(folio): scope all ax selectors through screen root nodes

Use desc:ScreenName > desc:element path queries so every selector is
rooted at the screen level. focusedInput stays unscoped since it lives
in the app root, outside any screen.

* fix(folio): guard newAccountBalanceIsZero against navigation false positives

Scoped selectors return [] when not on HomeScreen so accounts vanish and
reappear as apparently-new on each visit. Skip the check when prev was empty.

* chore(folio): link @sanderling/spec to local pkg/spec for IDE type checking

* feat(spec): add desc, class, clickable, enabled, checked, focused, selected to AccessibilityElement

Runtime fields set by the verifier were missing from the TypeScript type,
causing linting errors on el.desc and related accesses in specs.

* chore(folio): switch to bun, add tsconfig.json for IDE type checking

- Remove package-lock.json, add bun.lock
- Add tsconfig.json so VSCode resolves @sanderling/spec types
- Fix parseAccount/parseLedgerRow to accept string | undefined
2026-04-25 20:04:29 +07:00
pj 6c32fb0e1d feat(hierarchy): flat-to-tree + path query selector (#41)
* feat(hierarchy): add Node tree + path query support

Preserve parent-child relationships in a Node tree at parse time.
Extend Find/FindAll with " > " path operator for scoped queries,
e.g. id:LoginScreen > desc:EmailInput.

Flat Elements slice and all public signatures unchanged.

* test(hierarchy): add path query tests

Cover single-level, multi-level, mixed-type, not-found, wrong-subtree,
and FindAll across multiple roots.

* feat(folio): add modifier param to Screen composable

* feat(folio): testTag LoginScreen

* feat(folio): testTag HomeScreen

* feat(folio): testTag AddAccountScreen

* feat(folio): testTag LedgerScreen

* feat(folio): testTag AddTransactionScreen

* feat(folio): scope spec selectors to screen testTags via path queries

* fix(sidecar): align grpc-netty/grpc-okhttp with grpc-core version

Maestro 1.40.0 brings grpc-netty:1.50.2 which compiled against
AbstractManagedChannelImplBuilder, removed in grpc-core 1.64+.
Gradle was upgrading grpc-core to 1.68.0 while leaving grpc-netty at
1.50.2, causing NoClassDefFoundError at startup.

Force grpc-netty and grpc-okhttp to match grpcVersion so all gRPC
artifacts are binary-compatible.

* fix(sidecar): use ephemeral host port for AndroidDriver

Hard-coded port 7001 caused TcpForwarder.start() to fail with
TimeoutException when a previous sidecar process held the port open.
Pick a free port via ServerSocket(0) instead.

* chore(sidecar): bump maestro to 2.4.0, exclude graalvm from fat JAR

Maestro 2.4.0 still ships grpc-netty:1.50.2 so the grpc transport
resolution strategy is kept. GraalVM JS is excluded: unused by our
sidecar and causes shadow JAR expansion errors (pom treated as zip).

* fix(sidecar): update AndroidDriver calls for Maestro 2.4.0 API

launchApp no longer accepts a UUID argument. Third constructor param
changed from hostname to emulatorName so drop the "localhost" value.
2026-04-25 18:46:15 +07:00
pj 2a1b263b8c fix: WDA startup flakiness - warmup + connection drop message (#40)
* feat(ios): add simulator management package

* feat(testrun): add iOS platform path (simctl launch + direct TCP)

* feat(cli): add --ios-device flag and IosDevice option

* feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser)

* feat(folio-ios): wire SanderlingIos.start() in MainViewController

* feat(folio-ios): add test-ios justfile recipe

* fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings

* fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects

* feat(proto): add env map to LaunchRequest

* feat(driver): add env param to Launch interface + all implementations

* feat(testrun): launch iOS app via XCTest with env vars instead of simctl

* feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through

* feat(sidecar): wire env map in DriverService + IosDriverBackend in Main

* fix(sidecar): use LocalIOSDevice (WDA+simctl) + stop before relaunch

* fix(sidecar): include exception type in gRPC error description

* fix(sidecar): pick free WDA port instead of hardcoded 9100

Use SocketUtils.nextFreePort to pick a free port in the 22000-23000 range
rather than hardcoding 9100, which only worked if a previous WDA session
left a listener there.

* fix(sdk-ios): check semaphore wait result and throw on snapshot timeout

dispatch_semaphore_wait returns nonzero on timeout; ignoring the return
value caused pauseAndSnapshot to silently return an empty map, sending a
garbage empty STATE frame to the host. Now throws so the agent loop
reconnects instead.

* fix(folio-ios): register snapshot extractors before starting agent

SanderlingIos.start() was called before the snapshot objects were
initialized, so a PAUSE message arriving early produced an empty snapshot.
Move start() to after all extractors are registered.

* chore(ios): remove dead LaunchApp function

LaunchApp had no callers since bff3a49 switched iOS launch to go through
the sidecar driver. Remove the dead code and unused os import.

* test(ios): add unit tests for pickSimulator and iOS flag parsing

Tests for all pickSimulator branches (by name, by UDID, unknown, empty
list, iPhone preference, fallback to first). Also tests BootedUDID on a
canceled context and verifies --platform ios and --ios-device flags are
accepted by parseTestArgs.

* fix(ios): propagate error from BootedUDID instead of silently swallowing

* fix(sidecar): IosDriverBackend.healthy() returns true; WDA liveness checked in open()

* fix(sidecar): warm up WDA after health check to absorb startup race

* fix(runner): surface clear message on WDA connection drop

* docs(testrun): note WDA warmup location above WaitForHealth

* fix(sidecar): extract warmup + add one-shot WDA reconnect on IOException

* fix(runner): fatal on permanent WDA drop during hierarchy fetch

* fix(sidecar): walk cause chain in withReconnect to catch Maestro-wrapped IOException

* fix(sidecar): explicit Unit return in pressKey and waitForIdle withReconnect lambdas

* fix(sidecar): serialize WDA reconnect with ReentrantLock to prevent concurrent xcodebuild races

* fix web examples package config

* Revert "fix web examples package config"

This reverts commit 70c10ade27.

* chore: gitignore built sanderling binary

* fix(hierarchy): parse iOS [x1,y1][x2,y2] bounds + match iOS merged desc labels

* refactor(folio-spec): replace bloated spec with two focused properties

Login is opportunistic. Two concrete properties:
1. every new account starts with balance 0
2. every new txn changes ledger balance by exactly its signed amount

Actions: directed login -> addAccount -> addTxn -> back weighted flow.
2026-04-25 17:39:39 +07:00
pj 97154cf580 feat(ios): launch via XCTest with env vars for hierarchy/tap access (#39)
* feat(proto): add env map to LaunchRequest

* feat(driver): add env param to Launch interface and all implementations

* feat(sidecar): add IosDriverBackend using Maestro IOSDriver + env pass-through

* feat(testrun): launch iOS app via XCTest with env vars instead of simctl

* fix(ios): replace LaunchApp with BootedUDID; simctl launch moved to XCTest path

* test(cli): add tests for ios platform flag and ios-device flag parsing

* fix(sdk-ios): check semaphore wait result; resolve port from args and env; register extractors before start
2026-04-23 17:35:31 +07:00
pj 007dcddd69 feat(ios): iOS e2e support — Kotlin Native SDK + simulator driver + test-ios (#38)
* feat(ios): add simulator management package

* feat(testrun): add iOS platform path (simctl launch + direct TCP)

* feat(cli): add --ios-device flag and IosDevice option

* feat(sdk-ios): add Kotlin Native iOS SDK (TCP agent + POSIX socket + dispatch pauser)

* feat(folio-ios): wire SanderlingIos.start() in MainViewController

* feat(folio-ios): add test-ios justfile recipe

* fix(sdk-ios): remove unavailable C macros; manual byte swap + no-cast warnings

* fix(testrun): simctl-first launch order for iOS; Maestro init after SDK connects
2026-04-23 17:29:53 +07:00
pj 88db0cbea8 docs: web platform + clean URLs + dark/light mode (#37)
* chore(docs): replace d2 diagram pipeline with mermaid

Remove docs/_diagrams/ and d2 build step from Makefile. The HTML
template already initialises mermaid.js; diagrams are now inline
code fences rendered client-side.

* docs(architecture): add mermaid diagram + web/CDP platform docs

Replace SVG img tag with inline mermaid flowchart showing both native
(Maestro sidecar + in-app SDK) and web (Chrome CDP) paths. Update
Processes, Transports table, and per-step cycle sections.

* docs(design-principles): update principles 1-4 for web platform

Principles 1, 2, 3, and 4 referenced Maestro and native-only concepts.
Add web/CDP context and update driver-is-an-interface to name both
sidecar and chrome implementations.

* docs(manual): add web prerequisites and folio-web example

Update --platform flag to list android, ios, web. Add web prerequisites
section (Chrome, no SDK needed) and folio-web quick-start to
getting-started.

* chore(gitignore): untrack inspect dist/index.html build artifact

index.html is regenerated by vite on every build with a new content hash,
making it permanently dirty. Only .gitkeep is needed for //go:embed to
compile on a fresh checkout. Also remove duplicate dist/* line and stale
d2 diagram ignore entries.

* feat(docs): click-to-zoom for mermaid diagrams

* docs(architecture): change diagram layout from LR to TB

* docs(getting-started): link npm and Maven Central package headers

* update docs root

* docs(spec): rewrite npm package README

Update usage example to current API, drop stale version-compatibility
and license sections.

* build(docs): output pages as pagename/index.html for clean URLs

Split DOCS_OUT into INDEX_OUT (index.md files stay as index.html) and
PAGE_OUT (all other pages become pagename/index.html). The __ROOT__
depth computation already handles the extra directory level correctly.

* chore(docs): update sidebar links to directory-style URLs

* docs: update cross-links from .html to directory-style paths

* ci(docs): remove d2 install step

* feat(docs): dark/light mode toggle

Add theme toggle button (top-right, fixed). Persists preference in
localStorage; falls back to prefers-color-scheme. Flash-free via inline
script in <head> that sets data-theme before first paint.

* fix(docs): fix inspect image path broken by directory URL restructure

* feat(docs): click-to-fullscreen for all article images

* fix(inspect): allow AssetsFS override in ServerOptions; drop unused request param from serveIndex

* fix(inspect): use in-memory FS in tests so TestAssets_FallbackToIndexHTML passes without web build
2026-04-23 00:57:30 +07:00
pj af7b7e27b0 feat(folio-web): web sample app + CDP spec tests (#36)
* fix(runner): allow nil connection for web platform

* fix(testrun): skip SDK handshake for web platform

* feat(folio-web): add React/Vite web sample app

* feat(folio-web): add sanderling spec

* fix(chrome): use InsertText for multi-char text input

* feat(hierarchy): add Screen field populated from sanderling-screen attr

* fix(chrome): auto-detect viewport from CSS vars, fix InputText accumulation, expose route as screen

* fix(runner): fall back to hierarchy root screen when snapshot screen is empty

* fix(folio-web): broaden loggedIn extractor to all authenticated pages
2026-04-22 23:59:30 +07:00
pj eed99e58aa refactor: code organization cleanup (#35)
* chore: fix gitignore + decisions doc after web->inspect-ui rename

Update web/ references to inspect-ui/ in .gitignore and Makefile. Add
decisions.md tracking architectural decisions from code-org discussion.

* refactor: rename pkg/spec-api to pkg/spec

Aligns the directory name with the npm package name @sanderling/spec.
Updates Makefile, package.json directory field, and resolveSpecAPIPath.

* refactor(verifier): split bindings.go into types.go + bindings.go

Move shared public types (Action, ActionKind, LogEntry, Exception) to
types.go. bindings.go retains internal JS runtime wiring only.

* refactor(inspect): split runs.go into runs.go, runs_cache.go, runs_decode.go

runs.go: types (RunSummary, StepSummary, RunDetail, Run) and Scan.
runs_cache.go: Cache type, Open/Step/Detail methods, parseRun, scanSteps.
runs_decode.go: readMeta, tallyTrace, decodeStepSummary, validRunID.

* refactor: move android_env.go to internal/android/

Extracts Android device/AVD/adb logic into internal/android package.
Exports EnsureDevice, AdbReverse, AdbReverseRemove, EnvWithAndroidPlatformTools, AdbBinary.
Moves tests to internal/android/android_test.go. cmd/sanderling becomes a thin caller.

* refactor: extract test pipeline to internal/testrun/

runTestPipeline logic moves to testrun.Execute. buildDriver, resolveSpecAPIPath,
pickFreePort, and the progress logger move to internal/testrun/. cmd/sanderling/test_run.go
becomes a thin adapter. Tests follow their code.

* ci: update workflow paths after pkg/spec-api -> pkg/spec rename
2026-04-22 20:35:34 +07:00
pj 46abb28ed6 feat(sdk-android): snapshot property delegate + feature-scoped files + screen fix (#34)
* rename web to inspect ui

* feat(sdk-android): add camelToSnakeCase conversion

* feat(sdk-android): add snapshot() property delegate with tests

* feat(folio): add feature-scoped sanderling snapshot objects

* refactor(folio): slim FolioApplication to snapshot object references

* fix(folio): rename route snapshot to screen, update spec.ts
2026-04-22 19:55:45 +07:00
pj d5b6f6ccaf feat: Maestro driver integration + DeviceDriver architecture (#32)
* proto(driver): drop launcher_activity from LaunchRequest

* refactor(driver): rename Driver to DeviceDriver, drop launcherActivity from Launch

* refactor(driver): rename package maestro to sidecar

* feat(driver): add ChromeDriver backed by chromedp

* feat(hierarchy): replace XML parser with TreeNode JSON parser

* refactor(driver): update mock and runner to DeviceDriver, drop launcherActivity

* feat(runner): add platform routing for web vs sidecar

* test(verifier): update hierarchy fixtures from XML to TreeNode JSON

* feat(sidecar): extract readLogcat/readProcMetrics; add MaestroDriverBackend

- Extract readLogcat() and readProcMetrics() as internal package-level
  helpers parameterised by serial
- Add MaestroDriverBackend wiring maestro-client AndroidDriver
- Drop launcherActivity from DriverBackend interface and StubDriverBackend
- Add maestro-utils and micrometer-core as explicit compile deps

* refactor(sidecar): inject DriverBackend into DriverService; drop launcherActivity

- Remove serial and default backend from DriverService constructor
- Make backend a required parameter
- Drop launcherActivity from launch RPC handler
- Create MaestroDriverBackend in Main.kt when platform is android

* test(sidecar): update tests for dropped launcherActivity and required backend

* chore(sidecar): remove unnecessary micrometer-core direct dependency
2026-04-22 19:02:39 +07:00
pj b667abbbca perf: reduce per-step iteration time (~3.9s to ~2.3s) (#29)
* perf(sidecar): use exec-out + tmpfs for hierarchy dump

Avoids FUSE overhead on /sdcard and shell startup cost by using
exec-out with /data/local/tmp. Saves ~100ms per hierarchy fetch.

* perf(sidecar): replace Thread.sleep waitForIdle with real idle detection

Poll `dumpsys window -a` for mAnimating=true every 50ms instead of
blindly sleeping. Breaks early when device is idle, saving 500-800ms
per step since most settle in <200ms after an action.

* test(sidecar): add idle detection parsing tests

* perf(runner): parallelize hierarchy, metrics, and logs fetch

Run fetchHierarchy, captureMetrics, and collectLogs concurrently via
errgroup so metrics+logs (~150ms) hide behind the hierarchy fetch
(~2s) instead of running serially.

* perf(runner): pipeline post-action screenshot with next step

Defer the post-action screenshot from step N and run it concurrently
with step N+1's hierarchy/metrics/logs fetch. Saves ~335ms per step
by hiding screenshot latency behind the hierarchy fetch.

* test(runner): add tests for parallel fetch and pipelined screenshots

Verify that hierarchy, metrics, and logs are all called per step.
Verify post-action screenshots are written with correct step indices
when pipelined, including the final flush after the loop.

* perf(sidecar): grep mAnimating on-device instead of pulling full dump

The full `dumpsys window -a` output is ~88KB per poll. Running
grep on-device transfers only a count byte, cutting per-poll
overhead from ~63ms to ~56ms and eliminating 88KB of ADB transfer.
2026-04-22 17:43:27 +07:00
pj faebfe379d test(folio): replace tautological properties with 4 domain invariants (#28)
* test(folio): replace tautological properties with 4 domain invariants

Drop properties that can't fail (e.g. List.size >= 0, balances are Long
integers) or that just check the extractor itself (account_count equals
accounts.length). Keep auth routing liveness, error-clear liveness, and
reachability goals.

Add four properties that target the real code paths:
- balanceMatchesTransactionDelta: a new ledger row shifts the balance
  by exactly its signed amount (credit +, debit -).
- totalEqualsSumOfAccounts: home total equals sum of per-account
  balances at every state, not only on home.
- balanceChangeRequiresActiveAccount: an account's balance can only
  change while that account is the active one in the navigator.
- duplicateAccountNamesRejected: account names are unique under the
  app's actual dedup rule (case-insensitive after trim).

* chore(folio): upgrade sdk-android to io.github.priyanshujain.sanderling:0.0.1-rc4

Group id moved from io.github.priyanshujain to
io.github.priyanshujain.sanderling in the rc4 publish.

* chore(folio): pin @sanderling/spec to 0.0.1-rc4
2026-04-21 17:53:55 +07:00
pj 75780db3fa refactor(sdk-android): publish under io.github.priyanshujain.sanderling + ci docs fix (#27)
* refactor(sdk-android): publish under io.github.priyanshujain.sanderling

Move the published Maven coordinates to a sanderling sub-namespace so the
brand is visible in the dependency line (was io.github.priyanshujain:sdk-android).
Sub-namespace is auto-allowed by Sonatype under the verified parent groupId.

* chore: update sdk-android coordinates in folio + docs

Follow the groupId change to io.github.priyanshujain.sanderling:sdk-android.

* ci(docs): use --version flag for d2 install script

The d2 installer accepts --version vX.Y.Z, not --tag. The --tag form
was rejected as "unrecognized flag" on the docs workflow run after
PR #26 merged.
v0.0.1-rc4
2026-04-21 15:17:56 +07:00
pj 08288202cb docs+install: post-rename docs polish, install script, d2 architecture diagram (#26)
* docs: add sanderling bird artwork to README and docs index

* chore: add one-line install script for macOS and Linux

Detects os/arch, resolves latest (or pre-)release, verifies sha256,
and installs the binary into $HOME/.sanderling/bin.

* docs: use the one-line installer in getting-started

Replaces the broken `<version>` placeholder snippets with the
install.sh one-liner.

* docs: drop filler line under the install one-liner

* docs: rename index heading to Sanderling Manual

* docs(style): adopt JetBrains Mono and uppercase brand mark

* docs(getting-started): use justfile flow for folio sample

* docs(writing-specs): drop 'coming soon' notes for eventually and implies

* docs(inspect): rewrite layout for tabbed state panels and metrics chart

* docs(inspect): add UI screenshot

* docs: add Inspect to sidebar and index

* docs: restore mermaid bootstrap script in page template

* docs(style): constrain article images to content width

* docs(inspect): drop layout prose, keep what the screenshot doesn't show

* docs(architecture): add d2 source for architecture diagram

Replaces the in-page mermaid block with a d2-rendered SVG. Generated
outputs (svg/png) stay out of git; only the .d2 source is checked in.

* build(docs): render d2 diagrams into build/site/_assets/diagrams

* docs(architecture): swap mermaid block for rendered d2 svg

* ci(docs): install d2 before building the site

* docs(architecture): tighten layout and reroute label-crossing edges

Flip device/sidecar order so trace writer drops cleanly to runs/
without cutting through the JVM cell, right-align the inspect row
via a pad column, and tune grid gaps to keep gRPC and Unix socket
labels off the SANDERLING boundary.
2026-04-21 15:05:00 +07:00
pj 8ccf95c1cf refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling

Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.

* chore(proto): regenerate driverpb after module path rename

The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.

* refactor: rename CLI binary uatu -> sanderling

Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.

* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk

Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.

* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar

Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.

* refactor: rename Uatu API surface -> Sanderling

- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
  spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
  __uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.

* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling

Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.

* chore(build): rename gradle property + rootProject.name uatu -> sanderling

- Renames the uatu.version gradle property and all its -P references in
  Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.

* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1

Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.

* refactor: rename npm package @uatu/spec -> @sanderling/spec

Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.

* docs: rename uatu -> sanderling in README, docs, and URLs

- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
  'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
  internal/inspect/server.go.

* refactor: rename remaining internal uatu strings -> sanderling

- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
  localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
2026-04-21 11:57:49 +07:00
pj 13bb2feb82 feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI

Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.

* test(trace): cover EndedAt + new step fields round-trip

* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()

Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.

* feat(runner): stamp residuals, hierarchy, selector targets, ended_at

Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.

* feat(inspect): scaffold embed dist for SPA assets

Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).

* chore(web): ignore web/ build output in root .gitignore

* chore(web): add bun + vite + vitest scaffold config

* feat(web): monochrome design tokens, typography, app shell CSS

* chore(web): placeholder for self-hosted JetBrains Mono fonts

* feat(web): index.html entry with style links and root mount

* feat(web): typescript types mirroring run/step trace schema

* feat(web): typed fetchers for runs/steps/screenshots

* feat(web): App shell with router and run/step routes

* feat(inspect): runs scan, lazy step parse, mtime-aware cache

* feat(web): RunList route with table, loading, and error states

* feat(web): RunDetail route shell with three placeholder panels

* feat(inspect): fsnotify-backed runs watcher with debounce

* fix(web): use jest-dom/vitest entry so matchers register

* test(web): cover listRuns happy path and error response

* test(web): render RunList with mocked fetch and assert row

* chore(web): commit bun lockfile

* feat(inspect): http handlers for runs/steps/screenshots/SSE

* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy

* feat(cmd): add 'uatu inspect' subcommand

* fix(web): align TS types with snake_case wire format

Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.

* feat(web): add ActionList panel for run-detail step navigation

* feat(web): add SnapshotTable panel with diff highlighting

Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.

* test(web): cover SnapshotTable rendering and diff behavior

Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.

* feat(web): add Screenshot panel with bounds and tap overlays

Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.

* fix(web): guard scrollIntoView call for jsdom compatibility

* test(web): cover Screenshot panel rendering and overlays

* test(web): cover ActionList rendering, selection, keyboard, and markers

* feat(web): add ExceptionsPanel component

* test(web): add ExceptionsPanel tests

* feat(web): add Timeline panel with property swimlanes

Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.

* test(web): cover Timeline empty state, cells, status, click, highlight

* feat(web): add ResidualNode recursive AST renderer

* test(web): cover ResidualNode operators, predicate, and error chip

* feat(web): add ViolationsPanel with status badges and jump button

* test(web): cover ViolationsPanel rows, status grouping, and jump button

* test(web): register testing-library cleanup globally

All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.

* feat(web): hooks for url/keyboard/theme/sse

* feat(web): wire all panels into run-detail with phone-dominant grid

ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.

* test(web): add three reference run fixtures (clean, violation, exception)

* build: web targets in Makefile + bun in CI; docs(inspect)

- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
  install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
  vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.

* feat(runner): capture a screenshot per step

The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.

* feat(inspect): include action_label in StepSummary

Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.

* test(inspect): accept either #app or #root in SPA shell fallback

* feat(web): render action_label and screen in ActionList rows

Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.

* feat(sidecar): implement screencap for android driver backend

Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.

* feat(proto): add Metrics RPC for per-step CPU and memory capture

* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes

* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status

* feat(runner): capture metrics + before/after screenshots per step

Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.

* fix(runner,sidecar): measure CPU across step via /proc stat delta

'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.

* feat(web): add Metrics type for per-step cpu and memory

* refactor(web): replace --accent-change with --accent-positive token

* refactor(web): recolor chip-progress as neutral outlined chip

* refactor(web): use neutral border for changed snapshot rows

* feat(web): add MetricsChart panel with HEAP and CPU lanes

SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.

* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows

Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.

* fix(runner): stop copying Tap selector into action.text

The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.

* fix(web): use text-muted for swipe arrow after accent-change removal

* fix(web): snapshot values truncate with ellipsis + title tooltip

Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.

* feat(web): state-before/after columns + metrics chart at bottom

RunDetail now renders a four-column grid:
  actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.

* fix(web): skip zero-value ticks + add exception markers to metrics

HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.

* fix(web): let action body column shrink below its content

Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.

* feat(web): bigger state screenshots + single properties row

Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.

* feat(web): add minimal Tabs component

Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.

* feat(web): fold timeline into MetricsChart as STEPS lane

Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.

* refactor(web): tabbed state cards, drop standalone Timeline panel

State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.

* feat(web): lock app shell to 100vh with no page scroll

html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.

* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)

Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).

* feat(web): ViolationsPanel supports violationsOnly filter

* feat(web): ActionList arrow-key nav + listbox semantics + smaller font

Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.

* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets

Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.

* feat(web): fourth 'Violations' tab + wider actions + shorter metrics

Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.

* feat(web): compact RunDetail layout using 1px borders instead of panel padding

* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis

Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.

* fix(web): RunList rows no longer stretch to fill viewport height

Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.

* misc changes

* fix(web): hoist useState above early return in MetricsChart

Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.

* fix(web): subscribe to named SSE event instead of 'message'

Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.

* fix(inspect): unsubscribe SSE clients on disconnect

Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.

Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.

* fix(trace): rename resolvedBounds/tapPoint to snake_case

Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.

* chore(web): drop vitest and remove UI tests from CI

No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).

Fixes CI failure where `vitest run` exits 1 with no test files.

* chore(make): dedupe sidecar embed and drop recursive make

Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
2026-04-21 11:32:17 +07:00
pj 36188ca906 test+refactor: real sidecar test, deterministic test sleeps, slog step line (#22)
* test(sidecar): replace assertTrue(true) placeholder with real server test

MainTest.mainExists() always passed and inflated the green-check count.
DriverServiceTest covers RPCs, but SidecarServer start/stop had no
coverage. Drop the placeholder and add SidecarServerTest that binds to
port 0, asserts a real ephemeral port, and stops cleanly.

* test(agent): drop 50ms sleep before cancel in TestServer_AcceptCancelsOnContext

Accept's closeListenerOnCancel watcher closes the listener as soon as
ctx fires, regardless of whether the outer Accept has reached
listener.Accept() yet. The sleep was a CI-flake surface (50ms is not
enough on a slow runner), and dropping it still exercises the same
outcome — Accept returns with ctx.Err() after cancellation.

Stable across 50x -count runs.

* test(agent): replace 2s sleep with done-chan in TestConn_SnapshotTimesOutIfSDKSilent

The silent-SDK fake held the connection open via time.Sleep(2s), which
coupled the test's wall clock to the server's 200ms snapshot-timeout
assertion. Swap for a done channel closed by t.Cleanup — the goroutine
exits when the test ends, independent of timing.

* test(maestro): make WaitForHealth_PollsUntilReady deterministic

Replace the 50ms wall-clock sleep that flipped healthReady with a
healthReadyAfterCall counter in the fake server. The handler returns
ready=true once healthCalls reaches the threshold, so the test's
"at least 2 polls before ready" assertion is satisfied by call
count rather than a race between the flip goroutine and the 25ms
poll loop.

* refactor(runner): route per-step progress through slog instead of fmt.Printf

The runner already carries a *slog.Logger for warnings (logger.Warn on
decode failures, predicate errors). The per-step status line was the
outlier — a bare fmt.Printf that wrote to os.Stdout unconditionally,
bypassing both the injected logger and any caller-configured writer.

Switch it to logger.Info("step", "index", ..., "screen", ..., "nodes", ...).
The caller (cmd/uatu) is responsible for wiring a logger whose handler
renders to the right stream; the next commit adds that wiring.

* feat(cli): render runner progress via a thin slog handler on stdout

progressHandler writes Info records as "msg key=value ..." and prefixes
warnings/errors with their level, matching the prose style of the
surrounding CLI status prints. Wired into the runner via
runner.Options.Logger so the per-step status line still lands on stdout
without slog's default time= / level= framing.
2026-04-20 17:58:41 +07:00
pj 323878c34a fix(verifier): don't crash on throwing JS predicates (#21)
* fix(verifier): don't crash on throwing JS predicates

formulaThunk used to panic whenever goja returned an error from a
predicate callable, and nothing on the LTL -> runner path recovered, so
a malformed spec (e.g. a property whose body throws or touches an
undefined field) would kill the verifier process.

Latch the first error on formulaState, return false so LTL marks the
property violated, and expose PredicateError(name) that walks the
property's formula-spec tree and surfaces the latched cause.

* fix(runner): log predicate errors alongside violations

For each violated property, surface the verifier's latched predicate
error via logger.Warn so operators can distinguish a genuine false
verdict from a malformed spec. Add a runner-level test asserting that a
throwing predicate no longer crashes the run and that the error message
appears in the log.
2026-04-20 16:55:10 +07:00
pj a2e96af1af WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio

Directory-level rename and path references in Go tests, bundle-check,
top-level README, and getting-started docs. Package declarations,
Gradle config, iOS bundle IDs, and class names follow in later commits.

* refactor(folio): rename Kotlin package dev.uatu.sample to app.folio

Moves source dirs and sqldelight schema from dev/uatu/sample to
app/folio, updates package declarations and imports, and switches
Android namespace/applicationId, iOS binaryOption bundleId, and
sqldelight database packageName to the new identifier.

* refactor(folio): rename SampleApplication to FolioApplication

Android manifest now points at .FolioApplication with label 'Folio'
instead of 'Uatu Sample'.

* refactor(folio): set iOS bundle id and display name to Folio

bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio.
CFBundleName + CFBundleDisplayName -> 'Folio'.

* refactor(folio): update demo email to [email protected]

* refactor(folio): point justfile at app.folio bundle id

Updates xcrun simctl launch target, uatu test --bundle-id, and the
build/uninstall comments to reference folio instead of sample.

* test: update fixture package ids to app.folio

Sidecar activity-resolver test and verifier spec-integration XML
fixtures referenced the old dev.uatu.sample Android package. Updates
them to match the folio app's real package id so the tests stay
representative of what the CLI sees on-device.

* test(verifier): rename SampleApp identifiers to Folio

Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and
sampleAppHierarchyXML const (now loginHierarchyXML for consistency with
the other per-screen fixtures). Updates trailing sample-app mentions in
comments and assertion messages.

* refactor(folio): rename Gradle/npm/wasm project identifiers to folio

settings.gradle.kts rootProject.name, package.json + package-lock.json
name, and the WasmJS index.html <title> all still read 'uatu-sample' /
'Uatu Sample'. Realigns them with the Folio brand.

* docs(folio): rewrite README title + getting-started bundle id

examples/folio/README.md is now titled 'Folio' with the Kotlin source
paths corrected to app/folio. Getting-started example uses --bundle-id
app.folio. Harness launch message is now generic ('app under test')
since uatu-sample-harness is not specific to folio.

* chore(folio): drop trailing 'sample' reference in gradle.properties

* refactor(folio): rename LoginPage composable to LoginScreen

Align with KMP/Android industry convention (NowInAndroid, Cash App,
JetBrains samples use Screen, not Page).

* refactor(folio): rename HomePage composable to HomeScreen

* refactor(folio): rename AddAccountPage composable to AddAccountScreen

* refactor(folio): rename LedgerPage composable to LedgerScreen

* refactor(folio): rename AddTransactionPage composable to AddTransactionScreen

* refactor(folio): split Models.kt into app.folio.data package

Account, Transaction (with TxnType), and Session move into their own
files under app.folio.data, matching NowInAndroid-style per-type
organization.

* refactor(folio): move data layer into app.folio.data package

Repository, LedgerStore (expect + interface), SqlLedgerStore,
WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext,
and Snapshot move into app.folio.data. Update all consumer imports.

* refactor(folio): move Navigation into app.folio.navigation package

Split the former Navigation.kt into Route.kt (sealed interface) and
Navigator.kt (singleton). Update consumer imports across screens,
App.kt, and FolioApplication.

* refactor(folio): move Platform and Format into app.folio.platform

Both files carry expect declarations (Platform object, formatDate);
grouping them into a dedicated platform package makes the KMP seam
obvious and mirrors the structure used by JetBrains samples.

* refactor(folio): move login into feature/auth package

Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline
the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into
LoginScreen since it is the sole caller.

* refactor(folio): move HomeScreen into feature/home package

* refactor(folio): move account creation into feature/account package

AddAccountScreen gets its own AddAccountUiState colocated with the
screen, replacing the shared UiState.addAccountError.

* refactor(folio): move ledger screens into feature/ledger package

LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger
with AddTransactionUiState (txnError, txnFormType) colocated. The
former catch-all UiState.kt is removed now that each screen owns its
state alongside its UI.

* refactor(folio): split Theme.kt; move theme and icons to subpackages

Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and
Type.kt (typography) under app.folio.ui.theme. Icons moves to
app.folio.ui.icon. Update every consumer's imports to match.

* refactor(folio): split ui components into per-file under ui/component

Former Widgets.kt and Components.kt become 10 focused files: AppButton,
Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/
BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style
of one composable per file in a designsystem/component package.

* chore(folio): consolidate uatu testing files under uatu/ folder

Move spec.ts, package.json, package-lock.json into examples/folio/uatu
so all uatu-specific testing artifacts live in one place. runs/ and
node_modules/ follow the same convention (both remain gitignored).
Update justfile, README, and the two Go consumers (bundle-check tool +
verifier/trace tests) that referenced the old path.

* refactor(trace): drop folio path in writer test

Round-trip only needs a non-empty string; neutralize to keep the
library free of folio references.

* refactor(sidecar): neutralize ResolveActivity test fixtures

Swap app.folio for com.example.app in the fixture strings so the
sidecar tests don't reference the example app by name.

* refactor(bundle-check): take spec path as argument

Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a
positional spec path instead so the tool works for any example and
leaves no folio reference in the library surface.

* test(verifier): add neutral integration spec and hierarchy fixtures

Adds testdata/integration_spec.ts with two routes ("list", "form"),
an InputText on text_field, a Tap on primary/secondary_action, a
safety property (itemCountNonNegative), and a liveness property
(submitEventually). Adds hierarchies_test.go with matching XML
fixtures. Constants intentionally go in a _test.go at package root
rather than testdata/hierarchies.go because go skips .go files
under testdata/.

* refactor(verifier): replace folio integration tests with neutral ones

Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test*
entry points to TestIntegrationSpec*. Uses the synthetic spec and
hierarchies added in the previous commit so the library's test suite
no longer references examples/folio at all.

Folio-specific coverage remains covered by examples/folio/justfile's
'just test' which exercises the real spec on device/emulator.

* chore: remove cmd/uatu-sample-harness

Not referenced by Makefile, docs, CI, or any script. Duplicates the
adb reverse helpers already in cmd/uatu/test_run.go, and its name
implies ownership by the sample app which violates the library/
example decoupling. If a bare-protocol debugging tool is later
needed it belongs inside cmd/uatu/.

* docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax

The directory-tree Layout section rots faster than the code and
duplicates what ls shows for free. The expect/actual paragraph in
Stack was the same kind of filler. The README also claimed 'just
AVD=Pixel_7 test' but the justfile reads AVD as an env var via
env_var_or_default, so the correct invocation is 'AVD=Pixel_7
just test'.
2026-04-20 16:04:20 +07:00
pj c1de4bf57a fix(sample-app): address PR #18 review findings (#19)
- WebLedgerStore.accountExistsByName case-insensitive (matches SQL COLLATE NOCASE)
- drop SampleApplication.maybeInjectDebugError; sample spec now passes clean
- drop noLogcatErrors from sample spec (default matches system-wide E logs)
- openRandomAccount picks uniformly from findAll instead of first match
- clear txnError on any add-transaction interaction, not only valid amount input
- escapeForAdbInputText quotes shell metacharacters (quotes, backslash, etc.)
- extract buildClearKeyevents helper with empty/normal/cap tests
- broaden spec_integration_test.go to cover home, add-account, ledger,
  add-transaction action generators
- document FocusTracker single-focus invariant
2026-04-20 13:31:50 +07:00
pj 8381a98aaf feat(sample-app): ledger spec parity (#18)
* feat(sample-app): add FocusTracker

* feat(sample-app): support description on TextInput

* feat(sample-app): support description on AppButton and Segmented

* feat(sample-app): stable ids on login screen

* feat(sample-app): stable ids on add-account screen

* feat(sample-app): stable ids on add-transaction screen

* feat(sample-app): stable ids on home and ledger screens

* feat(sample-app): hoist error + form state into UiState

* feat(sample-app): register auth_status + accounts snapshots

* feat(sample-app): register ledger_rows + ledger_balance snapshots

* feat(sample-app): register error + focused_input snapshots

* refactor(sample-app): scaffold spec.ts extractors + safety

* feat(sample-app): spec accounting invariants

* feat(sample-app): spec state-machine monotonicity

* feat(sample-app): spec liveness properties

* feat(sample-app): spec auth + account action generators

* feat(sample-app): spec transaction action generators

* feat(sample-app): spec weighted workflow

* feat(sample-app): publish login form input snapshots

* feat(sample-app): publish account + txn input snapshots

* feat(sample-app): register form input snapshots

* fix(sample-app): sequence login + idempotent text inputs in spec

* fix(sample-app): clear UiState on page dispose

* fix(sample-app): tighten adversarial login, allow retype on error

* test(verifier): update sample-app spec integration for new selectors

* feat(sidecar): clear focused field before InputText types

* refactor(sample-app): simplify loginHelper to focus-driven sequencing

* refactor(sample-app): drop input-value UiState mirrors

* refactor(sample-app): drop input-value snapshot registrations

* test(verifier): adjust sample-app integration for replace-on-type
2026-04-20 13:28:41 +07:00
pj 7493945251 feat: LTL operators, sampling, and default generators (#17)
* feat(ltl): add Now/Next/Eventually/Implies/Or/And/Not formulas

Replace the fold-with-latch evaluator with a residual-formula reducer.
Each Observe() instantiates a fresh obligation from the root (stripping
an outer Always), reduces each pending obligation against current state,
latches Violated on first failure, and surfaces Pending verdicts for
deferred obligations. Existing Always/Pure/Thunk tests continue to pass.

* feat(ltl): support relative duration for eventually().within()

* feat(proto): add Swipe, PressKey, RecentLogs RPCs

* feat(verifier,runner): formula handles, new action kinds, rich state

- verifier: add formula-spec registry; bindNow/bindNext/bindEventually with
  chainable .implies/.or/.and/.not and .within(n,unit) on eventually; bindFrom
  for uniform sampling. bindAlways keeps accepting plain predicates.
- verifier: store lastTree, lastAction, step time, logs, exceptions on the
  Verifier; SnapshotInput replaces the (snapshots, tree) pair. stateObject now
  produces state.lastAction/time/logs/exceptions matching the TS State type.
- verifier: make taps/swipes/waitOnce/pressKey built-in generators actually
  fire; taps picks a clickable, enabled element from the last hierarchy.
- agent: add exceptions field to Message wire format.
- driver: add Swipe/PressKey/RecentLogs to Driver interface; wire maestro
  client and mock driver. LogEntry exposed for runner consumption.
- runner: apply Swipe/PressKey/Wait actions; collect logcat and exceptions;
  pass lastAction and step time into PushSnapshot.

* feat(spec-api): LTL operators, new actions, richer State

- ltl.ts exports now/next/eventually; always overload accepts a Formula
- types.ts: Formula gains implies/or/and/not; EventuallyFormula adds .within;
  State gains lastAction/time/logs/exceptions; Swipe/PressKey/Wait action types
- actions.ts: Swipe/PressKey/Wait/from constructors; waitOnce + pressKey
  default generators
- tests exercise the chaining, sampling, and new actions through a recorded
  fake runtime

* feat(sidecar): add swipe, pressKey, recentLogs RPC handlers

* feat(sdk-android): capture uncaught exceptions

Install a default uncaught handler on Uatu.start, chained with any
existing handler so Android's crash reporter still runs. Expose
Uatu.reportError for callers to forward caught throwables. A bounded
circular buffer (default 50) drains into each STATE message's new
exceptions field. Protocol.kt serializes/deserializes the field,
matching the Go wire format added to internal/agent/protocol.go.

* feat(spec-api): add @uatu/spec/defaults/properties bundle

* feat(sample-app): exercise new LTL operators + defaults

spec.ts now imports eventually/next/now/from from @uatu/spec and
noUncaughtExceptions from @uatu/spec/defaults/properties. It declares
three properties that exercise the new surface:

- accountCountNonNegative: plain always() safety
- addAccountAdvances: always(now(x).implies(next(y)))
- eventuallyLoggedIn: eventually(p).within(30, "seconds")
- noUncaughtExceptions: imported default

The weighted actions root uses from() for random phone/name sampling
and entries for taps/swipes/waitOnce/pressKey built-ins.

SampleApplication gains a debug hook gated on the system property
uatu.inject_error so the e2e run can synthesize an Uatu.reportError and
verify noUncaughtExceptions violates.

cmd/uatu/test_run.go adds a subpath alias so specs importing
"@uatu/spec/defaults/properties" resolve against the in-tree source
when running from the uatu checkout. The spec-integration tests swap
the old click-counter fixtures for the new login hierarchy.

* feat(trace): record swipe/key/wait details + exceptions

trace.Step gains an Exceptions array so the trace captures the
class/message/stackTrace for each SDK-reported throwable in a step.
trace.Action gains FromX/FromY/ToX/ToY/Key/DurationMillis so the full
payload of Swipe/PressKey/Wait actions is visible in trace.jsonl.

sample-app's debug error hook now gates on ApplicationInfo.DEBUGGABLE
instead of a system property (adb setprop fails on non-rooted
emulators).
2026-04-20 02:19:39 +07:00
pj 5b5594ee05 Port sample-app to KMP with Android, iOS, Web, and SQLite (#16)
* chore(sample-app): hoist gradle wrapper to sample-app root

* chore(sample-app): add KMP root gradle config

* chore(sample-app): add composeApp KMP module build config

* chore(sample-app): add Android manifest for composeApp

* feat(sample-app): add shared domain models and auth constants

* feat(sample-app): add shared number and date formatting

* feat(sample-app): add cross-platform storage, clock, and id

* feat(sample-app): add shared repository with file-backed state

* feat(sample-app): add shared in-memory navigator

* feat(sample-app): add Compose theme and design tokens

* feat(sample-app): add shared UI components (icons, screen, widgets)

* feat(sample-app): add Login and Home pages

* feat(sample-app): add AddAccount, Ledger, AddTransaction pages

* feat(sample-app): add App root composable with routing

* feat(sample-app): add Android Application and Activity hosting Compose UI

* chore(sample-app): remove legacy android module (replaced by composeApp)

* chore(sample-app): bump to latest stable Kotlin/AGP/Compose deps

* fix(sample-app): make iOS compile (drop @Volatile, set bundleId)

* feat(sample-app): add xcodegen spec, SwiftUI host, and iOS Info.plist

* chore(sample-app): ignore build artifacts and generated xcodeproj

* chore(sample-app): update justfile for composeApp layout, add ios target

* docs(sample-app): rewrite README for KMP + iOS flow

* refactor(sample-app): drop in-app status bar and formatClock

* feat(sample-app): add SQLDelight schema and per-platform drivers

* refactor(sample-app): back Repository with SQLite, drop file serializer

* feat(sample-app): wire native back on Android and iOS edge swipe

* feat(sample-app): semantic roles, labels, and a11y descriptions

* fix(sample-app): link libsqlite3 for iOS target

SQLDelight's native driver needs libsqlite3.tbd on iOS; without it the
linker fails with undefined _sqlite3_bind_blob and friends.

* refactor(sample-app): abstract storage behind LedgerStore interface

Platform-specific createLedgerStore() returns a SqlLedgerStore backed
by SQLDelight on Android + iOS. Opens the door for a pure in-memory
web implementation that does not require a SQLite driver.

* feat(sample-app): add wasmJs target with in-memory LedgerStore

Wires a Compose Multiplatform browser canvas entry point. The web
implementation of LedgerStore is an in-memory model with localStorage
persistence, so it does not need a SQLite driver. Back navigation maps
the browser back button to the same BackHandler contract Android and
iOS use.

* chore(sample-app): settings + gitignore for wasmJs dev run

Registers the Node.js distributions repository and switches
repositoriesMode to PREFER_PROJECT so the Kotlin wasmJs plugin can
download its toolchain. Adds kotlin-js-store (lockfile) and ignores
runs/, web screenshots, playwright-mcp scratch output.

* fix(sample-app): singularize transaction count on Home

Shows "1 transaction" not "1 transactions" for accounts with a single
transaction; falls back to "$count transactions" otherwise.
2026-04-19 15:56:33 +07:00