Commit Graph
28 Commits
Author SHA1 Message Date
pj 48e901bd9b Merge branch 'android-home-diagnostic' into trial-merge 2026-08-15 22:51:48 +05:30
pj 3af6791029 docs(cli): the android doctor checks are not path-only 2026-08-15 22:38:34 +05:30
pj 19707df04d docs(ci): the android step number describes a local emulator, not ci
the leg disables animations and the number was measured with them on. the
first real dispatch carries 4 transitional steps over 200, so the cross-fade
wait does still fire in ci, just far less often.
2026-08-15 22:25:54 +05:30
pj e45fcfd04b docs(ci): the ios leg does not convict on the runner, and a seed cannot fix it
seed 28 reproduced its walk on macos-15 and reached the bug at the step it
convicts at locally. it still could not be judged: the run did not return
home between step 19 and step 136, so the counting invariant saw a rise of
15 against a window of 37 submits.
2026-08-15 19:39:38 +05:30
pj 788c68c777 docs(ci): record what the ios calibration assumes and where it was measured 2026-08-15 18:52:59 +05:30
pj 11f72a722a follow-ups from the pr #73 review (#77)
* ci(folio): run gradle on jdk 21 for the metro plugin

the metro gradle plugin folio builds with publishes org.gradle.jvm.version 21
and java 21 class files, so every leg failed at the folio build on a 17
runtime. local builds pass on jdk 25, which is why only ci saw it.

* fix(build): clean pkg/spec/dist, not the dead spec-api path

* chore: point stale spec-api comments at pkg/spec

* fix(spec): publish src so an installed package carries the runtime entries

* fix(testrun): alias the installed spec package so one module graph loads

* fix(spec): export Direction, ScrollAction and LongPressAction from the entry

* docs(spec): cut the package readme to a description and doc links

* docs: say how the cli and spec package versions relate

* fix(verifier): report whether the last action was confirmed applied

Both hosts get applied: true when the runner saw the dispatch succeed and
applied: null when it could not, so an unconfirmed action stops arriving at
the spec as no action at all.

* fix(runner): an apply error leaves the action's fate unknown, not undone

A deadline that fires after the tap was dispatched leaves the effect
committed. Reporting nil made the spec see an effect with no action to cause
it, which is how the counting property convicts a healthy app.

* fix(release): stage the sidecar jar at the renamed embed path

* test(replay-ui): trace fixtures for the vacuity counts

one real green run, one run that rendered nothing, one that judges every property at least once.

* ci(replay-ui): count the steps each property judged

the exit code says no property returned false; it does not say any property was ever evaluated. this reads the trace and reports judged vs declined per property, and fails when the step page never rendered.

* test(replay-ui): cover the summary script from make test

* ci(replay-ui): summarise through the vacuity script

* docs(ci): explain the replay-ui judged/declined counts

* fix(verifier): encode element-valued extractors into the trace

An ax element exports with its find/findAll host functions attached, and
json.Marshal refuses the whole value over them: json: unsupported type:
func(goja.FunctionCall) goja.Value. The encoding failed, curr stayed nil,
and the goja hosts (ios, android) recorded null for every element-valued
extractor in both the per-step diff and the violation witness.

Apply the web host's sanitize rule before marshaling, so one rule encodes
an element on both hosts.

* test(verifier): pin element encoding to one rule on both hosts

* test(runner): assert an element reaches trace.jsonl and its witness

* feat(spec): give state.lastAction an applied field

Three states, not two: no action is a null lastAction, applied: true is an
action the runner confirmed, applied: null is one it dispatched and never
learned the fate of.

* fix(folio): do not attribute an effect to an unconfirmed action

submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both
convict by pinning an effect on the last action, so both decline unless the
runner saw it applied. The fixtures now say which fate they mean.

* test(folio): an unconfirmed submit belongs in the window

The count is an upper bound on the submits a window holds, so the tap that may
have landed counts and committedTransactionsExceedSubmits has nothing to
convict on.

* test(runner): a tap that lands under a failed apply is not a double submit

Drives the real folio counting predicates through the runner against a device
that commits the tap and then times out. The double-submit case is the control:
without it a green proves only that the property never fired.

* test(verifier): pin the three lastAction states on both hosts

The web page is handed the same applied field the goja object exposes, so a
property cannot read one thing on native and another on web.

* docs(spec-language): document the three lastAction states
2026-08-15 15:51:33 +05:30
pj b02e86b2e3 ci: dispatch workflows for folio and the replay ui (#73)
* feat(runner): stop the step loop at the first violation on request

* feat(testrun): report violations as a typed error under exit-on-violation

* feat(cli): add --exit-on-violation and exit 2 when it fires

* docs(cli): document --exit-on-violation, --max-steps, and exit codes

* fix(web): enumerate and query across shadow roots in both producers

* test(chrome): compare both producers on a shadow-dom parity page

* test(browser): drive a canvas-under-shadow-root fixture end to end

* fix(web): select the focused field inside a shadow root before typing

* fix(web): report the pathname as the screen when there is no hash route

* fix(web): settle on dom quiescence instead of returning at body ready

* feat(replay-ui): add data-testid hooks the dogfood spec drives

* feat(replay-ui): add the dogfood spec sanderling runs against the replay ui

* fix(replay-ui): scope the screenshot property to the named state panel

* chore(make): add per-platform sanderling build targets

* ci: add dispatch workflows for folio and the replay ui

* docs: describe the dispatch workflows and how to read a failure

* ci(folio): give the ios leg its jdk, android sdk, just, and a clean app start

* refactor(web): use max for the settle budget

* ci: pin calibrated seeds, skip the flaky ios reinstall, bound every job

* ci: authenticate and pin the buf setup step

the anonymous release download hit the shared runner ip rate limit and
failed the job with 'socket hang up' after three retries.

* docs: record that canvas apps need a dom proxy to be text-fuzzable

* fix(ios): bound lifecycle rpcs and claim the target device

a launch the simulator rejects sent the xctest session into a recovery
chain that answered minutes late or never, and the rpc had no deadline,
so the run hung with no trace and no error. also take a per-udid flock:
a second run's reinstall lands under the first's live automation session
and wedges it.

* docs(ci): correct the ios hang wording and note the device lock

* refactor(verifier): derive the lastAction shape from one field list

both hosts must show a spec the same lastAction. one ordered list now
feeds the goja object and the json the web host installs, so they
cannot drift.

* fix(web): install lastAction in the page before extractors read it

state.lastAction was hardcoded null on web, so every property reading it
was silently vacuous: a correct property passed without ever firing.

* fix(web): carry element identity on actions and fix findAll on paths

an action's target was coordinates only, so a property matching on which
element was acted upon could never fire. ax.findAll([a,b]) also returned
nothing on web.

* fix(chrome): wait out a route transition before sampling facts

the tree stays byte-identical and quiet across a cross-fade, so both the
quiet timer and the unchanged-tree escape called it settled mid-flight
and extractors read two screens at once.

* fix: bound the pre-run app launch

launch happens before the runner starts, so --duration never covered it
and a wedged driver hung with no trace and no error.

* fix(folio): read balances from merged cards and treat unreadable as unknown

compose for web merges the whole accountcard subtree, so the balance
child never exists there and every card parsed as 0. the property then
compared 0 to 0 and fired on any submit, which is a false positive
generator. unknown is now null and null is vacuously true.

* test(folio): cover merged-card parsing and unknown balances

* ci(folio): make web an expect-the-bug leg

the web runtime can observe the double submit now, so the health gate
understates it. seed 1 finds it at step 109, 3 runs out of 3.

* docs(ci): explain why a submit tap landing on home is the bug

* fix(ios): read a StaticText's label as its text

AXValue was the only source for text, but a StaticText carries its
string in AXLabel, so nothing on screen had .text on ios: a spec reading
it saw everything on android and nothing here.

* docs(ci): correct the calibrated step ranges

* fix(folio): stop convicting on arithmetic float64 cannot hold

past 2^53 cents the gap between representable values is 128, so a real
1600-cent move reads back as something else and the equality is false
for a healthy submit as readily as a double one. also match parseCents:
a sign or an oversized amount is rejected, not read as an amount.

* test(folio): pin the safe-integer guard and its boundary

* docs: stop teaching the zero-default that caused a false alarm

* docs: write down the silent-vacuity failure modes

* feat(folio): tag the home total and the card transaction count

the total was the only untagged node on the screen, so the spec had to
sum cards and a clipped card broke the sum.

* fix(folio): read the app's own total and refuse contaminated windows

summing cards went null when one was clipped, and the null poisoned the
carrier for the rest of the run. the balance window also spanned every
transaction since the last home visit, so the property convicted on
deltas it could not attribute: the old web witness was 3.16x the typed
amount, not 2x.

* test(folio): pin the window rules and the count invariant

* fix(folio): never read a frame that shows two screens

android dumps a cross-fade with both screens in the tree. the route said
add-transaction while an unscoped find said home, so the oracle took a
half-rendered total as fresh and convicted on a tap that committed
nothing. one function now decides the route and returns null when the
frame is ambiguous.

* test(folio): cover transition frames, card readings and creation

* fix(folio): only disambiguate counts that came from merged text

the equal-length digit rule exists because web merges the card and an
account named -1 makes '12' ambiguous. a dedicated count node has
nothing to disambiguate, so applying it there threw away real evidence.

* ci(folio): pin the recalibrated seeds and drop android to a health gate

web 3 and ios 7 convict 3 runs out of 3 with an exactly 2x witness.
android convicts 2 in 5 because the same seed does not walk the same
trajectory there, so it proves the app runs instead.

* docs(ci): describe the two properties and why android cannot convict

* fix(android): wait out a route cross-fade before snapshotting

the dump could hold two screens at once, and the runner refuses to act
on such a tree, so a quarter of android steps applied no action and the
count varied per run: the same seed never walked the same trajectory.
the ios companion and the chrome driver already do this.

* ci(folio): let the android leg run far enough to see its conviction

* docs: only the repo owner merges

* ci(folio): a thrown predicate is not a conviction

exit 2 means the run recorded a violation, and a predicate that throws
is recorded as one too. so was newAccountBalanceIsZero, an unrelated
property in the same spec. the gate read the exit code and went green
with detection dead.

* ci: install idb-companion from its tap and stop interpolating inputs

idb-companion is not in homebrew-core, so the ios leg died before it
built anything. replay-ui expanded dispatch inputs into the shell.

* docs: correct the snippets and numbers that drifted from the code

* test(sidecar): pin that a slow read counts toward the stability streak

* fix(web): read the page's extractors only on steps that count

the page advances the spec's carriers when it evaluates, but the runner
applied the result only on non-transitional steps. a discarded step
moved the window forward anyway, so the next accepted pair bracketed two
transactions while counting one submit, and convicted a healthy app.
extractor errors now fail the run instead of leaving goja's values in
current against v8's in previous.

* fix(chrome): anchor the transition deadline when the dom goes quiet

it was anchored at script start, so a page that churned past the window
reached the check already expired and returned mid cross-fade. the
driver now publishes the idle timeout it needs, since the caller's 1s
could never spend the 800ms window.

* fix(web): fail on a partial extractor override

same mixed-producer hazard as the install error: some extractors hold
the page's value and the rest hold goja's, and a property comparing
across that split fires on a healthy app.

* docs: six of seven, the seventh is the stock property

* fix(folio): drop a name two cards answer to

homeTxnCountsOf keyed on the account name and let the last card win, so
two accounts the fuzzer named the same collapsed into one entry. a
reading that saw one Travel card and a later one that saw both then
subtracted two different accounts' counts, and
submitCommitsOneTransactionPerAction convicted a healthy app of
double-submitting. it is a gated property in folio-run.sh, so that reads
as "found the submit bug" over a card scrolling into view.

same rule createdAccountHasNonZeroBalance already applies: a name
nothing can attribute is no evidence. counted over every card, since an
unreadable twin spoils the identity too.

* perf(folio): read each frame once

every extractor asked routeOf, and routeOf does five ax.find calls. on
web each find walks the document and every shadow root beneath it, so
the spec cost 110 tree walks a step; homeCards was parsed four times
over. now 5 and once.

keyed on the identity of the state object because both hosts build a new
one per step and hand that one object to every getter, so it cannot
outlive its frame. holding the reference is what keeps that true rather
than likely.

* fix(web): keep an undefined reading's index through JSON

json has no undefined, so an extractor whose getter returned one had its
whole index dropped by JSON.stringify. that index then kept goja's
dump-derived value while its neighbours held the page's, and a property
comparing previous to current across the split fires on a healthy app.
folio has nine on(route, tag) extractors, so this was most extractors on
most steps.

each reading is wrapped in a {value} envelope: the drop now happens
inside the entry, and an absent value means the getter returned
undefined, which is what the goja host records for the same getter. a
json null would instead claim it returned null and x.current ===
undefined would answer differently on the two hosts.

* feat(verifier): report the registered extractor count

the web path needs it to check the page sent one reading per extractor.

* fix(runner): fail when the page reports fewer readings than extractors

the comment here already claimed a partial override was fatal. it was
not: the skipped check only catches indices outside the extractor list,
so a page reporting values for some extractors and not others left the
rest holding goja's reading of the dump with nothing said.

* test(browser): drive an undefined reading through the whole web path

four layers carry it: the page's envelope, the driver's unwrap, the
runner's count check and the verifier's decode. each has a unit test and
only a run proves they compose. goes red both ways, decoding an absent
value as null and dropping the envelope.

* fix(web): offer the aria roles a user activates

only role=button was in the tappable set, so link, checkbox, radio,
switch, tab, option, the menuitems and treeitem were invisible to the
enumeration however plain the control looked. the replay ui builds its
step rows as <li role="option">, and the spec dogfooding it had to
hand-write an action to reach them because no default verb could see a
single row.

both producers build the set from the same role list, since the parity
test compares them element by element.

* test(browser): tap a role-based control end to end

every control on the page is an <li role="option">, the shape the
replay ui gives its step rows, and the spec carries no action of its
own: the property firing is the evidence the default enumeration offered
a tap on one.

* fix(web): read aria-disabled as disabled

the enabled fact came off the disabled property, which only real form
controls have. it reads undefined on the role-based controls the
tappable set now covers, so every one of them looked enabled however
plainly it was marked otherwise, and the fuzzer would spend actions on
inert ones.

both producers answer the same two ways, and the parity fixture carries
a disabled row so the comparison covers it: reverting one side alone
names the element and the fact.

* docs(replay-ui): the enumeration reaches step rows now

the comment said role="option" is not in the tappable selector set,
which stopped being true a few commits ago. selectAStep stays, for the
reason the tab weight below it stays: one row among the page's clickable
elements is a thin chance, and both step-facing properties go vacuous on
a run that never selects one.

* test(runner): bound the last-action test by steps, not wall clock

100ms of wall clock against an assertion that two steps ran fatals under
load with "the web path never installed it", which reads as a
regression. every sibling test in the package uses a long duration and
MaxSteps.

* ci: run the kotlin tests in make test

RouteTransitionTest and the stability poll cover the android settle and
nothing in ci ran them. :sidecar:test needs no android sdk, checked by
running it with ANDROID_HOME pointed at nothing.

* fix(sidecar): measure the stability streak as observed quiet

parameterising pollUntilStable also moved the clock to the start of the
read that opened a run of identical snapshots, so a read's own duration
counted as quiet. the pre-existing caller polls a real uiautomator dump:
at 400ms a read, 750ms of required quiet became 250ms of observed quiet
and the poll settled in two reads instead of four.

the parameters stay, the semantics go back.

* test(sidecar): pin the transition cap by driving it

it asserted 1500 >= 700 + 300, two constants, which can only fail if
someone edits a constant. it now drives awaitSettledTree against a fade
that lands after 700ms and asserts it hands back the settled tree before
the cap. cut the cap to 1000 and it goes red.

* ci: pin buf-setup-action to a commit

it takes a token now, so a floating tag is a token handed to whatever
that tag moves to. note v1 there is a branch, not a tag, so the ref
lookup that resolves it is matching-refs/heads/v1.

* ci: declare least-privilege permissions

none of the three declared any, so each got the repository default.
release.yml and docs.yml already do this. all three only check out,
build, test and upload artifacts.

* ci: fail fast when a server never comes up

the readiness loops fell through silently after 30 tries, so a server
that never started surfaced as an opaque driver failure minutes later.
each now says what did not answer and on which port.

* ci(folio): a missing trace is not a verdict

with no trace the android gate ran its grep against ./trace.jsonl and
reported "never reached AddTransactionScreen, so it never got past
login", which is not what happened. the web and ios branches had the
same misdiagnosis on exit 0.

same class, one line up: the classifier's own failure was swallowed, so
with the evidence reader dead the gate printed a healthy run and exited
0.

* ci(replay-ui): skip a run directory with no trace

the summarise step is if: always(), and under github's bash -eo pipefail
an unmatched glob stays literal, the redirect fails, pipefail carries it
into the assignment and -e kills the step. so a failed fuzz run went red
twice, once for the real reason.
2026-08-15 13:01:27 +05:30
pj 26b49b379a fix ltl semantics and unify action enumeration (#71)
* fix(ltl): give every thunk a construction identity

Two distinct unnamed predicates both described as "Thunk(...)", so obligation
collapse merged their residuals and could drop a live violation. Identity is
assigned at construction and the fields are unexported, so a thunk cannot be
built without one.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): reduce a thrown-predicate residual instead of panicking

The verifier substitutes an ErrorFormula for the residual of a property whose
predicate threw, and that residual is fed back in on the next step. reduce had
no case for it, so the run crashed. It re-reports the same failure now.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): make a bounded always the dual of a bounded eventually

G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still
pending when the window closed, so nnf's negation normal form was not semantics
preserving. Both sides now range over the observations at which their inner can
definitely resolve: the eventually keeps a pending inner as a disjunct, and the
always discharges vacuously at window close.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): arm a one-shot root once per run

A root that carries its own horizon is one obligation for the whole run, not one
per observation. Re-instantiating a top-level eventually monitored G F<=n(p)
instead of F<=n(p) and left one live obligation per step behind; a bounded
always restarted its window every step and never closed.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): stop wrapping a top-level eventually in always

`eventually(p).within(300, "seconds")` as a property meant "within 300 seconds
of every step", which spawned an obligation per step with its own resolved
deadline. A 553-step run carried 553 of them and serialized a 75 KB residual.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): serialize the resolved deadline of a bounded window

Two obligations spawned at different steps from one duration-bounded formula
differ only in the deadline the evaluator resolved for them, so they serialized
identically and the trace erased a distinction the evaluator makes. The authored
window stays in amount/unit; the resolved deadline rides alongside.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): split a witness's origin step from its detection step

A deferred obligation spans two steps: the one that armed it and the one whose
reduction failed. They were conflated under one index, so the extractor snapshot
(which is the detecting step's state) was reported against the origin step.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): record a witness's detection step in the trace

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* feat(replay-ui): show the step a violation was detected at

The witness evidence is the detecting step's state, so say which step that is
and let a reader jump to it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): record the extractor state the predicates actually read

On the web path extractor bodies are evaluated in V8 and injected here, but only
the goja value was replaced. The trace diff and the violation witness therefore
described a state no property ever saw.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* refactor(spec): one candidate producer over one target-eligibility rule

Both hosts routed verbs themselves and both policies enumerated their own
actions, and all four drifted. Web sent `swipes` to scrollable containers only,
so swipe-to-dismiss on a list row was reachable on native and unreachable on
web; the model policy folded gestures its own way and could not reach what the
seeded picker drew.

A host now reports facts about every element and never decides which verb may
act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates
is the single enumeration, and the model policy reads it through
__sanderlingEnumerateBuiltin__ instead of reimplementing it in Go.

Gesture verbs change with it: scrolls stay vertical over scrollable containers,
swipes go free-form in all four directions from any element with real bounds.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): name a builtin scroll by its drag origin

A builtin gesture carries endpoints and no selector, so every scroll rendered as
"Scroll down " in the prompt's recent-action memory and two scrollable regions
were indistinguishable.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): clear storage over cdp instead of scripting an opaque origin

Launch runs while the tab is still on about:blank, whose opaque origin denies
storage access, so localStorage.clear() threw SecurityError and every web run
died at launch. Storage.clearDataForOrigin needs no navigation. The exception
helper lands here because "Uncaught" is what hid this for so long.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): enable the swiftshader webgl fallback

Headless Chrome runs with --disable-gpu, and without this flag it refuses the
software WebGL backend: getContext returns null, so a canvas-rendered app paints
nothing and every screenshot is identical black.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(web): resolve testTag through data-testid or id

Compose Multiplatform emits its testTag into the element id, which the native
table already accepts via the resource-id alias. The two web selector tables
were the only place that rejected it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* test(spec): type-check the spec api as part of make test

The fake runtime in api.test.ts did not return a chainable handle from extract,
so the file had not type-checked since named() was added. Wiring the check into
make test stops it drifting again.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* docs(manual): one-shot eventually and the gesture verbs

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
2026-08-12 18:06:04 +05:30
pj 7343085614 llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker

* feat(spec): make llm marker inert on the JS picker

* feat(spec): expose __sanderlingSampleInput__ corpus draw

* feat(openrouter): minimal chat-completions client

* test(openrouter): cover request shape, parse, and errors

* feat(verifier): thread screenshot + capture corpus sampler

* feat(verifier): LLM accessors — candidates, config, sampler

* test(verifier): cover AllCandidates, LLMConfig, SampleInput

* feat(trace): record action Source and LLMReasoning

* feat(runner): thread step screenshot into PushSnapshot

* feat(runner): llmSource selects actions via OpenRouter

* feat(runner): wire llmSource selection and trace stamping

* test(runner): cover llmSource selection, mapping, downscale

* docs(folio): add llm action-backend example spec

* docs(folio): document the LLM action backend run

* feat(llmclient): support OPENAI_API_KEY, openrouter wins

* refactor(runner): rename openrouter package to llmclient

* docs: both api keys, example model gpt-5.4-nano

* docs: add pr style rules to claude.md

* fix(runner): explain action kinds in llm prompt to stop swipe loops

* feat(trace): record llm ranked list and chosen rank

* feat(runner): stamp llm ranked list and chosen rank on trace

* fix(runner): tap by selector to survive layout shift after observe

* revert(runner): drop selector-first tap; broke path/testTag selectors

* feat(spec): llm() accepts optional instructions

* feat(verifier): read llm instructions off config

* feat(runner): append spec instructions to llm system prompt

* docs(folio): describe app in llm spec instructions

* feat(bundler): map generator export to globalThis.generator

* feat(verifier): read llm config off globalThis.generator

* feat(runner): gate llm source on --generator flag

* feat(cmd): add --generator llm|seeded flag

* test: cover --generator flag parsing and pickSources gating

* feat(verifier): enumerate llm candidates by walking actionsRoot

collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.

* test(verifier): cover candidate enumeration walk

* feat(verifier): add SetupAction to walk setup without the seeded root

* test(verifier): cover SetupAction setup-only precedence

* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order

* feat(trace): record llm choice number and chosen_action echo

* feat(runner): llm picks one number from weighted candidates

drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).

* test(runner): cover choice schema, strict-skip, and setup precedence

* refactor(verifier): drop the superseded AllCandidates enumeration

* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts

* fix(verifier): label editable fields by hint, not the typed value

an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.

* test(runner): cover weight-suffixed echo and stripWeightSuffix

* fix(runner): accept chosen_action echo that carries the weight suffix

real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).

* fix(verifier): skip llm enumeration on cross-fade frames

a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.

* feat(folio): show current balance on the add-transaction screen

renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.

* fix(replay): derive device space from screen extent, not first node

the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.

* fix(folio): show balance as a compact one-line label

per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.

* fix(folio): move balance into the header, one compact line under the account name

* fix(replay): attribute deferred violations to the causing step, not detection

* fix(replay): show a step's own violations in both panels, no next-step bleed

* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check

* chore: ignore .playwright-mcp scratch output

* docs: document the llm generator and --generator flag

* docs(spec): correct the llm() comment; config reads off globalThis.generator

* docs: add pr description rules
2026-07-31 21:12:00 +05:30
pj 991c583eb9 docs update with case study (#63)
* docs(manual): add introduction page

* docs(manual): rewrite getting started as guided first run

* docs(manual): rewrite writing specs as a folio tutorial

* docs(manual): document missing spec API in reference

* docs(manual): plain-language rewrite of runs page

* docs: real introductions on index pages and README

* fix(docs): sibling links from directory-style pages need ../

* fix(docs): correct sampling and restart-cost claims to match implementation

* docs: nav lists Introduction and Case study; roadmap points to milestone

* docs(manual): make getting started target the reader's own app, not Folio

* docs(manual): add Folio case study page

* docs: point manual navigation at the case study

* docs(readme): lead with the case study, fix roadmap link

* docs: roadmap links to milestone, sync clear-data default and cross-links
2026-06-09 20:13:54 +05:30
pj 90224dfd06 Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers

The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.

* feat(ios): resolve physical devices from devicectl

ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.

* feat(ioscompanion): runner-only device driver mode

NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.

* test(ioscompanion): cover device wiring, routing, and shell-out argv

Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.

* feat(testrun): route physical-device iOS runs to the device driver

Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.

* feat(doctor): device prereqs replace java/sidecar for ios-device

iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.

* feat(conformance): device backend uses iphoneos app and tunnel orphan checks

The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).

* feat(folio): device build linking the iosArm64 framework

project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.

* docs(cli): document ios-device doctor checks and the device flags

The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.

* fix(ioscompanion): resolve signing key path to absolute

xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.

* fix(ioscompanion): re-enable signing for the device runner build

companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.

* fix(ioscompanion): key the device build cache on signing identity

The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.

* docs(getting-started): document physical iOS device setup

Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.

* feat(ios): native usbmux client and in-process tunnel forwarder

Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.

* refactor(ios): drive device tunnel via io.Closer seam

Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.

* refactor(ios): remove iproxy spawn from device runner

* test(ios): cover tunnel close via io.Closer not child process

* feat(doctor): check usbmuxd socket instead of iproxy on PATH

* chore(conformance): drop iproxy orphan check; tunnel is in-process

* docs(ios): device tunnel uses native usbmux, nothing to install

* chore: gitignore the signing keys directory

* feat(folio): add Android launcher icon (black bg, white dot)

* feat(folio): add iOS app icon (black bg, white dot)

* feat(folio): add web favicon (black bg, white dot)

* docs(ioscompanion): fix stale const comments

* refactor(ioscompanion): inline single-use devicectl argv builders

* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests

* refactor(ioscompanion): inline firstNonEmpty

* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam

* test(doctor): trim redundant signing-check test

* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env

* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
2026-06-09 18:38:52 +05:30
pj b44077afde replay ui fix (#56)
* refactor: rename inspect to replay across the codebase

Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/,
the CLI subcommand from `sanderling inspect` to `sanderling replay`, and
updates all references in docs, Makefile, README, and Go comments.

* feat(replay-ui): show spec filename with full path on hover

RunList and RunDetail now render the basename of spec_path (e.g.
login.spec.ts) with the full path available as a title tooltip.
2026-06-03 16:17:26 +05:30
pj c5bb176be8 UX refactor (#52)
* feat(ltl): bound fields on AlwaysFormula and named thunks

Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.

* feat(ltl): negation normal form pass

nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.

* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse

Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.

* test(ltl): property-based NNF laws

Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.

* test(ltl): Finalize, bounded eventually, latch, collapse

Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.

* feat(inspect): within clause on always residual node

A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.

* feat(ltl): witness violations and (bool,error) predicate thunks

* test(ltl): migrate thunk call sites to (bool,error)

* feat(ltl): flag thrown-predicate witnesses with IsError

* refactor(verifier): replace predicate err side-channel with violation witness

* test(verifier): witness API for thrown predicates

* feat(trace): witnesses map and skipped-verification marker on Step

* feat(runner): thread violation witnesses, finalize, skip marker into trace

* test(ltl): lock violation witness reason, IsError, and step

* test(verifier): finalize surfaces unmet eventually with witness

* fix(ltl): eliminate implies and bounded-always false-negatives

Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.

* test(ltl): lock implies and bounded-always false-negative regressions

* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick

* feat(testrun): inject seed into web bundle via SANDERLING_SEED define

* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring

* test(spec): add Go math/rand/v2 PCG oracle and golden fixture

* feat(spec): bit-exact PCG port of Go math/rand/v2

* test(spec): assert pcg.ts matches the PCG golden fixture

* feat(spec): shared input corpus and press-key pools

* feat(spec): action-tree types and Host interface

* feat(spec): verb support matrix and warn-once helper

* feat(spec): deterministic shared action picker

* test(spec): verb matrix and warn-once semantics

* test(spec): picker draw-order and determinism

* refactor(spec): actions.ts returns pure GeneratorNode data trees

* refactor(spec): wire from() sampling through the picker rng

* feat(spec): shared runtime-entry installs next-action over pick.ts

* feat(spec): export LongPress/Scroll/longPresses/scrolls factories

* test(spec): assert data-tree shapes for action factories

* test(spec): runtime-entry serializeAction wire-contract round-trip

* refactor(spec): bridge data-tree nodes to the legacy goja picker tags

* fix(spec): web runtime walks the spec's globalThis.actions data tree

* test(spec): tolerate legacy bridge fields on builtin nodes

* refactor(spec): installRuntime accepts a lazy root resolver

The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.

* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker

Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).

* test(spec): cover the WEB Host surface and seed precision

Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.

* refactor(spec): picker emits native selector + scroll endpoints, setup precedence

* feat(spec): goja runtime entry wires the shared picker over the Go host

* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin

* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker

* refactor(spec): drop the legacy goja bridge fields from action factories

* feat(spec): serialize selector-only string targets for the runner to re-resolve

* refactor(verifier): one DecodeAction reads the unified flat wire contract

* refactor(verifier): goja host + shared picker replace the duplicate Go picker

* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime

* test(verifier): author specs through the shared picker path

* test(runner): bundle authored specs with the goja runtime entry

* feat(verifier): collect unsupported verbs for the run report

* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource

* feat(testrun): surface unsupported verbs in run report

* test(verifier): cross-runtime goja/node parity gate on the shared picker

* test(verifier): unsupported verbs collected deduped in first-seen order

* test(runner): summary reports no unsupported verbs on a clean run

* test(spec): golden-fixture cross-runtime parity gate for the node picker

Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.

* test(verifier): assert goja picker against the same cross-runtime golden

Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.

* refactor(spec): rename pressKey generator export to pressKeys

* refactor(spec): update barrel re-exports for pressKeys

* test(spec): update pressKeys generator export name

* docs(spec): rename pressKey generator to pressKeys

* refactor(spec): extract samplerRng into shared sampler-rng module

* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)

* test(spec): cover fluent value generators determinism and chaining

* refactor(bundler): inject globalThis trailer from spec named exports

* refactor(bundler): reuse registration trailer in web bundler

* test(bundler): cover named-export globalThis registration

* feat(spec): add named() to Extracted handle type

* feat(web-runtime): named() and cross-extractor read guard

* feat(verifier): named() and cross-extractor read guard in goja

* test(verifier): cross-extractor read guard and named()

* test(web-runtime): export runtime and extractors for tests

* test(web-runtime): named() and cross-extractor read guard

* refactor(folio): drop manual globalThis trailer (bundler injects it)

* refactor(folio): seed txn amounts via integers().between(1,500)

* refactor(folio-web): drop manual globalThis trailer (bundler injects it)

* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs

* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts

* refactor(folio-web): name extractors so violation witnesses are readable

* fix(web-runtime): propagate extractor getter throws and unpoison locked global

Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.

* test(spec): install fake runtime via defineProperty to survive locked global

* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors

* feat(runner): add MaxSteps bound to Options

* test(runner): MaxSteps stops after exactly N steps

* test(driverpb): drop proto getter round-trip tautology

* test(sidecar): drop stub-mode placeholder tautology tests

* test(mock): drop default-field-value assertion test

* test(ltl): drop Verdict.String tautology tests

* refactor(runner): extract RenderSummary for snapshot testing

* test(runner): golden snapshots for trace stream and violation summary

* feat(web-runtime): capture uncaught errors into state.exceptions

* test(integration): add throwing and counter web fixtures

* test(integration): add specs for the web fixtures

* test(integration): drive web fixtures through the real pipeline in headless Chrome

* chore(make): add test-browser target for the Chrome-driven suite

* ci: run the Chrome-driven browser suite in a separate job

* refactor(test): relocate browser suite to test/browser

* refactor(permissions): delete dead internal/permissions package

* refactor(test): rename package to browser_test

* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets

* chore(make): point test-browser at test/browser

* docs(decisions): record internal/permissions deletion

* refactor(doctor): use sidecarassets package

* refactor(testrun): use sidecarassets package

* fix(test): resolve testdata relative to browser_test.go

* refactor(verifier): remove dead __sanderlingIndex compat alias

* refactor(bundler): use encoding/json for JS string literals

* docs(action-space): use vendor-neutral native driver wording

* refactor(hierarchy): scrub backend tool name from comments

* refactor(driver): scrub backend tool name from comments

* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver

* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap

* refactor(chrome): implement DoubleTap as two taps with the gap

* refactor(mock): record DoubleTap and DoubleTapSelector actions

* refactor(runner): delegate double-tap to driver, drop gesture timing

* test(runner): assert double-tap delegates to driver DoubleTap

* docs(cmd): add package docs to CLI and developer tools

* docs(driver): add package docs to driver interface and chrome backend

* docs(driver): add package docs to mock and sidecar backends

* docs(platform): add package docs to android and ios device prep

* docs: add package docs to bundler and inspect

* docs(ltl): add package doc to temporal logic evaluator

* docs: add package docs to runner and testrun pipeline

* docs: add package docs to trace and verifier

* docs(sidecarassets): add package doc for embedded JAR loader

* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI

* test(chrome): gate real-Chrome driver tests behind the browser tag

* chore(make): run chrome driver tests in the browser job

* fix(web-runtime): guard global error listeners for non-browser hosts

The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.

* ci(browser): re-enable unprivileged user namespaces for headless Chrome

ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.

* ci(browser): pin stable Chrome for the driver tests

setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.

* feat(defaults): add scroll and rebalance action weights

Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.

* feat(defaults): trim scroll action weight wiring

* fix(build): point sidecar jar ignore and embed paths at sidecarassets

* test(defaults): drop stale longPresses re-export assertion

longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.

* fix(chrome): raise DevTools websocket read timeout to 60s

Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.
2026-06-02 09:52:53 +05:30
pj 88db9653e5 refactoring default action layer (#51)
* feat(hierarchy): add editable signal with native derivation

* feat(chrome): emit editable flag in hierarchy dump

* feat(verifier): expose editable on ax element objects

* feat(spec): add editable to selector and element types

* feat(verifier): register typing builtin generator

* feat(verifier): typing builtin types edge-case corpus into editable fields

* feat(spec): export typing builtin generator

* feat(spec): add defaultActions bundle

* feat(spec): export @sanderling/spec/defaults subpath

* feat(folio): layer defaultActions breadth over targeted flows

* test(verifier): typing builtin targets editable fields, declines otherwise

* test(hierarchy): editable derivation and selector matching

* test(spec): defaultActions, typing, and defaults barrel resolve

* fix(testrun): alias @sanderling/spec/defaults for the bundler

* test(chrome): editable flag for inputs, textarea, contenteditable

* feat(spec): typing builtin for the web (V8) action path

* chore(folio): auto-boot a bootable AVD in just test/install when none connected

* feat(driver): add ForegroundChecker optional capability

* feat(android): detect foreground package via adb dumpsys

* feat(sidecar): implement ForegroundApp via adb for android

* feat(runner): relaunch app when foreground escapes during exploration

* fix(spec): drop hardware back from defaultActions to stay in-app

* feat(spec): add DoubleTap action type and constructor

* feat(spec): wire DoubleTap through web-runtime serializer

* feat(verifier): bind doubleTap and decode DoubleTap actions

* feat(runner): dispatch DoubleTap as two taps inside one step

* test(doubleTap): cover constructor, verifier round-trip, and runner dispatch

* feat(folio): add noDuplicateTxnPerStep invariant and doubleSubmitTxn action

* fix(folio): track ledger row count across non-ledger steps; pin reproducer seed

* feat(spec): add doubleTaps random-target builtin to defaultActions

* feat(verifier): add doubleTaps random-target generator

* refactor(folio): drop doubleSubmitTxn; fuzzer surfaces double-submit via defaultActions

* fix(folio): make ledgerRowsSeen monotonic to suppress transient-render false positives

* feat(verifier): track newly-violated property set per step

Sticky `always(P)` violations re-surfaced on every step after onset,
flooding traces and summaries with duplicate records. EvaluateProperties
now diffs against the prior verdict map and records the onset set; a new
NewlyViolatedProperties accessor exposes it so callers can emit each
violation exactly once at its onset step. The verdict-map return is
preserved for residual / current-verdict consumers.

* refactor(runner): emit onset-only violations to trace and summary

Switch the per-step violation list from the sticky verdict map to the
verifier's onset set. Each property now appears exactly once across a
run: at the step it first violates, not on every subsequent step where
the residual stays false. Removes the dead violationNames helper.

* style(verifier): use maps.Copy for verdict snapshot

* fix(folio): make login spec content-driven (idempotent across re-entries)

* fix(verifier): canonicalize selector strings

Object/chain JS selectors used to fall through to goja's default
stringification, producing "[object Object]" tags that surfaced as
garbage in trace.action.selector. Emit canonical "k:v" / " > "-joined
strings instead so the tag round-trips back through the hierarchy
selector grammar.

* refactor(folio): replace txn invariants with balanceMatchesAddedTxn

Collapse noDuplicateTxnPerStep and newTxnChangesBalance into a single
per-row property: every newly-appearing ledger row's signed amount must
match the ledger balance delta. A double-submit lands two rows whose
individual amounts cannot both equal the aggregate delta, so each row
fires the property, catching both the row-count and balance-math
classes of bug under one semantic invariant.

* refactor(trace): drop WriteScreenshotAfter

Only one screenshot per step is captured now (concurrently with
hierarchy after settle), so the -after.png variant is unused.

* refactor(runner): one concurrent screenshot per step

Move screenshot capture into the post-action errgroup so it observes
the same UI moment as the hierarchy fetch. Drop the pre-action and
deferred -after captures. Skip WaitForIdle when the action is Wait
since the wait itself provides settling time.

* refactor(inspect-ui): use next step's screenshot for state after

Each step now has one screenshot (the moment of observation). The
"state after" view of step N is the same moment as step (N+1)'s
observation, so reuse that file rather than expecting a separate
-after.png.

* feat(sidecar): structural-hash settle poll

Add pollUntilStable and structuralHash helpers; wire them into the
Stub, Maestro, and iOS backends' waitForIdle. The structural hash
ignores bounds-only flicker (measure passes) but trips on any change
in resource-id/class/content-desc/text, so a Compose cross-fade where
both source and destination composables are momentarily alive no
longer slips through Maestro's waitForAppToSettle and contaminates
the next hierarchy fetch.

* test(sidecar): cover pollUntilStable and structuralHash

Verify the poll returns on two equal snapshots, after transient
churn, and at the cap when never stable; assert the hash ignores
bounds-only flicker and detects content changes.

* feat(spec): accept optional name on extract()

Add an (name, getter) overload so each extractor handle carries a
debuggable label that future trace fields (per-step diffs) can key
off. The web-runtime falls back to extractor_\${index} when none is
supplied so existing call sites keep working unchanged.

* test(spec): cover extract name overload

Verify the runtime receives an undefined name in the legacy shape,
the supplied name in the (name, getter) shape, and that
extract("name") with no getter throws.

* feat(verifier): name extractors for diff surfacing

bindExtract accepts an optional name argument; falls back to
extractor_N when omitted. The name is stored on extractorState
alongside prev/curr value caches that the next change will use to
emit per-step diffs.

* chore(folio): name every extract() call

Give each extractor in the Folio spec a debuggable label so the
inspect UI can render extractor-value diffs at violation steps
keyed by intent (ledgerRows, route, ledgerBalance, ...) rather
than by registration index.

* feat(verifier): track extractor value transitions

Cache each extractor's prior and current JSON-encoded value during
PushSnapshot; expose ChangedExtractors to surface per-step diffs the
runner can emit into the trace. The first observation flushes every
non-null extractor as a change so the inspect UI shows initial state
breadcrumbs alongside later transitions.

* test(verifier): cover ChangedExtractors diffs

Verify initial snapshot reports both named and fallback-named
extractors, a subsequent change surfaces prev/curr, and a no-op
snapshot leaves the diff empty.

* feat(trace): emit extractor_changes per step

Add ExtractorChanges to trace.Step and a runner helper that converts
the verifier's diff map into the trace shape. The inspect UI keys
its violation breadcrumbs off this field.

* feat(inspect-ui): render extractor-change breadcrumbs at violations

Show prev -> curr for each extractor whose value changed on the
selected step, anchored under the violation row in ActionList.
Long values collapse into <details> so the inline diff stays
readable while the full payload is one click away.

* fix(sidecar): cap stability poll independently of settle budget

The previous shape halved durationMillis between waitForAppToSettle
and the structural poll, then hammered hierarchy() at 80ms intervals
- on Maestro this stacked enough RPCs that hierarchy fetches began
timing out under load and the run stalled. Pass the full budget to
waitForAppToSettle and cap the follow-up structural poll at 600ms
with a 120ms interval, so the device sees at most a handful of
extra hierarchy reads per step.

* feat(cli): default --clear-data on so runs start fresh

* feat(sidecar): streak-based settle with route-transition detection

Two changes layered into the stability poll:

1. stabilitySnapshot returns null while the tree carries more than one
   route-level Screen tag (resource-id / testTag / identifier ending
   in "Screen"), so the poll cannot declare a NavHost cross-fade
   stable. Apps following the Compose route convention get this
   detection for free; apps that don't fall through to the generic
   signal below.

2. pollUntilStable now requires an uninterrupted stable streak of at
   least MIN_STABLE_STREAK_MILLIS rather than just N consecutive
   matches. A late transition that fires after a brief calm window
   breaks the streak instead of slipping past. Interval widened to
   250ms so UiAutomation isn't hammered under fuzz load.

* test(sidecar): cover streak reset and route-transition rejection

Verify the poll honors MIN_STABLE_STREAK_MILLIS, that a transient
mid-stream change resets the streak, that null returns block streak
progress through a NavHost cross-fade, and that stabilitySnapshot
counts only route-level attribute keys when summing Screen tags.

* feat(runner): re-fetch on transitional hierarchy capture

Some actions trigger async work (DB write, ViewModel coroutine) whose
navigation transition begins after the sidecar settle poll has already
exited. Without intervention, the next iteration's hierarchy fetch
lands mid cross-fade and the verifier observes a partial extractor
state which then surfaces as a false-positive violation at the step
where the transition completes.

fetchSyncedState pairs hierarchy + screenshot in one goroutine and
retries the pair (up to 4 times, 200ms apart) while the captured tree
contains more than one route-level *Screen tag. Steps that observe
no transition get no added cost; steps that catch a transition pay
up to ~600ms extra wall time but record a tree that matches the
post-transition state the property language expects to compare.

* feat(runner): gate first action on app reaching foreground

* test(runner): cover startup foreground gate and back-press

* feat(verifier): scope random-action targets to app package

Random tap/doubleTap/type/swipe candidates now exclude nodes whose package differs from the app under test, so exploration never fuzzes the soft keyboard, system UI, or permission dialogs. An unset app package or an element with no package stays in scope, preserving behavior on iOS.

* feat(testrun): pass app package into verifier scope filter

* test(verifier): cover package-scoped target selection

* feat(hierarchy): derive package from resource-id prefix

The Android sidecar omits an explicit package attribute, so the verifier's package scope filter was a no-op and the keyboard still leaked into targets. Native nodes carry their package as the resource-id prefix; derive it there when the attribute is absent. Compose testTags are colon-less and stay empty, keeping them in scope.

* test(hierarchy): cover package derivation from resource-id

* chore: stop tracking inspect-ui/dist build artifacts

* feat(android): detect focused-window package via dumpsys window

* feat(driver): add FocusedWindowChecker capability

* fix(runner): gate first observe on the app window being drawn, not just resumed

* test(mock): add FocusedWindowApp with foreground mirroring

* test(runner): cover startup gate waiting for app window to draw

* feat(proto): add Snapshot RPC for atomic hierarchy+screenshot

Pairs hierarchy and screenshot in a single response so the runner can
capture both under a backend mutex, avoiding the cross-fade race where
the two reads describe different frames.

* feat(sidecar): add snapshot default on DriverBackend

Default impl calls hierarchy() then screenshot(). The service layer wraps
the call in a mutex so concurrent runners observe a serialized pair.

* feat(sidecar): wire Snapshot handler with serialization lock

Synchronizes backend.snapshot() so concurrent runners observe a
serialized hierarchy+screenshot pair, eliminating the cross-fade race
where two parallel reads describe different frames.

* test(sidecar): cover Snapshot wire path and serialization lock

SnapshotHandlerTest asserts both fields are populated, concurrent calls
are serialized, and the default impl runs hierarchy then screenshot.

* feat(driver): expose Snapshot on DeviceDriver and sidecar client

Snapshot wraps the new atomic-snapshot gRPC: the runner gets hierarchy
and screenshot from one round-trip whose two reads are serialized on
the sidecar side.

* feat(driver): add Snapshot to chrome and mock drivers

The chrome tab is single-threaded so its Snapshot pairs the two reads
without extra locking. The mock records ActionSnapshot so tests can
assert the runner reaches for the paired RPC.

* refactor(runner): observe each step via the atomic Snapshot RPC

fetchSyncedState now issues one Snapshot per attempt so hierarchy and
screenshot describe the same on-device frame. The transitional retry
stays: that case handles a fully-captured but mid cross-fade frame,
which atomic capture cannot fix.

* test(runner): assert step uses Snapshot, not raw hierarchy/screenshot

TestRunner_UsesAtomicSnapshot catches regressions to the two-goroutine
race, and the existing parallel-fetch test now keys off ActionSnapshot.

* test(driver): cover Snapshot in proto descriptor and sidecar client

Adds Snapshot to the descriptor allowlist and a sidecar-client test that
asserts both fields come back over the wire.

* feat(trace): add Transitional flag to Step

* fix(runner): skip verifier for transitional trees after retry budget

When fetchSyncedState exits its retry loop with a tree that still shows a NavHost cross-fade, the runner now marks the step transitional, writes the step + screenshot to the trace, and skips Verifier.PushSnapshot / EvaluateProperties / ChangedExtractors so the previous-to-current extractor advance is not poisoned by transient state. The next clean step's previous still references the prior clean state. NextAction continues to run so the loop never deadlocks on a never-stabilizing screen.

* test(runner): cover transitional step skips verifier and clean control

* refactor(trace): rename Step.Action to Step.NextAction

The trace step's action field is the action chosen FOR THE NEXT iteration
based on observing this step's hierarchy, not the action that produced
this step. Rename Step.Action to Step.NextAction and the JSON tag to
next_action to make causality explicit at the data level.

* refactor(runner): assign trace action to Step.NextAction field

Follows the rename of trace.Step.Action to Step.NextAction. The runner
already computed the next iteration's action here; only the field name
changes.

* refactor(inspect): decode trace step's next_action JSON field

Mirrors the trace schema rename of action to next_action. The summary
shape exposed to the SPA (action_kind/action_label) keeps its current
JSON tags since these are derived labels, not the raw next-action.

* test(inspect): update fixtures to use next_action trace field

Aligns inspect tests with the trace schema rename. Step constructors
now set NextAction and the JSONL fixtures use the next_action tag.

* refactor(inspect-ui): rename Step.action to Step.next_action

Aligns the SPA type and consumers with the trace schema rename. The
StepSummary.action_kind/action_label labels stay unchanged since they
are derived labels, not the raw next-action.

* fix(folio): extract balanceMatchesAddedSum predicate as testable helper

Move the ledger-balance-vs-added-rows predicate into a pure helper module
so the property's logic is unit-testable in isolation. Marks the sanderling
example as an ES module so cross-package ESM imports resolve under node.

* fix(folio): use sum-of-added-rows in balanceMatchesAddedTxn

The old predicate (every row's signed amount equals delta) silently passed
the double-submit bug because two same-amount rows each match the delta in
isolation. Switching to the sum check (addedSum === delta) catches both the
double-submit case and any future multi-row append whose total drifts from
the balance change.

* test(spec): cover balanceMatchesAddedSum single, sum-match, over, under cases

Pins the sum-based predicate: a single new row matching delta and two new
rows summing to delta both hold; two-row over-sum (double-submit) and
under-sum cases both violate.

* fix(build): rebuild sidecar JAR when Kotlin sources change

Without source-file deps on $(SIDECAR_JAR), make never re-ran shadowJar
after a Kotlin edit, so a stale embedded JAR shipped on every install
and the new sidecar code was silently absent at runtime.

* fix(chrome): launch with no-sandbox so headless Chrome starts in CI

* fix(sidecar): type text at cursor instead of clearing the field

InputText now appends at the focus caret, matching the native driver
and the standard mobile-input contract, instead of deleting existing
content first. Adds an injectable command runner so the behavior is
testable without a device.

* test(sidecar): assert InputText types at cursor without clearing

Captures the adb command stream and verifies a single input-text call
with no preceding delete keyevents, plus the adb escaping cases.

* feat(proto): add LongPress RPC

* chore(proto): regenerate Go stubs for LongPress

* feat(driver): add LongPress to DeviceDriver interface

* feat(sidecar): add LongPress client method

* feat(mock): record LongPress action

* feat(chrome): implement LongPress as press-and-hold

* feat(sidecar): implement longPress across backends

* feat(sidecar): dispatch LongPress RPC to backend

* test(sidecar): cover LongPress dispatch

* test(sidecar): implement longPress in snapshot test backend

* feat(verifier): add LongPress and Scroll action kinds

* feat(folio-spec): predicate that gates balance check on TxnSubmit tap

Replaces the row-sum predicate (which always held by construction since
balance is derived from rows in Folio) with one that compares the typed
amount to the actual balance delta after a tap on TxnSubmit. Catches the
planted double-submit bug.

* feat(folio-spec): wire submitMovesBalanceByTypedAmount property

Adds lastAction and totalBalance extractors and uses them in the new
property. Drops ledgerRows/ledgerBalance extractors since nothing else
referenced them.

* feat(verifier): wire longPresses and scrolls generators

* test(verifier): cover longPresses and scrolls generators

* test(folio-spec): unit tests for submitChangesBalanceByTypedAmount

Covers single vs double submit, the DoubleTap variant, vacuous cases
(null action, wrong kind, wrong target, zero typed), and selector-as-
object coercion.

* feat(spec): add LongPress and Scroll authoring surface

* feat(spec): no-op LongPress and Scroll in web runtime

* feat(spec): re-export longPresses and scrolls as opt-in generators

* test(spec): cover LongPress and Scroll runtime members

* test(proto): expect LongPress in service descriptor

* feat(runner): dispatch LongPress and Scroll actions

* test(runner): cover LongPress and Scroll dispatch

* docs(action-space): move LongPress, Scroll, DoubleTap to current actions

* fix(runner): mark nil/empty hierarchy as transitional

A failed or empty sidecar hierarchy fetch was pushed straight to the
verifier, letting spec extractors crash with "Cannot read property 'map'
of undefined" when findAll returned null. Treat that case like a
transitional capture: skip the verifier push, still record the step, and
keep the loop progressing.

* fix(verifier): populate Action.On when tap chooser picks an element

Coordinate-targeted Taps/DoubleTaps left On empty, so action-gated
properties reading lastAction.on couldn't tell which target was hit and
were vacuously skipped. Resolve the picked element to a stable
key:value selector (resource-id, testTag, text, desc) and validate it
resolves back to the same element so we don't accidentally redirect the
tap to a sibling that shares the identifier.

* fix(folio): add parseTypedAmount helper matching app's parseCents

Raw user input like "50" must become 5000 cents, not 50. The existing
parseDollarCents helper strips non-digits and so reads "50" as 50 cents,
which is correct for formatted balance text but off by 100x for raw
input from the amount field.

* fix(folio): parse raw amount input as cents in submit predicate

txnAmountField holds raw user keystrokes, not formatted balance text.
Route it through parseTypedAmount so "50" reads as $50, matching how
the app commits the transaction.

* fix(folio): carry forward total balance across off-screen transitions

AddTransactionScreen shows neither AccountCard nor LedgerBalance, so the
extractor used to report 0 at the step before submit. That made every
non-zero current balance look like the full delta and tripped the typed
amount property on every honest submit. Remember the last-seen sum and
return it whenever the current snapshot has no balance signal.

* test(folio): cover submit predicate with raw typed-amount inputs

Pipes realistic raw keystrokes through parseTypedAmount + the predicate
so single submits clear and double submits fire as expected.

* feat(folio): add computeHomeTotalBalance helper

Pure helper that tracks Home multi-account total only and carries the last
Home sum across off-Home steps. Ledger's single-account balance is excluded
because mixing it would corrupt cross-screen scale comparisons.

* fix(folio): totalBalance carrier tracks only Home, not Ledger

Home cardSum is a multi-account total; Ledger's LedgerBalance is a single
account on a different scale. Blending them in the carrier produced bogus
cross-screen deltas (prev from Ledger, curr from Home), triggering false
positives in submitMovesBalanceByTypedAmount. Restrict the carrier to
Home AccountCard totals via the computeHomeTotalBalance helper.

* test(spec): cover computeHomeTotalBalance carrier behaviour

Tests Home sums, carrier passthrough on off-Home steps, the Ledger
scale-mismatch case, and a Home > off-Home > Home sequence.

* feat(runner): treat transient apply errors as transitional steps

Sidecar input RPCs occasionally hang with DEADLINE_EXCEEDED or
UNAVAILABLE on long fuzzing runs. The per-step loop previously
propagated any applyAction error and killed the run after a single
flake. Detect transient gRPC failures via status.FromError, mark the
step transitional, skip the post-action idle poll, and continue to the
next step. Fatal errors (outer ctx cancellation, non-transient codes,
verifier crashes) still propagate.

* test(runner): cover transient apply error resilience

TestRunner_TransientApplyErrorMarksTransitional drives the runner
through a wrapper that fails the first TapSelector with a gRPC
DeadlineExceeded then succeeds. Asserts the run does not exit, the
failed step is marked transitional with no violations, and the next
step runs cleanly. TestIsTransientApplyError_Classification covers the
helper's matching rules directly so future code changes don't quietly
drop a transient case.

* fix(folio): gate submit-balance property on Home route landing

totalBalance is only freshly computed when AccountCards are visible on
Home; off-Home landings return the carrier and would false-fire the
property, latching always(next(F)) to false and masking the real
double-submit bug. Skip vacuously when route is not "home".

* test(spec): cover route gate in submit-balance predicate

Adds route arg to existing cases (all use "home") and adds five new
cases: ledger landing with stale carrier, add-transaction with
double-insert delta, null route, plus home-landing positive and
double-insert negative cases anchoring the gate's allow path.
2026-06-01 12:48:51 +05:30
pj f572c8ba66 WIP: docs: refresh after iOS + web support (#50)
* docs: README covers iOS + web, surface both example apps

* docs(cli): document --ios-device and per-platform doctor

* docs: tighten README, fold examples into Docs list

* docs(runs): correct --clear-data lifecycle wording

Default behavior no longer wipes app data between runs; --clear-data is now opt-in.

* docs(getting-started): add iOS path, separate folio and folio-web

Document just test-ios under examples/folio, and distinguish the KMP
sample from the React + Vite folio-web sample.

* docs(inspect): document the eight panels

Lists Screenshot, ActionList, Timeline, ViolationsPanel, HierarchyPanel,
SnapshotTable, MetricsChart, ExceptionsPanel. Cross-links HierarchyPanel
to the spec language reference.

* docs(writing-specs): document setup export, flag noLogcatErrors as android-only

Mirrors pkg/spec/README.md so the manual covers the runner's setup-first
fall-through. Marks noLogcatErrors as Android-only so iOS/web spec
authors know it silently no-ops.

* docs(folio): document web target and iOS sanderling test recipe

After the KMP refactor folio also runs on wasmJs and the justfile exposes
just web, just web-build, and just test-ios. Surface all three.

* docs(folio-web): add README

Covers prerequisites, demo credentials, just test recipe, and how the
React + Vite host exposes state to the sanderling spec via stable ids
and data-* attributes.

* docs: scrub driver-implementation name from user docs

Drop the implementation tool name from README, cli.md doctor table, and
spec-language.md. These docs should describe behaviour, not the specific
underlying tool the native sidecar wraps.

* docs(development): scrub driver-implementation name from dev docs

architecture, design-principles, decisions now describe the native
sidecar by role (gRPC surface over OS UI-test pipeline) rather than by
the specific tool it wraps.
2026-05-25 16:17:21 +05:30
pj dd54c24c4e feat: --clear-data flag + typed attribute selectors (#48)
* feat(test): add --clear-data flag to clear app data on launch

* test+docs: cover --clear-data flag in CLI parser test and reference

* feat(spec): type AttrSelector with known attribute names

Replace AttrSelector = Record<string, string> with KnownAttrSelectors
plus a string|boolean index signature, so authors get autocomplete and
type-checking on testTag / focused / clickable / etc. while raw driver
attributes still type-check via the fallback. Boolean state attributes
accept native booleans; goja stringifies them at the marshal boundary.

AccessibilityElement.attrs becomes RawAttrs (typed string-valued shape
of the same canonical names) so element.attrs.testTag autocompletes.

* test(verifier): native boolean selector value matches focused=true

* docs+folio: use native boolean for focused selector and document typed attrs
2026-04-27 01:02:21 +07:00
pj 28909d954b docs: spec language reference page + writing-specs rewrite (#46)
* docs(architecture): add device/emulator node and XCTest edge to diagram

* docs(manual): rewrite writing-specs with accurate API and updated examples

* docs(manual): add spec-language reference page

* docs(sidebar): add spec-language entry to sidebar nav
2026-04-26 13:07:31 +07:00
pj 776becdf4b Remove in-app SDK (#43)
* chore: delete internal/agent package

* chore(build): remove sdk-android from gradle settings

* chore(makefile): remove sdk-android targets

* chore(ci): remove release-android job from release workflow

* chore(folio): remove sdk-android dependency

* chore(folio): remove SDK initialization from FolioApplication

* chore(folio): delete snapshot extractor files

* feat(folio): add balance to account card content description

* feat(folio): add hierarchy content descriptions to LedgerScreen

* refactor(folio): rewrite spec.ts to use ax extractors

* docs: remove in-app SDK from README

* feat(folio): add focused_input indicator to App

* docs: remove in-app SDK from index

* refactor(runner): remove agent SDK connection and snapshot step

* test(runner): update tests for SDK removal

* docs: remove Android SDK section from getting-started

* refactor(testrun): remove agent SDK connection setup

* docs: remove snapshots from writing-specs

* docs: remove in-app SDK from architecture doc

* docs(folio): update README for SDK removal

* docs: update per-step cycle diagram in architecture doc

* fix(folio): detect screens from unique element presence, not id: selectors

testTag() in Compose is not exposed as resource-id without testTagsAsResourceId.
Use desc: selectors for elements unique to each screen instead of id: path queries.

* feat(folio): add screen root contentDescription for scoped ax selection

Each screen root gets semantics { contentDescription = "ScreenName" } so
sanderling specs can scope element lookups through the screen: desc:LoginScreen > desc:login_submit.

* fix(folio): scope all ax selectors through screen root nodes

Use desc:ScreenName > desc:element path queries so every selector is
rooted at the screen level. focusedInput stays unscoped since it lives
in the app root, outside any screen.

* fix(folio): guard newAccountBalanceIsZero against navigation false positives

Scoped selectors return [] when not on HomeScreen so accounts vanish and
reappear as apparently-new on each visit. Skip the check when prev was empty.

* chore(folio): link @sanderling/spec to local pkg/spec for IDE type checking

* feat(spec): add desc, class, clickable, enabled, checked, focused, selected to AccessibilityElement

Runtime fields set by the verifier were missing from the TypeScript type,
causing linting errors on el.desc and related accesses in specs.

* chore(folio): switch to bun, add tsconfig.json for IDE type checking

- Remove package-lock.json, add bun.lock
- Add tsconfig.json so VSCode resolves @sanderling/spec types
- Fix parseAccount/parseLedgerRow to accept string | undefined
2026-04-25 20:04:29 +07:00
pj 88db0cbea8 docs: web platform + clean URLs + dark/light mode (#37)
* chore(docs): replace d2 diagram pipeline with mermaid

Remove docs/_diagrams/ and d2 build step from Makefile. The HTML
template already initialises mermaid.js; diagrams are now inline
code fences rendered client-side.

* docs(architecture): add mermaid diagram + web/CDP platform docs

Replace SVG img tag with inline mermaid flowchart showing both native
(Maestro sidecar + in-app SDK) and web (Chrome CDP) paths. Update
Processes, Transports table, and per-step cycle sections.

* docs(design-principles): update principles 1-4 for web platform

Principles 1, 2, 3, and 4 referenced Maestro and native-only concepts.
Add web/CDP context and update driver-is-an-interface to name both
sidecar and chrome implementations.

* docs(manual): add web prerequisites and folio-web example

Update --platform flag to list android, ios, web. Add web prerequisites
section (Chrome, no SDK needed) and folio-web quick-start to
getting-started.

* chore(gitignore): untrack inspect dist/index.html build artifact

index.html is regenerated by vite on every build with a new content hash,
making it permanently dirty. Only .gitkeep is needed for //go:embed to
compile on a fresh checkout. Also remove duplicate dist/* line and stale
d2 diagram ignore entries.

* feat(docs): click-to-zoom for mermaid diagrams

* docs(architecture): change diagram layout from LR to TB

* docs(getting-started): link npm and Maven Central package headers

* update docs root

* docs(spec): rewrite npm package README

Update usage example to current API, drop stale version-compatibility
and license sections.

* build(docs): output pages as pagename/index.html for clean URLs

Split DOCS_OUT into INDEX_OUT (index.md files stay as index.html) and
PAGE_OUT (all other pages become pagename/index.html). The __ROOT__
depth computation already handles the extra directory level correctly.

* chore(docs): update sidebar links to directory-style URLs

* docs: update cross-links from .html to directory-style paths

* ci(docs): remove d2 install step

* feat(docs): dark/light mode toggle

Add theme toggle button (top-right, fixed). Persists preference in
localStorage; falls back to prefers-color-scheme. Flash-free via inline
script in <head> that sets data-theme before first paint.

* fix(docs): fix inspect image path broken by directory URL restructure

* feat(docs): click-to-fullscreen for all article images

* fix(inspect): allow AssetsFS override in ServerOptions; drop unused request param from serveIndex

* fix(inspect): use in-memory FS in tests so TestAssets_FallbackToIndexHTML passes without web build
2026-04-23 00:57:30 +07:00
pj eed99e58aa refactor: code organization cleanup (#35)
* chore: fix gitignore + decisions doc after web->inspect-ui rename

Update web/ references to inspect-ui/ in .gitignore and Makefile. Add
decisions.md tracking architectural decisions from code-org discussion.

* refactor: rename pkg/spec-api to pkg/spec

Aligns the directory name with the npm package name @sanderling/spec.
Updates Makefile, package.json directory field, and resolveSpecAPIPath.

* refactor(verifier): split bindings.go into types.go + bindings.go

Move shared public types (Action, ActionKind, LogEntry, Exception) to
types.go. bindings.go retains internal JS runtime wiring only.

* refactor(inspect): split runs.go into runs.go, runs_cache.go, runs_decode.go

runs.go: types (RunSummary, StepSummary, RunDetail, Run) and Scan.
runs_cache.go: Cache type, Open/Step/Detail methods, parseRun, scanSteps.
runs_decode.go: readMeta, tallyTrace, decodeStepSummary, validRunID.

* refactor: move android_env.go to internal/android/

Extracts Android device/AVD/adb logic into internal/android package.
Exports EnsureDevice, AdbReverse, AdbReverseRemove, EnvWithAndroidPlatformTools, AdbBinary.
Moves tests to internal/android/android_test.go. cmd/sanderling becomes a thin caller.

* refactor: extract test pipeline to internal/testrun/

runTestPipeline logic moves to testrun.Execute. buildDriver, resolveSpecAPIPath,
pickFreePort, and the progress logger move to internal/testrun/. cmd/sanderling/test_run.go
becomes a thin adapter. Tests follow their code.

* ci: update workflow paths after pkg/spec-api -> pkg/spec rename
2026-04-22 20:35:34 +07:00
pj 75780db3fa refactor(sdk-android): publish under io.github.priyanshujain.sanderling + ci docs fix (#27)
* refactor(sdk-android): publish under io.github.priyanshujain.sanderling

Move the published Maven coordinates to a sanderling sub-namespace so the
brand is visible in the dependency line (was io.github.priyanshujain:sdk-android).
Sub-namespace is auto-allowed by Sonatype under the verified parent groupId.

* chore: update sdk-android coordinates in folio + docs

Follow the groupId change to io.github.priyanshujain.sanderling:sdk-android.

* ci(docs): use --version flag for d2 install script

The d2 installer accepts --version vX.Y.Z, not --tag. The --tag form
was rejected as "unrecognized flag" on the docs workflow run after
PR #26 merged.
2026-04-21 15:17:56 +07:00
pj 08288202cb docs+install: post-rename docs polish, install script, d2 architecture diagram (#26)
* docs: add sanderling bird artwork to README and docs index

* chore: add one-line install script for macOS and Linux

Detects os/arch, resolves latest (or pre-)release, verifies sha256,
and installs the binary into $HOME/.sanderling/bin.

* docs: use the one-line installer in getting-started

Replaces the broken `<version>` placeholder snippets with the
install.sh one-liner.

* docs: drop filler line under the install one-liner

* docs: rename index heading to Sanderling Manual

* docs(style): adopt JetBrains Mono and uppercase brand mark

* docs(getting-started): use justfile flow for folio sample

* docs(writing-specs): drop 'coming soon' notes for eventually and implies

* docs(inspect): rewrite layout for tabbed state panels and metrics chart

* docs(inspect): add UI screenshot

* docs: add Inspect to sidebar and index

* docs: restore mermaid bootstrap script in page template

* docs(style): constrain article images to content width

* docs(inspect): drop layout prose, keep what the screenshot doesn't show

* docs(architecture): add d2 source for architecture diagram

Replaces the in-page mermaid block with a d2-rendered SVG. Generated
outputs (svg/png) stay out of git; only the .d2 source is checked in.

* build(docs): render d2 diagrams into build/site/_assets/diagrams

* docs(architecture): swap mermaid block for rendered d2 svg

* ci(docs): install d2 before building the site

* docs(architecture): tighten layout and reroute label-crossing edges

Flip device/sidecar order so trace writer drops cleanly to runs/
without cutting through the JVM cell, right-align the inspect row
via a pad column, and tune grid gaps to keep gRPC and Unix socket
labels off the SANDERLING boundary.
2026-04-21 15:05:00 +07:00
pj 8ccf95c1cf refactor: rename project uatu -> sanderling (#24)
* refactor: rename Go module path uatu -> sanderling

Module path github.com/priyanshujain/uatu -> github.com/priyanshujain/sanderling,
including all imports and the proto go_package option. Generated .pb.go files
rewritten in-place; safe to regenerate with protoc later.

* chore(proto): regenerate driverpb after module path rename

The previous sed-based module rename corrupted the embedded descriptor
byte lengths. buf generate rewrites them cleanly.

* refactor: rename CLI binary uatu -> sanderling

Updates Makefile target + UATU_BIN var, .goreleaser project/build IDs,
.gitignore comment, and all user-facing strings in the CLI help text,
error messages, and tests. Binary is now bin/sanderling.

* refactor(sdk): rename Kotlin package dev.uatu.sdk -> dev.sanderling.sdk

Moves sdk/android/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the Gradle namespace. Class
names (Uatu, UatuRuntime) are renamed in a follow-up commit.

* refactor(sidecar): rename Kotlin package dev.uatu.sidecar -> dev.sanderling.sidecar

Moves sidecar/src/{main,test}/kotlin/dev/uatu -> dev/sanderling and
rewrites package declarations, imports, and the application mainClass.

* refactor: rename Uatu API surface -> Sanderling

- Kotlin: Uatu -> Sanderling, UatuRuntime -> SanderlingRuntime (+ files).
- JS host binding: globalThis.__uatu__ -> __sanderling__ (Go verifier,
  spec-api, tests).
- TS interface: UatuRuntime -> SanderlingRuntime; internal tags
  __uatuFormula / __uatuActionGenerator -> __sanderling* variants.
- Go trace: UatuVersion field + uatu_version JSON tag renamed.
- Socket naming: uatu-agent / uatu-agent-reader -> sanderling-agent*.
- Sample app, docs, inline-JS test strings updated to match.

* refactor(examples): rename examples/folio/uatu -> examples/folio/sanderling

Renames the example spec directory; updates justfile paths + gitignore
entries accordingly. Package.json name/description and @uatu/spec
dependency are renamed in the npm + docs commits.

* chore(build): rename gradle property + rootProject.name uatu -> sanderling

- Renames the uatu.version gradle property and all its -P references in
  Makefile, build.gradle.kts files, and .github/workflows/release.yml.
- settings.gradle.kts rootProject.name = "sanderling".
- Renames .env.local.example header + release-cli workflow job name.

* refactor(proto): rename proto package uatu.driver.v1 -> sanderling.driver.v1

Updates the proto package and java_package, regenerates driver.pb.go +
driver_grpc.pb.go, rewrites Kotlin imports and the gRPC ServiceName
assertion in driver_test.go.

* refactor: rename npm package @uatu/spec -> @sanderling/spec

Renames package name in pkg/spec-api/package.json + lockfile, all
consumer imports (examples/folio spec, testdata, verifier tests), the
esbuild alias in cmd/sanderling/test_run.go, and related doc references.

* docs: rename uatu -> sanderling in README, docs, and URLs

- README + docs/{manual,development}/*: narrative + GitHub + Pages URLs.
- POM + npm package.json repo/homepage/bugs URLs.
- .gitignore + embed_stub + Makefile-comment references updated to
  'make sanderling'.
- Minor narrative comments in cmd/sanderling/test_run.go and
  internal/inspect/server.go.

* refactor: rename remaining internal uatu strings -> sanderling

- SANDERLING_TEST_PHONE/OTP env vars (cmd + bundler tests).
- sanderling-sidecar runtime tmp dir + extracted JAR filename.
- Inspect web UI: @sanderling/inspect-web package, title, theme
  localStorage key, RunList empty-state copy, uatu_version TS field.
- Sample app storage key sanderling.ledger.v1.
- Test data: sanderling_test AVD name + com.example.sanderling_test.
- Release docs tarball name template.
2026-04-21 11:57:49 +07:00
pj 13bb2feb82 feat: uatu inspect UI (web trace explorer) (#23)
* feat(trace): extend Step/Action/Meta schema for inspect UI

Add Step.Hierarchy, Step.Residuals, Action.Selector/ResolvedBounds/TapPoint,
Meta.EndedAt and JSON tags on hierarchy.Element/Bounds/Tree so trace.jsonl
can drive the upcoming uatu inspect web UI.

* test(trace): cover EndedAt + new step fields round-trip

* feat(ltl): MarshalJSON for Formula AST + Evaluator.Residual()

Each Formula concrete type now serializes to a closed-set residual node
(true/false/not/and/or/implies/always/now/next/eventually/predicate/error)
that mirrors the TS spec API surface. Evaluator.Residual() folds pending
obligations into a single Formula so the runner can stamp one residual
per property per step into trace.jsonl.

* feat(runner): stamp residuals, hierarchy, selector targets, ended_at

Each Step now carries the captured hierarchy, per-property residual ASTs,
and (for Tap/InputText) the selector + resolved bounds + tap point. The
test_run command writes meta.ended_at on graceful shutdown so the inspect
UI can distinguish completed runs from in-progress ones.

* feat(inspect): scaffold embed dist for SPA assets

Stage 2 stub for the inspect server. Real web bundle gets wired in
Stage 4 (Makefile copies web/dist into internal/inspect/dist).

* chore(web): ignore web/ build output in root .gitignore

* chore(web): add bun + vite + vitest scaffold config

* feat(web): monochrome design tokens, typography, app shell CSS

* chore(web): placeholder for self-hosted JetBrains Mono fonts

* feat(web): index.html entry with style links and root mount

* feat(web): typescript types mirroring run/step trace schema

* feat(web): typed fetchers for runs/steps/screenshots

* feat(web): App shell with router and run/step routes

* feat(inspect): runs scan, lazy step parse, mtime-aware cache

* feat(web): RunList route with table, loading, and error states

* feat(web): RunDetail route shell with three placeholder panels

* feat(inspect): fsnotify-backed runs watcher with debounce

* fix(web): use jest-dom/vitest entry so matchers register

* test(web): cover listRuns happy path and error response

* test(web): render RunList with mocked fetch and assert row

* chore(web): commit bun lockfile

* feat(inspect): http handlers for runs/steps/screenshots/SSE

* test(inspect): cover handlers, screenshot whitelist, SSE, dev proxy

* feat(cmd): add 'uatu inspect' subcommand

* fix(web): align TS types with snake_case wire format

Go inspect server serializes RunSummary, StepSummary, Step, Meta with
snake_case JSON tags (matching the on-disk trace.jsonl/meta.json). Update
the TS types and consumers to match so API responses parse without
runtime undefined fields. Action keeps resolvedBounds/tapPoint as camelCase
because those keys were defined that way in the trace schema.

* feat(web): add ActionList panel for run-detail step navigation

* feat(web): add SnapshotTable panel with diff highlighting

Renders snapshots dictionary as a flat sorted dotted-path tree.
Changed leaves get data-changed plus a hover title with the previous value.

* test(web): cover SnapshotTable rendering and diff behavior

Eight cases: empty state, sort order, dotted-path expansion,
changed/unchanged/missing-previous flagging, and inline-vs-expanded arrays.

* feat(web): add Screenshot panel with bounds and tap overlays

Center column of run-detail page. Renders the device screenshot
scaled to fit, with an SVG overlay drawing resolvedBounds as a
violation-colored rect, tapPoint as a contrast ring, and swipes
as an arrow. Falls back to a placeholder when src is missing or
the image fails to load.

* fix(web): guard scrollIntoView call for jsdom compatibility

* test(web): cover Screenshot panel rendering and overlays

* test(web): cover ActionList rendering, selection, keyboard, and markers

* feat(web): add ExceptionsPanel component

* test(web): add ExceptionsPanel tests

* feat(web): add Timeline panel with property swimlanes

Renders SVG swimlanes per property with violated/pending/holds cells,
action-marker dots, click-to-seek, and a selected-step highlight bar.

* test(web): cover Timeline empty state, cells, status, click, highlight

* feat(web): add ResidualNode recursive AST renderer

* test(web): cover ResidualNode operators, predicate, and error chip

* feat(web): add ViolationsPanel with status badges and jump button

* test(web): cover ViolationsPanel rows, status grouping, and jump button

* test(web): register testing-library cleanup globally

All six panel test files added local afterEach(cleanup); centralize it in
the shared setup so future tests inherit DOM isolation by default.

* feat(web): hooks for url/keyboard/theme/sse

* feat(web): wire all panels into run-detail with phone-dominant grid

ActionList left, Screenshot center, Snapshots/Properties/Exceptions
stacked right, Timeline bottom. URL-synced step index (useStep), keyboard
shortcuts (j/k/arrows/g/G/.), light+dark theme toggle stored in
localStorage, SSE auto-refresh on the run index.

* test(web): add three reference run fixtures (clean, violation, exception)

* build: web targets in Makefile + bun in CI; docs(inspect)

- Makefile: web-build/web-dev/inspect-dev/test-web targets; uatu and
  install now depend on web-build so the binary embeds the latest SPA.
- ci.yml: setup-bun + cache; existing make test now runs web typecheck +
  vitest as part of the full suite.
- docs/manual/inspect.md: panel reference, keyboard shortcuts, URLs.
- docs/manual/cli.md: document uatu inspect.
- README: link to inspect docs.

* feat(runner): capture a screenshot per step

The driver already exposes Screenshot(ctx), but the runner never called
it. Each step now writes <run>/screenshots/step-NNNNN.png right after
the trace line, using the same failure-is-a-warning posture as other
best-effort observability hooks. Makes the inspect UI's center panel
actually useful.

* feat(inspect): include action_label in StepSummary

Tap/InputText/Swipe/PressKey/Wait each get a short human-readable
label (selector, quoted text, swipe direction, key name, duration) so
the action list panel can render readable rows instead of just 'Tap'
with no target.

* test(inspect): accept either #app or #root in SPA shell fallback

* feat(web): render action_label and screen in ActionList rows

Step rows now show 'Tap id:save', 'InputText "alice"', 'Swipe up',
'PressKey back', etc. Steps with no action fall back to
'observe @ <screen>' so the list reads as a flow instead of a wall
of '--' placeholders.

* feat(sidecar): implement screencap for android driver backend

Was stubbed to return an empty byte array, which made the runner's
per-step screenshot capture a no-op. Shell out to 'adb exec-out
screencap -p' and stream the PNG bytes back. Width/height stay zero
because the PNG header carries them; the Go side can parse if needed.

* feat(proto): add Metrics RPC for per-step CPU and memory capture

* feat(driver): Metrics(bundleID) returns cpu_percent + heap/total bytes

* feat(sidecar): implement Metrics RPC via adb top + /proc/<pid>/status

* feat(runner): capture metrics + before/after screenshots per step

Each step now writes step-NNNNN.png (before applyAction) and
step-NNNNN-after.png (after the action + wait-for-idle). The runner
samples Driver.Metrics(bundleID) before writing the trace line and
stamps Step.Metrics with cpu_percent, heap_bytes, total_memory_bytes
so the inspect UI can chart CPU and heap over the run.

* fix(runner,sidecar): measure CPU across step via /proc stat delta

'top -d 0.3 -n 2' measures CPU in a 300ms window that coincides with
the SDK-paused app, always reporting 0%. Switch to reading
/proc/<pid>/stat utime+stime and computing the delta between successive
calls; the natural step cadence gives a 2-5s measurement window that
captures the action response and render cycle. Also moved the sample
to before snapshotStep so the delta starts before the SDK pause.

* feat(web): add Metrics type for per-step cpu and memory

* refactor(web): replace --accent-change with --accent-positive token

* refactor(web): recolor chip-progress as neutral outlined chip

* refactor(web): use neutral border for changed snapshot rows

* feat(web): add MetricsChart panel with HEAP and CPU lanes

SVG-based time-series chart rendering heap bytes and CPU percent per
step across two stacked lanes, with a shared step axis below. Lines are
monochrome; a vertical highlight marks the selected step; per-step hit
rects make any click seek to that step.

* feat(web): revamp ActionList with tag targets, elapsed time, and expandable rows

Render selector-based Tap actions as <tag/> markup, show zero-padded MM:SS.mmm
elapsed time per row, and expand the active row with Position/Content sub-rows
when a full Step is available. Adds formatActionRow/formatElapsed helpers and
covers both with unit tests.

* fix(runner): stop copying Tap selector into action.text

The 'Content' inspect row should show the user-supplied text for
InputText actions and stay empty for Taps. Previously the runner copied
action.On into traceAction.Text for both, so the inspect UI showed the
selector as the tap's 'Content'.

* fix(web): use text-muted for swipe arrow after accent-change removal

* fix(web): snapshot values truncate with ellipsis + title tooltip

Long JSON values were breaking one character per line due to
overflow-wrap:anywhere in a narrow column. Switch to single-line ellipsis
with the full value exposed via the title attribute on hover.

* feat(web): state-before/after columns + metrics chart at bottom

RunDetail now renders a four-column grid:
  actions | state-before | state-after | side (exceptions + timeline)
with MetricsChart spanning the bottom row. Each state column shows its
own screenshot (step-NNNNN.png vs step-NNNNN-after.png), snapshot table,
and violations panel. ActionList now receives runStartMillis and the
selected Step so the active row can expand Position/Content sub-rows.

* fix(web): skip zero-value ticks + add exception markers to metrics

HEAP '0B' and CPU '100%' labels overlapped at the lane boundary. Drop
the bottom-of-range tick on both lanes (baseline is implied) and widen
LANE_GAP so the remaining labels have breathing room. Accept an
exceptionStepIndices prop and draw a dashed red vertical line at each
to surface exception spikes directly on the CPU/heap chart.

* fix(web): let action body column shrink below its content

Required minmax(0, 1fr) so the row grid honours the column's min-size of
0 instead of the implicit 'auto', preventing the action-list from
overflowing its parent when the target string is long.

* feat(web): bigger state screenshots + single properties row

Collapse snapshots into a summary chip ('SNAPSHOTS · N violations') so
the screenshot fills its state card. Deduplicate ViolationsPanel —
show it once in a new full-width 'properties' row between the state
cards and the timeline. Drop the right sidebar; exceptions now surface
as dashed markers on the metrics chart with the ExceptionsPanel only
rendering when there are actual exceptions to report.

* feat(web): add minimal Tabs component

Monochrome tab strip with underline-on-active. Used by state-before
and state-after cards to swap between Screenshot, Snapshots, Properties.
Pane scrolls internally so the outer grid stays fixed-height.

* feat(web): fold timeline into MetricsChart as STEPS lane

Adds a thin per-step status row above HEAP showing violated (red),
pending (dim gray) or holds (green-tinted). Extends highlight +
exception markers to span the status lane. Frees a whole row in the
detail grid so the page can fit in 100vh.

* refactor(web): tabbed state cards, drop standalone Timeline panel

State-before/after now use Tabs (Screenshot / Snapshots / Properties,
default Screenshot). Removes the dedicated timeline row; status lane
lives on the metrics chart. Banner is gone from the shell.

* feat(web): lock app shell to 100vh with no page scroll

html/body/#root fill the viewport, body gets overflow:hidden, and the
detail grid uses minmax(0, 1fr) rows so inner panels own their scroll.
Tightens toolbar + panel padding for a denser feel.

* feat(web): arrow-key nav + badges on Tabs (WAI-ARIA tablist)

Roving tabindex, ArrowLeft/Right/Up/Down/Home/End navigation, explicit
aria-selected/aria-controls/id wiring, and support for an optional
badge inside each tab (used for violation counts).

* feat(web): ViolationsPanel supports violationsOnly filter

* feat(web): ActionList arrow-key nav + listbox semantics + smaller font

Promote the list to role=listbox with role=option rows; roving tabindex
lets ArrowUp/Down (and Home/End) seek between steps with focus. Font
size dropped to 11px and padding tightened so long selector-tag labels
fit in the 340px actions column.

* fix(web): useKeyboardNav yields arrow keys to tablist/listbox targets

Previously pressing ArrowRight on a focused tab switched tabs AND
advanced the step. Skip arrow handling when the event target is inside
an element with an arrow-owning ARIA role.

* feat(web): fourth 'Violations' tab + wider actions + shorter metrics

Adds a Violations tab to each state card showing only violated properties
(with count badge on the tab label when > 0). Actions column widened
from 280px to 340px, bottom metrics strip trimmed from 220px to 140px
with tighter lane heights, so the whole page still fits in 100vh with
no scrollbar.

* feat(web): compact RunDetail layout using 1px borders instead of panel padding

* refactor(inspect): simplify MetricsChart to HEAP+CPU with time axis

Drop the STEPS status lane and per-sample circle markers, switch the
x-axis from step indices to mm:ss clock time, trim y-axis ticks to
min/max with compact units, rotate lane labels into the left gutter,
and replace the thin playhead line with a wider dotted red band.
Traces stay grayscale; red appears only on the playhead pattern.

* fix(web): RunList rows no longer stretch to fill viewport height

Tables inherited flex: 1 1 auto from .app-main > * and distributed extra
vertical space across rows. Override with flex: 0 0 auto + align-self.

* misc changes

* fix(web): hoist useState above early return in MetricsChart

Calling useState after an unconditional early return violates React's
Rules of Hooks: the empty-samples branch renders 0 hooks while the
populated branch calls 1. On the initial null->loaded transition of
history the hook count changes and React throws.

* fix(web): subscribe to named SSE event instead of 'message'

Server emits 'event: runs.changed' frames; the WHATWG EventSource spec
dispatches those as events of type 'runs.changed', not 'message'. The
listener registered on 'message' was never fired, so RunList never
auto-refreshed on run create/finish/delete.

* fix(inspect): unsubscribe SSE clients on disconnect

Watcher.Subscribe appended to a slice with no matching removal path,
so every closed EventSource connection leaked its channel. Over a
long-running server the slice grew unbounded and every fs event paid
O(N) iterating dead channels. Add Unsubscribe + defer it in
handleEvents.

Unsubscribe does not close the channel: broadcast snapshots the
slice without holding the mutex, so a concurrent close would race
with its non-blocking send.

* fix(trace): rename resolvedBounds/tapPoint to snake_case

Every other json tag in the trace schema (from_x, duration_millis,
bundle_sha256, etc.) uses snake_case. The two new Action fields
introduced with the inspect UI broke that pattern. Rename them
before the format ships to external consumers.

* chore(web): drop vitest and remove UI tests from CI

No UI tests wanted in web. Removes vitest, jsdom, testing-library
devDeps and the vitest.setup.ts + vite.config.ts test block.
Makefile test-web becomes web-typecheck (typecheck only).

Fixes CI failure where `vitest run` exits 1 with no test files.

* chore(make): dedupe sidecar embed and drop recursive make

Make $(SIDECAR_JAR) the real recipe and $(SIDECAR_EMBED) a file
target, so uatu/install/inspect-dev share one copy step and
sidecar/release-cli just depend on the jar instead of re-invoking make.
2026-04-21 11:32:17 +07:00
pj a2e96af1af WIP: rename sample app to Folio (#20)
* refactor: rename examples/sample-app to examples/folio

Directory-level rename and path references in Go tests, bundle-check,
top-level README, and getting-started docs. Package declarations,
Gradle config, iOS bundle IDs, and class names follow in later commits.

* refactor(folio): rename Kotlin package dev.uatu.sample to app.folio

Moves source dirs and sqldelight schema from dev/uatu/sample to
app/folio, updates package declarations and imports, and switches
Android namespace/applicationId, iOS binaryOption bundleId, and
sqldelight database packageName to the new identifier.

* refactor(folio): rename SampleApplication to FolioApplication

Android manifest now points at .FolioApplication with label 'Folio'
instead of 'Uatu Sample'.

* refactor(folio): set iOS bundle id and display name to Folio

bundleIdPrefix + PRODUCT_BUNDLE_IDENTIFIER -> app.folio.
CFBundleName + CFBundleDisplayName -> 'Folio'.

* refactor(folio): update demo email to [email protected]

* refactor(folio): point justfile at app.folio bundle id

Updates xcrun simctl launch target, uatu test --bundle-id, and the
build/uninstall comments to reference folio instead of sample.

* test: update fixture package ids to app.folio

Sidecar activity-resolver test and verifier spec-integration XML
fixtures referenced the old dev.uatu.sample Android package. Updates
them to match the folio app's real package id so the tests stay
representative of what the CLI sees on-device.

* test(verifier): rename SampleApp identifiers to Folio

Renames TestSampleAppSpec* functions, bundleSampleAppSpec helper, and
sampleAppHierarchyXML const (now loginHierarchyXML for consistency with
the other per-screen fixtures). Updates trailing sample-app mentions in
comments and assertion messages.

* refactor(folio): rename Gradle/npm/wasm project identifiers to folio

settings.gradle.kts rootProject.name, package.json + package-lock.json
name, and the WasmJS index.html <title> all still read 'uatu-sample' /
'Uatu Sample'. Realigns them with the Folio brand.

* docs(folio): rewrite README title + getting-started bundle id

examples/folio/README.md is now titled 'Folio' with the Kotlin source
paths corrected to app/folio. Getting-started example uses --bundle-id
app.folio. Harness launch message is now generic ('app under test')
since uatu-sample-harness is not specific to folio.

* chore(folio): drop trailing 'sample' reference in gradle.properties

* refactor(folio): rename LoginPage composable to LoginScreen

Align with KMP/Android industry convention (NowInAndroid, Cash App,
JetBrains samples use Screen, not Page).

* refactor(folio): rename HomePage composable to HomeScreen

* refactor(folio): rename AddAccountPage composable to AddAccountScreen

* refactor(folio): rename LedgerPage composable to LedgerScreen

* refactor(folio): rename AddTransactionPage composable to AddTransactionScreen

* refactor(folio): split Models.kt into app.folio.data package

Account, Transaction (with TxnType), and Session move into their own
files under app.folio.data, matching NowInAndroid-style per-type
organization.

* refactor(folio): move data layer into app.folio.data package

Repository, LedgerStore (expect + interface), SqlLedgerStore,
WebLedgerStore, DriverFactory (expect + actuals), AndroidLedgerContext,
and Snapshot move into app.folio.data. Update all consumer imports.

* refactor(folio): move Navigation into app.folio.navigation package

Split the former Navigation.kt into Route.kt (sealed interface) and
Navigator.kt (singleton). Update consumer imports across screens,
App.kt, and FolioApplication.

* refactor(folio): move Platform and Format into app.folio.platform

Both files carry expect declarations (Platform object, formatDate);
grouping them into a dedicated platform package makes the KMP seam
obvious and mirrors the structure used by JetBrains samples.

* refactor(folio): move login into feature/auth package

Create app.folio.feature.auth with LoginScreen + LoginUiState. Inline
the former Auth.kt (DEMO_EMAIL, DEMO_PASSWORD, checkCredentials) into
LoginScreen since it is the sole caller.

* refactor(folio): move HomeScreen into feature/home package

* refactor(folio): move account creation into feature/account package

AddAccountScreen gets its own AddAccountUiState colocated with the
screen, replacing the shared UiState.addAccountError.

* refactor(folio): move ledger screens into feature/ledger package

LedgerScreen and AddTransactionScreen move into app.folio.feature.ledger
with AddTransactionUiState (txnError, txnFormType) colocated. The
former catch-all UiState.kt is removed now that each screen owns its
state alongside its UI.

* refactor(folio): split Theme.kt; move theme and icons to subpackages

Theme split into Theme.kt (tokens, layout dims, LedgerTheme) and
Type.kt (typography) under app.folio.ui.theme. Icons moves to
app.folio.ui.icon. Update every consumer's imports to match.

* refactor(folio): split ui components into per-file under ui/component

Former Widgets.kt and Components.kt become 10 focused files: AppButton,
Card, EmptyState, ErrorText, FieldLabel, Header, IconButton (w/
BackButton), Screen, Segmented, TextInput. Matches NowInAndroid style
of one composable per file in a designsystem/component package.

* chore(folio): consolidate uatu testing files under uatu/ folder

Move spec.ts, package.json, package-lock.json into examples/folio/uatu
so all uatu-specific testing artifacts live in one place. runs/ and
node_modules/ follow the same convention (both remain gitignored).
Update justfile, README, and the two Go consumers (bundle-check tool +
verifier/trace tests) that referenced the old path.

* refactor(trace): drop folio path in writer test

Round-trip only needs a non-empty string; neutralize to keep the
library free of folio references.

* refactor(sidecar): neutralize ResolveActivity test fixtures

Swap app.folio for com.example.app in the fixture strings so the
sidecar tests don't reference the example app by name.

* refactor(bundle-check): take spec path as argument

Previously the tool hardcoded examples/folio/uatu/spec.ts. Accept a
positional spec path instead so the tool works for any example and
leaves no folio reference in the library surface.

* test(verifier): add neutral integration spec and hierarchy fixtures

Adds testdata/integration_spec.ts with two routes ("list", "form"),
an InputText on text_field, a Tap on primary/secondary_action, a
safety property (itemCountNonNegative), and a liveness property
(submitEventually). Adds hierarchies_test.go with matching XML
fixtures. Constants intentionally go in a _test.go at package root
rather than testdata/hierarchies.go because go skips .go files
under testdata/.

* refactor(verifier): replace folio integration tests with neutral ones

Renames bundleFolioSpec -> bundleIntegrationSpec and the three Test*
entry points to TestIntegrationSpec*. Uses the synthetic spec and
hierarchies added in the previous commit so the library's test suite
no longer references examples/folio at all.

Folio-specific coverage remains covered by examples/folio/justfile's
'just test' which exercises the real spec on device/emulator.

* chore: remove cmd/uatu-sample-harness

Not referenced by Makefile, docs, CI, or any script. Duplicates the
adb reverse helpers already in cmd/uatu/test_run.go, and its name
implies ownership by the sample app which violates the library/
example decoupling. If a bare-protocol debugging tool is later
needed it belongs inside cmd/uatu/.

* docs(folio): drop Layout section and KMP layout paragraph; fix AVD override syntax

The directory-tree Layout section rots faster than the code and
duplicates what ls shows for free. The expect/actual paragraph in
Stack was the same kind of filler. The README also claimed 'just
AVD=Pixel_7 test' but the justfile reads AVD as an env var via
env_var_or_default, so the correct invocation is 'AVD=Pixel_7
just test'.
2026-04-20 16:04:20 +07:00
pj d0578dbaaa fix(runner): warn on malformed screen snapshot (#13)
* fix(runner): warn on malformed screen snapshot

screenFromSnapshot swallowed json.Unmarshal errors, so a non-string
screen value silently became "" in the step log and trace while the
verifier still saw the raw JSON. Return the error and warn at the
call site, matching the hierarchy warning pattern.

* docs: clarify --avd is optional for uatu test

The CLI accepts --avd as an empty-string default (cmd/uatu/main.go:49)
and only requires it when no device is connected and multiple AVDs
exist (cmd/uatu/android_env.go:63). Docs and examples that showed it
as required or always-passed were misleading.
2026-04-18 17:14:31 +07:00
pj 6b1c17328c fix(sample-app): make it standalone, no repo_root assumptions (#8)
* fix(cli): fall back to node_modules for @uatu/spec resolution

Drop the hard failure when the uatu source tree is not reachable from
the spec file. Users integrating uatu in their own app have @uatu/spec
installed via npm; esbuild now resolves it from node_modules.

* build(gradle): drop :sample-app include from root settings

The sample now has its own Gradle project in examples/sample-app/android.

* build(sample-app): vendor gradle wrapper

Users running the sample build the APK via ./gradlew from inside the
sample-app's own android/ directory, no repo-root wrapper required.

* build(sample-app): make gradle project standalone

Drop the project(':sdk-android') dependency in favor of the Maven
Central coordinate io.github.priyanshujain:sdk-android. The sample now
owns its settings.gradle.kts and gradle.properties, so it builds
without any pieces of the uatu source tree.

* chore(sample-app): declare @uatu/spec npm dependency

Mirrors what a downstream user would put in their own package.json.
Uses file: for pre-release development; becomes a normal semver pin
once @uatu/spec ships to npm.

* docs(sample-app): rewrite justfile and add README

Justfile drops repo_root; all recipes run against the local gradle
wrapper and uatu from PATH. README is scoped to what a user needs to
run the sample against their own device.

* docs(manual): update sample install steps to standalone layout

./gradlew :sample-app:installDebug no longer exists; the sample owns
its own wrapper under android/.

* feat(sdk): log when Uatu.start succeeds

Silent SDK start makes the "SDK didn't connect" failure mode
impossible to debug. One INFO line at start time is enough.

* fix(sidecar): launch via am start -W instead of monkey

monkey -p <pkg> -c LAUNCHER 1 is unreliable on API 36+: it reports no
error but silently fails to start the activity, so the SDK never runs
and the CLI times out on the SDK-accept handshake.

Resolve the launcher activity via `cmd package resolve-activity
--brief` and launch it with `am start -W -n`. -W makes the call block
until the activity is up, which also makes the subsequent SDK
accept timing deterministic.

* feat(cli): auto-resolve Android device; boot AVD if none connected

--avd becomes optional. Resolution order:
 - use any already-connected adb device;
 - else if --avd names an existing AVD, boot it and wait for boot;
 - else if --avd is missing or names no AVD, error with a clear message.

Falls back to $ANDROID_HOME/emulator/emulator when the binary is not on
PATH, so a standard Android SDK install works without extra shell setup.

* docs(sample-app): AVD is optional; document both paths

just test runs against any connected device. If none, pass AVD=<name>
to have uatu boot the emulator for you.

* feat(cli): auto-discover Android SDK; auto-pick the lone AVD

adb and emulator are looked up via PATH, then $ANDROID_HOME,
$ANDROID_SDK_ROOT, ~/Library/Android/sdk, ~/Android/Sdk, and the
Homebrew cask path. The discovered platform-tools directory is
prepended to the sidecar's PATH so its adb subprocess calls work too.

When --avd isn't passed and no device is connected, the CLI picks the
sole local AVD and boots it. Multiple AVDs → error listing them.

* build(make): add `make install` that go-installs uatu onto PATH

Puts `uatu` into $GOBIN (or $GOPATH/bin) so the sample and any local
dev flow can call it without PATH= prefixes.

* docs(sample-app): zero-config just test; dotenv-load for persistence

Drop the expectation that users prefix commands with PATH=, ANDROID_HOME=,
or AVD=. `just test` now works as-is; optional knobs can be pinned in a
.env file alongside the justfile.

* fix(sample-app): auto-detect ANDROID_HOME for Gradle tasks

The Go CLI finds the SDK itself, but AGP still needs ANDROID_HOME to
resolve `sdk.dir`. The justfile now resolves it from env or canonical
install paths before invoking ./gradlew, so `just install` works out
of the box on a standard Android SDK setup.

* gitignore runs directory for sample app
2026-04-18 15:31:47 +07:00
pj e62319e916 docs: pandoc-based site and v0.1.0 groundwork (#5)
* chore(prose): remove em-dashes from config files

* chore(prose): remove em-dashes from android sdk config

* docs(spec-api): remove em-dash from README

* fix(doctor): reword sidecar-jar error without em-dash

* test(sidecar): reword assertion message without em-dash

* docs: add CLAUDE.md with project conventions

* build: add docs target for pandoc site

* docs(site): add pandoc template and stylesheet

* docs(site): add pandoc build script

* docs(site): add landing pages

* docs(manual): add getting-started

* docs(manual): add writing-specs

* docs(manual): add runs

* docs(manual): add cli reference

* docs(dev): add design principles

* docs(dev): add architecture

* ci: deploy docs site to github pages

* docs: rewrite README as entry point to docs site
2026-04-18 14:00:57 +07:00