Commit Graph
228 Commits
Author SHA1 Message Date
pj 404573566d feat(verifier): expose extractor names and rebuilt property formulas
an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.
2026-08-16 17:44:52 +05:30
pj 3cdc0a1a2a feat(runner): bound every device call and record the actions that never reached the app
observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.
2026-08-16 17:44:40 +05:30
pj 1ea5ab0c59 feat(trace): version each step and record its logs, exceptions and navigations
a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.
2026-08-16 17:44:34 +05:30
pj 944d1e9ad9 feat(chrome): read the page's exceptions and navigations, and hold the picker state across them
a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.
2026-08-16 17:44:29 +05:30
pj 45dd52c5d9 fix(chrome): scroll a gesture point into view and dispatch trusted input
getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.
2026-08-16 17:44:23 +05:30
pj 05c8523762 fix(chrome): emit every markup attribute and read checked and selected off the property
the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.
2026-08-16 17:44:04 +05:30
pj de654334b0 feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag
a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.
2026-08-16 17:43:20 +05:30
pj ca667b1b77 fix(selectors): resolve text to the innermost match and scan the root in both forms
an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.
2026-08-16 17:43:10 +05:30
pj a245194120 fix(sidecar): map the driver's refusals onto the gesture errors
OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.
2026-08-16 17:42:00 +05:30
pj 433f786080 fix(sidecar): stop dropping gestures, selectors and keys in silence
a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.
2026-08-16 17:41:55 +05:30
pj 1b01bdbaf1 feat(ios): derive scrollable from the snapshot's tree depth
the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.
2026-08-16 17:41:49 +05:30
pj 0ab5c305e3 fix(ios): refuse a gesture the screen has no surface under
the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.
2026-08-16 17:41:38 +05:30
pj da3cc5f68c feat(driver): add escape to the pressKey surface
escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.
2026-08-16 17:41:33 +05:30
pj 4b7a3878ac feat(driver): declare undelivered-action errors and three optional capabilities
ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.
2026-08-16 17:41:11 +05:30
pj 2755643195 docs(claude): add delegation and record-keeping sections
delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.
2026-08-16 17:41:06 +05:30
pj 7d6ed32474 docs(spec-language): say what a trace records for an element-valued extractor 2026-08-15 18:59:27 +05:30
pj ff4666c126 test(runner): element-valued extractors reach the trace 2026-08-15 18:59:03 +05:30
pj 91ca574d67 test(verifier): an unrecordable extractor value is reported, not dropped 2026-08-15 18:59:03 +05:30
pj 9795c9ade3 fix(verifier): record an element-valued extractor instead of dropping it
An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.
2026-08-15 18:59:00 +05:30
pj bf2da973de fix(runner): an ambiguous name loses to the coordinates it was built from
Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 14:02:01 +05:30
pj fd7bf17256 test(runner): sibling taps reach the driver at their own coordinates
Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:59:30 +05:30
pj 03af6abe72 fix(web-runtime): hold the V8 host to the same naming gate
elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:59:17 +05:30
pj 57977cde0a fix(verifier): name an element only when the selector names it alone
ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:59:17 +05:30
pj 6d03c5788a fix(chrome): emit data-testid so both resolvers name the same element
The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:59:17 +05:30
pj 231a74b876 fix(runner): a selector the tree cannot resolve is not a focus failure
otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:59:03 +05:30
pj c8e6c6634c docs: every target runs on this machine, so start one rather than skip it
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:48:47 +05:30
pj 7066b23a22 fix(web-runtime): answer clickable for an element reached through ax
The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:45:30 +05:30
pj 9b9c7cf6e2 fix(chrome): name a web field by its hint, not its CSS class
visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:45:30 +05:30
pj ebed84afc3 feat(campaign): record the label source in the manifest
A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:32:05 +05:30
pj ff2de344e9 feat(campaign): make the label source a cell dimension
A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:32:05 +05:30
pj d987526e47 Merge origin/master into llm-recording-and-analysis
Brings in #73, which landed web fact-parity work overlapping this branch:
selector-tagged ax handles, aria-disabled in `enabled`, shadow-DOM traversal
and a `scrollable`/`editable` dump, plus a third folio property and two new
testTags on HomeScreen.

Six files conflicted. The option-carrying ones took the union of both sides'
fields, so `--label-source` and `--exit-on-violation` both reach the pipeline.
In web-runtime.ts both sides changed how an element is described: master gave
`elementHandle` a selector to tag handles with, this branch gave it raw
attribute names and a field hint. Both survive, and `enabled` now answers
through master's isEnabled while `editable` stays.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-15 13:28:11 +05:30
pj b02e86b2e3 ci: dispatch workflows for folio and the replay ui (#73)
* feat(runner): stop the step loop at the first violation on request

* feat(testrun): report violations as a typed error under exit-on-violation

* feat(cli): add --exit-on-violation and exit 2 when it fires

* docs(cli): document --exit-on-violation, --max-steps, and exit codes

* fix(web): enumerate and query across shadow roots in both producers

* test(chrome): compare both producers on a shadow-dom parity page

* test(browser): drive a canvas-under-shadow-root fixture end to end

* fix(web): select the focused field inside a shadow root before typing

* fix(web): report the pathname as the screen when there is no hash route

* fix(web): settle on dom quiescence instead of returning at body ready

* feat(replay-ui): add data-testid hooks the dogfood spec drives

* feat(replay-ui): add the dogfood spec sanderling runs against the replay ui

* fix(replay-ui): scope the screenshot property to the named state panel

* chore(make): add per-platform sanderling build targets

* ci: add dispatch workflows for folio and the replay ui

* docs: describe the dispatch workflows and how to read a failure

* ci(folio): give the ios leg its jdk, android sdk, just, and a clean app start

* refactor(web): use max for the settle budget

* ci: pin calibrated seeds, skip the flaky ios reinstall, bound every job

* ci: authenticate and pin the buf setup step

the anonymous release download hit the shared runner ip rate limit and
failed the job with 'socket hang up' after three retries.

* docs: record that canvas apps need a dom proxy to be text-fuzzable

* fix(ios): bound lifecycle rpcs and claim the target device

a launch the simulator rejects sent the xctest session into a recovery
chain that answered minutes late or never, and the rpc had no deadline,
so the run hung with no trace and no error. also take a per-udid flock:
a second run's reinstall lands under the first's live automation session
and wedges it.

* docs(ci): correct the ios hang wording and note the device lock

* refactor(verifier): derive the lastAction shape from one field list

both hosts must show a spec the same lastAction. one ordered list now
feeds the goja object and the json the web host installs, so they
cannot drift.

* fix(web): install lastAction in the page before extractors read it

state.lastAction was hardcoded null on web, so every property reading it
was silently vacuous: a correct property passed without ever firing.

* fix(web): carry element identity on actions and fix findAll on paths

an action's target was coordinates only, so a property matching on which
element was acted upon could never fire. ax.findAll([a,b]) also returned
nothing on web.

* fix(chrome): wait out a route transition before sampling facts

the tree stays byte-identical and quiet across a cross-fade, so both the
quiet timer and the unchanged-tree escape called it settled mid-flight
and extractors read two screens at once.

* fix: bound the pre-run app launch

launch happens before the runner starts, so --duration never covered it
and a wedged driver hung with no trace and no error.

* fix(folio): read balances from merged cards and treat unreadable as unknown

compose for web merges the whole accountcard subtree, so the balance
child never exists there and every card parsed as 0. the property then
compared 0 to 0 and fired on any submit, which is a false positive
generator. unknown is now null and null is vacuously true.

* test(folio): cover merged-card parsing and unknown balances

* ci(folio): make web an expect-the-bug leg

the web runtime can observe the double submit now, so the health gate
understates it. seed 1 finds it at step 109, 3 runs out of 3.

* docs(ci): explain why a submit tap landing on home is the bug

* fix(ios): read a StaticText's label as its text

AXValue was the only source for text, but a StaticText carries its
string in AXLabel, so nothing on screen had .text on ios: a spec reading
it saw everything on android and nothing here.

* docs(ci): correct the calibrated step ranges

* fix(folio): stop convicting on arithmetic float64 cannot hold

past 2^53 cents the gap between representable values is 128, so a real
1600-cent move reads back as something else and the equality is false
for a healthy submit as readily as a double one. also match parseCents:
a sign or an oversized amount is rejected, not read as an amount.

* test(folio): pin the safe-integer guard and its boundary

* docs: stop teaching the zero-default that caused a false alarm

* docs: write down the silent-vacuity failure modes

* feat(folio): tag the home total and the card transaction count

the total was the only untagged node on the screen, so the spec had to
sum cards and a clipped card broke the sum.

* fix(folio): read the app's own total and refuse contaminated windows

summing cards went null when one was clipped, and the null poisoned the
carrier for the rest of the run. the balance window also spanned every
transaction since the last home visit, so the property convicted on
deltas it could not attribute: the old web witness was 3.16x the typed
amount, not 2x.

* test(folio): pin the window rules and the count invariant

* fix(folio): never read a frame that shows two screens

android dumps a cross-fade with both screens in the tree. the route said
add-transaction while an unscoped find said home, so the oracle took a
half-rendered total as fresh and convicted on a tap that committed
nothing. one function now decides the route and returns null when the
frame is ambiguous.

* test(folio): cover transition frames, card readings and creation

* fix(folio): only disambiguate counts that came from merged text

the equal-length digit rule exists because web merges the card and an
account named -1 makes '12' ambiguous. a dedicated count node has
nothing to disambiguate, so applying it there threw away real evidence.

* ci(folio): pin the recalibrated seeds and drop android to a health gate

web 3 and ios 7 convict 3 runs out of 3 with an exactly 2x witness.
android convicts 2 in 5 because the same seed does not walk the same
trajectory there, so it proves the app runs instead.

* docs(ci): describe the two properties and why android cannot convict

* fix(android): wait out a route cross-fade before snapshotting

the dump could hold two screens at once, and the runner refuses to act
on such a tree, so a quarter of android steps applied no action and the
count varied per run: the same seed never walked the same trajectory.
the ios companion and the chrome driver already do this.

* ci(folio): let the android leg run far enough to see its conviction

* docs: only the repo owner merges

* ci(folio): a thrown predicate is not a conviction

exit 2 means the run recorded a violation, and a predicate that throws
is recorded as one too. so was newAccountBalanceIsZero, an unrelated
property in the same spec. the gate read the exit code and went green
with detection dead.

* ci: install idb-companion from its tap and stop interpolating inputs

idb-companion is not in homebrew-core, so the ios leg died before it
built anything. replay-ui expanded dispatch inputs into the shell.

* docs: correct the snippets and numbers that drifted from the code

* test(sidecar): pin that a slow read counts toward the stability streak

* fix(web): read the page's extractors only on steps that count

the page advances the spec's carriers when it evaluates, but the runner
applied the result only on non-transitional steps. a discarded step
moved the window forward anyway, so the next accepted pair bracketed two
transactions while counting one submit, and convicted a healthy app.
extractor errors now fail the run instead of leaving goja's values in
current against v8's in previous.

* fix(chrome): anchor the transition deadline when the dom goes quiet

it was anchored at script start, so a page that churned past the window
reached the check already expired and returned mid cross-fade. the
driver now publishes the idle timeout it needs, since the caller's 1s
could never spend the 800ms window.

* fix(web): fail on a partial extractor override

same mixed-producer hazard as the install error: some extractors hold
the page's value and the rest hold goja's, and a property comparing
across that split fires on a healthy app.

* docs: six of seven, the seventh is the stock property

* fix(folio): drop a name two cards answer to

homeTxnCountsOf keyed on the account name and let the last card win, so
two accounts the fuzzer named the same collapsed into one entry. a
reading that saw one Travel card and a later one that saw both then
subtracted two different accounts' counts, and
submitCommitsOneTransactionPerAction convicted a healthy app of
double-submitting. it is a gated property in folio-run.sh, so that reads
as "found the submit bug" over a card scrolling into view.

same rule createdAccountHasNonZeroBalance already applies: a name
nothing can attribute is no evidence. counted over every card, since an
unreadable twin spoils the identity too.

* perf(folio): read each frame once

every extractor asked routeOf, and routeOf does five ax.find calls. on
web each find walks the document and every shadow root beneath it, so
the spec cost 110 tree walks a step; homeCards was parsed four times
over. now 5 and once.

keyed on the identity of the state object because both hosts build a new
one per step and hand that one object to every getter, so it cannot
outlive its frame. holding the reference is what keeps that true rather
than likely.

* fix(web): keep an undefined reading's index through JSON

json has no undefined, so an extractor whose getter returned one had its
whole index dropped by JSON.stringify. that index then kept goja's
dump-derived value while its neighbours held the page's, and a property
comparing previous to current across the split fires on a healthy app.
folio has nine on(route, tag) extractors, so this was most extractors on
most steps.

each reading is wrapped in a {value} envelope: the drop now happens
inside the entry, and an absent value means the getter returned
undefined, which is what the goja host records for the same getter. a
json null would instead claim it returned null and x.current ===
undefined would answer differently on the two hosts.

* feat(verifier): report the registered extractor count

the web path needs it to check the page sent one reading per extractor.

* fix(runner): fail when the page reports fewer readings than extractors

the comment here already claimed a partial override was fatal. it was
not: the skipped check only catches indices outside the extractor list,
so a page reporting values for some extractors and not others left the
rest holding goja's reading of the dump with nothing said.

* test(browser): drive an undefined reading through the whole web path

four layers carry it: the page's envelope, the driver's unwrap, the
runner's count check and the verifier's decode. each has a unit test and
only a run proves they compose. goes red both ways, decoding an absent
value as null and dropping the envelope.

* fix(web): offer the aria roles a user activates

only role=button was in the tappable set, so link, checkbox, radio,
switch, tab, option, the menuitems and treeitem were invisible to the
enumeration however plain the control looked. the replay ui builds its
step rows as <li role="option">, and the spec dogfooding it had to
hand-write an action to reach them because no default verb could see a
single row.

both producers build the set from the same role list, since the parity
test compares them element by element.

* test(browser): tap a role-based control end to end

every control on the page is an <li role="option">, the shape the
replay ui gives its step rows, and the spec carries no action of its
own: the property firing is the evidence the default enumeration offered
a tap on one.

* fix(web): read aria-disabled as disabled

the enabled fact came off the disabled property, which only real form
controls have. it reads undefined on the role-based controls the
tappable set now covers, so every one of them looked enabled however
plainly it was marked otherwise, and the fuzzer would spend actions on
inert ones.

both producers answer the same two ways, and the parity fixture carries
a disabled row so the comparison covers it: reverting one side alone
names the element and the fact.

* docs(replay-ui): the enumeration reaches step rows now

the comment said role="option" is not in the tappable selector set,
which stopped being true a few commits ago. selectAStep stays, for the
reason the tab weight below it stays: one row among the page's clickable
elements is a thin chance, and both step-facing properties go vacuous on
a run that never selects one.

* test(runner): bound the last-action test by steps, not wall clock

100ms of wall clock against an assertion that two steps ran fatals under
load with "the web path never installed it", which reads as a
regression. every sibling test in the package uses a long duration and
MaxSteps.

* ci: run the kotlin tests in make test

RouteTransitionTest and the stability poll cover the android settle and
nothing in ci ran them. :sidecar:test needs no android sdk, checked by
running it with ANDROID_HOME pointed at nothing.

* fix(sidecar): measure the stability streak as observed quiet

parameterising pollUntilStable also moved the clock to the start of the
read that opened a run of identical snapshots, so a read's own duration
counted as quiet. the pre-existing caller polls a real uiautomator dump:
at 400ms a read, 750ms of required quiet became 250ms of observed quiet
and the poll settled in two reads instead of four.

the parameters stay, the semantics go back.

* test(sidecar): pin the transition cap by driving it

it asserted 1500 >= 700 + 300, two constants, which can only fail if
someone edits a constant. it now drives awaitSettledTree against a fade
that lands after 700ms and asserts it hands back the settled tree before
the cap. cut the cap to 1000 and it goes red.

* ci: pin buf-setup-action to a commit

it takes a token now, so a floating tag is a token handed to whatever
that tag moves to. note v1 there is a branch, not a tag, so the ref
lookup that resolves it is matching-refs/heads/v1.

* ci: declare least-privilege permissions

none of the three declared any, so each got the repository default.
release.yml and docs.yml already do this. all three only check out,
build, test and upload artifacts.

* ci: fail fast when a server never comes up

the readiness loops fell through silently after 30 tries, so a server
that never started surfaced as an opaque driver failure minutes later.
each now says what did not answer and on which port.

* ci(folio): a missing trace is not a verdict

with no trace the android gate ran its grep against ./trace.jsonl and
reported "never reached AddTransactionScreen, so it never got past
login", which is not what happened. the web and ios branches had the
same misdiagnosis on exit 0.

same class, one line up: the classifier's own failure was swallowed, so
with the evidence reader dead the gate printed a healthy run and exited
0.

* ci(replay-ui): skip a run directory with no trace

the summarise step is if: always(), and under github's bash -eo pipefail
an unmatched glob stays literal, the redirect fails, pipefail carries it
into the assignment and -e kills the step. so a failed fuzz run went red
twice, once for the real reason.
2026-08-15 13:01:27 +05:30
pj a97e5fef8f docs(manual): a step bound counts observations, not runner steps
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:37 +05:30
pj 3f0132c92a fix(folio-web): bound the reachability properties by steps
At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:37 +05:30
pj 349a0644ac test(spec): guard the step unit on the authoring surface
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:37 +05:30
pj a449c99f08 test(ltl): pin that a slow policy does not fail on time alone
Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:37 +05:30
pj a1e6bd7853 fix(ltl): keep the authored window on a step-bounded obligation
reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:19 +05:30
pj 0cc20539bc test(campaign): wait for the trap instead of racing it
The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:50:19 +05:30
pj eacf3fd11f fix(analyze): divide per-hour rates by time actually worked
A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:43:12 +05:30
pj c63a2e4897 fix(campaign): record both clocks a run was measured on
Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:43:12 +05:30
pj 95a474339e perf(runner): confirm focus only when another element holds it
Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:43:12 +05:30
pj dd17f831ef fix(campaign): signal a timed-out run so it reaps its sidecar
CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-14 17:27:13 +05:30
pj 9a2bbbf769 fix(runner): confirm focus moved before typing
InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 23:00:39 +05:30
pj 6fd3fef9fa docs(manual): attrs carries raw attribute names on web too
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 22:55:20 +05:30
pj 78e561ba7e fix(verifier): name a web handle by the same ladder as a tree element
The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 22:55:20 +05:30
pj 037a803f0e fix(spec): key web attrs by the names the markup writes
attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 22:55:20 +05:30
pj 739e921788 docs: minimal changes, self-documenting code, tests as first-class
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 22:28:18 +05:30
pj b6dabaeeca fix(folio-web): enumerate authored targets and values, declare the llm generator
The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 21:45:05 +05:30
pj 4c5932b4be fix(folio): enumerate authored targets and values instead of sampling
Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 21:45:05 +05:30
pj f307354bc7 docs(manual): value generators are refused under the model policy too
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:58:18 +05:30