Commit Graph
229 Commits
Author SHA1 Message Date
pj f307354bc7 docs(manual): value generators are refused under the model policy too
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:58:18 +05:30
pj 3cab483b96 test(verifier): setup still draws, and the seeded stream is unmoved
Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:58:17 +05:30
pj 0a71e05588 feat(spec): refuse a multi-value generator while enumerating
integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:58:17 +05:30
pj 7880e3e70f fix(runner): abort on a candidate enumeration that refused
Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:44:00 +05:30
pj 0094a7fc64 fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets
Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:44:00 +05:30
pj c76ba4b497 feat(spec): refuse a multi-item authored sampler while enumerating
from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-13 00:44:00 +05:30
pj 7dc23b137a docs(manual): document object-selector key rules 2026-08-13 00:41:38 +05:30
pj 8a98915118 test(chrome): drive the live-page parity test through both selector forms 2026-08-13 00:41:23 +05:30
pj cb880edf44 fix(spec): match a merged label by its leading name on web too
The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.
2026-08-13 00:41:23 +05:30
pj b8bb44cf61 test(spec): pin the unknown-key diagnostic to one text
The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.
2026-08-13 00:39:29 +05:30
pj 269574706d feat(spec): reject an unknown object-selector key in the web runtime
Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.
2026-08-13 00:38:40 +05:30
pj 1332471b72 feat(verifier): fail the spec on a selector key that cannot match
An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.
2026-08-13 00:36:05 +05:30
pj 35b0856e6c test(hierarchy): pin both selector forms and the unknown-key report 2026-08-13 00:35:58 +05:30
pj 73f32668e6 fix(hierarchy): give id and desc one meaning in both selector forms
The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.
2026-08-13 00:34:00 +05:30
pj c77b78d63b test(chrome): compare both selector matchers over one live page
Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.
2026-08-13 00:28:16 +05:30
pj bf3f230b1b fix(spec): read the injected seed per call
Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.
2026-08-13 00:27:10 +05:30
pj a5d358e24e feat(replay-ui): render idPrefix targets as a prefix tag 2026-08-13 00:05:18 +05:30
pj 0122f241dc docs(manual): document the idPrefix selector 2026-08-13 00:05:00 +05:30
pj b1a0678bf9 feat(sidecar): match idPrefix in the tap-by-selector path 2026-08-13 00:04:40 +05:30
pj a8444cc34f feat(spec): match idPrefix in the web runtime
The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.
2026-08-13 00:02:21 +05:30
pj 59975231ce feat(chrome): translate idPrefix to a starts-with id match 2026-08-13 00:00:57 +05:30
pj 1fcb7a9f7d feat(hierarchy): match identifiers by role prefix
idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.
2026-08-13 00:00:44 +05:30
pj b8431c3e13 test(runner): both policies must dispatch the same authored action
Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:51:01 +05:30
pj 5ad2b39e6f fix(spec): carry the container on an authored scroll
serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:51:01 +05:30
pj 70573a319e feat(verifier): decode a container-only scroll
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:51:01 +05:30
pj 73f387f8bd fix(verifier): lower authored actions the way the seeded arm does
The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:51:01 +05:30
pj 2cbe03c3fb fix(analyze): divide by actions that ran
Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:43:19 +05:30
pj f6d562e3dc feat(campaign): count dispatched actions, not steps
A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:43:19 +05:30
pj 78a56c44f5 test(runner): the echo guard admits a repeated description
Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:32:24 +05:30
pj bc4b9fb25b fix(runner): report every action that was chosen and never dispatched
applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:32:24 +05:30
pj 7f5cf610bd fix(verifier): dedup candidates by what they execute, not how they read
The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:29:16 +05:30
pj caa4b16866 feat(cli): add --label-source
Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:19:24 +05:30
pj ba8c4810cf feat(trace): record the label source as arm membership
Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:19:24 +05:30
pj de00412f67 feat(runner): thread the label source to the model picker
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:19:24 +05:30
pj 8bb20c4721 feat(verifier): select the candidate label source
Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:19:24 +05:30
pj 1f71e052d7 clean up dead jetbrains mono wiring in replay-ui (#70)
* fix(replay-ui): drop @font-face rules for fonts that were never shipped

* fix(replay-ui): drop unresolvable JetBrains Mono from --font-mono stack

* chore(replay-ui): remove vestigial empty public/fonts dir
2026-08-12 23:18:03 +05:30
pj 019d608f65 feat(analyze): survival analysis over campaign directories
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 23:03:03 +05:30
pj 71dffef2f2 docs(manual): document llm-calls.jsonl
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:57:06 +05:30
pj 2a3e419652 fix(runner): record when a chosen action was never dispatched
A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:57:06 +05:30
pj 0b70dd1659 fix(runner): a guard-skipped step is no longer a silent log line
The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:57:06 +05:30
pj a97c09f6dc feat(verifier): expose the step a snapshot was observed at
It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:56:58 +05:30
pj ec5872ac6e feat(trace): record one typed outcome per model-driven step
llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:56:58 +05:30
pj b9cdf7a571 feat(llmclient): parse usage and the served model
An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:56:58 +05:30
pj 76dce1a75e experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs

runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record arm membership and host in meta.json

meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --arm and populate run meta from it

Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): sweep seeds for one experiment cell

campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.

Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): no silent generator fallback, and llm on web

--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.

pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): make the hierarchy dump agree with the web runtime

Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.

scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.

The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(conformance): give the gate reproducible seeds

SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.

The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit editable as a plain boolean

editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): leave the head subtree out of the web target walk

collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(chrome): compare the facts both hosts derive from one DOM

The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.

This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* chore(make): run the browser packages one at a time

Both launch Chrome and launching two at once has failed with "Launch: context
canceled".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style: remove every em-dash and en-dash

Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): honor the caller context in Launch

Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.

The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(sidecarassets): publish the extracted jar through a rename

Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): kill a run that outlives --run-timeout

A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* style(test): gofmt browser_test.go

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
2026-08-12 22:20:31 +05:30
pj 26b49b379a fix ltl semantics and unify action enumeration (#71)
* fix(ltl): give every thunk a construction identity

Two distinct unnamed predicates both described as "Thunk(...)", so obligation
collapse merged their residuals and could drop a live violation. Identity is
assigned at construction and the fields are unexported, so a thunk cannot be
built without one.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): reduce a thrown-predicate residual instead of panicking

The verifier substitutes an ErrorFormula for the residual of a property whose
predicate threw, and that residual is fed back in on the next step. reduce had
no case for it, so the run crashed. It re-reports the same failure now.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): make a bounded always the dual of a bounded eventually

G<=n(f) and not F<=n(not f) disagreed on traces where the inner was still
pending when the window closed, so nnf's negation normal form was not semantics
preserving. Both sides now range over the observations at which their inner can
definitely resolve: the eventually keeps a pending inner as a disjunct, and the
always discharges vacuously at window close.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): arm a one-shot root once per run

A root that carries its own horizon is one obligation for the whole run, not one
per observation. Re-instantiating a top-level eventually monitored G F<=n(p)
instead of F<=n(p) and left one live obligation per step behind; a bounded
always restarted its window every step and never closed.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): stop wrapping a top-level eventually in always

`eventually(p).within(300, "seconds")` as a property meant "within 300 seconds
of every step", which spawned an obligation per step with its own resolved
deadline. A 553-step run carried 553 of them and serialized a 75 KB residual.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(ltl): serialize the resolved deadline of a bounded window

Two obligations spawned at different steps from one duration-bounded formula
differ only in the deadline the evaluator resolved for them, so they serialized
identically and the trace erased a distinction the evaluator makes. The authored
window stays in amount/unit; the resolved deadline rides alongside.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): split a witness's origin step from its detection step

A deferred obligation spans two steps: the one that armed it and the one whose
reduction failed. They were conflated under one index, so the extractor snapshot
(which is the detecting step's state) was reported against the origin step.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): record a witness's detection step in the trace

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* feat(replay-ui): show the step a violation was detected at

The witness evidence is the detecting step's state, so say which step that is
and let a reader jump to it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(verifier): record the extractor state the predicates actually read

On the web path extractor bodies are evaluated in V8 and injected here, but only
the goja value was replaced. The trace diff and the violation witness therefore
described a state no property ever saw.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* refactor(spec): one candidate producer over one target-eligibility rule

Both hosts routed verbs themselves and both policies enumerated their own
actions, and all four drifted. Web sent `swipes` to scrollable containers only,
so swipe-to-dismiss on a list row was reachable on native and unreachable on
web; the model policy folded gestures its own way and could not reach what the
seeded picker drew.

A host now reports facts about every element and never decides which verb may
act on it: targets.ts acceptsTarget owns that for both. pick.ts builtinCandidates
is the single enumeration, and the model policy reads it through
__sanderlingEnumerateBuiltin__ instead of reimplementing it in Go.

Gesture verbs change with it: scrolls stay vertical over scrollable containers,
swipes go free-form in all four directions from any element with real bounds.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(runner): name a builtin scroll by its drag origin

A builtin gesture carries endpoints and no selector, so every scroll rendered as
"Scroll down " in the prompt's recent-action memory and two scrollable regions
were indistinguishable.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): clear storage over cdp instead of scripting an opaque origin

Launch runs while the tab is still on about:blank, whose opaque origin denies
storage access, so localStorage.clear() threw SecurityError and every web run
died at launch. Storage.clearDataForOrigin needs no navigation. The exception
helper lands here because "Uncaught" is what hid this for so long.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(chrome): enable the swiftshader webgl fallback

Headless Chrome runs with --disable-gpu, and without this flag it refuses the
software WebGL backend: getContext returns null, so a canvas-rendered app paints
nothing and every screenshot is identical black.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* fix(web): resolve testTag through data-testid or id

Compose Multiplatform emits its testTag into the element id, which the native
table already accepts via the resource-id alias. The two web selector tables
were the only place that rejected it.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* test(spec): type-check the spec api as part of make test

The fake runtime in api.test.ts did not return a chainable handle from extract,
so the file had not type-checked since named() was added. Wiring the check into
make test stops it drifting again.

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J

* docs(manual): one-shot eventually and the gesture verbs

Claude-Session: https://claude.ai/code/session_01Fj4wJUikdABuMQEETwW55J
2026-08-12 18:06:04 +05:30
pj 7343085614 llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker

* feat(spec): make llm marker inert on the JS picker

* feat(spec): expose __sanderlingSampleInput__ corpus draw

* feat(openrouter): minimal chat-completions client

* test(openrouter): cover request shape, parse, and errors

* feat(verifier): thread screenshot + capture corpus sampler

* feat(verifier): LLM accessors — candidates, config, sampler

* test(verifier): cover AllCandidates, LLMConfig, SampleInput

* feat(trace): record action Source and LLMReasoning

* feat(runner): thread step screenshot into PushSnapshot

* feat(runner): llmSource selects actions via OpenRouter

* feat(runner): wire llmSource selection and trace stamping

* test(runner): cover llmSource selection, mapping, downscale

* docs(folio): add llm action-backend example spec

* docs(folio): document the LLM action backend run

* feat(llmclient): support OPENAI_API_KEY, openrouter wins

* refactor(runner): rename openrouter package to llmclient

* docs: both api keys, example model gpt-5.4-nano

* docs: add pr style rules to claude.md

* fix(runner): explain action kinds in llm prompt to stop swipe loops

* feat(trace): record llm ranked list and chosen rank

* feat(runner): stamp llm ranked list and chosen rank on trace

* fix(runner): tap by selector to survive layout shift after observe

* revert(runner): drop selector-first tap; broke path/testTag selectors

* feat(spec): llm() accepts optional instructions

* feat(verifier): read llm instructions off config

* feat(runner): append spec instructions to llm system prompt

* docs(folio): describe app in llm spec instructions

* feat(bundler): map generator export to globalThis.generator

* feat(verifier): read llm config off globalThis.generator

* feat(runner): gate llm source on --generator flag

* feat(cmd): add --generator llm|seeded flag

* test: cover --generator flag parsing and pickSources gating

* feat(verifier): enumerate llm candidates by walking actionsRoot

collect-walk the weighted action tree: recurse weighted branches
accumulating selection probability, call authored leaves once for
concrete actions, enumerate builtins per element. label controls by
visible text (borrowing descendant text), fold gestures into directional
scrolls over scrollable containers, drop disabled, dedup descriptions.

* test(verifier): cover candidate enumeration walk

* feat(verifier): add SetupAction to walk setup without the seeded root

* test(verifier): cover SetupAction setup-only precedence

* refactor(llmclient): make JSONSchema.Schema raw json for pinned field order

* feat(trace): record llm choice number and chosen_action echo

* feat(runner): llm picks one number from weighted candidates

drop the seeded-root call for a setup-only precedence path, render a
numbered weighted candidate list, pin a reasoning-first choice schema,
strict-skip when chosen_action does not echo the numbered entry, and let
the model supply typed values (corpus fallback when empty).

* test(runner): cover choice schema, strict-skip, and setup precedence

* refactor(verifier): drop the superseded AllCandidates enumeration

* feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts

* fix(verifier): label editable fields by hint, not the typed value

an editable field's own text is its transient content; prefer the hint
so the field is named by purpose and the label stays stable.

* test(runner): cover weight-suffixed echo and stripWeightSuffix

* fix(runner): accept chosen_action echo that carries the weight suffix

real runs showed the model copies the whole numbered line including the
trailing (w34) weight annotation, so strict-skip rejected ~91% of picks
and the llm was paralyzed. strip the weight suffix before comparing. also
nudge the prompt to stress-test repeated submissions (idempotency).

* fix(verifier): skip llm enumeration on cross-fade frames

a navhost mid-transition carries >1 route *Screen in a collapsed
coordinate space; acting on it taps garbage (soft keyboard). real runs
showed the llm acting on 44% of steps being such frames. skip them so the
llm re-observes a settled frame next step.

* feat(folio): show current balance on the add-transaction screen

renders the account's balance (testTag TxnCurrentBalance) below the
account name, above the credit/debit toggle, so before/after screenshots
carry comparison data.

* fix(replay): derive device space from screen extent, not first node

the first positive-bounds element is often a short status-bar node
(320x24 on android); using it gave a 320/24 aspect ratio that squashed
the screenshot overlay into a grey horizontal band. use the max extent
across elements (like the runner's screenBounds) instead.

* fix(folio): show balance as a compact one-line label

per review: one line, account-name-sized, e.g. "Balance: $0.00"
instead of a large balance card.

* fix(folio): move balance into the header, one compact line under the account name

* fix(replay): attribute deferred violations to the causing step, not detection

* fix(replay): show a step's own violations in both panels, no next-step bleed

* refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check

* chore: ignore .playwright-mcp scratch output

* docs: document the llm generator and --generator flag

* docs(spec): correct the llm() comment; config reads off globalThis.generator

* docs: add pr description rules
2026-07-31 21:12:00 +05:30
pj 6b0d6cb971 WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial

* feat(test): add --device flag to target a specific Android device by serial

* feat(folio): select Android device via ANDROID_DEVICE in justfile

* feat(conformance): add android backend to the gate suite

* feat(android): keep device awake and unlocked so the app stays foreground

* feat(conformance): prep physical android device (autofill/verifier/stayon)

* fix(android): make device prep best-effort so OEM-blocked commands don't abort the run

* fix(verifier): require positive bounds for swipe candidates

A zero-bounds element centers at (0,0); a downward swipe from the
top-left corner is the system gesture that pulls down the notification
shade, dragging the fuzzer out of the app. Swipes now require positive
bounds like every other verb.

* fix(runner): harden app-scope guard against launcher and overlays

The per-step guard now relaunches and waits until the app window is
actually drawn before proceeding, so a slow physical-device relaunch no
longer lets an observe or action land on the launcher. It also detects a
system overlay (notification shade) stealing window focus while the app
stays resumed, and dismisses it with back.

* feat(android): harden physical-device runs in device prep

Device prep now disables the AOSP cached-app freezer, phantom-process
killer, and Doze (and exempts the driver) so OEM background management
stops suspending the driver mid-run. Adds ReinstallApp for clear-state on
ROMs that deny pm clear, and teaches focus detection to report the
notification shade as systemui so the scope guard can dismiss it.

* feat(driver): clear-state via APK reinstall when pm clear is blocked

When an APK path is set, Android clear-state resets the app by
uninstalling and reinstalling instead of asking the sidecar to pm clear,
which hardened OEM builds (ColorOS) deny even to the adb shell user.
Falls back to the sidecar clear path when no APK path is provided.

* feat(cli): add --android-app-path for clear-state reinstall

Wires the APK path from the test command through to the sidecar client so
Android clear-state can reset apps on OEM builds that deny pm clear.

* chore(folio): pass --android-app-path in just test

* fix(runner): clamp swipe/scroll origin out of edge gesture zones

A gesture starting in the top status-bar strip pulls down the
notification shade; the bottom and side strips are the home and back
gestures. Any of them drags the fuzzer out of the app. Swipe and scroll
origins are now clamped into a safe inner area sized from the maximum
element extent (the Android hierarchy root reports zero bounds, so the
extent is the reliable screen size). Calibrated on device: origins below
~7% of height no longer open the shade.

* perf(sidecar): faster Android text input and drop redundant settle poll

inputText now uses adb `input text` for short shell-safe ASCII (~5x
faster than the driver's per-character path) and falls back to the driver
for unicode, injection payloads, and overflow-length strings. waitForIdle
drops the structural-hash poll that followed waitForAppToSettle: each
hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per
mutating step for marginal benefit, and the runner already re-fetches
transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4
still pass.

* fix(verifier): exclude soft-keyboard region from action candidates

The fuzzer was tapping Gboard's "Settings" key, navigating out of the
app. That key is a bare FrameLayout with a content-desc and no package or
resource-id, so the package-based scope filter missed it. Candidates whose
center falls in the keyboard region (derived from the IME elements' bounds)
are now dropped, so no tap or long-press lands on a key. Opt-in with app
scoping; unscoped runs keep every node.

* perf(runner): replace focus-tap settle with a brief wait

The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText
step on a physical device while the keyboard animated in. The tap registers
focus immediately and text is injected into the focused view, so a short
fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay
green.

* chore(conformance): platform-aware G5 p95 budget for android

The 2500ms ceiling was calibrated on the iOS simulator. A physical Android
device drives every step over USB (snapshot + settle + adb round-trips), so
its per-step floor is several times higher; holding it to 2500ms would force
removing the settle/retry logic the correctness gates depend on. The android
backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500.

* fix(sidecar): retry maestro android driver startup

The maestro Android driver's dadb.open() occasionally misses its startup
deadline (its instrumentation host is slow to come up right after a reboot
or per-run reinstall), which aborted the whole run. Retry the open a few
times with a short backoff so a transient timeout recovers.

* chore(conformance): widen android G5 budget to 5500ms

Physical-device p95 swung 3209-4612ms across sessions (cold runs right
after a reboot are slower). 4500ms was too tight for that jitter; 5500ms
covers the observed ceiling with headroom.

* web replay fix

* feat(android): force 3-button nav during runs to prevent app drift

On gesture navigation a fuzzer swipe can trigger swipe-up-home or
edge-back and fling the app off screen. Device-prep now switches to
3-button navigation for the run (no edge gestures; the nav bar's buttons
are systemui-owned and already excluded from action candidates) and
restores the original navigation mode when the run ends. Best effort:
leaves nav untouched if the overlay command is unavailable.

* fix(android): target the selected device in adb reads; don't strand nav mode

Review fixes:
- ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the
  foreground/scope guard works when several devices are attached (the --device
  path). Previously they ran bare `adb shell`, which errors with multiple
  devices, silently disabling app-scope enforcement. The sidecar client passes
  its serial through.
- Extract an adbArgs helper and route every adb call through it, removing four
  duplicated serial-arg builders.
- ForceThreeButtonNav now decides what to restore before changing anything: if
  the current mode is unknown or already 3-button it leaves nav untouched,
  instead of switching and then stranding the device in 3-button. Logic split
  into the pure navModeToRestore, now unit tested.

* fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds

Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was
stranded above screenBounds by an insertion). Extend the clamp test to assert an
off-screen destination is clamped onto the screen and that the origin lands
exactly on the margin.

* test(verifier): cover keyboardRegionTop, including the decor-view guard

The full-screen IME decor view rejection had no test; removing it left the
suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in
favor of the real keyboard line, and decor-only -> sentinel.

* style(cli): gofmt testOptions field alignment

* fix(sidecar): keep a leading dash off the fast input path

A value starting with '-' could be read as an option by `adb input text`, so
the fast-path regex now requires a non-dash first character; such values fall
back to the driver. Also cover the dadb-target branch where a colon precedes a
non-numeric port (a USB serial, not host:port).

* refactor(verifier): scope action candidates by window ownership

Replaces the leaky per-element package check and the keyboard-region Y
heuristic with one rule: walk the window tree propagating each node's owning
package (empty and the neutral android framework package are transparent); a
node is in scope only when no concrete foreign package owns it (the app's own
window carries no package on Compose apps) or the owner is the app package.

This drops whole foreign windows (soft keyboard, system UI, launcher) AND
their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which
the old empty-package-is-in-scope rule admitted and which navigated out of the
app. Deletes keyboardRegionTop/isInputMethodElement.

* fix(runner): re-check foreground at apply time, skip stale actions

ensureForeground runs before observe, but the app can leave between observe and
apply (a prior gesture settling late); swipes/keys then fire stale coordinates
onto whatever screen is now up. Re-check foreground immediately before applying
and, when the app is gone, skip the action and log it (making the escape
visible) so the next step's guard relaunches instead.

* fix(android): type long ASCII via fast guarded path to stop keystroke escape

A 4096-char corpus string exceeded the fast input cap and fell to the
per-character driver path, which takes ~120s. During that uninterruptible
window focus could leave the app and the remaining keystrokes sprayed into
the launcher search box. Route shell-safe ASCII of any length through adb
input text, chunked, re-checking the foreground app between chunks and
stopping if it changed.

* chore: ignore gate artifacts and local scratch files

* refactor(runner): narrow gesture clamp to the top shade strip

3-button nav (forced for every run) disables the side back and bottom home
gestures at the OS level. On-device probing confirmed side and bottom swipe
origins no longer drift, leaving the notification shade as the only edge
gesture a swipe can trigger. Clamp only the top strip; keep origin and
destination on screen otherwise.

* chore(format): add .editorconfig enforcing 80-column limit

* chore(format): add prettier config with 80-char printWidth

* chore(deps): add prettier devDependency to replay-ui

* chore(deps): add prettier devDependency to folio-web

* chore(deps): add prettier devDependency to spec package

* chore(format): add swift-format config with 80-char lineLength

* feat(format): add make fmt targets for per-language 80-col formatting

* fix(runner): translate gesture to safe area so near-top scrolls keep direction

Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp.

* fix(runner): apply-time guard consults focused window, not just resumed activity

ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay.

* test(runner): cover apply-time foreground skip and appIsForeground table

Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised.

* fix(sidecar): harden android driver open, input guard, pressKey, foreground marker

- openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF).
- pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract.
- the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner.
- foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard.

* fix(conformance): pin self-test p95 budget and score install failures as run failures

self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue.

* fix(android): require --device when several devices are connected

With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous.

* refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands

The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test.

* perf(verifier): memoize scopedElements per tree

scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot.

* fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear

Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail.

* test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms

The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup.

* refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty

bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true.

* perf(sidecar): reuse a single Jackson ObjectMapper

structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val.

* refactor(android): remove unused AdbReverse/AdbReverseRemove

No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed.

* style(runner): trim non-load-bearing comments from this PR's runner code and tests

* style(sidecar): trim non-load-bearing comments from this PR's driver code and tests
2026-06-11 10:10:05 +05:30
pj 991c583eb9 docs update with case study (#63)
* docs(manual): add introduction page

* docs(manual): rewrite getting started as guided first run

* docs(manual): rewrite writing specs as a folio tutorial

* docs(manual): document missing spec API in reference

* docs(manual): plain-language rewrite of runs page

* docs: real introductions on index pages and README

* fix(docs): sibling links from directory-style pages need ../

* fix(docs): correct sampling and restart-cost claims to match implementation

* docs: nav lists Introduction and Case study; roadmap points to milestone

* docs(manual): make getting started target the reader's own app, not Folio

* docs(manual): add Folio case study page

* docs: point manual navigation at the case study

* docs(readme): lead with the case study, fix roadmap link

* docs: roadmap links to milestone, sync clear-data default and cross-links
2026-06-09 20:13:54 +05:30
pj 90224dfd06 Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers

The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.

* feat(ios): resolve physical devices from devicectl

ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.

* feat(ioscompanion): runner-only device driver mode

NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.

* test(ioscompanion): cover device wiring, routing, and shell-out argv

Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.

* feat(testrun): route physical-device iOS runs to the device driver

Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.

* feat(doctor): device prereqs replace java/sidecar for ios-device

iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.

* feat(conformance): device backend uses iphoneos app and tunnel orphan checks

The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).

* feat(folio): device build linking the iosArm64 framework

project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.

* docs(cli): document ios-device doctor checks and the device flags

The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.

* fix(ioscompanion): resolve signing key path to absolute

xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.

* fix(ioscompanion): re-enable signing for the device runner build

companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.

* fix(ioscompanion): key the device build cache on signing identity

The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.

* docs(getting-started): document physical iOS device setup

Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.

* feat(ios): native usbmux client and in-process tunnel forwarder

Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.

* refactor(ios): drive device tunnel via io.Closer seam

Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.

* refactor(ios): remove iproxy spawn from device runner

* test(ios): cover tunnel close via io.Closer not child process

* feat(doctor): check usbmuxd socket instead of iproxy on PATH

* chore(conformance): drop iproxy orphan check; tunnel is in-process

* docs(ios): device tunnel uses native usbmux, nothing to install

* chore: gitignore the signing keys directory

* feat(folio): add Android launcher icon (black bg, white dot)

* feat(folio): add iOS app icon (black bg, white dot)

* feat(folio): add web favicon (black bg, white dot)

* docs(ioscompanion): fix stale const comments

* refactor(ioscompanion): inline single-use devicectl argv builders

* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests

* refactor(ioscompanion): inline firstNonEmpty

* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam

* test(doctor): trim redundant signing-check test

* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env

* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
2026-06-09 18:38:52 +05:30
pj e04631d4c9 Remove the JVM sidecar's iOS backend (#65)
* refactor(sidecar): drop IosDriverBackend

* refactor(sidecar): route ios platform off the iOS backend

* chore(sidecar): remove maestro ios dependencies

* feat(testrun): reject physical iOS with a clear message

* refactor(sidecar): drop iOS hierarchy helpers and their test

* build(sidecar): strip iOS runner bundles and classes from the fat jar

* test(testrun): cover physical-iOS rejection
2026-06-08 20:29:35 +05:30