mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
d5f3937338455e5cdb1cca72f4265486319c6c82
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9b4ff5f247 |
record what the model picker did, and make both policies see the same actions (#74)
* feat(llmclient): parse usage and the served model An LLM-in-the-loop evaluation has to report tokens per action and cost per defect, and the client discarded both counters. Served model is recorded separately from the requested one because a router can substitute a differently-priced variant. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record one typed outcome per model-driven step llm-calls.jsonl carries the prompts as sent, the candidate list as the model saw it, the screenshot reference, the raw response, tokens, latency and how the step ended. It sits beside trace.jsonl rather than inside it because every trace line already carries a full hierarchy and both the replay server and the campaign summarizer scan all of them; folding prompts in would grow the lines those readers parse for data neither reads. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): expose the step a snapshot was observed at It lags the runner's current step whenever a transitional tree caused an observation to be skipped, which is exactly when the model is shown an older screen than the step it is choosing for. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): a guard-skipped step is no longer a silent log line The strict echo-skip left only a logger.Warn, so a step the guard discarded was indistinguishable in the trace from a picker that legitimately declined. Any yield or actions-per-hour figure computed from model traces mixed the two. Every path that ends a step without a model-chosen action now records its own outcome. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): record when a chosen action was never dispatched A step could carry a next_action that the foreground guard or an apply error stopped from running, and nothing said so. An executed-action count read off trace.jsonl included actions that acted on nothing. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): document llm-calls.jsonl Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(analyze): survival analysis over campaign directories Steps to first violation with clean runs right-censored at the budget, since per-run yield is a binary at 11 to 45 percent and separating two arms on it would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum with Vargha-Delaney A12, Holm within each family. A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator and would be believed, so every statistic is validated against a published worked example with the source named in the test: R survdiff on aml, Freireich 6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and the tie-corrected variance, checked against an exact permutation variance. Failed and timed-out runs are excluded as missing data and counted by reason, never treated as censored observations, which would bias the result. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): select the candidate label source Candidates takes the label source as an argument rather than storing it, which is what keeps the asymmetry structural: the seeded picker selects by index and never calls Candidates, so the mode cannot reach it. That asymmetry is load-bearing, because it makes the two seeded cells of the factorial a manipulation check with identical draw streams. The identifier ladder deliberately has no text rung. A fallback that reached for text would silently turn one arm back into the other. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(runner): thread the label source to the model picker Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record the label source as arm membership Recorded for seeded runs too, unlike model and instructions. Without it the two seeded cells are indistinguishable in the artifact and the manipulation check cannot be grouped. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(cli): add --label-source Unknown values are rejected at parse time rather than falling back to the default, matching the generator check: a campaign that completes with the wrong arm and a correct-looking output directory is worse than one that fails. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): dedup candidates by what they execute, not how they read The dedup key was the rendered description, which embeds the label, so two distinct controls sharing a visible label collapsed to one entry and the survivor carried the first one's action. The second control was not mislabelled, it was absent from the candidate list, so no policy could reach it. Two scrollable containers collapsed the same way, leaving the second unscrollable. The key is now the executable Action struct itself plus whether the model supplies the typed text, so a new Action field cannot silently fall out of it. Descriptions may now repeat; the numbering disambiguates and the echo guard is index-anchored, not description-anchored. This also makes the label source a pure observation-channel change. It was not one before: the label fed the dedup key, so the two arms of the labelling factor enumerated different-sized candidate lists, in both directions depending on the screen. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): report every action that was chosen and never dispatched applyAction could return nil without calling the driver, so the trace showed an action that looked executed and acted on nothing. Six paths did it: a tap, double-tap or long-press whose coordinates do not resolve and which carries no selector, a long-press whose selector is stale, an empty key press, and a zero-duration wait. It now reports whether it dispatched, and the runner records the reason and clears lastAction so the verifier never attributes the next state to an action that did not run. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): the echo guard admits a repeated description Descriptions can now repeat after candidates dedup by what they execute. The guard is index-anchored, so this pins that a repeated string cannot make it misfire in either direction. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): count dispatched actions, not steps A step where the policy declined has no action, and a step whose action was never dispatched did nothing. Both were being counted as actions by everything downstream. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(analyze): divide by actions that ran Defects per thousand actions counted every step, including steps that chose nothing and steps whose action was never dispatched. The inflation is policy-dependent, so it does not cancel between arms: on the fixture campaign the model arm's yield was reported at 60.3 per thousand against a true 120.7, because half its steps did nothing. A runs.jsonl without the count is refused by name and line rather than read as zero actions, which would report every per-action rate wrongly. The report also carries steps beside actions now, so the gap is visible rather than folded into a denominator. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): lower authored actions the way the seeded arm does The authored descriptor path had no parity guard and diverged from the wire format on almost every verb. A Wait lost its duration and was skipped as a zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe to (0,0) instead of being dropped. An authored target object with no x property panicked the whole run at candidate enumeration: ToInteger was called on a nil goja.Value. A target on the screen origin is still kept, so the drop rule cannot swallow it. Builtins were never affected. They serialize through the same path the seeded arm uses, which the existing policy parity test covers. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): decode a container-only scroll Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): carry the container on an authored scroll serializeAction sent the container's own point as both endpoints, so an authored Scroll({in, direction}) reached the driver as a drag from a point to itself and did nothing, on the seeded arm. The wire now carries the selector and leaves the drag to the runner, which sizes it from the container's bounds and has always had tested support for it that nothing could produce. No rng runs in the serializer, which lowers an already-drawn action, so the draw stream does not move. Builtin scrolls compute both endpoints and their bytes are unchanged. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): both policies must dispatch the same authored action Compares the recorded driver calls across 13 authored shapes. The builtin path had a parity guard and the authored path had none, which is why it drifted on almost every verb. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(hierarchy): match identifiers by role prefix idPrefix: is id: with starts-with in place of equality, so a list whose rows are named <role>_<record id> is reachable by the durable half. The Android package prefix is skipped the same way id: skips it. Routing both prefix kinds through matchAttr also makes the object form work: {descPrefix: ...} matched nothing on the native side while the web runtime honoured it. * feat(chrome): translate idPrefix to a starts-with id match * feat(spec): match idPrefix in the web runtime The DOM has no package prefix, so the native rule reduces to [id^=]. Both prefix kinds now go through the one key table, which drops the separate descPrefix branch that string and object selectors each carried. * feat(sidecar): match idPrefix in the tap-by-selector path * docs(manual): document the idPrefix selector * feat(replay-ui): render idPrefix targets as a prefix tag * fix(spec): read the injected seed per call Binding it at module scope bound it to whenever the module was first imported, so a test file that imported the runtime before setting SANDERLING_SEED froze the seed at zero for every file after it. The bundler still replaces the expression with a literal. * test(chrome): compare both selector matchers over one live page Selector matching is written once per runtime: internal/hierarchy over the dump, web-runtime.ts over the DOM. Nothing made the two agree, and a selector that resolves on one and not the other is silent, since an empty match yields no action and the run still passes. * fix(hierarchy): give id and desc one meaning in both selector forms The object form fell through to the raw attribute map, which carries no id or desc key on any platform, so {id: "save"} matched nothing while "id:save" matched. The repo's own web spec uses the object form thirty times. Both forms now resolve through one switch. Adds the accepted-key list and UnknownSelectorKeys with it, since the same silence hides any mistyped key. A key some element carries is always accepted, so raw driver attributes stay reachable. * test(hierarchy): pin both selector forms and the unknown-key report * feat(verifier): fail the spec on a selector key that cannot match An empty match is indistinguishable from a screen with no such element, so a mistyped key generates no action for the whole run and the campaign finishes clean having explored nothing. The goja boundary now throws, naming the key and the accepted list. * feat(spec): reject an unknown object-selector key in the web runtime Same rule and the same message as the native side: a key no element can carry throws instead of matching nothing. The accepted list is one list, committed as a fixture both suites assert, so a spec cannot be accepted by one runtime and rejected by the other. * test(spec): pin the unknown-key diagnostic to one text The two runtimes each claimed to raise the other's message and nothing checked it. Both now render the committed text for the committed key. * fix(spec): match a merged label by its leading name on web too The native desc rule accepts the label or the label at the head of an iOS merged label; both web translators compared the whole string, so the same selector matched natively and missed on web. The live-page parity test caught it. * test(chrome): drive the live-page parity test through both selector forms * docs(manual): document object-selector key rules * feat(spec): refuse a multi-item authored sampler while enumerating from().generate() draws from the picker's rng, which exists only inside walkActions. The model policy enumerates authored leaves outside that walk, so the sampler silently yielded its first item on every step: measured over 30 draws the seeded arm reached three targets in roughly equal proportion and the model was offered only the first. The two policies had different action spaces and nothing said so. A single-item sampler short-circuits before the rng, so both policies get the same value and it is not refused. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets Candidates returns an error now. The refusal is thrown at the draw and wrapped with the source of the leaf that made it, since generate() cannot know which leaf it is inside. Only that marked refusal is fatal: this walk calls every leaf on every step, so promoting the rest would kill model runs the seeded arm survives. Authored actions on a disabled target are no longer dropped from the model's candidate list. The seeded picker executes whatever the leaf authored, and a control the application forgot to re-enable is exactly where boundary defects live, so a policy that cannot attempt it cannot find them. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): abort on a candidate enumeration that refused Recorded as candidates_failed before the run stops, so the trace says why. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(spec): refuse a multi-value generator while enumerating integers, strings, emails and edgeCaseText read the same rng from() does, so under the model policy an authored InputText typed the same value on every step while the seeded arm varied it. That is a silently different experiment, not just a silently different action space. Single-valued spans are exempt, because both policies then get the same value: between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is still refused, since the length is pinned but each character is drawn from 62. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(verifier): setup still draws, and the seeded stream is unmoved Setup runs through the picker with the rng under both policies, so a generator there is legitimate and must keep working. Interleaving enumeration and setup catches the flag leaking out of the model's walk. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): value generators are refused under the model policy too Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio): enumerate authored targets and values instead of sampling Sampling inside an authored leaf is refused under the model policy now, because the draw collapses to its first item there. Each sampled leaf offers one action per value instead. Lists are short, three rather than five, because the two form leaves also carry their submit and the seeded picker splits a leaf's probability across the actions it returns. The doubleTaps path that reaches the planted defect is unchanged at 5.88 percent, since no root or defaults weight moved. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio-web): enumerate authored targets and values, declare the llm generator The two edge-case typing leaves become the typing builtin at their combined weight: that text is deliberately not domain-specific, so naming the field and leaving the text to the policy is the designed path, and it keeps the seeded arm on the corpus while the model writes its own. Total weight is unchanged at 165, so every surviving branch keeps its share and submitTxn stays at 9.70 percent. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs: minimal changes, self-documenting code, tests as first-class Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): key web attrs by the names the markup writes attrs was spread from element.dataset, whose DOMStringMap keys are camelCase, so a spec reading attrs["data-cents"] the way every native host reports it read undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently zero: someTransactionExists could never be satisfied, balanceMatchesTransaction Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and passed vacuously. Three properties reported nothing because the harness was blind, not because the application was correct. The handle also fills hintText and editable now, so an authored InputText on web names its field the way the same action names it on Android instead of rendering as Type "12.34" into "". Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): name a web handle by the same ladder as a tree element The handle fallback read only text, which is textContent and therefore always empty for an input, so the model could not tell the amount field from the note field. It now mirrors visibleLabel's ladder rather than introducing a second naming scheme. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): attrs carries raw attribute names on web too Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): confirm focus moved before typing InputText tapped its target, slept, then typed. Android and web both inject into whatever holds focus, so a tap that missed sent the whole string somewhere else and nothing reported it. On an emulator with a floating keyboard panel parked over the password field, the tap pressed the keyboard's emoji key and every step appended the password to the email instead, forever, because the setup leaf is guarded on the password being empty. The hierarchy is re-read after the tap and the target, or something in its subtree, must hold focus. Platforms whose hierarchy carries no focused attribute skip the read, so they pay nothing. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): signal a timed-out run so it reaps its sidecar CommandContext kills outright, so a run stopped by --run-timeout never ran its own shutdown and left a sidecar holding a port and a quarter gigabyte, reparented to init and deaf to SIGTERM. The timeout exists for unattended hosts, which is exactly where nobody is watching to reap what it leaves. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * perf(runner): confirm focus only when another element holds it Measured over 717 InputText steps: nothing was focused before the tap 23.8 percent of the time, the target already held focus 60.4 percent, and a different element held it 15.8 percent. Silent corruption is only reachable from that third class, and all four real rejections observed came from it. Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy reads, and recovers about 8 percent of Android run time. The pre-tap and post-tap conditions are now the same predicate stated once. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): record both clocks a run was measured on Duration came from the monotonic clock, which does not advance while a host sleeps: one calibration run under-reported by about 15 minutes. A run now carries monotonic_millis for how long it worked and wall_clock_millis for how much time passed, which is what makes a sleep visible at all. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(analyze): divide per-hour rates by time actually worked A host asleep mid-run tested nothing, and charging that sleep to an arm reports it slower for a reason unrelated to the arm. The legend also claimed wall clock while the number was monotonic. Campaigns written before the split are still read through the old field name so their run hours do not silently zero. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(campaign): wait for the trap instead of racing it The reaping test gave the wedged script one second to install its TERM trap, so a loaded machine signalled it first and the test failed for a reason it does not test. It now waits for the script to say the trap exists, then cancels. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(ltl): keep the authored window on a step-bounded obligation reduce decremented StepBound into the residual, so the trace reported the remaining window rather than the authored one: a within(1915, "steps") showed up as 1875 after 40 steps, and the replay UI renders that string verbatim. The duration case was fixed when bounded windows were made to serialize their resolved deadline; the step case was not, and withinFor's comment claimed otherwise. The window is now immutable and the closing observation is resolved once, which mirrors Deadline exactly. A step counts observations the evaluator reduced, not steps the runner executed, because a skipped step gave the property no chance to discharge and transitional-step rate is itself policy-dependent. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(ltl): pin that a slow policy does not fail on time alone Same 300-observation trace at two cadences: a 300 second bound holds for the seeded arm and violates for the model arm eight observations before the predicate fires, while a step bound holds for both. Green before and after, because the step unit already worked; this pins the property rather than fixing it. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(spec): guard the step unit on the authoring surface Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio-web): bound the reachability properties by steps At one model call per step the model arm takes 359 seconds where the seeded arm takes 47, so a second-based deadline reported violations that were the arm's speed rather than the application's behaviour. The three cross-arm reachability properties now bound by steps, derived at the measured 6.383 steps per second. The two auth-transition properties keep seconds: a user waits through those regardless of which policy is driving. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): a step bound counts observations, not runner steps Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): make the label source a cell dimension A 2x2 of policy against labelling needs the runner to express both factors. It could only express the policy, so half the factorial had to go through --extra, where the manifest would not record what was actually run. Rejected at parse rather than on dispatch: a sweep that finds the bad value on run 1 of 40 has already spent a cell's worth of device time. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): record the label source in the manifest A finished sweep should say which cell it ran without anyone having to remember the invocation. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): name a web field by its hint, not its CSS class visibleLabel reads hintText first for an editable element. The dump never emitted it, so an empty web input fell through text, description and descendant text to its class name, and the model was shown an identifier no user can read on exactly the fields a labelling experiment varies. Same ladder as fieldHint in web-runtime.ts, so one field is named one way on both hosts. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(web-runtime): answer clickable for an element reached through ax The handle hardcoded true, so every text node and container a spec reached through state.ax claimed to be a tap target while the enumeration and the hierarchy dump both resolved it through the tappable selector. The parity test now compares the handle against the enumeration element by element in a real browser, which is where the three answers have to agree. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs: every target runs on this machine, so start one rather than skip it Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): a selector the tree cannot resolve is not a focus failure otherElementHoldsFocus answered true when FindNode returned nothing, so an unresolvable target read as "another element holds focus". confirmFocus then re-dumped, resolved nothing again, and errored unconditionally. Three of those in a row abort the run. Not knowing where the target is says nothing about where the text would land. The guard's real case, a resolved target with focus outside its subtree, still errors exactly as before. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): emit data-testid so both resolvers name the same element The V8 host names a web target by data-testid and TapSelector translates that selector into a CSS attribute match, but the dump carried no such attribute and no alias could supply one, since an alias only redirects to a key that already holds the value. tree.Find was therefore always nil for exactly the selectors examples/folio-web tags with. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): name an element only when the selector names it alone ax.findAll stamped every result with the query selector, and resolveCoordinates prefers the tree lookup over the element's own coordinates, so N sibling candidates all executed on the first match. On folio's Home screen the fuzzer could never open any account but the first. The gate tests identity rather than cardinality: no node other than this one answers to the rendered string, checked with the same lookup the runner runs. A rendered object selector can resolve somewhere the query never matched, so counting the query would call that unique. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(web-runtime): hold the V8 host to the same naming gate elementHandle stamped the query selector on every result the same way, so the merge carried the sibling collision onto web for authored ax targets. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): sibling taps reach the driver at their own coordinates Drives 40 real draws from a spec that taps each card, through the picker, the serializer and DecodeAction, and asserts on the points the driver saw. Against the shared-selector bug all 40 landed on the first card. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): an ambiguous name loses to the coordinates it was built from Attribute values match by substring, so a selector that named one element where the candidate was built can name several in the tree it resolves against, and the lookup sent every one of them to the first match. The host gates blank an ambiguous tag at enumeration time; this closes the gap between that moment and the action. A bare-string target carries no coordinates, so the first match stays the answer there rather than dropping an authored action. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): record an element-valued extractor instead of dropping it An element carries find/findAll host functions, so json.Marshal refused the whole value and the encoder answered nil. ChangedExtractors then emitted no entry: no error, no warning, no value. Project the value the way the web host already does (functions dropped, cycles and over-deep branches null, non-finite numbers null) and turn whatever is still beyond JSON into an error the author sees, rather than a missing extractor. * test(verifier): an unrecordable extractor value is reported, not dropped * test(runner): element-valued extractors reach the trace * docs(spec-language): say what a trace records for an element-valued extractor * docs(claude): add delegation and record-keeping sections delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository. * feat(driver): declare undelivered-action errors and three optional capabilities ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner. * feat(driver): add escape to the pressKey surface escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all. * fix(ios): refuse a gesture the screen has no surface under the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error. * feat(ios): derive scrollable from the snapshot's tree depth the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one. * fix(sidecar): stop dropping gestures, selectors and keys in silence a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing. * fix(sidecar): map the driver's refusals onto the gesture errors OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak. * fix(selectors): resolve text to the innermost match and scan the root in both forms an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree. * feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump. * fix(chrome): emit every markup attribute and read checked and selected off the property the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever. * fix(chrome): scroll a gesture point into view and dispatch trusted input getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires. * feat(chrome): read the page's exceptions and navigations, and hold the picker state across them a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision. * feat(trace): version each step and record its logs, exceptions and navigations a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report. * feat(runner): bound every device call and record the actions that never reached the app observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one. * feat(verifier): expose extractor names and rebuilt property formulas an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does. * feat(testrun): expose the seeded bundle a run loaded BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity. * feat(tracecorpus): load recorded runs for offline measures reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl. * refactor(seedspec): move seed spec parsing out of the campaign command the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change. * feat(analyze): time an event at the step it was detected and report the quartiles an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median. * feat(analyze): add the seed-paired signed-rank comparison and record the holm family --paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct. * test(analyze): recover planted effects through the tool's own entry point a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk. * feat(label-coverage): report the addressable share of an app's interactive surface reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression. * feat(exploration-reach): count the distinct structural states a stored run visited the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay. * feat(defect-identity): count distinct defects across stored runs a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed. * feat(oracle-reduction): replay stored traces under four reduced oracles re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding. * feat(implementation-sweep): run one campaign against every implementation of a requirement installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed. * feat(corpus-sweep): run one specification against a served corpus of implementations same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them. * docs(manual): document innermost text matching, escape and the web scroll verb text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what. * test(browser): assert an uncaught page exception reaches the trace the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary. * feat(trace): a step can name the precondition it could not meet A step that never had the app under test in front of it observed something else, and nothing in the trace said so. Index 0 carries the startup gate's verdict, so a run that never started is a trace holding that record and nothing else rather than a run that explored and found nothing. * fix(runner): budget the foreground gate in time, not in polls Eight polls is not a budget. Each poll costs whatever the driver's idle wait happens to take, so the same launch cleared the gate on one device and exhausted it on another: across 80 runs of one app, the gate reported "app never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16 runs, and it was wrong every time. On API 34 settleForForeground returned in ~100ms, so the eight polls gave up 1.2s into a launch whose window drew at ~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14 runs then spent their first step on the launch animation instead of the app, which is the one-step offset that came out of that campaign looking like a platform difference. The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same duration on every device, and a verdict of "not in front" ends the run instead of warning and carrying on: a run that never got its app on screen holds no evidence about the app, and the trace records why at step 0. * test(runner): the gate keeps looking until its budget runs out Locks the three facts the campaign was missing: a window that draws after more polls than the old count allowed still clears the gate, an app that never comes forward ends the run with a typed error, and both the startup verdict and every mid-run step the guard could not recover are readable off trace.jsonl. * feat(campaign): count the runs that were never in the app A run that failed its precondition has zero steps and no violations, which is what a short clean run looks like too. The summary now counts the trace records naming an unmet precondition, so a campaign directory answers "how many of these were never in the app" without grepping any log. * docs(triage): name the trace field a run that never started leaves * fix(selectors): tag names the whole tag, not a substring of it matchSelectorKind had no case for tag, so it fell through to the raw attribute path and matched by substring. web-runtime.ts compiles tag to a CSS type selector, so tag:li resolved to <todo-list> on the Go side and to nothing on the web side. * test(chrome): both resolvers agree on tag where a container's name contains its child's * fix(make): build the binary instead of matching the build directory build/ exists at the repo root, so make build was satisfied by the directory and left a stale bin/sanderling in place. * feat(verifier): expose the property names a loaded spec registered * feat(testrun): refuse a run against a spec that registers no properties A spec with no properties drove the app and reported no violations, which is indistinguishable from a spec that judged something and found nothing. Execute now aborts after loading the spec unless the run asks for the opt-out by name. * feat(cli): --allow-no-properties opts a run out of the refusal * docs(cli): document --allow-no-properties * feat(bundle-check): fail a spec that bundles but registers no properties * test(bundle-check): cover the zero-property refusal and pin the reported bundle * feat(folio-web): predicates for counting commits against submit actions * feat(folio-web): judge one commit per submit over a home-card window Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta, which compared two consecutive steps on one screen and so could not see a double submission that lands across a navigation. * fix(folio-web): keep submit live for 400ms after saving Defers the navigation back so the button is tappable while the label reads Saved, widening the double-submit window the counting property is there to catch. * feat(confusion-matrix): score the checker against a blind reviewer Cross-tabulates the properties that fired against the human verdict, one cell per implementation, over a sweep whose implementations all passed their own generated tests. An implementation that failed to build, has no usable run, or carries no filed verdict is listed as missing data rather than counted as a clean cell. Landing the package in one commit because the intermediate splits would not link. * test(confusion-matrix): reject malformed inputs and keep missing data out of the cells * test(confusion-matrix): cover cell assignment, precision and recall * fix(chrome): focus descends into the shadow root document.activeElement names the host, not the node focused inside it, so a Compose-for-wasm app that mounts its tree in a shadow root reported focus on div#app forever. confirmFocus could never be satisfied and every InputText step aborted the run after three tries. selectAllScript already descends the boundary; the tree builder did not. * test(implementation-sweep): supply the binaries the missing-binary test does not test resolveBinaries ranges a map, so with more than one binary absent the error named whichever it reached first. The test passed locally only because bun and sanderling were on PATH; on CI it was a three-way coin flip. * fix(replay-ui): read data-* attributes by their markup names The web runtime now publishes raw markup attribute names, so attrs["step"] read nothing where the markup writes data-step. Three properties went vacuous and exactlyOneStepIsSelected reported false against a UI that was fine. The test also fails if a dataOf key gains no matching attribute, or if an attribute it derives is rendered nowhere. * fix(web-runtime): focus descends into the shadow root here too The Go driver already descends the boundary; the V8 host did not, so the two enumerations disagreed about focus on any shadow-mounted app. The harness now answers activeElement the way a real root does: a root names a node of its own tree, so only the shadow root itself names the field. * fix(implementation-sweep): name every missing binary, in flag order Ranging a map returned at the first failure, so an operator missing three binaries was told about one, fixed it, reran, and was told about the next. The function exists to stop the sweep once rather than fail per implementation and seed. Two identical runs also printed different errors, which is why this reached master as a flake instead of a clean red. * fix(chrome): focus follows the caret to the field it types into Compose for wasm never focuses the semantics node carrying the testTag. It proxies keystrokes through a hidden 1px backing input that is a sibling of the a11y tree, so the node the runner tapped never held focus and confirmFocus refused to type into every Compose text field. Focus is re-attributed to the smallest editable whose box holds the caret's centre. Centre-point rather than full containment because the caret's height comes from the text style and the field's from its layout box, so a taller font would silently drop back to refusing. * fix(corpus-sweep): name every missing binary, in flag order Same map-ranging bug as the sibling tool, and this copy had no test on the missing-binary path at all. * fix(web-runtime): a handle answers editable for itself, not its container isContentEditable is inherited, so every span inside a contenteditable div called itself typeable. collectTargets and the chrome dump both require the element itself to match; the handle was the one that did not. * test(chrome): a hinted field is not named by its css class The fixture inputs carried no class at all, so the test could not fail the way the bug did. They now carry folio-web-shaped classes, and the test asserts the editable gate the hint is read behind. * test(chrome): the handle and the enumeration agree on editable too The helper compared clickable alone, so the inherited-contenteditable bug was caught by unit test only and never in a real browser. * fix(web-runtime): focus follows the caret to the field it types into Mirrors the driver, so the two hosts agree about focus on a Compose page. The harness inherits custom properties down the parent chain the way CSS does, so an implementation matching the inline style attribute fails. * fix(campaign): name every missing required flag, in flag order Five required flags ranged as a map, so omitting three told the operator about one, chosen at random. * fix(corpus-sweep): name every missing required flag, in flag order * fix(implementation-sweep): name every missing required flag, in flag order * fix(confusion-matrix): name every missing required flag, in flag order * ci: pin the idb-companion tap to the formula the companion is staged from The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and prepare.sh stages bin/ and Frameworks/ as siblings because the binary resolves through @rpath. Floating on it also made the hard-coded companion-1.1.8 output name a lie. The ios-assets cache does not cover this: it restores and make rebuilds anyway, because checkout stamps prepare.sh newer than the archived tarball. Master was green only because its last run predated the bump. * fix(campaign): refuse to start on a device that is not there A sweep launched at six serials, three of which had been deleted from the host. 19 of 20 runs were lost, and not because half the devices were wrong: a worker on a dead serial fails in about 31 seconds and immediately pulls another seed, so three bad workers drained sixteen seeds while the three good workers were still inside their first run. Fast failure is more dangerous than slow failure, because the fast failure consumes the resource the slow one would have left alone. Preflight names every missing serial before the first seed is dispatched. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): quarantine a device that keeps failing fast Preflight cannot catch a device that disappears mid-sweep, which is what happened: the serials were alive the previous day. Three consecutive failures under two minutes, with no run that worked in between, is a property of the device and not a coincidence. The manifest records which device was quarantined and which seeds have no result, so an aborted sweep says so in its own artefact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record the device a run executed on meta.json carried the host but not the device, so a trace could not say what hardware produced it without the campaign manifest beside it. An experiment splitting cells across api levels could only join them through that manifest. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * ci: let a restored ios bundle survive make's mtime check The cache restored and the build ran anyway: a restored tarball keeps the mtime it was archived with while checkout stamps the sources, so make read every bundle as stale. Both logged Cache hit and rebuilt regardless. Dating the bundles after their sources fixes the lie where it is told. Order-only prerequisites would have fixed it in make, but a laptop has no cache key, so editing prepare.sh would silently embed the previous tarball. The formula version joins the key because a hit now decides what gets embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2 runs shared a key. * fix(confusion-matrix): a campaign that died is missing data, not a true negative The sweep-level loop excluded a run on launch_error alone, while excludedBecause already checked the campaign process's exit code. An interrupted campaign wrote exit_code -1 with an empty launch_error, so its one completed seed scored the implementation as a clean cell on a tenth of the planned evidence. The fixture builder wrote one exit code into both the sweep record and the campaign run record, which is why no test could tell the two levels apart. * fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets A run stops at whichever comes first, the step budget or --duration, so a clean run that reached the wall clock exited with fewer steps than the budget and was still credited with the whole of it. The model arm pays a network call and a screenshot per step, so it reaches the wall sooner and was handed exposure it never had. Nothing checked that two arms shared a budget either. Thirty identical clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14 from the rank-sum while the log-rank in the same report read p 1.0000. groupArms already refused this within one arm. The claims the old convention left in comments and report lines are corrected rather than left standing beside the new behaviour. * fix(runner): a source that was asked and handed nothing says so NextAction returning ErrNoAction left the step with no skip reason, so a run whose every model call failed on transport, a non-2xx, an empty choices array or an echo mismatch printed no violations and exited 0. Only llm-calls.jsonl knew it had never touched the app. The reason now travels the path the other five already take, so it reaches the trace, the summary, and the campaign's dispatched-action exclusion. A held step never asks and keeps carrying nothing. * feat(testrun): refuse a run that dispatched none of its actions Same argument as the zero-property refusal: an instrument that drove nothing must not report a clean result. A first-screen violation still wins under --exit-on-violation, --allow-no-properties exempts the extraction sweeps that measure reach rather than judge, and one dispatched action is enough, so a generator quiet on some screens is untouched. * docs(cli): document --label-source * docs(spec-language): name the hintText selector's host divergence The line said the key matches placeholder alone, which is true of the web runtime and not of the tree, where it resolves against the derived attribute. A spec author reading it wrote a selector that matched on one host and not the other. * feat(bundle-check): --allow-no-properties opts out of the refusal The run path grew the opt-out and the freeze gate did not, so a spec the extraction and portability sweeps register nothing for on purpose could be run but never frozen. The refusal now names the flag the way the runner's does. * test(verifier): an unreadable committed fixture fails, it does not skip The comment said the round trip always runs. A skip on a fixture that is committed turns a missing or truncated file into a green. * fix(testrun): the refusal asks whether the generator drove, not whether anything did A dead provider against folio exited 0 on a real emulator: the login setup dispatched three actions before the generator was consulted, so DispatchedActions was 3 and the gate never fired while the generator drove the app zero times across 83 steps. Any spec with a login setup was immune, which is the normal case. Summary counts generator actions separately and the refusal reads that. NoActionsDispatchedError becomes NoGeneratorActionsError, because a run that dispatched three login taps was lying in the old name. * feat(runner): the summary says how many steps the generator drove A green llm run carried no evidence of how much the generator actually drove: the count was inferable only from llm-calls.jsonl outcomes, and the number the refusal turns on was invisible in the run's own output. * fix(testrun): an ios run records the simulator it executed on Device was read from --device, which only an android run sets, so every ios meta.json left the field empty and the trace could not say what hardware produced it. * fix(campaign): the action count leaves the setup's login out on a model run Defects per thousand actions divided by every dispatched step, so a spec whose setup logs in inflated the denominator by however many steps that took. It is the same error the run gate had, and it does not cancel between arms. A model run is separable because only an llm-selected action stamps next_action.source. A seeded run is not: its setup returns through the same entry with no marker, and 11261 dispatched steps across the 169 recorded runs carry no source at all, so excluding on it blind would report every seeded run as having explored nothing. The seeded arm counts as before and a test pins that. * feat(hierarchy): an element reports whether it masks what is typed into it ios reads it off SecureTextField, which the companion already sent and nothing read; web reads input[type=password]. Android cannot: the native tree mapper drops the password attribute before the sidecar sees it, so the fact is three-valued and null there rather than a false that would read as "not secure". * fix(verifier): a secure field's typed value never reaches the record A folio login run wrote the account email and password in cleartext into llm-calls.jsonl, 166 times in one run, beside screenshots of the same screens. Three sites rendered it: the recent-action memory, the candidate list, and the trace. One helper now covers all three so a fourth cannot bypass it, and the driver still receives the real text. Android redacts every typed value because it cannot tell a secure field from any other. That asymmetry is deliberate and documented: safe by default on the target that cannot tell. * fix(runner): a secure field's value does not reach state.lastAction either folio extracts lastAction, and extractor values are persisted as extractor_changes, so the password still reached the run directory through the spec after the three render sites were closed. The wrap sits in the runner rather than in lastActionFields because the hosts hold the next step's tree, not the one the action was chosen against: a field that stops being secure between the two would publish what the trace withheld. Live and replay now agree byte for byte. * fix(trace): an action names the generator that produced it The setup exclusion landed for the model arm only, because only a model pick stamped a source. A seeded run returned setup's action through the same entry with no marker, so its denominator still counted the login while the model arm's did not, and the two are compared. serializeAction names setup and seeded on the wire, so both arms are counted by one rule. An already-recorded trace names nothing and keeps exactly the count it was reported with; unattributed_actions counts those steps so the old denominator cannot pass as the new one. TraceVersion is deliberately unbumped: oracle-reduction refuses a differing version, and a bump would make all 169 recorded runs unreplayable. * fix(defect-identity): degrade a redacted origin action to its selector The full action key read the typed value straight from the trace, where redaction renders every value typed into one field as the same string, so two runs that typed different values there collapsed into one identity and the report said nothing about it. The key now drops a redacted value, falls back to the selector for that action, and counts the rows it did that to, so the undercount reads as an undercount. * fix(campaign): a record always says how many actions named no producer An omitted count reads the same as a run recorded before actions carried a source, so the two cannot be told apart by anything downstream. * fix(analyze): read how much of a record's action count names no producer A runs.jsonl written before actions named one has no field, and its whole count is of unknown provenance rather than none of it. * fix(analyze): refuse to compare attributed and unattributed denominators One arm's actions may include the login the spec's setup drove and the other's cannot, so a per-action rate over the two divides by different things and the tests rank the bookkeeping. * fix(analyze): mark an action count of unknown provenance in the report * docs(manual): what an action count with no producer means for a rate * fix(folio): install through adb so a remote adb server works Gradle's install task talks to adb through ddmlib, which reads only ANDROID_ADB_SERVER_PORT and dials the loopback address, so ADB_SERVER_SOCKET never reaches it and `just test` could not touch a remote emulator. Gradle now only assembles the APK and adb does the install, which picks up the same server every other call in the run talks to. * docs(folio): say how to point just test at a remote adb server * test(conformance): the g4 fixture holds what a redacted android trace holds Android reports no secure fact for any field, so every InputText it records writes the redaction placeholder rather than the typed value. The fixture still carried the real value, which is the only reason the gate reported itself as catching the doubling. Two more fixtures come with it: a repeated-character corpus value that reads as its own doubling and must not fail, and a backend that does record the typed value. Red at this commit: G4 reports PASS on a doubled field it cannot see. * fix(testrun): a recorded violation outranks the dead-run refusal A campaign never passes --exit-on-violation, so the refusal was discarding runs that had found something: exit_code 1 in the record and the analysis drops them as missing data. A run that recorded a violation holds a verdict, which is the whole reason the refusal exists. * fix(testrun): the dead-run refusal gets its own opt-out --allow-no-properties was waiving two unrelated refusals, so a sweep passing it for the property-free reason silently lost a detector it never asked to disable, and a run with properties could only get the dead-run exemption by claiming one it did not want. * feat(cli): --allow-no-generator-actions The flag the dead-run refusal names, wired through to the pipeline. The property-free flag goes back to meaning what it says. * refactor(analyze): open the log-rank up to a weight on the risk set The log-rank is one member of a family that differs only in how much each event time counts. Nothing else changes: the counts it reports stay counts whatever the weight, and the published-dataset results are unmoved. * feat(analyze): add the gehan generalized wilcoxon test The rank-sum carried over to right-censored samples: every pair of runs is scored by which one outlived the other, and a pair censoring cannot order counts as half rather than as a difference neither run supports. The effect size and the p-value are the same statistic, and with nothing censored both are exactly what the rank-sum reports. * fix(analyze): compare arms on censored runs, not on flattened step counts stepTimes threw the censoring flag away and handed the rank-sum a plain number per run, so a run the wall clock stopped at step 12 was ranked as one that violated at step 12. That was defensible while every clean run sat at the budget, the largest value any run could take, and it stopped being defensible when a clean run started being censored where it stopped. Twenty runs clean at step 12 against twenty violations at step 100 read a12 0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank reading p 1.0000. The pairwise comparison is now the Gehan test over the observations themselves, and the report says how many run pairs censoring left with no order between them, which is how much of the effect size is the null value rather than an observation. * fix(conformance): g4 reads a doubling off the observed field value The typed value stopped reaching the trace on any target that reports no secure fact for the field, which on android is every field, so the gate was comparing the redaction placeholder against itself and passing whatever the driver did. The observed value is not redacted, and a field holding one string twice over is the doubling itself. A value that is a single character repeated stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither can be told apart from its own doubling. The recorded-value check stays for the targets that do record it, where it also catches a doubling appended to content the field already held. * fix(spec): a secure selector names the password field on web secure is derived from the field type, not written by the markup, so matching it as a raw attribute reached nothing: the key is accepted, no unknown-key error fires, and find answered undefined on web for the field it answers with on ios. false is every editable field that is not a password entry, since an element that is no field reports null and answers to neither value. * test(chrome): resolve the secure selector on both matchers the fixture covers the password entry, the three shapes of editable field that are not one, and a checkbox that is no field at all. * test(chrome): compare the secure fact across both producers it is the fourth fact the dump and the web runtime derive independently, and the one that decides whether a typed value is written into the shared record. three-valued, so the fixture guard requires all three states rather than both polarities. * docs(manual): state what a secure selector matches * test(conformance): g4 keeps checking past an input typed at coordinates An InputText that names no field aborts the analyzer, so the gate reports the whole run as failed and checks none of the steps after it. 129 of the 485 recorded traces hold such a step. Red at this commit: jq stops on a null selector and the gate reports FAIL. * fix(conformance): g4 skips an input that names no field jq splits an empty string into no segments, so reading the last one off an action typed at coordinates threw and took the rest of the run's steps with it. Such a step names nothing to check; the gate now passes over it and keeps checking the ones that do. * test(browser): the exit code a dead run and a violated one actually leave Drives the built binary against a page with nothing to tap and reads the process status, then the same run through campaign to pin what lands in runs.jsonl: exit_code 1 there is a detection the analysis drops as missing data. * fix(spec): keep a secure selector valid beside another key a multi-key object selector concatenates its parts into one compound, and a type selector is valid only at the head of one, so {id, secure} built '[id="pwd"]input[type="password"]' and querySelectorAll threw. * fix(analyze): score a seed pair by which run outlived the other The paired path had the same defect as the unpaired one: it subtracted two step counts and handed the differences to the signed-rank test, so a pair holding a run the wall clock stopped at step 12 entered as a difference neither run supports. Twenty seeds where the first arm was still clean at step 12 and the second violated at step 5 in six of them read sign -1 and p 0.0011, pointing at the arm that never violated. A pair is now scored the way the unpaired comparison scores one and tested by the exact sign test over the pairs whose order censoring determines, which is what the log-rank stratified by seed reduces to here. The signed-rank goes with the differences it needed: a magnitude-based paired test wants a difference from every pair, and the arms censor on different clocks. The median difference stays, over the pairs where both runs violated, and says so. * docs(analyze): name the tests the tool actually runs The --paired flag advertised the signed-rank, two comments and a test message still said rank-sum, and nothing said what rankSum is doing in the tree now that no campaign reaches it. * docs(manual): exit 1 also means a run that holds no verdict And the flag the dead-run refusal now names, which --allow-no-properties used to double as. * docs(skills): quote the summary line the runner prints now The setup skill's empty-page claim was the stale one that mattered: that run records no_action_produced on every step and exits 1, it does not sit at exit 0 with no violations. Numbers remeasured against the counter and throwing fixtures. * test(conformance): g4 sees a doubling appended to what the field held Redaction cost the gate this shape on android: the driver typed the value twice onto existing content, so the whole value is not its own doubling and the typed value is not in the trace to compare against. The recorded-value check still catches it on the backends that record one. Red at this commit: G4 reports PASS on a field that grew by one string twice. * refactor(analyze): hoist the sign test's loop bound * fix(conformance): g4 reads a doubling out of what the field grew by The whole-value check misses a driver that typed the value twice onto content the field already held, which is the append-vs-replace shape the recorded value used to catch before it was redacted. What the field grew by over the snapshot the action was chosen against is the same signal and needs no typed value. Checked against every recorded trace under conformance/runs: 485 traces, 299 of them carrying an InputText, none newly failing. * fix(analyze): write an undefined paired p-value as null, not as NaN A paired contrast where censoring orders no pair has no p-value, and JSON has no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN' and wrote no summary at all after printing a complete report. The two fields join the medians and the rates already carried as pointers, undefined reading as null in the summary and n/a in the report. Reachable since a clean run started being censored where it stopped: an arm the wall clock stops before its partner ever violates orders nothing. * fix(spec): a boolean state selector names what the live element reports clickable, enabled, focused, checked and selected are derived from the element rather than written by the markup, so matching them as raw attributes built [clickable="true"] and reached nothing: the keys are accepted, no unknown-key error fires, and the worked example in docs/manual/spec-language.md found no element on web and passed having checked nothing. Each key is answered by the same function elementHandle derives the fact with, since no CSS says what any of them says: :focus names the shadow host of a focused field as well, :checked misses a checked custom element and answers for a selected option besides, and [checked] is the state the page loaded with rather than the one the user left it in. * fix(spec): keep a tag selector valid beside another key a multi-key object selector concatenates its parts into one compound, and a type selector is valid only at the head of one, so {id, tag} built '[id="amount"]input' and querySelectorAll threw. whether a spec got an exception or an element depended on the order its author wrote the keys in. * fix(chrome): state every boolean flag the dump can state internal/hierarchy writes the attribute a selector matches on only where the producer stated the flag, so a state emitted as null is one no selector can ask about: {clickable: false} and {enabled: false} matched nothing at all against a web dump while matching on android, which states every flag both ways. only secure stays three-valued. * test(chrome): resolve the five state selectors on both matchers the fixture differs one state at a time: a disabled button and an aria-disabled role control, a box ticked by script with no checked attribute beside one cleared by script that has it, and a select whose first option is selected without the markup saying so anywhere. half the states are asked inside one container, because a state the whole page has an opinion about answers with most of the document and a want list nobody can check by reading. * test(chrome): compare checked, selected and focused across both producers the target enumeration carries none of the three, so they reach a spec through the ax handle alone, and a selector naming one of them resolves against that same reading. the shadow fixture holds the focused control inside its shadow root, where document.activeElement names the mount element and only a producer that descends finds the field. * fix(spec): keep a selector out of the head subtree the head renders nothing, so the hierarchy dump drops it and so does the enumeration the picker walks, but a selector still resolved into it: a whole-page findAll answered with <head> and <title> here and with neither on the goja host, which is a divergence the moment a state selector asks a question every element has an answer to. * docs(manual): state what the other boolean state selectors match * fix(folio): refuse to install and fuzz a device nobody named adb falls through to the local server when ADB_SERVER_SOCKET is unset, and claims the only device attached there. That could be a personal handset, and a run installs the app, clears its state and fuzzes it. Every recipe that touches a device now resolves the target through _require-device, which only picks on its own when a single local emulator is all adb sees. * docs(folio): state that android recipes need ANDROID_DEVICE * fix(spec): and text with the keys written beside it a compound object selector dropped text and matched on the other keys alone, so {testTag: "Row", text: "Alice"} selected every row carrying the tag where internal/hierarchy selects the one row the author named. matching more than the spec said is silent: the find lands on a row nobody wrote and every property over it still passes. text is answered against the element the way the boolean states are, since css cannot ask what an element's text says and the xpath that can cannot ask about the rest, and the innermost rule now holds over what the whole selector matched, where internal/hierarchy holds it. a text-only selector still compiles to the same innermost xpath. * test(spec): pin text against the key beside it in either order object keys iterate in insertion order, so the order the author wrote them in decided what a compound selector meant. the innermost rule is pinned over the whole selector's matches: a row whose badge carries the class and the text both is dropped, one whose badge carries the text alone is kept, and a state key is anded before either. * test(chrome): compare a compound text selector across both matchers one page, both resolvers, text written before and after the key beside it. the object form now encodes its keys in the order the filters state them rather than the order a map iterates, so both orders are asked. the row and the badge under it share a class so the innermost rule has something to drop, and {text, clickable} pins that text is anded before that rule runs: the innermost element carrying "January" is the option, and the select is the only element that is both. * docs(manual): state how text combines with the key beside it the object selector section said every pair must match without saying where the innermost rule then lands. * fix(hierarchy): reach the class attribute through className className is an accepted selector key that no producer writes: android reports the view class, ios the element type and the chrome dump el.className, all of them under `class`. With no alias onto that key the selector matched NOTHING here on every platform while the web runtime resolved it against the live DOM, so {className: "status"} named the row and the badge on one host and no element at all on the other. The failure is silent: the key is accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. * test(chrome): compare className across both matchers one page, both resolvers, the two names for the one attribute. class is asked beside className so the pair is pinned to the same elements rather than each to itself: the row and the badge under it both carry it. * test(spec): pin className and class on the same elements this host answers both names against the live DOM and internal/hierarchy now aliases the second onto the first, so a name dropped from the table here would match nothing on web while the dump still answers it. * docs(manual): list className among the cross-platform aliases the key was already typed on the spec surface and already resolved on web, and the alias table said nothing about which attribute it reads. * fix(hierarchy): reach the accessible label through every name for it label and accessibilityLabel aliased onto accessibilityText alone, which only the ios sidecar writes, and alias expansion is ONE level: the hop from accessibilityText to content-desc was never taken, so both keys matched nothing on android and on the chrome dump, which write the fact under content-desc. ariaLabel and contentDescription aliased onto nothing at all and matched nothing anywhere. The web runtime resolves all four against the live DOM, so a selector naming a field this way found it on one host and no element at all on the other. The keys are accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. Each name lists both keys rather than chaining through accessibilityText: transitive expansion would silently widen every existing key at once. * fix(hierarchy): reach a web test tag through testTag and testID Compose for Web writes a test tag as data-testid, which is what the web runtime resolves both names against. testTag aliased onto the three identifier keys and not that one, and testID aliased onto nothing at all, so a tag the web runtime found on every row of a list named no element here and every property over it passed vacuously. * fix(spec): resolve the identifier, label and class aliases against the DOM identifier, accessibilityIdentifier, accessibilityText and elementType are the names ios writes four facts under, and internal/hierarchy aliases each onto the key the other producers write. This table listed none of them, so each fell through to a raw attribute lookup and built [accessibilityIdentifier="summary_card"], which no element carries. Every one of them resolved against the dump on the goja host and named nothing here. The keys are accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. * fix(spec): an editable or scrollable selector names what this host derives Both facts are derived from the live element rather than written by the markup, and matching them as attributes built [editable="true"], which no page carries. Both resolve against the dump on the goja host, so a spec naming a field or a scroll container that way found it there and no element at all here, with no unknown-key error to say so. Each reads the same function the fact is derived with, so a selector cannot name an element this host calls something else: the handle, the picker's target list and the editable selector all go through isEditable, and scrollable reads the overflow test collectTargets reads. scrollable false names nothing rather than every element that does not scroll: both producers state the fact only where it holds, the way an element that is no field at all answers to neither value of secure. * test(chrome): compare the alias keys and the two derived facts one page, both resolvers, the ten names that resolved on one host only. each alias is asked beside the key it resolves through, so the pair is pinned to the same elements rather than each to itself. the page grows a container that overflows its box and a neighbour that does not, because scrollable is derived from the box: without one the only scrolling element on the page is the document root, whose answer moves with the window. * fix(hierarchy): bounds is a raw attribute, not a cross-platform key Every native dump writes the rectangle out as a string under bounds, and no DOM element carries an attribute of that name, so the key resolved against the dump and matched nothing on web on every page there is. It is accepted, so no unknown-key error said so, and no mapping can be invented for it: there is no DOM fact to map it to. Off the accepted list the web runtime raises the unknown-key error instead of matching nothing in silence, and the key still resolves wherever a producer writes it, through the escape hatch every other raw attribute already uses: a key some element carries is a key that can match, on both sides. * docs(manual): state which attribute each alias reads, and what bounds is the table listed neither name for the accessible label that a web page writes, nor the key a web test tag lands on, and said nothing about elementType. editable and scrollable are boolean states like the rest, and scrollable is the one of them the platforms state only where it holds. bounds is a raw driver attribute rather than an accepted key. * docs(hierarchy): the package doc names every key an alias reaches it described the alias table as it stood before the label and test-tag names reached the keys android and web write, and said nothing about expansion being one level deep, which is why each name has to list every key rather than hop through another alias. * fix(spec): a hint selector names the ladder both producers derive hintText and placeholderValue are the accessible-name ladder, derived from the live element, and compiling them to [placeholder="..."] made them name the wrong field or none at all. A field labelled by an aria-label or a bound <label> carries no placeholder, so it resolved against the dump on the goja host and reached nothing here; one carrying both answered to its placeholder here where the dump answers to its aria-label, which lands a find on an element nobody named. Both keys read the same fieldHint elementHandle and the hierarchy dump (internal/driver/chrome/driver.go) derive the fact with, so a selector cannot name a field this host calls something else. An empty hint names nothing rather than everything that is no field: both producers write the fact only where the ladder answered. placeholder stays the attribute the markup writes, which is what the dump carries under that name too, so a field whose hint is something else still answers to it on both hosts. * fix(chrome): a hint target is not tapped by the placeholder attribute TapSelector is a third resolver, and it built [placeholder="..."] for hintText and placeholderValue too. Now that both matchers read the accessible-name ladder, that CSS names a field whose hint is its aria-label and whose placeholder happens to carry the value, which is an element neither matcher named. No CSS says what the ladder says, so both keys fall through to a match that reaches nothing and the step fails naming the selector, the way every other derived key in this file already does. A selector reaches here only where the dump resolved it to no coordinates at all. * test(chrome): compare the hint keys and placeholder across both matchers The page gains four fields that differ one rung at a time: a bound label, a placeholder, a placeholder an aria-label outranks, and the name the form gives the field. Only the placeholder rung was reachable before, so hintText and placeholderValue named a field on the goja host and no element at all on web for the other three, and named the field here and nothing there for the rung the ladder passed over. placeholder was measured empty on both hosts because nothing on the page carried the attribute, which said nothing about it. It now names the field the markup wrote it on and not the field whose hint is its aria-label. The third resolver reads the same selectors: what TranslateStringSelector builds for a hint key has to match nothing over CDP rather than the field carrying the value as a placeholder. * docs(manual): both hosts read the hint ladder, placeholder is the attribute The web section said the hintText key does not read the ladder on both hosts and told authors to select such a field by attrs.hintText instead. Both hosts read it now, so that instruction is gone rather than left standing beside a newer sentence. placeholder is stated as the attribute the markup writes and nothing more, the tap path is stated as failing by name where no CSS says what the ladder says, and the alias table gains the row it was missing. |
||
|
|
76dce1a75e |
experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs runner.Options.MaxSteps already worked but was unreachable from the command line. A step budget is what makes two generators comparable: one making a model call per step and one drawing from a PRNG are not comparable per second. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record arm membership and host in meta.json meta.json recorded the seed but not which picker ran, how it was configured, what budget it was given, or which machine produced it. A directory of runs cannot be attributed to an experiment cell without those, which makes any factorial computed from such a directory unanalysable after the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(cli): add --arm and populate run meta from it Model and instructions are recorded only when the LLM picker is the one that will actually run, so a spec declaring generator = llm() that is run under the seeded picker does not label its trace with a model it never called. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): sweep seeds for one experiment cell campaign.json lists the seeds a sweep intended to run and is written before the first run, so a host that dropped runs shows up as missing seeds rather than as a smaller sample. Seed 0 is rejected: sanderling test reads it as "derive a seed from the clock", which is why conformance/gates.sh controls nothing today. Each run contributes one runs.jsonl line carrying steps to first violation by origin step, the step that armed the failed obligation, so the survival analysis never reopens a trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): no silent generator fallback, and llm on web --generator llm against a spec declaring no generator = llm(...) logged a warning and ran the seeded picker. For a comparison campaign that is silent arm corruption: the run completes, the directory looks correct, and the wrong policy drove it. It is now fatal. pickSources also returned the V8 source for both action and extractor on web before it looked at the generator, so the llm policy was unreachable there. The two axes are now independent: the driver picks the extractor source, the flag picks the action source, and llmSource composes with either because the runner populates the candidate list and screenshot on every platform. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): make the hierarchy dump agree with the web runtime Three facts differed between the dump the goja host reads and the DOM the V8 host reads, so the two enumerated different candidates on one page. scrollable was never emitted, and worker.go reads exactly that attribute while targets.ts requires it for scrolls, so the goja host could not offer a single web scroll. clickable tested el.onclick, which React assigns to its root container for event delegation, making the whole viewport a tap target here and in no other enumeration. Both now resolve through the selector sets in pkg/spec/src/web-runtime.ts. The dump also rooted at body while collectTargets walks querySelectorAll("*"), so the goja host never saw html, where page-level scrolling lives. It now roots at documentElement and skips the head subtree, which is all zero-bounds and would otherwise carry script and title text into the trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(conformance): give the gate reproducible seeds SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from the clock", so the tunable controlled nothing and a gate failure could not be re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the results table so a failing row names its stream. The five runs stay on five different streams: a gate that scored one path five times would catch less than one that scores five. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): emit editable as a plain boolean editable was emitted as `isEditable || null`, and an absent field sends internal/hierarchy into the native fallback, which reads any class name containing "EditText" as an Android text widget. On web that is just a CSS class, so a page styling a div with it was editable to the goja host and not to the web runtime, and the model policy could be offered typing into a div. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): leave the head subtree out of the web target walk collectTargets walked querySelectorAll("*") while the hierarchy dump skips head, so the two hosts enumerated different element sets on every page with a <head>. No candidate changes: builtinCandidates pushes only for targets acceptsTarget admits, and head elements have no positive bounds, so the list the draw ranges over is untouched. What changes is that targetIndex now means the same thing on both hosts. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(chrome): compare the facts both hosts derive from one DOM The existing parity harness hand-authors the facts on both sides, so it proves that given identical facts both hosts select identical candidates, and says nothing about the two code paths that derive those facts from a real page. Four divergences lived in that blind spot and it passed throughout. This drives one real page and compares clickable, enabled, editable, scrollable and positiveBounds element by element, plus the element sets themselves, which is what catches a host that omits html or includes head. Reverting any of the four fixes makes it fail naming the element and the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * chore(make): run the browser packages one at a time Both launch Chrome and launching two at once has failed with "Launch: context canceled". Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style: remove every em-dash and en-dash Eighteen occurrences across fourteen files. Each sentence was repunctuated to suit what the dash was doing rather than swapped for a hyphen, which produces comma splices. The minus sign in folio-web's ledger is a minus sign and stays. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): honor the caller context in Launch Launch and clearState ran against d.tabCtx, so a target that accepts the connection and never answers wedged the process past its own --duration and through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for the rest of the sweep with no diagnostic. The browser is still allocated against d.tabCtx first, because chromedp starts Chrome under whichever context calls Run first and allocating under a caller deadline would kill the browser when Launch returns. Everything after allocation goes through runCtx. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(sidecarassets): publish the extracted jar through a rename Extract wrote a 96 MB jar with a plain WriteFile into a temp path every sanderling process on the host shares. On a cold host several concurrent workers all miss the checksum and all write the same path, and O_TRUNC lets one spawn a JVM against another's half-written archive. A fresh experiment host is exactly a cold host. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): kill a run that outlives --run-timeout A wedged run holds its worker for the rest of the sweep, and on an unattended host nothing else will send it a signal. Defaults to three times --duration and must exceed it. A killed run is recorded as timed_out rather than as a generic failure, so the analysis can tell a lost cell from a real crash. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style(test): gofmt browser_test.go Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX |
||
|
|
6b0d6cb971 |
WIP: Drive physical Android devices over USB (#67)
* feat(sidecar): reach USB devices via the adb server by serial * feat(test): add --device flag to target a specific Android device by serial * feat(folio): select Android device via ANDROID_DEVICE in justfile * feat(conformance): add android backend to the gate suite * feat(android): keep device awake and unlocked so the app stays foreground * feat(conformance): prep physical android device (autofill/verifier/stayon) * fix(android): make device prep best-effort so OEM-blocked commands don't abort the run * fix(verifier): require positive bounds for swipe candidates A zero-bounds element centers at (0,0); a downward swipe from the top-left corner is the system gesture that pulls down the notification shade, dragging the fuzzer out of the app. Swipes now require positive bounds like every other verb. * fix(runner): harden app-scope guard against launcher and overlays The per-step guard now relaunches and waits until the app window is actually drawn before proceeding, so a slow physical-device relaunch no longer lets an observe or action land on the launcher. It also detects a system overlay (notification shade) stealing window focus while the app stays resumed, and dismisses it with back. * feat(android): harden physical-device runs in device prep Device prep now disables the AOSP cached-app freezer, phantom-process killer, and Doze (and exempts the driver) so OEM background management stops suspending the driver mid-run. Adds ReinstallApp for clear-state on ROMs that deny pm clear, and teaches focus detection to report the notification shade as systemui so the scope guard can dismiss it. * feat(driver): clear-state via APK reinstall when pm clear is blocked When an APK path is set, Android clear-state resets the app by uninstalling and reinstalling instead of asking the sidecar to pm clear, which hardened OEM builds (ColorOS) deny even to the adb shell user. Falls back to the sidecar clear path when no APK path is provided. * feat(cli): add --android-app-path for clear-state reinstall Wires the APK path from the test command through to the sidecar client so Android clear-state can reset apps on OEM builds that deny pm clear. * chore(folio): pass --android-app-path in just test * fix(runner): clamp swipe/scroll origin out of edge gesture zones A gesture starting in the top status-bar strip pulls down the notification shade; the bottom and side strips are the home and back gestures. Any of them drags the fuzzer out of the app. Swipe and scroll origins are now clamped into a safe inner area sized from the maximum element extent (the Android hierarchy root reports zero bounds, so the extent is the reliable screen size). Calibrated on device: origins below ~7% of height no longer open the shade. * perf(sidecar): faster Android text input and drop redundant settle poll inputText now uses adb `input text` for short shell-safe ASCII (~5x faster than the driver's per-character path) and falls back to the driver for unicode, injection payloads, and overflow-length strings. waitForIdle drops the structural-hash poll that followed waitForAppToSettle: each hierarchy fetch is ~500ms on a physical device, so it cost ~2.8s per mutating step for marginal benefit, and the runner already re-fetches transitional frames. Cuts p95 step latency from ~6.5s to ~5.1s; G1-G4 still pass. * fix(verifier): exclude soft-keyboard region from action candidates The fuzzer was tapping Gboard's "Settings" key, navigating out of the app. That key is a bare FrameLayout with a content-desc and no package or resource-id, so the package-based scope filter missed it. Candidates whose center falls in the keyboard region (derived from the IME elements' bounds) are now dropped, so no tap or long-press lands on a key. Opt-in with app scoping; unscoped runs keep every node. * perf(runner): replace focus-tap settle with a brief wait The full WaitForIdle after a field-focus tap cost ~0.5-1s per InputText step on a physical device while the keyboard animated in. The tap registers focus immediately and text is injected into the focused view, so a short fixed wait suffices. Drops p95 step latency ~5.1s to ~4.0s; G1-G4 stay green. * chore(conformance): platform-aware G5 p95 budget for android The 2500ms ceiling was calibrated on the iOS simulator. A physical Android device drives every step over USB (snapshot + settle + adb round-trips), so its per-step floor is several times higher; holding it to 2500ms would force removing the settle/retry logic the correctness gates depend on. The android backend now defaults to 4500ms (override with P95_LIMIT_MS); iOS stays 2500. * fix(sidecar): retry maestro android driver startup The maestro Android driver's dadb.open() occasionally misses its startup deadline (its instrumentation host is slow to come up right after a reboot or per-run reinstall), which aborted the whole run. Retry the open a few times with a short backoff so a transient timeout recovers. * chore(conformance): widen android G5 budget to 5500ms Physical-device p95 swung 3209-4612ms across sessions (cold runs right after a reboot are slower). 4500ms was too tight for that jitter; 5500ms covers the observed ceiling with headroom. * web replay fix * feat(android): force 3-button nav during runs to prevent app drift On gesture navigation a fuzzer swipe can trigger swipe-up-home or edge-back and fling the app off screen. Device-prep now switches to 3-button navigation for the run (no edge gestures; the nav bar's buttons are systemui-owned and already excluded from action candidates) and restores the original navigation mode when the run ends. Best effort: leaves nav untouched if the overlay command is unavailable. * fix(android): target the selected device in adb reads; don't strand nav mode Review fixes: - ForegroundPackage/FocusedWindowPackage now take a serial and pass -s, so the foreground/scope guard works when several devices are attached (the --device path). Previously they ran bare `adb shell`, which errors with multiple devices, silently disabling app-scope enforcement. The sidecar client passes its serial through. - Extract an adbArgs helper and route every adb call through it, removing four duplicated serial-arg builders. - ForceThreeButtonNav now decides what to restore before changing anything: if the current mode is unknown or already 3-button it leaves nav untouched, instead of switching and then stranding the device in 3-button. Logic split into the pure navModeToRestore, now unit tested. * fix(runner): restore scrollBounds doc; cover destination clamp and screenBounds Review fixes: move the scrollBounds doc comment back onto scrollBounds (it was stranded above screenBounds by an insertion). Extend the clamp test to assert an off-screen destination is clamped onto the screen and that the origin lands exactly on the margin. * test(verifier): cover keyboardRegionTop, including the decor-view guard The full-screen IME decor view rejection had no test; removing it left the suite green. Add direct cases: no keyboard -> sentinel, decor view ignored in favor of the real keyboard line, and decor-only -> sentinel. * style(cli): gofmt testOptions field alignment * fix(sidecar): keep a leading dash off the fast input path A value starting with '-' could be read as an option by `adb input text`, so the fast-path regex now requires a non-dash first character; such values fall back to the driver. Also cover the dadb-target branch where a colon precedes a non-numeric port (a USB serial, not host:port). * refactor(verifier): scope action candidates by window ownership Replaces the leaky per-element package check and the keyboard-region Y heuristic with one rule: walk the window tree propagating each node's owning package (empty and the neutral android framework package are transparent); a node is in scope only when no concrete foreign package owns it (the app's own window carries no package on Compose apps) or the owner is the app package. This drops whole foreign windows (soft keyboard, system UI, launcher) AND their empty-package child wrappers -- e.g. a keyboard's 'Settings' key, which the old empty-package-is-in-scope rule admitted and which navigated out of the app. Deletes keyboardRegionTop/isInputMethodElement. * fix(runner): re-check foreground at apply time, skip stale actions ensureForeground runs before observe, but the app can leave between observe and apply (a prior gesture settling late); swipes/keys then fire stale coordinates onto whatever screen is now up. Re-check foreground immediately before applying and, when the app is gone, skip the action and log it (making the escape visible) so the next step's guard relaunches instead. * fix(android): type long ASCII via fast guarded path to stop keystroke escape A 4096-char corpus string exceeded the fast input cap and fell to the per-character driver path, which takes ~120s. During that uninterruptible window focus could leave the app and the remaining keystrokes sprayed into the launcher search box. Route shell-safe ASCII of any length through adb input text, chunked, re-checking the foreground app between chunks and stopping if it changed. * chore: ignore gate artifacts and local scratch files * refactor(runner): narrow gesture clamp to the top shade strip 3-button nav (forced for every run) disables the side back and bottom home gestures at the OS level. On-device probing confirmed side and bottom swipe origins no longer drift, leaving the notification shade as the only edge gesture a swipe can trigger. Clamp only the top strip; keep origin and destination on screen otherwise. * chore(format): add .editorconfig enforcing 80-column limit * chore(format): add prettier config with 80-char printWidth * chore(deps): add prettier devDependency to replay-ui * chore(deps): add prettier devDependency to folio-web * chore(deps): add prettier devDependency to spec package * chore(format): add swift-format config with 80-char lineLength * feat(format): add make fmt targets for per-language 80-col formatting * fix(runner): translate gesture to safe area so near-top scrolls keep direction Clamping the swipe origin to the top margin while leaving the destination on the full screen used two reference frames: a scrollable container pinned in the top strip had its origin pushed past the destination, reversing the gesture. Translate the whole from->to segment down by the same delta so the origin clears the shade strip without flipping direction. Adds a scroll-near-top test that fails under the old origin-only clamp. * fix(runner): apply-time guard consults focused window, not just resumed activity ensureForeground detects a system overlay (notification shade) owning the focused window while the app stays the resumed activity, but appIsForeground only queried ForegroundApp. A swipe that pulls the shade over the app between observe and apply then fired onto the shade. Mirror the focus check at apply time so the action skips and the next step dismisses the overlay. * test(runner): cover apply-time foreground skip and appIsForeground table Adds a Run-level test asserting no tap reaches the driver while a system overlay holds focus (guards against the skip branch being dead-coded), plus a decision-table test for appIsForeground. Adds ForegroundErr/FocusedWindowErr to the mock driver so the guard's transient-read paths are exercised. * fix(sidecar): harden android driver open, input guard, pressKey, foreground marker - openWithRetry rebuilt a closed AndroidDriver, whose gRPC channel is final and shut down by close(); the retry then ran against a dead channel. Build a fresh driver per attempt and extract a unit-tested retryOpen helper (named DRIVER_OPEN_ATTEMPTS/BACKOFF). - pressKey on the Maestro backend did KEY_MAP[key] (no lowercase, no throw), silently dropping unknown or wrong-case keys; route through a pure maestroKeyFor that lowercases and rejects unknown keys like the Stub contract. - the mid-type foreground guard (typeShellSafe) was untested; extract a pure typeChunks and cover stop-on-foreground-change, always-send-first-chunk, and unknown-owner. - foreground detection required the literal topResumedActivity=ActivityRecord; align parseResumedPackage to the same *ResumedActivity marker set Go reads so OEM wording does not disable the guard. * fix(conformance): pin self-test p95 budget and score install failures as run failures self_test reused the backend-dependent P95_LIMIT_MS, so under BACKEND=android the 4000ms slow fixture rated PASS and the offline analyzer check failed from an env var; pin it to 2500. A per-run adb install failure ran unguarded under set -e and aborted the whole harness; guard it, record the run as a G1 failure, and continue. * fix(android): require --device when several devices are connected With no serial requested and more than one device online, pickDevice silently returned connected[0], but that serial is never threaded into the per-step adb calls, so every later bare adb command failed with "more than one device". Error instead and ask for --device, mirroring pickAVD; a single device stays unambiguous. * refactor(android): move PrepareDevice doc onto it; extract tested wakeCommands The PrepareDevice doc block was stranded above adbArgs, leaving the exported function undocumented under godoc. Move it back and split the wake/keyguard tuples into wakeCommands so they have a unit test. * perf(verifier): memoize scopedElements per tree scopedElements rebuilt a full tree walk plus map on every candidatesForVerb call (~16 per step). Cache the result keyed on lastTree and invalidate it in PushSnapshot. * fix(sidecar): default reinstallApp in SetClearStateReinstall; cover non-android clear Only Dial set reinstallApp, so a Client built another way would nil-deref on Android clear-state. Default it in SetClearStateReinstall too. Add a non-android test so the platform guard has negative coverage: dropping the android check would now fail. * test(runner): make focusTapSettle injectable so apply tests don't sleep 250ms The focus-tap settle was a const, so five InputText apply tests each blocked the full 250ms. Make it a package var and shorten it per-test with cleanup. * refactor(runner,android): drop unused bringToForeground return; grep no-match yields empty bringToForeground's bool return was read by no caller. FocusedWindowPackage's on-device grep exited 1 on no match, surfacing as an error instead of the documented ""; add || true. * perf(sidecar): reuse a single Jackson ObjectMapper structuralHash, countRouteScreens, and hierarchy each built a fresh ObjectMapper per call inside the stability poll; the instance is thread-safe and meant to be reused. Hoist one shared val. * refactor(android): remove unused AdbReverse/AdbReverseRemove No callers anywhere in the tree; they were also the only adb calls bypassing adbArgs. Dead code, removed. * style(runner): trim non-load-bearing comments from this PR's runner code and tests * style(sidecar): trim non-load-bearing comments from this PR's driver code and tests |
||
|
|
90224dfd06 |
Physical-device iOS support (#64) (#66)
* feat(companion): add appState, eraseText, pressKey runner handlers
The Go runner transport already calls these methods; the in-device runner
implemented them only latently. They become load-bearing on the device
path, where the hybrid's legacy-companion fallback is absent. Backward
compatible: the simulator hybrid never calls them.
* feat(ios): resolve physical devices from devicectl
ResolveDevice parses xcrun devicectl list devices into Device{Name,
HardwareUDID, CoreDeviceID}: the hardware UDID feeds xcodebuild/iproxy
and the CoreDevice id feeds devicectl install. Matches by name or either
id; errors list candidates on none/ambiguous. Fixes the stale sidecar
comment on ResolveTarget.
* feat(ioscompanion): runner-only device driver mode
NewDevice reuses Driver with d.companion set to the runner dialed over an
iproxy usbmux tunnel, hybrid=false, runnerClient=nil. The existing accessor
seams then route launch/snapshot/text/gesture to the runner with no new
DeviceDriver methods. Device seams swap clear-state to a devicectl
reinstall, container reset to a warn-once no-op, and paste grant to a no-op.
realSpawnDeviceRunner builds and signs the runner at run time via the App
Store Connect API key (no Xcode UI), caching on a source hash.
* test(ioscompanion): cover device wiring, routing, and shell-out argv
Seam-driven NewDevice wiring + gesture/text routing (asserting no keyboard
HID), devicectl/build/test/iproxy argv builders, xctestrun test-target dict
name parsing, signing-credential env checks, and source-hash cache keying.
* feat(testrun): route physical-device iOS runs to the device driver
Execute resolves a non-simulator iOS target through ios.ResolveDevice into
its hardware UDID and CoreDevice id; buildDriver constructs NewDevice via a
seam instead of rejecting the device. Generalizes the --ios-device and
--ios-app-path help to cover the device path; signing stays env-read, never
a flag.
* feat(doctor): device prereqs replace java/sidecar for ios-device
iosDeviceChecks now verifies devicectl, iproxy on PATH, a connected+paired
device (via ios.ConnectedDevices), and App Store Connect signing creds (via
ioscompanion.VerifyDeviceSigning). The retired JVM sidecar checks stay only
under android.
* feat(conformance): device backend uses iphoneos app and tunnel orphan checks
The device backend now builds via just ios-device, points --ios-app-path at
the Debug-iphoneos bundle, and reinstalls each run for clear-state. The G5
orphan scan replaces the retired sidecar.jar check with lingering iproxy and
device test-without-building sessions (destination platform=iOS,id=).
* feat(folio): device build linking the iosArm64 framework
project.yml selects the Kotlin framework slice by SDK (iosArm64 for
iphoneos, iosSimulatorArm64 for simulator) and links via -framework Shared
on the SDK-conditional search path. New ios-device/test-ios-device recipes
mirror ios/test-ios, signing the Debug-iphoneos build with the .env API key.
* docs(cli): document ios-device doctor checks and the device flags
The --ios-device flag now also selects a connected device; --ios-app-path
covers the device install; the doctor gains an ios-device platform whose
checks are devicectl, iproxy, a paired device, and signing credentials.
Corrects the --clear-data default to true.
* fix(ioscompanion): resolve signing key path to absolute
xcodebuild's -authenticationKeyPath requires an absolute path, but .env
files commonly carry a repo-relative one. Resolve it against the working
directory before the stat so a relative ASC_API_KEY_PATH still signs.
* fix(ioscompanion): re-enable signing for the device runner build
companion/project.yml disables code signing for the simulator build, so
the device build inherited it and produced an unsigned runner that the
device rejected at install (0xe8008018). build-for-testing now forces
CODE_SIGNING_ALLOWED/REQUIRED=YES so automatic provisioning signs it.
* fix(ioscompanion): key the device build cache on signing identity
The cache marker hashed only sources, so switching signing team or key
reused a runner signed with the stale identity, which the device rejects at
install (0xe8008018). Fold team + key id into the cache key so a signing
change forces a rebuild.
* docs(getting-started): document physical iOS device setup
Lists the iproxy requirement and the App Store Connect signing env vars
(SANDERLING_IOS_TEAM, ASC_API_*) a device run needs, plus the
test-ios-device recipe and the doctor check.
* feat(ios): native usbmux client and in-process tunnel forwarder
Talk to macOS usbmuxd directly instead of shelling out to iproxy, so the
device path depends on nothing beyond macOS + Xcode.
* refactor(ios): drive device tunnel via io.Closer seam
Replace the tunnelChild *exec.Cmd and spawnTunnel seam with a tunnel
io.Closer and startTunnel seam backed by the in-process usbmux forwarder.
* refactor(ios): remove iproxy spawn from device runner
* test(ios): cover tunnel close via io.Closer not child process
* feat(doctor): check usbmuxd socket instead of iproxy on PATH
* chore(conformance): drop iproxy orphan check; tunnel is in-process
* docs(ios): device tunnel uses native usbmux, nothing to install
* chore: gitignore the signing keys directory
* feat(folio): add Android launcher icon (black bg, white dot)
* feat(folio): add iOS app icon (black bg, white dot)
* feat(folio): add web favicon (black bg, white dot)
* docs(ioscompanion): fix stale const comments
* refactor(ioscompanion): inline single-use devicectl argv builders
* refactor(ioscompanion): inline xcodegenArgs, drop tautological argv tests
* refactor(ioscompanion): inline firstNonEmpty
* refactor(doctor): dedup usbmuxd socket path via ioscompanion seam
* test(doctor): trim redundant signing-check test
* refactor(ioscompanion): deliver COMPANION_PORT via TEST_RUNNER_ env
* fix(testrun): seam preflight so iOS routing tests pass on CI without xcrun
|
||
|
|
406b7516b3 |
iOS simulator driver: Go-native companion-backed backend (#62)
* perf(ios): use prebuilt XCTest runner to cut startup * chore(ioscompanion): add companion asset prepare script * feat(ioscompanion): embed and extract simulator companion bundle * test(ioscompanion): cover companion stub and embedded extraction * docs: add third party notices for vendored companion * chore: ignore vendored companion bundle artifact * build(proto): pin simulator companion proto v1.1.8 * build(proto): add dedicated buf module and gen template for pinned proto * build(proto): exclude pinned companion proto from root buf workspace * feat(ioscompanion): commit generated companion gRPC stubs * feat(ioscompanion): map flat companion describe dump to TreeNode JSON * test(ioscompanion): add hierarchy-map golden and unit tests * feat(ioscompanion): port screen-settle stability polling to Go * test(ioscompanion): cover settle transitional, hash, streak, and cap rules * feat(ioscompanion): add USB HID keymap module * test(ioscompanion): cover keymap branches and paste-chord constants * build: embed companion assets via withcompanion tag * feat(ioscompanion): add transport companion interface * feat(ioscompanion): add HID event wrapper and builders * feat(ioscompanion): wire gRPC companion client and Dial * test(ioscompanion): cover HID builders and unit conversions * test(ioscompanion): cover Dial, process-state mapping, and install archive * test(ioscompanion): add gated simulator integration smoke test * feat(ioscompanion): text input and gesture HID composition with pasteboard fallback * test(ioscompanion): cover input composers, paste dialog loop, and pure helpers * feat(ioscompanion): add Describe to companion transport * feat(ioscompanion): implement DeviceDriver with companion supervision * test(ioscompanion): unit tests with fake companion transport * test(ioscompanion): gated companion smoke test * feat(ios): add ResolveTarget for simulator vs physical-device routing * feat(testrun): route iOS simulators through the native companion driver * refactor(testrun): defer the java preflight check to the physical-device path * feat(cli): add --ios-app-path flag * feat(doctor): split iOS checks into simulator and physical-device paths * test(folio): add gate-analyzer fixtures for G1-G5 * feat(folio): add iOS conformance gate script * chore(folio): wire gates recipe, app path, and ignore gate output * style: gofmt struct alignment drift * fix(doctor): probe simctl via xcrun instead of PATH lookup * fix(ioscompanion): spawn companion under driver-lifetime context * test(ioscompanion): prove companion child outlives startup context * fix(ioscompanion): chunk install payload under companion message cap * test(ioscompanion): cover install payload chunking * fix(ioscompanion): reinstall via simctl and sanitize companion env * fix(ioscompanion): wait out unresolved accessibility values after launch * perf(ioscompanion): paste long text for atomic landing * test(ioscompanion): cover paste threshold, retry flow, and sentinel detection * fix(ioscompanion): treat unresolved bridge values as transitional, never as content * fix(ioscompanion): accept masked secure-field values as paste landing * test(ioscompanion): cover sentinel mapping and masked-field landing * fix(ioscompanion): atomic erase and single-send paste to prevent doubling * test(ioscompanion): cover atomic erase, single chord, unverifiable field * fix(ioscompanion): verify paste on a time budget that outlasts the bridge blackout * test(ioscompanion): cover bridge-blackout paste verification * fix(ioscompanion): drop unresolved-value settle gate that never let empty-field screens settle * refactor(ioscompanion): name the empty-editable-field sentinel for what it is * perf(ioscompanion): tighten settle streak for the fast companion transport * feat(ioscompanion): pre-grant pasteboard access so unicode input skips the OS prompt * refactor(ioscompanion): drop paste warm-up now that the grant suppresses the prompt * test(ioscompanion): cover pasteboard grant on launch, drop warm-up tests * fix(ioscompanion): retry describe past transient collapsed accessibility dumps * test(ioscompanion): cover collapsed-dump detection * perf(ioscompanion): split raw and retrying describe so settle does not double-wait collapses * perf(ioscompanion): tighten settle now that collapses are handled separately * fix(ioscompanion): replace field content on input so blackout-skipped erase cannot accumulate text * test(ioscompanion): cover replace-on-input and TextReplacer capability * refactor(ioscompanion): neutralize HID events behind the transport seam * feat(companion): add simulator runner project skeleton * feat(companion): serve accessibility snapshots over the wire protocol * feat(companion): synthesize timestamped touch gestures * feat(companion): type text with replace semantics * feat(companion): serve the wire protocol from a parked runner * feat(ioscompanion): add TextEditor capability and unavailable sentinel to the transport seam * feat(ioscompanion): route text input through a text-editing companion when available * fix(companion): bind listener by port and source screen size from snapshot * feat(ioscompanion): add runner companion JSON transport * test(ioscompanion): cover runner transport protocol mapping * fix(companion): synthesize gestures synchronously to avoid the async completion crash * fix(companion): type on the main thread and recover from focus assertions * fix(companion): keep serving after an automation failure * refactor(companion): tidy snapshot serialization * fix(companion): honor sequential tap gaps and survive synthesis exceptions * feat(ioscompanion): expose native typing with an explicit replace flag * chore(companion): add runner asset prepare script * feat(ioscompanion): embed and extract the runner test bundle * test(ioscompanion): cover runner asset extraction * build(ioscompanion): commit runner asset archive * feat(ioscompanion): pair the legacy companion with the in-simulator runner * test(ioscompanion): cover hybrid routing, paste-grant skip, and port binding * fix(ioscompanion): reconnect after interrupted runner calls instead of restarting * fix(ioscompanion): route hybrid lifecycle through the runner and harden restarts * feat(companion): launch and terminate apps through the automation session * build(ioscompanion): refresh runner asset with session lifecycle * fix(ioscompanion): classify connection deadline expiry as caller budget * fix(companion): capture snapshots on the main thread inside the catch bridge * build(ioscompanion): refresh runner asset with main-thread snapshots * perf(ioscompanion): count read spans toward settle and capture snapshots concurrently * feat(ioscompanion): make the hybrid simulator companion the default * test(folio): cover runner-session orphans in the gate harness * test(ioscompanion): pin the child-lifetime test to the legacy path * fix(ioscompanion): keep mappable text on one HID stream and verify unicode clears * fix(ioscompanion): pause the clear chord so selection applies before the delete * fix(companion): prune the keyboard subtree from snapshots * build(ioscompanion): refresh runner asset without keyboard elements * fix(ioscompanion): capture the screenshot transport before a recovery can reassign it * fix(companion): pin the runner listener to loopback * fix(companion): size the replace delete prefix to cover any focused field * build(ioscompanion): refresh runner asset with loopback bind and replace fix * fix(cli): cancel the run context on SIGINT so spawned children are reaped * fix(testrun): point the device java preflight hint at the ios-device doctor * fix(folio): word-bound the G2 ERROR scan and drop the dead objc allowlist glob * test(ioscompanion): cover stopProcess, restart, and failed bring-up supervision * chore: add test-companion target for the withcompanion-tagged suite * chore(ioscompanion): stop tracking the runner archive build artifact * build: produce the runner archive from source like the companion bundle * refactor(conformance): move the gate harness out of examples/folio * chore(folio): drop the gate harness wiring from the example app |