Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.
Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.
The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.
The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.
Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.
Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.
Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.
Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.
A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.
Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.
An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.
The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.
Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.
Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.
Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.
The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.
idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.
Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.
Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.
No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.
An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.
Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.
A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.
The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.
This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.
The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.
A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.
Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX