The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.
The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.
One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.
The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.
The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.
serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.
Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.
A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.
The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.
A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.
Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.
The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.
The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.
The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.
Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.
The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.
Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.
Preflight names every missing serial before the first seed is dispatched.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.
Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.
resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.
Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.
Landing the package in one commit because the intermediate splits would
not link.
A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.
same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.
installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.
re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.
a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.
the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.
reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.
a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.
--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.
an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.
the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.
A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.
Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.
A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.
A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.
Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(cli): add --max-steps for step-bounded runs
runner.Options.MaxSteps already worked but was unreachable from the command
line. A step budget is what makes two generators comparable: one making a
model call per step and one drawing from a PRNG are not comparable per second.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(trace): record arm membership and host in meta.json
meta.json recorded the seed but not which picker ran, how it was configured,
what budget it was given, or which machine produced it. A directory of runs
cannot be attributed to an experiment cell without those, which makes any
factorial computed from such a directory unanalysable after the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(cli): add --arm and populate run meta from it
Model and instructions are recorded only when the LLM picker is the one that
will actually run, so a spec declaring generator = llm() that is run under the
seeded picker does not label its trace with a model it never called.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): sweep seeds for one experiment cell
campaign.json lists the seeds a sweep intended to run and is written before
the first run, so a host that dropped runs shows up as missing seeds rather
than as a smaller sample. Seed 0 is rejected: sanderling test reads it as
"derive a seed from the clock", which is why conformance/gates.sh controls
nothing today.
Each run contributes one runs.jsonl line carrying steps to first violation by
origin step, the step that armed the failed obligation, so the survival
analysis never reopens a trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(runner): no silent generator fallback, and llm on web
--generator llm against a spec declaring no generator = llm(...) logged a
warning and ran the seeded picker. For a comparison campaign that is silent
arm corruption: the run completes, the directory looks correct, and the wrong
policy drove it. It is now fatal.
pickSources also returned the V8 source for both action and extractor on web
before it looked at the generator, so the llm policy was unreachable there.
The two axes are now independent: the driver picks the extractor source, the
flag picks the action source, and llmSource composes with either because the
runner populates the candidate list and screenshot on every platform.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): make the hierarchy dump agree with the web runtime
Three facts differed between the dump the goja host reads and the DOM the V8
host reads, so the two enumerated different candidates on one page.
scrollable was never emitted, and worker.go reads exactly that attribute while
targets.ts requires it for scrolls, so the goja host could not offer a single
web scroll. clickable tested el.onclick, which React assigns to its root
container for event delegation, making the whole viewport a tap target here and
in no other enumeration. Both now resolve through the selector sets in
pkg/spec/src/web-runtime.ts.
The dump also rooted at body while collectTargets walks querySelectorAll("*"),
so the goja host never saw html, where page-level scrolling lives. It now roots
at documentElement and skips the head subtree, which is all zero-bounds and
would otherwise carry script and title text into the trace.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(conformance): give the gate reproducible seeds
SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from
the clock", so the tunable controlled nothing and a gate failure could not be
re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the
results table so a failing row names its stream.
The five runs stay on five different streams: a gate that scored one path five
times would catch less than one that scores five.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): emit editable as a plain boolean
editable was emitted as `isEditable || null`, and an absent field sends
internal/hierarchy into the native fallback, which reads any class name
containing "EditText" as an Android text widget. On web that is just a CSS
class, so a page styling a div with it was editable to the goja host and not to
the web runtime, and the model policy could be offered typing into a div.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(spec): leave the head subtree out of the web target walk
collectTargets walked querySelectorAll("*") while the hierarchy dump skips head,
so the two hosts enumerated different element sets on every page with a <head>.
No candidate changes: builtinCandidates pushes only for targets acceptsTarget
admits, and head elements have no positive bounds, so the list the draw ranges
over is untouched. What changes is that targetIndex now means the same thing on
both hosts.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* test(chrome): compare the facts both hosts derive from one DOM
The existing parity harness hand-authors the facts on both sides, so it proves
that given identical facts both hosts select identical candidates, and says
nothing about the two code paths that derive those facts from a real page. Four
divergences lived in that blind spot and it passed throughout.
This drives one real page and compares clickable, enabled, editable, scrollable
and positiveBounds element by element, plus the element sets themselves, which
is what catches a host that omits html or includes head. Reverting any of the
four fixes makes it fail naming the element and the fact.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* chore(make): run the browser packages one at a time
Both launch Chrome and launching two at once has failed with "Launch: context
canceled".
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style: remove every em-dash and en-dash
Eighteen occurrences across fourteen files. Each sentence was repunctuated to
suit what the dash was doing rather than swapped for a hyphen, which produces
comma splices. The minus sign in folio-web's ledger is a minus sign and stays.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(chrome): honor the caller context in Launch
Launch and clearState ran against d.tabCtx, so a target that accepts the
connection and never answers wedged the process past its own --duration and
through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for
the rest of the sweep with no diagnostic.
The browser is still allocated against d.tabCtx first, because chromedp starts
Chrome under whichever context calls Run first and allocating under a caller
deadline would kill the browser when Launch returns. Everything after
allocation goes through runCtx.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* fix(sidecarassets): publish the extracted jar through a rename
Extract wrote a 96 MB jar with a plain WriteFile into a temp path every
sanderling process on the host shares. On a cold host several concurrent
workers all miss the checksum and all write the same path, and O_TRUNC lets one
spawn a JVM against another's half-written archive. A fresh experiment host is
exactly a cold host.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(campaign): kill a run that outlives --run-timeout
A wedged run holds its worker for the rest of the sweep, and on an unattended
host nothing else will send it a signal. Defaults to three times --duration and
must exceed it. A killed run is recorded as timed_out rather than as a generic
failure, so the analysis can tell a lost cell from a real crash.
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* style(test): gofmt browser_test.go
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
* feat(ltl): bound fields on AlwaysFormula and named thunks
Add StepBound/Duration/Deadline to AlwaysFormula as the dual of bounded
Eventually, give ThunkFormula a Name for stable identity, add ThunkNamed,
and surface both in describe() and MarshalJSON.
* feat(ltl): negation normal form pass
nnf/pushNot rewrite a formula so every Not wraps only a Thunk or Error
leaf, dualizing Always<->Eventually and preserving bounds.
* feat(ltl): NNF in NewEvaluator, bounded-always, Finalize, collapse
Apply nnf on construction, reduce bounded Always symmetric to bounded
Eventually (vacuous holds once the window closes), add Finalize to
resolve undischarged liveness obligations to Violated at run end, and
collapse structurally-identical pending obligations.
* test(ltl): property-based NNF laws
Lock double-negation identity, Always/Eventually duality with bound
preservation, leaf pushdown, and not(always true) reaching Violated.
* test(ltl): Finalize, bounded eventually, latch, collapse
Property tests for monotonic violation latch and eventually-within
violating iff n consecutive false, plus Finalize and collapse cases.
* feat(inspect): within clause on always residual node
A negated bounded eventually serializes as a bounded always; render its
bound instead of dropping it.
* feat(ltl): witness violations and (bool,error) predicate thunks
* test(ltl): migrate thunk call sites to (bool,error)
* feat(ltl): flag thrown-predicate witnesses with IsError
* refactor(verifier): replace predicate err side-channel with violation witness
* test(verifier): witness API for thrown predicates
* feat(trace): witnesses map and skipped-verification marker on Step
* feat(runner): thread violation witnesses, finalize, skip marker into trace
* test(ltl): lock violation witness reason, IsError, and step
* test(verifier): finalize surfaces unmet eventually with witness
* fix(ltl): eliminate implies and bounded-always false-negatives
Rewrite a -> b to (not a) or b in NNF so a pending temporal antecedent
can no longer defer the whole implication and drop a consequent that was
false at the current step. Carry a pending inner past a bounded-Always
window close instead of dropping it to holds, so a deferred obligation is
resolved by a later step or Finalize.
* test(ltl): lock implies and bounded-always false-negative regressions
* fix(web-runtime): seed PRNG for reproducible runs and align weighted pick
* feat(testrun): inject seed into web bundle via SANDERLING_SEED define
* test: cover web-runtime seeded PRNG, weighted pick, and seed define wiring
* test(spec): add Go math/rand/v2 PCG oracle and golden fixture
* feat(spec): bit-exact PCG port of Go math/rand/v2
* test(spec): assert pcg.ts matches the PCG golden fixture
* feat(spec): shared input corpus and press-key pools
* feat(spec): action-tree types and Host interface
* feat(spec): verb support matrix and warn-once helper
* feat(spec): deterministic shared action picker
* test(spec): verb matrix and warn-once semantics
* test(spec): picker draw-order and determinism
* refactor(spec): actions.ts returns pure GeneratorNode data trees
* refactor(spec): wire from() sampling through the picker rng
* feat(spec): shared runtime-entry installs next-action over pick.ts
* feat(spec): export LongPress/Scroll/longPresses/scrolls factories
* test(spec): assert data-tree shapes for action factories
* test(spec): runtime-entry serializeAction wire-contract round-trip
* refactor(spec): bridge data-tree nodes to the legacy goja picker tags
* fix(spec): web runtime walks the spec's globalThis.actions data tree
* test(spec): tolerate legacy bridge fields on builtin nodes
* refactor(spec): installRuntime accepts a lazy root resolver
The web bundle imports the runtime before the spec, so the action root
on globalThis.actions only exists after the spec evaluates. Accept a
function form so the goja and web hosts resolve the root per tick.
* refactor(spec): web-runtime becomes the WEB Host, delegates to shared picker
Delete the duplicate picker (resolveGenerator/pickWeighted/randomTap/
randomInput/randomSwipe/randomPressKey/pickFromArray, the mulberry32 PRNG,
and the snake_case serializeAction) plus the __sanderling__ action factory
binds. web-runtime now implements Host (platform/seedHi/seedLo from the
injected 64-bit seed via BigInt, queryCandidates over the live DOM with a
per-tick cache, reportUnsupported) and calls installRuntime so both engines
run pick.ts over the same Pcg. Swipe/longPress/scroll follow the verbs.ts
matrix instead of silently returning null. Keeps the DOM helpers (selector
translation, queryElement, elementHandle, buildState, sanitize, extractors)
and the global locking. Net -214 lines (741 -> 527).
* test(spec): cover the WEB Host surface and seed precision
Replace the deleted-picker tests with Host coverage: platform()==web,
seedHi() parsing a 64-bit seed without Number precision loss, seedLo()==0,
reportUnsupported warning, the installed next-action/extractor globals, and
queryCandidates verb routing + per-tick caching over a querySelectorAll stub.
* refactor(spec): picker emits native selector + scroll endpoints, setup precedence
* feat(spec): goja runtime entry wires the shared picker over the Go host
* feat(bundler): optional RuntimeFile prepends a runtime-entry import via stdin
* feat(testrun): bundle the goja runtime entry so the verifier runs the shared picker
* refactor(spec): drop the legacy goja bridge fields from action factories
* feat(spec): serialize selector-only string targets for the runner to re-resolve
* refactor(verifier): one DecodeAction reads the unified flat wire contract
* refactor(verifier): goja host + shared picker replace the duplicate Go picker
* refactor(runner): decode V8 actions via the unified DecodeAction; wire goja runtime
* test(verifier): author specs through the shared picker path
* test(runner): bundle authored specs with the goja runtime entry
* feat(verifier): collect unsupported verbs for the run report
* refactor(runner): collapse WebDriver forks behind ActionSource/ExtractorSource
* feat(testrun): surface unsupported verbs in run report
* test(verifier): cross-runtime goja/node parity gate on the shared picker
* test(verifier): unsupported verbs collected deduped in first-seen order
* test(runner): summary reports no unsupported verbs on a clean run
* test(spec): golden-fixture cross-runtime parity gate for the node picker
Replace the env-driven parity harness with a shared scenario module and a
committed golden the node picker asserts independently. The goja side asserts
the same golden, so neither runtime invokes the other at test time.
* test(verifier): assert goja picker against the same cross-runtime golden
Drop the node-subprocess coupling: the goja side now installs a stub
__sanderlingHost__ with the fixed candidate list and asserts the committed
golden, matching pkg/spec/test/parity.test.ts.
* refactor(spec): rename pressKey generator export to pressKeys
* refactor(spec): update barrel re-exports for pressKeys
* test(spec): update pressKeys generator export name
* docs(spec): rename pressKey generator to pressKeys
* refactor(spec): extract samplerRng into shared sampler-rng module
* feat(spec): add fluent seeded value generators (strings/integers/emails/edgeCaseText)
* test(spec): cover fluent value generators determinism and chaining
* refactor(bundler): inject globalThis trailer from spec named exports
* refactor(bundler): reuse registration trailer in web bundler
* test(bundler): cover named-export globalThis registration
* feat(spec): add named() to Extracted handle type
* feat(web-runtime): named() and cross-extractor read guard
* feat(verifier): named() and cross-extractor read guard in goja
* test(verifier): cross-extractor read guard and named()
* test(web-runtime): export runtime and extractors for tests
* test(web-runtime): named() and cross-extractor read guard
* refactor(folio): drop manual globalThis trailer (bundler injects it)
* refactor(folio): seed txn amounts via integers().between(1,500)
* refactor(folio-web): drop manual globalThis trailer (bundler injects it)
* fix(folio-web): seed card/txn-type selection via from().generate() for reproducible runs
* refactor(folio-web): weight valid generators against edgeCaseText for names/amounts
* refactor(folio-web): name extractors so violation witnesses are readable
* fix(web-runtime): propagate extractor getter throws and unpoison locked global
Stop swallowing getter errors in evaluateExtractors so the cross-extractor read guard aborts loudly, matching goja's PushSnapshot. Make the __sanderling__ lock configurable (still non-writable) so a shared test process can reinstall a fake.
* test(spec): install fake runtime via defineProperty to survive locked global
* test(web-runtime): assert uncaught cross-extractor read aborts evaluateExtractors
* feat(runner): add MaxSteps bound to Options
* test(runner): MaxSteps stops after exactly N steps
* test(driverpb): drop proto getter round-trip tautology
* test(sidecar): drop stub-mode placeholder tautology tests
* test(mock): drop default-field-value assertion test
* test(ltl): drop Verdict.String tautology tests
* refactor(runner): extract RenderSummary for snapshot testing
* test(runner): golden snapshots for trace stream and violation summary
* feat(web-runtime): capture uncaught errors into state.exceptions
* test(integration): add throwing and counter web fixtures
* test(integration): add specs for the web fixtures
* test(integration): drive web fixtures through the real pipeline in headless Chrome
* chore(make): add test-browser target for the Chrome-driven suite
* ci: run the Chrome-driven browser suite in a separate job
* refactor(test): relocate browser suite to test/browser
* refactor(permissions): delete dead internal/permissions package
* refactor(test): rename package to browser_test
* refactor(sidecarassets): rename internal/sidecar to internal/sidecarassets
* chore(make): point test-browser at test/browser
* docs(decisions): record internal/permissions deletion
* refactor(doctor): use sidecarassets package
* refactor(testrun): use sidecarassets package
* fix(test): resolve testdata relative to browser_test.go
* refactor(verifier): remove dead __sanderlingIndex compat alias
* refactor(bundler): use encoding/json for JS string literals
* docs(action-space): use vendor-neutral native driver wording
* refactor(hierarchy): scrub backend tool name from comments
* refactor(driver): scrub backend tool name from comments
* refactor(driver): add DoubleTap and DoubleTapSelector to DeviceDriver
* refactor(sidecar): implement DoubleTap with the sub-100ms inter-tap gap
* refactor(chrome): implement DoubleTap as two taps with the gap
* refactor(mock): record DoubleTap and DoubleTapSelector actions
* refactor(runner): delegate double-tap to driver, drop gesture timing
* test(runner): assert double-tap delegates to driver DoubleTap
* docs(cmd): add package docs to CLI and developer tools
* docs(driver): add package docs to driver interface and chrome backend
* docs(driver): add package docs to mock and sidecar backends
* docs(platform): add package docs to android and ios device prep
* docs: add package docs to bundler and inspect
* docs(ltl): add package doc to temporal logic evaluator
* docs: add package docs to runner and testrun pipeline
* docs: add package docs to trace and verifier
* docs(sidecarassets): add package doc for embedded JAR loader
* fix(chrome): add disable-dev-shm-usage so Chrome starts in CI
* test(chrome): gate real-Chrome driver tests behind the browser tag
* chore(make): run chrome driver tests in the browser job
* fix(web-runtime): guard global error listeners for non-browser hosts
The module registered window error/unhandledrejection listeners at top
level, which threw under Node (the spec-api test runner) where
globalThis.addEventListener is absent. Register only when the API exists;
the real browser run is unaffected.
* ci(browser): re-enable unprivileged user namespaces for headless Chrome
ubuntu-latest moved to 24.04, whose AppArmor restriction on unprivileged
user namespaces stops headless Chrome from opening its DevTools socket
even with --no-sandbox, surfacing as the driver's 'websocket url timeout'.
Relax the sysctl for the job and add a direct launch check so a future
breakage shows Chrome's own stderr rather than an opaque driver timeout.
* ci(browser): pin stable Chrome for the driver tests
setup-chrome's default latest pulled a dev Chromium (150) whose remote
debugging socket never came up under chromedp, while plain --dump-dom
worked. Pin the stable channel, which the driver is tested against.
* feat(defaults): add scroll and rebalance action weights
Use relative-integer weights (taps/typing co-primary 100, scrolls 50,
swipes 25, doubleTaps 10); the picker normalizes by their total. Adds
scrolls to defaultActions as a first-class reveal behavior.
* feat(defaults): trim scroll action weight wiring
* fix(build): point sidecar jar ignore and embed paths at sidecarassets
* test(defaults): drop stale longPresses re-export assertion
longPresses is opt-in vocabulary, no longer re-exported from
defaults/actions.ts since e0d3b20; its builtin resolution is already
covered by api.test.ts. Trim the defaults test to scrolls, which is an
actual default export.
* fix(chrome): raise DevTools websocket read timeout to 60s
Chrome cold-start on a loaded CI runner can exceed chromedp's 20s
default for reading the DevTools websocket URL, flaking the browser
tests with "websocket url timeout reached". Give launch more headroom.