record what the model picker did, and make both policies see the same actions (#74)

* feat(llmclient): parse usage and the served model

An LLM-in-the-loop evaluation has to report tokens per action and cost per
defect, and the client discarded both counters. Served model is recorded
separately from the requested one because a router can substitute a
differently-priced variant.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record one typed outcome per model-driven step

llm-calls.jsonl carries the prompts as sent, the candidate list as the model
saw it, the screenshot reference, the raw response, tokens, latency and how the
step ended. It sits beside trace.jsonl rather than inside it because every trace
line already carries a full hierarchy and both the replay server and the
campaign summarizer scan all of them; folding prompts in would grow the lines
those readers parse for data neither reads.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): expose the step a snapshot was observed at

It lags the runner's current step whenever a transitional tree caused an
observation to be skipped, which is exactly when the model is shown an older
screen than the step it is choosing for.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a guard-skipped step is no longer a silent log line

The strict echo-skip left only a logger.Warn, so a step the guard discarded was
indistinguishable in the trace from a picker that legitimately declined. Any
yield or actions-per-hour figure computed from model traces mixed the two.
Every path that ends a step without a model-chosen action now records its own
outcome.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): record when a chosen action was never dispatched

A step could carry a next_action that the foreground guard or an apply error
stopped from running, and nothing said so. An executed-action count read off
trace.jsonl included actions that acted on nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): document llm-calls.jsonl

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(analyze): survival analysis over campaign directories

Steps to first violation with clean runs right-censored at the budget, since
per-run yield is a binary at 11 to 45 percent and separating two arms on it
would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum
with Vargha-Delaney A12, Holm within each family.

A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator
and would be believed, so every statistic is validated against a published
worked example with the source named in the test: R survdiff on aml, Freireich
6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for
Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and
the tie-corrected variance, checked against an exact permutation variance.

Failed and timed-out runs are excluded as missing data and counted by reason,
never treated as censored observations, which would bias the result.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): select the candidate label source

Candidates takes the label source as an argument rather than storing it, which
is what keeps the asymmetry structural: the seeded picker selects by index and
never calls Candidates, so the mode cannot reach it. That asymmetry is
load-bearing, because it makes the two seeded cells of the factorial a
manipulation check with identical draw streams.

The identifier ladder deliberately has no text rung. A fallback that reached
for text would silently turn one arm back into the other.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(runner): thread the label source to the model picker

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the label source as arm membership

Recorded for seeded runs too, unlike model and instructions. Without it the two
seeded cells are indistinguishable in the artifact and the manipulation check
cannot be grouped.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(cli): add --label-source

Unknown values are rejected at parse time rather than falling back to the
default, matching the generator check: a campaign that completes with the wrong
arm and a correct-looking output directory is worse than one that fails.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): dedup candidates by what they execute, not how they read

The dedup key was the rendered description, which embeds the label, so two
distinct controls sharing a visible label collapsed to one entry and the
survivor carried the first one's action. The second control was not mislabelled,
it was absent from the candidate list, so no policy could reach it. Two
scrollable containers collapsed the same way, leaving the second unscrollable.

The key is now the executable Action struct itself plus whether the model
supplies the typed text, so a new Action field cannot silently fall out of it.
Descriptions may now repeat; the numbering disambiguates and the echo guard is
index-anchored, not description-anchored.

This also makes the label source a pure observation-channel change. It was not
one before: the label fed the dedup key, so the two arms of the labelling
factor enumerated different-sized candidate lists, in both directions depending
on the screen.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): report every action that was chosen and never dispatched

applyAction could return nil without calling the driver, so the trace showed an
action that looked executed and acted on nothing. Six paths did it: a tap,
double-tap or long-press whose coordinates do not resolve and which carries no
selector, a long-press whose selector is stale, an empty key press, and a
zero-duration wait. It now reports whether it dispatched, and the runner records
the reason and clears lastAction so the verifier never attributes the next state
to an action that did not run.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): the echo guard admits a repeated description

Descriptions can now repeat after candidates dedup by what they execute. The
guard is index-anchored, so this pins that a repeated string cannot make it
misfire in either direction.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): count dispatched actions, not steps

A step where the policy declined has no action, and a step whose action was
never dispatched did nothing. Both were being counted as actions by everything
downstream.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide by actions that ran

Defects per thousand actions counted every step, including steps that chose
nothing and steps whose action was never dispatched. The inflation is
policy-dependent, so it does not cancel between arms: on the fixture campaign
the model arm's yield was reported at 60.3 per thousand against a true 120.7,
because half its steps did nothing.

A runs.jsonl without the count is refused by name and line rather than read as
zero actions, which would report every per-action rate wrongly. The report also
carries steps beside actions now, so the gap is visible rather than folded into
a denominator.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): lower authored actions the way the seeded arm does

The authored descriptor path had no parity guard and diverged from the wire
format on almost every verb. A Wait lost its duration and was skipped as a
zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that
resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe
to (0,0) instead of being dropped.

An authored target object with no x property panicked the whole run at
candidate enumeration: ToInteger was called on a nil goja.Value. A target on
the screen origin is still kept, so the drop rule cannot swallow it.

Builtins were never affected. They serialize through the same path the seeded
arm uses, which the existing policy parity test covers.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(verifier): decode a container-only scroll

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): carry the container on an authored scroll

serializeAction sent the container's own point as both endpoints, so an
authored Scroll({in, direction}) reached the driver as a drag from a point to
itself and did nothing, on the seeded arm. The wire now carries the selector
and leaves the drag to the runner, which sizes it from the container's bounds
and has always had tested support for it that nothing could produce.

No rng runs in the serializer, which lowers an already-drawn action, so the
draw stream does not move. Builtin scrolls compute both endpoints and their
bytes are unchanged.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): both policies must dispatch the same authored action

Compares the recorded driver calls across 13 authored shapes. The builtin path
had a parity guard and the authored path had none, which is why it drifted on
almost every verb.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(hierarchy): match identifiers by role prefix

idPrefix: is id: with starts-with in place of equality, so a list whose
rows are named <role>_<record id> is reachable by the durable half. The
Android package prefix is skipped the same way id: skips it.

Routing both prefix kinds through matchAttr also makes the object form
work: {descPrefix: ...} matched nothing on the native side while the web
runtime honoured it.

* feat(chrome): translate idPrefix to a starts-with id match

* feat(spec): match idPrefix in the web runtime

The DOM has no package prefix, so the native rule reduces to [id^=]. Both
prefix kinds now go through the one key table, which drops the separate
descPrefix branch that string and object selectors each carried.

* feat(sidecar): match idPrefix in the tap-by-selector path

* docs(manual): document the idPrefix selector

* feat(replay-ui): render idPrefix targets as a prefix tag

* fix(spec): read the injected seed per call

Binding it at module scope bound it to whenever the module was first
imported, so a test file that imported the runtime before setting
SANDERLING_SEED froze the seed at zero for every file after it. The
bundler still replaces the expression with a literal.

* test(chrome): compare both selector matchers over one live page

Selector matching is written once per runtime: internal/hierarchy over the
dump, web-runtime.ts over the DOM. Nothing made the two agree, and a
selector that resolves on one and not the other is silent, since an empty
match yields no action and the run still passes.

* fix(hierarchy): give id and desc one meaning in both selector forms

The object form fell through to the raw attribute map, which carries no
id or desc key on any platform, so {id: "save"} matched nothing while
"id:save" matched. The repo's own web spec uses the object form thirty
times. Both forms now resolve through one switch.

Adds the accepted-key list and UnknownSelectorKeys with it, since the
same silence hides any mistyped key. A key some element carries is always
accepted, so raw driver attributes stay reachable.

* test(hierarchy): pin both selector forms and the unknown-key report

* feat(verifier): fail the spec on a selector key that cannot match

An empty match is indistinguishable from a screen with no such element,
so a mistyped key generates no action for the whole run and the campaign
finishes clean having explored nothing. The goja boundary now throws,
naming the key and the accepted list.

* feat(spec): reject an unknown object-selector key in the web runtime

Same rule and the same message as the native side: a key no element can
carry throws instead of matching nothing. The accepted list is one list,
committed as a fixture both suites assert, so a spec cannot be accepted
by one runtime and rejected by the other.

* test(spec): pin the unknown-key diagnostic to one text

The two runtimes each claimed to raise the other's message and nothing
checked it. Both now render the committed text for the committed key.

* fix(spec): match a merged label by its leading name on web too

The native desc rule accepts the label or the label at the head of an iOS
merged label; both web translators compared the whole string, so the same
selector matched natively and missed on web. The live-page parity test
caught it.

* test(chrome): drive the live-page parity test through both selector forms

* docs(manual): document object-selector key rules

* feat(spec): refuse a multi-item authored sampler while enumerating

from().generate() draws from the picker's rng, which exists only inside
walkActions. The model policy enumerates authored leaves outside that walk, so
the sampler silently yielded its first item on every step: measured over 30
draws the seeded arm reached three targets in roughly equal proportion and the
model was offered only the first. The two policies had different action spaces
and nothing said so.

A single-item sampler short-circuits before the rng, so both policies get the
same value and it is not refused.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets

Candidates returns an error now. The refusal is thrown at the draw and wrapped
with the source of the leaf that made it, since generate() cannot know which
leaf it is inside. Only that marked refusal is fatal: this walk calls every
leaf on every step, so promoting the rest would kill model runs the seeded arm
survives.

Authored actions on a disabled target are no longer dropped from the model's
candidate list. The seeded picker executes whatever the leaf authored, and a
control the application forgot to re-enable is exactly where boundary defects
live, so a policy that cannot attempt it cannot find them.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): abort on a candidate enumeration that refused

Recorded as candidates_failed before the run stops, so the trace says why.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(spec): refuse a multi-value generator while enumerating

integers, strings, emails and edgeCaseText read the same rng from() does, so
under the model policy an authored InputText typed the same value on every
step while the seeded arm varied it. That is a silently different experiment,
not just a silently different action space.

Single-valued spans are exempt, because both policies then get the same value:
between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is
still refused, since the length is pinned but each character is drawn from 62.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(verifier): setup still draws, and the seeded stream is unmoved

Setup runs through the picker with the rng under both policies, so a generator
there is legitimate and must keep working. Interleaving enumeration and setup
catches the flag leaking out of the model's walk.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): value generators are refused under the model policy too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio): enumerate authored targets and values instead of sampling

Sampling inside an authored leaf is refused under the model policy now, because
the draw collapses to its first item there. Each sampled leaf offers one action
per value instead.

Lists are short, three rather than five, because the two form leaves also carry
their submit and the seeded picker splits a leaf's probability across the
actions it returns. The doubleTaps path that reaches the planted defect is
unchanged at 5.88 percent, since no root or defaults weight moved.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): enumerate authored targets and values, declare the llm generator

The two edge-case typing leaves become the typing builtin at their combined
weight: that text is deliberately not domain-specific, so naming the field and
leaving the text to the policy is the designed path, and it keeps the seeded
arm on the corpus while the model writes its own.

Total weight is unchanged at 165, so every surviving branch keeps its share and
submitTxn stays at 9.70 percent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: minimal changes, self-documenting code, tests as first-class

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(spec): key web attrs by the names the markup writes

attrs was spread from element.dataset, whose DOMStringMap keys are camelCase,
so a spec reading attrs["data-cents"] the way every native host reports it read
undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently
zero: someTransactionExists could never be satisfied, balanceMatchesTransaction
Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and
passed vacuously. Three properties reported nothing because the harness was
blind, not because the application was correct.

The handle also fills hintText and editable now, so an authored InputText on
web names its field the way the same action names it on Android instead of
rendering as Type "12.34" into "".

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name a web handle by the same ladder as a tree element

The handle fallback read only text, which is textContent and therefore always
empty for an input, so the model could not tell the amount field from the note
field. It now mirrors visibleLabel's ladder rather than introducing a second
naming scheme.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): attrs carries raw attribute names on web too

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): confirm focus moved before typing

InputText tapped its target, slept, then typed. Android and web both inject
into whatever holds focus, so a tap that missed sent the whole string somewhere
else and nothing reported it. On an emulator with a floating keyboard panel
parked over the password field, the tap pressed the keyboard's emoji key and
every step appended the password to the email instead, forever, because the
setup leaf is guarded on the password being empty.

The hierarchy is re-read after the tap and the target, or something in its
subtree, must hold focus. Platforms whose hierarchy carries no focused
attribute skip the read, so they pay nothing.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): signal a timed-out run so it reaps its sidecar

CommandContext kills outright, so a run stopped by --run-timeout never ran its
own shutdown and left a sidecar holding a port and a quarter gigabyte,
reparented to init and deaf to SIGTERM. The timeout exists for unattended
hosts, which is exactly where nobody is watching to reap what it leaves.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* perf(runner): confirm focus only when another element holds it

Measured over 717 InputText steps: nothing was focused before the tap 23.8
percent of the time, the target already held focus 60.4 percent, and a
different element held it 15.8 percent. Silent corruption is only reachable
from that third class, and all four real rejections observed came from it.

Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy
reads, and recovers about 8 percent of Android run time. The pre-tap and
post-tap conditions are now the same predicate stated once.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): record both clocks a run was measured on

Duration came from the monotonic clock, which does not advance while a host
sleeps: one calibration run under-reported by about 15 minutes. A run now
carries monotonic_millis for how long it worked and wall_clock_millis for how
much time passed, which is what makes a sleep visible at all.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(analyze): divide per-hour rates by time actually worked

A host asleep mid-run tested nothing, and charging that sleep to an arm reports
it slower for a reason unrelated to the arm. The legend also claimed wall clock
while the number was monotonic. Campaigns written before the split are still
read through the old field name so their run hours do not silently zero.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(campaign): wait for the trap instead of racing it

The reaping test gave the wedged script one second to install its TERM trap,
so a loaded machine signalled it first and the test failed for a reason it does
not test. It now waits for the script to say the trap exists, then cancels.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(ltl): keep the authored window on a step-bounded obligation

reduce decremented StepBound into the residual, so the trace reported the
remaining window rather than the authored one: a within(1915, "steps") showed
up as 1875 after 40 steps, and the replay UI renders that string verbatim. The
duration case was fixed when bounded windows were made to serialize their
resolved deadline; the step case was not, and withinFor's comment claimed
otherwise.

The window is now immutable and the closing observation is resolved once, which
mirrors Deadline exactly. A step counts observations the evaluator reduced,
not steps the runner executed, because a skipped step gave the property no
chance to discharge and transitional-step rate is itself policy-dependent.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(ltl): pin that a slow policy does not fail on time alone

Same 300-observation trace at two cadences: a 300 second bound holds for the
seeded arm and violates for the model arm eight observations before the
predicate fires, while a step bound holds for both. Green before and after,
because the step unit already worked; this pins the property rather than
fixing it.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(spec): guard the step unit on the authoring surface

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(folio-web): bound the reachability properties by steps

At one model call per step the model arm takes 359 seconds where the seeded arm
takes 47, so a second-based deadline reported violations that were the arm's
speed rather than the application's behaviour. The three cross-arm reachability
properties now bound by steps, derived at the measured 6.383 steps per second.

The two auth-transition properties keep seconds: a user waits through those
regardless of which policy is driving.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs(manual): a step bound counts observations, not runner steps

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): make the label source a cell dimension

A 2x2 of policy against labelling needs the runner to express both factors.
It could only express the policy, so half the factorial had to go through
--extra, where the manifest would not record what was actually run.

Rejected at parse rather than on dispatch: a sweep that finds the bad value
on run 1 of 40 has already spent a cell's worth of device time.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(campaign): record the label source in the manifest

A finished sweep should say which cell it ran without anyone having to
remember the invocation.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): name a web field by its hint, not its CSS class

visibleLabel reads hintText first for an editable element. The dump never
emitted it, so an empty web input fell through text, description and
descendant text to its class name, and the model was shown an identifier no
user can read on exactly the fields a labelling experiment varies.

Same ladder as fieldHint in web-runtime.ts, so one field is named one way on
both hosts.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): answer clickable for an element reached through ax

The handle hardcoded true, so every text node and container a spec reached
through state.ax claimed to be a tap target while the enumeration and the
hierarchy dump both resolved it through the tappable selector.

The parity test now compares the handle against the enumeration element by
element in a real browser, which is where the three answers have to agree.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* docs: every target runs on this machine, so start one rather than skip it

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): a selector the tree cannot resolve is not a focus failure

otherElementHoldsFocus answered true when FindNode returned nothing, so an
unresolvable target read as "another element holds focus". confirmFocus then
re-dumped, resolved nothing again, and errored unconditionally. Three of those
in a row abort the run.

Not knowing where the target is says nothing about where the text would land.
The guard's real case, a resolved target with focus outside its subtree, still
errors exactly as before.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(chrome): emit data-testid so both resolvers name the same element

The V8 host names a web target by data-testid and TapSelector translates that
selector into a CSS attribute match, but the dump carried no such attribute
and no alias could supply one, since an alias only redirects to a key that
already holds the value. tree.Find was therefore always nil for exactly the
selectors examples/folio-web tags with.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): name an element only when the selector names it alone

ax.findAll stamped every result with the query selector, and resolveCoordinates
prefers the tree lookup over the element's own coordinates, so N sibling
candidates all executed on the first match. On folio's Home screen the fuzzer
could never open any account but the first.

The gate tests identity rather than cardinality: no node other than this one
answers to the rendered string, checked with the same lookup the runner runs.
A rendered object selector can resolve somewhere the query never matched, so
counting the query would call that unique.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(web-runtime): hold the V8 host to the same naming gate

elementHandle stamped the query selector on every result the same way, so the
merge carried the sibling collision onto web for authored ax targets.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* test(runner): sibling taps reach the driver at their own coordinates

Drives 40 real draws from a spec that taps each card, through the picker, the
serializer and DecodeAction, and asserts on the points the driver saw. Against
the shared-selector bug all 40 landed on the first card.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(runner): an ambiguous name loses to the coordinates it was built from

Attribute values match by substring, so a selector that named one element
where the candidate was built can name several in the tree it resolves
against, and the lookup sent every one of them to the first match. The host
gates blank an ambiguous tag at enumeration time; this closes the gap between
that moment and the action.

A bare-string target carries no coordinates, so the first match stays the
answer there rather than dropping an authored action.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(verifier): record an element-valued extractor instead of dropping it

An element carries find/findAll host functions, so json.Marshal refused the
whole value and the encoder answered nil. ChangedExtractors then emitted no
entry: no error, no warning, no value. Project the value the way the web host
already does (functions dropped, cycles and over-deep branches null, non-finite
numbers null) and turn whatever is still beyond JSON into an error the author
sees, rather than a missing extractor.

* test(verifier): an unrecordable extractor value is reported, not dropped

* test(runner): element-valued extractors reach the trace

* docs(spec-language): say what a trace records for an element-valued extractor

* docs(claude): add delegation and record-keeping sections

delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository.

* feat(driver): declare undelivered-action errors and three optional capabilities

ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner.

* feat(driver): add escape to the pressKey surface

escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all.

* fix(ios): refuse a gesture the screen has no surface under

the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error.

* feat(ios): derive scrollable from the snapshot's tree depth

the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one.

* fix(sidecar): stop dropping gestures, selectors and keys in silence

a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing.

* fix(sidecar): map the driver's refusals onto the gesture errors

OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak.

* fix(selectors): resolve text to the innermost match and scan the root in both forms

an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree.

* feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag

a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump.

* fix(chrome): emit every markup attribute and read checked and selected off the property

the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever.

* fix(chrome): scroll a gesture point into view and dispatch trusted input

getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires.

* feat(chrome): read the page's exceptions and navigations, and hold the picker state across them

a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision.

* feat(trace): version each step and record its logs, exceptions and navigations

a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report.

* feat(runner): bound every device call and record the actions that never reached the app

observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one.

* feat(verifier): expose extractor names and rebuilt property formulas

an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does.

* feat(testrun): expose the seeded bundle a run loaded

BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity.

* feat(tracecorpus): load recorded runs for offline measures

reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl.

* refactor(seedspec): move seed spec parsing out of the campaign command

the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change.

* feat(analyze): time an event at the step it was detected and report the quartiles

an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median.

* feat(analyze): add the seed-paired signed-rank comparison and record the holm family

--paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct.

* test(analyze): recover planted effects through the tool's own entry point

a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk.

* feat(label-coverage): report the addressable share of an app's interactive surface

reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression.

* feat(exploration-reach): count the distinct structural states a stored run visited

the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay.

* feat(defect-identity): count distinct defects across stored runs

a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed.

* feat(oracle-reduction): replay stored traces under four reduced oracles

re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding.

* feat(implementation-sweep): run one campaign against every implementation of a requirement

installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed.

* feat(corpus-sweep): run one specification against a served corpus of implementations

same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them.

* docs(manual): document innermost text matching, escape and the web scroll verb

text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what.

* test(browser): assert an uncaught page exception reaches the trace

the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary.

* feat(trace): a step can name the precondition it could not meet

A step that never had the app under test in front of it observed something
else, and nothing in the trace said so. Index 0 carries the startup gate's
verdict, so a run that never started is a trace holding that record and nothing
else rather than a run that explored and found nothing.

* fix(runner): budget the foreground gate in time, not in polls

Eight polls is not a budget. Each poll costs whatever the driver's idle wait
happens to take, so the same launch cleared the gate on one device and
exhausted it on another: across 80 runs of one app, the gate reported "app
never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16
runs, and it was wrong every time. On API 34 settleForForeground returned in
~100ms, so the eight polls gave up 1.2s into a launch whose window drew at
~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14
runs then spent their first step on the launch animation instead of the app,
which is the one-step offset that came out of that campaign looking like a
platform difference.

The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same
duration on every device, and a verdict of "not in front" ends the run instead
of warning and carrying on: a run that never got its app on screen holds no
evidence about the app, and the trace records why at step 0.

* test(runner): the gate keeps looking until its budget runs out

Locks the three facts the campaign was missing: a window that draws after more
polls than the old count allowed still clears the gate, an app that never comes
forward ends the run with a typed error, and both the startup verdict and every
mid-run step the guard could not recover are readable off trace.jsonl.

* feat(campaign): count the runs that were never in the app

A run that failed its precondition has zero steps and no violations, which is
what a short clean run looks like too. The summary now counts the trace records
naming an unmet precondition, so a campaign directory answers "how many of
these were never in the app" without grepping any log.

* docs(triage): name the trace field a run that never started leaves

* fix(selectors): tag names the whole tag, not a substring of it

matchSelectorKind had no case for tag, so it fell through to the raw
attribute path and matched by substring. web-runtime.ts compiles tag to a
CSS type selector, so tag:li resolved to <todo-list> on the Go side and to
nothing on the web side.

* test(chrome): both resolvers agree on tag where a container's name contains its child's

* fix(make): build the binary instead of matching the build directory

build/ exists at the repo root, so make build was satisfied by the
directory and left a stale bin/sanderling in place.

* feat(verifier): expose the property names a loaded spec registered

* feat(testrun): refuse a run against a spec that registers no properties

A spec with no properties drove the app and reported no violations,
which is indistinguishable from a spec that judged something and found
nothing. Execute now aborts after loading the spec unless the run asks
for the opt-out by name.

* feat(cli): --allow-no-properties opts a run out of the refusal

* docs(cli): document --allow-no-properties

* feat(bundle-check): fail a spec that bundles but registers no properties

* test(bundle-check): cover the zero-property refusal and pin the reported bundle

* feat(folio-web): predicates for counting commits against submit actions

* feat(folio-web): judge one commit per submit over a home-card window

Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta,
which compared two consecutive steps on one screen and so could not see a
double submission that lands across a navigation.

* fix(folio-web): keep submit live for 400ms after saving

Defers the navigation back so the button is tappable while the label
reads Saved, widening the double-submit window the counting property
is there to catch.

* feat(confusion-matrix): score the checker against a blind reviewer

Cross-tabulates the properties that fired against the human verdict, one
cell per implementation, over a sweep whose implementations all passed
their own generated tests. An implementation that failed to build, has no
usable run, or carries no filed verdict is listed as missing data rather
than counted as a clean cell.

Landing the package in one commit because the intermediate splits would
not link.

* test(confusion-matrix): reject malformed inputs and keep missing data out of the cells

* test(confusion-matrix): cover cell assignment, precision and recall

* fix(chrome): focus descends into the shadow root

document.activeElement names the host, not the node focused inside it, so
a Compose-for-wasm app that mounts its tree in a shadow root reported
focus on div#app forever. confirmFocus could never be satisfied and every
InputText step aborted the run after three tries.

selectAllScript already descends the boundary; the tree builder did not.

* test(implementation-sweep): supply the binaries the missing-binary test does not test

resolveBinaries ranges a map, so with more than one binary absent the
error named whichever it reached first. The test passed locally only
because bun and sanderling were on PATH; on CI it was a three-way coin
flip.

* fix(replay-ui): read data-* attributes by their markup names

The web runtime now publishes raw markup attribute names, so attrs["step"]
read nothing where the markup writes data-step. Three properties went
vacuous and exactlyOneStepIsSelected reported false against a UI that was
fine.

The test also fails if a dataOf key gains no matching attribute, or if an
attribute it derives is rendered nowhere.

* fix(web-runtime): focus descends into the shadow root here too

The Go driver already descends the boundary; the V8 host did not, so the
two enumerations disagreed about focus on any shadow-mounted app.

The harness now answers activeElement the way a real root does: a root
names a node of its own tree, so only the shadow root itself names the
field.

* fix(implementation-sweep): name every missing binary, in flag order

Ranging a map returned at the first failure, so an operator missing three
binaries was told about one, fixed it, reran, and was told about the next.
The function exists to stop the sweep once rather than fail per
implementation and seed.

Two identical runs also printed different errors, which is why this
reached master as a flake instead of a clean red.

* fix(chrome): focus follows the caret to the field it types into

Compose for wasm never focuses the semantics node carrying the testTag.
It proxies keystrokes through a hidden 1px backing input that is a
sibling of the a11y tree, so the node the runner tapped never held focus
and confirmFocus refused to type into every Compose text field.

Focus is re-attributed to the smallest editable whose box holds the
caret's centre. Centre-point rather than full containment because the
caret's height comes from the text style and the field's from its layout
box, so a taller font would silently drop back to refusing.

* fix(corpus-sweep): name every missing binary, in flag order

Same map-ranging bug as the sibling tool, and this copy had no test on
the missing-binary path at all.

* fix(web-runtime): a handle answers editable for itself, not its container

isContentEditable is inherited, so every span inside a contenteditable
div called itself typeable. collectTargets and the chrome dump both
require the element itself to match; the handle was the one that did not.

* test(chrome): a hinted field is not named by its css class

The fixture inputs carried no class at all, so the test could not fail
the way the bug did. They now carry folio-web-shaped classes, and the
test asserts the editable gate the hint is read behind.

* test(chrome): the handle and the enumeration agree on editable too

The helper compared clickable alone, so the inherited-contenteditable bug
was caught by unit test only and never in a real browser.

* fix(web-runtime): focus follows the caret to the field it types into

Mirrors the driver, so the two hosts agree about focus on a Compose page.
The harness inherits custom properties down the parent chain the way CSS
does, so an implementation matching the inline style attribute fails.

* fix(campaign): name every missing required flag, in flag order

Five required flags ranged as a map, so omitting three told the operator
about one, chosen at random.

* fix(corpus-sweep): name every missing required flag, in flag order

* fix(implementation-sweep): name every missing required flag, in flag order

* fix(confusion-matrix): name every missing required flag, in flag order

* ci: pin the idb-companion tap to the formula the companion is staged from

The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and
prepare.sh stages bin/ and Frameworks/ as siblings because the binary
resolves through @rpath. Floating on it also made the hard-coded
companion-1.1.8 output name a lie.

The ios-assets cache does not cover this: it restores and make rebuilds
anyway, because checkout stamps prepare.sh newer than the archived
tarball. Master was green only because its last run predated the bump.

* fix(campaign): refuse to start on a device that is not there

A sweep launched at six serials, three of which had been deleted from the
host. 19 of 20 runs were lost, and not because half the devices were wrong:
a worker on a dead serial fails in about 31 seconds and immediately pulls
another seed, so three bad workers drained sixteen seeds while the three good
workers were still inside their first run.

Fast failure is more dangerous than slow failure, because the fast failure
consumes the resource the slow one would have left alone.

Preflight names every missing serial before the first seed is dispatched.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* fix(campaign): quarantine a device that keeps failing fast

Preflight cannot catch a device that disappears mid-sweep, which is what
happened: the serials were alive the previous day. Three consecutive failures
under two minutes, with no run that worked in between, is a property of the
device and not a coincidence.

The manifest records which device was quarantined and which seeds have no
result, so an aborted sweep says so in its own artefact.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* feat(trace): record the device a run executed on

meta.json carried the host but not the device, so a trace could not say what
hardware produced it without the campaign manifest beside it. An experiment
splitting cells across api levels could only join them through that manifest.

Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX

* ci: let a restored ios bundle survive make's mtime check

The cache restored and the build ran anyway: a restored tarball keeps the
mtime it was archived with while checkout stamps the sources, so make read
every bundle as stale. Both logged Cache hit and rebuilt regardless.

Dating the bundles after their sources fixes the lie where it is told.
Order-only prerequisites would have fixed it in make, but a laptop has no
cache key, so editing prepare.sh would silently embed the previous tarball.

The formula version joins the key because a hit now decides what gets
embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2
runs shared a key.

* fix(confusion-matrix): a campaign that died is missing data, not a true negative

The sweep-level loop excluded a run on launch_error alone, while
excludedBecause already checked the campaign process's exit code. An
interrupted campaign wrote exit_code -1 with an empty launch_error, so
its one completed seed scored the implementation as a clean cell on a
tenth of the planned evidence.

The fixture builder wrote one exit code into both the sweep record and
the campaign run record, which is why no test could tell the two levels
apart.

* fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets

A run stops at whichever comes first, the step budget or --duration, so
a clean run that reached the wall clock exited with fewer steps than the
budget and was still credited with the whole of it. The model arm pays a
network call and a screenshot per step, so it reaches the wall sooner and
was handed exposure it never had.

Nothing checked that two arms shared a budget either. Thirty identical
clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14
from the rank-sum while the log-rank in the same report read p 1.0000.
groupArms already refused this within one arm.

The claims the old convention left in comments and report lines are
corrected rather than left standing beside the new behaviour.

* fix(runner): a source that was asked and handed nothing says so

NextAction returning ErrNoAction left the step with no skip reason, so a
run whose every model call failed on transport, a non-2xx, an empty
choices array or an echo mismatch printed no violations and exited 0.
Only llm-calls.jsonl knew it had never touched the app.

The reason now travels the path the other five already take, so it
reaches the trace, the summary, and the campaign's dispatched-action
exclusion. A held step never asks and keeps carrying nothing.

* feat(testrun): refuse a run that dispatched none of its actions

Same argument as the zero-property refusal: an instrument that drove
nothing must not report a clean result. A first-screen violation still
wins under --exit-on-violation, --allow-no-properties exempts the
extraction sweeps that measure reach rather than judge, and one
dispatched action is enough, so a generator quiet on some screens is
untouched.

* docs(cli): document --label-source

* docs(spec-language): name the hintText selector's host divergence

The line said the key matches placeholder alone, which is true of the web
runtime and not of the tree, where it resolves against the derived
attribute. A spec author reading it wrote a selector that matched on one
host and not the other.

* feat(bundle-check): --allow-no-properties opts out of the refusal

The run path grew the opt-out and the freeze gate did not, so a spec the
extraction and portability sweeps register nothing for on purpose could
be run but never frozen. The refusal now names the flag the way the
runner's does.

* test(verifier): an unreadable committed fixture fails, it does not skip

The comment said the round trip always runs. A skip on a fixture that is
committed turns a missing or truncated file into a green.

* fix(testrun): the refusal asks whether the generator drove, not whether anything did

A dead provider against folio exited 0 on a real emulator: the login
setup dispatched three actions before the generator was consulted, so
DispatchedActions was 3 and the gate never fired while the generator
drove the app zero times across 83 steps. Any spec with a login setup
was immune, which is the normal case.

Summary counts generator actions separately and the refusal reads that.
NoActionsDispatchedError becomes NoGeneratorActionsError, because a run
that dispatched three login taps was lying in the old name.

* feat(runner): the summary says how many steps the generator drove

A green llm run carried no evidence of how much the generator actually
drove: the count was inferable only from llm-calls.jsonl outcomes, and
the number the refusal turns on was invisible in the run's own output.

* fix(testrun): an ios run records the simulator it executed on

Device was read from --device, which only an android run sets, so every
ios meta.json left the field empty and the trace could not say what
hardware produced it.

* fix(campaign): the action count leaves the setup's login out on a model run

Defects per thousand actions divided by every dispatched step, so a
spec whose setup logs in inflated the denominator by however many steps
that took. It is the same error the run gate had, and it does not cancel
between arms.

A model run is separable because only an llm-selected action stamps
next_action.source. A seeded run is not: its setup returns through the
same entry with no marker, and 11261 dispatched steps across the 169
recorded runs carry no source at all, so excluding on it blind would
report every seeded run as having explored nothing. The seeded arm
counts as before and a test pins that.

* feat(hierarchy): an element reports whether it masks what is typed into it

ios reads it off SecureTextField, which the companion already sent and
nothing read; web reads input[type=password]. Android cannot: the native
tree mapper drops the password attribute before the sidecar sees it, so
the fact is three-valued and null there rather than a false that would
read as "not secure".

* fix(verifier): a secure field's typed value never reaches the record

A folio login run wrote the account email and password in cleartext into
llm-calls.jsonl, 166 times in one run, beside screenshots of the same
screens. Three sites rendered it: the recent-action memory, the candidate
list, and the trace. One helper now covers all three so a fourth cannot
bypass it, and the driver still receives the real text.

Android redacts every typed value because it cannot tell a secure field
from any other. That asymmetry is deliberate and documented: safe by
default on the target that cannot tell.

* fix(runner): a secure field's value does not reach state.lastAction either

folio extracts lastAction, and extractor values are persisted as
extractor_changes, so the password still reached the run directory
through the spec after the three render sites were closed.

The wrap sits in the runner rather than in lastActionFields because the
hosts hold the next step's tree, not the one the action was chosen
against: a field that stops being secure between the two would publish
what the trace withheld. Live and replay now agree byte for byte.

* fix(trace): an action names the generator that produced it

The setup exclusion landed for the model arm only, because only a model
pick stamped a source. A seeded run returned setup's action through the
same entry with no marker, so its denominator still counted the login
while the model arm's did not, and the two are compared.

serializeAction names setup and seeded on the wire, so both arms are
counted by one rule. An already-recorded trace names nothing and keeps
exactly the count it was reported with; unattributed_actions counts those
steps so the old denominator cannot pass as the new one. TraceVersion is
deliberately unbumped: oracle-reduction refuses a differing version, and
a bump would make all 169 recorded runs unreplayable.

* fix(defect-identity): degrade a redacted origin action to its selector

The full action key read the typed value straight from the trace, where
redaction renders every value typed into one field as the same string, so
two runs that typed different values there collapsed into one identity and
the report said nothing about it. The key now drops a redacted value, falls
back to the selector for that action, and counts the rows it did that to, so
the undercount reads as an undercount.

* fix(campaign): a record always says how many actions named no producer

An omitted count reads the same as a run recorded before actions carried a
source, so the two cannot be told apart by anything downstream.

* fix(analyze): read how much of a record's action count names no producer

A runs.jsonl written before actions named one has no field, and its whole
count is of unknown provenance rather than none of it.

* fix(analyze): refuse to compare attributed and unattributed denominators

One arm's actions may include the login the spec's setup drove and the
other's cannot, so a per-action rate over the two divides by different
things and the tests rank the bookkeeping.

* fix(analyze): mark an action count of unknown provenance in the report

* docs(manual): what an action count with no producer means for a rate

* fix(folio): install through adb so a remote adb server works

Gradle's install task talks to adb through ddmlib, which reads only
ANDROID_ADB_SERVER_PORT and dials the loopback address, so
ADB_SERVER_SOCKET never reaches it and `just test` could not touch a
remote emulator. Gradle now only assembles the APK and adb does the
install, which picks up the same server every other call in the run
talks to.

* docs(folio): say how to point just test at a remote adb server

* test(conformance): the g4 fixture holds what a redacted android trace holds

Android reports no secure fact for any field, so every InputText it records
writes the redaction placeholder rather than the typed value. The fixture still
carried the real value, which is the only reason the gate reported itself as
catching the doubling. Two more fixtures come with it: a repeated-character
corpus value that reads as its own doubling and must not fail, and a backend
that does record the typed value.

Red at this commit: G4 reports PASS on a doubled field it cannot see.

* fix(testrun): a recorded violation outranks the dead-run refusal

A campaign never passes --exit-on-violation, so the refusal was discarding
runs that had found something: exit_code 1 in the record and the analysis
drops them as missing data. A run that recorded a violation holds a verdict,
which is the whole reason the refusal exists.

* fix(testrun): the dead-run refusal gets its own opt-out

--allow-no-properties was waiving two unrelated refusals, so a sweep passing
it for the property-free reason silently lost a detector it never asked to
disable, and a run with properties could only get the dead-run exemption by
claiming one it did not want.

* feat(cli): --allow-no-generator-actions

The flag the dead-run refusal names, wired through to the pipeline. The
property-free flag goes back to meaning what it says.

* refactor(analyze): open the log-rank up to a weight on the risk set

The log-rank is one member of a family that differs only in how much each
event time counts. Nothing else changes: the counts it reports stay counts
whatever the weight, and the published-dataset results are unmoved.

* feat(analyze): add the gehan generalized wilcoxon test

The rank-sum carried over to right-censored samples: every pair of runs is
scored by which one outlived the other, and a pair censoring cannot order
counts as half rather than as a difference neither run supports. The effect
size and the p-value are the same statistic, and with nothing censored both
are exactly what the rank-sum reports.

* fix(analyze): compare arms on censored runs, not on flattened step counts

stepTimes threw the censoring flag away and handed the rank-sum a plain
number per run, so a run the wall clock stopped at step 12 was ranked as one
that violated at step 12. That was defensible while every clean run sat at
the budget, the largest value any run could take, and it stopped being
defensible when a clean run started being censored where it stopped.

Twenty runs clean at step 12 against twenty violations at step 100 read a12
0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank
reading p 1.0000. The pairwise comparison is now the Gehan test over the
observations themselves, and the report says how many run pairs censoring
left with no order between them, which is how much of the effect size is the
null value rather than an observation.

* fix(conformance): g4 reads a doubling off the observed field value

The typed value stopped reaching the trace on any target that reports no
secure fact for the field, which on android is every field, so the gate was
comparing the redaction placeholder against itself and passing whatever the
driver did. The observed value is not redacted, and a field holding one string
twice over is the doubling itself. A value that is a single character repeated
stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither
can be told apart from its own doubling.

The recorded-value check stays for the targets that do record it, where it also
catches a doubling appended to content the field already held.

* fix(spec): a secure selector names the password field on web

secure is derived from the field type, not written by the markup, so
matching it as a raw attribute reached nothing: the key is accepted, no
unknown-key error fires, and find answered undefined on web for the
field it answers with on ios. false is every editable field that is not
a password entry, since an element that is no field reports null and
answers to neither value.

* test(chrome): resolve the secure selector on both matchers

the fixture covers the password entry, the three shapes of editable
field that are not one, and a checkbox that is no field at all.

* test(chrome): compare the secure fact across both producers

it is the fourth fact the dump and the web runtime derive independently,
and the one that decides whether a typed value is written into the
shared record. three-valued, so the fixture guard requires all three
states rather than both polarities.

* docs(manual): state what a secure selector matches

* test(conformance): g4 keeps checking past an input typed at coordinates

An InputText that names no field aborts the analyzer, so the gate reports the
whole run as failed and checks none of the steps after it. 129 of the 485
recorded traces hold such a step.

Red at this commit: jq stops on a null selector and the gate reports FAIL.

* fix(conformance): g4 skips an input that names no field

jq splits an empty string into no segments, so reading the last one off an
action typed at coordinates threw and took the rest of the run's steps with it.
Such a step names nothing to check; the gate now passes over it and keeps
checking the ones that do.

* test(browser): the exit code a dead run and a violated one actually leave

Drives the built binary against a page with nothing to tap and reads the
process status, then the same run through campaign to pin what lands in
runs.jsonl: exit_code 1 there is a detection the analysis drops as missing
data.

* fix(spec): keep a secure selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, secure}
built '[id="pwd"]input[type="password"]' and querySelectorAll threw.

* fix(analyze): score a seed pair by which run outlived the other

The paired path had the same defect as the unpaired one: it subtracted two
step counts and handed the differences to the signed-rank test, so a pair
holding a run the wall clock stopped at step 12 entered as a difference
neither run supports. Twenty seeds where the first arm was still clean at
step 12 and the second violated at step 5 in six of them read sign -1 and
p 0.0011, pointing at the arm that never violated.

A pair is now scored the way the unpaired comparison scores one and tested by
the exact sign test over the pairs whose order censoring determines, which is
what the log-rank stratified by seed reduces to here. The signed-rank goes
with the differences it needed: a magnitude-based paired test wants a
difference from every pair, and the arms censor on different clocks. The
median difference stays, over the pairs where both runs violated, and says so.

* docs(analyze): name the tests the tool actually runs

The --paired flag advertised the signed-rank, two comments and a test message
still said rank-sum, and nothing said what rankSum is doing in the tree now
that no campaign reaches it.

* docs(manual): exit 1 also means a run that holds no verdict

And the flag the dead-run refusal now names, which --allow-no-properties
used to double as.

* docs(skills): quote the summary line the runner prints now

The setup skill's empty-page claim was the stale one that mattered: that run
records no_action_produced on every step and exits 1, it does not sit at
exit 0 with no violations. Numbers remeasured against the counter and
throwing fixtures.

* test(conformance): g4 sees a doubling appended to what the field held

Redaction cost the gate this shape on android: the driver typed the value twice
onto existing content, so the whole value is not its own doubling and the typed
value is not in the trace to compare against. The recorded-value check still
catches it on the backends that record one.

Red at this commit: G4 reports PASS on a field that grew by one string twice.

* refactor(analyze): hoist the sign test's loop bound

* fix(conformance): g4 reads a doubling out of what the field grew by

The whole-value check misses a driver that typed the value twice onto content
the field already held, which is the append-vs-replace shape the recorded value
used to catch before it was redacted. What the field grew by over the snapshot
the action was chosen against is the same signal and needs no typed value.

Checked against every recorded trace under conformance/runs: 485 traces, 299 of
them carrying an InputText, none newly failing.

* fix(analyze): write an undefined paired p-value as null, not as NaN

A paired contrast where censoring orders no pair has no p-value, and JSON has
no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN'
and wrote no summary at all after printing a complete report. The two fields
join the medians and the rates already carried as pointers, undefined reading
as null in the summary and n/a in the report.

Reachable since a clean run started being censored where it stopped: an arm
the wall clock stops before its partner ever violates orders nothing.

* fix(spec): a boolean state selector names what the live element reports

clickable, enabled, focused, checked and selected are derived from the
element rather than written by the markup, so matching them as raw
attributes built [clickable="true"] and reached nothing: the keys are
accepted, no unknown-key error fires, and the worked example in
docs/manual/spec-language.md found no element on web and passed having
checked nothing.

Each key is answered by the same function elementHandle derives the fact
with, since no CSS says what any of them says: :focus names the shadow
host of a focused field as well, :checked misses a checked custom
element and answers for a selected option besides, and [checked] is the
state the page loaded with rather than the one the user left it in.

* fix(spec): keep a tag selector valid beside another key

a multi-key object selector concatenates its parts into one compound,
and a type selector is valid only at the head of one, so {id, tag} built
'[id="amount"]input' and querySelectorAll threw. whether a spec got an
exception or an element depended on the order its author wrote the keys
in.

* fix(chrome): state every boolean flag the dump can state

internal/hierarchy writes the attribute a selector matches on only where
the producer stated the flag, so a state emitted as null is one no
selector can ask about: {clickable: false} and {enabled: false} matched
nothing at all against a web dump while matching on android, which
states every flag both ways. only secure stays three-valued.

* test(chrome): resolve the five state selectors on both matchers

the fixture differs one state at a time: a disabled button and an
aria-disabled role control, a box ticked by script with no checked
attribute beside one cleared by script that has it, and a select whose
first option is selected without the markup saying so anywhere.

half the states are asked inside one container, because a state the
whole page has an opinion about answers with most of the document and a
want list nobody can check by reading.

* test(chrome): compare checked, selected and focused across both producers

the target enumeration carries none of the three, so they reach a spec
through the ax handle alone, and a selector naming one of them resolves
against that same reading. the shadow fixture holds the focused control
inside its shadow root, where document.activeElement names the mount
element and only a producer that descends finds the field.

* fix(spec): keep a selector out of the head subtree

the head renders nothing, so the hierarchy dump drops it and so does the
enumeration the picker walks, but a selector still resolved into it: a
whole-page findAll answered with <head> and <title> here and with
neither on the goja host, which is a divergence the moment a state
selector asks a question every element has an answer to.

* docs(manual): state what the other boolean state selectors match

* fix(folio): refuse to install and fuzz a device nobody named

adb falls through to the local server when ADB_SERVER_SOCKET is unset, and
claims the only device attached there. That could be a personal handset, and a
run installs the app, clears its state and fuzzes it. Every recipe that touches
a device now resolves the target through _require-device, which only picks on
its own when a single local emulator is all adb sees.

* docs(folio): state that android recipes need ANDROID_DEVICE

* fix(spec): and text with the keys written beside it

a compound object selector dropped text and matched on the other keys
alone, so {testTag: "Row", text: "Alice"} selected every row carrying
the tag where internal/hierarchy selects the one row the author named.
matching more than the spec said is silent: the find lands on a row
nobody wrote and every property over it still passes.

text is answered against the element the way the boolean states are,
since css cannot ask what an element's text says and the xpath that can
cannot ask about the rest, and the innermost rule now holds over what
the whole selector matched, where internal/hierarchy holds it. a
text-only selector still compiles to the same innermost xpath.

* test(spec): pin text against the key beside it in either order

object keys iterate in insertion order, so the order the author wrote
them in decided what a compound selector meant. the innermost rule is
pinned over the whole selector's matches: a row whose badge carries the
class and the text both is dropped, one whose badge carries the text
alone is kept, and a state key is anded before either.

* test(chrome): compare a compound text selector across both matchers

one page, both resolvers, text written before and after the key beside
it. the object form now encodes its keys in the order the filters state
them rather than the order a map iterates, so both orders are asked.

the row and the badge under it share a class so the innermost rule has
something to drop, and {text, clickable} pins that text is anded before
that rule runs: the innermost element carrying "January" is the option,
and the select is the only element that is both.

* docs(manual): state how text combines with the key beside it

the object selector section said every pair must match without saying
where the innermost rule then lands.

* fix(hierarchy): reach the class attribute through className

className is an accepted selector key that no producer writes: android
reports the view class, ios the element type and the chrome dump
el.className, all of them under `class`. With no alias onto that key the
selector matched NOTHING here on every platform while the web runtime
resolved it against the live DOM, so {className: "status"} named the row
and the badge on one host and no element at all on the other.

The failure is silent: the key is accepted, so no unknown-key error
fires, and a property over the element that was never found passes
having checked nothing.

* test(chrome): compare className across both matchers

one page, both resolvers, the two names for the one attribute. class is
asked beside className so the pair is pinned to the same elements rather
than each to itself: the row and the badge under it both carry it.

* test(spec): pin className and class on the same elements

this host answers both names against the live DOM and internal/hierarchy
now aliases the second onto the first, so a name dropped from the table
here would match nothing on web while the dump still answers it.

* docs(manual): list className among the cross-platform aliases

the key was already typed on the spec surface and already resolved on
web, and the alias table said nothing about which attribute it reads.

* fix(hierarchy): reach the accessible label through every name for it

label and accessibilityLabel aliased onto accessibilityText alone, which
only the ios sidecar writes, and alias expansion is ONE level: the hop
from accessibilityText to content-desc was never taken, so both keys
matched nothing on android and on the chrome dump, which write the fact
under content-desc. ariaLabel and contentDescription aliased onto
nothing at all and matched nothing anywhere.

The web runtime resolves all four against the live DOM, so a selector
naming a field this way found it on one host and no element at all on
the other. The keys are accepted, so no unknown-key error fires, and a
property over the element that was never found passes having checked
nothing.

Each name lists both keys rather than chaining through accessibilityText:
transitive expansion would silently widen every existing key at once.

* fix(hierarchy): reach a web test tag through testTag and testID

Compose for Web writes a test tag as data-testid, which is what the web
runtime resolves both names against. testTag aliased onto the three
identifier keys and not that one, and testID aliased onto nothing at
all, so a tag the web runtime found on every row of a list named no
element here and every property over it passed vacuously.

* fix(spec): resolve the identifier, label and class aliases against the DOM

identifier, accessibilityIdentifier, accessibilityText and elementType
are the names ios writes four facts under, and internal/hierarchy
aliases each onto the key the other producers write. This table listed
none of them, so each fell through to a raw attribute lookup and built
[accessibilityIdentifier="summary_card"], which no element carries.

Every one of them resolved against the dump on the goja host and named
nothing here. The keys are accepted, so no unknown-key error fires, and
a property over the element that was never found passes having checked
nothing.

* fix(spec): an editable or scrollable selector names what this host derives

Both facts are derived from the live element rather than written by the
markup, and matching them as attributes built [editable="true"], which
no page carries. Both resolve against the dump on the goja host, so a
spec naming a field or a scroll container that way found it there and no
element at all here, with no unknown-key error to say so.

Each reads the same function the fact is derived with, so a selector
cannot name an element this host calls something else: the handle, the
picker's target list and the editable selector all go through
isEditable, and scrollable reads the overflow test collectTargets reads.

scrollable false names nothing rather than every element that does not
scroll: both producers state the fact only where it holds, the way an
element that is no field at all answers to neither value of secure.

* test(chrome): compare the alias keys and the two derived facts

one page, both resolvers, the ten names that resolved on one host only.
each alias is asked beside the key it resolves through, so the pair is
pinned to the same elements rather than each to itself.

the page grows a container that overflows its box and a neighbour that
does not, because scrollable is derived from the box: without one the
only scrolling element on the page is the document root, whose answer
moves with the window.

* fix(hierarchy): bounds is a raw attribute, not a cross-platform key

Every native dump writes the rectangle out as a string under bounds, and
no DOM element carries an attribute of that name, so the key resolved
against the dump and matched nothing on web on every page there is. It
is accepted, so no unknown-key error said so, and no mapping can be
invented for it: there is no DOM fact to map it to.

Off the accepted list the web runtime raises the unknown-key error
instead of matching nothing in silence, and the key still resolves
wherever a producer writes it, through the escape hatch every other raw
attribute already uses: a key some element carries is a key that can
match, on both sides.

* docs(manual): state which attribute each alias reads, and what bounds is

the table listed neither name for the accessible label that a web page
writes, nor the key a web test tag lands on, and said nothing about
elementType. editable and scrollable are boolean states like the rest,
and scrollable is the one of them the platforms state only where it
holds. bounds is a raw driver attribute rather than an accepted key.

* docs(hierarchy): the package doc names every key an alias reaches

it described the alias table as it stood before the label and test-tag
names reached the keys android and web write, and said nothing about
expansion being one level deep, which is why each name has to list every
key rather than hop through another alias.

* fix(spec): a hint selector names the ladder both producers derive

hintText and placeholderValue are the accessible-name ladder, derived
from the live element, and compiling them to [placeholder="..."] made
them name the wrong field or none at all. A field labelled by an
aria-label or a bound <label> carries no placeholder, so it resolved
against the dump on the goja host and reached nothing here; one carrying
both answered to its placeholder here where the dump answers to its
aria-label, which lands a find on an element nobody named.

Both keys read the same fieldHint elementHandle and the hierarchy dump
(internal/driver/chrome/driver.go) derive the fact with, so a selector
cannot name a field this host calls something else. An empty hint names
nothing rather than everything that is no field: both producers write
the fact only where the ladder answered.

placeholder stays the attribute the markup writes, which is what the
dump carries under that name too, so a field whose hint is something
else still answers to it on both hosts.

* fix(chrome): a hint target is not tapped by the placeholder attribute

TapSelector is a third resolver, and it built [placeholder="..."] for
hintText and placeholderValue too. Now that both matchers read the
accessible-name ladder, that CSS names a field whose hint is its
aria-label and whose placeholder happens to carry the value, which is an
element neither matcher named.

No CSS says what the ladder says, so both keys fall through to a match
that reaches nothing and the step fails naming the selector, the way
every other derived key in this file already does. A selector reaches
here only where the dump resolved it to no coordinates at all.

* test(chrome): compare the hint keys and placeholder across both matchers

The page gains four fields that differ one rung at a time: a bound
label, a placeholder, a placeholder an aria-label outranks, and the name
the form gives the field. Only the placeholder rung was reachable
before, so hintText and placeholderValue named a field on the goja host
and no element at all on web for the other three, and named the field
here and nothing there for the rung the ladder passed over.

placeholder was measured empty on both hosts because nothing on the page
carried the attribute, which said nothing about it. It now names the
field the markup wrote it on and not the field whose hint is its
aria-label.

The third resolver reads the same selectors: what TranslateStringSelector
builds for a hint key has to match nothing over CDP rather than the field
carrying the value as a placeholder.

* docs(manual): both hosts read the hint ladder, placeholder is the attribute

The web section said the hintText key does not read the ladder on both
hosts and told authors to select such a field by attrs.hintText instead.
Both hosts read it now, so that instruction is gone rather than left
standing beside a newer sentence.

placeholder is stated as the attribute the markup writes and nothing
more, the tap path is stated as failing by name where no CSS says what
the ladder says, and the alias table gains the row it was missing.
This commit is contained in:
pj authored and GitHub committed 2026-08-19 11:07:49 +05:30
1 parent 95a19fcc1e
commit 9b4ff5f247
243 files changed
+33472 -1062

No files matched your search

+318
View File
@@ -0,0 +1,318 @@
package main
import (
"fmt"
"math"
"slices"
"time"
)
type armSummary struct {
Arm string `json:"arm"`
Generator string `json:"generator,omitempty"`
Platform string `json:"platform,omitempty"`
StepBudget int `json:"step_budget"`
Directories []string `json:"directories"`
Recorded int `json:"recorded_runs"`
Usable int `json:"usable_runs"`
Violated int `json:"violated_runs"`
Censored int `json:"censored_runs"`
Excluded int `json:"excluded_runs"`
ExcludedByReason map[string]int `json:"excluded_by_reason,omitempty"`
MissingSeeds []int64 `json:"missing_seeds,omitempty"`
EventsHeldAtBudget int `json:"events_held_at_budget"`
EventsDetectedAfterOrigin int `json:"events_detected_after_origin"`
MedianStepsToFirstViolation *float64 `json:"median_steps_to_first_violation"`
FirstQuartileSteps *float64 `json:"first_quartile_steps_to_first_violation"`
ThirdQuartileSteps *float64 `json:"third_quartile_steps_to_first_violation"`
SurvivalCurve []survivalPoint `json:"survival_curve,omitempty"`
ViolationRate *float64 `json:"violation_rate"`
TotalSteps int `json:"total_steps"`
TotalActions int `json:"total_actions"`
UnattributedActions int `json:"unattributed_actions"`
TotalRunHours float64 `json:"total_run_hours"`
Detections int `json:"detections"`
DefectsPerThousandActions *float64 `json:"defects_per_thousand_actions"`
DefectsPerHour *float64 `json:"defects_per_hour"`
DistinctDefects int `json:"distinct_defects"`
SingletonDefects int `json:"singleton_defects"`
SingletonFraction *float64 `json:"singleton_fraction"`
DefectRunCounts map[string]int `json:"defect_run_counts,omitempty"`
}
type pairwiseResult struct {
First string `json:"first"`
Second string `json:"second"`
FirstSize int `json:"first_size"`
SecondSize int `json:"second_size"`
Statistic float64 `json:"u"`
A12 float64 `json:"a12"`
// Unordered is how many of the run pairs behind U and A12 have no order
// between them, because they were tied or because censoring stopped one run
// before the other violated. Each of them counts as half, so it is also how
// much of the effect size is the null value rather than an observation.
Unordered int `json:"unordered_pairs"`
PValue float64 `json:"p_value"`
HolmPValue float64 `json:"holm_p_value"`
}
type analysis struct {
GeneratedAt time.Time `json:"generated_at"`
Outcome string `json:"outcome"`
// Question names the family Holm corrects within. The correction is applied
// across the comparisons of one research question and never across the
// paper, so the family a p-value was adjusted in has to be recorded next to
// it rather than left to the reader to reconstruct.
Question string `json:"question,omitempty"`
HolmFamilySize int `json:"holm_family_size"`
Arms []armSummary `json:"arms"`
LogRank *logRankResult `json:"log_rank"`
Pairwise []pairwiseResult `json:"pairwise"`
Paired *pairedComparison `json:"paired,omitempty"`
Notes []string `json:"notes,omitempty"`
}
const outcomeDescription = "steps to first violation, right-censored at the last step a clean run reached"
func analyse(arms []arm, now time.Time) (analysis, error) {
result, testable, err := baseAnalysis(arms, now)
if err != nil {
return analysis{}, err
}
if len(testable) >= 2 {
result.Pairwise = comparePairs(testable)
result.HolmFamilySize = countCorrected(result.Pairwise)
}
return result, nil
}
// analysePaired is the seed-matched design of the actuation ablation: two arms
// running the same seeds, contrasted seed by seed rather than as two
// independent samples.
func analysePaired(arms []arm, now time.Time) (analysis, error) {
result, testable, err := baseAnalysis(arms, now)
if err != nil {
return analysis{}, err
}
if len(testable) != 2 {
return analysis{}, fmt.Errorf("a paired comparison needs exactly two arms with usable runs, found %d", len(testable))
}
comparison, err := pairArms(testable[0], testable[1])
if err != nil {
return analysis{}, err
}
if comparison.Pairs == 0 {
return analysis{}, fmt.Errorf("arms %q and %q share no seed with a usable run in both",
testable[0].Name, testable[1].Name)
}
if comparison.PValue != nil {
adjusted := holm([]float64{*comparison.PValue})[0]
comparison.HolmPValue = &adjusted
result.HolmFamilySize = 1
}
result.Paired = &comparison
return result, nil
}
func baseAnalysis(arms []arm, now time.Time) (analysis, []arm, error) {
result := analysis{GeneratedAt: now, Outcome: outcomeDescription}
for _, current := range arms {
result.Arms = append(result.Arms, summarize(current))
}
var testable []arm
for _, current := range arms {
if len(current.observations()) > 0 {
testable = append(testable, current)
}
}
if len(testable) < len(arms) {
result.Notes = append(result.Notes,
"arms with no usable runs are reported but left out of the log-rank test and the pairwise comparisons")
}
if err := sameBudget(testable); err != nil {
return analysis{}, nil, err
}
if err := sameAttribution(testable); err != nil {
return analysis{}, nil, err
}
if len(testable) >= 2 {
names := make([]string, len(testable))
groups := make([][]observation, len(testable))
for index, current := range testable {
names[index] = current.Name
groups[index] = current.observations()
}
test := logRank(names, groups)
result.LogRank = &test
}
return result, testable, nil
}
// sameBudget refuses arms that were given different exposure. A clean run is
// censored somewhere at or below its arm's budget, so the arm with the larger
// budget carries censored runs the smaller arm could not have produced, and
// every test that ranks the two against each other reads that as the arm
// surviving longer. It is the cross-arm form of what groupArms already refuses
// within one arm.
func sameBudget(arms []arm) error {
for index := 1; index < len(arms); index++ {
if arms[index].Budget != arms[0].Budget {
return fmt.Errorf("arm %q has step budget %d and arm %q has %d: "+
"runs censored at different budgets cannot be compared",
arms[0].Name, arms[0].Budget, arms[index].Name, arms[index].Budget)
}
}
return nil
}
// sameAttribution refuses arms whose actions were counted against different
// denominators. An arm recorded before an action named its producer counts
// whatever the spec's setup dispatched among its actions, and an arm recorded
// after leaves the login out, so the same rate over the two divides by
// different things and the tests rank a bookkeeping difference. Two arms of the
// same unknown provenance are diluted alike and compare; one of each does not.
func sameAttribution(arms []arm) error {
for index := 1; index < len(arms); index++ {
unknown, attributed := arms[0], arms[index]
if (unknown.unattributedActions() == 0) == (attributed.unattributedActions() == 0) {
continue
}
if unknown.unattributedActions() == 0 {
unknown, attributed = attributed, unknown
}
return fmt.Errorf("arm %q counts %d action(s) of unknown provenance and arm %q counts none: "+
"a denominator that may include the spec's setup cannot be compared against one that excludes it",
unknown.Name, unknown.unattributedActions(), attributed.Name)
}
return nil
}
func countCorrected(pairs []pairwiseResult) int {
corrected := 0
for _, pair := range pairs {
if !math.IsNaN(pair.PValue) {
corrected++
}
}
return corrected
}
func comparePairs(arms []arm) []pairwiseResult {
var pairs []pairwiseResult
for first := 0; first < len(arms); first++ {
for second := first + 1; second < len(arms); second++ {
test := gehanTest(arms[first].observations(), arms[second].observations())
pairs = append(pairs, pairwiseResult{
First: arms[first].Name,
Second: arms[second].Name,
FirstSize: test.FirstSize,
SecondSize: test.SecondSize,
Statistic: test.Statistic,
A12: test.A12,
Unordered: test.Unordered,
PValue: test.PValue,
HolmPValue: math.NaN(),
})
}
}
// Holm runs over this one family of comparisons. A comparison whose p-value
// could not be computed is not part of the family and does not shrink the
// correction the others receive.
var family []int
var raw []float64
for index, pair := range pairs {
if math.IsNaN(pair.PValue) {
continue
}
family = append(family, index)
raw = append(raw, pair.PValue)
}
for position, adjusted := range holm(raw) {
pairs[family[position]].HolmPValue = adjusted
}
return pairs
}
func summarize(current arm) armSummary {
summary := armSummary{
Arm: current.Name,
Generator: current.Generator,
Platform: current.Platform,
StepBudget: current.Budget,
Directories: current.Directories,
Recorded: len(current.Runs),
MissingSeeds: current.MissingSeeds,
UnattributedActions: current.unattributedActions(),
}
runsPerDefect := map[string]int{}
for _, item := range current.Runs {
if item.ExcludedBecause != "" {
summary.Excluded++
if summary.ExcludedByReason == nil {
summary.ExcludedByReason = map[string]int{}
}
summary.ExcludedByReason[item.ExcludedBecause]++
continue
}
summary.Usable++
summary.TotalSteps += item.Steps
// Steps and actions differ by the steps that chose no action, the steps
// whose action was never dispatched, and the steps the spec's setup
// drove into position. Only what the action generator dispatched
// explored the app, so only that belongs in a per-action rate.
summary.TotalActions += item.Actions
summary.TotalRunHours += float64(item.MonotonicMillis) / float64(time.Hour/time.Millisecond)
if item.ClampedToBudget {
summary.EventsHeldAtBudget++
}
if item.Violated && item.EventStep > item.OriginStep {
summary.EventsDetectedAfterOrigin++
}
if item.Violated {
summary.Violated++
} else {
summary.Censored++
}
distinct := slices.Compact(slices.Sorted(slices.Values(item.ViolatedProperties)))
summary.Detections += len(distinct)
for _, property := range distinct {
runsPerDefect[property]++
}
}
summary.SurvivalCurve = kaplanMeier(current.observations())
if median, ok := medianSurvival(summary.SurvivalCurve); ok {
summary.MedianStepsToFirstViolation = &median
}
if lower, ok := quantileSurvival(summary.SurvivalCurve, 0.25); ok {
summary.FirstQuartileSteps = &lower
}
if upper, ok := quantileSurvival(summary.SurvivalCurve, 0.75); ok {
summary.ThirdQuartileSteps = &upper
}
if summary.Usable > 0 {
rate := float64(summary.Violated) / float64(summary.Usable)
summary.ViolationRate = &rate
}
if summary.TotalActions > 0 {
perThousand := 1000 * float64(summary.Detections) / float64(summary.TotalActions)
summary.DefectsPerThousandActions = &perThousand
}
if summary.TotalRunHours > 0 {
perHour := float64(summary.Detections) / summary.TotalRunHours
summary.DefectsPerHour = &perHour
}
if len(runsPerDefect) > 0 {
summary.DefectRunCounts = runsPerDefect
summary.DistinctDefects = len(runsPerDefect)
for _, count := range runsPerDefect {
if count == 1 {
summary.SingletonDefects++
}
}
fraction := float64(summary.SingletonDefects) / float64(summary.DistinctDefects)
summary.SingletonFraction = &fraction
}
return summary
}
+422
View File
@@ -0,0 +1,422 @@
package main
import (
"io"
"math"
"path/filepath"
"strings"
"testing"
"time"
)
func violatingRun(seed int64, steps, origin int, properties ...string) classifiedRun {
return classifiedRun{
Seed: seed,
Steps: steps,
Actions: steps,
MonotonicMillis: 60_000,
OriginStep: origin,
EventStep: origin,
Violated: true,
ViolatedProperties: properties,
}
}
func cleanRun(seed int64, steps int) classifiedRun {
return classifiedRun{Seed: seed, Steps: steps, Actions: steps, MonotonicMillis: 60_000}
}
func TestSummarize_ArmWhereNoRunViolated(t *testing.T) {
summary := summarize(arm{
Name: "quiet",
Budget: 40,
Runs: []classifiedRun{cleanRun(1, 40), cleanRun(2, 40), cleanRun(3, 40)},
})
if summary.Usable != 3 || summary.Censored != 3 || summary.Violated != 0 {
t.Errorf("summary %+v, want three censored runs", summary)
}
if summary.MedianStepsToFirstViolation != nil {
t.Errorf("median %v, want undefined", *summary.MedianStepsToFirstViolation)
}
if summary.ViolationRate == nil || *summary.ViolationRate != 0 {
t.Errorf("violation rate %v, want 0", summary.ViolationRate)
}
if summary.DistinctDefects != 0 || summary.SingletonFraction != nil {
t.Errorf("defects %d singleton fraction %v, want none", summary.DistinctDefects, summary.SingletonFraction)
}
}
func TestSummarize_ArmWhereEveryRunViolated(t *testing.T) {
summary := summarize(arm{
Name: "loud",
Budget: 40,
Runs: []classifiedRun{
violatingRun(1, 5, 5, "cartTotal"),
violatingRun(2, 9, 9, "cartTotal"),
violatingRun(3, 11, 11, "cartTotal", "backNavigation"),
},
})
if summary.Violated != 3 || summary.Censored != 0 {
t.Errorf("summary %+v, want three events", summary)
}
if summary.MedianStepsToFirstViolation == nil || *summary.MedianStepsToFirstViolation != 9 {
t.Errorf("median %v, want 9", summary.MedianStepsToFirstViolation)
}
if *summary.ViolationRate != 1 {
t.Errorf("violation rate %v, want 1", *summary.ViolationRate)
}
if summary.Detections != 4 || summary.DistinctDefects != 2 {
t.Errorf("detections %d distinct %d, want 4 and 2", summary.Detections, summary.DistinctDefects)
}
// backNavigation appears in one run of three, cartTotal in all three.
if summary.SingletonDefects != 1 || math.Abs(*summary.SingletonFraction-0.5) > 1e-12 {
t.Errorf("singletons %d fraction %v, want 1 and 0.5", summary.SingletonDefects, summary.SingletonFraction)
}
// 25 actions over 3 minutes.
if math.Abs(*summary.DefectsPerThousandActions-160) > 1e-9 {
t.Errorf("defects per thousand actions %v, want 160", *summary.DefectsPerThousandActions)
}
if math.Abs(*summary.DefectsPerHour-80) > 1e-9 {
t.Errorf("defects per hour %v, want 80", *summary.DefectsPerHour)
}
}
// A step that chose no action, and a step whose action was never dispatched,
// left the app untouched. Counting them would inflate the denominator of every
// per-action rate, and the inflation differs by arm so it does not cancel.
func TestSummarize_CountsDispatchedActionsNotSteps(t *testing.T) {
summary := summarize(arm{
Name: "declines",
Budget: 40,
Runs: []classifiedRun{
{Seed: 1, Steps: 40, Actions: 10, MonotonicMillis: 3_600_000,
Violated: true, OriginStep: 12, ViolatedProperties: []string{"cartTotal"}},
{Seed: 2, Steps: 40, Actions: 6, MonotonicMillis: 3_600_000},
},
})
if summary.TotalSteps != 80 {
t.Errorf("total steps %d, want 80", summary.TotalSteps)
}
if summary.TotalActions != 16 {
t.Errorf("total actions %d, want 16 dispatched of 80 steps", summary.TotalActions)
}
if summary.DefectsPerThousandActions == nil {
t.Fatal("no defects per thousand actions")
}
expected := 1000.0 / 16.0
if math.Abs(*summary.DefectsPerThousandActions-expected) > 1e-9 {
t.Errorf("defects per thousand actions %v, want %v", *summary.DefectsPerThousandActions, expected)
}
}
func TestSummarize_ArmThatDispatchedNothingHasNoPerActionRate(t *testing.T) {
summary := summarize(arm{
Name: "inert",
Budget: 20,
Runs: []classifiedRun{
{Seed: 1, Steps: 20, Actions: 0, MonotonicMillis: 3_600_000,
Violated: true, OriginStep: 3, ViolatedProperties: []string{"cartTotal"}},
},
})
if summary.TotalActions != 0 || summary.TotalSteps != 20 {
t.Errorf("steps %d actions %d, want 20 and 0", summary.TotalSteps, summary.TotalActions)
}
if summary.DefectsPerThousandActions != nil {
t.Errorf("defects per thousand actions %v, want none with nothing dispatched", *summary.DefectsPerThousandActions)
}
if summary.DefectsPerHour == nil || *summary.DefectsPerHour != 1 {
t.Errorf("defects per hour %v, want 1: the run still consumed an hour", summary.DefectsPerHour)
}
}
func runHoursFor(t *testing.T, record map[string]any) float64 {
t.Helper()
directory := filepath.Join(t.TempDir(), "campaign")
record["seed"] = 1
writeCampaign(t, directory,
map[string]any{"arm": "seeded", "max_steps": 50, "seeds": []int{1}},
[]map[string]any{record})
arms, err := groupArms([]string{directory})
if err != nil {
t.Fatal(err)
}
return summarize(arms[0]).TotalRunHours
}
// A host asleep mid-run advanced the wall clock while testing nothing, so the
// sleep has no place in the denominator of a per-hour rate.
func TestSummarize_RunHoursCountTheTimeWorkedNotTheTimeThatPassed(t *testing.T) {
hours := runHoursFor(t, map[string]any{
"exit_code": 0, "steps": 50, "actions": 50,
"monotonic_millis": 3_600_000, "wall_clock_millis": 5_400_000,
})
if hours != 1 {
t.Errorf("run hours %v, want the 1 hour worked rather than the 1.5 hours that passed", hours)
}
}
// Campaigns recorded before the two clocks were split gave the same monotonic
// reading the name duration_millis, and their hours still have to count.
func TestSummarize_RunHoursReadCampaignsWrittenBeforeTheClocksWereSplit(t *testing.T) {
hours := runHoursFor(t, map[string]any{
"exit_code": 0, "steps": 50, "actions": 50, "duration_millis": 3_600_000,
})
if hours != 1 {
t.Errorf("run hours %v, want 1 from the older duration_millis field", hours)
}
}
func TestSummarize_ArmWithNoUsableRunsAfterExclusions(t *testing.T) {
summary := summarize(arm{
Name: "broken",
Budget: 40,
Runs: []classifiedRun{
{Seed: 1, ExcludedBecause: reasonTimedOut},
{Seed: 2, ExcludedBecause: reasonNonzeroExit},
{Seed: 3, ExcludedBecause: reasonNonzeroExit},
},
})
if summary.Usable != 0 || summary.Excluded != 3 {
t.Errorf("summary %+v, want no usable runs and three exclusions", summary)
}
if summary.ExcludedByReason[reasonNonzeroExit] != 2 || summary.ExcludedByReason[reasonTimedOut] != 1 {
t.Errorf("exclusions %v", summary.ExcludedByReason)
}
if summary.ViolationRate != nil || summary.MedianStepsToFirstViolation != nil {
t.Error("reported a rate or a median for an arm with nothing in it")
}
if summary.DefectsPerThousandActions != nil || summary.DefectsPerHour != nil {
t.Error("reported a yield rate with no actions and no time")
}
if len(summary.SurvivalCurve) != 0 {
t.Errorf("survival curve %v, want empty", summary.SurvivalCurve)
}
}
func TestSummarize_SingleRunArm(t *testing.T) {
summary := summarize(arm{Name: "one", Budget: 40, Runs: []classifiedRun{violatingRun(1, 6, 6, "cartTotal")}})
if summary.Usable != 1 || summary.Violated != 1 {
t.Errorf("summary %+v", summary)
}
if summary.MedianStepsToFirstViolation == nil || *summary.MedianStepsToFirstViolation != 6 {
t.Errorf("median %v, want 6", summary.MedianStepsToFirstViolation)
}
if summary.SingletonDefects != 1 || *summary.SingletonFraction != 1 {
t.Errorf("singletons %d fraction %v, want 1 and 1", summary.SingletonDefects, summary.SingletonFraction)
}
}
// Excluded runs must not reach the survival data at all, and the counts must
// keep them visible.
func TestAnalyse_ExcludedRunsNeverBecomeObservations(t *testing.T) {
current := arm{
Name: "mixed",
Budget: 30,
Runs: []classifiedRun{
violatingRun(1, 8, 8, "cartTotal"),
cleanRun(2, 30),
{Seed: 3, ExcludedBecause: reasonTimedOut},
},
}
observations := current.observations()
if len(observations) != 2 {
t.Fatalf("%d observations, want 2", len(observations))
}
summary := summarize(current)
if summary.Usable != 2 || summary.Excluded != 1 || summary.Violated != 1 || summary.Censored != 1 {
t.Errorf("summary %+v", summary)
}
}
func TestAnalyse_ArmWithNoUsableRunsIsReportedButNotTested(t *testing.T) {
result, err := analyse([]arm{
{Name: "a", Budget: 30, Runs: []classifiedRun{violatingRun(1, 4, 4), violatingRun(2, 6, 6)}},
{Name: "b", Budget: 30, Runs: []classifiedRun{cleanRun(1, 30), cleanRun(2, 30)}},
{Name: "c", Budget: 30, Runs: []classifiedRun{{Seed: 1, ExcludedBecause: reasonNonzeroExit}}},
}, time.Unix(0, 0).UTC())
if err != nil {
t.Fatal(err)
}
if len(result.Arms) != 3 {
t.Fatalf("%d arms reported, want all 3", len(result.Arms))
}
if result.LogRank == nil || len(result.LogRank.Groups) != 2 {
t.Fatalf("log-rank %+v, want the two testable arms", result.LogRank)
}
if len(result.Pairwise) != 1 {
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
}
if len(result.Notes) == 0 {
t.Error("no note explaining the dropped arm")
}
}
// With a single testable arm there is nothing to compare against, and the tool
// must say so instead of producing a statistic.
func TestAnalyse_SingleArmHasNoTests(t *testing.T) {
result, err := analyse([]arm{
{Name: "a", Budget: 30, Runs: []classifiedRun{violatingRun(1, 4, 4)}},
}, time.Unix(0, 0).UTC())
if err != nil {
t.Fatal(err)
}
if result.LogRank != nil || len(result.Pairwise) != 0 {
t.Errorf("log-rank %+v pairwise %v, want neither", result.LogRank, result.Pairwise)
}
}
// Holm is applied within the family of pairwise comparisons, so with three arms
// the smallest raw p-value is multiplied by three.
func TestComparePairs_AppliesHolmWithinTheFamily(t *testing.T) {
arms := []arm{
{Name: "a", Budget: 40, Runs: manyRuns(12, 4)},
{Name: "b", Budget: 40, Runs: manyRuns(12, 20)},
{Name: "c", Budget: 40, Runs: manyRuns(12, 36)},
}
pairs := comparePairs(arms)
if len(pairs) != 3 {
t.Fatalf("%d comparisons, want 3", len(pairs))
}
raw := make([]float64, len(pairs))
for index, pair := range pairs {
raw[index] = pair.PValue
if pair.HolmPValue < pair.PValue-1e-12 {
t.Errorf("%s vs %s: holm p %v below raw p %v", pair.First, pair.Second, pair.HolmPValue, pair.PValue)
}
}
expected := holm(raw)
for index, pair := range pairs {
if math.Abs(pair.HolmPValue-expected[index]) > 1e-12 {
t.Errorf("comparison %d holm p %v, want %v", index, pair.HolmPValue, expected[index])
}
}
}
// a12 above one half means the first arm needed more steps before its first
// violation, so the arm that finds defects sooner sits below one half.
func TestComparePairs_A12DirectionFollowsStepCounts(t *testing.T) {
pairs := comparePairs([]arm{
{Name: "slow", Budget: 40, Runs: manyRuns(6, 30)},
{Name: "fast", Budget: 40, Runs: manyRuns(6, 4)},
})
if pairs[0].A12 <= 0.5 {
t.Errorf("a12 %v for the slower arm listed first, want above 0.5", pairs[0].A12)
}
}
func writeCleanCampaign(t *testing.T, directory, name string, budget, steps, runs int) {
t.Helper()
seeds := make([]int, 0, runs)
records := make([]map[string]any, 0, runs)
for seed := 1; seed <= runs; seed++ {
seeds = append(seeds, seed)
records = append(records, map[string]any{
"seed": seed, "exit_code": 0, "steps": steps, "actions": steps, "monotonic_millis": 60_000,
})
}
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": budget, "seeds": seeds}, records)
}
// Arms censored at different budgets are not on the same clock: every clean run
// of the wider arm outranks every clean run of the narrower one whatever the
// app did, so the pairwise and the paired test reach a foregone conclusion the
// log-rank in the same report contradicts. groupArms already refuses this
// within one arm, and comparing across arms is the same hazard.
func TestRun_RefusesToCompareArmsCensoredAtDifferentBudgets(t *testing.T) {
cases := []struct {
name string
wideSteps int
arguments []string
}{
{name: "identical runs under different budgets", wideSteps: 100},
{name: "each arm run to its own budget", wideSteps: 400},
{name: "paired", wideSteps: 400, arguments: []string{"--paired"}},
}
for _, test := range cases {
root := t.TempDir()
wide := filepath.Join(root, "wide")
narrow := filepath.Join(root, "narrow")
writeCleanCampaign(t, wide, "wide", 400, test.wideSteps, 30)
writeCleanCampaign(t, narrow, "narrow", 100, 100, 30)
err := run(append(test.arguments, wide, narrow), io.Discard, io.Discard)
if err == nil {
t.Fatalf("%s: arms censored at 400 and at 100 steps were compared without complaint", test.name)
}
for _, fragment := range []string{"wide", "400", "narrow", "100", "different budgets"} {
if !strings.Contains(err.Error(), fragment) {
t.Errorf("%s: error %q is missing %q", test.name, err, fragment)
}
}
}
}
func writeSourcedCampaign(t *testing.T, directory, name string, steps, unattributed int) {
t.Helper()
const budget, runs, actions = 40, 6, 20
seeds := make([]int, 0, runs)
records := make([]map[string]any, 0, runs)
for seed := 1; seed <= runs; seed++ {
seeds = append(seeds, seed)
records = append(records, map[string]any{
"seed": seed, "exit_code": 0, "steps": steps, "actions": actions,
"monotonic_millis": 60_000, "unattributed_actions": unattributed,
})
}
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": budget, "seeds": seeds}, records)
}
// An arm whose actions name no producer counts whatever the spec's setup
// dispatched in the denominator of every per-action rate, and an arm whose
// actions name one leaves the login out of it. The two denominators measure
// different things, so a test that ranks one arm against the other reads a
// difference in what was counted as a difference in what the arms found.
func TestRun_RefusesToCompareArmsWhoseActionsWereCountedDifferently(t *testing.T) {
cases := []struct {
name string
arguments []string
}{
{name: "independent samples"},
{name: "paired", arguments: []string{"--paired"}},
}
for _, test := range cases {
root := t.TempDir()
attributed := filepath.Join(root, "attributed")
legacy := filepath.Join(root, "legacy")
writeSourcedCampaign(t, attributed, "attributed", 40, 0)
writeSourcedCampaign(t, legacy, "legacy", 30, 20)
err := run(append(test.arguments, attributed, legacy), io.Discard, io.Discard)
if err == nil {
t.Fatalf("%s: an arm of unknown provenance was tested against an attributed one without complaint", test.name)
}
for _, fragment := range []string{"attributed", "legacy", "120", "unknown provenance"} {
if !strings.Contains(err.Error(), fragment) {
t.Errorf("%s: error %q is missing %q", test.name, err, fragment)
}
}
}
}
// Two arms recorded before actions named their producer are on the same
// denominator as each other, diluted the same way, so they compare.
func TestRun_ComparesTwoArmsThatBothNameNoProducer(t *testing.T) {
root := t.TempDir()
first := filepath.Join(root, "first")
second := filepath.Join(root, "second")
writeSourcedCampaign(t, first, "first", 40, 20)
writeSourcedCampaign(t, second, "second", 30, 20)
if err := run([]string{first, second}, io.Discard, io.Discard); err != nil {
t.Fatalf("two arms of the same unknown provenance were refused: %v", err)
}
}
func manyRuns(count, originStep int) []classifiedRun {
runs := make([]classifiedRun, 0, count)
for index := 0; index < count; index++ {
runs = append(runs, violatingRun(int64(index), originStep+index, originStep+index, "cartTotal"))
}
return runs
}
@@ -0,0 +1,148 @@
package main
import (
"bytes"
"io"
"math"
"path/filepath"
"strings"
"testing"
)
// A run stops at whichever comes first, the step budget or the campaign's wall
// clock, so an arm that spends more wall clock per step leaves runs censored far
// below the budget. Those runs are not observations of a violation at that step,
// and the comparison between arms has to read them as the bounds they are.
type wallClockRun struct {
steps int
violated bool
}
func stoppedShort(count, steps int) []wallClockRun {
runs := make([]wallClockRun, 0, count)
for index := 0; index < count; index++ {
runs = append(runs, wallClockRun{steps: steps})
}
return runs
}
func violatedAt(count, steps int) []wallClockRun {
runs := make([]wallClockRun, 0, count)
for index := 0; index < count; index++ {
runs = append(runs, wallClockRun{steps: steps, violated: true})
}
return runs
}
func writeWallClockCampaign(t *testing.T, directory, name string, budget int, runs []wallClockRun) string {
t.Helper()
seeds := make([]int, 0, len(runs))
records := make([]map[string]any, 0, len(runs))
for index, run := range runs {
seed := index + 1
seeds = append(seeds, seed)
record := map[string]any{
"seed": seed, "exit_code": 0, "steps": run.steps, "actions": run.steps,
"monotonic_millis": 60_000,
}
if run.violated {
record["first_violation_origin_step"] = run.steps
record["violated_properties"] = []string{"plantedProperty"}
}
records = append(records, record)
}
writeCampaign(t, directory, map[string]any{
"arm": name, "generator": "seeded", "platform": "android",
"max_steps": budget, "seeds": seeds,
}, records)
return directory
}
func wallClockArms(t *testing.T, early, late []wallClockRun) (string, string) {
t.Helper()
root := t.TempDir()
return writeWallClockCampaign(t, filepath.Join(root, "early"), "early", 400, early),
writeWallClockCampaign(t, filepath.Join(root, "late"), "late", 400, late)
}
// Twenty runs stopped clean at step 12 against twenty violations at step 100.
// Nothing in the first arm was observed past step 12, so no pair of runs across
// the arms has a determined order and there is no difference to report.
func TestRun_RunsStoppedBeforeEveryEventCarryNoComparison(t *testing.T) {
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), violatedAt(20, 100))
pair := analyseCampaigns(t, earlyDirectory, lateDirectory).Pairwise[0]
if pair.First != "early" || pair.Second != "late" {
t.Fatalf("comparison %s vs %s, want early vs late", pair.First, pair.Second)
}
if math.Abs(pair.A12-0.5) > 1e-12 {
t.Errorf("a12 %.4f between an arm censored at 12 and one violating at 100, want 0.5: "+
"a run that stopped at step 12 never reached step 100", pair.A12)
}
if pair.PValue < 0.05 {
t.Errorf("p %.3e, want no significant difference: the arms were never observed over the same steps",
pair.PValue)
}
}
// Where the two arms were observed together, over the first twelve steps, the
// arm the wall clock stopped is the one that did not violate. The effect size
// has to follow that and not the step counts the flattening reads.
func TestRun_EffectSizeFollowsWhatCensoringDetermines(t *testing.T) {
late := append(violatedAt(6, 5), violatedAt(14, 100)...)
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), late)
pair := analyseCampaigns(t, earlyDirectory, lateDirectory).Pairwise[0]
if pair.A12 <= 0.5 {
t.Errorf("a12 %.4f, want above 0.5: six of the late arm's runs violated by step 5 "+
"and none of the early arm's twenty had violated by step 12", pair.A12)
}
}
// The seed-matched contrast reads the same censored runs and reaches the same
// conclusion or it is not measuring the same thing.
func TestRun_PairedContrastFollowsWhatCensoringDetermines(t *testing.T) {
late := append(violatedAt(6, 5), violatedAt(14, 100)...)
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), late)
result := analyseCampaigns(t, "--paired", earlyDirectory, lateDirectory)
paired := *result.Paired
if paired.First != "early" || paired.Second != "late" {
t.Fatalf("paired %s minus %s, want early minus late", paired.First, paired.Second)
}
if paired.Sign != 1 {
t.Errorf("sign %+d, want +1: the late arm is the one seen to violate first, in the six pairs "+
"where the order is determined at all", paired.Sign)
}
if paired.A12 <= 0.5 {
t.Errorf("a12 within pairs %.4f, want above 0.5", paired.A12)
}
}
// Nothing orders any pair here, so there is no test to report. The summary has
// to say that rather than failing to write a number that does not exist: a
// NaN p-value is not JSON and the whole summary went unwritten behind it.
func TestRun_PairedContrastWithNoOrderedPairSaysSo(t *testing.T) {
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), violatedAt(20, 100))
result := analyseCampaigns(t, "--paired", earlyDirectory, lateDirectory)
paired := *result.Paired
if paired.Pairs != 20 || paired.Unordered != 20 {
t.Fatalf("paired %+v, want twenty pairs and all of them unordered", paired)
}
if paired.PValue != nil || paired.HolmPValue != nil {
t.Errorf("p %v and holm p %v, want both undefined", paired.PValue, paired.HolmPValue)
}
if result.HolmFamilySize != 0 {
t.Errorf("holm family of %d, want none where nothing was tested", result.HolmFamilySize)
}
var stdout bytes.Buffer
if err := run([]string{"--paired", earlyDirectory, lateDirectory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
if !strings.Contains(stdout.String(), "sign test over the 0 ordered pair(s), p n/a, holm p n/a") {
t.Errorf("report does not say the test was not run:\n%s", stdout.String())
}
}
@@ -0,0 +1,109 @@
package main
import (
"bytes"
"io"
"path/filepath"
"strings"
"testing"
)
// These records are the shape a real folio-web campaign produced: an
// `eventually` obligation armed on step 1, never satisfied, and reported when
// the run ended at the step budget. Timing that event by the step that armed it
// puts the first violation on step 1 and reports a median of one step to first
// violation for an arm that spent its whole budget before it could know.
func liveness(seed int64, budget int, properties ...string) map[string]any {
return map[string]any{
"seed": seed, "exit_code": 0, "steps": budget, "actions": budget - 1,
"monotonic_millis": 15000,
"first_violation_origin_step": 1,
"first_violation_detected_step": budget,
"first_violation_reason": "eventually never satisfied",
"violated_properties": properties,
}
}
func TestClassify_ObligationReportedAtTheRunEndIsTimedAtItsDetection(t *testing.T) {
detected := 40
item := classify(runRecord{
Seed: 2, Steps: 40, Actions: stepPointer(39),
FirstViolationOriginStep: stepPointer(1),
FirstViolationDetectedStep: &detected,
ViolatedProperties: []string{"someTransactionExists"},
}, 40)
if !item.Violated {
t.Fatal("run not marked as violated")
}
if item.OriginStep != 1 {
t.Errorf("origin step %d, want the step that armed the obligation", item.OriginStep)
}
if item.EventStep != 40 {
t.Errorf("event step %d, want the step the run could know, 40", item.EventStep)
}
current := arm{Budget: 40, Runs: []classifiedRun{item}}
observations := current.observations()
if len(observations) != 1 || !observations[0].Event || observations[0].Steps != 40 {
t.Errorf("observations %+v, want one event at 40", observations)
}
}
// A safety property that trips under its own action is detected on the step
// that armed it, so nothing about the existing outcome moves.
func TestClassify_SafetyViolationKeepsItsOriginStep(t *testing.T) {
detected := 12
item := classify(runRecord{
Steps: 12, FirstViolationOriginStep: stepPointer(12), FirstViolationDetectedStep: &detected,
}, 400)
if item.EventStep != 12 || item.OriginStep != 12 {
t.Errorf("run %+v, want an event at 12", item)
}
}
// A campaign written before the detected step was recorded still reads, and
// keeps timing its events at the origin.
func TestClassify_MissingDetectedStepKeepsTheOrigin(t *testing.T) {
item := classify(runRecord{Steps: 30, FirstViolationOriginStep: stepPointer(18)}, 400)
if item.EventStep != 18 {
t.Errorf("event step %d, want the origin step 18", item.EventStep)
}
}
// The whole-pipeline form of the same thing, on the record shape a real web
// campaign wrote. Before the outcome was timed at detection this reported a
// median of 1 step to first violation.
func TestRun_LivenessFlushedAtTheBudgetDoesNotReportAOneStepMedian(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "seeded-web")
writeCampaign(t, directory, map[string]any{
"arm": "seeded-web", "generator": "seeded", "platform": "web",
"max_steps": 40, "seeds": []int{1, 2, 3, 4, 5},
}, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 38, "monotonic_millis": 13512},
liveness(2, 40, "accountCreationReachable", "someTransactionExists"),
liveness(3, 40, "someTransactionExists"),
liveness(4, 40, "accountCreationReachable", "someTransactionExists"),
{"seed": 5, "exit_code": 0, "steps": 40, "actions": 40, "monotonic_millis": 15433},
})
summary := armByName(t, analyseCampaigns(t, directory), "seeded-web")
if summary.MedianStepsToFirstViolation == nil {
t.Fatal("median undefined, want it at the budget")
}
if *summary.MedianStepsToFirstViolation != 40 {
t.Errorf("median %v steps to first violation, want 40: no run could know before the budget",
*summary.MedianStepsToFirstViolation)
}
if summary.EventsDetectedAfterOrigin != 3 {
t.Errorf("%d events detected after their origin, want 3", summary.EventsDetectedAfterOrigin)
}
var stdout bytes.Buffer
if err := run([]string{directory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
if !strings.Contains(stdout.String(), "timed 3 violation(s) at the step they were detected") {
t.Errorf("the report does not say the events were timed at detection\n%s", stdout.String())
}
}
@@ -0,0 +1,83 @@
package main
import "math"
// standardNormalUpperTail is P(Z > z) for a standard normal Z.
func standardNormalUpperTail(z float64) float64 {
return 0.5 * math.Erfc(z/math.Sqrt2)
}
// chiSquareUpperTail is P(X > x) for a chi-square variate with the given
// degrees of freedom, which is the regularized upper incomplete gamma
// Q(degreesOfFreedom/2, x/2).
func chiSquareUpperTail(x float64, degreesOfFreedom int) float64 {
if degreesOfFreedom <= 0 || math.IsNaN(x) {
return math.NaN()
}
if x <= 0 {
return 1
}
return regularizedUpperGamma(float64(degreesOfFreedom)/2, x/2)
}
const (
gammaIterationLimit = 2000
gammaTolerance = 1e-15
gammaTiny = 1e-300
)
// regularizedUpperGamma is Q(shape, x). The series is used below the crossover
// and the continued fraction above it, as in Numerical Recipes in C, 2nd ed.,
// section 6.2 (gammp/gammq).
func regularizedUpperGamma(shape, x float64) float64 {
if x < shape+1 {
return 1 - lowerGammaSeries(shape, x)
}
return upperGammaContinuedFraction(shape, x)
}
func lowerGammaSeries(shape, x float64) float64 {
term := 1 / shape
sum := term
for iteration := 1; iteration < gammaIterationLimit; iteration++ {
term *= x / (shape + float64(iteration))
sum += term
if math.Abs(term) < math.Abs(sum)*gammaTolerance {
break
}
}
return sum * math.Exp(-x+shape*math.Log(x)-logGamma(shape))
}
// upperGammaContinuedFraction evaluates Q(shape, x) with the modified Lentz
// algorithm, Numerical Recipes in C, 2nd ed., section 5.2.
func upperGammaContinuedFraction(shape, x float64) float64 {
b := x + 1 - shape
c := 1 / gammaTiny
d := 1 / b
h := d
for iteration := 1; iteration < gammaIterationLimit; iteration++ {
numerator := -float64(iteration) * (float64(iteration) - shape)
b += 2
d = numerator*d + b
if math.Abs(d) < gammaTiny {
d = gammaTiny
}
c = b + numerator/c
if math.Abs(c) < gammaTiny {
c = gammaTiny
}
d = 1 / d
delta := d * c
h *= delta
if math.Abs(delta-1) < gammaTolerance {
break
}
}
return h * math.Exp(-x+shape*math.Log(x)-logGamma(shape))
}
func logGamma(x float64) float64 {
value, _ := math.Lgamma(x)
return value
}
@@ -0,0 +1,74 @@
package main
import (
"math"
"testing"
)
// Chi-square critical values are the standard published table entries: the
// upper-tail probability of each of these statistics is the stated alpha in any
// chi-square table, for example Pearson and Hartley, Biometrika Tables for
// Statisticians, Table 8.
func TestChiSquareUpperTail_MatchesPublishedCriticalValues(t *testing.T) {
cases := []struct {
statistic float64
degreesOfFreedom int
expected float64
}{
{3.841459, 1, 0.05},
{6.634897, 1, 0.01},
{10.827566, 1, 0.001},
{5.991465, 2, 0.05},
{9.210340, 2, 0.01},
{7.814728, 3, 0.05},
{11.344867, 3, 0.01},
{9.487729, 4, 0.05},
{18.307038, 10, 0.05},
}
for _, test := range cases {
got := chiSquareUpperTail(test.statistic, test.degreesOfFreedom)
if math.Abs(got-test.expected) > 1e-6 {
t.Errorf("chiSquareUpperTail(%v, %d) = %v, want %v", test.statistic, test.degreesOfFreedom, got, test.expected)
}
}
}
// For one degree of freedom the upper tail has the closed form erfc(sqrt(x/2)),
// which is an independent check on the incomplete gamma routine.
func TestChiSquareUpperTail_AgreesWithClosedFormAtOneDegreeOfFreedom(t *testing.T) {
for _, statistic := range []float64{0.1, 1, 3.4, 16.79, 40, 120} {
expected := math.Erfc(math.Sqrt(statistic / 2))
got := chiSquareUpperTail(statistic, 1)
if math.Abs(got-expected) > 1e-12*math.Max(1, expected) {
t.Errorf("chiSquareUpperTail(%v, 1) = %v, want %v", statistic, got, expected)
}
}
}
func TestChiSquareUpperTail_ZeroStatisticIsCertain(t *testing.T) {
if got := chiSquareUpperTail(0, 1); got != 1 {
t.Errorf("chiSquareUpperTail(0, 1) = %v, want 1", got)
}
}
// Standard normal quantiles from any published normal table.
func TestStandardNormalUpperTail_MatchesPublishedQuantiles(t *testing.T) {
cases := []struct {
z float64
expected float64
}{
{1.281552, 0.10},
{1.644854, 0.05},
{1.959964, 0.025},
{2.326348, 0.01},
{2.575829, 0.005},
{3.090232, 0.001},
{0, 0.5},
}
for _, test := range cases {
got := standardNormalUpperTail(test.z)
if math.Abs(got-test.expected) > 1e-6 {
t.Errorf("standardNormalUpperTail(%v) = %v, want %v", test.z, got, test.expected)
}
}
}
@@ -0,0 +1,314 @@
package main
import (
"bytes"
"encoding/json"
"io"
"math"
"os"
"path/filepath"
"strings"
"testing"
)
// buildFixtureCampaign writes a campaign directory shaped exactly like the one
// the campaign tool emits: campaign.json plus one runs.jsonl line per seed.
func buildFixtureCampaign(t *testing.T, directory, armName string, budget int, records []map[string]any) {
t.Helper()
seeds := make([]int, 0, len(records))
for _, record := range records {
seeds = append(seeds, record["seed"].(int))
}
writeCampaign(t, directory, map[string]any{
"arm": armName,
"generator": "seeded",
"platform": "web",
"spec_path": "/specs/folio.ts",
"bundle_id": "app.folio",
"max_steps": budget,
"seeds": seeds,
"host": "experiment-host",
"started_at": "2026-08-12T00:00:00Z",
"argument_temp": nil,
}, records)
}
func seededArmRecords() []map[string]any {
// Ten runs: two violate early, one violates late, six run the budget clean,
// one times out and is missing data rather than a censored observation.
return []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 2, "exit_code": 0, "steps": 18, "actions": 16, "duration_millis": 120000,
"first_violation_origin_step": 14, "violated_properties": []string{"cartTotalMatches"}},
{"seed": 3, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 4, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 5, "exit_code": 0, "steps": 44, "actions": 42, "duration_millis": 240000,
"first_violation_origin_step": 41, "violated_properties": []string{"cartTotalMatches", "backLeavesApp"}},
{"seed": 6, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 7, "exit_code": -1, "timed_out": true, "actions": 0, "duration_millis": 900000},
{"seed": 8, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 9, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
{"seed": 10, "exit_code": 0, "steps": 21, "actions": 19, "duration_millis": 130000,
"first_violation_origin_step": 19, "violated_properties": []string{"cartTotalMatches"}},
}
}
func llmArmRecords() []map[string]any {
// Eight runs: six violate, one clean, one failed to launch. This arm
// dispatches an action on about half its steps, which is the asymmetry the
// per-action denominator has to survive.
return []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 7, "actions": 3, "duration_millis": 400000,
"first_violation_origin_step": 5, "violated_properties": []string{"cartTotalMatches"}},
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 5, "duration_millis": 420000,
"first_violation_origin_step": 8, "violated_properties": []string{"backLeavesApp"}},
{"seed": 3, "exit_code": 0, "steps": 60, "actions": 30, "duration_millis": 1800000, "first_violation_origin_step": nil},
{"seed": 4, "exit_code": 0, "steps": 5, "actions": 2, "duration_millis": 380000,
"first_violation_origin_step": 3, "violated_properties": []string{"cartTotalMatches"}},
{"seed": 5, "exit_code": 0, "steps": 13, "actions": 7, "duration_millis": 500000,
"first_violation_origin_step": 11, "violated_properties": []string{"cartTotalMatches", "priceNeverNegative"}},
{"seed": 6, "exit_code": -1, "actions": 0, "launch_error": "fork/exec sanderling: no such file or directory"},
{"seed": 7, "exit_code": 0, "steps": 6, "actions": 3, "duration_millis": 390000,
"first_violation_origin_step": 6, "violated_properties": []string{"cartTotalMatches"}},
{"seed": 8, "exit_code": 0, "steps": 16, "actions": 8, "duration_millis": 520000,
"first_violation_origin_step": 15, "violated_properties": []string{"backLeavesApp"}},
}
}
func TestRun_EndToEndOverFixtureCampaignDirectories(t *testing.T) {
root := t.TempDir()
seededDirectory := filepath.Join(root, "seeded-web")
llmDirectory := filepath.Join(root, "llm-web")
buildFixtureCampaign(t, seededDirectory, "seeded", 60, seededArmRecords())
buildFixtureCampaign(t, llmDirectory, "llm", 60, llmArmRecords())
summaryPath := filepath.Join(root, "analysis.json")
var stdout bytes.Buffer
if err := run([]string{"--json", summaryPath, seededDirectory, llmDirectory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
text := stdout.String()
for _, fragment := range []string{
"steps to first violation, right-censored at the last step a clean run reached",
"log-rank across 2 arms",
"pairwise gehan generalized wilcoxon",
"llm vs seeded",
"excluded 1 run(s) as missing data",
} {
if !strings.Contains(text, fragment) {
t.Errorf("stdout is missing %q\n%s", fragment, text)
}
}
body, err := os.ReadFile(summaryPath)
if err != nil {
t.Fatal(err)
}
var result analysis
if err := json.Unmarshal(body, &result); err != nil {
t.Fatal(err)
}
if len(result.Arms) != 2 {
t.Fatalf("%d arms, want 2", len(result.Arms))
}
byName := map[string]armSummary{}
for _, summary := range result.Arms {
byName[summary.Arm] = summary
}
seeded := byName["seeded"]
if seeded.Usable != 9 || seeded.Violated != 3 || seeded.Censored != 6 || seeded.Excluded != 1 {
t.Errorf("seeded arm %+v, want 9 usable, 3 violated, 6 censored, 1 excluded", seeded)
}
if seeded.ExcludedByReason[reasonTimedOut] != 1 {
t.Errorf("seeded exclusions %v, want one timeout", seeded.ExcludedByReason)
}
if seeded.MedianStepsToFirstViolation != nil {
t.Errorf("seeded median %v, want undefined with 3 of 9 violating",
*seeded.MedianStepsToFirstViolation)
}
if math.Abs(*seeded.ViolationRate-3.0/9.0) > 1e-12 {
t.Errorf("seeded violation rate %v, want 1/3", *seeded.ViolationRate)
}
// cartTotalMatches in 3 runs, backLeavesApp in 1 of 2 distinct defects.
if seeded.DistinctDefects != 2 || seeded.SingletonDefects != 1 {
t.Errorf("seeded defects %d singletons %d, want 2 and 1", seeded.DistinctDefects, seeded.SingletonDefects)
}
if seeded.TotalSteps != 443 || seeded.TotalActions != 425 {
t.Errorf("seeded steps %d actions %d, want 443 and 425", seeded.TotalSteps, seeded.TotalActions)
}
llm := byName["llm"]
if llm.Usable != 7 || llm.Violated != 6 || llm.Censored != 1 || llm.Excluded != 1 {
t.Errorf("llm arm %+v, want 7 usable, 6 violated, 1 censored, 1 excluded", llm)
}
if llm.ExcludedByReason[reasonLaunchError] != 1 {
t.Errorf("llm exclusions %v, want one launch error", llm.ExcludedByReason)
}
if llm.MedianStepsToFirstViolation == nil || *llm.MedianStepsToFirstViolation != 8 {
t.Errorf("llm median %v, want 8", llm.MedianStepsToFirstViolation)
}
// The arm dispatches an action on half its steps, so counting steps would
// halve its yield per thousand actions and flatter it against seeded.
if llm.TotalSteps != 116 || llm.TotalActions != 58 {
t.Errorf("llm steps %d actions %d, want 116 and 58", llm.TotalSteps, llm.TotalActions)
}
if llm.Detections != 7 {
t.Fatalf("llm detections %d, want 7", llm.Detections)
}
expected := 7000.0 / 58.0
if math.Abs(*llm.DefectsPerThousandActions-expected) > 1e-9 {
t.Errorf("llm defects per thousand actions %v, want %v", *llm.DefectsPerThousandActions, expected)
}
if result.LogRank == nil {
t.Fatal("no log-rank result")
}
if result.LogRank.DegreesOfFreedom != 1 {
t.Errorf("log-rank df %d, want 1", result.LogRank.DegreesOfFreedom)
}
if result.LogRank.PValue > 0.05 {
t.Errorf("log-rank p %v, want the two clearly different arms to separate", result.LogRank.PValue)
}
if len(result.Pairwise) != 1 {
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
}
pair := result.Pairwise[0]
if pair.First != "llm" || pair.Second != "seeded" {
t.Errorf("comparison %s vs %s, want arms in sorted order", pair.First, pair.Second)
}
if pair.A12 >= 0.5 {
t.Errorf("a12 %v, want the arm that violates sooner below 0.5", pair.A12)
}
if pair.HolmPValue != pair.PValue {
t.Errorf("holm p %v differs from raw p %v in a family of one", pair.HolmPValue, pair.PValue)
}
// Six seeded runs and one llm run ran the budget clean, and two censored
// runs have no order between them whatever step either stopped on.
if pair.Unordered < 6 {
t.Errorf("%d unordered pair(s), want at least the six pairs of censored runs", pair.Unordered)
}
}
func TestRun_ReportsBothArmsWhenOneHasNothingUsable(t *testing.T) {
root := t.TempDir()
good := filepath.Join(root, "good")
broken := filepath.Join(root, "broken")
buildFixtureCampaign(t, good, "good", 30, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28},
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 1000,
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
})
buildFixtureCampaign(t, broken, "broken", 30, []map[string]any{
{"seed": 1, "exit_code": 3, "actions": 0},
{"seed": 2, "timed_out": true, "exit_code": -1, "actions": 0},
})
var stdout bytes.Buffer
if err := run([]string{good, broken}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
text := stdout.String()
if !strings.Contains(text, "broken") {
t.Errorf("the arm with nothing usable is not reported\n%s", text)
}
if strings.Contains(text, "log-rank across") {
t.Errorf("ran a log-rank with only one testable arm\n%s", text)
}
if !strings.Contains(text, "arms with no usable runs are reported but left out") {
t.Errorf("no note about the dropped arm\n%s", text)
}
}
func TestRun_JsonToStdout(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "only")
buildFixtureCampaign(t, directory, "only", 20, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 20, "actions": 17},
})
var stdout bytes.Buffer
if err := run([]string{"--json", "-", "--campaign", directory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
start := strings.Index(stdout.String(), "{")
if start < 0 {
t.Fatalf("no json in stdout\n%s", stdout.String())
}
var result analysis
if err := json.Unmarshal([]byte(stdout.String()[start:]), &result); err != nil {
t.Fatalf("json: %v", err)
}
if len(result.Arms) != 1 || result.Arms[0].Arm != "only" {
t.Errorf("arms %+v", result.Arms)
}
}
func TestRun_RejectsTheSameDirectoryTwice(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "one")
buildFixtureCampaign(t, directory, "one", 20, []map[string]any{{"seed": 1, "exit_code": 0, "steps": 20, "actions": 20}})
err := run([]string{directory, directory}, io.Discard, io.Discard)
if err == nil || !strings.Contains(err.Error(), "twice") {
t.Fatalf("error %v, want a refusal to double count", err)
}
}
func TestRun_RequiresACampaignDirectory(t *testing.T) {
if err := run(nil, io.Discard, io.Discard); err == nil {
t.Fatal("expected an error with no campaign directories")
}
}
// A campaign recorded before actions named their producer cannot say whether
// the login taps of the spec's setup are inside its per-action denominator, and
// the report has to say so where that denominator is read rather than leave the
// reader to date the file.
func TestRun_MarksAnArmWhoseActionsNameNoProducer(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "before-source")
buildFixtureCampaign(t, directory, "before-source", 30, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28, "duration_millis": 60000},
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 60000,
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
})
var stdout bytes.Buffer
if err := run([]string{"--json", "-", directory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
text := stdout.String()
if !strings.Contains(text, "37 (37 unattributed)") {
t.Errorf("the actions cell does not carry the unattributed count\n%s", text)
}
if !strings.Contains(text, "unknown provenance") {
t.Errorf("the report does not say the denominator's provenance is unknown\n%s", text)
}
var result analysis
if err := json.Unmarshal([]byte(text[strings.Index(text, "{"):]), &result); err != nil {
t.Fatalf("json: %v", err)
}
if len(result.Arms) != 1 || result.Arms[0].UnattributedActions != 37 {
t.Errorf("arms %+v, want 37 unattributed actions", result.Arms)
}
}
func TestRun_LeavesAnArmWhoseActionsAllNameAProducerUnmarked(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "sourced")
buildFixtureCampaign(t, directory, "sourced", 30, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28, "duration_millis": 60000, "unattributed_actions": 0},
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 60000, "unattributed_actions": 0,
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
})
var stdout bytes.Buffer
if err := run([]string{directory}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
if strings.Contains(stdout.String(), "unattributed") || strings.Contains(stdout.String(), "unknown provenance") {
t.Errorf("an arm whose every action names a producer was marked\n%s", stdout.String())
}
}
@@ -0,0 +1,55 @@
package main
// Published right-censored datasets whose log-rank and Kaplan-Meier results are
// reported in the survival-analysis literature and in R's survival package, so
// every expected number in these tests can be checked against a source rather
// than against this tool's own output.
// gehanSixMercaptopurine and gehanPlacebo are remission times in weeks from the
// 6-MP versus placebo trial in acute leukaemia, Freireich et al. (1963). This is
// the dataset R's survival literature calls gehan. A trailing plus in the
// published listing marks a censored time.
//
// 6-MP: 6, 6, 6, 6+, 7, 9+, 10, 10+, 11+, 13, 16, 17+, 19+, 20+, 22, 23, 25+, 32+, 32+, 34+, 35+
// placebo: 1, 1, 2, 2, 3, 4, 4, 5, 5, 8, 8, 8, 8, 11, 11, 12, 12, 15, 17, 22, 23
var (
gehanSixMercaptopurine = []observation{
{6, true}, {6, true}, {6, true}, {6, false},
{7, true}, {9, false}, {10, true}, {10, false},
{11, false}, {13, true}, {16, true}, {17, false},
{19, false}, {20, false}, {22, true}, {23, true},
{25, false}, {32, false}, {32, false}, {34, false}, {35, false},
}
gehanPlacebo = []observation{
{1, true}, {1, true}, {2, true}, {2, true}, {3, true},
{4, true}, {4, true}, {5, true}, {5, true}, {8, true},
{8, true}, {8, true}, {8, true}, {11, true}, {11, true},
{12, true}, {12, true}, {15, true}, {17, true}, {22, true}, {23, true},
}
)
// amlMaintained and amlNonmaintained are the acute myelogenous leukaemia
// survival times in weeks from Miller (1997), shipped as the aml dataset in R's
// survival package. Five subjects are censored, at 13, 16, 28, 45 and 161 weeks.
//
// maintained: 9, 13, 13+, 18, 23, 28+, 31, 34, 45+, 48, 161+
// nonmaintained: 5, 5, 8, 8, 12, 16+, 23, 27, 30, 33, 43, 45
var (
amlMaintained = []observation{
{9, true}, {13, true}, {13, false}, {18, true}, {23, true},
{28, false}, {31, true}, {34, true}, {45, false}, {48, true}, {161, false},
}
amlNonmaintained = []observation{
{5, true}, {5, true}, {8, true}, {8, true}, {12, true}, {16, false},
{23, true}, {27, true}, {30, true}, {33, true}, {43, true}, {45, true},
}
)
// chorioamnionTerm and chorioamnionEarly are permeability constants of the human
// chorioamnion at term and between 12 and 26 weeks gestational age, Hollander
// and Wolfe (1973), 69f. R's wilcox.test help page uses exactly these vectors as
// its two-sample example.
var (
chorioamnionTerm = []float64{0.80, 0.83, 1.89, 1.04, 1.45, 1.38, 1.91, 1.64, 0.73, 1.46}
chorioamnionEarly = []float64{1.15, 0.88, 0.90, 0.74, 1.21}
)
+97
View File
@@ -0,0 +1,97 @@
package main
import "math"
// gehanResult is one pairwise comparison of two arms of right-censored runs: an
// effect size, the run pairs that have no order between them, and the test of
// the same statistic against the null of equal hazards.
type gehanResult struct {
FirstSize int
SecondSize int
Statistic float64
A12 float64
Unordered int
PValue float64
}
// outlives orders two runs the only way right-censoring allows. A run censored
// at step t violated at no step up to t and stopped for a reason of its own, so
// it outlives a violation at or before t and nothing orders it against a
// violation after t or against another censored run. Comparing the two step
// counts as plain numbers instead reads a run the wall clock stopped at step 12
// as one that violated at step 12.
func outlives(left, right observation) int {
switch {
case left.Event && right.Event:
switch {
case left.Steps > right.Steps:
return 1
case left.Steps < right.Steps:
return -1
}
case left.Event:
if right.Steps >= left.Steps {
return -1
}
case right.Event:
if left.Steps >= right.Steps {
return 1
}
}
return 0
}
// atRiskWeight is Gehan's weight: an event counts for as many runs as were still
// at risk when it happened. It is what makes the weighted log-rank statistic the
// same quantity as the pairwise count below, so the effect size and the p-value
// are one statistic rather than two that can disagree.
func atRiskWeight(atRisk float64) float64 { return atRisk }
// gehanTest is the Gehan-Breslow generalized Wilcoxon test: the rank-sum
// carried over to right-censored samples by scoring every pair of runs by which
// one outlived the other and leaving the pairs censoring cannot order out of the
// count. Gehan (1965), "A Generalized Wilcoxon Test for Comparing Arbitrarily
// Singly-Censored Samples", Biometrika 52(1-2), 203-223; Breslow (1970).
//
// Statistic is that count, U, and A12 is it over the number of pairs: the share
// of run pairs in which the first arm survived longer, an unordered pair
// counting as half. With nothing censored the two are exactly the Mann-Whitney U
// and the Vargha-Delaney A12 the uncensored rank-sum reports. Where censoring
// leaves a pair unordered, the half it contributes is the null value, so an
// unordered pair can only pull the effect size toward 0.5 and can never
// manufacture a direction.
//
// The p-value is the same statistic standardized: the weighted log-rank with
// Gehan's weight has this U for its statistic, and its variance is the
// conditional hypergeometric one summed over event times, which is what keeps
// the test honest when the arms censor on different schedules. The permutation
// variance Gehan originally paired with the statistic does not.
func gehanTest(first, second []observation) gehanResult {
result := gehanResult{
FirstSize: len(first),
SecondSize: len(second),
Statistic: math.NaN(),
A12: math.NaN(),
PValue: math.NaN(),
}
if len(first) == 0 || len(second) == 0 {
return result
}
outlived := 0.0
for _, left := range first {
for _, right := range second {
switch outlives(left, right) {
case 1:
outlived++
case 0:
outlived += 0.5
result.Unordered++
}
}
}
result.Statistic = outlived
result.A12 = outlived / float64(len(first)*len(second))
test := weightedLogRank([]string{"first", "second"}, [][]observation{first, second}, atRiskWeight)
result.PValue = test.PValue
return result
}
+256
View File
@@ -0,0 +1,256 @@
package main
import (
"math"
"testing"
)
func tiedPairs(first, second []float64) int {
tied := 0
for _, left := range first {
for _, right := range second {
if left == right {
tied++
}
}
}
return tied
}
func events(values []float64) []observation {
items := make([]observation, 0, len(values))
for _, value := range values {
items = append(items, observation{Steps: value, Event: true})
}
return items
}
// Every ordering censoring supports and every ordering it does not.
func TestOutlives_OrdersOnlyWhatTheCensoringSupports(t *testing.T) {
cases := []struct {
name string
left, right observation
want int
}{
{"two violations", observation{30, true}, observation{10, true}, 1},
{"two violations the other way", observation{10, true}, observation{30, true}, -1},
{"two violations at the same step", observation{10, true}, observation{10, true}, 0},
{"censored after the violation", observation{30, false}, observation{10, true}, 1},
{"censored on the violation's own step", observation{10, false}, observation{10, true}, 1},
{"censored before the violation", observation{10, false}, observation{30, true}, 0},
{"violation before the censoring", observation{10, true}, observation{30, false}, -1},
{"violation after the censoring", observation{30, true}, observation{10, false}, 0},
{"both censored", observation{10, false}, observation{30, false}, 0},
}
for _, test := range cases {
if got := outlives(test.left, test.right); got != test.want {
t.Errorf("%s: %v against %v ordered %+d, want %+d", test.name, test.left, test.right, got, test.want)
}
}
}
// With nothing censored the test is the rank-sum, so its statistic and effect
// size have to be the ones the rank-sum reports on the same numbers, ties
// included.
func TestGehanTest_ReducesToTheRankSumWhenNothingIsCensored(t *testing.T) {
cases := [][2][]float64{
{chorioamnionTerm, chorioamnionEarly},
{{1, 2, 3, 4}, {3, 4, 5, 6}},
{{40, 40, 40}, {40, 40, 40, 40}},
{{5, 6, 7}, {1, 2}},
{{3}, {9}},
}
for _, test := range cases {
result := gehanTest(events(test[0]), events(test[1]))
reference := rankSum(test[0], test[1])
if result.Statistic != reference.Statistic {
t.Errorf("u %v over %v and %v, want the rank-sum's %v",
result.Statistic, test[0], test[1], reference.Statistic)
}
if result.A12 != reference.A12 {
t.Errorf("a12 %v over %v and %v, want the rank-sum's %v",
result.A12, test[0], test[1], reference.A12)
}
if want := tiedPairs(test[0], test[1]); result.Unordered != want {
t.Errorf("%d unordered pair(s) over %v and %v, want the %d tied ones and no others",
result.Unordered, test[0], test[1], want)
}
}
}
// The failure the flattening produced: an arm the wall clock stopped at step 12
// says nothing about step 100, so there is no difference to find and no
// direction to report.
func TestGehanTest_RunsStoppedBeforeEveryViolationOrderNothing(t *testing.T) {
stopped := []observation{{12, false}, {12, false}, {12, false}, {12, false}}
violated := []observation{{100, true}, {100, true}, {100, true}}
result := gehanTest(stopped, violated)
if result.Unordered != 12 || result.A12 != 0.5 {
t.Errorf("%d of 12 pairs unordered, a12 %v, want all of them and 0.5", result.Unordered, result.A12)
}
if result.PValue < 0.05 {
t.Errorf("p %v, want no difference between arms never observed over the same steps", result.PValue)
}
}
// A censored run outliving a violation is evidence, and it is the only kind the
// wall-clock case leaves: four runs still clean at step 12 against three
// violations by step 5.
func TestGehanTest_CensoringLeavesTheEvidenceItDoesSupport(t *testing.T) {
stopped := []observation{{12, false}, {12, false}, {12, false}, {12, false}}
violated := []observation{{5, true}, {5, true}, {5, true}}
result := gehanTest(stopped, violated)
if result.Unordered != 0 || result.Statistic != 12 || result.A12 != 1 {
t.Errorf("result %+v, want every pair ordered for the arm that had not violated", result)
}
if result.PValue > 0.05 {
t.Errorf("p %v, want the arms to separate", result.PValue)
}
}
func riskAndDeaths(group []observation, steps float64) (float64, float64) {
atRisk, deaths := 0.0, 0.0
for _, item := range group {
if item.Steps >= steps {
atRisk++
}
if item.Steps == steps && item.Event {
deaths++
}
}
return atRisk, deaths
}
// gehanReference is Gehan's statistic and its conditional variance written
// straight from the definitions,
//
// S = sum over event times of (Y2*d1 - Y1*d2)
// V = sum over event times of d(Y-d)/(Y-1) * Y1*Y2
//
// which is an independent calculation rather than a second call into the code
// under test.
func gehanReference(first, second []observation) (float64, float64) {
pooled := append(append([]observation{}, first...), second...)
statistic, variance := 0.0, 0.0
for _, steps := range distinctSteps(pooled) {
firstAtRisk, firstDeaths := riskAndDeaths(first, steps)
secondAtRisk, secondDeaths := riskAndDeaths(second, steps)
deaths := firstDeaths + secondDeaths
if deaths == 0 {
continue
}
atRisk := firstAtRisk + secondAtRisk
statistic += secondAtRisk*firstDeaths - firstAtRisk*secondDeaths
if atRisk > 1 {
variance += deaths * (atRisk - deaths) / (atRisk - 1) * firstAtRisk * secondAtRisk
}
}
return statistic, variance
}
// The effect size and the p-value have to be the same statistic seen twice, or
// the report can carry a direction its p-value does not support. Counting run
// pairs and accumulating over risk sets are two routes to Gehan's statistic, and
// they are tied by S = mn - 2U.
func TestGehanTest_PairCountAndRiskSetAgreeOnOneStatistic(t *testing.T) {
cases := []struct {
name string
first, second []observation
}{
{"6-mp against placebo", gehanSixMercaptopurine, gehanPlacebo},
{"maintained against nonmaintained", amlMaintained, amlNonmaintained},
{"stopped short against violating late", []observation{{12, false}, {12, false}, {14, false}},
[]observation{{5, true}, {100, true}, {100, true}}},
{"censoring tied with an event", []observation{{20, false}, {20, true}, {35, true}},
[]observation{{20, true}, {20, false}, {9, true}}},
}
for _, test := range cases {
result := gehanTest(test.first, test.second)
statistic, variance := gehanReference(test.first, test.second)
pairs := float64(len(test.first) * len(test.second))
if got := pairs - 2*result.Statistic; math.Abs(got-statistic) > 1e-9 {
t.Errorf("%s: pair count gives a statistic of %v, the risk sets give %v", test.name, got, statistic)
}
expected := chiSquareUpperTail(statistic*statistic/variance, 1)
if math.Abs(result.PValue-expected) > 1e-12 {
t.Errorf("%s: p %v, want %v from statistic %v over variance %v",
test.name, result.PValue, expected, statistic, variance)
}
}
}
// The 6-MP trial is the dataset the test is named for. Its log-rank result is
// checked elsewhere against the published one; here the generalized Wilcoxon
// has to reach the same conclusion, with the maintained arm outliving the
// placebo arm on both routes.
func TestGehanTest_SeparatesThePublishedLeukaemiaTrial(t *testing.T) {
result := gehanTest(gehanSixMercaptopurine, gehanPlacebo)
if result.A12 <= 0.5 {
t.Errorf("a12 %v, want the 6-mp arm to outlive the placebo arm", result.A12)
}
if result.PValue > 0.001 {
t.Errorf("p %v, want the arms to separate as the log-rank has them separate", result.PValue)
}
logRankResult := logRank([]string{"6-mp", "placebo"},
[][]observation{gehanSixMercaptopurine, gehanPlacebo})
if logRankResult.PValue > 0.001 {
t.Fatalf("log-rank p %v: the comparison being made is not the one this test assumes", logRankResult.PValue)
}
}
func censoredAt(count int, steps float64) []observation {
items := make([]observation, 0, count)
for index := 0; index < count; index++ {
items = append(items, observation{Steps: steps})
}
return items
}
// A specification whose violations are all obligations reported when the run
// ends puts every event on one step, and the comparison collapses to a single
// two-by-two table of violated against clean. The statistic there is the
// Mantel-Haenszel chi-square of that table,
//
// (N-1)(ad-bc)^2 / ((a+b)(c+d)(a+c)(b+d))
//
// and the (Y-d)/(Y-1) term in the variance is what carries the (N-1)/N that
// separates it from the Pearson chi-square. Nothing else in the pipeline
// exercises that term hard, because tied events are otherwise rare.
func TestGehanTest_EveryViolationOnOneStepIsTheMantelHaenszelTable(t *testing.T) {
cases := [][4]int{{30, 20, 10, 40}, {25, 25, 15, 35}, {20, 20, 12, 28}}
for _, test := range cases {
firstEvents, firstCensored, secondEvents, secondCensored := test[0], test[1], test[2], test[3]
first := append(events(repeated(firstEvents, 400)), censoredAt(firstCensored, 400)...)
second := append(events(repeated(secondEvents, 400)), censoredAt(secondCensored, 400)...)
total := float64(firstEvents + firstCensored + secondEvents + secondCensored)
crossProduct := float64(firstEvents*secondCensored - firstCensored*secondEvents)
chiSquare := (total - 1) * crossProduct * crossProduct /
(float64(firstEvents+firstCensored) * float64(secondEvents+secondCensored) *
float64(firstEvents+secondEvents) * float64(firstCensored+secondCensored))
got := gehanTest(first, second).PValue
want := chiSquareUpperTail(chiSquare, 1)
if math.Abs(got-want) > 1e-12 {
t.Errorf("%v: p %v, want the table's %v from chi-square %v", test, got, want, chiSquare)
}
}
}
func repeated(count int, value float64) []float64 {
values := make([]float64, 0, count)
for index := 0; index < count; index++ {
values = append(values, value)
}
return values
}
func TestGehanTest_EmptyArmHasNoComparison(t *testing.T) {
result := gehanTest(nil, []observation{{5, true}})
if !math.IsNaN(result.PValue) || !math.IsNaN(result.A12) || !math.IsNaN(result.Statistic) {
t.Errorf("result %+v, want everything undefined", result)
}
}
+36
View File
@@ -0,0 +1,36 @@
package main
import (
"math"
"slices"
)
// holm applies the Holm (1979) step-down correction within one family of
// comparisons, enforcing monotonicity across the sorted p-values the way R's
// p.adjust does. Holm, "A Simple Sequentially Rejective Multiple Test
// Procedure", Scandinavian Journal of Statistics 6(2), 65-70.
func holm(pValues []float64) []float64 {
count := len(pValues)
adjusted := make([]float64, count)
order := make([]int, count)
for index := range order {
order[index] = index
}
slices.SortStableFunc(order, func(left, right int) int {
switch {
case pValues[left] < pValues[right]:
return -1
case pValues[left] > pValues[right]:
return 1
default:
return 0
}
})
running := 0.0
for position, index := range order {
scaled := float64(count-position) * pValues[index]
running = math.Max(running, scaled)
adjusted[index] = math.Min(running, 1)
}
return adjusted
}
+54
View File
@@ -0,0 +1,54 @@
package main
import (
"math"
"testing"
)
// Both cases are printed R output in Eve Slavich, "Four strategies for dealing
// with multiple comparisons", UNSW Stats Central, slides 9 and 10:
//
// pValues = c(0.01, 0.2, 0.08, 0.03)
// p.adjust(pValues, method = "holm")
// ## [1] 0.04 0.20 0.16 0.09
//
// pValues = c(0.01, 0.2, 0.08, 0.03, 0.02, 0.01)
// p.adjust(pValues, method = "holm")
// ## [1] 0.06 0.20 0.16 0.09 0.08 0.06
//
// The second case exercises the monotonicity step: sorted p are
// .01 .01 .02 .03 .08 .20, scaled by 6 5 4 3 2 1 to .06 .05 .08 .09 .16 .20,
// and the running maximum lifts the second back to .06.
func TestHolm_MatchesPublishedAdjustment(t *testing.T) {
cases := []struct {
raw []float64
expected []float64
}{
{[]float64{0.01, 0.2, 0.08, 0.03}, []float64{0.04, 0.20, 0.16, 0.09}},
{[]float64{0.01, 0.2, 0.08, 0.03, 0.02, 0.01}, []float64{0.06, 0.20, 0.16, 0.09, 0.08, 0.06}},
}
for _, test := range cases {
adjusted := holm(test.raw)
for index, want := range test.expected {
if math.Abs(adjusted[index]-want) > 1e-12 {
t.Errorf("holm(%v)[%d] = %v, want %v", test.raw, index, adjusted[index], want)
}
}
}
}
func TestHolm_CapsAtOneAndKeepsOrder(t *testing.T) {
adjusted := holm([]float64{0.4, 0.5, 0.9})
for index, value := range adjusted {
if value != 1 {
t.Errorf("adjusted[%d] = %v, want 1", index, value)
}
}
single := holm([]float64{0.03})
if len(single) != 1 || single[0] != 0.03 {
t.Errorf("single comparison adjusted to %v, want 0.03 unchanged", single)
}
if got := holm(nil); len(got) != 0 {
t.Errorf("holm(nil) = %v, want empty", got)
}
}
@@ -0,0 +1,155 @@
package main
import (
"bytes"
"io"
"path/filepath"
"strings"
"testing"
)
// A campaign that produced nothing still has a manifest, and the manifest is
// what makes the difference between an arm that ran nothing and an arm that was
// never scheduled. Reading one must report every intended seed as missing rather
// than an arm with a small clean sample.
func TestRun_EmptyCampaignIsReportedAsEverySeedMissing(t *testing.T) {
root := t.TempDir()
empty := filepath.Join(root, "empty")
writeCampaign(t, empty, map[string]any{
"arm": "seeded", "max_steps": 400, "seeds": []int{1, 2, 3, 4, 5},
}, nil)
var stdout bytes.Buffer
if err := run([]string{empty}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
result := analyseCampaigns(t, empty)
summary := armByName(t, result, "seeded")
if summary.Recorded != 0 || summary.Usable != 0 {
t.Errorf("recorded %d usable %d, want none", summary.Recorded, summary.Usable)
}
if len(summary.MissingSeeds) != 5 {
t.Errorf("missing seeds %v, want all five", summary.MissingSeeds)
}
if summary.MedianStepsToFirstViolation != nil || summary.ViolationRate != nil {
t.Errorf("median %v violation rate %v, want neither from no runs",
summary.MedianStepsToFirstViolation, summary.ViolationRate)
}
if result.LogRank != nil || len(result.Pairwise) != 0 {
t.Errorf("log-rank %+v pairwise %v, want no tests", result.LogRank, result.Pairwise)
}
if !strings.Contains(stdout.String(), "seeded") {
t.Errorf("the empty arm is not in the report\n%s", stdout.String())
}
}
// A host that stopped part way through leaves a directory that looks complete.
// The seeds it never reached are the difference between a partial campaign and
// a smaller one, and the report has to carry that count.
func TestRun_PartialCampaignCountsTheSeedsTheHostNeverReached(t *testing.T) {
root := t.TempDir()
partial := filepath.Join(root, "partial")
var records []map[string]any
for seed := 1; seed <= 6; seed++ {
records = append(records, map[string]any{
"seed": seed, "exit_code": 0, "steps": 400, "actions": 380, "monotonic_millis": 360000,
})
}
writeCampaign(t, partial, map[string]any{
"arm": "seeded", "max_steps": 400,
"seeds": []int{1, 2, 3, 4, 5, 6, 7, 8, 9, 10},
}, records)
summary := armByName(t, analyseCampaigns(t, partial), "seeded")
if summary.Usable != 6 {
t.Errorf("%d usable runs, want the 6 that landed", summary.Usable)
}
if len(summary.MissingSeeds) != 4 {
t.Errorf("missing seeds %v, want the 4 the host never reached", summary.MissingSeeds)
}
var stdout bytes.Buffer
if err := run([]string{partial}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
if !strings.Contains(stdout.String(), "missing") {
t.Errorf("no missing-seed column in the report\n%s", stdout.String())
}
}
// The ablation is seed-matched, so a run lost on one arm removes its partner
// from the comparison too. Those seeds are named because a paired sample that
// silently shrinks is how a campaign reports a difference between two arms that
// were not in fact matched.
func TestRun_PairedComparisonNamesSeedsLostOnOneArm(t *testing.T) {
root := t.TempDir()
declared := []int{1, 2, 3, 4, 5, 6}
pre := filepath.Join(root, "pre")
post := filepath.Join(root, "post")
writeCampaign(t, pre, map[string]any{"arm": "pre", "max_steps": 400, "seeds": declared}, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 300, "actions": 290, "monotonic_millis": 1000, "first_violation_origin_step": 300, "violated_properties": []string{"p"}},
{"seed": 2, "exit_code": 0, "steps": 400, "actions": 390, "monotonic_millis": 1000},
{"seed": 3, "exit_code": 0, "steps": 250, "actions": 240, "monotonic_millis": 1000, "first_violation_origin_step": 250, "violated_properties": []string{"p"}},
{"seed": 4, "exit_code": -1, "timed_out": true, "actions": 0},
})
writeCampaign(t, post, map[string]any{"arm": "post", "max_steps": 400, "seeds": declared}, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 38, "monotonic_millis": 1000, "first_violation_origin_step": 40, "violated_properties": []string{"p"}},
{"seed": 2, "exit_code": 0, "steps": 60, "actions": 55, "monotonic_millis": 1000, "first_violation_origin_step": 60, "violated_properties": []string{"p"}},
{"seed": 3, "exit_code": 0, "steps": 50, "actions": 47, "monotonic_millis": 1000, "first_violation_origin_step": 50, "violated_properties": []string{"p"}},
{"seed": 4, "exit_code": 0, "steps": 70, "actions": 66, "monotonic_millis": 1000, "first_violation_origin_step": 70, "violated_properties": []string{"p"}},
})
result := analyseCampaigns(t, "--paired", pre, post)
if result.Paired == nil {
t.Fatal("no paired comparison")
}
if result.Paired.Pairs != 3 {
t.Errorf("%d pairs, want the 3 seeds usable on both arms", result.Paired.Pairs)
}
if len(result.Paired.UnpairedSeeds) != 1 || result.Paired.UnpairedSeeds[0] != 4 {
t.Errorf("unpaired seeds %v, want [4]", result.Paired.UnpairedSeeds)
}
for _, summary := range result.Arms {
if len(summary.MissingSeeds) != 2 {
t.Errorf("arm %s missing seeds %v, want seeds 5 and 6", summary.Arm, summary.MissingSeeds)
}
}
var stdout bytes.Buffer
if err := run([]string{"--paired", pre, post}, &stdout, io.Discard); err != nil {
t.Fatal(err)
}
if !strings.Contains(stdout.String(), "usable in one arm only") {
t.Errorf("the report does not name the lost seed\n%s", stdout.String())
}
}
func TestRun_PairedRefusesAnythingOtherThanTwoArms(t *testing.T) {
root := t.TempDir()
var directories []string
for _, name := range []string{"a", "b", "c"} {
directory := filepath.Join(root, name)
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": 40, "seeds": []int{1}},
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
directories = append(directories, directory)
}
err := run(append([]string{"--paired"}, directories...), io.Discard, io.Discard)
if err == nil || !strings.Contains(err.Error(), "exactly two arms") {
t.Fatalf("error %v, want a refusal to pair three arms", err)
}
}
func TestRun_PairedRefusesArmsThatShareNoSeed(t *testing.T) {
root := t.TempDir()
north := filepath.Join(root, "north")
south := filepath.Join(root, "south")
writeCampaign(t, north, map[string]any{"arm": "north", "max_steps": 40, "seeds": []int{1, 2}},
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
writeCampaign(t, south, map[string]any{"arm": "south", "max_steps": 40, "seeds": []int{3, 4}},
[]map[string]any{{"seed": 3, "exit_code": 0, "steps": 40, "actions": 40}})
err := run([]string{"--paired", north, south}, io.Discard, io.Discard)
if err == nil || !strings.Contains(err.Error(), "share no seed") {
t.Fatalf("error %v, want a refusal to pair arms that ran different seeds", err)
}
}
+330
View File
@@ -0,0 +1,330 @@
package main
import (
"bufio"
"encoding/json"
"fmt"
"os"
"path/filepath"
"slices"
"strings"
)
const (
manifestFileName = "campaign.json"
recordsFileName = "runs.jsonl"
maxRecordBytes = 4 * 1024 * 1024
)
// manifest mirrors the fields analyze reads from campaign.json. The step budget
// lives here rather than in any run, because it is the exposure every run in an
// arm was given and the ceiling a clean run can be censored at.
type manifest struct {
Arm string `json:"arm"`
Generator string `json:"generator"`
Platform string `json:"platform"`
MaxSteps int `json:"max_steps"`
Seeds []int64 `json:"seeds"`
Host string `json:"host"`
}
// runRecord mirrors the fields analyze reads from one line of runs.jsonl.
type runRecord struct {
Seed int64 `json:"seed"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error"`
TimedOut bool `json:"timed_out"`
// MonotonicMillis is how long the run worked, and it is what every
// per-hour rate here divides by: a host asleep mid-run tested nothing, so
// charging that time to the arm would report it as slower for a reason
// that has nothing to do with the arm. The wall clock the campaign also
// records answers the other question, how much time passed.
MonotonicMillis int64 `json:"monotonic_millis"`
// DurationMillis is the name campaigns written before the two clocks were
// split gave the same monotonic reading, so those files still read.
DurationMillis int64 `json:"duration_millis"`
TraceError string `json:"trace_error"`
Steps int `json:"steps"`
FirstViolationOriginStep *int `json:"first_violation_origin_step"`
// FirstViolationDetectedStep is the step the violation was reported on,
// which is the origin step for a safety property tripping under its own
// action and the end of the budget for an obligation that never discharged.
// It is what the survival analysis times the event by; see eventStep.
FirstViolationDetectedStep *int `json:"first_violation_detected_step"`
ViolatedProperties []string `json:"violated_properties"`
// Actions is the count of steps on which the action generator dispatched an
// action, which excludes the steps the spec's setup drove before the
// generator was consulted wherever the trace names them. It is a pointer so
// that a runs.jsonl written before the campaign tool counted them is refused
// rather than read as an arm that acted zero times. The campaign tool always
// emits the field, so its absence dates the file.
Actions *int `json:"actions"`
// UnattributedActions is how many of the run's actions name no producer, so
// nothing can say whether the spec's setup drove them. It is a pointer
// because a runs.jsonl written before actions named a producer at all has no
// field, and reading that silence as none would let a denominator of unknown
// provenance pass for one that excludes the setup's login.
UnattributedActions *int `json:"unattributed_actions"`
}
// Exclusion reasons. A run that failed or timed out is missing data, not a
// censored observation: it broke off, so its step count is not exposure the app
// survived and counting it as one would bias the survival estimate downward.
const (
reasonLaunchError = "launch error"
reasonTimedOut = "timed out"
reasonNonzeroExit = "nonzero exit"
reasonTraceError = "unreadable trace"
reasonMalformedStep = "violation step outside the budget"
)
type classifiedRun struct {
Seed int64
Steps int
Actions int
UnattributedActions int
MonotonicMillis int64
OriginStep int
// EventStep is when the run could know, and it is what the survival
// analysis measures. It is the origin step whenever the two agree.
EventStep int
Violated bool
ClampedToBudget bool
ViolatedProperties []string
ExcludedBecause string
}
type arm struct {
Name string
Budget int
Generator string
Platform string
Directories []string
Runs []classifiedRun
MissingSeeds []int64
}
func loadCampaign(directory string) (manifest, []runRecord, error) {
body, err := os.ReadFile(filepath.Join(directory, manifestFileName))
if err != nil {
return manifest{}, nil, fmt.Errorf("read %s: %w", manifestFileName, err)
}
var declared manifest
if err := json.Unmarshal(body, &declared); err != nil {
return manifest{}, nil, fmt.Errorf("parse %s in %s: %w", manifestFileName, directory, err)
}
if declared.Arm == "" {
return manifest{}, nil, fmt.Errorf("%s in %s has no arm", manifestFileName, directory)
}
if declared.MaxSteps <= 0 {
return manifest{}, nil, fmt.Errorf("%s in %s has max_steps %d: clean runs have nothing to be censored at",
manifestFileName, directory, declared.MaxSteps)
}
file, err := os.Open(filepath.Join(directory, recordsFileName))
if err != nil {
return manifest{}, nil, fmt.Errorf("read %s: %w", recordsFileName, err)
}
defer file.Close()
var records []runRecord
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
lineNumber := 0
for scanner.Scan() {
lineNumber++
raw := strings.TrimSpace(scanner.Text())
if raw == "" {
continue
}
var record runRecord
if err := json.Unmarshal([]byte(raw), &record); err != nil {
return manifest{}, nil, fmt.Errorf("%s line %d in %s: %w", recordsFileName, lineNumber, directory, err)
}
if record.Actions == nil {
return manifest{}, nil, fmt.Errorf("%s line %d in %s has no actions count: it was written before "+
"dispatched actions were counted, and reading the missing count as zero would report every "+
"per-action rate wrongly; re-run the campaign to produce it",
recordsFileName, lineNumber, directory)
}
records = append(records, record)
}
if err := scanner.Err(); err != nil {
return manifest{}, nil, fmt.Errorf("read %s in %s: %w", recordsFileName, directory, err)
}
return declared, records, nil
}
// unattributedActions is how many of the record's actions carry no producer. A
// campaign written before the field existed carries none at all, so its whole
// action count is of unknown provenance rather than none of it.
func unattributedActions(record runRecord) int {
if record.UnattributedActions != nil {
return *record.UnattributedActions
}
if record.Actions == nil {
return 0
}
return *record.Actions
}
func (r runRecord) workingMillis() int64 {
if r.MonotonicMillis != 0 {
return r.MonotonicMillis
}
return r.DurationMillis
}
// classify turns one record into the run the analysis works with, deciding
// whether it is usable and, if it is, whether it is an event or censored.
func classify(record runRecord, budget int) classifiedRun {
item := classifiedRun{
Seed: record.Seed,
Steps: record.Steps,
MonotonicMillis: record.workingMillis(),
ViolatedProperties: slices.Clone(record.ViolatedProperties),
}
if record.Actions != nil {
item.Actions = *record.Actions
}
item.UnattributedActions = unattributedActions(record)
switch {
case record.LaunchError != "":
item.ExcludedBecause = reasonLaunchError
return item
case record.TimedOut:
item.ExcludedBecause = reasonTimedOut
return item
case record.ExitCode != 0:
item.ExcludedBecause = reasonNonzeroExit
return item
case record.TraceError != "":
item.ExcludedBecause = reasonTraceError
return item
}
if record.FirstViolationOriginStep == nil {
if len(record.ViolatedProperties) > 0 {
item.ExcludedBecause = reasonMalformedStep
}
return item
}
origin := *record.FirstViolationOriginStep
if origin < 1 {
item.ExcludedBecause = reasonMalformedStep
return item
}
item.Violated = true
item.OriginStep = origin
item.EventStep = eventStep(record, origin)
if item.EventStep > budget {
// The run-end finalize line reports obligations that never discharged
// at an index one past the last executed step. That is a real detection
// but not a real step, so it is held at the budget and counted.
item.EventStep = budget
item.ClampedToBudget = true
}
return item
}
// eventStep is the step at which the run could know it had violated. A safety
// property tripping under its own action is detected on the step that armed it
// and the two agree. An obligation that never discharges is reported when the
// run ends, and timing that at the step that armed it would record a liveness
// failure flushed at the budget as a violation found on the first step, which
// is a number the run cannot support and which no censored run can be compared
// against: the end of the run is the clock the clean runs are censored on, so
// the events have to be on it too. A campaign written before the field existed carries no
// detected step and keeps the origin.
func eventStep(record runRecord, origin int) int {
if record.FirstViolationDetectedStep == nil {
return origin
}
if detected := *record.FirstViolationDetectedStep; detected > origin {
return detected
}
return origin
}
// groupArms folds every campaign directory into its arm. Two directories with
// the same arm label are pooled, which is how a campaign split across hosts is
// analysed, but they must agree on the step budget.
func groupArms(directories []string) ([]arm, error) {
byName := map[string]*arm{}
var order []string
for _, directory := range directories {
declared, records, err := loadCampaign(directory)
if err != nil {
return nil, err
}
current, seen := byName[declared.Arm]
if !seen {
current = &arm{
Name: declared.Arm,
Budget: declared.MaxSteps,
Generator: declared.Generator,
Platform: declared.Platform,
}
byName[declared.Arm] = current
order = append(order, declared.Arm)
}
if current.Budget != declared.MaxSteps {
return nil, fmt.Errorf("arm %q has step budget %d in an earlier campaign and %d in %s: "+
"runs censored at different budgets cannot be pooled",
declared.Arm, current.Budget, declared.MaxSteps, directory)
}
current.Directories = append(current.Directories, directory)
present := map[int64]bool{}
for _, record := range records {
present[record.Seed] = true
current.Runs = append(current.Runs, classify(record, declared.MaxSteps))
}
for _, seed := range declared.Seeds {
if !present[seed] {
current.MissingSeeds = append(current.MissingSeeds, seed)
}
}
}
slices.Sort(order)
arms := make([]arm, 0, len(order))
for _, name := range order {
arms = append(arms, *byName[name])
}
return arms, nil
}
// unattributedActions is how much of the arm's per-action denominator has no
// producer behind it, counted over the runs the denominator is built from.
func (a arm) unattributedActions() int {
total := 0
for _, item := range a.Runs {
if item.ExcludedBecause != "" {
continue
}
total += item.UnattributedActions
}
return total
}
// observations returns the usable runs as survival data: an event at the step
// that armed the first violation, or a censored observation at the last step
// the run reached.
func (a arm) observations() []observation {
var result []observation
for _, item := range a.Runs {
if item.ExcludedBecause != "" {
continue
}
result = append(result, observationOf(item, a.Budget))
}
return result
}
func observationOf(item classifiedRun, budget int) observation {
if item.Violated {
return observation{Steps: float64(item.EventStep), Event: true}
}
// A run that hit the campaign's wall clock stopped short of the budget, and
// the steps it never ran are not exposure it survived.
return observation{Steps: float64(min(item.Steps, budget)), Event: false}
}
+218
View File
@@ -0,0 +1,218 @@
package main
import (
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
)
func stepPointer(value int) *int { return &value }
func TestClassify_FailedAndTimedOutRunsAreMissingDataNotCensored(t *testing.T) {
cases := []struct {
name string
record runRecord
reason string
}{
{"launch error", runRecord{LaunchError: "fork/exec: no such file"}, reasonLaunchError},
{"timed out", runRecord{TimedOut: true, ExitCode: -1}, reasonTimedOut},
{"nonzero exit", runRecord{ExitCode: 3}, reasonNonzeroExit},
{"unreadable trace", runRecord{TraceError: "no run directory with meta.json"}, reasonTraceError},
{"violation at step zero", runRecord{FirstViolationOriginStep: stepPointer(0)}, reasonMalformedStep},
{"violation without a step", runRecord{ViolatedProperties: []string{"cartTotal"}}, reasonMalformedStep},
}
for _, test := range cases {
item := classify(test.record, 50)
if item.ExcludedBecause != test.reason {
t.Errorf("%s: excluded because %q, want %q", test.name, item.ExcludedBecause, test.reason)
}
}
}
// A run also ends when the campaign's wall clock does, so a clean run can stop
// well short of the budget. Censoring it at the budget would credit it with
// steps it never ran, and a slower arm loses fewer steps in the same wall clock
// than a fast one, so the credit does not cancel between arms.
func TestClassify_CleanRunIsCensoredAtTheStepsItRan(t *testing.T) {
cases := []struct {
name string
steps int
budget int
censored float64
}{
{"stopped by the wall clock short of the budget", 12, 400, 12},
{"ran the whole budget", 400, 400, 400},
{"recorded more steps than the manifest budget", 420, 400, 400},
}
for _, test := range cases {
item := classify(runRecord{Seed: 4, Steps: test.steps, DurationMillis: 1000}, test.budget)
if item.ExcludedBecause != "" || item.Violated {
t.Fatalf("%s: run %+v, want a usable clean run", test.name, item)
}
current := arm{Budget: test.budget, Runs: []classifiedRun{item}}
observations := current.observations()
if len(observations) != 1 || observations[0].Event || observations[0].Steps != test.censored {
t.Errorf("%s: observations %+v, want one censored observation at %v",
test.name, observations, test.censored)
}
}
}
func TestClassify_ViolationIsAnEventAtTheOriginStep(t *testing.T) {
item := classify(runRecord{Seed: 5, Steps: 12, FirstViolationOriginStep: stepPointer(7)}, 50)
if !item.Violated || item.OriginStep != 7 || item.ClampedToBudget {
t.Fatalf("run %+v, want an unclamped event at step 7", item)
}
current := arm{Budget: 50, Runs: []classifiedRun{item}}
observations := current.observations()
if len(observations) != 1 || !observations[0].Event || observations[0].Steps != 7 {
t.Errorf("observations %+v, want one event at 7", observations)
}
}
// The run-end finalize line reports at an index one past the last executed step,
// so an origin past the budget is held at the budget and counted rather than
// silently turned into a censored run.
func TestClassify_ViolationPastTheBudgetIsHeldAtTheBudget(t *testing.T) {
item := classify(runRecord{FirstViolationOriginStep: stepPointer(51)}, 50)
if !item.Violated || item.EventStep != 50 || !item.ClampedToBudget {
t.Errorf("run %+v, want a clamped event at 50", item)
}
}
func writeCampaign(t *testing.T, directory string, declared map[string]any, records []map[string]any) {
t.Helper()
if err := os.MkdirAll(directory, 0o755); err != nil {
t.Fatal(err)
}
body, err := json.MarshalIndent(declared, "", " ")
if err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(directory, manifestFileName), append(body, '\n'), 0o644); err != nil {
t.Fatal(err)
}
var lines strings.Builder
for _, record := range records {
line, err := json.Marshal(record)
if err != nil {
t.Fatal(err)
}
lines.Write(line)
lines.WriteByte('\n')
}
if err := os.WriteFile(filepath.Join(directory, recordsFileName), []byte(lines.String()), 0o644); err != nil {
t.Fatal(err)
}
}
func TestClassify_CarriesTheDispatchedActionCount(t *testing.T) {
actions := 7
item := classify(runRecord{Seed: 4, Steps: 30, Actions: &actions}, 50)
if item.Steps != 30 || item.Actions != 7 {
t.Errorf("run %+v, want 30 steps and 7 actions", item)
}
}
func TestGroupArms_PoolsDirectoriesSharingAnArmAndReportsMissingSeeds(t *testing.T) {
root := t.TempDir()
writeCampaign(t, filepath.Join(root, "north"), map[string]any{
"arm": "seeded", "max_steps": 40, "seeds": []int{1, 2, 3},
}, []map[string]any{
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 33},
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 8, "first_violation_origin_step": 9},
})
writeCampaign(t, filepath.Join(root, "south"), map[string]any{
"arm": "seeded", "max_steps": 40, "seeds": []int{4},
}, []map[string]any{
{"seed": 4, "exit_code": 0, "steps": 40, "actions": 40},
})
arms, err := groupArms([]string{filepath.Join(root, "north"), filepath.Join(root, "south")})
if err != nil {
t.Fatal(err)
}
if len(arms) != 1 {
t.Fatalf("%d arms, want 1", len(arms))
}
if len(arms[0].Runs) != 3 {
t.Errorf("%d runs, want 3", len(arms[0].Runs))
}
if len(arms[0].MissingSeeds) != 1 || arms[0].MissingSeeds[0] != 3 {
t.Errorf("missing seeds %v, want [3]", arms[0].MissingSeeds)
}
if len(arms[0].Directories) != 2 {
t.Errorf("directories %v, want both", arms[0].Directories)
}
}
func TestGroupArms_RejectsDisagreeingStepBudgets(t *testing.T) {
root := t.TempDir()
writeCampaign(t, filepath.Join(root, "a"), map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}},
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
writeCampaign(t, filepath.Join(root, "b"), map[string]any{"arm": "seeded", "max_steps": 80, "seeds": []int{2}},
[]map[string]any{{"seed": 2, "exit_code": 0, "steps": 80, "actions": 80}})
_, err := groupArms([]string{filepath.Join(root, "a"), filepath.Join(root, "b")})
if err == nil || !strings.Contains(err.Error(), "different budgets") {
t.Fatalf("error %v, want a refusal to pool different budgets", err)
}
}
func TestGroupArms_RejectsAMissingStepBudget(t *testing.T) {
root := t.TempDir()
writeCampaign(t, filepath.Join(root, "a"), map[string]any{"arm": "seeded", "seeds": []int{1}},
[]map[string]any{{"seed": 1, "exit_code": 0}})
_, err := groupArms([]string{filepath.Join(root, "a")})
if err == nil || !strings.Contains(err.Error(), "censored at") {
t.Fatalf("error %v, want a complaint about max_steps", err)
}
}
func TestGroupArms_ReportsBadRecordLines(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "a")
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}}, nil)
if err := os.WriteFile(filepath.Join(directory, recordsFileName),
[]byte("{\"seed\":1,\"actions\":0}\nnot json\n"), 0o644); err != nil {
t.Fatal(err)
}
_, err := groupArms([]string{directory})
if err == nil || !strings.Contains(err.Error(), "line 2") {
t.Fatalf("error %v, want the offending line number", err)
}
}
// A runs.jsonl written before the campaign counted dispatched actions has no
// such field. Reading the absence as zero would divide by zero, so the whole
// campaign is refused instead.
func TestGroupArms_RefusesRecordsWithoutADispatchedActionCount(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "old-format")
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1, 2}},
[]map[string]any{
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40},
{"seed": 2, "exit_code": 0, "steps": 40},
})
_, err := groupArms([]string{directory})
if err == nil {
t.Fatal("a runs.jsonl without an action count was accepted")
}
for _, fragment := range []string{"line 2", "actions", directory} {
if !strings.Contains(err.Error(), fragment) {
t.Errorf("error %q is missing %q", err, fragment)
}
}
}
func TestGroupArms_RefusesAnExcludedRecordWithoutAnActionCount(t *testing.T) {
root := t.TempDir()
directory := filepath.Join(root, "old-format")
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}},
[]map[string]any{{"seed": 1, "exit_code": 3}})
if _, err := groupArms([]string{directory}); err == nil {
t.Fatal("an old-format record was accepted because the run was excluded anyway")
}
}
+121
View File
@@ -0,0 +1,121 @@
// Command analyze reduces campaign directories to the statistics the
// evaluation reports. The primary outcome is steps to first violation, with
// clean runs right-censored where they stopped rather than discarded: defect
// yield per run is a binary that would need on the order of eighty runs an arm
// to separate, while survival analysis uses every run, including the clean ones.
package main
import (
"encoding/json"
"errors"
"flag"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"time"
)
const usage = `analyze reports the statistics of a sanderling evaluation from campaign directories.
Usage:
analyze [--json <path>] <campaign-dir> [<campaign-dir> ...]
Each directory is one produced by the campaign tool and must hold campaign.json
and runs.jsonl. Directories sharing an arm label are pooled and must agree on
the step budget, and arms compared against each other must agree on it too.
One invocation is one research question: Holm corrects across the comparisons it
produces and across nothing else.
`
type stringList []string
func (list *stringList) String() string { return strings.Join(*list, ",") }
func (list *stringList) Set(value string) error {
if strings.TrimSpace(value) == "" {
return errors.New("empty campaign directory")
}
*list = append(*list, value)
return nil
}
func run(arguments []string, stdout, stderr io.Writer) error {
flagSet := flag.NewFlagSet("analyze", flag.ContinueOnError)
flagSet.SetOutput(stderr)
flagSet.Usage = func() {
fmt.Fprint(stderr, usage)
flagSet.PrintDefaults()
}
var directories stringList
var jsonPath string
var question string
var paired bool
flagSet.Var(&directories, "campaign", "campaign directory to read; repeat for more, or pass them as arguments")
flagSet.StringVar(&jsonPath, "json", "", "write the machine-readable summary here, or - for stdout")
flagSet.StringVar(&question, "question", "", "the research question these campaigns answer; Holm corrects within one invocation, and this records which family that was")
flagSet.BoolVar(&paired, "paired", false, "the two arms ran the same seeds: contrast them seed by seed with the sign test instead of pooling them into two independent samples")
if err := flagSet.Parse(arguments); err != nil {
return err
}
directories = append(directories, flagSet.Args()...)
if len(directories) == 0 {
return errors.New("no campaign directories given")
}
seen := map[string]bool{}
for _, directory := range directories {
resolved, err := filepath.Abs(directory)
if err != nil {
return fmt.Errorf("resolve %s: %w", directory, err)
}
if seen[resolved] {
return fmt.Errorf("campaign directory %s given twice: its runs would be counted twice", directory)
}
seen[resolved] = true
}
arms, err := groupArms(directories)
if err != nil {
return err
}
var result analysis
if paired {
result, err = analysePaired(arms, time.Now().UTC())
if err != nil {
return err
}
} else {
result, err = analyse(arms, time.Now().UTC())
if err != nil {
return err
}
}
result.Question = question
writeReport(result, stdout)
if jsonPath == "" {
return nil
}
body, err := json.MarshalIndent(result, "", " ")
if err != nil {
return fmt.Errorf("marshal summary: %w", err)
}
body = append(body, '\n')
if jsonPath == "-" {
_, err = stdout.Write(body)
return err
}
return os.WriteFile(jsonPath, body, 0o644)
}
func main() {
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
if errors.Is(err, flag.ErrHelp) {
return
}
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
}
+174
View File
@@ -0,0 +1,174 @@
package main
import (
"fmt"
"math"
"slices"
)
// signTest is the exact two-sided sign test over matched pairs. Under the null
// that neither arm reaches its first violation sooner, a pair whose order the
// censoring determines falls either way with probability one half, so the count
// is binomial and the two-sided p-value doubles the smaller tail. Pairs left
// with no order carry no information and are not trials.
//
// It is the seed-matched form of the comparison the unpaired test makes, and
// it is what the log-rank stratified by seed reduces to with one run per arm in
// each stratum. The magnitude-based alternatives are not available: a
// difference in steps needs both runs to have violated, and a paired test built
// on scores of censored times, the paired Prentice-Wilcoxon among them, is
// centred at zero under the null only when the two arms censor alike, which is
// exactly what the wall clock stops them from doing.
func signTest(favouringFirst, favouringSecond int) float64 {
trials := favouringFirst + favouringSecond
if trials == 0 {
return math.NaN()
}
smaller := min(favouringFirst, favouringSecond)
tail := 0.0
for count := 0; count <= smaller; count++ {
tail += math.Exp(logBinomialCoefficient(trials, count) - float64(trials)*math.Ln2)
}
return math.Min(2*tail, 1)
}
func logBinomialCoefficient(trials, chosen int) float64 {
all, _ := math.Lgamma(float64(trials + 1))
picked, _ := math.Lgamma(float64(chosen + 1))
rest, _ := math.Lgamma(float64(trials-chosen) + 1)
return all - picked - rest
}
// pairedComparison is the seed-matched contrast the actuation ablation reports.
// A pair is scored the way the unpaired comparison scores one, by which run
// outlived the other, so Sign is +1 when the second arm is the one seen to
// violate sooner across the pairs whose order censoring determines.
type pairedComparison struct {
First string `json:"first"`
Second string `json:"second"`
Pairs int `json:"pairs"`
UnpairedSeeds []int64 `json:"unpaired_seeds,omitempty"`
Sign int `json:"sign"`
FirstSooner int `json:"first_sooner"`
SecondSooner int `json:"second_sooner"`
// Unordered is the pairs the censoring leaves in no order, either because
// both runs ended clean or because the run that stopped first stopped before
// the other violated. They are not evidence either way and are not trials.
Unordered int `json:"unordered_pairs"`
// MedianDifference is in steps and is undefined unless some pair has both
// runs violating, which is the only shape a difference in steps can be read
// off. BothViolated says how many pairs it summarizes, because it describes
// those pairs and not the sample.
MedianDifference *float64 `json:"median_step_difference"`
BothViolated int `json:"both_violated_pairs"`
// A12 is the within-pair form of the Vargha-Delaney effect size, the share
// of matched seeds on which the first arm took more steps, an unordered pair
// counting as half. A matched design has no reason to compare the two arms
// as pooled bags of runs when each seed has a partner.
A12 float64 `json:"a12_within_pairs"`
// PValue is undefined, and null in the summary, when censoring left no pair
// ordered: there is nothing for the test to be a test of, and JSON has no
// way to write the number that is not there.
PValue *float64 `json:"p_value"`
HolmPValue *float64 `json:"holm_p_value"`
}
// pairArms matches the two arms by seed and contrasts them pair by pair. A seed
// usable in one arm and not the other is named rather than dropped silently,
// because that is a host that lost a run and it is what the campaign manifest
// exists to make visible.
func pairArms(first, second arm) (pairedComparison, error) {
firstBySeed, err := usableBySeed(first)
if err != nil {
return pairedComparison{}, err
}
secondBySeed, err := usableBySeed(second)
if err != nil {
return pairedComparison{}, err
}
comparison := pairedComparison{First: first.Name, Second: second.Name, A12: math.NaN()}
var differences []float64
for _, seed := range sortedSeeds(firstBySeed, secondBySeed) {
left, inFirst := firstBySeed[seed]
right, inSecond := secondBySeed[seed]
if !inFirst || !inSecond {
comparison.UnpairedSeeds = append(comparison.UnpairedSeeds, seed)
continue
}
leftRun := observationOf(left, first.Budget)
rightRun := observationOf(right, second.Budget)
comparison.Pairs++
switch outlives(leftRun, rightRun) {
case 1:
comparison.SecondSooner++
case -1:
comparison.FirstSooner++
default:
comparison.Unordered++
}
if leftRun.Event && rightRun.Event {
comparison.BothViolated++
differences = append(differences, leftRun.Steps-rightRun.Steps)
}
}
if comparison.Pairs == 0 {
return comparison, nil
}
if len(differences) > 0 {
median := medianOf(differences)
comparison.MedianDifference = &median
}
switch {
case comparison.SecondSooner > comparison.FirstSooner:
comparison.Sign = 1
case comparison.FirstSooner > comparison.SecondSooner:
comparison.Sign = -1
}
comparison.A12 = (float64(comparison.SecondSooner) + 0.5*float64(comparison.Unordered)) / float64(comparison.Pairs)
if tested := signTest(comparison.FirstSooner, comparison.SecondSooner); !math.IsNaN(tested) {
comparison.PValue = &tested
}
return comparison, nil
}
func usableBySeed(current arm) (map[int64]classifiedRun, error) {
bySeed := map[int64]classifiedRun{}
for _, item := range current.Runs {
if item.ExcludedBecause != "" {
continue
}
if _, repeated := bySeed[item.Seed]; repeated {
return nil, fmt.Errorf("arm %q has more than one usable run for seed %d: a seed-matched "+
"comparison cannot choose between them", current.Name, item.Seed)
}
bySeed[item.Seed] = item
}
return bySeed, nil
}
func sortedSeeds(sets ...map[int64]classifiedRun) []int64 {
var seeds []int64
seen := map[int64]bool{}
for _, set := range sets {
for seed := range set {
if seen[seed] {
continue
}
seen[seed] = true
seeds = append(seeds, seed)
}
}
slices.Sort(seeds)
return seeds
}
func medianOf(values []float64) float64 {
sorted := slices.Sorted(slices.Values(values))
middle := len(sorted) / 2
if len(sorted)%2 == 1 {
return sorted[middle]
}
return (sorted[middle-1] + sorted[middle]) / 2
}
+190
View File
@@ -0,0 +1,190 @@
package main
import (
"math"
"testing"
)
// The two-sided sign test is R's binom.test(k, n) at p = 0.5, which is the
// doubled tail of a symmetric binomial and can be worked out by hand from the
// coefficients: 2 * sum(C(n, i), i <= min(k, n-k)) / 2^n.
func TestSignTest_MatchesTheBinomialTail(t *testing.T) {
cases := []struct {
first, second int
want float64
}{
{0, 10, 2.0 / 1024},
{1, 9, 2 * 11.0 / 1024},
{3, 7, 2 * 176.0 / 1024},
{5, 5, 1},
{0, 1, 1},
{2, 0, 0.5},
}
for _, test := range cases {
got := signTest(test.first, test.second)
if math.Abs(got-test.want) > 1e-12 {
t.Errorf("sign test on %d against %d gives %v, want %v", test.first, test.second, got, test.want)
}
if reversed := signTest(test.second, test.first); math.Abs(reversed-got) > 1e-12 {
t.Errorf("sign test on %d against %d gives %v reversed and %v forward",
test.first, test.second, reversed, got)
}
}
}
// A campaign runs tens of seeds, not tens of thousands, but the tail is summed
// through log-gamma rather than through factorials so that a lopsided family
// stays a number rather than becoming an overflow.
func TestSignTest_LargeCountsStayFinite(t *testing.T) {
if got := signTest(0, 200); got <= 0 || got > 1e-59 {
t.Errorf("sign test on 0 against 200 gives %v, want a positive value around 2^-199", got)
}
if got := signTest(100, 100); math.Abs(got-1) > 1e-12 {
t.Errorf("sign test on an even split gives %v, want 1", got)
}
}
func TestSignTest_NoOrderedPairHasNoTest(t *testing.T) {
if got := signTest(0, 0); !math.IsNaN(got) {
t.Errorf("sign test with nothing to test gives %v, want undefined", got)
}
}
func TestPairArms_ScoresEachPairByWhichRunOutlivedTheOther(t *testing.T) {
pre := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
violatingRun(1, 30, 30, "doubleTapCharges"),
cleanRun(2, 40),
violatingRun(3, 25, 25, "doubleTapCharges"),
}}
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{
violatingRun(1, 10, 10, "doubleTapCharges"),
violatingRun(2, 12, 12, "doubleTapCharges"),
violatingRun(3, 25, 25, "doubleTapCharges"),
}}
comparison, err := pairArms(pre, post)
if err != nil {
t.Fatal(err)
}
if comparison.Pairs != 3 {
t.Fatalf("%d pairs, want 3", comparison.Pairs)
}
// Seed 1 violated at 30 against 10 and seed 2 was still clean at 40 when its
// partner violated at 12, so both go to the second arm; seed 3 violated on
// the same step in both and has no order.
if comparison.SecondSooner != 2 || comparison.FirstSooner != 0 || comparison.Unordered != 1 {
t.Errorf("counts %+v, want two favouring the second arm and one unordered", comparison)
}
if comparison.Sign != 1 {
t.Errorf("sign %d, want +1 for the arm that violated later", comparison.Sign)
}
if math.Abs(comparison.A12-2.5/3) > 1e-12 {
t.Errorf("a12 within pairs %v, want %v", comparison.A12, 2.5/3)
}
// Only seeds 1 and 3 have a difference in steps to take a median of, 20 and
// 0: the pair holding a clean run has no difference either arm supports.
if comparison.BothViolated != 2 || comparison.MedianDifference == nil || *comparison.MedianDifference != 10 {
t.Errorf("median difference %v over %d pair(s), want 10 over 2",
comparison.MedianDifference, comparison.BothViolated)
}
if want := signTest(0, 2); comparison.PValue == nil || *comparison.PValue != want {
t.Errorf("p %v, want the sign test's %v over the two ordered pairs", comparison.PValue, want)
}
}
// Two clean runs are two runs that were still going when they stopped, whatever
// step each stopped on, so the pair says nothing and is not a trial.
func TestPairArms_PairsOfCleanRunsAreNotEvidence(t *testing.T) {
early := arm{Name: "early", Budget: 400, Runs: []classifiedRun{cleanRun(1, 12), cleanRun(2, 14)}}
late := arm{Name: "late", Budget: 400, Runs: []classifiedRun{cleanRun(1, 400), cleanRun(2, 380)}}
comparison, err := pairArms(early, late)
if err != nil {
t.Fatal(err)
}
if comparison.Unordered != 2 || comparison.Sign != 0 {
t.Errorf("comparison %+v, want both pairs unordered and no direction", comparison)
}
if comparison.PValue != nil {
t.Errorf("p %v, want undefined with no ordered pair", *comparison.PValue)
}
if comparison.MedianDifference != nil {
t.Errorf("median difference %v, want undefined where no pair has two violations",
*comparison.MedianDifference)
}
if comparison.A12 != 0.5 {
t.Errorf("a12 within pairs %v, want 0.5", comparison.A12)
}
}
// A run excluded as missing data cannot be paired against anything, and the
// seed it came from has to be named rather than silently shrinking the sample.
func TestPairArms_NamesSeedsUsableInOneArmOnly(t *testing.T) {
pre := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
violatingRun(1, 30, 30, "p"),
{Seed: 2, ExcludedBecause: reasonTimedOut},
violatingRun(3, 20, 20, "p"),
}}
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{
violatingRun(1, 10, 10, "p"),
violatingRun(2, 11, 11, "p"),
}}
comparison, err := pairArms(pre, post)
if err != nil {
t.Fatal(err)
}
if comparison.Pairs != 1 {
t.Fatalf("%d pairs, want 1", comparison.Pairs)
}
if len(comparison.UnpairedSeeds) != 2 || comparison.UnpairedSeeds[0] != 2 || comparison.UnpairedSeeds[1] != 3 {
t.Errorf("unpaired seeds %v, want [2 3]", comparison.UnpairedSeeds)
}
}
func TestPairArms_RefusesTwoUsableRunsForOneSeed(t *testing.T) {
pooled := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
violatingRun(1, 30, 30, "p"),
violatingRun(1, 12, 12, "p"),
}}
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{violatingRun(1, 10, 10, "p")}}
if _, err := pairArms(pooled, post); err == nil {
t.Fatal("paired two arms where one seed ran twice")
}
}
// The paired comparison is the ablation's decision rule, so the direction it
// reports has to survive the arms being passed the other way round.
func TestPairArms_DirectionReversesWithTheArms(t *testing.T) {
pre := arm{Name: "pre", Budget: 400, Runs: []classifiedRun{
cleanRun(1, 400), cleanRun(2, 400), violatingRun(3, 380, 380, "p"),
cleanRun(4, 400), violatingRun(5, 350, 350, "p"),
}}
post := arm{Name: "post", Budget: 400, Runs: []classifiedRun{
violatingRun(1, 40, 40, "p"), violatingRun(2, 90, 90, "p"), violatingRun(3, 60, 60, "p"),
violatingRun(4, 120, 120, "p"), violatingRun(5, 30, 30, "p"),
}}
forward, err := pairArms(pre, post)
if err != nil {
t.Fatal(err)
}
reversed, err := pairArms(post, pre)
if err != nil {
t.Fatal(err)
}
if forward.Sign != 1 || reversed.Sign != -1 {
t.Errorf("signs %+d and %+d, want +1 then -1", forward.Sign, reversed.Sign)
}
if *forward.MedianDifference != -*reversed.MedianDifference {
t.Errorf("median differences %v and %v, want opposites",
*forward.MedianDifference, *reversed.MedianDifference)
}
if math.Abs(*forward.PValue-*reversed.PValue) > 1e-12 {
t.Errorf("p-values %v and %v, want the same two-sided value", *forward.PValue, *reversed.PValue)
}
if math.Abs(forward.A12+reversed.A12-1) > 1e-12 {
t.Errorf("a12 %v and %v, want them to sum to 1", forward.A12, reversed.A12)
}
}
+676
View File
@@ -0,0 +1,676 @@
package main
import (
"encoding/json"
"fmt"
"io"
"math"
"math/rand"
"os"
"path/filepath"
"testing"
)
// A pipeline exercised only on data whose answer nobody knows reports that it
// runs, not that it is right. These tests plant effects whose value is known in
// advance from the model that generated the data, and require the pipeline to
// recover them from campaign directories it reads off disk through its own
// entry point.
//
// plantedModel is the generator: a run that has not yet violated does so at
// every step with probability Hazard, and a run reaching Budget without
// violating is right-censored there. Steps to first violation is therefore
// geometric, truncated at the budget, and every quantity the pipeline reports
// about it has a closed form below.
type plantedModel struct {
Hazard float64
Budget int
}
// recordedValues is the distribution of what one run contributes to the
// analysis: the violation step when it violates, and the budget when it does
// not, since a censored run is held at the budget. Index t carries P(value = t).
func (model plantedModel) recordedValues() []float64 {
probabilities := make([]float64, model.Budget+1)
survival := 1.0
for step := 1; step < model.Budget; step++ {
probabilities[step] = survival * model.Hazard
survival *= 1 - model.Hazard
}
probabilities[model.Budget] = survival
return probabilities
}
func (model plantedModel) survivalAt(step int) float64 {
return math.Pow(1-model.Hazard, float64(step))
}
func (model plantedModel) censoredShare() float64 {
return model.survivalAt(model.Budget)
}
// medianSteps is the smallest step at which the true survival function falls to
// or below one half, which is what the Kaplan-Meier median estimates.
func (model plantedModel) medianSteps() (int, bool) {
for step := 1; step <= model.Budget; step++ {
if model.survivalAt(step) <= 0.5 {
return step, true
}
}
return 0, false
}
// populationA12 is the Vargha-Delaney effect size between two planted models,
// computed from the models rather than from any sample: the probability that a
// run of the first takes more steps than a run of the second, counting a tie as
// half. Every run held at the budget ties with every other, which is the term a
// pipeline mishandling censoring gets wrong.
func populationA12(first, second plantedModel) float64 {
left, right := first.recordedValues(), second.recordedValues()
total := 0.0
for value, probability := range left {
if probability == 0 {
continue
}
for other, otherProbability := range right {
switch {
case value > other:
total += probability * otherProbability
case value == other:
total += 0.5 * probability * otherProbability
}
}
}
return total
}
type plantedRun struct {
seed int64
steps int
violated bool
}
func (model plantedModel) draw(seed int64, source *rand.Rand) plantedRun {
for step := 1; step <= model.Budget; step++ {
if source.Float64() < model.Hazard {
return plantedRun{seed: seed, steps: step, violated: true}
}
}
return plantedRun{seed: seed, steps: model.Budget}
}
func (model plantedModel) drawCampaign(source *rand.Rand, seeds int) []plantedRun {
runs := make([]plantedRun, 0, seeds)
for seed := int64(1); seed <= int64(seeds); seed++ {
runs = append(runs, model.draw(seed, source))
}
return runs
}
// writePlantedCampaign emits the campaign directory shape the campaign tool
// writes, so the planted data reaches the statistics through the same loader,
// classifier and censoring rules as a real sweep.
func writePlantedCampaign(t *testing.T, directory, armName string, budget int, runs []plantedRun) string {
t.Helper()
seeds := make([]int64, 0, len(runs))
records := make([]map[string]any, 0, len(runs))
for _, run := range runs {
seeds = append(seeds, run.seed)
record := map[string]any{
"seed": run.seed,
"exit_code": 0,
"steps": run.steps,
"actions": run.steps,
"monotonic_millis": int64(run.steps) * 900,
}
if run.violated {
record["first_violation_origin_step"] = run.steps
record["violated_properties"] = []string{"plantedProperty"}
} else {
record["first_violation_origin_step"] = nil
}
records = append(records, record)
}
writeCampaign(t, directory, map[string]any{
"arm": armName,
"generator": "seeded",
"platform": "android",
"max_steps": budget,
"seeds": seeds,
"host": "planted",
}, records)
return directory
}
// analyseCampaigns runs the tool exactly as the command line does and returns
// the machine-readable summary it wrote.
func analyseCampaigns(t *testing.T, arguments ...string) analysis {
t.Helper()
summaryPath := filepath.Join(t.TempDir(), "analysis.json")
if err := run(append([]string{"--json", summaryPath}, arguments...), io.Discard, io.Discard); err != nil {
t.Fatal(err)
}
body, err := os.ReadFile(summaryPath)
if err != nil {
t.Fatal(err)
}
var result analysis
if err := json.Unmarshal(body, &result); err != nil {
t.Fatal(err)
}
return result
}
func armByName(t *testing.T, result analysis, name string) armSummary {
t.Helper()
for _, summary := range result.Arms {
if summary.Arm == name {
return summary
}
}
t.Fatalf("no arm %q in %v", name, result.Arms)
return armSummary{}
}
// plantTwoArms writes two campaign directories drawn from the given models and
// returns them in the order they were written.
func plantTwoArms(t *testing.T, sourceSeed int64, seeds int, first, second plantedModel) (string, string) {
t.Helper()
root := t.TempDir()
source := rand.New(rand.NewSource(sourceSeed))
firstDirectory := writePlantedCampaign(t, filepath.Join(root, "first"), "first", first.Budget,
first.drawCampaign(source, seeds))
secondDirectory := writePlantedCampaign(t, filepath.Join(root, "second"), "second", second.Budget,
second.drawCampaign(source, seeds))
return firstDirectory, secondDirectory
}
// The effect size is the number the paper reports as the size of a difference,
// so it is checked against the value the generating models define rather than
// against anything this tool produced. Four independent draws are used because
// a tolerance that holds for one lucky sample is not a check.
func TestPlanted_RecoversTheKnownEffectSize(t *testing.T) {
slow := plantedModel{Hazard: 0.005, Budget: 400}
fast := plantedModel{Hazard: 0.02, Budget: 400}
expected := populationA12(slow, fast)
if expected < 0.6 {
t.Fatalf("planted effect %v is too small to be worth checking", expected)
}
for _, sourceSeed := range []int64{1, 2, 3, 4} {
slowDirectory, fastDirectory := plantTwoArms(t, sourceSeed, 400, slow, fast)
result := analyseCampaigns(t, slowDirectory, fastDirectory)
if len(result.Pairwise) != 1 {
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
}
pair := result.Pairwise[0]
if pair.First != "first" || pair.Second != "second" {
t.Fatalf("comparison %s vs %s, want first vs second", pair.First, pair.Second)
}
if math.Abs(pair.A12-expected) > 0.03 {
t.Errorf("seed %d: a12 %.4f, want the planted %.4f", sourceSeed, pair.A12, expected)
}
if pair.PValue > 1e-6 {
t.Errorf("seed %d: p-value %v for an effect this large", sourceSeed, pair.PValue)
}
}
}
// A12 above one half has to mean the first arm took longer, and a sign flip
// anywhere in the pipeline is the failure that produces a plausible wrong
// answer rather than an obvious one.
func TestPlanted_EffectSizeDirectionFollowsTheArmOrder(t *testing.T) {
slow := plantedModel{Hazard: 0.005, Budget: 400}
fast := plantedModel{Hazard: 0.02, Budget: 400}
source := rand.New(rand.NewSource(7))
slowRuns := slow.drawCampaign(source, 300)
fastRuns := fast.drawCampaign(source, 300)
// Comparisons are ordered by arm label, so the same two samples are
// analysed twice with the labels swapped.
root := t.TempDir()
forward := analyseCampaigns(t,
writePlantedCampaign(t, filepath.Join(root, "slow-first"), "a", slow.Budget, slowRuns),
writePlantedCampaign(t, filepath.Join(root, "fast-second"), "b", fast.Budget, fastRuns)).Pairwise[0]
reversed := analyseCampaigns(t,
writePlantedCampaign(t, filepath.Join(root, "slow-second"), "b", slow.Budget, slowRuns),
writePlantedCampaign(t, filepath.Join(root, "fast-first"), "a", fast.Budget, fastRuns)).Pairwise[0]
if forward.A12 <= 0.5 {
t.Errorf("a12 %.4f with the slower arm first, want above 0.5", forward.A12)
}
if math.Abs(forward.A12+reversed.A12-1) > 1e-12 {
t.Errorf("a12 %.4f and %.4f, want them to sum to 1", forward.A12, reversed.A12)
}
if math.Abs(forward.PValue-reversed.PValue) > 1e-12 {
t.Errorf("p-values %v and %v, want the same two-sided value", forward.PValue, reversed.PValue)
}
}
// The median and the quartiles are read off the Kaplan-Meier curve, and the
// curve itself is checked point by point against the survival function that
// generated the data. A pipeline counting a censored run as a violation drives
// this estimate down hard.
func TestPlanted_SurvivalCurveTracksTheGeneratingHazard(t *testing.T) {
model := plantedModel{Hazard: 0.005, Budget: 400}
if model.censoredShare() < 0.1 {
t.Fatalf("planted censoring is %.3f, too little to exercise the estimator", model.censoredShare())
}
directory, _ := plantTwoArms(t, 11, 600, model, model)
summary := armByName(t, analyseCampaigns(t, directory), "first")
for _, step := range []int{20, 50, 100, 200, 300} {
estimated, ok := survivalAtStep(summary.SurvivalCurve, float64(step))
if !ok {
t.Fatalf("curve has no point at or before step %d", step)
}
if want := model.survivalAt(step); math.Abs(estimated-want) > 0.05 {
t.Errorf("survival at %d is %.4f, want the planted %.4f", step, estimated, want)
}
}
median, ok := model.medianSteps()
if !ok {
t.Fatal("the planted model has no median to recover")
}
if summary.MedianStepsToFirstViolation == nil {
t.Fatal("median undefined, want the planted median")
}
if math.Abs(*summary.MedianStepsToFirstViolation-float64(median)) > 10 {
t.Errorf("median %v, want the planted %d", *summary.MedianStepsToFirstViolation, median)
}
assertQuartile(t, summary.FirstQuartileSteps, model, 0.25)
assertQuartile(t, summary.ThirdQuartileSteps, model, 0.75)
}
func assertQuartile(t *testing.T, got *float64, model plantedModel, fraction float64) {
t.Helper()
if got == nil {
t.Fatalf("quantile %v undefined, want a value", fraction)
}
if want := model.survivalAt(int(*got)); want > 1-fraction+0.05 {
t.Errorf("quantile %v at step %v, where the planted survival is %.4f", fraction, *got, want)
}
}
func survivalAtStep(curve []survivalPoint, step float64) (float64, bool) {
value, found := 0.0, false
for _, point := range curve {
if point.Steps > step {
break
}
value, found = point.Survival, true
}
return value, found
}
// A true null must not be called significant more often than the test's own
// level allows. This is the check that catches a wrong variance term, which
// leaves every point estimate looking reasonable and only shows up in how often
// the pipeline claims a difference that is not there.
func TestPlanted_TrueNullIsNotCalledSignificantAboveItsLevel(t *testing.T) {
const replicates = 300
model := plantedModel{Hazard: 0.01, Budget: 400}
rejectedByGehan, rejectedByLogRank := 0, 0
for replicate := 0; replicate < replicates; replicate++ {
first, second := plantTwoArms(t, int64(1000+replicate), 30, model, model)
result := analyseCampaigns(t, first, second)
if result.Pairwise[0].PValue < 0.05 {
rejectedByGehan++
}
if result.LogRank.PValue < 0.05 {
rejectedByLogRank++
}
}
// Three standard errors around 0.05 at 300 replicates is 0.05 +/- 0.038.
// The lower bound is asserted too: a test that never rejects has bought its
// level by losing the power the experiment is sized for.
for name, rejected := range map[string]int{"gehan": rejectedByGehan, "log-rank": rejectedByLogRank} {
rate := float64(rejected) / replicates
if rate > 0.09 || rate < 0.015 {
t.Errorf("%s called a true null significant in %.1f%% of %d replicates, want about 5%%",
name, 100*rate, replicates)
}
}
}
// The same null with most runs censored, where nearly every observation is tied
// at the budget and the tie correction is what keeps the level honest.
func TestPlanted_TrueNullUnderHeavyCensoringKeepsItsLevel(t *testing.T) {
const replicates = 300
model := plantedModel{Hazard: 0.0004, Budget: 400}
if model.censoredShare() < 0.8 {
t.Fatalf("planted censoring is %.3f, want most runs censored", model.censoredShare())
}
rejected := 0
for replicate := 0; replicate < replicates; replicate++ {
first, second := plantTwoArms(t, int64(9000+replicate), 40, model, model)
if analyseCampaigns(t, first, second).Pairwise[0].PValue < 0.05 {
rejected++
}
}
if rate := float64(rejected) / replicates; rate > 0.09 {
t.Errorf("gehan called a true null significant in %.1f%% of %d replicates under heavy censoring, want at most about 5%%",
100*rate, replicates)
}
}
// Most runs never violating is the expected shape of an arm at a budget chosen
// for a harder subject, and it is the case where an implementation that drops
// censored runs instead of holding them at the budget still produces a number.
func TestPlanted_HeavyCensoringLeavesTheMedianUndefinedAndKeepsTheEffect(t *testing.T) {
quiet := plantedModel{Hazard: 0.00025, Budget: 400}
loud := plantedModel{Hazard: 0.002, Budget: 400}
if quiet.censoredShare() < 0.85 {
t.Fatalf("planted censoring is %.3f, want most runs censored", quiet.censoredShare())
}
if _, ok := quiet.medianSteps(); ok {
t.Fatal("the heavily censored model has a median, so the test is not testing what it says")
}
quietDirectory, loudDirectory := plantTwoArms(t, 21, 400, quiet, loud)
result := analyseCampaigns(t, quietDirectory, loudDirectory)
summary := armByName(t, result, "first")
if summary.MedianStepsToFirstViolation != nil {
t.Errorf("median %v, want undefined where fewer than half the runs violate",
*summary.MedianStepsToFirstViolation)
}
if summary.ThirdQuartileSteps != nil {
t.Errorf("third quartile %v, want undefined", *summary.ThirdQuartileSteps)
}
if summary.Usable != 400 {
t.Errorf("%d usable runs, want all 400: a censored run is data", summary.Usable)
}
censoredShare := float64(summary.Censored) / float64(summary.Usable)
if math.Abs(censoredShare-quiet.censoredShare()) > 0.04 {
t.Errorf("censored share %.3f, want the planted %.3f", censoredShare, quiet.censoredShare())
}
expected := populationA12(quiet, loud)
pair := result.Pairwise[0]
if math.Abs(pair.A12-expected) > 0.03 {
t.Errorf("a12 %.4f under heavy censoring, want the planted %.4f", pair.A12, expected)
}
if pair.PValue > 0.01 {
t.Errorf("p-value %v, want the effect to survive the censoring", pair.PValue)
}
if result.LogRank.Observed[0] >= result.LogRank.Expected[0] {
t.Errorf("log-rank observed %v against expected %v for the quiet arm, want fewer than expected",
result.LogRank.Observed[0], result.LogRank.Expected[0])
}
}
// An arm at a real budget is mostly one enormous group of runs censored
// together at it. No exact null distribution covers that, so the conditional
// variance carries the p-value, and it is checked against a permutation p-value
// over the pipeline's own two samples.
func TestPlanted_CensoringAtTheBudgetTracksThePermutationPValue(t *testing.T) {
cases := []struct {
sourceSeed int64
quiet, loud float64
seeds int
}{
{33, 0.0008, 0.0025, 60},
{51, 0.0008, 0.0014, 50},
{52, 0.0006, 0.0012, 50},
{53, 0.0010, 0.0016, 50},
}
discriminating := 0
for _, test := range cases {
quiet := plantedModel{Hazard: test.quiet, Budget: 400}
loud := plantedModel{Hazard: test.loud, Budget: 400}
source := rand.New(rand.NewSource(test.sourceSeed))
quietRuns := quiet.drawCampaign(source, test.seeds)
loudRuns := loud.drawCampaign(source, test.seeds)
root := t.TempDir()
quietDirectory := writePlantedCampaign(t, filepath.Join(root, "quiet"), "quiet", quiet.Budget, quietRuns)
loudDirectory := writePlantedCampaign(t, filepath.Join(root, "loud"), "loud", loud.Budget, loudRuns)
pair := analyseCampaigns(t, loudDirectory, quietDirectory).Pairwise[0]
permuted := permutationGehanTwoSided(plantedObservations(loudRuns), plantedObservations(quietRuns))
if permuted > 0.01 {
discriminating++
}
// The conditional variance and the permutation variance are two
// estimators and not one, so they are near rather than equal. Over these
// four samples the gap is at most 0.011, at a p-value of 0.49 where
// nothing is decided; where a decision is made it is under 0.005. A
// variance wrong by a factor moves the p-value across orders of
// magnitude, which is what this catches.
if math.Abs(pair.PValue-permuted) > 0.015 {
t.Errorf("hazards %v against %v: p-value %.4f, want near the permutation p-value %.4f",
test.quiet, test.loud, pair.PValue, permuted)
}
}
// A permutation p-value already at the floor cannot show a variance moving
// by a fraction, so the check has to include cases that are not overwhelming.
if discriminating < 3 {
t.Fatalf("%d of %d cases carry a p-value large enough to discriminate", discriminating, len(cases))
}
}
func plantedObservations(runs []plantedRun) []observation {
items := make([]observation, 0, len(runs))
for _, run := range runs {
items = append(items, observation{Steps: float64(run.steps), Event: run.violated})
}
return items
}
// permutationGehanTwoSided is the randomization p-value of Gehan's statistic:
// relabel the pooled runs many times, censoring and all, and count how often the
// statistic lands at least as far from zero as the observed one. Both arms here
// censor on the same schedule, which is what makes relabelling them a null.
func permutationGehanTwoSided(first, second []observation) float64 {
pooled := append(append([]observation{}, first...), second...)
observed, _ := gehanReference(first, second)
source := rand.New(rand.NewSource(99))
const shuffles = 20000
extreme := 0
for shuffle := 0; shuffle < shuffles; shuffle++ {
source.Shuffle(len(pooled), func(left, right int) {
pooled[left], pooled[right] = pooled[right], pooled[left]
})
statistic, _ := gehanReference(pooled[:len(first)], pooled[len(first):])
if math.Abs(statistic) >= math.Abs(observed)-1e-9 {
extreme++
}
}
return float64(extreme) / shuffles
}
// Holm has to change an answer somewhere, or its presence in the pipeline is
// decoration. Six arms drawn from hazards close together produce a family where
// uncorrected p-values call several comparisons significant and the correction
// withdraws some of them.
func TestPlanted_HolmWithdrawsAConclusionUncorrectedPValuesReach(t *testing.T) {
budget := 400
hazards := []float64{0.0100, 0.0125, 0.0150, 0.0175, 0.0200, 0.0225}
root := t.TempDir()
source := rand.New(rand.NewSource(4242))
var directories []string
for index, hazard := range hazards {
model := plantedModel{Hazard: hazard, Budget: budget}
name := fmt.Sprintf("arm%d", index)
directories = append(directories, writePlantedCampaign(t,
filepath.Join(root, name), name, budget, model.drawCampaign(source, 40)))
}
result := analyseCampaigns(t, directories...)
if len(result.Pairwise) != 15 {
t.Fatalf("%d comparisons, want 15 over six arms", len(result.Pairwise))
}
if result.HolmFamilySize != 15 {
t.Errorf("holm family size %d, want 15", result.HolmFamilySize)
}
uncorrected, corrected, withdrawn := 0, 0, 0
for _, pair := range result.Pairwise {
if pair.PValue < 0.05 {
uncorrected++
}
if pair.HolmPValue < 0.05 {
corrected++
}
if pair.PValue < 0.05 && pair.HolmPValue >= 0.05 {
withdrawn++
}
}
if withdrawn == 0 {
t.Fatalf("holm withdrew no conclusion: %d of 15 comparisons significant before and after", uncorrected)
}
if corrected >= uncorrected {
t.Errorf("%d comparisons significant uncorrected and %d after holm, want fewer", uncorrected, corrected)
}
}
// Holm is applied within one research question and not across the paper, so the
// same two campaigns analysed on their own must keep their raw p-value while the
// same comparison inside a larger family is corrected against that family.
func TestPlanted_HolmFamilyIsTheInvocationAndNotEveryComparisonEverMade(t *testing.T) {
budget := 400
root := t.TempDir()
source := rand.New(rand.NewSource(777))
var directories []string
for index, hazard := range []float64{0.010, 0.014, 0.018, 0.022} {
model := plantedModel{Hazard: hazard, Budget: budget}
name := fmt.Sprintf("arm%d", index)
directories = append(directories, writePlantedCampaign(t,
filepath.Join(root, name), name, budget, model.drawCampaign(source, 40)))
}
whole := analyseCampaigns(t, append([]string{"--question", "RQ4"}, directories...)...)
alone := analyseCampaigns(t, "--question", "RQ4a", directories[0], directories[1])
if whole.Question != "RQ4" || alone.Question != "RQ4a" {
t.Errorf("questions %q and %q, want them recorded next to the corrected values", whole.Question, alone.Question)
}
if whole.HolmFamilySize != 6 || alone.HolmFamilySize != 1 {
t.Fatalf("family sizes %d and %d, want 6 and 1", whole.HolmFamilySize, alone.HolmFamilySize)
}
if alone.Pairwise[0].HolmPValue != alone.Pairwise[0].PValue {
t.Errorf("holm p %v against raw p %v in a family of one, want them equal",
alone.Pairwise[0].HolmPValue, alone.Pairwise[0].PValue)
}
var inFamily pairwiseResult
for _, pair := range whole.Pairwise {
if pair.First == alone.Pairwise[0].First && pair.Second == alone.Pairwise[0].Second {
inFamily = pair
}
}
if math.Abs(inFamily.PValue-alone.Pairwise[0].PValue) > 1e-12 {
t.Fatalf("the same comparison has raw p %v in one family and %v in the other",
inFamily.PValue, alone.Pairwise[0].PValue)
}
if inFamily.HolmPValue <= alone.Pairwise[0].HolmPValue {
t.Errorf("holm p %v inside a family of six, want it above the %v it carries alone",
inFamily.HolmPValue, alone.Pairwise[0].HolmPValue)
}
}
// The ablation's outcome is a per-seed difference with a sign, so the planted
// shift is applied seed by seed and the pipeline has to return that shift with
// that sign from the campaign directories alone.
func TestPlanted_PairedComparisonRecoversTheShiftAndItsSign(t *testing.T) {
budget := 400
base := plantedModel{Hazard: 0.01, Budget: budget}
const shift = 60
source := rand.New(rand.NewSource(5150))
var post, pre []plantedRun
// Arms are contrasted in the order their labels sort, so the plant is stated
// the same way: post-repair against pre-repair. A seed where the unshifted
// run violated is a pair the shifted run is known to have outlived, whether
// it violated later or ran on clean; a seed where neither violated is a pair
// with no order.
firstSooner, bothViolated, unordered := 0, 0, 0
for seed := int64(1); seed <= 30; seed++ {
fast := base.draw(seed, source)
slow := plantedRun{seed: seed, steps: fast.steps + shift, violated: fast.violated}
if !fast.violated || slow.steps > budget {
slow = plantedRun{seed: seed, steps: budget}
}
post = append(post, fast)
pre = append(pre, slow)
switch {
case !fast.violated:
unordered++
case slow.violated:
bothViolated++
firstSooner++
default:
firstSooner++
}
}
if bothViolated == 0 {
t.Fatal("no pair has two violations, so the planted shift is nowhere the analysis can read it")
}
root := t.TempDir()
preDirectory := writePlantedCampaign(t, filepath.Join(root, "pre"), "pre-repair", budget, pre)
postDirectory := writePlantedCampaign(t, filepath.Join(root, "post"), "post-repair", budget, post)
result := analyseCampaigns(t, "--paired", "--question", "RQ4 ablation", preDirectory, postDirectory)
if result.Paired == nil {
t.Fatal("no paired comparison, want one")
}
if len(result.Pairwise) != 0 {
t.Errorf("%d unpaired comparisons alongside the paired one, want none in a seed-matched design",
len(result.Pairwise))
}
paired := *result.Paired
if paired.First != "post-repair" || paired.Second != "pre-repair" {
t.Fatalf("paired %s minus %s, want post-repair minus pre-repair", paired.First, paired.Second)
}
if paired.Pairs != 30 {
t.Errorf("%d pairs, want 30", paired.Pairs)
}
// Every pair where both runs violated was shifted by exactly the plant, so
// the median over them is the plant itself rather than a mixture of it with
// the step counts censored runs never reached.
if paired.BothViolated != bothViolated {
t.Errorf("%d pair(s) with two violations, want %d", paired.BothViolated, bothViolated)
}
if paired.MedianDifference == nil || *paired.MedianDifference != -shift {
t.Errorf("median difference %v, want the planted %d", paired.MedianDifference, -shift)
}
if paired.Sign != -1 {
t.Errorf("sign %+d, want -1 for the arm that violated sooner", paired.Sign)
}
if paired.FirstSooner != firstSooner || paired.SecondSooner != 0 || paired.Unordered != unordered {
t.Errorf("counts %+v, want %d favouring the shifted arm, none the other way and %d unordered",
paired, firstSooner, unordered)
}
if want := 0.5 * float64(unordered) / 30; paired.A12 != want {
t.Errorf("a12 within pairs %v, want %v where no pair favours the first arm", paired.A12, want)
}
if paired.PValue == nil || *paired.PValue > 0.001 {
t.Errorf("p-value %v for a shift planted in every pair", paired.PValue)
}
if *paired.HolmPValue != *paired.PValue {
t.Errorf("holm p %v in a family of one, want the raw %v", *paired.HolmPValue, *paired.PValue)
}
}
// The paired test is the ablation's decision rule, so a null there has to stay a
// null: two arms drawn from the same model, matched by seed, must not report a
// difference more often than the level allows.
func TestPlanted_PairedNullIsNotCalledSignificantAboveItsLevel(t *testing.T) {
model := plantedModel{Hazard: 0.01, Budget: 400}
const replicates = 200
rejected := 0
for replicate := 0; replicate < replicates; replicate++ {
first, second := plantTwoArms(t, int64(5000+replicate), 30, model, model)
result := analyseCampaigns(t, "--paired", first, second)
if *result.Paired.PValue < 0.05 {
rejected++
}
}
if rate := float64(rejected) / replicates; rate > 0.10 {
t.Errorf("the paired test called a true null significant in %.1f%% of %d replicates, want about 5%%",
100*rate, replicates)
}
}
+17
View File
@@ -0,0 +1,17 @@
package main
// quantileSurvival is the smallest step count at which the product-limit
// estimate falls to or below 1-fraction, which is the fraction-th quantile of
// steps to first violation. It is undefined whenever the curve never falls that
// far, which is what an arm where most runs exhaust the budget produces, and the
// second return value says so rather than substituting a number the data does
// not contain.
func quantileSurvival(curve []survivalPoint, fraction float64) (float64, bool) {
threshold := 1 - fraction
for _, point := range curve {
if point.Survival <= threshold {
return point.Steps, true
}
}
return 0, false
}
+225
View File
@@ -0,0 +1,225 @@
package main
import (
"math"
"slices"
)
type rankSumResult struct {
FirstSize int `json:"first_size"`
SecondSize int `json:"second_size"`
Statistic float64 `json:"mann_whitney_u"`
A12 float64 `json:"a12"`
PValue float64 `json:"p_value"`
Exact bool `json:"exact"`
}
// exactRankSumLimit matches R's wilcox.test: the exact null distribution is
// used only when both samples are below this size and nothing is tied.
const exactRankSumLimit = 50
// vargaDelaneyA12 is the probability that a value drawn from first exceeds one
// drawn from second, counting a tie as half:
//
// A = P(X > Y) + 0.5 * P(X = Y)
//
// Vargha and Delaney (2000), "A Critique and Improvement of the CL Common
// Language Effect Size Statistics of McGraw and Wong", Journal of Educational
// and Behavioral Statistics 25(2), 101-132.
func vargaDelaneyA12(first, second []float64) float64 {
if len(first) == 0 || len(second) == 0 {
return math.NaN()
}
total := 0.0
for _, left := range first {
for _, right := range second {
switch {
case left > right:
total++
case left == right:
total += 0.5
}
}
}
return total / float64(len(first)*len(second))
}
// rankSum is the two-sided Wilcoxon rank-sum (Mann-Whitney) test. The reported
// statistic is U for the first sample, the same quantity R's wilcox.test calls
// W. The exact null distribution is used when there are no ties and both
// samples are small; otherwise the normal approximation is used with the
// continuity correction and the tie correction to the variance.
//
// It takes plain numbers, so it cannot be given campaign runs, where a clean one
// carries a bound and not a value. What it is here for is the uncensored case
// the pipeline's comparison has to reproduce: this implementation is checked
// against R on a published sample, and gehanTest is checked against this one.
func rankSum(first, second []float64) rankSumResult {
firstSize, secondSize := len(first), len(second)
result := rankSumResult{
FirstSize: firstSize,
SecondSize: secondSize,
Statistic: math.NaN(),
A12: math.NaN(),
PValue: math.NaN(),
}
if firstSize == 0 || secondSize == 0 {
return result
}
pooled := make([]float64, 0, firstSize+secondSize)
pooled = append(pooled, first...)
pooled = append(pooled, second...)
ranks, tieGroups := midRanks(pooled)
rankTotal := 0.0
for index := 0; index < firstSize; index++ {
rankTotal += ranks[index]
}
statistic := rankTotal - float64(firstSize)*float64(firstSize+1)/2
result.Statistic = statistic
result.A12 = vargaDelaneyA12(first, second)
if len(tieGroups) == 0 && firstSize < exactRankSumLimit && secondSize < exactRankSumLimit {
result.Exact = true
result.PValue = exactRankSumTwoSided(statistic, firstSize, secondSize)
return result
}
result.PValue = normalRankSumTwoSided(statistic, firstSize, secondSize, tieGroups)
return result
}
// normalRankSumTwoSided follows the large-sample branch of R's wilcox.test:
//
// sigma^2 = (m*n/12) * ((N+1) - sum(t^3 - t) / (N*(N-1)))
//
// where t runs over the sizes of the tied groups. The 0.5 shift toward the null
// mean is the continuity correction.
func normalRankSumTwoSided(statistic float64, firstSize, secondSize int, tieGroups []int) float64 {
sizeProduct := float64(firstSize) * float64(secondSize)
variance := rankSumVariance(firstSize, secondSize, tieGroups)
if variance <= 0 {
return 1
}
centered := statistic - sizeProduct/2
correction := 0.0
switch {
case centered > 0:
correction = 0.5
case centered < 0:
correction = -0.5
}
z := (centered - correction) / math.Sqrt(variance)
tail := math.Min(standardNormalUpperTail(z), standardNormalUpperTail(-z))
return math.Min(2*tail, 1)
}
func rankSumVariance(firstSize, secondSize int, tieGroups []int) float64 {
sizeProduct := float64(firstSize) * float64(secondSize)
total := float64(firstSize + secondSize)
tieAdjustment := 0.0
for _, size := range tieGroups {
count := float64(size)
tieAdjustment += count*count*count - count
}
return (sizeProduct / 12) * ((total + 1) - tieAdjustment/(total*(total-1)))
}
// exactRankSumTwoSided doubles the smaller exact tail, as R's wilcox.test does.
func exactRankSumTwoSided(statistic float64, firstSize, secondSize int) float64 {
counts := exactRankSumCounts(firstSize, secondSize)
total := 0.0
for _, count := range counts {
total += count
}
value := int(math.Round(statistic))
tail := 0.0
if statistic > float64(firstSize*secondSize)/2 {
for u := value; u < len(counts); u++ {
tail += counts[u]
}
} else {
for u := 0; u <= value && u < len(counts); u++ {
tail += counts[u]
}
}
return math.Min(2*tail/total, 1)
}
// exactRankSumUpperTail is P(U >= statistic) under the null with no ties.
func exactRankSumUpperTail(statistic float64, firstSize, secondSize int) float64 {
counts := exactRankSumCounts(firstSize, secondSize)
total, tail := 0.0, 0.0
for u, count := range counts {
total += count
if float64(u) >= statistic {
tail += count
}
}
return tail / total
}
// exactRankSumCounts returns the number of untied assignments producing each
// value of U from 0 to firstSize*secondSize. U equals the sum of the zero-based
// pooled positions held by the first sample, less firstSize*(firstSize-1)/2, so
// the count is a subset-sum tally over those positions.
func exactRankSumCounts(firstSize, secondSize int) []float64 {
maximum := firstSize * secondSize
offset := firstSize * (firstSize - 1) / 2
high := maximum + offset
table := make([][]float64, firstSize+1)
for index := range table {
table[index] = make([]float64, high+1)
}
table[0][0] = 1
for position := 0; position < firstSize+secondSize; position++ {
for chosen := min(position+1, firstSize); chosen >= 1; chosen-- {
row, previous := table[chosen], table[chosen-1]
for sum := high; sum >= position; sum-- {
if previous[sum-position] != 0 {
row[sum] += previous[sum-position]
}
}
}
}
counts := make([]float64, maximum+1)
for u := range counts {
counts[u] = table[firstSize][u+offset]
}
return counts
}
// midRanks ranks values from 1, averaging the ranks within a tied group, and
// also returns the size of every group of size two or more.
func midRanks(values []float64) ([]float64, []int) {
order := make([]int, len(values))
for index := range order {
order[index] = index
}
slices.SortStableFunc(order, func(left, right int) int {
switch {
case values[left] < values[right]:
return -1
case values[left] > values[right]:
return 1
default:
return 0
}
})
ranks := make([]float64, len(values))
var tieGroups []int
for start := 0; start < len(order); {
end := start + 1
for end < len(order) && values[order[end]] == values[order[start]] {
end++
}
shared := float64(start+1+end) / 2
for index := start; index < end; index++ {
ranks[order[index]] = shared
}
if end-start > 1 {
tieGroups = append(tieGroups, end-start)
}
start = end
}
return ranks, tieGroups
}
+211
View File
@@ -0,0 +1,211 @@
package main
import (
"math"
"testing"
)
// R's wilcox.test on the Hollander and Wolfe (1973), 69f chorioamnion data
// reports W = 35 with an exact two-sided p-value of 0.2544; the one-sided
// greater alternative that the help page uses reports the same W with
// p-value = 0.1272.
func TestRankSum_MatchesPublishedChorioamnionResult(t *testing.T) {
result := rankSum(chorioamnionTerm, chorioamnionEarly)
if result.Statistic != 35 {
t.Errorf("statistic %v, want 35", result.Statistic)
}
if !result.Exact {
t.Error("expected the exact null distribution for untied samples this small")
}
if math.Abs(result.PValue-0.2544) > 5e-5 {
t.Errorf("two-sided p-value %.6f, want 0.2544", result.PValue)
}
upper := exactRankSumUpperTail(35, len(chorioamnionTerm), len(chorioamnionEarly))
if math.Abs(upper-0.1272) > 5e-5 {
t.Errorf("one-sided p-value %.6f, want 0.1272", upper)
}
}
// A12 is P(X > Y) + 0.5 P(X = Y), which is U/(mn). With the published W = 35
// and sample sizes 10 and 5 the effect size is 35/50 = 0.70. The expected value
// is therefore the published Mann-Whitney statistic combined with the published
// definition in Vargha and Delaney (2000), not a number this tool produced.
func TestVargaDelaneyA12_MatchesPublishedChorioamnionStatistic(t *testing.T) {
got := vargaDelaneyA12(chorioamnionTerm, chorioamnionEarly)
if math.Abs(got-0.70) > 1e-12 {
t.Errorf("A12 = %v, want 0.70", got)
}
if reversed := vargaDelaneyA12(chorioamnionEarly, chorioamnionTerm); math.Abs(reversed-0.30) > 1e-12 {
t.Errorf("reversed A12 = %v, want 0.30", reversed)
}
}
// Identical samples are stochastically equal, which Vargha and Delaney define
// as A = 0.5, and complete separation gives 1 and 0.
func TestVargaDelaneyA12_BoundaryCases(t *testing.T) {
same := []float64{1, 2, 3, 4}
if got := vargaDelaneyA12(same, same); got != 0.5 {
t.Errorf("A12 of a sample against itself = %v, want 0.5", got)
}
if got := vargaDelaneyA12([]float64{5, 6, 7}, []float64{1, 2}); got != 1 {
t.Errorf("A12 with complete dominance = %v, want 1", got)
}
if got := vargaDelaneyA12([]float64{1, 2}, []float64{5, 6, 7}); got != 0 {
t.Errorf("A12 with complete subordination = %v, want 0", got)
}
}
// The counting definition and the rank-sum route must agree, including when the
// samples are tied against each other, which is the case the evaluation data is
// usually in because censored runs pile up on the step they stopped at.
func TestVargaDelaneyA12_AgreesWithRankSumStatistic(t *testing.T) {
cases := [][2][]float64{
{{1, 2, 3}, {2, 3, 4}},
{{40, 40, 40, 12}, {40, 7, 3}},
{{5}, {5, 5, 5}},
{{9, 9, 9}, {9, 9, 9}},
}
for _, test := range cases {
result := rankSum(test[0], test[1])
expected := result.Statistic / float64(len(test[0])*len(test[1]))
if math.Abs(result.A12-expected) > 1e-12 {
t.Errorf("A12 %v for %v vs %v, want U/(mn) = %v", result.A12, test[0], test[1], expected)
}
}
}
// The tie-corrected variance is checked against the exact permutation variance
// of the statistic, computed here by enumerating every way to split the pooled
// midranks. That is an independent calculation, not a second call into the
// implementation under test.
func TestRankSumVariance_MatchesExactPermutationVariance(t *testing.T) {
cases := [][]float64{
{1, 2, 3, 4, 5, 6, 7, 8},
{40, 40, 40, 40, 12, 7, 3, 3},
{5, 5, 5, 5, 5, 5, 5, 9},
{2, 2, 3, 3, 3, 4, 9, 9, 9},
}
for _, pooled := range cases {
firstSize := len(pooled) / 2
ranks, tieGroups := midRanks(pooled)
mean, variance := permutationMomentsOfRankSum(ranks, firstSize)
expectedMean := float64(firstSize*(len(pooled)-firstSize)) / 2
if math.Abs(mean-expectedMean) > 1e-9 {
t.Errorf("%v: permutation mean %v, want %v", pooled, mean, expectedMean)
}
got := rankSumVariance(firstSize, len(pooled)-firstSize, tieGroups)
if math.Abs(got-variance) > 1e-9 {
t.Errorf("%v: tie-corrected variance %v, want the permutation variance %v", pooled, got, variance)
}
}
}
// permutationMomentsOfRankSum enumerates every subset of the given size and
// returns the mean and variance of the Mann-Whitney statistic over them.
func permutationMomentsOfRankSum(ranks []float64, firstSize int) (float64, float64) {
offset := float64(firstSize) * float64(firstSize+1) / 2
var values []float64
chosen := make([]int, 0, firstSize)
var walk func(start int)
walk = func(start int) {
if len(chosen) == firstSize {
total := 0.0
for _, index := range chosen {
total += ranks[index]
}
values = append(values, total-offset)
return
}
for index := start; index < len(ranks); index++ {
chosen = append(chosen, index)
walk(index + 1)
chosen = chosen[:len(chosen)-1]
}
}
walk(0)
mean := 0.0
for _, value := range values {
mean += value
}
mean /= float64(len(values))
variance := 0.0
for _, value := range values {
variance += (value - mean) * (value - mean)
}
return mean, variance / float64(len(values))
}
func TestMidRanks_AveragesTiedGroups(t *testing.T) {
ranks, tieGroups := midRanks([]float64{3, 1, 3, 2, 3})
expected := []float64{4, 1, 4, 2, 4}
for index, want := range expected {
if ranks[index] != want {
t.Errorf("rank %d = %v, want %v", index, ranks[index], want)
}
}
if len(tieGroups) != 1 || tieGroups[0] != 3 {
t.Errorf("tie groups %v, want [3]", tieGroups)
}
}
func TestRankSum_TiedSamplesUseTheNormalApproximation(t *testing.T) {
result := rankSum([]float64{1, 2, 3, 4}, []float64{3, 4, 5, 6})
if result.Exact {
t.Error("used the exact null distribution despite ties")
}
if math.IsNaN(result.PValue) || result.PValue < 0 || result.PValue > 1 {
t.Errorf("p-value %v", result.PValue)
}
}
// Every observation identical carries no information, and the test must say so
// rather than dividing by a zero variance.
func TestRankSum_AllValuesIdentical(t *testing.T) {
result := rankSum([]float64{40, 40, 40}, []float64{40, 40, 40, 40})
if result.PValue != 1 {
t.Errorf("p-value %v, want 1", result.PValue)
}
if result.A12 != 0.5 {
t.Errorf("A12 %v, want 0.5", result.A12)
}
}
func TestRankSum_SingleObservationPerSample(t *testing.T) {
result := rankSum([]float64{3}, []float64{9})
if result.Statistic != 0 {
t.Errorf("statistic %v, want 0", result.Statistic)
}
if result.A12 != 0 {
t.Errorf("A12 %v, want 0", result.A12)
}
if math.IsNaN(result.PValue) || result.PValue > 1 {
t.Errorf("p-value %v", result.PValue)
}
}
func TestRankSum_EmptySampleHasNoStatistic(t *testing.T) {
result := rankSum(nil, []float64{1, 2, 3})
if !math.IsNaN(result.PValue) || !math.IsNaN(result.A12) {
t.Errorf("result %+v, want everything undefined", result)
}
}
// The exact null distribution must be a proper distribution: the counts sum to
// the binomial coefficient and the distribution is symmetric about mn/2.
func TestExactRankSumCounts_FormAProperSymmetricDistribution(t *testing.T) {
counts := exactRankSumCounts(4, 6)
total := 0.0
for _, count := range counts {
total += count
}
if total != 210 {
t.Errorf("counts sum to %v, want C(10,4) = 210", total)
}
for index := range counts {
mirrored := counts[len(counts)-1-index]
if counts[index] != mirrored {
t.Errorf("count at %d is %v but %v at the mirrored point", index, counts[index], mirrored)
}
}
}
+228
View File
@@ -0,0 +1,228 @@
package main
import (
"fmt"
"io"
"maps"
"math"
"slices"
"strconv"
"strings"
"text/tabwriter"
)
func writeReport(result analysis, out io.Writer) {
fmt.Fprintf(out, "primary outcome: %s\n\n", result.Outcome)
writeTable(out, []string{"arm", "runs", "violated", "censored", "excluded", "missing", "median steps", "iqr steps", "violation rate"},
func(add func(...string)) {
for _, summary := range result.Arms {
add(
summary.Arm,
strconv.Itoa(summary.Usable),
strconv.Itoa(summary.Violated),
strconv.Itoa(summary.Censored),
strconv.Itoa(summary.Excluded),
strconv.Itoa(len(summary.MissingSeeds)),
formatMedian(summary.MedianStepsToFirstViolation),
formatMedian(summary.FirstQuartileSteps)+" to "+formatMedian(summary.ThirdQuartileSteps),
formatRatio(summary.ViolationRate, 3),
)
}
})
fmt.Fprintln(out)
fmt.Fprintln(out, "a detection is one distinct property violated in one run; run hours sum the time the runs worked,")
fmt.Fprintln(out, "on the monotonic clock, so a host that slept mid-run is not charged for the sleep")
fmt.Fprintln(out, "actions count the steps the generator dispatched one on; the rest chose nothing, had the choice")
fmt.Fprintln(out, "thrown away, or were the spec's setup driving the app into position before the generator ran")
writeTable(out, []string{"arm", "steps", "actions", "run hours", "detections", "defects/1k actions", "defects/hour", "distinct defects", "found in one run"},
func(add func(...string)) {
for _, summary := range result.Arms {
add(
summary.Arm,
strconv.Itoa(summary.TotalSteps),
formatActions(summary),
fmt.Sprintf("%.2f", summary.TotalRunHours),
strconv.Itoa(summary.Detections),
formatRatio(summary.DefectsPerThousandActions, 2),
formatRatio(summary.DefectsPerHour, 2),
strconv.Itoa(summary.DistinctDefects),
formatSingletons(summary),
)
}
})
for _, summary := range result.Arms {
if len(summary.ExcludedByReason) == 0 {
continue
}
var parts []string
for _, reason := range sortedKeys(summary.ExcludedByReason) {
parts = append(parts, fmt.Sprintf("%s=%d", reason, summary.ExcludedByReason[reason]))
}
fmt.Fprintf(out, "\n%s excluded %d run(s) as missing data, not as censored observations: %s",
summary.Arm, summary.Excluded, strings.Join(parts, ", "))
}
for _, summary := range result.Arms {
if summary.UnattributedActions > 0 {
fmt.Fprintf(out, "\n%s counts %d action(s) of unknown provenance, recorded before an action named its producer: "+
"its per-action denominator may include the login the spec's setup drove",
summary.Arm, summary.UnattributedActions)
}
}
for _, summary := range result.Arms {
if summary.EventsHeldAtBudget > 0 {
fmt.Fprintf(out, "\n%s held %d violation(s) reported past the budget at %d steps",
summary.Arm, summary.EventsHeldAtBudget, summary.StepBudget)
}
}
for _, summary := range result.Arms {
if summary.EventsDetectedAfterOrigin > 0 {
fmt.Fprintf(out, "\n%s timed %d violation(s) at the step they were detected rather than the step that armed them, "+
"which is what an obligation reported only when the run ended looks like",
summary.Arm, summary.EventsDetectedAfterOrigin)
}
}
if len(result.Arms) > 0 {
fmt.Fprintln(out)
}
if result.LogRank != nil {
fmt.Fprintf(out, "\nlog-rank across %d arms: chi-square %.4f on %d df, p %s\n",
len(result.LogRank.Groups), result.LogRank.ChiSquare, result.LogRank.DegreesOfFreedom,
formatPValue(result.LogRank.PValue))
writeTable(out, []string{"arm", "n", "observed", "expected"}, func(add func(...string)) {
for index, name := range result.LogRank.Groups {
add(name,
strconv.Itoa(result.LogRank.Sizes[index]),
fmt.Sprintf("%.0f", result.LogRank.Observed[index]),
fmt.Sprintf("%.2f", result.LogRank.Expected[index]))
}
})
}
if result.Paired != nil {
writePaired(out, *result.Paired)
}
if len(result.Pairwise) > 0 {
fmt.Fprintln(out, "\npairwise gehan generalized wilcoxon, which reads a censored run as the bound it is")
fmt.Fprintln(out, "a12 above 0.5 means the first arm takes more steps to its first violation; u counts the run")
fmt.Fprintln(out, "pairs the first arm outlived, and the unordered pairs count as half in both u and a12")
writeTable(out, []string{"comparison", "n1", "n2", "u", "a12", "unordered", "p", "holm p"}, func(add func(...string)) {
for _, pair := range result.Pairwise {
add(
pair.First+" vs "+pair.Second,
strconv.Itoa(pair.FirstSize),
strconv.Itoa(pair.SecondSize),
fmt.Sprintf("%.1f", pair.Statistic),
fmt.Sprintf("%.3f", pair.A12),
fmt.Sprintf("%d of %d", pair.Unordered, pair.FirstSize*pair.SecondSize),
formatPValue(pair.PValue),
formatPValue(pair.HolmPValue),
)
}
})
}
if result.HolmFamilySize > 0 {
family := "this invocation"
if result.Question != "" {
family = result.Question
}
fmt.Fprintf(out, "\nholm correction applied within %s, over %d comparison(s)\n", family, result.HolmFamilySize)
}
for _, note := range result.Notes {
fmt.Fprintf(out, "\nnote: %s\n", note)
}
}
func writePaired(out io.Writer, comparison pairedComparison) {
fmt.Fprintf(out, "\npaired per-seed contrast, %s against %s, each pair scored by which run outlived the other\n",
comparison.First, comparison.Second)
fmt.Fprintf(out, "%d seed pair(s): %s sooner in %d, %s sooner in %d, left in no order by censoring in %d\n",
comparison.Pairs, comparison.First, comparison.FirstSooner,
comparison.Second, comparison.SecondSooner, comparison.Unordered)
fmt.Fprintf(out, "median difference %s over the %d pair(s) where both runs violated, sign %+d, a12 within pairs %.3f\n",
formatStepDifference(comparison.MedianDifference), comparison.BothViolated, comparison.Sign, comparison.A12)
fmt.Fprintf(out, "sign test over the %d ordered pair(s), p %s, holm p %s\n",
comparison.FirstSooner+comparison.SecondSooner,
formatOptionalPValue(comparison.PValue), formatOptionalPValue(comparison.HolmPValue))
if len(comparison.UnpairedSeeds) > 0 {
fmt.Fprintf(out, "%d seed(s) usable in one arm only and left out of the pairing: %v\n",
len(comparison.UnpairedSeeds), comparison.UnpairedSeeds)
}
}
func sortedKeys(counts map[string]int) []string {
return slices.Sorted(maps.Keys(counts))
}
func writeTable(out io.Writer, header []string, rows func(add func(...string))) {
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
fmt.Fprintln(writer, strings.Join(header, "\t"))
rows(func(cells ...string) {
fmt.Fprintln(writer, strings.Join(cells, "\t"))
})
writer.Flush()
}
// formatMedian says undefined rather than substituting a mean, because a curve
// that never reaches one half has no median to report.
func formatMedian(value *float64) string {
if value == nil {
return "undefined"
}
return strconv.FormatFloat(*value, 'f', -1, 64)
}
func formatStepDifference(value *float64) string {
if value == nil {
return "undefined"
}
return fmt.Sprintf("%+.1f steps", *value)
}
// formatActions marks a denominator with actions whose producer nothing names,
// because the rate beside it then divides by a count that may include the
// login the spec's setup drove.
func formatActions(summary armSummary) string {
if summary.UnattributedActions == 0 {
return strconv.Itoa(summary.TotalActions)
}
return fmt.Sprintf("%d (%d unattributed)", summary.TotalActions, summary.UnattributedActions)
}
func formatRatio(value *float64, digits int) string {
if value == nil {
return "n/a"
}
return strconv.FormatFloat(*value, 'f', digits, 64)
}
func formatSingletons(summary armSummary) string {
if summary.SingletonFraction == nil {
return "n/a"
}
return fmt.Sprintf("%d/%d (%.3f)", summary.SingletonDefects, summary.DistinctDefects, *summary.SingletonFraction)
}
func formatOptionalPValue(value *float64) string {
if value == nil {
return "n/a"
}
return formatPValue(*value)
}
func formatPValue(value float64) string {
switch {
case math.IsNaN(value):
return "n/a"
case value < 1e-4:
return fmt.Sprintf("%.3e", value)
default:
return fmt.Sprintf("%.4f", value)
}
}
+247
View File
@@ -0,0 +1,247 @@
package main
import (
"math"
"slices"
)
// observation is one run reduced to what the survival analysis needs: the step
// count at which it left the risk set, and whether it left because a violation
// was found (an event) or because the run ended without one (right-censored). A run that failed or timed out is neither and never
// reaches this type.
type observation struct {
Steps float64
Event bool
}
type survivalPoint struct {
Steps float64 `json:"steps"`
AtRisk int `json:"at_risk"`
Events int `json:"events"`
Censored int `json:"censored"`
Survival float64 `json:"survival"`
}
// kaplanMeier is the product-limit estimate of Kaplan and Meier (1958), one row
// per distinct observed step count. Runs censored at a step count tied with an
// event are counted in the risk set for that event, which is the standard
// convention.
func kaplanMeier(observations []observation) []survivalPoint {
if len(observations) == 0 {
return nil
}
remaining := len(observations)
survival := 1.0
var curve []survivalPoint
for _, steps := range distinctSteps(observations) {
events, censored := 0, 0
for _, item := range observations {
if item.Steps != steps {
continue
}
if item.Event {
events++
} else {
censored++
}
}
atRisk := remaining
if events > 0 {
survival *= 1 - float64(events)/float64(atRisk)
}
curve = append(curve, survivalPoint{
Steps: steps,
AtRisk: atRisk,
Events: events,
Censored: censored,
Survival: survival,
})
remaining -= events + censored
}
return curve
}
// medianSurvival is the smallest step count at which the estimate falls to or
// below one half. It is undefined whenever fewer than half the runs violate,
// and the second return value says so: substituting a mean there would report a
// number the data does not contain.
func medianSurvival(curve []survivalPoint) (float64, bool) {
return quantileSurvival(curve, 0.5)
}
func distinctSteps(observations []observation) []float64 {
steps := make([]float64, 0, len(observations))
for _, item := range observations {
steps = append(steps, item.Steps)
}
slices.Sort(steps)
return slices.Compact(steps)
}
type logRankResult struct {
Groups []string `json:"groups"`
Sizes []int `json:"sizes"`
Observed []float64 `json:"observed"`
Expected []float64 `json:"expected"`
ChiSquare float64 `json:"chi_square"`
DegreesOfFreedom int `json:"degrees_of_freedom"`
PValue float64 `json:"p_value"`
}
// logRank is the k-sample Mantel-Haenszel log-rank test, the member of the
// weighted family that counts every event time alike. Mantel (1966); Peto and
// Peto (1972).
func logRank(names []string, groups [][]observation) logRankResult {
return weightedLogRank(names, groups, func(atRisk float64) float64 { return 1 })
}
// weightedLogRank is the family the log-rank belongs to. At every distinct event
// time it contrasts observed with expected events under the null of equal
// hazards, weights that difference by weight(atRisk), and combines the k-1
// independent weighted differences through their covariance matrix:
// chi-square = U' V^-1 U on k-1 degrees of freedom. Observed and Expected stay
// event counts whatever the weight, because a weighted count is not one.
// Klein and Moeschberger, Survival Analysis, 2nd ed., section 7.3.
func weightedLogRank(names []string, groups [][]observation, weight func(atRisk float64) float64) logRankResult {
var keptNames []string
var kept [][]observation
for index, group := range groups {
if len(group) == 0 {
continue
}
keptNames = append(keptNames, names[index])
kept = append(kept, group)
}
names, groups = keptNames, kept
count := len(groups)
result := logRankResult{
Groups: names,
Sizes: make([]int, count),
Observed: make([]float64, count),
Expected: make([]float64, count),
DegreesOfFreedom: count - 1,
PValue: math.NaN(),
}
if count < 2 {
return result
}
var pooled []observation
for index, group := range groups {
result.Sizes[index] = len(group)
pooled = append(pooled, group...)
}
covariance := make([][]float64, count)
for index := range covariance {
covariance[index] = make([]float64, count)
}
atRisk := make([]float64, count)
deaths := make([]float64, count)
weightedDifference := make([]float64, count)
for _, steps := range distinctSteps(pooled) {
totalAtRisk, totalDeaths := 0.0, 0.0
for index, group := range groups {
atRisk[index], deaths[index] = 0, 0
for _, item := range group {
if item.Steps >= steps {
atRisk[index]++
}
if item.Steps == steps && item.Event {
deaths[index]++
}
}
totalAtRisk += atRisk[index]
totalDeaths += deaths[index]
}
if totalDeaths == 0 {
continue
}
weightAtStep := weight(totalAtRisk)
for index := range groups {
expected := totalDeaths * atRisk[index] / totalAtRisk
result.Observed[index] += deaths[index]
result.Expected[index] += expected
weightedDifference[index] += weightAtStep * (deaths[index] - expected)
}
if totalAtRisk <= 1 {
continue
}
scale := weightAtStep * weightAtStep * totalDeaths * (totalAtRisk - totalDeaths) / (totalAtRisk - 1)
for row := range groups {
share := atRisk[row] / totalAtRisk
covariance[row][row] += scale * share * (1 - share)
for column := range groups {
if column == row {
continue
}
covariance[row][column] -= scale * share * atRisk[column] / totalAtRisk
}
}
}
reduced := make([][]float64, count-1)
difference := make([]float64, count-1)
for row := 0; row < count-1; row++ {
reduced[row] = make([]float64, count-1)
copy(reduced[row], covariance[row][:count-1])
difference[row] = weightedDifference[row]
}
solution, ok := solveLinearSystem(reduced, difference)
if !ok {
result.ChiSquare = 0
result.PValue = 1
return result
}
statistic := 0.0
for index := range difference {
statistic += difference[index] * solution[index]
}
if statistic < 0 || math.IsNaN(statistic) {
statistic = 0
}
result.ChiSquare = statistic
result.PValue = chiSquareUpperTail(statistic, result.DegreesOfFreedom)
return result
}
// solveLinearSystem solves matrix*x = vector by Gaussian elimination with
// partial pivoting, reporting failure rather than a value when the matrix is
// singular, which is what a group with no events at all produces.
func solveLinearSystem(matrix [][]float64, vector []float64) ([]float64, bool) {
size := len(vector)
work := make([][]float64, size)
for row := range work {
work[row] = make([]float64, size+1)
copy(work[row], matrix[row])
work[row][size] = vector[row]
}
for column := 0; column < size; column++ {
pivot := column
for row := column + 1; row < size; row++ {
if math.Abs(work[row][column]) > math.Abs(work[pivot][column]) {
pivot = row
}
}
if math.Abs(work[pivot][column]) < 1e-12 {
return nil, false
}
work[column], work[pivot] = work[pivot], work[column]
for row := column + 1; row < size; row++ {
factor := work[row][column] / work[column][column]
for next := column; next <= size; next++ {
work[row][next] -= factor * work[column][next]
}
}
}
solution := make([]float64, size)
for row := size - 1; row >= 0; row-- {
total := work[row][size]
for column := row + 1; column < size; column++ {
total -= work[row][column] * solution[column]
}
solution[row] = total / work[row][row]
}
return solution, true
}
+323
View File
@@ -0,0 +1,323 @@
package main
import (
"math"
"math/rand"
"slices"
"strconv"
"testing"
)
// The product-limit estimates for the 6-MP arm of Freireich et al. (1963) are
// the worked example reproduced in Collett, Modelling Survival Data in Medical
// Research, and in the standard course treatments of the gehan data:
//
// t: 6 7 10 13 16 22 23
// S(t): 0.857 0.807 0.753 0.690 0.627 0.538 0.448
func TestKaplanMeier_MatchesPublishedGehanEstimates(t *testing.T) {
curve := kaplanMeier(gehanSixMercaptopurine)
expected := map[float64]float64{
6: 0.857, 7: 0.807, 10: 0.753, 13: 0.690, 16: 0.627, 22: 0.538, 23: 0.448,
}
seen := 0
for _, point := range curve {
want, ok := expected[point.Steps]
if !ok {
continue
}
seen++
if math.Abs(point.Survival-want) > 5e-4 {
t.Errorf("S(%v) = %.4f, want %v", point.Steps, point.Survival, want)
}
}
if seen != len(expected) {
t.Fatalf("matched %d of %d published times", seen, len(expected))
}
}
// The risk set at each time is the count of runs still under observation, with
// runs censored at a tied time counted as at risk for that event.
func TestKaplanMeier_RiskSetHandlesTiesAndCensoring(t *testing.T) {
curve := kaplanMeier(gehanSixMercaptopurine)
expected := map[float64]struct {
atRisk int
events int
censored int
}{
6: {21, 3, 1},
7: {17, 1, 0},
9: {16, 0, 1},
10: {15, 1, 1},
13: {12, 1, 0},
23: {6, 1, 0},
}
for _, point := range curve {
want, ok := expected[point.Steps]
if !ok {
continue
}
if point.AtRisk != want.atRisk || point.Events != want.events || point.Censored != want.censored {
t.Errorf("at %v: risk=%d events=%d censored=%d, want risk=%d events=%d censored=%d",
point.Steps, point.AtRisk, point.Events, point.Censored, want.atRisk, want.events, want.censored)
}
}
}
// Published medians: 23 weeks for 6-MP against 8 weeks for placebo (Gehan and
// Freireich, "The 6-MP versus placebo clinical trial in acute leukemia",
// Clinical Trials 8(3), 2011), and 31 against 23 weeks for the aml arms as
// reported by survfit in R's survival package.
func TestMedianSurvival_MatchesPublishedMedians(t *testing.T) {
cases := []struct {
name string
observations []observation
expected float64
}{
{"gehan 6-MP", gehanSixMercaptopurine, 23},
{"gehan placebo", gehanPlacebo, 8},
{"aml maintained", amlMaintained, 31},
{"aml nonmaintained", amlNonmaintained, 23},
}
for _, test := range cases {
median, ok := medianSurvival(kaplanMeier(test.observations))
if !ok {
t.Errorf("%s: median undefined, want %v", test.name, test.expected)
continue
}
if median != test.expected {
t.Errorf("%s: median %v, want %v", test.name, median, test.expected)
}
}
}
// R's survival package documents this log-rank on the aml data:
//
// N Observed Expected (O-E)^2/E (O-E)^2/V
// x=Maintained 11 7 10.69 1.27 3.4
// x=Nonmaintained 12 11 7.31 1.86 3.4
// Chisq= 3.4 on 1 degrees of freedom, p= 0.0653
func TestLogRank_MatchesPublishedAmlResult(t *testing.T) {
result := logRank([]string{"maintained", "nonmaintained"}, [][]observation{amlMaintained, amlNonmaintained})
if result.Observed[0] != 7 || result.Observed[1] != 11 {
t.Errorf("observed %v, want [7 11]", result.Observed)
}
if math.Abs(result.Expected[0]-10.69) > 5e-3 || math.Abs(result.Expected[1]-7.31) > 5e-3 {
t.Errorf("expected %v, want [10.69 7.31]", result.Expected)
}
for index, want := range []float64{1.27, 1.86} {
difference := result.Observed[index] - result.Expected[index]
got := difference * difference / result.Expected[index]
if math.Abs(got-want) > 5e-3 {
t.Errorf("(O-E)^2/E for group %d = %.4f, want %v", index, got, want)
}
}
if math.Abs(result.ChiSquare-3.4) > 5e-2 {
t.Errorf("chi-square %.4f, want 3.4", result.ChiSquare)
}
if math.Abs(result.PValue-0.0653) > 5e-4 {
t.Errorf("p-value %.6f, want 0.0653", result.PValue)
}
if result.DegreesOfFreedom != 1 {
t.Errorf("degrees of freedom %d, want 1", result.DegreesOfFreedom)
}
}
// The log-rank on the Freireich 6-MP trial is the textbook worked example:
// observed 9 against 19.25 expected in the treated arm and 21 against 10.75 in
// the control arm, Mantel-Haenszel chi-square 16.79 on 1 degree of freedom,
// p = 4.17e-05. Reported for instance in Rodriguez, Kaplan-Meier and
// Mantel-Haenszel, https://grodri.github.io/survival/gehan, and in Collett.
func TestLogRank_MatchesPublishedGehanResult(t *testing.T) {
result := logRank([]string{"6-MP", "placebo"}, [][]observation{gehanSixMercaptopurine, gehanPlacebo})
if result.Observed[0] != 9 || result.Observed[1] != 21 {
t.Errorf("observed %v, want [9 21]", result.Observed)
}
if math.Abs(result.Expected[0]-19.25) > 5e-3 || math.Abs(result.Expected[1]-10.75) > 5e-3 {
t.Errorf("expected %v, want [19.25 10.75]", result.Expected)
}
if math.Abs(result.ChiSquare-16.79) > 5e-3 {
t.Errorf("chi-square %.4f, want 16.79", result.ChiSquare)
}
if math.Abs(result.PValue-4.17e-5) > 5e-8 {
t.Errorf("p-value %.3e, want 4.17e-05", result.PValue)
}
}
// An arm with no usable runs contributes nothing and must not consume a degree
// of freedom or make the covariance matrix singular.
func TestLogRank_EmptyGroupIsDropped(t *testing.T) {
two := logRank([]string{"a", "b"}, [][]observation{amlMaintained, amlNonmaintained})
three := logRank([]string{"a", "b", "c"}, [][]observation{amlMaintained, amlNonmaintained, nil})
if math.Abs(two.ChiSquare-three.ChiSquare) > 1e-12 {
t.Errorf("chi-square %.10f with an empty third group, want %.10f", three.ChiSquare, two.ChiSquare)
}
if three.DegreesOfFreedom != 1 {
t.Errorf("degrees of freedom %d, want 1", three.DegreesOfFreedom)
}
if len(three.Groups) != 2 {
t.Errorf("groups %v, want the empty arm dropped", three.Groups)
}
}
// No published multi-arm dataset with a printed log-rank chi-square was found
// small enough to embed, so the k-group covariance algebra is checked against
// its own null distribution instead: under the null of equal hazards the
// statistic is asymptotically chi-square on k-1 degrees of freedom, so its mean
// over random relabellings has to sit near k-1. A wrong variance term or a wrong
// degrees-of-freedom count moves this badly.
func TestLogRank_NullMeanTracksDegreesOfFreedom(t *testing.T) {
pooled := make([]observation, 0, 36)
for index := 0; index < 36; index++ {
pooled = append(pooled, observation{Steps: float64(index%17 + 1), Event: index%5 != 0})
}
for _, groupCount := range []int{2, 3, 4} {
generator := rand.New(rand.NewSource(20260812))
total := 0.0
const replicates = 4000
for replicate := 0; replicate < replicates; replicate++ {
shuffled := slices.Clone(pooled)
generator.Shuffle(len(shuffled), func(left, right int) {
shuffled[left], shuffled[right] = shuffled[right], shuffled[left]
})
names := make([]string, groupCount)
groups := make([][]observation, groupCount)
for index, item := range shuffled {
groups[index%groupCount] = append(groups[index%groupCount], item)
}
for index := range names {
names[index] = strconv.Itoa(index)
}
total += logRank(names, groups).ChiSquare
}
mean := total / replicates
expected := float64(groupCount - 1)
if math.Abs(mean-expected) > 0.15*expected {
t.Errorf("%d groups: null mean chi-square %.3f, want near %v", groupCount, mean, expected)
}
}
}
// A three-group split of one homogeneous sample must not look significant, and
// the statistic must be finite on 2 degrees of freedom.
func TestLogRank_ThreeIdenticalGroupsAreNotSignificant(t *testing.T) {
group := []observation{{4, true}, {7, true}, {9, false}, {12, true}, {20, false}}
result := logRank([]string{"a", "b", "c"}, [][]observation{group, group, group})
if result.DegreesOfFreedom != 2 {
t.Fatalf("degrees of freedom %d, want 2", result.DegreesOfFreedom)
}
if result.ChiSquare > 1e-9 {
t.Errorf("chi-square %.10f for three identical groups, want 0", result.ChiSquare)
}
if math.Abs(result.PValue-1) > 1e-9 {
t.Errorf("p-value %v, want 1", result.PValue)
}
}
func TestKaplanMeier_EveryObservationCensored(t *testing.T) {
observations := []observation{{40, false}, {40, false}, {40, false}}
curve := kaplanMeier(observations)
if len(curve) != 1 {
t.Fatalf("curve has %d points, want 1", len(curve))
}
if curve[0].Survival != 1 || curve[0].Events != 0 || curve[0].Censored != 3 {
t.Errorf("point %+v, want survival 1 with 3 censored", curve[0])
}
if _, ok := medianSurvival(curve); ok {
t.Error("median defined for an arm where nothing violated")
}
}
func TestKaplanMeier_EveryObservationAnEvent(t *testing.T) {
curve := kaplanMeier([]observation{{2, true}, {4, true}, {6, true}, {8, true}})
last := curve[len(curve)-1]
if last.Survival != 0 {
t.Errorf("final survival %v, want 0", last.Survival)
}
median, ok := medianSurvival(curve)
if !ok || median != 4 {
t.Errorf("median %v ok=%v, want 4", median, ok)
}
}
func TestKaplanMeier_TiedEventTimesDropOnce(t *testing.T) {
curve := kaplanMeier([]observation{{5, true}, {5, true}, {5, true}, {9, true}})
if len(curve) != 2 {
t.Fatalf("curve has %d points, want 2", len(curve))
}
if curve[0].Events != 3 || math.Abs(curve[0].Survival-0.25) > 1e-12 {
t.Errorf("first point %+v, want 3 events and survival 0.25", curve[0])
}
if curve[1].Survival != 0 {
t.Errorf("second point %+v, want survival 0", curve[1])
}
}
func TestKaplanMeier_SingleObservation(t *testing.T) {
event := kaplanMeier([]observation{{11, true}})
if len(event) != 1 || event[0].Survival != 0 {
t.Fatalf("single event curve %+v", event)
}
median, ok := medianSurvival(event)
if !ok || median != 11 {
t.Errorf("median %v ok=%v, want 11", median, ok)
}
censored := kaplanMeier([]observation{{11, false}})
if len(censored) != 1 || censored[0].Survival != 1 {
t.Fatalf("single censored curve %+v", censored)
}
if _, ok := medianSurvival(censored); ok {
t.Error("median defined for a single censored observation")
}
}
func TestKaplanMeier_NoObservations(t *testing.T) {
if curve := kaplanMeier(nil); curve != nil {
t.Errorf("curve %v for no observations, want nil", curve)
}
if _, ok := medianSurvival(nil); ok {
t.Error("median defined for an empty curve")
}
}
func TestLogRank_SingleObservationPerGroup(t *testing.T) {
result := logRank([]string{"a", "b"}, [][]observation{{{3, true}}, {{9, true}}})
if math.IsNaN(result.ChiSquare) || result.ChiSquare < 0 {
t.Errorf("chi-square %v", result.ChiSquare)
}
if math.IsNaN(result.PValue) || result.PValue > 1 || result.PValue < 0 {
t.Errorf("p-value %v", result.PValue)
}
}
// With no events anywhere the covariance matrix is singular and there is
// nothing to test, which must report no difference rather than a divide by zero.
func TestLogRank_NoEventsAnywhere(t *testing.T) {
result := logRank([]string{"a", "b"}, [][]observation{
{{40, false}, {40, false}},
{{40, false}, {40, false}, {40, false}},
})
if result.ChiSquare != 0 || result.PValue != 1 {
t.Errorf("chi-square %v p-value %v, want 0 and 1", result.ChiSquare, result.PValue)
}
}
// The weighted family reports event counts, not weighted ones: the report
// prints observed against expected as counts of violations, and a weight that
// changes the statistic must leave those alone.
func TestWeightedLogRank_WeightsTheStatisticAndNotTheCounts(t *testing.T) {
names := []string{"6-mp", "placebo"}
groups := [][]observation{gehanSixMercaptopurine, gehanPlacebo}
plain := logRank(names, groups)
weighted := weightedLogRank(names, groups, atRiskWeight)
if !slices.Equal(plain.Observed, weighted.Observed) || !slices.Equal(plain.Expected, weighted.Expected) {
t.Errorf("weighted observed %v expected %v, want the counts %v and %v",
weighted.Observed, weighted.Expected, plain.Observed, plain.Expected)
}
if math.Abs(weighted.ChiSquare-plain.ChiSquare) < 1e-9 {
t.Errorf("chi-square %v under Gehan's weight and %v unweighted, want the weight to reach the statistic",
weighted.ChiSquare, plain.ChiSquare)
}
}
+71 -8
View File
@@ -1,14 +1,25 @@
// Command bundle-check is a developer tool that bundles a spec file to confirm it compiles.
// Command bundle-check is a developer tool that bundles a spec file and loads
// it into the evaluator to confirm it compiles and registers properties.
package main
import (
"errors"
"flag"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"github.com/priyanshujain/sanderling/internal/bundler"
"github.com/priyanshujain/sanderling/internal/testrun"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// checkSeed keeps the load deterministic. The bundle it seeds is only loaded,
// never hashed or reported, so the value is arbitrary.
const checkSeed = 1
func bundleSpec(specSrc, entryFile string) (bundler.Result, error) {
return bundler.Bundle(bundler.Options{
EntryFile: entryFile,
@@ -20,13 +31,67 @@ func bundleSpec(specSrc, entryFile string) (bundler.Result, error) {
})
}
// registeredProperties bundles the spec the way a run bundles it, with the
// runtime entry that assigns globalThis.properties, and loads it into the real
// evaluator. Bundling alone proves nothing about registration: a spec that
// registers no property compiles perfectly and then judges nothing.
func registeredProperties(entryFile string) ([]string, error) {
bundle, err := testrun.BundleSpec(entryFile, checkSeed)
if err != nil {
return nil, fmt.Errorf("bundle with runtime entry: %w", err)
}
evaluator, err := verifier.New()
if err != nil {
return nil, fmt.Errorf("evaluator: %w", err)
}
if err := evaluator.Load(string(bundle.JavaScript)); err != nil {
return nil, fmt.Errorf("load spec: %w", err)
}
return evaluator.PropertyNames(), nil
}
func check(specSrc, entryFile string, stdout io.Writer) error {
return checkWithOptions(specSrc, entryFile, false, stdout)
}
func checkWithOptions(specSrc, entryFile string, allowNoProperties bool, stdout io.Writer) error {
result, err := bundleSpec(specSrc, entryFile)
if err != nil {
return fmt.Errorf("bundle: %w", err)
}
fmt.Fprintf(stdout, "bundled: %d bytes, sha256=%s\n", len(result.JavaScript), result.SHA256)
names, err := registeredProperties(entryFile)
if err != nil {
return err
}
if len(names) == 0 && !allowNoProperties {
return errors.New("the spec bundles and loads cleanly but registers no properties: " +
"nothing is wrong with the source, and a run against it would check nothing " +
"and report no violations. Pass --allow-no-properties for a spec that measures " +
"what it extracts or where the generator reaches")
}
fmt.Fprintf(stdout, "properties registered: %d (%s)\n", len(names), strings.Join(names, ", "))
return nil
}
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: bundle-check <spec.ts>")
flagSet := flag.NewFlagSet("bundle-check", flag.ExitOnError)
allowNoProperties := flagSet.Bool("allow-no-properties", false,
"accept a spec that registers no properties, for a pre-registration that measures what the spec extracts or where the generator reaches")
flagSet.Usage = func() {
fmt.Fprintln(flagSet.Output(), "usage: bundle-check [--allow-no-properties] <spec.ts>")
flagSet.PrintDefaults()
}
if err := flagSet.Parse(os.Args[1:]); err != nil {
os.Exit(1)
}
if flagSet.NArg() != 1 {
flagSet.Usage()
os.Exit(1)
}
entryFile, err := filepath.Abs(os.Args[1])
entryFile, err := filepath.Abs(flagSet.Arg(0))
if err != nil {
fmt.Fprintf(os.Stderr, "resolve spec path: %v\n", err)
os.Exit(1)
@@ -38,10 +103,8 @@ func main() {
os.Exit(1)
}
result, err := bundleSpec(filepath.Join(repoRoot, "pkg/spec/src"), entryFile)
if err != nil {
fmt.Fprintf(os.Stderr, "bundle: %v\n", err)
if err := checkWithOptions(filepath.Join(repoRoot, "pkg/spec/src"), entryFile, *allowNoProperties, os.Stdout); err != nil {
fmt.Fprintf(os.Stderr, "%v\n", err)
os.Exit(1)
}
fmt.Printf("bundled: %d bytes, sha256=%s\n", len(result.JavaScript), result.SHA256)
}
@@ -1,8 +1,10 @@
package main
import (
"bytes"
"os"
"path/filepath"
"strings"
"testing"
)
@@ -55,3 +57,72 @@ func TestBundleSpec_ResolvesSpecAliases(t *testing.T) {
t.Errorf("unstable bundle hash: %s vs %s", result.SHA256, second.SHA256)
}
}
func testdataSpec(t *testing.T, name string) string {
t.Helper()
path, err := filepath.Abs(filepath.Join("testdata", name))
if err != nil {
t.Fatal(err)
}
return path
}
func TestCheck_RejectsSpecThatRegistersNoProperties(t *testing.T) {
var stdout bytes.Buffer
err := check(repoSpecSrc(t), testdataSpec(t, "no-properties.ts"), &stdout)
if err == nil {
t.Fatal("a spec registering zero properties passed the gate: a run against it reports no violations while checking nothing")
}
if !strings.Contains(err.Error(), "registers no properties") {
t.Errorf("failure must name the missing registration, got: %v", err)
}
if !strings.Contains(stdout.String(), "bundled: ") {
t.Errorf("bundle size and hash must still be reported for a spec awaiting its properties, got: %q", stdout.String())
}
}
// The extraction and portability sweeps freeze a spec that registers nothing
// on purpose, and run it with --allow-no-properties. Without the same opt-out
// here the gate that is supposed to freeze those pre-registrations is the one
// thing that cannot accept them.
func TestCheck_RunsTheZeroPropertySpecUnderTheOptOut(t *testing.T) {
var stdout bytes.Buffer
if err := checkWithOptions(repoSpecSrc(t), testdataSpec(t, "no-properties.ts"), true, &stdout); err != nil {
t.Fatalf("the opt-out did not admit a spec that registers nothing: %v", err)
}
if !strings.Contains(stdout.String(), "properties registered: 0") {
t.Errorf("the report must still say nothing was registered, got: %q", stdout.String())
}
if !strings.Contains(stdout.String(), "bundled: ") {
t.Errorf("bundle size and hash must still be reported, got: %q", stdout.String())
}
}
func TestCheck_ReportsRegisteredPropertyCountAndNames(t *testing.T) {
var stdout bytes.Buffer
if err := check(repoSpecSrc(t), testdataSpec(t, "spec.ts"), &stdout); err != nil {
t.Fatalf("check: %v", err)
}
if !strings.Contains(stdout.String(), "properties registered: 1 (noUncaughtExceptions)") {
t.Errorf("expected the registered count and name, got: %q", stdout.String())
}
}
// TestCheck_ReportsUnchangedBundleHash pins the reported bundle of a fixture
// that imports nothing, so the hash moves only when the bundler itself does.
// Frozen pre-registrations record these hashes, so bundle-check reporting a
// different bundle (the runtime-entry one, say) silently invalidates them.
func TestCheck_ReportsUnchangedBundleHash(t *testing.T) {
const (
plainBytes = 57
plainSHA256 = "bd757084b3f29c04c68ba2aa4c7b93e63b6dc6dd08ea3f726e0d1878be845888"
)
result, err := bundleSpec(repoSpecSrc(t), testdataSpec(t, "plain.ts"))
if err != nil {
t.Fatal(err)
}
if len(result.JavaScript) != plainBytes || result.SHA256 != plainSHA256 {
t.Errorf("reported bundle changed: %d bytes sha256=%s, want %d bytes sha256=%s",
len(result.JavaScript), result.SHA256, plainBytes, plainSHA256)
}
}
@@ -0,0 +1,7 @@
import { Tap, actions } from "@sanderling/spec";
export const properties = {};
export const actionsRoot = actions((state) => {
const button = state.ax.find("desc:primary");
return button ? [Tap({ on: button })] : [];
});
+1
View File
@@ -0,0 +1 @@
export const answer = 42;
+209 -22
View File
@@ -10,10 +10,14 @@ import (
"os"
"os/exec"
"path/filepath"
"slices"
"strconv"
"strings"
"sync"
"syscall"
"time"
"github.com/priyanshujain/sanderling/internal/android"
)
// commandExecutor runs one sanderling invocation and returns its exit code.
@@ -25,6 +29,12 @@ func executeCommand(ctx context.Context, binary string, arguments []string, outp
command := exec.CommandContext(ctx, binary, arguments...)
command.Stdout = output
command.Stderr = output
// SIGTERM rather than the default kill: a run killed outright never runs its
// own shutdown, and its sidecar survives holding a port and a quarter
// gigabyte. The run timeout exists for unattended hosts, which is exactly
// where nobody is watching to reap what it leaves.
command.Cancel = func() error { return command.Process.Signal(syscall.SIGTERM) }
command.WaitDelay = runShutdownGrace
err := command.Run()
if err == nil {
return 0, nil
@@ -33,31 +43,135 @@ func executeCommand(ctx context.Context, binary string, arguments []string, outp
if errors.As(err, &exitError) {
return exitError.ExitCode(), nil
}
if ctx.Err() != nil {
return -1, nil
}
return -1, err
}
// runRecord is one line of runs.jsonl.
// connectedDevices lists the devices the host currently has. A variable so a
// preflight test runs without a device farm attached.
var connectedDevices = android.ConnectedDevices
// A failure that came back in less than fastFailureThreshold never did the
// work the run was asked to do: a step-budgeted run takes tens of minutes,
// while a worker pointed at a device that is gone gives up in about half a
// minute. fastFailuresBeforeQuarantine of those in a row, with no run that
// worked in between to reset the count, is a property of the device rather
// than a flake, and it is where the cost of being wrong (one worker's
// throughput, since the seeds stay on the shared queue) is still smaller than
// the cost of being right one seed later.
const (
fastFailureThreshold = 2 * time.Minute
fastFailuresBeforeQuarantine = 3
)
// runShutdownGrace bounds how long a signalled run gets to stop its sidecar
// before it is killed. It exceeds the sidecar's own 15s shutdown grace, or the
// escalation would land while the run was still doing what it was asked.
const runShutdownGrace = 30 * time.Second
// runRecord is one line of runs.jsonl. MonotonicMillis is how long the run
// worked and WallClockMillis is how much time passed; they answer different
// questions and differ by however long the host slept mid-run, which an
// unattended overnight sweep is exactly where to expect.
type runRecord struct {
Seed int64 `json:"seed"`
Device string `json:"device,omitempty"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error,omitempty"`
TimedOut bool `json:"timed_out,omitempty"`
StartedAt time.Time `json:"started_at"`
DurationMillis int64 `json:"duration_millis"`
RunDirectory string `json:"run_directory,omitempty"`
TraceError string `json:"trace_error,omitempty"`
Seed int64 `json:"seed"`
Device string `json:"device,omitempty"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error,omitempty"`
TimedOut bool `json:"timed_out,omitempty"`
StartedAt time.Time `json:"started_at"`
MonotonicMillis int64 `json:"monotonic_millis"`
WallClockMillis int64 `json:"wall_clock_millis"`
RunDirectory string `json:"run_directory,omitempty"`
TraceError string `json:"trace_error,omitempty"`
traceSummary
}
// clocks reads the two measures a run is timed on. The monotonic clock does
// not advance while the host is asleep, so on its own it reports a run that
// slept through a quarter of an hour as a quarter of an hour shorter than it
// was; the wall clock advances but can be stepped by the host.
type clocks struct {
monotonicNow func() time.Time
wallClockNow func() time.Time
}
func systemClocks() clocks {
return clocks{
monotonicNow: time.Now,
// Round(0) drops the monotonic reading time.Now carries, so
// subtracting two of these readings uses the wall clock.
wallClockNow: func() time.Time { return time.Now().Round(0) },
}
}
type campaign struct {
configuration config
executor commandExecutor
stdout io.Writer
records io.Writer
clocks clocks
mutex sync.Mutex
failures int
unreadable int
quarantined []quarantinedDevice
// unrunSeeds are the seeds left without a trustworthy result: those a
// quarantined device consumed on its way out, and those still queued when
// the last worker stopped.
unrunSeeds []int64
}
// failureStreak is one worker's run of fast failures, and the seeds they cost.
type failureStreak struct {
fastFailures int
consumedSeeds []int64
}
// failedFast reports a run that came back non-zero too quickly to have done the
// work it was given. A killed run is excluded: it outlived the run timeout,
// which is the opposite of failing fast.
func failedFast(record runRecord) bool {
return record.ExitCode != 0 && !record.TimedOut &&
time.Duration(record.MonotonicMillis)*time.Millisecond < fastFailureThreshold
}
// preflightDevices refuses to start until every serial in --devices is present.
// A worker aimed at a serial that is gone fails in seconds and pulls the next
// seed, so a handful of dead serials drain the queue while the healthy workers
// are still inside their first run. Discovering that on run 1 of 20 is already
// too late: the sweep is spent, and its output does not say why.
func preflightDevices(ctx context.Context, configuration config) error {
// Android is the platform whose worker names a device this can enumerate: a
// web worker is a label with nothing behind it, and an --ios-device is
// resolved by simctl or by devicectl depending on whether it names a
// simulator or a paired phone.
if configuration.platform != "android" || len(configuration.devices) == 0 {
return nil
}
present, err := connectedDevices(ctx)
if err != nil {
return fmt.Errorf("list android devices: %w", err)
}
var missing []string
for _, device := range configuration.devices {
if !slices.Contains(present, device) {
missing = append(missing, device)
}
}
if len(missing) == 0 {
return nil
}
return fmt.Errorf("not starting: %d of %d --devices not connected: %s (adb reports %s)",
len(missing), len(configuration.devices), strings.Join(missing, ", "), presentDevices(present))
}
func presentDevices(present []string) string {
if len(present) == 0 {
return "no devices"
}
return strings.Join(present, ", ")
}
func runCampaign(ctx context.Context, configuration config, executor commandExecutor, stdout io.Writer) error {
@@ -65,6 +179,9 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
return fmt.Errorf("%s already exists in %s: pick a fresh --output so two campaigns do not share a directory",
manifestFileName, configuration.outputDirectory)
}
if err := preflightDevices(ctx, configuration); err != nil {
return err
}
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
return fmt.Errorf("create campaign dir: %w", err)
}
@@ -79,7 +196,8 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
return err
}
host, _ := os.Hostname()
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, host, binaryPath, version, time.Now().UTC())); err != nil {
intended := buildManifest(configuration, host, binaryPath, version, time.Now().UTC())
if err := writeManifest(configuration.outputDirectory, intended); err != nil {
return fmt.Errorf("write %s: %w", manifestFileName, err)
}
@@ -90,13 +208,26 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
}
defer recordsFile.Close()
sweep := &campaign{configuration: configuration, executor: executor, stdout: stdout, records: recordsFile}
sweep := &campaign{configuration: configuration, executor: executor, stdout: stdout, records: recordsFile, clocks: systemClocks()}
fmt.Fprintf(stdout, "campaign %s: %d seeds, %d worker(s), %s\n",
configuration.arm, len(configuration.seeds), len(workerDevices(configuration.devices)), configuration.outputDirectory)
sweep.sweep(ctx)
fmt.Fprintf(stdout, "campaign complete: %d of %d runs failed, %d produced an unreadable trace\n",
sweep.failures, len(configuration.seeds), sweep.unreadable)
if len(sweep.quarantined) > 0 || len(sweep.unrunSeeds) > 0 {
intended.Quarantined = sweep.quarantined
intended.UnrunSeeds = sweep.unrunSeeds
if err := writeManifest(configuration.outputDirectory, intended); err != nil {
fmt.Fprintf(stdout, "warning: rewrite %s: %v\n", manifestFileName, err)
}
fmt.Fprintf(stdout, "%d of %d seeds have no result: %v\n",
len(sweep.unrunSeeds), len(configuration.seeds), sweep.unrunSeeds)
}
if len(sweep.quarantined) > 0 && len(sweep.quarantined) == len(workerDevices(configuration.devices)) {
return fmt.Errorf("every device was quarantined (%s); %d of %d seeds have no result",
strings.Join(quarantinedNames(sweep.quarantined), ", "), len(sweep.unrunSeeds), len(configuration.seeds))
}
if sweep.failures > 0 {
return fmt.Errorf("%d of %d runs failed", sweep.failures, len(configuration.seeds))
}
@@ -142,15 +273,68 @@ func (c *campaign) sweep(ctx context.Context) {
waitGroup.Add(1)
go func(device string) {
defer waitGroup.Done()
for seed := range queue {
if ctx.Err() != nil {
return
}
c.report(c.runSeed(ctx, seed, device))
}
c.work(ctx, device, queue)
}(device)
}
waitGroup.Wait()
for seed := range queue {
c.unrunSeeds = append(c.unrunSeeds, seed)
}
slices.Sort(c.unrunSeeds)
}
// work runs seeds on one device until the queue is empty or the device has
// failed fast often enough in a row to be quarantined. The streak is worker
// local: one worker drives one device, and any run that did its work clears it.
func (c *campaign) work(ctx context.Context, device string, queue <-chan int64) {
var streak failureStreak
for seed := range queue {
if ctx.Err() != nil {
return
}
record := c.runSeed(ctx, seed, device)
c.report(record)
// A cancelled campaign fails every run in flight in seconds, which is
// the shutdown and not the device.
if ctx.Err() != nil {
return
}
if !failedFast(record) {
streak = failureStreak{}
continue
}
streak.fastFailures++
streak.consumedSeeds = append(streak.consumedSeeds, seed)
if streak.fastFailures >= fastFailuresBeforeQuarantine {
c.quarantine(device, streak)
return
}
}
}
// quarantine takes a device out of the sweep. The seeds it burned are reported
// unrun rather than requeued: each already holds that attempt's log, and a
// second record for the same seed would make one seed count as two runs.
func (c *campaign) quarantine(device string, streak failureStreak) {
c.mutex.Lock()
defer c.mutex.Unlock()
c.quarantined = append(c.quarantined, quarantinedDevice{
Device: device,
FastFailures: streak.fastFailures,
ConsumedSeeds: streak.consumedSeeds,
})
c.unrunSeeds = append(c.unrunSeeds, streak.consumedSeeds...)
fmt.Fprintf(c.stdout, "quarantined device %q: %d runs in a row failed in under %s; seeds %v have no result\n",
device, streak.fastFailures, fastFailureThreshold, streak.consumedSeeds)
}
func quarantinedNames(devices []quarantinedDevice) []string {
names := make([]string, 0, len(devices))
for _, device := range devices {
names = append(names, device.Device)
}
return names
}
func (c *campaign) runSeed(ctx context.Context, seed int64, device string) runRecord {
@@ -173,12 +357,14 @@ func (c *campaign) runSeed(ctx context.Context, seed int64, device string) runRe
runCtx, cancelRun := context.WithTimeout(ctx, c.configuration.runTimeout)
defer cancelRun()
start := time.Now()
monotonicStart := c.clocks.monotonicNow()
wallClockStart := c.clocks.wallClockNow()
exitCode, runErr := c.executor(runCtx, c.configuration.sanderlingPath, runArguments(c.configuration, seedText, device), logFile)
if runCtx.Err() != nil && ctx.Err() == nil {
record.TimedOut = true
}
record.DurationMillis = time.Since(start).Milliseconds()
record.MonotonicMillis = c.clocks.monotonicNow().Sub(monotonicStart).Milliseconds()
record.WallClockMillis = c.clocks.wallClockNow().Sub(wallClockStart).Milliseconds()
record.ExitCode = exitCode
if runErr != nil {
record.LaunchError = runErr.Error()
@@ -207,9 +393,10 @@ func (c *campaign) report(record runRecord) {
if err := json.NewEncoder(c.records).Encode(record); err != nil {
fmt.Fprintf(c.stdout, "warning: seed %d record: %v\n", record.Seed, err)
}
fmt.Fprintf(c.stdout, "seed=%d device=%q outcome=%s steps=%d exit=%d duration=%s\n",
fmt.Fprintf(c.stdout, "seed=%d device=%q outcome=%s steps=%d exit=%d monotonic=%s wall_clock=%s\n",
record.Seed, record.Device, outcome(record), record.Steps, record.ExitCode,
time.Duration(record.DurationMillis)*time.Millisecond)
time.Duration(record.MonotonicMillis)*time.Millisecond,
time.Duration(record.WallClockMillis)*time.Millisecond)
}
func outcome(record runRecord) string {
@@ -69,6 +69,91 @@ func writeFakeRun(t *testing.T, arguments []string, steps []trace.Step) {
writeRunDirectory(t, argumentValue(arguments, "--output"), "20260101-000000", steps)
}
// readings hands out the given instants in turn, so a test can script a clock
// that jumps across a host sleep independently of one that stops through it.
func readings(instants ...time.Time) func() time.Time {
var index int
return func() time.Time {
instant := instants[min(index, len(instants)-1)]
index++
return instant
}
}
// A sleeping host stops the monotonic clock and not the wall clock, so a run
// timed on the monotonic clock alone reports the sleep as time that never
// passed. The record carries both, named for the clock each came from.
func TestRunSeed_RecordsTheTimeWorkedAndTheTimeThatPassed(t *testing.T) {
directory := t.TempDir()
startedAt := time.Date(2026, 8, 14, 2, 0, 0, 0, time.UTC)
var records bytes.Buffer
sweep := &campaign{
configuration: testConfiguration(t, directory, "--seeds", "1"),
executor: func(context.Context, string, []string, io.Writer) (int, error) {
return 0, nil
},
stdout: io.Discard,
records: &records,
clocks: clocks{
monotonicNow: readings(startedAt, startedAt.Add(2*time.Minute)),
wallClockNow: readings(startedAt, startedAt.Add(17*time.Minute)),
},
}
sweep.report(sweep.runSeed(context.Background(), 1, ""))
var written map[string]any
if err := json.Unmarshal(records.Bytes(), &written); err != nil {
t.Fatalf("decode %q: %v", records.String(), err)
}
if written["monotonic_millis"] != float64((2 * time.Minute).Milliseconds()) {
t.Errorf("monotonic_millis %v, want the two minutes of work", written["monotonic_millis"])
}
if written["wall_clock_millis"] != float64((17 * time.Minute).Milliseconds()) {
t.Errorf("wall_clock_millis %v, want the seventeen minutes that passed", written["wall_clock_millis"])
}
}
func TestRunCampaign_RecordsDispatchedActionsNotSteps(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1")
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
writeFakeRun(t, arguments, []trace.Step{
actingStep(1), observedStep(2), skippedActionStep(3, "unresolved_selector"), actingStep(4),
})
return 0, nil
})
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
t.Fatal(err)
}
records := readRecords(t, directory)
if len(records) != 1 {
t.Fatalf("records: got %d, want 1", len(records))
}
if records[0].Steps != 4 || records[0].Actions != 2 {
t.Errorf("steps %d actions %d, want 4 and 2", records[0].Steps, records[0].Actions)
}
}
// A run that never produced a readable trace still has to carry the field, so
// analysis can tell a zero-action run from a file written before the count. The
// same holds for the unattributed count: a run whose every action named its
// producer says so with a zero, and a file that says nothing is one recorded
// before actions named one at all.
func TestRunRecord_AlwaysCarriesTheActionCounts(t *testing.T) {
body, err := json.Marshal(runRecord{Seed: 7, TraceError: "no run directory with meta.json"})
if err != nil {
t.Fatal(err)
}
for _, field := range []string{`"actions":0`, `"unattributed_actions":0`} {
if !strings.Contains(string(body), field) {
t.Errorf("record %s omits %s", body, field)
}
}
}
func TestRunCampaign_WritesManifestBeforeAnyRun(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory)
@@ -172,6 +257,7 @@ func TestRunCampaign_RecordsPerRunSummary(t *testing.T) {
func TestRunCampaign_DistributesSeedsAcrossDeviceWorkers(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1-9", "--devices", "device-a,device-b,device-c")
devicesPresent(t, "device-a", "device-b", "device-c")
var mutex sync.Mutex
assignments := map[int64]string{}
@@ -263,6 +349,28 @@ func TestRunCampaign_ContinuesAfterFailingRun(t *testing.T) {
}
}
func TestRunCampaign_GivesEveryRunTheCellsLabelSource(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1-3", "--label-source", "resource-id")
var mutex sync.Mutex
var dispatched []string
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
mutex.Lock()
dispatched = append(dispatched, argumentValue(arguments, "--label-source"))
mutex.Unlock()
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
return 0, nil
})
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
t.Fatal(err)
}
if !slices.Equal(dispatched, []string{"resource-id", "resource-id", "resource-id"}) {
t.Errorf("--label-source reaching sanderling: got %v, want resource-id on every run", dispatched)
}
}
func TestRunCampaign_RefusesToReuseACampaignDirectory(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory)
@@ -360,3 +468,208 @@ func TestParseArguments_RunTimeoutDefaultsToThreeTimesDuration(t *testing.T) {
t.Errorf("run timeout default: got %s, want 12m", configuration.runTimeout)
}
}
// devicesPresent points the preflight at a fixed set of serials, so a campaign
// can be preflighted on a host with no device farm attached.
func devicesPresent(t *testing.T, present ...string) {
t.Helper()
original := connectedDevices
connectedDevices = func(context.Context) ([]string, error) { return present, nil }
t.Cleanup(func() { connectedDevices = original })
}
// A worker aimed at a serial that no longer exists fails in seconds and pulls
// the next seed, so a few dead serials drain the queue while the healthy
// workers are still inside their first run. That has to be caught before the
// first seed is dispatched: a sweep that discovers it on run 1 of 20 has
// already been destroyed, and its output does not say so.
func TestRunCampaign_RefusesToStartWhenADeviceIsMissing(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory,
"--seeds", "1-20", "--devices", "emulator-5554,emulator-5564,emulator-5556")
devicesPresent(t, "emulator-5554", "emulator-5556")
executor := func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
t.Errorf("the campaign dispatched %v despite a missing device", arguments)
return 0, nil
}
err := runCampaign(context.Background(), configuration, executor, io.Discard)
if err == nil || !strings.Contains(err.Error(), "not connected: emulator-5564") {
t.Fatalf("the error must name the missing serial, got %v", err)
}
if _, statErr := os.Stat(filepath.Join(directory, manifestFileName)); statErr == nil {
t.Error("a campaign that cannot run wrote a manifest")
}
}
func TestRunCampaign_RunsWhenEveryDeviceIsPresent(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1-2", "--devices", "emulator-5554,emulator-5556")
devicesPresent(t, "emulator-5556", "emulator-5580", "emulator-5554")
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
return 0, nil
})
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
t.Fatal(err)
}
if got := len(readRecords(t, directory)); got != 2 {
t.Errorf("records: got %d, want 2", got)
}
}
// Preflight only means something where the worker names a device. A campaign
// without --devices has one worker and no serial, and a web worker is a label
// with no device behind it; neither may be blocked by a device check.
func TestRunCampaign_PreflightsOnlyWhereWorkersNameADevice(t *testing.T) {
for _, testCase := range []struct {
name string
extra []string
}{
{"no devices", []string{"--seeds", "1"}},
{"web workers", []string{"--seeds", "1", "--platform", "web", "--devices", "worker-a,worker-b"}},
} {
t.Run(testCase.name, func(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, testCase.extra...)
original := connectedDevices
connectedDevices = func(context.Context) ([]string, error) {
t.Error("preflight looked for devices where the workers name none")
return nil, fmt.Errorf("no adb server")
}
t.Cleanup(func() { connectedDevices = original })
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
return 0, nil
})
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
t.Fatal(err)
}
})
}
}
// Preflight cannot catch a device that disappears mid-campaign. A device that
// keeps failing in seconds is drained of seeds by exactly the speed of its
// failure, so it has to stop being given any.
func TestRunCampaign_QuarantinesADeviceThatKeepsFailingFast(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1-8", "--devices", "device-a,device-b")
devicesPresent(t, "device-a", "device-b")
failedEnough := make(chan struct{})
var closeOnce sync.Once
var timedOut atomic.Bool
var mutex sync.Mutex
dispatched := map[string][]int64{}
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, output io.Writer) (int, error) {
device := argumentValue(arguments, "--device")
seed, err := strconv.ParseInt(argumentValue(arguments, "--seed"), 10, 64)
if err != nil {
t.Errorf("seed argument: %v", err)
}
mutex.Lock()
dispatched[device] = append(dispatched[device], seed)
failures := len(dispatched["device-a"])
mutex.Unlock()
if device == "device-a" {
fmt.Fprintln(output, "device 'device-a' not found")
if failures >= fastFailuresBeforeQuarantine {
closeOnce.Do(func() { close(failedEnough) })
}
return 1, nil
}
// The healthy worker holds its seed until the sick one has failed its
// way to quarantine, so the split of the queue is the scheduler's
// decision and not a race between two equally fast fakes.
select {
case <-failedEnough:
case <-time.After(5 * time.Second):
timedOut.Store(true)
}
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
return 0, nil
})
var stdout bytes.Buffer
err := runCampaign(context.Background(), configuration, executor, &stdout)
if timedOut.Load() {
t.Fatal("device-a never reached quarantine")
}
if err == nil || !strings.Contains(err.Error(), "3 of 8 runs failed") {
t.Fatalf("expected the three fast failures to be reported, got %v", err)
}
if got := len(dispatched["device-a"]); got != fastFailuresBeforeQuarantine {
t.Errorf("seeds sent to the failing device: got %d, want %d", got, fastFailuresBeforeQuarantine)
}
if got := len(dispatched["device-b"]); got != 8-fastFailuresBeforeQuarantine {
t.Errorf("seeds sent to the healthy device: got %d, want %d", got, 8-fastFailuresBeforeQuarantine)
}
ran := append(append([]int64{}, dispatched["device-a"]...), dispatched["device-b"]...)
slices.Sort(ran)
if !slices.Equal(ran, []int64{1, 2, 3, 4, 5, 6, 7, 8}) {
t.Errorf("every seed must be dispatched exactly once: got %v", ran)
}
if !strings.Contains(stdout.String(), `quarantined device "device-a"`) {
t.Errorf("the quarantine must be reported in the campaign output:\n%s", stdout.String())
}
recorded := readManifest(t, directory)
if len(recorded.Quarantined) != 1 || recorded.Quarantined[0].Device != "device-a" {
t.Fatalf("the manifest must record the quarantine: %+v", recorded.Quarantined)
}
burned := slices.Clone(dispatched["device-a"])
slices.Sort(burned)
if !slices.Equal(recorded.Quarantined[0].ConsumedSeeds, burned) {
t.Errorf("consumed seeds: got %v, want %v", recorded.Quarantined[0].ConsumedSeeds, burned)
}
if !slices.Equal(recorded.UnrunSeeds, burned) {
t.Errorf("the seeds the quarantined device burned have no result and must be reported unrun: got %v, want %v",
recorded.UnrunSeeds, burned)
}
}
// Every device gone is not a campaign that should keep pulling seeds: the rest
// of the queue would be spent producing the same failure.
func TestRunCampaign_AbortsWhenEveryDeviceIsQuarantined(t *testing.T) {
directory := t.TempDir()
configuration := testConfiguration(t, directory, "--seeds", "1-10", "--devices", "device-a,device-b")
devicesPresent(t, "device-a", "device-b")
var dispatches atomic.Int32
executor := versionAnswering(func(_ context.Context, _ string, _ []string, output io.Writer) (int, error) {
dispatches.Add(1)
fmt.Fprintln(output, "sidecar health check: context deadline exceeded")
return 1, nil
})
var stdout bytes.Buffer
err := runCampaign(context.Background(), configuration, executor, &stdout)
if err == nil || !strings.Contains(err.Error(), "every device was quarantined") {
t.Fatalf("a campaign with no device left must abort with a clear error, got %v", err)
}
if got := dispatches.Load(); got != int32(2*fastFailuresBeforeQuarantine) {
t.Errorf("runs dispatched: got %d, want %d: the sweep must stop rather than spin through the queue",
got, 2*fastFailuresBeforeQuarantine)
}
recorded := readManifest(t, directory)
if !slices.Equal(recorded.UnrunSeeds, []int64{1, 2, 3, 4, 5, 6, 7, 8, 9, 10}) {
t.Errorf("every seed is either burned by a quarantined device or never dispatched: got %v", recorded.UnrunSeeds)
}
}
func readManifest(t *testing.T, campaignDirectory string) manifest {
t.Helper()
body, err := os.ReadFile(filepath.Join(campaignDirectory, manifestFileName))
if err != nil {
t.Fatal(err)
}
var recorded manifest
if err := json.Unmarshal(body, &recorded); err != nil {
t.Fatal(err)
}
return recorded
}
@@ -2,6 +2,7 @@ package main
import (
"bytes"
"context"
"encoding/json"
"io"
"os"
@@ -9,6 +10,7 @@ import (
"slices"
"strings"
"testing"
"time"
)
// stubSanderling answers `version`, writes a run directory shaped like the one
@@ -142,3 +144,53 @@ func TestRun_EndToEndAgainstStubBinary(t *testing.T) {
t.Errorf("progress output: %q", stdout.String())
}
}
func TestExecuteCommand_RunTimeoutSignalsSoTheRunReapsItsChildren(t *testing.T) {
directory := t.TempDir()
marker := filepath.Join(directory, "child-reaped")
trapped := filepath.Join(directory, "trap-installed")
script := filepath.Join(directory, "wedged")
body := "#!/bin/sh\n" +
"sleep 300 &\n" +
"child=$!\n" +
"trap 'kill $child; echo reaped > " + marker + "; exit 143' TERM\n" +
"echo installed > " + trapped + "\n" +
"wait $child\n"
if err := os.WriteFile(script, []byte(body), 0o755); err != nil {
t.Fatal(err)
}
// Cancel only once the script has installed its trap. A fixed deadline
// races the shell under a loaded machine, and a run signalled before its
// trap exists fails this test for a reason it does not test.
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
finished := make(chan error, 1)
go func() {
_, err := executeCommand(ctx, script, nil, io.Discard)
finished <- err
}()
waitForFile(t, trapped)
cancel()
if err := <-finished; err != nil {
t.Fatalf("execute: %v", err)
}
if _, err := os.Stat(marker); err != nil {
t.Fatal("the run timeout killed the run outright, so it never reaped its own children: " +
"an unattended sweep leaks one sidecar per wedged run")
}
}
func waitForFile(t *testing.T, path string) {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
for time.Now().Before(deadline) {
if _, err := os.Stat(path); err == nil {
return
}
time.Sleep(5 * time.Millisecond)
}
t.Fatalf("%s never appeared", path)
}
+30 -9
View File
@@ -15,6 +15,8 @@ import (
"strings"
"syscall"
"time"
"github.com/priyanshujain/sanderling/internal/seedspec"
)
type config struct {
@@ -23,6 +25,7 @@ type config struct {
platform string
arm string
generator string
labelSource string
maxSteps int
duration time.Duration
seeds []int64
@@ -58,6 +61,7 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
flagSet.StringVar(&configuration.platform, "platform", "android", "target platform: android, ios, web")
flagSet.StringVar(&configuration.arm, "arm", "", "experiment cell label recorded on every run (required)")
flagSet.StringVar(&configuration.generator, "generator", "seeded", "action generator: seeded or llm")
flagSet.StringVar(&configuration.labelSource, "label-source", "visible-text", "how candidates are named to the llm generator: visible-text (what a user reads) or resource-id (the identifier the app assigned). The seeded generator picks by index and ignores this")
flagSet.IntVar(&configuration.maxSteps, "max-steps", 0, "per-run step budget (required, must be positive)")
flagSet.DurationVar(&configuration.duration, "duration", 5*time.Minute, "per-run wall-clock ceiling")
flagSet.StringVar(&seedSpecification, "seeds", "", "seeds to run: ranges and lists, e.g. 1-10,20,30-32 (required)")
@@ -70,17 +74,26 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
}
configuration.extraArguments = flagSet.Args()
for name, value := range map[string]string{
"--spec": configuration.specPath,
"--bundle-id": configuration.bundleID,
"--arm": configuration.arm,
"--seeds": seedSpecification,
"--output": configuration.outputDirectory,
// Every missing flag is named together, in flag order: stopping at the
// first turns one rerun into one rerun per missing flag.
var missing []error
for _, required := range []struct {
name string
value string
}{
{"--spec", configuration.specPath},
{"--bundle-id", configuration.bundleID},
{"--arm", configuration.arm},
{"--seeds", seedSpecification},
{"--output", configuration.outputDirectory},
} {
if value == "" {
return config{}, fmt.Errorf("%s is required", name)
if required.value == "" {
missing = append(missing, fmt.Errorf("%s is required", required.name))
}
}
if err := errors.Join(missing...); err != nil {
return config{}, err
}
switch configuration.platform {
case "android", "ios", "web":
default:
@@ -91,6 +104,13 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
default:
return config{}, fmt.Errorf("unsupported generator: %q (seeded, llm)", configuration.generator)
}
// Rejected before the first run, because a sweep that discovers the bad
// value on run 1 of 40 has already spent the device time of a whole cell.
switch configuration.labelSource {
case "visible-text", "resource-id":
default:
return config{}, fmt.Errorf("unsupported label source: %q (visible-text, resource-id)", configuration.labelSource)
}
if configuration.maxSteps <= 0 {
// Steps to first violation is right-censored at the budget, so a
// campaign without one has nothing to censor its clean runs at.
@@ -109,7 +129,7 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
return config{}, fmt.Errorf("--run-timeout %s must exceed --duration %s, or every run is killed before it finishes",
configuration.runTimeout, configuration.duration)
}
seeds, err := parseSeeds(seedSpecification)
seeds, err := seedspec.Parse(seedSpecification)
if err != nil {
return config{}, fmt.Errorf("--seeds: %w", err)
}
@@ -170,6 +190,7 @@ func runArguments(configuration config, seed, device string) []string {
"--platform", configuration.platform,
"--arm", configuration.arm,
"--generator", configuration.generator,
"--label-source", configuration.labelSource,
"--max-steps", strconv.Itoa(configuration.maxSteps),
"--duration", configuration.duration.String(),
"--seed", seed,
+40
View File
@@ -54,6 +54,7 @@ func TestParseArguments_Rejections(t *testing.T) {
{"missing output", []string{"--spec", "s", "--bundle-id", "a", "--arm", "b", "--seeds", "1", "--max-steps", "10"}, "--output is required"},
{"bad platform", append(baseArguments(), "--platform", "windows"), "unsupported platform"},
{"bad generator", append(baseArguments(), "--generator", "vibes"), "unsupported generator"},
{"bad label source", append(baseArguments(), "--label-source", "resource_id"), `unsupported label source: "resource_id"`},
{"zero max steps", append(baseArguments(), "--max-steps", "0"), "--max-steps must be positive"},
{"seed zero", append(baseArguments(), "--seeds", "0-2"), "not reproducible"},
{"duplicate device", append(baseArguments(), "--devices", "a,a"), "duplicate device"},
@@ -70,6 +71,35 @@ func TestParseArguments_Rejections(t *testing.T) {
}
}
// Three flags missing is one rerun, not three: the operator is told about all
// of them at once, in flag order, whatever order the check happened to walk.
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
_, err := parseArguments(
[]string{"--bundle-id", "a", "--seeds", "1", "--max-steps", "10"},
io.Discard,
)
if err == nil {
t.Fatal("got no error, want every missing flag named")
}
message := err.Error()
previous := -1
for _, name := range []string{"--spec", "--arm", "--output"} {
at := strings.Index(message, name)
if at < 0 {
t.Fatalf("got %q, want %s named", message, name)
}
if at < previous {
t.Errorf("got %q, want the flags named in flag order", message)
}
previous = at
}
for _, supplied := range []string{"--bundle-id", "--seeds"} {
if strings.Contains(message, supplied) {
t.Errorf("got %q, want the supplied %s left out", message, supplied)
}
}
}
func TestRunArguments_PlatformDeviceFlagAndPassthrough(t *testing.T) {
cases := []struct {
platform string
@@ -113,6 +143,16 @@ func TestRunArguments_PlatformDeviceFlagAndPassthrough(t *testing.T) {
}
}
func TestRunArguments_LabelSourceDefaultsToVisibleText(t *testing.T) {
configuration, err := parseArguments(baseArguments(), io.Discard)
if err != nil {
t.Fatal(err)
}
if got := argumentValue(runArguments(configuration, "7", ""), "--label-source"); got != "visible-text" {
t.Errorf("--label-source = %q, want visible-text", got)
}
}
func argumentValue(arguments []string, name string) string {
index := slices.Index(arguments, name)
if index < 0 || index+1 >= len(arguments) {
+18
View File
@@ -21,6 +21,7 @@ const (
type manifest struct {
Arm string `json:"arm"`
Generator string `json:"generator"`
LabelSource string `json:"label_source"`
Platform string `json:"platform"`
SpecPath string `json:"spec_path"`
BundleID string `json:"bundle_id"`
@@ -34,6 +35,22 @@ type manifest struct {
SanderlingVersion string `json:"sanderling_version"`
StartedAt time.Time `json:"started_at"`
ArgumentTemplate []string `json:"argument_template"`
// Quarantined and UnrunSeeds are filled in when the sweep ends, so the
// artifact says a device dropped out and which seeds have no result rather
// than leaving both to be inferred from a thin runs.jsonl.
Quarantined []quarantinedDevice `json:"quarantined,omitempty"`
UnrunSeeds []int64 `json:"unrun_seeds,omitempty"`
}
// quarantinedDevice is a device the sweep stopped assigning seeds to, with the
// seeds it consumed on the way out. Those seeds are reported unrun rather than
// requeued: their directories already hold the failed attempt's log, and a
// second record for the same seed would make a seed count two runs.
type quarantinedDevice struct {
Device string `json:"device"`
FastFailures int `json:"fast_failures"`
ConsumedSeeds []int64 `json:"consumed_seeds"`
}
func buildManifest(configuration config, host, binaryPath, version string, startedAt time.Time) manifest {
@@ -44,6 +61,7 @@ func buildManifest(configuration config, host, binaryPath, version string, start
return manifest{
Arm: configuration.arm,
Generator: configuration.generator,
LabelSource: configuration.labelSource,
Platform: configuration.platform,
SpecPath: configuration.specPath,
BundleID: configuration.bundleID,
@@ -41,6 +41,30 @@ func TestBuildManifest_RecordsIntendedRunsAndTemplate(t *testing.T) {
}
}
func TestWriteManifest_RecordsTheLabelSourceCell(t *testing.T) {
configuration, err := parseArguments(append(baseArguments(), "--label-source", "resource-id"), io.Discard)
if err != nil {
t.Fatal(err)
}
directory := t.TempDir()
if err := writeManifest(directory, buildManifest(configuration, "host", "sanderling", "dev", time.Now())); err != nil {
t.Fatal(err)
}
body, err := os.ReadFile(filepath.Join(directory, manifestFileName))
if err != nil {
t.Fatal(err)
}
var decoded struct {
LabelSource string `json:"label_source"`
}
if err := json.Unmarshal(body, &decoded); err != nil {
t.Fatal(err)
}
if decoded.LabelSource != "resource-id" {
t.Errorf("label_source: got %q, want resource-id", decoded.LabelSource)
}
}
func TestBuildManifest_EmptyDeviceListSerializesAsArray(t *testing.T) {
configuration, err := parseArguments(baseArguments(), io.Discard)
if err != nil {
-79
View File
@@ -1,79 +0,0 @@
package main
import (
"fmt"
"strconv"
"strings"
)
// parseSeeds expands a seed specification such as "1-10,20,30-32" into the
// explicit seed list a campaign intends to run.
func parseSeeds(specification string) ([]int64, error) {
trimmed := strings.TrimSpace(specification)
if trimmed == "" {
return nil, fmt.Errorf("empty seed spec")
}
var seeds []int64
seen := map[int64]bool{}
for _, part := range strings.Split(trimmed, ",") {
part = strings.TrimSpace(part)
if part == "" {
return nil, fmt.Errorf("empty seed in %q", specification)
}
expanded, err := expandSeedPart(part)
if err != nil {
return nil, err
}
for _, seed := range expanded {
if seen[seed] {
return nil, fmt.Errorf("duplicate seed %d in %q", seed, specification)
}
seen[seed] = true
seeds = append(seeds, seed)
}
}
return seeds, nil
}
func expandSeedPart(part string) ([]int64, error) {
start, end, isRange := strings.Cut(part, "-")
if !isRange {
seed, err := parseSeed(part)
if err != nil {
return nil, err
}
return []int64{seed}, nil
}
first, err := parseSeed(strings.TrimSpace(start))
if err != nil {
return nil, fmt.Errorf("seed range %q: %w", part, err)
}
last, err := parseSeed(strings.TrimSpace(end))
if err != nil {
return nil, fmt.Errorf("seed range %q: %w", part, err)
}
if first > last {
return nil, fmt.Errorf("seed range %q: start %d is above end %d", part, first, last)
}
seeds := make([]int64, 0, last-first+1)
for seed := first; seed <= last; seed++ {
seeds = append(seeds, seed)
}
return seeds, nil
}
func parseSeed(text string) (int64, error) {
seed, err := strconv.ParseInt(text, 10, 64)
if err != nil {
return 0, fmt.Errorf("invalid seed %q: want a positive integer", text)
}
if seed == 0 {
// `sanderling test` reads --seed 0 as "derive a seed from the clock",
// so a campaign listing seed 0 records a run nobody can reproduce.
return 0, fmt.Errorf("seed 0 is not reproducible: sanderling test derives a random seed when --seed is 0, so list explicit non-zero seeds")
}
if seed < 0 {
return 0, fmt.Errorf("invalid seed %d: want a positive integer", seed)
}
return seed, nil
}
-49
View File
@@ -1,49 +0,0 @@
package main
import (
"slices"
"strings"
"testing"
)
func TestParseSeeds_RangesAndLists(t *testing.T) {
cases := []struct {
specification string
want []int64
}{
{"1-5", []int64{1, 2, 3, 4, 5}},
{"1,5,9", []int64{1, 5, 9}},
{"1-3,20,30-32", []int64{1, 2, 3, 20, 30, 31, 32}},
{" 7 , 8 ", []int64{7, 8}},
{"4-4", []int64{4}},
}
for _, testCase := range cases {
got, err := parseSeeds(testCase.specification)
if err != nil {
t.Fatalf("%q: %v", testCase.specification, err)
}
if !slices.Equal(got, testCase.want) {
t.Errorf("%q: got %v, want %v", testCase.specification, got, testCase.want)
}
}
}
func TestParseSeeds_RejectsSeedZero(t *testing.T) {
for _, specification := range []string{"0", "1,0,2", "0-3"} {
_, err := parseSeeds(specification)
if err == nil {
t.Fatalf("%q: expected rejection of seed 0", specification)
}
if !strings.Contains(err.Error(), "not reproducible") {
t.Errorf("%q: error should explain why seed 0 is rejected: %v", specification, err)
}
}
}
func TestParseSeeds_RejectsMalformed(t *testing.T) {
for _, specification := range []string{"", " ", "abc", "1,,2", "5-1", "1-", "-5", "1-2-3", "1.5", "2,2"} {
if seeds, err := parseSeeds(specification); err == nil {
t.Errorf("%q: expected error, got %v", specification, seeds)
}
}
}
+100 -6
View File
@@ -19,22 +19,86 @@ const maxTraceLineBytes = 16 * 1024 * 1024
// traceSummary is everything the analysis needs from one run, so it never has
// to open trace.jsonl again.
type traceSummary struct {
Steps int `json:"steps"`
Steps int `json:"steps"`
// Actions counts only the steps on which the action generator both chose an
// action and dispatched it. A step whose policy declined to act, a step
// whose chosen action the runner threw away, and a step the spec's setup
// drove before the generator was ever consulted all explored nothing, so
// counting any of them as an action inflates the denominator of every
// per-action rate. The inflation is policy-dependent, so it does not cancel
// between arms.
Actions int `json:"actions"`
// UnattributedActions counts the dispatched steps whose action names no
// producer, which only a trace recorded before actions carried one can do.
// Such a run's Actions is the count it was already reported with rather than
// a setup-excluding one, and this is what says so. It is written even when
// it is zero, because a record that omits it is one the analysis has to read
// as unattributable and a run where every action named a producer is the
// opposite of that.
UnattributedActions int `json:"unattributed_actions"`
FirstViolationOriginStep *int `json:"first_violation_origin_step"`
FirstViolationDetectedStep *int `json:"first_violation_detected_step"`
FirstViolationProperties []string `json:"first_violation_properties,omitempty"`
FirstViolationReason string `json:"first_violation_reason,omitempty"`
FirstViolationIsError bool `json:"first_violation_is_error,omitempty"`
ViolatedProperties []string `json:"violated_properties,omitempty"`
// PreconditionFailures counts the trace records naming a precondition the
// run could not meet: the startup gate's verdict at step 0, and every later
// step the scope guard could not bring the app back for. A run with one of
// these and no steps never started, and counting it as a run that explored
// and found nothing puts a harness failure in the same column as evidence.
PreconditionFailures int `json:"precondition_failures,omitempty"`
}
type traceLine struct {
Index int `json:"step"`
// Hierarchy is read only for its presence: the run-end finalize line is the
// one line carrying violations without an observed hierarchy.
Hierarchy json.RawMessage `json:"hierarchy"`
Violations []string `json:"violations"`
Witnesses map[string]trace.Witness `json:"witnesses"`
Hierarchy json.RawMessage `json:"hierarchy"`
// NextAction is nil on a step that chose no action. Its Source names the
// backend that chose the action, which is the only thing in the trace
// separating a step the spec's setup drove from one the generator drove.
NextAction *actionLine `json:"next_action"`
ActionSkipped string `json:"action_skipped"`
PreconditionFailure string `json:"precondition_failure"`
Violations []string `json:"violations"`
Witnesses map[string]trace.Witness `json:"witnesses"`
}
type actionLine struct {
Source string `json:"source"`
}
// generatorDispatched reports whether this step drove the app on the action
// generator's behalf, which is the exposure a per-action rate divides by. A
// spec's setup puts the app into its starting position before the generator is
// consulted, so its login taps are dispatched actions that explored nothing.
// Both arms are read the same way, off the source the action names, because a
// denominator that excludes setup on one arm and includes it on the other makes
// the two rates incomparable.
//
// An action naming no source at all is one recorded before the distinction
// existed and cannot be attributed now, so each arm keeps the count it was
// already reported with: everything the seeded picker dispatched, and only what
// the model stamped. summarizeTrace counts those steps separately so a
// pre-source run cannot pass its denominator off as a setup-excluding one.
func generatorDispatched(line traceLine, generator string) bool {
if line.ActionSkipped != "" || line.NextAction == nil {
return false
}
switch line.NextAction.Source {
case "":
return generator != trace.ActionSourceModel
case trace.ActionSourceSetup:
return false
default:
return true
}
}
// actionUnattributed reports a dispatched step whose action names no producer.
func actionUnattributed(line traceLine) bool {
return line.ActionSkipped == "" && line.NextAction != nil && line.NextAction.Source == ""
}
// findRunDirectory returns the run directory `sanderling test` created inside
@@ -68,14 +132,35 @@ func summarizeRun(seedDirectory string) (string, traceSummary, error) {
if err != nil {
return "", traceSummary{}, err
}
summary, err := summarizeTrace(filepath.Join(seedDirectory, name, "trace.jsonl"))
directory := filepath.Join(seedDirectory, name)
generator, err := runGenerator(directory)
if err != nil {
return name, traceSummary{}, err
}
summary, err := summarizeTrace(filepath.Join(directory, "trace.jsonl"), generator)
if err != nil {
return name, traceSummary{}, err
}
return name, summary, nil
}
func summarizeTrace(tracePath string) (traceSummary, error) {
// runGenerator reads which picker drove the run. A trace recorded before
// actions named their source cannot say whether an unstamped action came from
// the seeded picker or from the spec's setup under the model picker, and the
// two count differently.
func runGenerator(runDirectory string) (string, error) {
body, err := os.ReadFile(filepath.Join(runDirectory, "meta.json"))
if err != nil {
return "", fmt.Errorf("read meta: %w", err)
}
var meta trace.Meta
if err := json.Unmarshal(body, &meta); err != nil {
return "", fmt.Errorf("decode meta: %w", err)
}
return meta.Generator, nil
}
func summarizeTrace(tracePath, generator string) (traceSummary, error) {
file, err := os.Open(tracePath)
if err != nil {
return traceSummary{}, fmt.Errorf("open trace: %w", err)
@@ -103,6 +188,15 @@ func summarizeTrace(tracePath string) (traceSummary, error) {
if !synthetic && line.Index > summary.Steps {
summary.Steps = line.Index
}
if !synthetic && generatorDispatched(line, generator) {
summary.Actions++
}
if !synthetic && actionUnattributed(line) {
summary.UnattributedActions++
}
if line.PreconditionFailure != "" {
summary.PreconditionFailures++
}
for _, property := range line.Violations {
violated[property] = true
recordViolation(&summary, line, property)
+304 -1
View File
@@ -22,13 +22,70 @@ func observedStep(index int) trace.Step {
}
}
func actingStep(index int) trace.Step {
step := observedStep(index)
step.NextAction = &trace.Action{Kind: "tap", X: 12, Y: 34}
return step
}
func skippedActionStep(index int, reason string) trace.Step {
step := actingStep(index)
step.ActionSkipped = reason
return step
}
func setupStep(index int) trace.Step {
step := observedStep(index)
step.NextAction = &trace.Action{
Kind: "InputText",
Selector: "testTag:LoginScreen > testTag:LoginEmail",
Text: "[email protected]",
Source: trace.ActionSourceSetup,
}
return step
}
func seededStep(index int) trace.Step {
step := actingStep(index)
step.NextAction.Source = trace.ActionSourceSeeded
return step
}
func modelStep(index int) trace.Step {
step := actingStep(index)
step.NextAction.Source = trace.ActionSourceModel
return step
}
func skippedModelStep(index int, reason string) trace.Step {
step := modelStep(index)
step.ActionSkipped = reason
return step
}
func writeRunDirectory(t *testing.T, seedDirectory, name string, steps []trace.Step) string {
t.Helper()
return writeRunDirectoryWithMeta(
t,
seedDirectory,
name,
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline"},
steps,
)
}
func writeRunDirectoryWithMeta(
t *testing.T,
seedDirectory, name string,
declared trace.Meta,
steps []trace.Step,
) string {
t.Helper()
directory := filepath.Join(seedDirectory, name)
if err := os.MkdirAll(directory, 0o755); err != nil {
t.Fatal(err)
}
meta, err := json.Marshal(trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline"})
meta, err := json.Marshal(declared)
if err != nil {
t.Fatal(err)
}
@@ -72,6 +129,206 @@ func TestSummarizeRun_CleanRunIsCensored(t *testing.T) {
}
}
func TestSummarizeRun_CountsOnlyStepsThatDispatchedAnAction(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
actingStep(1),
observedStep(2),
actingStep(3),
skippedActionStep(4, "no_target"),
skippedActionStep(5, "unresolved_selector"),
skippedActionStep(6, "missing_key"),
skippedActionStep(7, "zero_duration_wait"),
skippedActionStep(8, "app_left_foreground"),
skippedActionStep(9, "apply_error"),
actingStep(10),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Steps != 10 {
t.Errorf("steps: got %d, want 10", summary.Steps)
}
if summary.Actions != 3 {
t.Errorf("actions: got %d, want 3 (one step chose nothing and six were never dispatched)", summary.Actions)
}
}
func TestSummarizeRun_ModelRunLeavesTheSetupsLoginOutOfTheActionCount(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
trace.Meta{Seed: 11, Platform: "web", Arm: "llm-identifier", Generator: "llm"},
[]trace.Step{
setupStep(1),
setupStep(2),
setupStep(3),
modelStep(4),
skippedModelStep(5, "unresolved_selector"),
modelStep(6),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Steps != 6 {
t.Errorf("steps: got %d, want 6", summary.Steps)
}
if summary.Actions != 2 {
t.Errorf("actions: got %d, want 2 (three login steps were setup's and one generator action was thrown away)", summary.Actions)
}
}
func TestSummarizeRun_ModelRunWithoutSetupCountsEveryDispatchedStep(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
trace.Meta{Seed: 11, Platform: "web", Arm: "llm-identifier", Generator: "llm"},
[]trace.Step{modelStep(1), modelStep(2), modelStep(3), modelStep(4)})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 4 {
t.Errorf("actions: got %d, want 4", summary.Actions)
}
}
func TestSummarizeRun_SeededRunLeavesTheSetupsLoginOutOfTheActionCount(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline", Generator: "seeded"},
[]trace.Step{
setupStep(1),
setupStep(2),
setupStep(3),
seededStep(4),
seededStep(5),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 2 {
t.Errorf("actions: got %d, want 2 (three login steps were setup's), which is the same "+
"rule the model arm is counted by", summary.Actions)
}
if summary.UnattributedActions != 0 {
t.Errorf("unattributed actions: got %d, want 0: every step named its source",
summary.UnattributedActions)
}
}
func TestSummarizeRun_SeededRunWithoutSetupCountsEveryDispatchedStep(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline", Generator: "seeded"},
[]trace.Step{seededStep(1), seededStep(2), seededStep(3), seededStep(4)})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 4 {
t.Errorf("actions: got %d, want 4", summary.Actions)
}
}
// Traces recorded before actions named their source cannot be re-attributed
// after the fact, so each arm keeps the count it was already reported with: the
// seeded arm counts every dispatched step, the model arm counts only what the
// model stamped. UnattributedActions is how such a run says so rather than
// passing its old denominator off as a setup-excluding one.
func TestSummarizeRun_TraceWithoutSourcesKeepsTheCountItWasReportedWith(t *testing.T) {
for _, testCase := range []struct {
generator string
actions int
unattributedActions int
}{
{generator: "seeded", actions: 5, unattributedActions: 5},
{generator: "llm", actions: 0, unattributedActions: 5},
} {
t.Run(testCase.generator, func(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
trace.Meta{Seed: 11, Platform: "web", Generator: testCase.generator},
[]trace.Step{
actingStep(1), actingStep(2), actingStep(3), actingStep(4), actingStep(5),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != testCase.actions {
t.Errorf("actions: got %d, want %d", summary.Actions, testCase.actions)
}
if summary.UnattributedActions != testCase.unattributedActions {
t.Errorf("unattributed actions: got %d, want %d",
summary.UnattributedActions, testCase.unattributedActions)
}
})
}
}
func TestSummarizeRun_SkipReasonsEachSuppressTheAction(t *testing.T) {
for _, reason := range []string{
"no_target", "unresolved_selector", "missing_key",
"zero_duration_wait", "app_left_foreground", "apply_error",
} {
seedDirectory := t.TempDir()
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
actingStep(1), skippedActionStep(2, reason),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 1 {
t.Errorf("%s: actions got %d, want 1", reason, summary.Actions)
}
}
}
func TestSummarizeTrace_NullActionIsNoAction(t *testing.T) {
seedDirectory := t.TempDir()
directory := writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{observedStep(1)})
lines := "{\"step\":1,\"hierarchy\":{},\"next_action\":null}\n" +
"{\"step\":2,\"hierarchy\":{},\"next_action\":{\"kind\":\"tap\"},\"action_skipped\":\"\"}\n"
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte(lines), 0o644); err != nil {
t.Fatal(err)
}
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 1 {
t.Errorf("actions: got %d, want 1", summary.Actions)
}
}
func TestSummarizeRun_FinalizeLineIsNotAnAction(t *testing.T) {
seedDirectory := t.TempDir()
finalize := trace.Step{
Index: 3,
Timestamp: time.Now().UTC(),
NextAction: &trace.Action{Kind: "tap"},
Violations: []string{"eventuallySettles"},
}
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{actingStep(1), actingStep(2), finalize})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.Actions != 2 {
t.Errorf("actions: got %d, want 2", summary.Actions)
}
}
func TestSummarizeRun_UsesWitnessOriginNotDetectionStep(t *testing.T) {
seedDirectory := t.TempDir()
violating := observedStep(9)
@@ -198,3 +455,49 @@ func TestSummarizeTrace_MalformedLine(t *testing.T) {
t.Fatal("expected an error for a malformed trace line")
}
}
// A run whose app never came to the foreground has to be countable off the
// summary. Without it the campaign row for a run that never started is a row of
// zero steps and no violations, which is what a clean short run looks like too.
func TestSummarizeRun_CountsPreconditionFailures(t *testing.T) {
seedDirectory := t.TempDir()
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
{Index: 0, PreconditionFailure: "app_not_in_foreground"},
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.PreconditionFailures != 1 {
t.Errorf("precondition failures: got %d, want 1", summary.PreconditionFailures)
}
if summary.Steps != 0 {
t.Errorf("steps: got %d, want 0; the run never observed anything", summary.Steps)
}
}
// The same fact mid-run: steps the scope guard could not bring the app back for
// are still steps, and they are counted separately from the ones that explored.
func TestSummarizeRun_CountsMidRunPreconditionFailures(t *testing.T) {
seedDirectory := t.TempDir()
outsideApp := func(index int) trace.Step {
step := actingStep(index)
step.PreconditionFailure = "app_not_in_foreground"
return step
}
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
actingStep(1), outsideApp(2), outsideApp(3), actingStep(4),
})
_, summary, err := summarizeRun(seedDirectory)
if err != nil {
t.Fatal(err)
}
if summary.PreconditionFailures != 2 {
t.Errorf("precondition failures: got %d, want 2", summary.PreconditionFailures)
}
if summary.Steps != 4 {
t.Errorf("steps: got %d, want 4", summary.Steps)
}
}
@@ -0,0 +1,55 @@
package main
import (
"fmt"
"os"
"regexp"
"strings"
)
// The three models the sample is drawn from, in the capability order
// model-implementations.md fixed before any implementation was generated.
var capabilityOrder = []string{"Sonnet 5", "Opus 5", "Fable 5"}
var (
implementationName = regexp.MustCompile(`^impl-\d+$`)
nonAlphanumeric = regexp.MustCompile(`[^a-z0-9]`)
)
func loadAssignment(path string) (map[string]string, error) {
body, err := os.ReadFile(path)
if err != nil {
return nil, fmt.Errorf("read assignment: %w", err)
}
assignments := map[string]string{}
for _, row := range parseTableRows(string(body)) {
name := row.cell(0)
if !implementationName.MatchString(name) {
continue
}
model, ok := canonicalModel(row.cell(1))
if !ok {
return nil, fmt.Errorf("%s line %d: %s is assigned model %q, which is none of %s",
path, row.Line, name, row.cell(1), strings.Join(capabilityOrder, ", "))
}
if existing, seen := assignments[name]; seen && existing != model {
return nil, fmt.Errorf("%s line %d: %s is assigned to both %s and %s",
path, row.Line, name, existing, model)
}
assignments[name] = model
}
if len(assignments) == 0 {
return nil, fmt.Errorf("%s maps no implementation to a model", path)
}
return assignments, nil
}
func canonicalModel(value string) (string, bool) {
key := nonAlphanumeric.ReplaceAllString(strings.ToLower(value), "")
for _, model := range capabilityOrder {
if key == nonAlphanumeric.ReplaceAllString(strings.ToLower(model), "") {
return model, true
}
}
return "", false
}
@@ -0,0 +1,345 @@
package main
import (
"bufio"
"encoding/json"
"fmt"
"maps"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
)
const (
sweepManifestFileName = "sweep.json"
sweepRecordsFileName = "implementations.jsonl"
campaignManifestFileName = "campaign.json"
campaignRecordsFileName = "runs.jsonl"
traceFileName = "trace.jsonl"
surfacesExtractor = "locatableSurfaces"
maxRecordBytes = 4 * 1024 * 1024
maxTraceLineBytes = 16 * 1024 * 1024
)
// Run exclusion reasons, kept in the vocabulary analyze already uses so the two
// tools describe the same run the same way.
const (
reasonLaunchError = "launch error"
reasonTimedOut = "timed out"
reasonNonzeroExit = "nonzero exit"
reasonTraceError = "unreadable trace"
)
type sweepManifest struct {
SpecPath string `json:"spec_path"`
Implementations []struct {
Name string `json:"name"`
} `json:"implementations"`
}
type sweepRunRecord struct {
Seed int64 `json:"seed"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error"`
CampaignDirectory string `json:"campaign_directory"`
}
type sweepImplementationRecord struct {
Name string `json:"implementation"`
FailedStage string `json:"failed_stage"`
Error string `json:"error"`
Runs []sweepRunRecord `json:"runs"`
}
type campaignRunRecord struct {
Seed int64 `json:"seed"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error"`
TimedOut bool `json:"timed_out"`
TraceError string `json:"trace_error"`
RunDirectory string `json:"run_directory"`
ViolatedProperties []string `json:"violated_properties"`
}
// checkerVerdict is one implementation's whole checker side, pooled across the
// seeds it was swept at.
type checkerVerdict struct {
Implementation string
FailedStage string
FailedError string
RunsRecorded int
RunsUsable int
ExcludedByReason map[string]int
FiredProperties []string
// SurfacesObserved holds every locatable surface seen true on at least one
// step of at least one usable run. A surface missing from it was never
// located across the whole sweep of this implementation.
SurfacesObserved map[string]bool
SurfacesKnown bool
TraceErrors []string
}
func (v checkerVerdict) fired() bool { return len(v.FiredProperties) > 0 }
type checkerSide struct {
Directory string
SpecPath string
Planned []string
Verdicts []checkerVerdict
}
func loadChecker(directory string) (checkerSide, error) {
body, err := os.ReadFile(filepath.Join(directory, sweepManifestFileName))
if err != nil {
return checkerSide{}, fmt.Errorf("read %s: %w", sweepManifestFileName, err)
}
var declared sweepManifest
if err := json.Unmarshal(body, &declared); err != nil {
return checkerSide{}, fmt.Errorf("parse %s in %s: %w", sweepManifestFileName, directory, err)
}
side := checkerSide{Directory: directory, SpecPath: declared.SpecPath}
for _, planned := range declared.Implementations {
side.Planned = append(side.Planned, planned.Name)
}
records, err := readSweepRecords(filepath.Join(directory, sweepRecordsFileName))
if err != nil {
return checkerSide{}, err
}
for _, record := range records {
verdict, err := readImplementation(directory, record)
if err != nil {
return checkerSide{}, err
}
side.Verdicts = append(side.Verdicts, verdict)
}
slices.SortFunc(side.Verdicts, func(a, b checkerVerdict) int {
return strings.Compare(a.Implementation, b.Implementation)
})
return side, nil
}
func readSweepRecords(path string) ([]sweepImplementationRecord, error) {
file, err := os.Open(path)
if err != nil {
return nil, fmt.Errorf("read %s: %w", sweepRecordsFileName, err)
}
defer file.Close()
var records []sweepImplementationRecord
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
lineNumber := 0
seen := map[string]bool{}
for scanner.Scan() {
lineNumber++
raw := strings.TrimSpace(scanner.Text())
if raw == "" {
continue
}
var record sweepImplementationRecord
if err := json.Unmarshal([]byte(raw), &record); err != nil {
return nil, fmt.Errorf("%s line %d: %w", sweepRecordsFileName, lineNumber, err)
}
if record.Name == "" {
return nil, fmt.Errorf("%s line %d names no implementation", sweepRecordsFileName, lineNumber)
}
if seen[record.Name] {
return nil, fmt.Errorf("%s line %d records %s a second time: its runs would be pooled twice",
sweepRecordsFileName, lineNumber, record.Name)
}
seen[record.Name] = true
records = append(records, record)
}
if err := scanner.Err(); err != nil {
return nil, fmt.Errorf("read %s: %w", sweepRecordsFileName, err)
}
return records, nil
}
func readImplementation(sweepDirectory string, record sweepImplementationRecord) (checkerVerdict, error) {
verdict := checkerVerdict{
Implementation: record.Name,
FailedStage: record.FailedStage,
FailedError: record.Error,
ExcludedByReason: map[string]int{},
SurfacesObserved: map[string]bool{},
}
fired := map[string]bool{}
for _, run := range record.Runs {
verdict.RunsRecorded++
if reason := sweepRunExcludedBecause(run); reason != "" {
verdict.ExcludedByReason[reason]++
continue
}
directory := resolveCampaignDirectory(sweepDirectory, record.Name, run)
campaignRuns, err := readCampaignRuns(directory)
if err != nil {
return checkerVerdict{}, err
}
for _, campaignRun := range campaignRuns {
if reason := excludedBecause(campaignRun); reason != "" {
verdict.ExcludedByReason[reason]++
continue
}
verdict.RunsUsable++
for _, property := range campaignRun.ViolatedProperties {
fired[property] = true
}
observed, err := readObservedSurfaces(filepath.Join(directory, campaignRun.RunDirectory))
if err != nil {
verdict.TraceErrors = append(verdict.TraceErrors, err.Error())
continue
}
verdict.SurfacesKnown = true
for surface, seen := range observed {
if seen {
verdict.SurfacesObserved[surface] = true
}
}
}
}
if len(fired) > 0 {
verdict.FiredProperties = slices.Sorted(maps.Keys(fired))
}
if len(verdict.ExcludedByReason) == 0 {
verdict.ExcludedByReason = nil
}
return verdict, nil
}
// resolveCampaignDirectory prefers the path the sweep recorded and falls back to
// the layout it names, so a sweep directory read on another machine than the one
// that wrote its absolute paths still resolves.
func resolveCampaignDirectory(sweepDirectory, name string, run sweepRunRecord) string {
if run.CampaignDirectory != "" {
if _, err := os.Stat(run.CampaignDirectory); err == nil {
return run.CampaignDirectory
}
}
return filepath.Join(sweepDirectory, name, "seed-"+strconv.FormatInt(run.Seed, 10))
}
func readCampaignRuns(directory string) ([]campaignRunRecord, error) {
if _, err := os.Stat(filepath.Join(directory, campaignManifestFileName)); err != nil {
return nil, fmt.Errorf("read %s in %s: %w", campaignManifestFileName, directory, err)
}
file, err := os.Open(filepath.Join(directory, campaignRecordsFileName))
if err != nil {
return nil, fmt.Errorf("read %s: %w", campaignRecordsFileName, err)
}
defer file.Close()
var records []campaignRunRecord
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
lineNumber := 0
for scanner.Scan() {
lineNumber++
raw := strings.TrimSpace(scanner.Text())
if raw == "" {
continue
}
var record campaignRunRecord
if err := json.Unmarshal([]byte(raw), &record); err != nil {
return nil, fmt.Errorf("%s line %d in %s: %w", campaignRecordsFileName, lineNumber, directory, err)
}
records = append(records, record)
}
if err := scanner.Err(); err != nil {
return nil, fmt.Errorf("read %s in %s: %w", campaignRecordsFileName, directory, err)
}
return records, nil
}
// sweepRunExcludedBecause reads the outcome of the campaign process itself,
// which the records inside its directory cannot report. One sweep run is one
// campaign of one seed, so a campaign that died left a runs.jsonl that is
// partial or empty, and scoring the runs it did write reads the seeds it never
// reached as agreement.
func sweepRunExcludedBecause(record sweepRunRecord) string {
switch {
case record.LaunchError != "":
return reasonLaunchError
case record.ExitCode != 0:
return reasonNonzeroExit
default:
return ""
}
}
func excludedBecause(record campaignRunRecord) string {
switch {
case record.LaunchError != "":
return reasonLaunchError
case record.TimedOut:
return reasonTimedOut
case record.ExitCode != 0:
return reasonNonzeroExit
case record.TraceError != "":
return reasonTraceError
default:
return ""
}
}
type traceLine struct {
ExtractorChanges map[string]struct {
Curr json.RawMessage `json:"curr"`
} `json:"extractor_changes"`
}
// readObservedSurfaces replays one run's locatableSurfaces readings. The
// verifier emits every extractor as a change on the first snapshot and only on
// a difference afterwards, so a surface true on any recorded change was located
// at least once, and one absent from every change was never located at all.
func readObservedSurfaces(runDirectory string) (map[string]bool, error) {
path := filepath.Join(runDirectory, traceFileName)
file, err := os.Open(path)
if err != nil {
return nil, fmt.Errorf("read %s: %w", path, err)
}
defer file.Close()
observed := map[string]bool{}
found := false
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 64*1024), maxTraceLineBytes)
lineNumber := 0
for scanner.Scan() {
lineNumber++
raw := strings.TrimSpace(scanner.Text())
if raw == "" {
continue
}
var line traceLine
if err := json.Unmarshal([]byte(raw), &line); err != nil {
return nil, fmt.Errorf("%s line %d: %w", path, lineNumber, err)
}
change, present := line.ExtractorChanges[surfacesExtractor]
if !present || len(change.Curr) == 0 {
continue
}
var reading map[string]bool
if err := json.Unmarshal(change.Curr, &reading); err != nil {
return nil, fmt.Errorf("%s line %d: %s is not an object of booleans: %w",
path, lineNumber, surfacesExtractor, err)
}
found = true
for surface, located := range reading {
if located {
observed[surface] = true
}
}
}
if err := scanner.Err(); err != nil {
return nil, fmt.Errorf("read %s: %w", path, err)
}
if !found {
return nil, fmt.Errorf("%s records no %s reading: this run cannot say whether a surface was located",
path, surfacesExtractor)
}
return observed, nil
}
@@ -0,0 +1,308 @@
package main
import (
"bytes"
"encoding/json"
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
)
// fixtureRun is one campaign of one implementation at one seed, in the shape
// implementation-sweep and campaign write together.
type fixtureRun struct {
Seed int64
ExitCode int
// CampaignExitCode is the campaign process's own exit status, which the
// sweep records beside the campaign directory. It is not the exit status of
// the run inside that campaign.
CampaignExitCode int
TimedOut bool
Violated []string
// Surfaces is the locatableSurfaces reading the trace records. A nil map
// with NoTrace false still writes a reading of every surface false.
Surfaces map[string]bool
NoTrace bool
}
type fixtureReview struct {
Overall string
Clauses map[string]string
Minutes int
}
type fixtureImplementation struct {
Name string
Model string
FailedStage string
Runs []fixtureRun
Review *fixtureReview
RawReview string
Adjudicated map[string]string
}
type fixture struct {
Sweep string
Reviews string
Assignment string
Mapping string
}
const defaultMapping = `
| property | clauses | surfaces |
| --- | --- | --- |
| sentOnlyAfterConfirmation | R5 R15 | stateWords |
| serverHoldsEachMessageOnce | R15 | none |
| unsentReachesZero | R14 | unsentCount |
| surface | never observed | note |
| --- | --- | --- |
| composer | unlocatable | R1 obliges one |
| stateWords | unlocatable | R4 obliges one on every composed message |
| unsentCount | inconclusive | R6 hides the count at zero |
`
func writeFixture(t *testing.T, implementations []fixtureImplementation, mappingBody string) fixture {
t.Helper()
root := t.TempDir()
built := fixture{
Sweep: filepath.Join(root, "sweep"),
Reviews: filepath.Join(root, "reviews"),
Assignment: filepath.Join(root, "assignment.md"),
Mapping: filepath.Join(root, "property-clauses.md"),
}
mustMkdir(t, built.Sweep)
mustMkdir(t, built.Reviews)
mustWrite(t, built.Mapping, mappingBody)
var planned []map[string]any
var assignmentRows []string
assignmentRows = append(assignmentRows, "| implementation | model |", "| --- | --- |")
var records []string
for _, implementation := range implementations {
planned = append(planned, map[string]any{"name": implementation.Name})
if implementation.Model != "" {
assignmentRows = append(assignmentRows,
fmt.Sprintf("| %s | %s |", implementation.Name, implementation.Model))
}
records = append(records, writeImplementation(t, built, implementation))
writeReviewFiles(t, built.Reviews, implementation)
}
mustWrite(t, built.Assignment, strings.Join(assignmentRows, "\n")+"\n")
mustWriteJSON(t, filepath.Join(built.Sweep, sweepManifestFileName), map[string]any{
"generator": "seeded",
"platform": "web",
"spec_path": "paper/experiments/e4/spec/spec.ts",
"max_steps": 400,
"seeds": []int{1, 2},
"host": "anton",
"implementations": planned,
})
mustWrite(t, filepath.Join(built.Sweep, sweepRecordsFileName), strings.Join(records, "\n")+"\n")
return built
}
func writeImplementation(t *testing.T, built fixture, implementation fixtureImplementation) string {
t.Helper()
record := map[string]any{"implementation": implementation.Name}
if implementation.FailedStage != "" {
record["failed_stage"] = implementation.FailedStage
record["error"] = "bun run build exited 1"
return encode(t, record)
}
var runs []map[string]any
for _, run := range implementation.Runs {
seedText := strconv.FormatInt(run.Seed, 10)
campaignDirectory := filepath.Join(built.Sweep, implementation.Name, "seed-"+seedText)
mustMkdir(t, campaignDirectory)
mustWriteJSON(t, filepath.Join(campaignDirectory, campaignManifestFileName), map[string]any{
"arm": implementation.Name,
"generator": "seeded",
"platform": "web",
"max_steps": 400,
"seeds": []int64{run.Seed},
})
runDirectory := filepath.Join("seed-"+seedText, "20260820T110000Z")
campaignRun := map[string]any{
"seed": run.Seed,
"exit_code": run.ExitCode,
"steps": 400,
"actions": 380,
"run_directory": runDirectory,
}
if run.TimedOut {
campaignRun["timed_out"] = true
}
if len(run.Violated) > 0 {
campaignRun["violated_properties"] = run.Violated
}
mustWrite(t, filepath.Join(campaignDirectory, campaignRecordsFileName), encode(t, campaignRun)+"\n")
if !run.NoTrace {
writeTrace(t, filepath.Join(campaignDirectory, runDirectory), run.Surfaces)
}
runs = append(runs, map[string]any{
"seed": run.Seed,
"exit_code": run.CampaignExitCode,
"campaign_directory": campaignDirectory,
})
}
record["runs"] = runs
return encode(t, record)
}
// writeTrace writes the one line the verifier emits on the first snapshot, when
// every extractor is reported as a change from null.
func writeTrace(t *testing.T, directory string, surfaces map[string]bool) {
t.Helper()
mustMkdir(t, directory)
reading := map[string]bool{}
for _, surface := range []string{"appRoot", "composer", "submit", "stateWords",
"unsentCount", "pendingIndicator", "offlineIndicator", "retryControl"} {
reading[surface] = surfaces[surface]
}
line := map[string]any{
"step": 1,
"extractor_changes": map[string]any{
surfacesExtractor: map[string]any{"prev": nil, "curr": reading},
},
}
mustWrite(t, filepath.Join(directory, traceFileName), encode(t, line)+"\n")
}
func writeReviewFiles(t *testing.T, directory string, implementation fixtureImplementation) {
t.Helper()
if implementation.RawReview != "" {
mustWrite(t, filepath.Join(directory, implementation.Name+".md"), implementation.RawReview)
return
}
if implementation.Review == nil {
return
}
mustWrite(t, filepath.Join(directory, implementation.Name+".md"), renderReview(*implementation.Review))
if len(implementation.Adjudicated) == 0 {
return
}
rows := []string{"| clause | verdict | why |", "| --- | --- | --- |"}
for _, clause := range allClauses() {
label, resolved := implementation.Adjudicated[clause]
if !resolved {
continue
}
rows = append(rows, fmt.Sprintf("| %s | %s | joint reread |", clause, label))
}
mustWrite(t, filepath.Join(directory, implementation.Name+"-adjudication.md"),
strings.Join(rows, "\n")+"\n")
}
func renderReview(review fixtureReview) string {
minutes := review.Minutes
if minutes == 0 {
minutes = 45
}
lines := []string{
"# review",
"",
"reviewer: Jane",
"date: 2026-09-02",
fmt.Sprintf("minutes: %d", minutes),
"",
"| clause | verdict | justification | steps |",
"| --- | --- | --- | --- |",
}
for _, clause := range allClauses() {
label := review.Clauses[clause]
if label == "" {
label = clauseMeets
}
lines = append(lines, fmt.Sprintf("| %s | %s | seen by hand | offline, compose, online |", clause, label))
}
lines = append(lines, "", "overall: "+review.Overall, "")
return strings.Join(lines, "\n")
}
func runTool(t *testing.T, built fixture) (result, string) {
t.Helper()
var stdout, stderr bytes.Buffer
jsonPath := filepath.Join(t.TempDir(), "matrix.json")
err := run([]string{
"--sweep", built.Sweep,
"--reviews", built.Reviews,
"--assignment", built.Assignment,
"--property-clauses", built.Mapping,
"--json", jsonPath,
}, &stdout, &stderr)
if err != nil {
t.Fatalf("run: %v\nstderr: %s", err, stderr.String())
}
body, err := os.ReadFile(jsonPath)
if err != nil {
t.Fatalf("read emitted summary: %v", err)
}
var emitted result
if err := json.Unmarshal(body, &emitted); err != nil {
t.Fatalf("emitted summary is not valid JSON: %v", err)
}
return emitted, stdout.String()
}
func cleanRun(seed int64, violated ...string) fixtureRun {
return fixtureRun{
Seed: seed,
Violated: violated,
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": true},
}
}
func mustMkdir(t *testing.T, path string) {
t.Helper()
if err := os.MkdirAll(path, 0o755); err != nil {
t.Fatal(err)
}
}
func mustWrite(t *testing.T, path, body string) {
t.Helper()
mustMkdir(t, filepath.Dir(path))
if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
}
func mustWriteJSON(t *testing.T, path string, value any) {
t.Helper()
mustWrite(t, path, encode(t, value)+"\n")
}
func encode(t *testing.T, value any) string {
t.Helper()
body, err := json.Marshal(value)
if err != nil {
t.Fatal(err)
}
return string(body)
}
func outcomeFor(t *testing.T, emitted result, name string) implementationOutcome {
t.Helper()
for _, row := range emitted.Outcomes {
if row.Implementation == name {
return row
}
}
t.Fatalf("%s carries no cell; excluded as %v", name, emitted.Excluded)
return implementationOutcome{}
}
func exclusionFor(t *testing.T, emitted result, name string) exclusion {
t.Helper()
for _, entry := range emitted.Excluded {
if entry.Implementation == name {
return entry
}
}
t.Fatalf("%s was not excluded; it scored %v", name, emitted.Outcomes)
return exclusion{}
}
+120
View File
@@ -0,0 +1,120 @@
// Command confusion-matrix cross-tabulates e4's checker verdicts against the
// blind human review, which is the measure model-implementations.md
// pre-registers: an implementation whose own suite passed, scored on whether a
// property fired and on whether the reviewer found a defect.
package main
import (
"encoding/json"
"errors"
"flag"
"fmt"
"io"
"os"
"time"
)
const usage = `confusion-matrix cross-tabulates the e4 checker against the blind human review.
Usage:
confusion-matrix --sweep <dir> --reviews <dir> --assignment <path> --property-clauses <path> [--json <path>]
--sweep is the directory implementation-sweep wrote: sweep.json, implementations.jsonl
and the per-implementation campaign directories under it.
--reviews holds one impl-NN.md per implementation in the shape review-protocol.md
fixes, plus any impl-NN-adjudication.md whose resolved labels replace the first
rater's for the clauses it names.
--assignment is implementations/assignment.md, the blinded implementation-to-model
mapping, opened only after the last verdict is filed.
--property-clauses declares which requirement clauses each property covers and which
locatable surfaces it reads. Without it a fired property cannot be scored against a
clause and a portability miss cannot be told from a clean run.
`
func run(arguments []string, stdout, stderr io.Writer) error {
flagSet := flag.NewFlagSet("confusion-matrix", flag.ContinueOnError)
flagSet.SetOutput(stderr)
flagSet.Usage = func() {
fmt.Fprint(stderr, usage)
flagSet.PrintDefaults()
}
var sweepDirectory string
var reviewsDirectory string
var assignmentPath string
var mappingPath string
var jsonPath string
flagSet.StringVar(&sweepDirectory, "sweep", "", "directory implementation-sweep wrote")
flagSet.StringVar(&reviewsDirectory, "reviews", "", "directory holding impl-NN.md verdict forms")
flagSet.StringVar(&assignmentPath, "assignment", "", "implementations/assignment.md, the implementation-to-model mapping")
flagSet.StringVar(&mappingPath, "property-clauses", "", "the declared property-to-clause and surface mapping")
flagSet.StringVar(&jsonPath, "json", "", "write the machine-readable summary here, or - for stdout")
if err := flagSet.Parse(arguments); err != nil {
return err
}
// Every missing flag is named together, in flag order: stopping at the
// first turns one rerun into one rerun per missing flag.
var missing []error
for _, required := range []struct {
name string
value string
}{
{"--sweep", sweepDirectory},
{"--reviews", reviewsDirectory},
{"--assignment", assignmentPath},
{"--property-clauses", mappingPath},
} {
if required.value == "" {
missing = append(missing, fmt.Errorf("%s is required", required.name))
}
}
if err := errors.Join(missing...); err != nil {
return err
}
mapping, err := loadMapping(mappingPath)
if err != nil {
return err
}
assignments, err := loadAssignment(assignmentPath)
if err != nil {
return err
}
checker, err := loadChecker(sweepDirectory)
if err != nil {
return err
}
reviews, err := loadReviews(reviewsDirectory)
if err != nil {
return err
}
result := crossTabulate(checker, reviews, assignments, mapping, time.Now().UTC())
writeReport(result, stdout)
if jsonPath == "" {
return nil
}
body, err := json.MarshalIndent(result, "", " ")
if err != nil {
return fmt.Errorf("marshal summary: %w", err)
}
body = append(body, '\n')
if jsonPath == "-" {
_, err = stdout.Write(body)
return err
}
return os.WriteFile(jsonPath, body, 0o644)
}
func main() {
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
if errors.Is(err, flag.ErrHelp) {
return
}
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
}
@@ -0,0 +1,232 @@
package main
import (
"fmt"
"os"
"regexp"
"slices"
"strings"
)
// clauseCount is R1 to R20, the twenty clauses requirement.md numbers and the
// twenty rows review-protocol.md requires on every verdict form.
const clauseCount = 20
const (
surfaceUnlocatable = "unlocatable"
surfaceInconclusive = "inconclusive"
)
const todoMarker = "todo"
// propertyMapping is one row of the property table: the clauses a property is
// the oracle for, and the locatable surfaces it reads.
type propertyMapping struct {
Property string
Clauses []string
Surfaces []string
}
// surfaceMapping says what a surface never observed across an implementation's
// whole sweep means. Only a surface the requirement obliges every
// implementation to show at all times can be read as unlocatable; a surface
// that is legitimately absent when nothing is in that state is inconclusive
// and can never mark a property unevaluated.
type surfaceMapping struct {
Surface string
NeverObserved string
Note string
}
type mapping struct {
Path string
Properties []propertyMapping
Surfaces []surfaceMapping
PropertyTodoRows int
byProperty map[string]propertyMapping
coveringProperty map[string][]string
unlocatableSurface map[string]bool
}
var clausePattern = regexp.MustCompile(`^[Rr]([0-9]{1,2})$`)
func canonicalClause(value string) (string, bool) {
match := clausePattern.FindStringSubmatch(strings.TrimSpace(value))
if match == nil {
return "", false
}
number := match[1]
trimmed := strings.TrimLeft(number, "0")
if trimmed == "" {
return "", false
}
clause := "R" + trimmed
if !slices.Contains(allClauses(), clause) {
return "", false
}
return clause, true
}
func allClauses() []string {
clauses := make([]string, 0, clauseCount)
for index := 1; index <= clauseCount; index++ {
clauses = append(clauses, fmt.Sprintf("R%d", index))
}
return clauses
}
func loadMapping(path string) (mapping, error) {
body, err := os.ReadFile(path)
if err != nil {
return mapping{}, fmt.Errorf("read property-clause mapping: %w", err)
}
result := mapping{
Path: path,
byProperty: map[string]propertyMapping{},
coveringProperty: map[string][]string{},
unlocatableSurface: map[string]bool{},
}
section := ""
for _, row := range parseTableRows(string(body)) {
head := strings.ToLower(row.cell(0))
switch head {
case "property":
section = "property"
continue
case "surface":
section = "surface"
continue
}
switch section {
case "property":
if err := result.addProperty(row); err != nil {
return mapping{}, fmt.Errorf("%s line %d: %w", path, row.Line, err)
}
case "surface":
if err := result.addSurface(row); err != nil {
return mapping{}, fmt.Errorf("%s line %d: %w", path, row.Line, err)
}
default:
return mapping{}, fmt.Errorf("%s line %d: table row before any header naming property or surface", path, row.Line)
}
}
if len(result.Surfaces) == 0 {
return mapping{}, fmt.Errorf("%s declares no surfaces: a portability miss cannot be told from a clean run without them", path)
}
for _, property := range result.Properties {
for _, surface := range property.Surfaces {
if _, declared := result.surfaceByName(surface); !declared {
return mapping{}, fmt.Errorf("%s: property %q reads surface %q, which the surface table does not declare",
path, property.Property, surface)
}
}
}
return result, nil
}
func (m *mapping) addProperty(row tableRow) error {
name := row.cell(0)
if name == "" {
return nil
}
if strings.EqualFold(name, todoMarker) {
m.PropertyTodoRows++
return nil
}
if _, seen := m.byProperty[name]; seen {
return fmt.Errorf("property %q is mapped twice", name)
}
entry := propertyMapping{Property: name}
for _, item := range splitList(row.cell(1)) {
if strings.EqualFold(item, todoMarker) || item == "-" || strings.EqualFold(item, "none") {
continue
}
clause, ok := canonicalClause(item)
if !ok {
return fmt.Errorf("property %q names clause %q, which is not one of R1 to R%d", name, item, clauseCount)
}
if slices.Contains(entry.Clauses, clause) {
continue
}
entry.Clauses = append(entry.Clauses, clause)
}
for _, item := range splitList(row.cell(2)) {
if strings.EqualFold(item, todoMarker) || item == "-" || strings.EqualFold(item, "none") {
continue
}
if !slices.Contains(entry.Surfaces, item) {
entry.Surfaces = append(entry.Surfaces, item)
}
}
m.Properties = append(m.Properties, entry)
m.byProperty[name] = entry
for _, clause := range entry.Clauses {
m.coveringProperty[clause] = append(m.coveringProperty[clause], name)
}
return nil
}
func (m *mapping) addSurface(row tableRow) error {
name := row.cell(0)
if name == "" || strings.EqualFold(name, todoMarker) {
return nil
}
meaning := strings.ToLower(row.cell(1))
if meaning != surfaceUnlocatable && meaning != surfaceInconclusive {
return fmt.Errorf("surface %q says %q for never observed, want %s or %s",
name, row.cell(1), surfaceUnlocatable, surfaceInconclusive)
}
if _, seen := m.surfaceByName(name); seen {
return fmt.Errorf("surface %q is declared twice", name)
}
m.Surfaces = append(m.Surfaces, surfaceMapping{Surface: name, NeverObserved: meaning, Note: row.cell(2)})
if meaning == surfaceUnlocatable {
m.unlocatableSurface[name] = true
}
return nil
}
func (m mapping) surfaceByName(name string) (surfaceMapping, bool) {
for _, surface := range m.Surfaces {
if surface.Surface == name {
return surface, true
}
}
return surfaceMapping{}, false
}
// unevaluable reports whether a property could not be evaluated against this
// implementation because a surface it reads was never located. A surface the
// mapping calls inconclusive never makes a property unevaluable, however many
// steps failed to observe it.
func (m mapping) unevaluable(property string, observed map[string]bool) bool {
entry, known := m.byProperty[property]
if !known {
return false
}
for _, surface := range entry.Surfaces {
if m.unlocatableSurface[surface] && !observed[surface] {
return true
}
}
return false
}
func (m mapping) covering(clause string) []string {
return m.coveringProperty[clause]
}
func (m mapping) knows(property string) bool {
_, known := m.byProperty[property]
return known
}
func (m mapping) unlocatableSurfaces() []string {
var names []string
for _, surface := range m.Surfaces {
if surface.NeverObserved == surfaceUnlocatable {
names = append(names, surface.Surface)
}
}
return names
}
@@ -0,0 +1,510 @@
package main
import (
"fmt"
"maps"
"slices"
"strings"
"time"
)
// The four cells of the pre-registered measure, named as
// model-implementations.md describes them.
const (
cellTruePositive = "checker fired, review confirmed"
cellFalsePositive = "checker fired, review found nothing"
cellFalseNegative = "checker silent, review found a defect"
cellTrueNegative = "checker silent, review found nothing"
)
// Reasons an implementation is missing data rather than a cell of the matrix.
const (
missingSweepStage = "sweep stopped before any run"
missingNoUsableRun = "no usable run"
missingNoSweepRecord = "no sweep record"
missingNoReview = "no verdict filed"
missingMalformed = "malformed verdict form"
missingSurfaces = "surface locatability unknown"
missingNoModel = "not in the assignment mapping"
)
type matrix struct {
Unit string `json:"unit"`
Scored int `json:"scored"`
TruePositive int `json:"true_positive"`
FalsePositive int `json:"false_positive"`
FalseNegative int `json:"false_negative"`
TrueNegative int `json:"true_negative"`
Precision *float64 `json:"precision"`
Recall *float64 `json:"recall"`
}
func (m *matrix) add(checkerPositive, humanPositive bool) {
m.Scored++
switch {
case checkerPositive && humanPositive:
m.TruePositive++
case checkerPositive:
m.FalsePositive++
case humanPositive:
m.FalseNegative++
default:
m.TrueNegative++
}
}
func (m *matrix) finish() {
m.Precision = ratio(m.TruePositive, m.TruePositive+m.FalsePositive)
m.Recall = ratio(m.TruePositive, m.TruePositive+m.FalseNegative)
}
func ratio(numerator, denominator int) *float64 {
if denominator == 0 {
return nil
}
value := float64(numerator) / float64(denominator)
return &value
}
// clauseMatrix scores one implementation-clause pair. Its three side buckets
// hold the pairs that are not evidence either way: a clause the reviewer could
// not judge, a clause whose every covering property was unevaluable because a
// surface was never located, and, tracked but still scored, a clause no
// property covers at all.
type clauseMatrix struct {
matrix
CannotTell int `json:"cannot_tell"`
UnevaluatedSurfaceMissed int `json:"unevaluated_surface_missed"`
UncoveredScored int `json:"uncovered_clause_pairs_scored"`
UncoveredFalseNegative int `json:"uncovered_clause_false_negatives"`
}
type implementationOutcome struct {
Implementation string `json:"implementation"`
Model string `json:"model"`
Cell string `json:"cell"`
FiredProperties []string `json:"fired_properties,omitempty"`
ViolatedClauses []string `json:"violated_clauses,omitempty"`
CannotTellClauses []string `json:"cannot_tell_clauses,omitempty"`
UnlocatableSurfaces []string `json:"unlocatable_surfaces,omitempty"`
UnevaluatedProperties []string `json:"unevaluated_properties,omitempty"`
// DefectOnlyOnUnevaluatedClauses marks a false negative the checker was
// never in a position to catch: every clause the reviewer faulted is
// covered only by properties a missing surface left unevaluable. It is a
// portability miss reported beside the matrix, never inside it.
DefectOnlyOnUnevaluatedClauses bool `json:"defect_only_on_unevaluated_clauses,omitempty"`
RunsUsable int `json:"runs_usable"`
ReviewMinutes int `json:"review_minutes,omitempty"`
}
type exclusion struct {
Implementation string `json:"implementation"`
Model string `json:"model,omitempty"`
Reason string `json:"reason"`
Detail string `json:"detail,omitempty"`
}
type modelBreakdown struct {
Model string `json:"model"`
Implementations matrix `json:"implementation_matrix"`
Clauses clauseMatrix `json:"clause_matrix"`
Excluded int `json:"excluded"`
DefectOnlyOnUnevaluatedClauses int `json:"defect_only_on_unevaluated_clauses"`
}
type portability struct {
Scored int `json:"implementations_scored"`
WithUnlocatableSurface int `json:"implementations_with_an_unlocatable_surface"`
BySurface map[string]int `json:"implementations_by_unlocatable_surface,omitempty"`
SurfacesReadAsInconclusive []string `json:"surfaces_a_miss_cannot_be_read_from,omitempty"`
}
type coverage struct {
MappedProperties int `json:"mapped_properties"`
MappingTodoRows int `json:"mapping_todo_rows"`
ClausesCovered []string `json:"clauses_covered,omitempty"`
ClausesUncovered []string `json:"clauses_no_property_covers,omitempty"`
FiredPropertiesNotMapped []string `json:"fired_properties_not_in_the_mapping,omitempty"`
}
type result struct {
GeneratedAt time.Time `json:"generated_at"`
SweepDirectory string `json:"sweep_directory"`
ReviewsDirectory string `json:"reviews_directory"`
MappingPath string `json:"property_clause_mapping"`
SpecPath string `json:"spec_path,omitempty"`
Implementations matrix `json:"implementation_matrix"`
Clauses clauseMatrix `json:"clause_matrix"`
ByModel []modelBreakdown `json:"by_model"`
Outcomes []implementationOutcome `json:"outcomes"`
Excluded []exclusion `json:"excluded,omitempty"`
Portability portability `json:"portability"`
ReviewMinutes int `json:"review_minutes_over_scored_implementations"`
Coverage coverage `json:"clause_coverage"`
Notes []string `json:"notes,omitempty"`
}
func crossTabulate(
checker checkerSide,
reviews reviewSide,
assignments map[string]string,
declared mapping,
now time.Time,
) result {
outcome := result{
GeneratedAt: now,
SweepDirectory: checker.Directory,
ReviewsDirectory: reviews.Directory,
MappingPath: declared.Path,
SpecPath: checker.SpecPath,
Implementations: matrix{Unit: "implementation"},
Clauses: clauseMatrix{matrix: matrix{Unit: "implementation-clause pair"}},
Portability: portability{BySurface: map[string]int{}},
}
reviewByName := map[string]reviewVerdict{}
for _, verdict := range reviews.Verdicts {
reviewByName[verdict.Implementation] = verdict
}
malformedByName := map[string]malformedReview{}
for _, entry := range reviews.Malformed {
malformedByName[entry.Implementation] = entry
}
checkerByName := map[string]checkerVerdict{}
for _, verdict := range checker.Verdicts {
checkerByName[verdict.Implementation] = verdict
}
byModel := map[string]*modelBreakdown{}
modelOf := func(name string) string { return assignments[name] }
breakdown := func(model string) *modelBreakdown {
current, seen := byModel[model]
if !seen {
current = &modelBreakdown{
Model: model,
Implementations: matrix{Unit: "implementation"},
Clauses: clauseMatrix{matrix: matrix{Unit: "implementation-clause pair"}},
}
byModel[model] = current
}
return current
}
unmappedFired := map[string]bool{}
for _, name := range implementationNames(checker, reviews, assignments) {
model := modelOf(name)
verdict, swept := checkerByName[name]
review, reviewed := reviewByName[name]
if reason, detail := missingData(name, model, verdict, swept, reviewed, malformedByName); reason != "" {
outcome.Excluded = append(outcome.Excluded, exclusion{
Implementation: name, Model: model, Reason: reason, Detail: detail,
})
if model != "" {
breakdown(model).Excluded++
}
continue
}
for _, property := range verdict.FiredProperties {
if !declared.knows(property) {
unmappedFired[property] = true
}
}
row := scoreImplementation(verdict, review, declared)
row.Model = model
outcome.Outcomes = append(outcome.Outcomes, row)
modelRow := breakdown(model)
outcome.Implementations.add(verdict.fired(), review.defective())
modelRow.Implementations.add(verdict.fired(), review.defective())
scoreClauses(&outcome.Clauses, &modelRow.Clauses, verdict, review, declared)
if row.DefectOnlyOnUnevaluatedClauses {
modelRow.DefectOnlyOnUnevaluatedClauses++
}
outcome.Portability.Scored++
outcome.ReviewMinutes += review.Minutes
if len(row.UnlocatableSurfaces) > 0 {
outcome.Portability.WithUnlocatableSurface++
for _, surface := range row.UnlocatableSurfaces {
outcome.Portability.BySurface[surface]++
}
}
}
outcome.Implementations.finish()
outcome.Clauses.finish()
for _, model := range capabilityOrder {
current, seen := byModel[model]
if !seen {
continue
}
current.Implementations.finish()
current.Clauses.finish()
outcome.ByModel = append(outcome.ByModel, *current)
}
outcome.Coverage = describeCoverage(declared, unmappedFired)
outcome.Portability.SurfacesReadAsInconclusive = inconclusiveSurfaces(declared)
outcome.Notes = buildNotes(outcome, declared, reviews)
return outcome
}
func implementationNames(checker checkerSide, reviews reviewSide, assignments map[string]string) []string {
names := map[string]bool{}
for _, planned := range checker.Planned {
names[planned] = true
}
for _, verdict := range checker.Verdicts {
names[verdict.Implementation] = true
}
for _, verdict := range reviews.Verdicts {
names[verdict.Implementation] = true
}
for _, entry := range reviews.Malformed {
names[entry.Implementation] = true
}
for name := range assignments {
names[name] = true
}
return slices.Sorted(maps.Keys(names))
}
// missingData names why an implementation carries no cell. A build that never
// finished, a run that never produced a usable campaign and a verdict that was
// never filed are all absent evidence: scoring any of them as a clean run would
// read the gap as agreement.
func missingData(
name string,
model string,
verdict checkerVerdict,
swept bool,
reviewed bool,
malformed map[string]malformedReview,
) (string, string) {
if model == "" {
return missingNoModel, ""
}
if !swept {
return missingNoSweepRecord, ""
}
if verdict.FailedStage != "" {
return missingSweepStage, fmt.Sprintf("%s: %s", verdict.FailedStage, verdict.FailedError)
}
if verdict.RunsUsable == 0 {
return missingNoUsableRun, excludedSummary(verdict)
}
if entry, broken := malformed[name]; broken {
return missingMalformed, entry.Reason
}
if !reviewed {
return missingNoReview, ""
}
if !verdict.SurfacesKnown {
return missingSurfaces, strings.Join(verdict.TraceErrors, "; ")
}
return "", ""
}
func excludedSummary(verdict checkerVerdict) string {
if len(verdict.ExcludedByReason) == 0 {
return ""
}
var parts []string
for _, reason := range slices.Sorted(maps.Keys(verdict.ExcludedByReason)) {
parts = append(parts, fmt.Sprintf("%s=%d", reason, verdict.ExcludedByReason[reason]))
}
return strings.Join(parts, ", ")
}
func scoreImplementation(verdict checkerVerdict, review reviewVerdict, declared mapping) implementationOutcome {
row := implementationOutcome{
Implementation: verdict.Implementation,
Cell: cellOf(verdict.fired(), review.defective()),
FiredProperties: verdict.FiredProperties,
ViolatedClauses: review.violatedClauses(),
RunsUsable: verdict.RunsUsable,
ReviewMinutes: review.Minutes,
}
for _, clause := range allClauses() {
if review.Clauses[clause] == clauseCannotTell {
row.CannotTellClauses = append(row.CannotTellClauses, clause)
}
}
for _, surface := range declared.unlocatableSurfaces() {
if !verdict.SurfacesObserved[surface] {
row.UnlocatableSurfaces = append(row.UnlocatableSurfaces, surface)
}
}
for _, property := range declared.Properties {
if declared.unevaluable(property.Property, verdict.SurfacesObserved) {
row.UnevaluatedProperties = append(row.UnevaluatedProperties, property.Property)
}
}
row.DefectOnlyOnUnevaluatedClauses = attributableToUnevaluated(verdict, review, declared)
return row
}
// attributableToUnevaluated reports a false negative the checker could not have
// caught: it fired nothing, the reviewer faulted at least one clause, and every
// clause the reviewer faulted is covered only by properties a missing surface
// left unevaluable.
func attributableToUnevaluated(verdict checkerVerdict, review reviewVerdict, declared mapping) bool {
if verdict.fired() || !review.defective() {
return false
}
violated := review.violatedClauses()
if len(violated) == 0 {
return false
}
for _, clause := range violated {
covering := declared.covering(clause)
if len(covering) == 0 {
return false
}
for _, property := range covering {
if !declared.unevaluable(property, verdict.SurfacesObserved) {
return false
}
}
}
return true
}
func cellOf(checkerPositive, humanPositive bool) string {
switch {
case checkerPositive && humanPositive:
return cellTruePositive
case checkerPositive:
return cellFalsePositive
case humanPositive:
return cellFalseNegative
default:
return cellTrueNegative
}
}
func scoreClauses(overall, model *clauseMatrix, verdict checkerVerdict, review reviewVerdict, declared mapping) {
fired := map[string]bool{}
for _, property := range verdict.FiredProperties {
fired[property] = true
}
for _, clause := range allClauses() {
label := review.Clauses[clause]
if label == clauseCannotTell {
overall.CannotTell++
model.CannotTell++
continue
}
covering := declared.covering(clause)
checkerPositive := false
for _, property := range covering {
if fired[property] {
checkerPositive = true
break
}
}
if !checkerPositive && len(covering) > 0 && allUnevaluable(covering, verdict, declared) {
overall.UnevaluatedSurfaceMissed++
model.UnevaluatedSurfaceMissed++
continue
}
humanPositive := label == clauseViolates
overall.add(checkerPositive, humanPositive)
model.add(checkerPositive, humanPositive)
if len(covering) == 0 {
overall.UncoveredScored++
model.UncoveredScored++
if humanPositive {
overall.UncoveredFalseNegative++
model.UncoveredFalseNegative++
}
}
}
}
func allUnevaluable(covering []string, verdict checkerVerdict, declared mapping) bool {
for _, property := range covering {
if !declared.unevaluable(property, verdict.SurfacesObserved) {
return false
}
}
return true
}
func describeCoverage(declared mapping, unmappedFired map[string]bool) coverage {
result := coverage{
MappedProperties: len(declared.Properties),
MappingTodoRows: declared.PropertyTodoRows,
}
for _, clause := range allClauses() {
if len(declared.covering(clause)) > 0 {
result.ClausesCovered = append(result.ClausesCovered, clause)
continue
}
result.ClausesUncovered = append(result.ClausesUncovered, clause)
}
if len(unmappedFired) > 0 {
result.FiredPropertiesNotMapped = slices.Sorted(maps.Keys(unmappedFired))
}
return result
}
func inconclusiveSurfaces(declared mapping) []string {
var names []string
for _, surface := range declared.Surfaces {
if surface.NeverObserved == surfaceInconclusive {
names = append(names, surface.Surface)
}
}
return names
}
func buildNotes(outcome result, declared mapping, reviews reviewSide) []string {
var notes []string
if len(declared.Properties) == 0 {
notes = append(notes, fmt.Sprintf(
"%s maps no property to a clause, so the clause matrix is empty and no portability miss can be detected; "+
"the implementation matrix below stands on its own", declared.Path))
}
if declared.PropertyTodoRows > 0 {
notes = append(notes, fmt.Sprintf("%s still carries %d TODO row(s) in its property table",
declared.Path, declared.PropertyTodoRows))
}
if len(outcome.Coverage.FiredPropertiesNotMapped) > 0 {
notes = append(notes, fmt.Sprintf(
"%d fired propert(ies) are absent from the mapping and could not be attributed to a clause: %s",
len(outcome.Coverage.FiredPropertiesNotMapped),
strings.Join(outcome.Coverage.FiredPropertiesNotMapped, ", ")))
}
for _, verdict := range reviews.Verdicts {
if !verdict.defective() && len(verdict.violatedClauses()) > 0 {
notes = append(notes, fmt.Sprintf(
"%s files %d violating clause(s) under an overall verdict of %s; the overall verdict is what the matrix scores",
verdict.Implementation, len(verdict.violatedClauses()), overallNotDefective))
}
if len(verdict.Adjudicated) > 0 {
notes = append(notes, fmt.Sprintf("%s uses the adjudicated label for %s",
verdict.Implementation, strings.Join(verdict.Adjudicated, ", ")))
}
}
if count := attributedFalseNegatives(outcome); count > 0 {
notes = append(notes, fmt.Sprintf(
"%d false negative(s) fault only clauses whose every property a missing surface left unevaluable: "+
"those are portability misses, not blind spots", count))
}
notes = append(notes, "second-rater agreement and Cohen's kappa are not computed here; "+
"an impl-NN-adjudication.md is read and its resolved labels replace the first rater's")
return notes
}
func attributedFalseNegatives(outcome result) int {
count := 0
for _, row := range outcome.Outcomes {
if row.DefectOnlyOnUnevaluatedClauses {
count++
}
}
return count
}
@@ -0,0 +1,461 @@
package main
import (
"strings"
"testing"
)
func TestCrossTabulateScoresEachImplementationIntoOneCell(t *testing.T) {
tests := []struct {
name string
implementation fixtureImplementation
wantCell string
wantExcluded string
}{
{
name: "checker and reviewer agree",
implementation: fixtureImplementation{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce"), cleanRun(2)},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
},
wantCell: cellTruePositive,
},
{
name: "checker fired and the reviewer found nothing",
implementation: fixtureImplementation{
Name: "impl-02", Model: "Sonnet 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
Review: &fixtureReview{Overall: overallNotDefective},
},
wantCell: cellFalsePositive,
},
{
name: "reviewer found a defect no property fired on",
implementation: fixtureImplementation{
Name: "impl-03", Model: "Fable 5",
Runs: []fixtureRun{cleanRun(1), cleanRun(2)},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
},
wantCell: cellFalseNegative,
},
{
name: "both clean",
implementation: fixtureImplementation{
Name: "impl-04", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1), cleanRun(2)},
Review: &fixtureReview{Overall: overallNotDefective},
},
wantCell: cellTrueNegative,
},
{
name: "the build never finished",
implementation: fixtureImplementation{
Name: "impl-05", Model: "Sonnet 5", FailedStage: "build",
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R11": clauseViolates}},
},
wantExcluded: missingSweepStage,
},
{
name: "every run timed out",
implementation: fixtureImplementation{
Name: "impl-06", Model: "Fable 5",
Runs: []fixtureRun{{Seed: 1, ExitCode: -1, TimedOut: true}},
Review: &fixtureReview{Overall: overallNotDefective},
},
wantExcluded: missingNoUsableRun,
},
{
name: "no verdict was filed",
implementation: fixtureImplementation{
Name: "impl-07", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
},
wantExcluded: missingNoReview,
},
{
name: "the verdict form is missing a clause",
implementation: fixtureImplementation{
Name: "impl-08", Model: "Sonnet 5",
Runs: []fixtureRun{cleanRun(1)},
RawReview: "reviewer: Jane\n\n| clause | verdict |\n| --- | --- |\n" +
"| R1 | meets |\n| R2 | meets |\n\noverall: not defective\n",
},
wantExcluded: missingMalformed,
},
{
name: "no trace says whether a surface was located",
implementation: fixtureImplementation{
Name: "impl-09", Model: "Fable 5",
Runs: []fixtureRun{{Seed: 1, NoTrace: true}},
Review: &fixtureReview{Overall: overallNotDefective},
},
wantExcluded: missingSurfaces,
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{test.implementation}, defaultMapping))
if test.wantExcluded != "" {
entry := exclusionFor(t, emitted, test.implementation.Name)
if entry.Reason != test.wantExcluded {
t.Fatalf("%s excluded as %q, want %q", test.implementation.Name, entry.Reason, test.wantExcluded)
}
if emitted.Implementations.Scored != 0 {
t.Errorf("missing data scored %d implementation(s), want none: absent evidence is not a clean run",
emitted.Implementations.Scored)
}
if !strings.Contains(stdout, "carry no cell") {
t.Errorf("the report never separates the missing data from the matrix:\n%s", stdout)
}
return
}
row := outcomeFor(t, emitted, test.implementation.Name)
if row.Cell != test.wantCell {
t.Fatalf("%s landed in %q, want %q", test.implementation.Name, row.Cell, test.wantCell)
}
if emitted.Implementations.Scored != 1 {
t.Errorf("scored %d implementation(s), want 1", emitted.Implementations.Scored)
}
})
}
}
// TestACampaignThatDiedIsMissingDataNotATrueNegative covers the sweep it was
// interrupted on: the one seed the campaign got through wrote a clean run
// before the process died, and scoring the implementation on it reads the nine
// seeds that never ran as agreement between the checker and the reviewer.
func TestACampaignThatDiedIsMissingDataNotATrueNegative(t *testing.T) {
interrupted := cleanRun(1)
interrupted.CampaignExitCode = -1
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-07", Model: "Opus 5",
Runs: []fixtureRun{interrupted},
Review: &fixtureReview{Overall: overallNotDefective},
}}, defaultMapping))
entry := exclusionFor(t, emitted, "impl-07")
if entry.Reason != missingNoUsableRun {
t.Fatalf("impl-07 excluded as %q, want %q", entry.Reason, missingNoUsableRun)
}
if !strings.Contains(entry.Detail, reasonNonzeroExit) {
t.Errorf("exclusion detail %q does not name %q", entry.Detail, reasonNonzeroExit)
}
if emitted.Implementations.TrueNegative != 0 || emitted.Implementations.Scored != 0 {
t.Errorf("implementation matrix = %+v, want no cell: a dead campaign is absent evidence",
emitted.Implementations)
}
if !strings.Contains(stdout, "carry no cell") {
t.Errorf("the report never separates the dead campaign from the matrix:\n%s", stdout)
}
}
// TestUnlocatableSurfaceIsNeitherAPositiveNorANegative pins the rule the
// pre-registration turns on: a clause whose every covering property a
// never-located surface left unrunnable is a portability miss. Counting it as a
// true positive credits the oracle for a property that never ran, and counting
// it as a false negative charges the oracle for a defect it was never in a
// position to see.
func TestUnlocatableSurfaceIsNeitherAPositiveNorANegative(t *testing.T) {
built := writeFixture(t, []fixtureImplementation{{
Name: "impl-01",
Model: "Opus 5",
Runs: []fixtureRun{{
Seed: 1,
Violated: []string{"serverHoldsEachMessageOnce"},
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
}},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{
"R5": clauseViolates,
"R15": clauseViolates,
}},
}}, defaultMapping)
emitted, stdout := runTool(t, built)
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1 for R5 behind an unlocated stateWords", got)
}
if got := emitted.Clauses.TruePositive; got != 1 {
t.Errorf("clause matrix reports %d true positive(s), want 1: only R15 had a property that ran and fired", got)
}
if got := emitted.Clauses.FalsePositive; got != 0 {
t.Errorf("clause matrix reports %d false positive(s), want 0", got)
}
if got := emitted.Clauses.FalseNegative; got != 0 {
t.Errorf("clause matrix reports %d false negative(s), want 0: R5 was never evaluated, so it was not missed", got)
}
if got := emitted.Clauses.Scored; got != clauseCount-1 {
t.Errorf("clause matrix scored %d pair(s), want %d: the unevaluated pair is excluded, not scored", got, clauseCount-1)
}
if total := emitted.Clauses.TruePositive + emitted.Clauses.FalsePositive +
emitted.Clauses.FalseNegative + emitted.Clauses.TrueNegative; total != emitted.Clauses.Scored {
t.Errorf("the four cells sum to %d against %d scored: an excluded pair leaked into a cell", total, emitted.Clauses.Scored)
}
row := outcomeFor(t, emitted, "impl-01")
if want := []string{"stateWords"}; !equalStrings(row.UnlocatableSurfaces, want) {
t.Errorf("impl-01 reports unlocatable surfaces %v, want %v", row.UnlocatableSurfaces, want)
}
if want := []string{"sentOnlyAfterConfirmation"}; !equalStrings(row.UnevaluatedProperties, want) {
t.Errorf("impl-01 reports unevaluated properties %v, want %v", row.UnevaluatedProperties, want)
}
if emitted.Portability.WithUnlocatableSurface != 1 {
t.Errorf("portability counts %d implementation(s) needing a locating adaptation, want 1",
emitted.Portability.WithUnlocatableSurface)
}
if !strings.Contains(stdout, "needed a locating adaptation") {
t.Errorf("the report never names the portability count:\n%s", stdout)
}
}
// TestUnlocatableSurfaceOnAMetClauseIsNotATrueNegative is the other half: a
// property that never ran is not evidence the implementation met the clause.
func TestUnlocatableSurfaceOnAMetClauseIsNotATrueNegative(t *testing.T) {
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01",
Model: "Sonnet 5",
Runs: []fixtureRun{{
Seed: 1,
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
}},
Review: &fixtureReview{Overall: overallNotDefective},
}}, defaultMapping))
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1: R5 is covered only by a property "+
"stateWords left unrunnable", got)
}
if got := emitted.Clauses.TrueNegative; got != clauseCount-1 {
t.Errorf("clause matrix reports %d true negative(s), want %d: R5 is excluded rather than credited", got, clauseCount-1)
}
if got := emitted.Clauses.FalsePositive; got != 0 {
t.Errorf("clause matrix reports %d false positive(s), want 0: a property that never ran cannot have fired", got)
}
}
// TestFiredPropertyOutranksAnUnevaluableSibling keeps the exclusion narrow: a
// clause is unevaluated only when every property covering it was unrunnable.
func TestFiredPropertyOutranksAnUnevaluableSibling(t *testing.T) {
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01",
Model: "Fable 5",
Runs: []fixtureRun{{
Seed: 1,
Violated: []string{"serverHoldsEachMessageOnce"},
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
}},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
}}, defaultMapping))
row := outcomeFor(t, emitted, "impl-01")
if row.Cell != cellTruePositive {
t.Fatalf("impl-01 landed in %q, want %q", row.Cell, cellTruePositive)
}
if got := emitted.Clauses.TruePositive; got != 1 {
t.Errorf("R15 produced %d true positive(s), want 1: one covering property ran and fired", got)
}
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1 for R5 alone", got)
}
}
// TestFalseNegativeBehindAnUnlocatedSurfaceIsReportedApart keeps a portability
// miss out of the blind-spot story it would otherwise be read as.
func TestFalseNegativeBehindAnUnlocatedSurfaceIsReportedApart(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01",
Model: "Opus 5",
Runs: []fixtureRun{{
Seed: 1,
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
}},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R5": clauseViolates}},
}}, defaultMapping))
row := outcomeFor(t, emitted, "impl-01")
if row.Cell != cellFalseNegative {
t.Fatalf("impl-01 landed in %q, want %q at the implementation unit", row.Cell, cellFalseNegative)
}
if !row.DefectOnlyOnUnevaluatedClauses {
t.Error("impl-01 faults only R5, whose one property never ran, and is not marked as a portability miss")
}
if got := emitted.Clauses.FalseNegative; got != 0 {
t.Errorf("clause matrix reports %d false negative(s), want 0", got)
}
if !strings.Contains(stdout, "portability misses, not blind spots") {
t.Errorf("the report never separates the portability miss from the blind spot:\n%s", stdout)
}
}
func TestPrecisionRecallAndPerModelBreakdown(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{
{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
},
{
Name: "impl-02", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
Review: &fixtureReview{Overall: overallNotDefective},
},
{
Name: "impl-03", Model: "Sonnet 5",
Runs: []fixtureRun{cleanRun(1)},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
},
{
Name: "impl-04", Model: "Fable 5",
Runs: []fixtureRun{cleanRun(1)},
Review: &fixtureReview{Overall: overallNotDefective},
},
}, defaultMapping))
if emitted.Implementations.Scored != 4 {
t.Fatalf("scored %d implementation(s), want 4", emitted.Implementations.Scored)
}
wantCells := map[string]int{"tp": 1, "fp": 1, "fn": 1, "tn": 1}
got := map[string]int{
"tp": emitted.Implementations.TruePositive,
"fp": emitted.Implementations.FalsePositive,
"fn": emitted.Implementations.FalseNegative,
"tn": emitted.Implementations.TrueNegative,
}
for cell, want := range wantCells {
if got[cell] != want {
t.Errorf("implementation matrix %s = %d, want %d", cell, got[cell], want)
}
}
if emitted.Implementations.Precision == nil || *emitted.Implementations.Precision != 0.5 {
t.Errorf("precision = %v, want 0.5", emitted.Implementations.Precision)
}
if emitted.Implementations.Recall == nil || *emitted.Implementations.Recall != 0.5 {
t.Errorf("recall = %v, want 0.5", emitted.Implementations.Recall)
}
if len(emitted.ByModel) != 3 {
t.Fatalf("broke down %d model(s), want 3", len(emitted.ByModel))
}
if order := []string{emitted.ByModel[0].Model, emitted.ByModel[1].Model, emitted.ByModel[2].Model}; !equalStrings(order, capabilityOrder) {
t.Errorf("models reported in %v, want the pre-registered capability order %v", order, capabilityOrder)
}
for _, model := range emitted.ByModel {
switch model.Model {
case "Opus 5":
if model.Implementations.TruePositive != 1 || model.Implementations.FalsePositive != 1 {
t.Errorf("Opus 5 = %+v, want one true positive and one false positive", model.Implementations)
}
case "Sonnet 5":
if model.Implementations.FalseNegative != 1 {
t.Errorf("Sonnet 5 = %+v, want one false negative", model.Implementations)
}
case "Fable 5":
if model.Implementations.TrueNegative != 1 {
t.Errorf("Fable 5 = %+v, want one true negative", model.Implementations)
}
}
}
if emitted.ReviewMinutes != 180 {
t.Errorf("review cost = %d minutes, want 180: four forms recording 45 each", emitted.ReviewMinutes)
}
for _, want := range []string{"precision", "recall", "Sonnet 5", "Opus 5", "Fable 5",
"review cost over the scored implementations: 180 minutes"} {
if !strings.Contains(stdout, want) {
t.Errorf("the report never prints %q:\n%s", want, stdout)
}
}
}
func TestCannotTellIsEvidenceNeitherWay(t *testing.T) {
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1)},
Review: &fixtureReview{Overall: overallNotDefective, Clauses: map[string]string{
"R17": clauseCannotTell,
"R18": clauseCannotTell,
}},
}}, defaultMapping))
if got := emitted.Clauses.CannotTell; got != 2 {
t.Fatalf("clause matrix reports %d cannot-tell pair(s), want 2", got)
}
if got := emitted.Clauses.Scored; got != clauseCount-2 {
t.Errorf("clause matrix scored %d pair(s), want %d: a specification error is not a true negative", got, clauseCount-2)
}
row := outcomeFor(t, emitted, "impl-01")
if want := []string{"R17", "R18"}; !equalStrings(row.CannotTellClauses, want) {
t.Errorf("impl-01 reports cannot-tell clauses %v, want %v", row.CannotTellClauses, want)
}
}
func TestUncoveredClauseMissIsSeparatedFromACoveredOne(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Fable 5",
Runs: []fixtureRun{cleanRun(1)},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
}}, defaultMapping))
if got := emitted.Clauses.FalseNegative; got != 1 {
t.Fatalf("clause matrix reports %d false negative(s), want 1 for R17", got)
}
if got := emitted.Clauses.UncoveredFalseNegative; got != 1 {
t.Errorf("clause matrix reports %d false negative(s) on a clause no property covers, want 1", got)
}
if want := []string{"R5", "R14", "R15"}; !equalStrings(emitted.Coverage.ClausesCovered, want) {
t.Errorf("coverage reports %v covered, want %v", emitted.Coverage.ClausesCovered, want)
}
if !strings.Contains(stdout, "no property covers R1, R2") {
t.Errorf("the report never names the clauses no property covers:\n%s", stdout)
}
}
func TestAdjudicatedLabelReplacesTheFirstRaters(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseCannotTell}},
Adjudicated: map[string]string{"R15": clauseViolates},
}}, defaultMapping))
if got := emitted.Clauses.TruePositive; got != 1 {
t.Fatalf("clause matrix reports %d true positive(s), want 1: R15 resolved to %s", got, clauseViolates)
}
if got := emitted.Clauses.CannotTell; got != 0 {
t.Errorf("clause matrix reports %d cannot-tell pair(s), want 0: the adjudicated label replaces it", got)
}
if !strings.Contains(stdout, "uses the adjudicated label for R15") {
t.Errorf("the report never says an adjudicated label was used:\n%s", stdout)
}
}
func TestFiredPropertyOutsideTheMappingIsReported(t *testing.T) {
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Sonnet 5",
Runs: []fixtureRun{cleanRun(1, "orderingHoldsAtTheServer")},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
}}, defaultMapping))
if want := []string{"orderingHoldsAtTheServer"}; !equalStrings(emitted.Coverage.FiredPropertiesNotMapped, want) {
t.Fatalf("coverage reports unmapped fired properties %v, want %v", emitted.Coverage.FiredPropertiesNotMapped, want)
}
if row := outcomeFor(t, emitted, "impl-01"); row.Cell != cellTruePositive {
t.Errorf("impl-01 landed in %q, want %q: an unmapped property still fired", row.Cell, cellTruePositive)
}
if !strings.Contains(stdout, "could not be attributed to a clause") {
t.Errorf("the report never flags the unmapped property:\n%s", stdout)
}
}
func equalStrings(got, want []string) bool {
if len(got) != len(want) {
return false
}
for index := range got {
if got[index] != want[index] {
return false
}
}
return true
}
@@ -0,0 +1,247 @@
package main
import (
"bytes"
"os"
"path/filepath"
"strings"
"testing"
)
func runExpectingError(t *testing.T, built fixture) string {
t.Helper()
var stdout, stderr bytes.Buffer
err := run([]string{
"--sweep", built.Sweep,
"--reviews", built.Reviews,
"--assignment", built.Assignment,
"--property-clauses", built.Mapping,
}, &stdout, &stderr)
if err == nil {
t.Fatalf("the tool reported success on input it must refuse:\n%s", stdout.String())
}
return err.Error()
}
func scoredFixture(name, model string) fixtureImplementation {
return fixtureImplementation{
Name: name, Model: model,
Runs: []fixtureRun{cleanRun(1)},
Review: &fixtureReview{Overall: overallNotDefective},
}
}
// Three flags missing is one rerun, not three: the operator is told about all
// of them at once, in flag order, whatever order the check happened to walk.
func TestRunNamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
var stdout, stderr bytes.Buffer
err := run([]string{"--sweep", "s"}, &stdout, &stderr)
if err == nil {
t.Fatal("got no error, want every missing flag named")
}
message := err.Error()
previous := -1
for _, name := range []string{"--reviews", "--assignment", "--property-clauses"} {
at := strings.Index(message, name)
if at < 0 {
t.Fatalf("got %q, want %s named", message, name)
}
if at < previous {
t.Errorf("got %q, want the flags named in flag order", message)
}
previous = at
}
if strings.Contains(message, "--sweep") {
t.Errorf("got %q, want the supplied --sweep left out", message)
}
}
func TestMappingRefusesInputTheMatrixCannotBeScoredFrom(t *testing.T) {
tests := []struct {
name string
mapping string
want string
}{
{
name: "a property reads a surface the surface table never declares",
mapping: "| property | clauses | surfaces |\n| p | R1 | badgeRow |\n" +
"| surface | never observed | note |\n| composer | unlocatable | |\n",
want: "the surface table does not declare",
},
{
name: "no surface is declared at all",
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n",
want: "declares no surfaces",
},
{
name: "a property names a clause outside the requirement",
mapping: "| property | clauses | surfaces |\n| p | R21 | none |\n" +
"| surface | never observed | note |\n| composer | unlocatable | |\n",
want: "which is not one of R1 to R20",
},
{
name: "a surface says something other than what a miss means",
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n" +
"| surface | never observed | note |\n| composer | maybe | |\n",
want: "want unlocatable or inconclusive",
},
{
name: "one property is mapped twice",
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n| p | R2 | none |\n" +
"| surface | never observed | note |\n| composer | unlocatable | |\n",
want: "is mapped twice",
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, test.mapping)
if got := runExpectingError(t, built); !strings.Contains(got, test.want) {
t.Fatalf("error %q does not name the problem %q", got, test.want)
}
})
}
}
func TestAssignmentRefusesAModelTheSampleWasNotDrawnFrom(t *testing.T) {
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, defaultMapping)
mustWrite(t, built.Assignment, "| implementation | model |\n| impl-01 | Opus 4 |\n")
if got := runExpectingError(t, built); !strings.Contains(got, "which is none of Sonnet 5, Opus 5, Fable 5") {
t.Fatalf("error %q does not refuse the unknown model", got)
}
}
func TestSweepRecordingAnImplementationTwiceIsRefused(t *testing.T) {
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, defaultMapping)
path := filepath.Join(built.Sweep, sweepRecordsFileName)
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
mustWrite(t, path, string(body)+string(body))
if got := runExpectingError(t, built); !strings.Contains(got, "a second time") {
t.Fatalf("error %q does not refuse the repeated implementation", got)
}
}
func TestMalformedVerdictFormsAreExcludedWithTheirReason(t *testing.T) {
tests := []struct {
name string
body string
want string
}{
{
name: "a clause carries a word that is not a verdict",
body: reviewWithRow("| R7 | probably fine | |"),
want: "want meets, violates or cannot tell",
},
{
name: "the form closes with no overall verdict",
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
"overall: not defective", "", 1),
want: "no overall verdict",
},
{
name: "the overall verdict is neither answer",
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
"overall: not defective", "overall: mostly ok", 1),
want: "is neither defective nor not defective",
},
{
name: "a clause is filed twice with two labels",
body: renderReview(fixtureReview{Overall: overallNotDefective}) +
"\n| R3 | violates | filed again |\n",
want: "is filed twice",
},
{
name: "a clause row is missing",
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
"| R12 | meets | seen by hand | offline, compose, online |\n", "", 1),
want: "no row for clause R12",
},
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
built := writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1)},
RawReview: test.body,
}}, defaultMapping)
emitted, stdout := runTool(t, built)
entry := exclusionFor(t, emitted, "impl-01")
if entry.Reason != missingMalformed {
t.Fatalf("impl-01 excluded as %q, want %q", entry.Reason, missingMalformed)
}
if !strings.Contains(entry.Detail, test.want) {
t.Errorf("exclusion detail %q does not name %q", entry.Detail, test.want)
}
if !strings.Contains(stdout, missingMalformed) {
t.Errorf("the report never prints the malformed form:\n%s", stdout)
}
if emitted.Implementations.Scored != 0 {
t.Errorf("a malformed form scored %d implementation(s), want none", emitted.Implementations.Scored)
}
})
}
}
func TestAnUnfilledMappingStillReportsTheImplementationMatrix(t *testing.T) {
unfilled := "| property | clauses | surfaces |\n| TODO | TODO | TODO |\n" +
"| surface | never observed | note |\n| stateWords | unlocatable | R4 obliges one |\n"
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Opus 5",
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
}}, unfilled))
if emitted.Implementations.TruePositive != 1 {
t.Fatalf("implementation matrix = %+v, want one true positive", emitted.Implementations)
}
if emitted.Clauses.TruePositive != 0 || emitted.Clauses.FalseNegative != 1 {
t.Errorf("clause matrix = %+v, want every clause uncovered", emitted.Clauses)
}
if emitted.Coverage.MappingTodoRows != 1 {
t.Errorf("coverage reports %d TODO row(s), want 1", emitted.Coverage.MappingTodoRows)
}
if !strings.Contains(stdout, "maps no property to a clause") {
t.Errorf("the report never says the mapping is unfilled:\n%s", stdout)
}
}
func TestAnImplementationOutsideTheAssignmentIsMissingData(t *testing.T) {
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{
scoredFixture("impl-01", "Opus 5"),
{
Name: "impl-02",
Runs: []fixtureRun{cleanRun(1)}, Review: &fixtureReview{Overall: overallNotDefective},
},
}, defaultMapping))
if entry := exclusionFor(t, emitted, "impl-02"); entry.Reason != missingNoModel {
t.Fatalf("impl-02 excluded as %q, want %q", entry.Reason, missingNoModel)
}
}
func TestSurfacesArePooledAcrossTheSeedsOneImplementationWasSweptAt(t *testing.T) {
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
Name: "impl-01", Model: "Fable 5",
Runs: []fixtureRun{
{Seed: 1, Surfaces: map[string]bool{"composer": true}},
{Seed: 2, Surfaces: map[string]bool{"stateWords": true}},
},
Review: &fixtureReview{Overall: overallNotDefective},
}}, defaultMapping))
row := outcomeFor(t, emitted, "impl-01")
if len(row.UnlocatableSurfaces) != 0 {
t.Fatalf("impl-01 reports %v unlocatable, want none: each surface was located on one seed or the other",
row.UnlocatableSurfaces)
}
if emitted.Clauses.UnevaluatedSurfaceMissed != 0 {
t.Errorf("clause matrix reports %d unevaluated pair(s), want none", emitted.Clauses.UnevaluatedSurfaceMissed)
}
}
func reviewWithRow(row string) string {
body := renderReview(fixtureReview{Overall: overallNotDefective})
return strings.Replace(body, "| R7 | meets | seen by hand | offline, compose, online |", row, 1)
}
@@ -0,0 +1,169 @@
package main
import (
"fmt"
"io"
"maps"
"slices"
"strconv"
"strings"
"text/tabwriter"
)
// labelledMatrix is one row of a printed matrix: the whole sample, or one model.
type labelledMatrix struct {
Label string
Value clauseMatrix
Excluded int
}
func writeReport(outcome result, out io.Writer) {
fmt.Fprintln(out, "unit: implementation whose own suite passed, scored on whether a property fired and on the reviewer's overall verdict")
fmt.Fprintln(out, "an implementation with no cell is missing data and is listed separately, never counted as a clean run")
fmt.Fprintln(out)
writeTable(out, []string{"group", "scored", "true pos", "false pos", "false neg", "true neg", "precision", "recall", "no cell"},
func(add func(...string)) {
for _, row := range implementationRows(outcome) {
add(
row.Label,
strconv.Itoa(row.Value.Scored),
strconv.Itoa(row.Value.TruePositive),
strconv.Itoa(row.Value.FalsePositive),
strconv.Itoa(row.Value.FalseNegative),
strconv.Itoa(row.Value.TrueNegative),
formatRatio(row.Value.Precision),
formatRatio(row.Value.Recall),
strconv.Itoa(row.Excluded),
)
}
})
fmt.Fprintln(out)
fmt.Fprintln(out, "clause pairs are an implementation against one of R1 to R20, scored through the declared property-to-clause mapping")
fmt.Fprintln(out, "cannot tell is the reviewer's specification-error answer and is evidence neither way")
fmt.Fprintln(out, "unevaluated is a clause whose every covering property a never-located surface left unrunnable: a portability miss, not a cell")
writeTable(out, []string{"group", "scored", "true pos", "false pos", "false neg", "true neg",
"precision", "recall", "cannot tell", "unevaluated", "uncovered", "uncovered false neg"},
func(add func(...string)) {
for _, row := range clauseRows(outcome) {
add(
row.Label,
strconv.Itoa(row.Value.Scored),
strconv.Itoa(row.Value.TruePositive),
strconv.Itoa(row.Value.FalsePositive),
strconv.Itoa(row.Value.FalseNegative),
strconv.Itoa(row.Value.TrueNegative),
formatRatio(row.Value.Precision),
formatRatio(row.Value.Recall),
strconv.Itoa(row.Value.CannotTell),
strconv.Itoa(row.Value.UnevaluatedSurfaceMissed),
strconv.Itoa(row.Value.UncoveredScored),
strconv.Itoa(row.Value.UncoveredFalseNegative),
)
}
})
fmt.Fprintln(out)
writeTable(out, []string{"implementation", "model", "cell", "usable runs", "fired", "violated clauses", "cannot tell"},
func(add func(...string)) {
for _, row := range outcome.Outcomes {
add(
row.Implementation,
row.Model,
row.Cell,
strconv.Itoa(row.RunsUsable),
joinOrDash(row.FiredProperties),
joinOrDash(row.ViolatedClauses),
strconv.Itoa(len(row.CannotTellClauses)),
)
}
})
if len(outcome.Excluded) > 0 {
fmt.Fprintf(out, "\n%d implementation(s) carry no cell: missing data, never a clean run\n", len(outcome.Excluded))
writeTable(out, []string{"implementation", "model", "reason", "detail"}, func(add func(...string)) {
for _, entry := range outcome.Excluded {
add(entry.Implementation, orDash(entry.Model), entry.Reason, orDash(entry.Detail))
}
})
}
fmt.Fprintf(out, "\nspecification portability over the %d scored implementation(s): %d needed a locating adaptation\n",
outcome.Portability.Scored, outcome.Portability.WithUnlocatableSurface)
if len(outcome.Portability.BySurface) > 0 {
writeTable(out, []string{"surface never located", "implementations"}, func(add func(...string)) {
for _, surface := range slices.Sorted(maps.Keys(outcome.Portability.BySurface)) {
add(surface, strconv.Itoa(outcome.Portability.BySurface[surface]))
}
})
}
if len(outcome.Portability.SurfacesReadAsInconclusive) > 0 {
fmt.Fprintf(out, "a miss on %s cannot be told from that surface being legitimately absent, so neither counts against portability\n",
strings.Join(outcome.Portability.SurfacesReadAsInconclusive, ", "))
}
fmt.Fprintf(out, "\nclause coverage: %d mapped propert(ies) cover %d of %d clauses\n",
outcome.Coverage.MappedProperties, len(outcome.Coverage.ClausesCovered), clauseCount)
if len(outcome.Coverage.ClausesUncovered) > 0 {
fmt.Fprintf(out, "no property covers %s\n", strings.Join(outcome.Coverage.ClausesUncovered, ", "))
}
fmt.Fprintf(out, "review cost over the scored implementations: %d minutes\n", outcome.ReviewMinutes)
for _, note := range outcome.Notes {
fmt.Fprintf(out, "\nnote: %s\n", note)
}
}
func implementationRows(outcome result) []labelledMatrix {
rows := []labelledMatrix{{
Label: "all",
Value: clauseMatrix{matrix: outcome.Implementations},
Excluded: len(outcome.Excluded),
}}
for _, model := range outcome.ByModel {
rows = append(rows, labelledMatrix{
Label: model.Model,
Value: clauseMatrix{matrix: model.Implementations},
Excluded: model.Excluded,
})
}
return rows
}
func clauseRows(outcome result) []labelledMatrix {
rows := []labelledMatrix{{Label: "all", Value: outcome.Clauses}}
for _, model := range outcome.ByModel {
rows = append(rows, labelledMatrix{Label: model.Model, Value: model.Clauses})
}
return rows
}
func writeTable(out io.Writer, header []string, rows func(add func(...string))) {
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
fmt.Fprintln(writer, strings.Join(header, "\t"))
rows(func(cells ...string) {
fmt.Fprintln(writer, strings.Join(cells, "\t"))
})
writer.Flush()
}
func formatRatio(value *float64) string {
if value == nil {
return "n/a"
}
return strconv.FormatFloat(*value, 'f', 3, 64)
}
func joinOrDash(values []string) string {
if len(values) == 0 {
return "-"
}
return strings.Join(values, " ")
}
func orDash(value string) string {
if value == "" {
return "-"
}
return value
}
@@ -0,0 +1,237 @@
package main
import (
"fmt"
"os"
"path/filepath"
"regexp"
"slices"
"strconv"
"strings"
)
const (
clauseMeets = "meets"
clauseViolates = "violates"
clauseCannotTell = "cannot tell"
)
const (
overallDefective = "defective"
overallNotDefective = "not defective"
)
// reviewVerdict is one filed verdict form: the twenty clause rows and the
// overall verdict review-protocol.md requires, with any adjudicated label
// already substituted for the first rater's.
type reviewVerdict struct {
Implementation string
Path string
Reviewer string
Date string
Minutes int
Overall string
Clauses map[string]string
Adjudicated []string
}
func (v reviewVerdict) defective() bool { return v.Overall == overallDefective }
func (v reviewVerdict) violatedClauses() []string {
var violated []string
for _, clause := range allClauses() {
if v.Clauses[clause] == clauseViolates {
violated = append(violated, clause)
}
}
return violated
}
type malformedReview struct {
Implementation string
Path string
Reason string
}
type reviewSide struct {
Directory string
Verdicts []reviewVerdict
Malformed []malformedReview
}
var (
reviewFileName = regexp.MustCompile(`^(impl-\d+)\.md$`)
adjudicationName = regexp.MustCompile(`^(impl-\d+)-adjudication\.md$`)
secondRaterName = regexp.MustCompile(`^impl-\d+-r2\.md$`)
keyValueLine = regexp.MustCompile(`(?i)^\s*[-*]?\s*(reviewer|date|minutes|overall verdict|overall)\s*:\s*(.+?)\s*$`)
leadingWholeNumbers = regexp.MustCompile(`^\d+`)
)
func loadReviews(directory string) (reviewSide, error) {
entries, err := os.ReadDir(directory)
if err != nil {
return reviewSide{}, fmt.Errorf("read reviews: %w", err)
}
side := reviewSide{Directory: directory}
adjudications := map[string]map[string]string{}
var forms []struct {
name string
path string
}
for _, entry := range entries {
if entry.IsDir() {
continue
}
path := filepath.Join(directory, entry.Name())
switch {
case secondRaterName.MatchString(entry.Name()):
continue
case adjudicationName.MatchString(entry.Name()):
name := adjudicationName.FindStringSubmatch(entry.Name())[1]
resolved, err := readAdjudication(path)
if err != nil {
side.Malformed = append(side.Malformed, malformedReview{
Implementation: name, Path: path, Reason: err.Error(),
})
continue
}
adjudications[name] = resolved
case reviewFileName.MatchString(entry.Name()):
forms = append(forms, struct {
name string
path string
}{reviewFileName.FindStringSubmatch(entry.Name())[1], path})
}
}
for _, form := range forms {
verdict, err := readReview(form.name, form.path)
if err != nil {
side.Malformed = append(side.Malformed, malformedReview{
Implementation: form.name, Path: form.path, Reason: err.Error(),
})
continue
}
for clause, label := range adjudications[form.name] {
if verdict.Clauses[clause] != label {
verdict.Adjudicated = append(verdict.Adjudicated, clause)
}
verdict.Clauses[clause] = label
}
slices.Sort(verdict.Adjudicated)
side.Verdicts = append(side.Verdicts, verdict)
}
slices.SortFunc(side.Verdicts, func(a, b reviewVerdict) int {
return strings.Compare(a.Implementation, b.Implementation)
})
slices.SortFunc(side.Malformed, func(a, b malformedReview) int {
return strings.Compare(a.Path, b.Path)
})
return side, nil
}
func readReview(name, path string) (reviewVerdict, error) {
body, err := os.ReadFile(path)
if err != nil {
return reviewVerdict{}, err
}
verdict := reviewVerdict{Implementation: name, Path: path, Clauses: map[string]string{}}
for _, line := range strings.Split(string(body), "\n") {
match := keyValueLine.FindStringSubmatch(normalizeCell(line))
if match == nil {
continue
}
value := strings.TrimSpace(match[2])
switch strings.ToLower(match[1]) {
case "reviewer":
verdict.Reviewer = value
case "date":
verdict.Date = value
case "minutes":
if digits := leadingWholeNumbers.FindString(value); digits != "" {
verdict.Minutes, _ = strconv.Atoi(digits)
}
case "overall", "overall verdict":
overall, ok := canonicalOverall(value)
if !ok {
return reviewVerdict{}, fmt.Errorf("overall verdict %q is neither %s nor %s", value, overallDefective, overallNotDefective)
}
verdict.Overall = overall
}
}
clauses, err := readClauseRows(string(body))
if err != nil {
return reviewVerdict{}, err
}
verdict.Clauses = clauses
for _, clause := range allClauses() {
if _, filed := verdict.Clauses[clause]; !filed {
return reviewVerdict{}, fmt.Errorf("no row for clause %s: the form must carry all %d", clause, clauseCount)
}
}
if verdict.Overall == "" {
return reviewVerdict{}, fmt.Errorf("no overall verdict: the matrix scores the reviewer's own %s or %s", overallDefective, overallNotDefective)
}
return verdict, nil
}
func readClauseRows(body string) (map[string]string, error) {
clauses := map[string]string{}
for _, row := range parseTableRows(body) {
clause, ok := canonicalClause(row.cell(0))
if !ok {
continue
}
label, ok := canonicalClauseVerdict(row.cell(1))
if !ok {
return nil, fmt.Errorf("clause %s has verdict %q, want %s, %s or %s",
clause, row.cell(1), clauseMeets, clauseViolates, clauseCannotTell)
}
if existing, seen := clauses[clause]; seen && existing != label {
return nil, fmt.Errorf("clause %s is filed twice, as %q and %q", clause, existing, label)
}
clauses[clause] = label
}
return clauses, nil
}
func readAdjudication(path string) (map[string]string, error) {
body, err := os.ReadFile(path)
if err != nil {
return nil, err
}
resolved, err := readClauseRows(string(body))
if err != nil {
return nil, err
}
if len(resolved) == 0 {
return nil, fmt.Errorf("no resolved clause rows")
}
return resolved, nil
}
func canonicalClauseVerdict(value string) (string, bool) {
switch strings.ToLower(strings.TrimSpace(value)) {
case clauseMeets, "meet", "met":
return clauseMeets, true
case clauseViolates, "violate", "violated":
return clauseViolates, true
case clauseCannotTell, "cannot-tell", "cannot_tell", "can't tell", "cant tell":
return clauseCannotTell, true
default:
return "", false
}
}
func canonicalOverall(value string) (string, bool) {
cleaned := strings.ToLower(strings.TrimSpace(value))
cleaned = strings.TrimSuffix(cleaned, ".")
switch cleaned {
case overallDefective:
return overallDefective, true
case overallNotDefective, "not-defective", "no defect", "clean":
return overallNotDefective, true
default:
return "", false
}
}
@@ -0,0 +1,79 @@
package main
import (
"regexp"
"strings"
)
// tableRow is one pipe-delimited markdown row with its 1-based line number, so
// a malformed cell can name the line the author has to go and fix.
type tableRow struct {
Line int
Cells []string
}
var separatorCell = regexp.MustCompile(`^:?-{2,}:?$`)
func parseTableRows(body string) []tableRow {
var rows []tableRow
for index, raw := range strings.Split(body, "\n") {
line := strings.TrimSpace(raw)
if !strings.HasPrefix(line, "|") {
continue
}
cells := splitCells(line)
if len(cells) == 0 || isSeparatorRow(cells) {
continue
}
rows = append(rows, tableRow{Line: index + 1, Cells: cells})
}
return rows
}
func splitCells(line string) []string {
trimmed := strings.Trim(line, "|")
parts := strings.Split(trimmed, "|")
cells := make([]string, 0, len(parts))
for _, part := range parts {
cells = append(cells, normalizeCell(part))
}
return cells
}
func isSeparatorRow(cells []string) bool {
for _, cell := range cells {
if !separatorCell.MatchString(cell) {
return false
}
}
return true
}
func normalizeCell(value string) string {
cleaned := strings.ReplaceAll(value, "**", "")
cleaned = strings.ReplaceAll(cleaned, "`", "")
return strings.Join(strings.Fields(cleaned), " ")
}
func (row tableRow) cell(index int) string {
if index >= len(row.Cells) {
return ""
}
return row.Cells[index]
}
var listSeparators = regexp.MustCompile(`[,\s]+`)
func splitList(value string) []string {
trimmed := strings.TrimSpace(value)
if trimmed == "" {
return nil
}
var items []string
for _, item := range listSeparators.Split(trimmed, -1) {
if item != "" {
items = append(items, item)
}
}
return items
}
+174
View File
@@ -0,0 +1,174 @@
package main
import (
"fmt"
"os"
"path/filepath"
"slices"
"strings"
)
// The corpus is tastejs/todomvc. Its examples/ directory holds 48
// implementations of one requirement; five are excluded and the 43 that remain
// are this experiment's population.
//
// The population is named here rather than read from whatever directories
// happen to exist, so a corpus at the wrong commit fails verifyCorpus instead
// of quietly sweeping a different sample and reporting it as this one.
var includedImplementations = []string{
"angular-dart", "angular2", "angular2_es2015", "angularjs", "angularjs_require",
"aurelia", "backbone", "backbone_marionette", "backbone_require", "binding-scala",
"canjs", "canjs_require", "closure", "dijon", "dojo",
"duel", "elm", "emberjs", "enyo_backbone", "exoskeleton",
"jquery", "js_of_ocaml", "jsblocks", "knockback", "knockoutjs",
"knockoutjs_require", "kotlin-react", "lavaca_require", "mithril", "polymer",
"ractive", "react", "react-alt", "react-backbone", "reagent",
"riotjs", "scalajs-react", "typescript-angular", "typescript-backbone", "typescript-react",
"vanilla-es6", "vanillajs", "vue",
}
// excludedImplementations are the five directories under examples/ that the
// corpus survey dropped. They are listed rather than merely omitted so
// verifyCorpus can insist the corpus holds exactly these 48 names: an example
// added or renamed upstream then stops the sweep rather than silently shrinking
// or growing the sample.
var excludedImplementations = []string{
"cujo", "emberjs_require", "firebase-angular", "gwt", "react-hooks",
}
const examplesDirectory = "examples"
// documentPath names the served document for the implementations that do not
// keep an index.html at the root of their example directory.
var documentPath = map[string]string{
"angular-dart": "examples/angular-dart/web/index.html",
"duel": "examples/duel/www/index.html",
}
func documentFor(name string) string {
if override, ok := documentPath[name]; ok {
return override
}
return examplesDirectory + "/" + name + "/index.html"
}
// implementation is one member of the population, with the port it owns for the
// whole sweep. The port is what keeps implementations apart: every one is
// served on its own, so every one is its own web origin and localStorage keeps
// their records in separate partitions. Four pairs in this corpus write the
// same key, and one origin between them is one record between them.
type implementation struct {
Name string
Document string
Port int
}
func (i implementation) Origin() string {
return fmt.Sprintf("http://127.0.0.1:%d", i.Port)
}
func (i implementation) URL() string {
return i.Origin() + "/" + i.Document
}
// verifyCorpus insists the corpus holds exactly the 48 example directories this
// population was drawn from. A sweep against a different checkout would still
// run, and its 43 arms would still be labelled, which is why the check is here
// and not left to whoever reads the results.
func verifyCorpus(corpusRoot string) error {
entries, err := os.ReadDir(filepath.Join(corpusRoot, examplesDirectory))
if err != nil {
return fmt.Errorf("--corpus: %w", err)
}
var found []string
for _, entry := range entries {
if entry.IsDir() {
found = append(found, entry.Name())
}
}
expected := slices.Concat(includedImplementations, excludedImplementations)
slices.Sort(expected)
slices.Sort(found)
if slices.Equal(expected, found) {
return nil
}
var missing, unexpected []string
for _, name := range expected {
if !slices.Contains(found, name) {
missing = append(missing, name)
}
}
for _, name := range found {
if !slices.Contains(expected, name) {
unexpected = append(unexpected, name)
}
}
return fmt.Errorf(
"%s/%s is not the corpus this population was drawn from: missing %v, unexpected %v",
corpusRoot,
examplesDirectory,
missing,
unexpected,
)
}
// selectImplementations resolves --implementations against the population. An
// empty selection is the whole population, which is what a real sweep runs; a
// named subset is for smoke runs and is recorded in the manifest like any other
// intent.
func selectImplementations(selection string) ([]string, error) {
if strings.TrimSpace(selection) == "" {
return slices.Clone(includedImplementations), nil
}
var names []string
seen := map[string]bool{}
for _, part := range strings.Split(selection, ",") {
name := strings.TrimSpace(part)
if name == "" {
return nil, fmt.Errorf("empty implementation in %q", selection)
}
if !slices.Contains(includedImplementations, name) {
return nil, fmt.Errorf(
"%q is not one of the %d implementations in this population",
name,
len(includedImplementations),
)
}
if seen[name] {
return nil, fmt.Errorf("duplicate implementation %q", name)
}
seen[name] = true
names = append(names, name)
}
return names, nil
}
// planImplementations gives every selected implementation its document and its
// own port, in population order so the manifest can name the URL each arm was
// served from before anything has been served.
func planImplementations(
corpusRoot string,
names []string,
basePort int,
) ([]implementation, error) {
if basePort+len(names)-1 > 65535 {
return nil, fmt.Errorf(
"--base-port %d leaves no room for %d implementations",
basePort,
len(names),
)
}
planned := make([]implementation, 0, len(names))
for index, name := range names {
document := documentFor(name)
if _, err := os.Stat(filepath.Join(corpusRoot, filepath.FromSlash(document))); err != nil {
return nil, fmt.Errorf("%s: %w", name, err)
}
planned = append(planned, implementation{
Name: name,
Document: document,
Port: basePort + index,
})
}
return planned, nil
}
@@ -0,0 +1,161 @@
package main
import (
"os"
"path/filepath"
"slices"
"strings"
"testing"
)
// writeCorpus builds a tree with the shape the sweep expects: every example
// directory the population was drawn from, each holding the document that
// implementation is served at.
func writeCorpus(t *testing.T, names []string) string {
t.Helper()
root := t.TempDir()
for _, name := range names {
if err := os.MkdirAll(filepath.Join(root, examplesDirectory, name), 0o755); err != nil {
t.Fatal(err)
}
document := filepath.Join(root, filepath.FromSlash(documentFor(name)))
if err := os.MkdirAll(filepath.Dir(document), 0o755); err != nil {
t.Fatal(err)
}
body := "<!doctype html><title>" + name + "</title><ul class=\"todo-list\"></ul>"
if err := os.WriteFile(document, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
}
return root
}
func wholeCorpus(t *testing.T) string {
t.Helper()
return writeCorpus(
t,
slices.Concat(includedImplementations, excludedImplementations),
)
}
func TestPopulation_IsFortyThreeNamesDisjointFromTheExclusions(t *testing.T) {
if len(includedImplementations) != 43 {
t.Errorf(
"population size: got %d, want 43",
len(includedImplementations),
)
}
seen := map[string]bool{}
for _, name := range includedImplementations {
if seen[name] {
t.Errorf("%q appears twice in the population", name)
}
seen[name] = true
}
for _, name := range excludedImplementations {
if seen[name] {
t.Errorf("%q is both included and excluded", name)
}
}
}
func TestVerifyCorpus_RejectsATreeThatIsNotThisCorpus(t *testing.T) {
if err := verifyCorpus(wholeCorpus(t)); err != nil {
t.Fatalf(
"the corpus this population was drawn from should verify: %v",
err,
)
}
short := writeCorpus(
t,
slices.Concat(includedImplementations[1:], excludedImplementations),
)
err := verifyCorpus(short)
if err == nil {
t.Fatal(
"a corpus missing an implementation swept 42 arms and reported them as 43",
)
}
if !strings.Contains(err.Error(), includedImplementations[0]) {
t.Errorf("the error should name what is missing: %v", err)
}
extra := writeCorpus(
t,
slices.Concat(
includedImplementations,
excludedImplementations,
[]string{"svelte"},
),
)
err = verifyCorpus(extra)
if err == nil {
t.Fatal(
"a corpus with an implementation this population never drew from verified",
)
}
if !strings.Contains(err.Error(), "svelte") {
t.Errorf("the error should name what is unexpected: %v", err)
}
}
func TestDocumentFor_ServesTheOverriddenPathForImplementationsWithoutARootIndex(
t *testing.T,
) {
for name, want := range map[string]string{
"angular-dart": "examples/angular-dart/web/index.html",
"duel": "examples/duel/www/index.html",
"vanillajs": "examples/vanillajs/index.html",
"react": "examples/react/index.html",
} {
if got := documentFor(name); got != want {
t.Errorf("%s document: got %q, want %q", name, got, want)
}
}
}
func TestPlanImplementations_RefusesAnImplementationWhoseDocumentIsMissing(
t *testing.T,
) {
root := wholeCorpus(t)
if err := os.Remove(filepath.Join(root, filepath.FromSlash(documentFor("duel")))); err != nil {
t.Fatal(err)
}
_, err := planImplementations(root, []string{"duel"}, 5400)
if err == nil {
t.Fatal(
"an implementation whose document is missing would be swept as a 404 page",
)
}
if !strings.Contains(err.Error(), "duel") {
t.Errorf("the error should name the implementation: %v", err)
}
}
func TestSelectImplementations_DefaultsToThePopulationAndRejectsNamesOutsideIt(
t *testing.T,
) {
all, err := selectImplementations("")
if err != nil {
t.Fatal(err)
}
if !slices.Equal(all, includedImplementations) {
t.Errorf(
"an empty selection should be the whole population, got %d names",
len(all),
)
}
subset, err := selectImplementations("react, angular2_es2015")
if err != nil {
t.Fatal(err)
}
if !slices.Equal(subset, []string{"react", "angular2_es2015"}) {
t.Errorf("subset: got %v", subset)
}
for _, selection := range []string{"cujo", "svelte", "react,react"} {
if _, err := selectImplementations(selection); err == nil {
t.Errorf("selection %q should have been rejected", selection)
}
}
}
@@ -0,0 +1,468 @@
package main
import (
"bufio"
"bytes"
"encoding/json"
"fmt"
"io"
"math/rand"
"net"
"net/http"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
)
// The test binary doubles as the fetcher the stub campaign uses, so the URL a
// campaign is handed is really requested while the sweep is serving, and what
// came back is on disk for the test to read.
func TestMain(m *testing.M) {
if url := os.Getenv("CORPUS_SWEEP_TEST_FETCH_URL"); url != "" {
recordFetch(url, os.Getenv("CORPUS_SWEEP_TEST_FETCH_LOG"))
return
}
os.Exit(m.Run())
}
func recordFetch(url, logPath string) {
line := ""
response, err := http.Get(url)
if err != nil {
line = fmt.Sprintf("%s -> error %v\n", url, err)
} else {
body, _ := io.ReadAll(response.Body)
response.Body.Close()
line = fmt.Sprintf("%s -> %d %s\n", url, response.StatusCode, body)
}
logFile, err := os.OpenFile(
logPath,
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
0o644,
)
if err != nil {
return
}
logFile.WriteString(line)
logFile.Close()
}
// stubCampaign records the argv it was handed and whether the sweep manifest
// was already on disk when it ran, fetches the URL it was told to drive, writes
// the campaign directory the real tool would write, and fails one seed.
const stubCampaign = `#!/bin/sh
output=""
seed=""
url=""
previous=""
for argument in "$@"; do
case "$previous" in
--output) output="$argument" ;;
--seeds) seed="$argument" ;;
--bundle-id) url="$argument" ;;
esac
previous="$argument"
done
manifest=missing
if [ -f "%[1]s" ]; then manifest=present; fi
echo "manifest=$manifest argv: $*" >> "%[2]s"
CORPUS_SWEEP_TEST_FETCH_URL="$url" CORPUS_SWEEP_TEST_FETCH_LOG="%[3]s" "%[4]s"
mkdir -p "$output"
printf '{"seeds":[%%s]}\n' "$seed" > "$output/campaign.json"
echo "stub campaign seed=$seed url=$url"
case "$output" in
*angular2_es2015/seed-5) exit 1 ;;
esac
exit 0
`
func TestRun_EndToEndAgainstAServedCorpusAndAStubCampaign(t *testing.T) {
root := t.TempDir()
corpus := wholeCorpus(t)
// dojo's document is served but cannot be read, so it stalls at serve and
// the two implementations either side of it still have to reach the
// campaign tool with both seeds.
unreadable := filepath.Join(corpus, filepath.FromSlash(documentFor("dojo")))
if err := os.Chmod(unreadable, 0o000); err != nil {
t.Fatal(err)
}
t.Cleanup(func() { os.Chmod(unreadable, 0o644) })
specPath := filepath.Join(root, "todo.ts")
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
t.Fatal(err)
}
output := filepath.Join(root, "campaigns")
campaignLog := filepath.Join(root, "campaign.log")
fetchLog := filepath.Join(root, "fetch.log")
testBinary, err := filepath.Abs(os.Args[0])
if err != nil {
t.Fatal(err)
}
campaignPath := writeScript(
t,
filepath.Join(root, "stub-campaign"),
fmt.Sprintf(
stubCampaign,
filepath.Join(output, manifestFileName),
campaignLog,
fetchLog,
testBinary,
),
)
sanderlingPath := writeScript(
t,
filepath.Join(root, "stub-sanderling"),
"#!/bin/sh\nexit 0\n",
)
selected := []string{"angular2", "angular2_es2015", "dojo"}
basePort := freePortRange(t, len(selected))
var stdout bytes.Buffer
err = run([]string{
"--corpus", corpus,
"--spec", specPath,
"--implementations", strings.Join(selected, ","),
"--seeds", "4-5",
"--max-steps", "40",
"--duration", "30s",
"--concurrency", "2",
"--base-port", fmt.Sprint(basePort),
"--output", output,
"--campaign", campaignPath,
"--sanderling", sanderlingPath,
}, &stdout, io.Discard)
if err == nil ||
!strings.Contains(err.Error(), "1 of 3 implementations never ran") {
t.Fatalf(
"expected the stalled implementation and the failed campaign to be reported, got %v",
err,
)
}
recorded := readManifest(t, filepath.Join(output, manifestFileName))
if !slices.Equal(recorded.Seeds, []int64{4, 5}) {
t.Errorf("intended seeds: got %v", recorded.Seeds)
}
if len(recorded.Implementations) != len(selected) {
t.Fatalf("intended implementations: got %v", recorded.Implementations)
}
for index, planned := range recorded.Implementations {
if planned.Name != selected[index] {
t.Errorf(
"implementation %d: got %q, want %q",
index,
planned.Name,
selected[index],
)
}
wantPort := basePort + index
wantURL := fmt.Sprintf(
"http://127.0.0.1:%d/examples/%s/index.html",
wantPort,
planned.Name,
)
if planned.Port != wantPort || planned.URL != wantURL {
t.Errorf(
"%s: port %d url %q, want %d %q",
planned.Name,
planned.Port,
planned.URL,
wantPort,
wantURL,
)
}
}
if recorded.Generator != "seeded" || recorded.Platform != "web" ||
recorded.MaxSteps != 40 {
t.Errorf(
"manifest generator/platform/budget: %q/%q/%d",
recorded.Generator,
recorded.Platform,
recorded.MaxSteps,
)
}
if recorded.CorpusRoot != corpus {
t.Errorf(
"manifest corpus root: got %q, want %q",
recorded.CorpusRoot,
corpus,
)
}
campaignLines := readLines(t, campaignLog)
if len(campaignLines) != 4 {
t.Fatalf(
"campaign invocations: got %d, want 4:\n%s",
len(campaignLines),
strings.Join(campaignLines, "\n"),
)
}
seen := map[string]bool{}
for _, line := range campaignLines {
if !strings.HasPrefix(line, "manifest=present") {
t.Errorf(
"a campaign ran before the sweep manifest was written: %q",
line,
)
}
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
arm := argumentValue(arguments, "--arm")
seed := argumentValue(arguments, "--seeds")
port := basePort + slices.Index(selected, arm)
wantBundle := fmt.Sprintf(
"http://127.0.0.1:%d/examples/%s/index.html",
port,
arm,
)
if got := argumentValue(arguments, "--bundle-id"); got != wantBundle {
t.Errorf(
"%s seed %s: bundle id %q, want %q",
arm,
seed,
got,
wantBundle,
)
}
if got := argumentValue(arguments, "--output"); got != filepath.Join(
output,
arm,
"seed-"+seed,
) {
t.Errorf("%s seed %s: campaign output %q", arm, seed, got)
}
if got := argumentValue(arguments, "--sanderling"); got != sanderlingPath {
t.Errorf("%s seed %s: sanderling path %q", arm, seed, got)
}
seen[arm+"/"+seed] = true
}
for _, want := range []string{"angular2/4", "angular2/5", "angular2_es2015/4", "angular2_es2015/5"} {
if !seen[want] {
t.Errorf("%s never reached the campaign tool", want)
}
}
// What the served page actually returned: each port answered with its own
// implementation's document, so no arm was driven against another's.
fetched := readLines(t, fetchLog)
for index, name := range []string{"angular2", "angular2_es2015"} {
want := fmt.Sprintf(
"http://127.0.0.1:%d/examples/%s/index.html -> 200 <!doctype html><title>%s</title>",
basePort+index,
name,
name,
)
matched := 0
for _, line := range fetched {
if strings.HasPrefix(line, want) {
matched++
}
}
if matched != 2 {
t.Errorf(
"%s was served its own document %d times, want 2:\n%s",
name,
matched,
strings.Join(fetched, "\n"),
)
}
}
records := readRecords(t, filepath.Join(output, recordsFileName))
if len(records) != 3 {
t.Fatalf("implementation records: got %d, want 3", len(records))
}
byName := map[string]implementationRecord{}
for _, record := range records {
byName[record.Name] = record
}
stalled := byName["dojo"]
if stalled.FailedStage != stageServe || len(stalled.Runs) != 0 {
t.Errorf(
"dojo: stage %q with %d runs, want a serve failure and no runs",
stalled.FailedStage,
len(stalled.Runs),
)
}
for _, name := range []string{"angular2", "angular2_es2015"} {
record := byName[name]
if record.FailedStage != "" || len(record.Runs) != 2 {
t.Errorf(
"%s: stage %q with %d runs, want no failure and 2 runs",
name,
record.FailedStage,
len(record.Runs),
)
}
if record.MonotonicMillis <= 0 {
t.Errorf(
"%s took %d ms, so nothing timed how long it worked",
name,
record.MonotonicMillis,
)
}
}
if exit := byName["angular2_es2015"].Runs[1].ExitCode; exit != 1 {
t.Errorf("angular2_es2015 seed 5 exit code: got %d, want 1", exit)
}
if exit := byName["angular2_es2015"].Runs[0].ExitCode; exit != 0 {
t.Errorf("angular2_es2015 seed 4 exit code: got %d, want 0", exit)
}
for _, name := range []string{"angular2", "angular2_es2015"} {
for _, seed := range []string{"4", "5"} {
directory := filepath.Join(output, name, "seed-"+seed)
if _, err := os.Stat(filepath.Join(directory, "campaign.json")); err != nil {
t.Errorf(
"%s seed %s: no campaign directory: %v",
name,
seed,
err,
)
}
log, err := os.ReadFile(filepath.Join(directory, "campaign.log"))
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(log), "stub campaign seed="+seed) {
t.Errorf(
"%s seed %s: campaign output was not captured: %q",
name,
seed,
log,
)
}
}
}
client := &http.Client{Timeout: 2 * time.Second}
for offset := range selected {
if response, err := client.Get(fmt.Sprintf("http://127.0.0.1:%d/", basePort+offset)); err == nil {
response.Body.Close()
t.Errorf(
"port %d is still served after the sweep finished",
basePort+offset,
)
}
}
if !strings.Contains(stdout.String(), "failed at serve") {
t.Errorf(
"progress output does not name the stalled implementation: %q",
stdout.String(),
)
}
}
func TestRun_RefusesAnOutputDirectoryThatAlreadyHoldsASweep(t *testing.T) {
root := t.TempDir()
output := filepath.Join(root, "campaigns")
if err := os.MkdirAll(output, 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(output, manifestFileName), []byte("{}"), 0o644); err != nil {
t.Fatal(err)
}
specPath := filepath.Join(root, "todo.ts")
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
t.Fatal(err)
}
err := run([]string{
"--corpus", wholeCorpus(t), "--spec", specPath, "--seeds", "1",
"--max-steps", "10", "--output", output,
}, io.Discard, io.Discard)
if err == nil || !strings.Contains(err.Error(), manifestFileName) {
t.Fatalf(
"two sweeps sharing a directory would interleave their records, got %v",
err,
)
}
}
func writeScript(t *testing.T, path, body string) string {
t.Helper()
if err := os.WriteFile(path, []byte(body), 0o755); err != nil {
t.Fatal(err)
}
return path
}
func readLines(t *testing.T, path string) []string {
t.Helper()
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
var lines []string
scanner := bufio.NewScanner(strings.NewReader(string(body)))
for scanner.Scan() {
if line := strings.TrimSpace(scanner.Text()); line != "" {
lines = append(lines, line)
}
}
return lines
}
func readManifest(t *testing.T, path string) manifest {
t.Helper()
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
var recorded manifest
if err := json.Unmarshal(body, &recorded); err != nil {
t.Fatal(err)
}
return recorded
}
func readRecords(t *testing.T, path string) []implementationRecord {
t.Helper()
var records []implementationRecord
for _, line := range readLines(t, path) {
var record implementationRecord
if err := json.Unmarshal([]byte(line), &record); err != nil {
t.Fatalf("%s: %v", line, err)
}
records = append(records, record)
}
return records
}
func argumentValue(arguments []string, name string) string {
index := slices.Index(arguments, name)
if index < 0 || index+1 >= len(arguments) {
return ""
}
return arguments[index+1]
}
// freePortRange finds count consecutive free ports, which is what the sweep
// hands out: one port per implementation from --base-port upwards.
func freePortRange(t *testing.T, count int) int {
t.Helper()
for range 100 {
base := 20000 + rand.Intn(20000)
if portsAreFree(base, count) {
return base
}
}
t.Fatalf("no run of %d free ports", count)
return 0
}
func portsAreFree(base, count int) bool {
for offset := range count {
listener, err := net.Listen(
"tcp",
fmt.Sprintf("127.0.0.1:%d", base+offset),
)
if err != nil {
return false
}
listener.Close()
}
return true
}
+287
View File
@@ -0,0 +1,287 @@
// Command corpus-sweep runs one specification against every implementation in a
// served corpus of independent implementations of the same requirement. It
// serves each one on its own port, which is what keeps them apart: the corpus
// holds pairs that write the same localStorage key, and one origin shared
// between two of them is one stored record shared between them.
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"os/signal"
"path/filepath"
"strconv"
"syscall"
"time"
"github.com/priyanshujain/sanderling/internal/seedspec"
)
// The generator and the platform are fixed rather than exposed: the
// pre-registration runs the seeded policy against the served corpus, and a
// sweep that could quietly run something else records a comparison nobody made.
const (
generator = "seeded"
platform = "web"
)
// defaultConcurrency is how many implementations are swept at once. fleet.md
// measured eight concurrent web campaigns clean at about 1.1 GB resident each,
// parallel efficiency 0.83 at eight against 0.87 at six, and said to re-measure
// before trusting anything above eight on a contended host. Six sits at the
// better efficiency and leaves the emulator farm that shares the host its
// slots. Serving costs nothing here: the corpus needs no build and no separate
// server process, so a worker is one browser.
const defaultConcurrency = 6
// defaultBasePort starts above the range the model-implementation sweep hands
// out, so the two can run on one host without either being served the other's
// pages.
const defaultBasePort = 5400
type config struct {
corpusRoot string
specPath string
outputDirectory string
implementations []string
seeds []int64
maxSteps int
duration time.Duration
concurrency int
basePort int
campaignPath string
sanderlingPath string
extraArguments []string
}
const usage = `corpus-sweep runs one seeded campaign per implementation of a served corpus.
Usage:
corpus-sweep --corpus <dir> --spec <path> --seeds <spec>
--max-steps <n> --output <dir> [flags]
[-- <sanderling test flags>]
Every implementation is served on its own port, so every one is its own origin
and none can read another's stored state, and every one is swept with the same
seeds and the same step budget. Each run's arm is the implementation's name.
Everything after a bare -- reaches every sanderling test call through the
campaign tool.
`
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
flagSet := flag.NewFlagSet("corpus-sweep", flag.ContinueOnError)
flagSet.SetOutput(stderr)
flagSet.Usage = func() {
fmt.Fprint(stderr, usage)
flagSet.PrintDefaults()
}
var configuration config
var seedSpecification string
var selection string
flagSet.StringVar(
&configuration.corpusRoot,
"corpus",
"",
"root of the checked-out corpus, holding examples/ (required)",
)
flagSet.StringVar(
&configuration.specPath,
"spec",
"",
"path to the property set every implementation is run against (required)",
)
flagSet.StringVar(
&seedSpecification,
"seeds",
"",
"seeds every implementation runs: ranges and lists, e.g. 1-10,20 (required)",
)
flagSet.StringVar(
&selection,
"implementations",
"",
"comma-separated subset of the population to sweep (default: all of it)",
)
flagSet.IntVar(
&configuration.maxSteps,
"max-steps",
0,
"per-run step budget, identical across implementations (required, must be positive)",
)
flagSet.DurationVar(
&configuration.duration,
"duration",
5*time.Minute,
"per-run wall-clock ceiling passed to each campaign",
)
flagSet.IntVar(
&configuration.concurrency,
"concurrency",
defaultConcurrency,
"implementations swept at once",
)
flagSet.IntVar(
&configuration.basePort,
"base-port",
defaultBasePort,
"first port served; each implementation takes the next one in population order",
)
flagSet.StringVar(
&configuration.outputDirectory,
"output",
"",
"campaign tree to create (required)",
)
flagSet.StringVar(
&configuration.campaignPath,
"campaign",
"campaign",
"campaign binary to invoke per implementation and seed",
)
flagSet.StringVar(
&configuration.sanderlingPath,
"sanderling",
"sanderling",
"sanderling binary each campaign invokes",
)
if err := flagSet.Parse(arguments); err != nil {
return config{}, err
}
configuration.extraArguments = flagSet.Args()
// Every missing flag is named together, in flag order: stopping at the
// first turns one rerun into one rerun per missing flag.
var missing []error
for _, required := range []struct {
name string
value string
}{
{"--corpus", configuration.corpusRoot},
{"--spec", configuration.specPath},
{"--seeds", seedSpecification},
{"--output", configuration.outputDirectory},
} {
if required.value == "" {
missing = append(
missing,
fmt.Errorf("%s is required", required.name),
)
}
}
if err := errors.Join(missing...); err != nil {
return config{}, err
}
if configuration.maxSteps <= 0 {
return config{}, fmt.Errorf(
"--max-steps must be positive: every implementation needs the same step budget",
)
}
if configuration.duration <= 0 {
return config{}, fmt.Errorf(
"--duration must be positive: %s",
configuration.duration,
)
}
if configuration.concurrency <= 0 {
return config{}, fmt.Errorf(
"--concurrency must be positive: %d",
configuration.concurrency,
)
}
if configuration.basePort < 1024 || configuration.basePort > 65535 {
return config{}, fmt.Errorf(
"--base-port %d is outside 1024-65535",
configuration.basePort,
)
}
seeds, err := seedspec.Parse(seedSpecification)
if err != nil {
return config{}, fmt.Errorf("--seeds: %w", err)
}
configuration.seeds = seeds
names, err := selectImplementations(selection)
if err != nil {
return config{}, fmt.Errorf("--implementations: %w", err)
}
configuration.implementations = names
for name, value := range map[string]*string{
"--corpus": &configuration.corpusRoot,
"--spec": &configuration.specPath,
"--output": &configuration.outputDirectory,
} {
absolute, err := filepath.Abs(*value)
if err != nil {
return config{}, fmt.Errorf("%s: %w", name, err)
}
*value = absolute
}
return configuration, nil
}
func campaignDirectory(
configuration config,
target implementation,
seed string,
) string {
return filepath.Join(
configuration.outputDirectory,
target.Name,
"seed-"+seed,
)
}
// campaignArguments builds one campaign invocation. The arm is the
// implementation's name, so every run in the tree can be attributed to the
// implementation it came from without reading back which port served it.
func campaignArguments(
configuration config,
target implementation,
seed string,
) []string {
arguments := []string{
"--spec", configuration.specPath,
"--bundle-id", target.URL(),
"--platform", platform,
"--arm", target.Name,
"--generator", generator,
"--max-steps", strconv.Itoa(configuration.maxSteps),
"--duration", configuration.duration.String(),
"--seeds", seed,
"--sanderling", configuration.sanderlingPath,
"--output", campaignDirectory(configuration, target, seed),
}
if len(configuration.extraArguments) > 0 {
arguments = append(arguments, "--")
arguments = append(arguments, configuration.extraArguments...)
}
return arguments
}
func run(arguments []string, stdout, stderr io.Writer) error {
configuration, err := parseArguments(arguments, stderr)
if err != nil {
return err
}
ctx, cancel := signal.NotifyContext(
context.Background(),
os.Interrupt,
syscall.SIGTERM,
)
defer cancel()
return runSweep(ctx, configuration, stdout)
}
func main() {
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
if errors.Is(err, flag.ErrHelp) {
return
}
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
}
@@ -0,0 +1,34 @@
package main
import (
"io"
"strings"
"testing"
)
// Three flags missing is one rerun, not three: the operator is told about all
// of them at once, in flag order, whatever order the check happened to walk.
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
_, err := parseArguments(
[]string{"--spec", "s", "--max-steps", "10"},
io.Discard,
)
if err == nil {
t.Fatal("got no error, want every missing flag named")
}
message := err.Error()
previous := -1
for _, name := range []string{"--corpus", "--seeds", "--output"} {
at := strings.Index(message, name)
if at < 0 {
t.Fatalf("got %q, want %s named", message, name)
}
if at < previous {
t.Errorf("got %q, want the flags named in flag order", message)
}
previous = at
}
if strings.Contains(message, "--spec") {
t.Errorf("got %q, want the supplied --spec left out", message)
}
}
+132
View File
@@ -0,0 +1,132 @@
package main
import (
"encoding/json"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"time"
)
const (
manifestFileName = "sweep.json"
recordsFileName = "implementations.jsonl"
)
// plannedImplementation is one implementation the sweep intends to run, with
// the origin that keeps its stored state its own and the URL every seed is
// driven at.
type plannedImplementation struct {
Name string `json:"name"`
Document string `json:"document"`
Port int `json:"port"`
Origin string `json:"origin"`
URL string `json:"url"`
}
// manifest is sweep.json: what the sweep INTENDED to run, written before the
// first run so a host that dropped an implementation or a seed shows up as a
// missing run rather than as a smaller sample.
type manifest struct {
Generator string `json:"generator"`
Platform string `json:"platform"`
SpecPath string `json:"spec_path"`
CorpusRoot string `json:"corpus_root"`
CorpusCommit string `json:"corpus_commit"`
MaxSteps int `json:"max_steps"`
DurationMillis int64 `json:"duration_millis"`
Seeds []int64 `json:"seeds"`
Implementations []plannedImplementation `json:"implementations"`
Concurrency int `json:"concurrency"`
Host string `json:"host"`
CampaignPath string `json:"campaign_path"`
SanderlingPath string `json:"sanderling_path"`
StartedAt time.Time `json:"started_at"`
}
func buildManifest(
configuration config,
implementations []implementation,
host string,
startedAt time.Time,
) manifest {
planned := make([]plannedImplementation, 0, len(implementations))
for _, target := range implementations {
planned = append(planned, plannedImplementation{
Name: target.Name,
Document: target.Document,
Port: target.Port,
Origin: target.Origin(),
URL: target.URL(),
})
}
return manifest{
Generator: generator,
Platform: platform,
SpecPath: configuration.specPath,
CorpusRoot: configuration.corpusRoot,
CorpusCommit: corpusCommit(configuration.corpusRoot),
MaxSteps: configuration.maxSteps,
DurationMillis: configuration.duration.Milliseconds(),
Seeds: configuration.seeds,
Implementations: planned,
Concurrency: configuration.concurrency,
Host: host,
CampaignPath: configuration.campaignPath,
SanderlingPath: configuration.sanderlingPath,
StartedAt: startedAt,
}
}
// corpusCommit records which checkout was swept. It is empty rather than fatal
// for a corpus that is not a git working tree, because the population check has
// already established the corpus holds the right examples.
func corpusCommit(corpusRoot string) string {
command := exec.Command("git", "-C", corpusRoot, "rev-parse", "HEAD")
output, err := command.Output()
if err != nil {
return ""
}
return strings.TrimSpace(string(output))
}
func writeManifest(directory string, value manifest) error {
body, err := json.MarshalIndent(value, "", " ")
if err != nil {
return fmt.Errorf("marshal manifest: %w", err)
}
return os.WriteFile(
filepath.Join(directory, manifestFileName),
append(body, '\n'),
0o644,
)
}
// runRecord is one campaign, which is one implementation at one seed.
type runRecord struct {
Seed int64 `json:"seed"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error,omitempty"`
CampaignDirectory string `json:"campaign_directory"`
MonotonicMillis int64 `json:"monotonic_millis"`
}
// implementationRecord is one line of implementations.jsonl. FailedStage names
// the step that stopped this implementation, and one that never got served
// carries no runs at all.
type implementationRecord struct {
Name string `json:"implementation"`
Document string `json:"document"`
Port int `json:"port"`
Origin string `json:"origin"`
URL string `json:"url"`
FailedStage string `json:"failed_stage,omitempty"`
Error string `json:"error,omitempty"`
StartedAt time.Time `json:"started_at"`
MonotonicMillis int64 `json:"monotonic_millis"`
Runs []runRecord `json:"runs"`
}
const stageServe = "serve"
@@ -0,0 +1,194 @@
package main
import (
"fmt"
"io"
"net/url"
"os"
"path/filepath"
"strings"
"testing"
)
// collisionPairs are implementations in this corpus that write the same
// localStorage key as each other. Served from one origin they share one stored
// record: in the corpus survey angular2_es2015 crashed at bootstrap on a record
// angular2 had written, which is a violation belonging to no implementation.
var collisionPairs = [][2]string{
{"angular2", "angular2_es2015"},
{"backbone", "backbone_require"},
{"canjs", "canjs_require"},
{"react", "typescript-react"},
}
func originOf(t *testing.T, rawURL string) string {
t.Helper()
parsed, err := url.Parse(rawURL)
if err != nil {
t.Fatalf("parse %q: %v", rawURL, err)
}
return parsed.Scheme + "://" + parsed.Host
}
// The whole population is swept so the assertion covers every implementation
// rather than the four pairs already known to collide: the survey found those
// four, and an unexamined fifth would be just as damaging.
func TestSweep_GivesEveryImplementationAnOriginNoOtherImplementationShares(
t *testing.T,
) {
root := t.TempDir()
corpus := wholeCorpus(t)
specPath := filepath.Join(root, "todo.ts")
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
t.Fatal(err)
}
output := filepath.Join(root, "campaigns")
campaignLog := filepath.Join(root, "campaign.log")
fetchLog := filepath.Join(root, "fetch.log")
testBinary, err := filepath.Abs(os.Args[0])
if err != nil {
t.Fatal(err)
}
campaignPath := writeScript(
t,
filepath.Join(root, "stub-campaign"),
fmt.Sprintf(
stubCampaign,
filepath.Join(output, manifestFileName),
campaignLog,
fetchLog,
testBinary,
),
)
sanderlingPath := writeScript(
t,
filepath.Join(root, "stub-sanderling"),
"#!/bin/sh\nexit 0\n",
)
err = run([]string{
"--corpus", corpus,
"--spec", specPath,
"--seeds", "1",
"--max-steps", "10",
"--duration", "30s",
"--concurrency", "8",
"--base-port", fmt.Sprint(freePortRange(t, len(includedImplementations))),
"--output", output,
"--campaign", campaignPath,
"--sanderling", sanderlingPath,
}, io.Discard, io.Discard)
if err != nil {
t.Fatalf("sweep: %v", err)
}
// What the sweep declared it would serve each implementation from.
recorded := readManifest(t, filepath.Join(output, manifestFileName))
if len(recorded.Implementations) != len(includedImplementations) {
t.Fatalf(
"intended implementations: got %d, want %d",
len(recorded.Implementations),
len(includedImplementations),
)
}
plannedOrigin := map[string]string{}
ownerOfOrigin := map[string]string{}
for _, planned := range recorded.Implementations {
if owner, taken := ownerOfOrigin[planned.Origin]; taken {
t.Errorf(
"%s and %s are both served from %s, so they share one localStorage",
owner,
planned.Name,
planned.Origin,
)
}
ownerOfOrigin[planned.Origin] = planned.Name
plannedOrigin[planned.Name] = planned.Origin
if got := originOf(t, planned.URL); got != planned.Origin {
t.Errorf(
"%s: url %q is not under the origin %q the manifest claims",
planned.Name,
planned.URL,
planned.Origin,
)
}
}
for _, pair := range collisionPairs {
if plannedOrigin[pair[0]] == plannedOrigin[pair[1]] {
t.Errorf(
"%s and %s write the same localStorage key and are both served from %s",
pair[0],
pair[1],
plannedOrigin[pair[0]],
)
}
}
// What actually reached the driver. The web driver parses the origin to
// clear out of --bundle-id, so two arms sharing a bundle-id origin clear
// and repopulate one another's storage however carefully they are labelled.
drivenOrigin := map[string]string{}
for _, line := range readLines(t, campaignLog) {
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
arm := argumentValue(arguments, "--arm")
if arm == "" {
t.Fatalf(
"a campaign ran with no arm, so its runs cannot be attributed: %q",
line,
)
}
drivenOrigin[arm] = originOf(t, argumentValue(arguments, "--bundle-id"))
}
if len(drivenOrigin) != len(includedImplementations) {
t.Fatalf(
"arms that reached the campaign tool: got %d, want %d",
len(drivenOrigin),
len(includedImplementations),
)
}
armOfOrigin := map[string]string{}
for arm, origin := range drivenOrigin {
if other, taken := armOfOrigin[origin]; taken {
t.Errorf(
"arms %s and %s were both driven at %s",
other,
arm,
origin,
)
}
armOfOrigin[origin] = arm
if origin != plannedOrigin[arm] {
t.Errorf(
"%s was driven at %s but the manifest promised %s",
arm,
origin,
plannedOrigin[arm],
)
}
}
// And what each origin answered with, which is the check that the ports
// are not merely distinct but each carries its own implementation.
served := map[string]string{}
for _, line := range readLines(t, fetchLog) {
requested, body, found := strings.Cut(line, " -> 200 ")
if !found {
t.Errorf("a served page did not answer: %q", line)
continue
}
served[originOf(t, requested)] = body
}
for arm, origin := range drivenOrigin {
if want := "<title>" + arm + "</title>"; !strings.Contains(
served[origin],
want,
) {
t.Errorf(
"%s at %s was served %q, which is not its own document",
arm,
origin,
served[origin],
)
}
}
}
+178
View File
@@ -0,0 +1,178 @@
package main
import (
"bytes"
"context"
"errors"
"fmt"
"io"
"io/fs"
"net"
"net/http"
"path"
"path/filepath"
"regexp"
"strings"
"time"
)
const (
readinessTimeout = 5 * time.Second
readinessInterval = 100 * time.Millisecond
)
// stubScriptPath is served from every implementation's own origin in place of a
// dependency that no longer answers.
const stubScriptPath = "/__corpus-sweep__/stub.js"
// deadDependency is polyfill.io, which was shut down after this corpus was
// pinned. One implementation loads it from its index.html; the request fails
// and the application works regardless, but it fails slowly, once per run, on
// every run. The corpus is served from this process, so the reference is
// rewritten to a local no-op on the way out.
var deadDependency = regexp.MustCompile(`https?://polyfill\.io/[^"'\s>]*`)
// staticServer serves the whole corpus tree on one implementation's port. The
// tree rather than the implementation's own directory, because an example that
// asks for a path above itself has to resolve the same way it would in the
// upstream repository. Only one implementation is ever visited on this port, so
// the origin still belongs to it alone.
type staticServer struct {
implementation implementation
listener net.Listener
server *http.Server
}
func startStaticServer(
corpusRoot string,
target implementation,
) (*staticServer, error) {
listener, err := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", target.Port))
if err != nil {
return nil, fmt.Errorf("serve %s: %w", target.Name, err)
}
server := &http.Server{Handler: corpusHandler(corpusRoot)}
running := &staticServer{
implementation: target,
listener: listener,
server: server,
}
go server.Serve(listener)
return running, nil
}
func (s *staticServer) stop() {
s.server.Close()
}
// waitReady confirms the served document answers before any run drives it. A
// wrong document path would otherwise reach the driver as a 404 page, and a
// sweep of 404 pages produces clean runs for every implementation.
//
// Only a transport error is retried. The listener is bound before the sweep
// starts, so a status that is not 200 is the server's answer about this
// document and waiting will not change it.
func (s *staticServer) waitReady(ctx context.Context) error {
url := s.implementation.URL()
client := &http.Client{Timeout: 5 * time.Second}
deadline := time.Now().Add(readinessTimeout)
var lastErr error
for {
if ctx.Err() != nil {
return ctx.Err()
}
response, err := client.Get(url)
if err == nil {
io.Copy(io.Discard, response.Body)
response.Body.Close()
if response.StatusCode == http.StatusOK {
return nil
}
return fmt.Errorf(
"%s answered %d, not the document",
url,
response.StatusCode,
)
}
lastErr = err
if time.Now().After(deadline) {
return fmt.Errorf(
"%s did not answer within %s: %w",
url,
readinessTimeout,
lastErr,
)
}
time.Sleep(readinessInterval)
}
}
func corpusHandler(corpusRoot string) http.Handler {
files := http.FileServer(http.Dir(corpusRoot))
return http.HandlerFunc(
func(writer http.ResponseWriter, request *http.Request) {
cleaned := path.Clean(
"/" + strings.TrimPrefix(request.URL.Path, "/"),
)
if cleaned == stubScriptPath {
writer.Header().Set("Content-Type", "application/javascript")
io.WriteString(
writer,
"/* corpus-sweep: dependency removed upstream */\n",
)
return
}
if !strings.HasSuffix(cleaned, ".html") {
files.ServeHTTP(writer, request)
return
}
body, modified, err := readDocument(corpusRoot, cleaned)
if err != nil {
// Never the file server's fallback: it answers a document it
// cannot read with a 200 directory listing, and a run driven at a
// listing explores nothing and comes back clean.
http.Error(writer, err.Error(), documentStatus(err))
return
}
rewritten := deadDependency.ReplaceAll(body, []byte(stubScriptPath))
http.ServeContent(
writer,
request,
path.Base(cleaned),
modified,
bytes.NewReader(rewritten),
)
},
)
}
func documentStatus(err error) int {
switch {
case errors.Is(err, fs.ErrNotExist):
return http.StatusNotFound
case errors.Is(err, fs.ErrPermission):
return http.StatusForbidden
default:
return http.StatusInternalServerError
}
}
func readDocument(corpusRoot, cleaned string) ([]byte, time.Time, error) {
file, err := http.Dir(corpusRoot).Open(cleaned)
if err != nil {
return nil, time.Time{}, err
}
defer file.Close()
info, err := file.Stat()
if err != nil || info.IsDir() {
return nil, time.Time{}, fmt.Errorf(
"%s is not a document",
filepath.FromSlash(cleaned),
)
}
body, err := io.ReadAll(file)
if err != nil {
return nil, time.Time{}, err
}
return body, info.ModTime(), nil
}
@@ -0,0 +1,214 @@
package main
import (
"context"
"io"
"net/http"
"os"
"path/filepath"
"strings"
"testing"
"time"
)
func get(t *testing.T, url string) (int, string) {
t.Helper()
client := &http.Client{Timeout: 5 * time.Second}
response, err := client.Get(url)
if err != nil {
t.Fatalf("GET %s: %v", url, err)
}
defer response.Body.Close()
body, err := io.ReadAll(response.Body)
if err != nil {
t.Fatal(err)
}
return response.StatusCode, string(body)
}
func TestStaticServer_ServesEachImplementationsDocumentFromItsOwnPort(
t *testing.T,
) {
root := wholeCorpus(t)
planned, err := planImplementations(
root,
[]string{"angular-dart", "duel", "react"},
freePortRange(t, 3),
)
if err != nil {
t.Fatal(err)
}
for _, target := range planned {
server, err := startStaticServer(root, target)
if err != nil {
t.Fatal(err)
}
defer server.stop()
if err := server.waitReady(context.Background()); err != nil {
t.Fatalf("%s: %v", target.Name, err)
}
status, body := get(t, target.URL())
if status != http.StatusOK ||
!strings.Contains(body, "<title>"+target.Name+"</title>") {
t.Errorf(
"%s at %s: got %d %q",
target.Name,
target.URL(),
status,
body,
)
}
}
}
func TestStaticServer_ReplacesTheDependencyThatNoLongerAnswers(t *testing.T) {
root := wholeCorpus(t)
document := filepath.Join(root, filepath.FromSlash(documentFor("aurelia")))
body := `<!doctype html><script src="https://polyfill.io/v3/polyfill.min.js?features=Promise"></script>`
if err := os.WriteFile(document, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
planned, err := planImplementations(
root,
[]string{"aurelia"},
freePortRange(t, 1),
)
if err != nil {
t.Fatal(err)
}
server, err := startStaticServer(root, planned[0])
if err != nil {
t.Fatal(err)
}
defer server.stop()
status, served := get(t, planned[0].URL())
if status != http.StatusOK {
t.Fatalf("document answered %d", status)
}
if strings.Contains(served, "polyfill.io") {
t.Errorf(
"a request to a host that no longer answers reaches the browser once per run: %q",
served,
)
}
if !strings.Contains(served, `src="`+stubScriptPath+`"`) {
t.Errorf(
"the reference was removed rather than pointed at a local no-op: %q",
served,
)
}
stubStatus, stub := get(t, planned[0].Origin()+stubScriptPath)
if stubStatus != http.StatusOK || stub == "" {
t.Errorf("stub script: got %d %q", stubStatus, stub)
}
}
// An example that asks for a path above its own directory has to resolve the
// same way it does upstream, which is why the whole tree is served rather than
// the one directory.
func TestStaticServer_ServesPathsAboveTheImplementationsOwnDirectory(
t *testing.T,
) {
root := wholeCorpus(t)
shared := filepath.Join(root, "node_modules", "todomvc-common")
if err := os.MkdirAll(shared, 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(shared, "base.css"), []byte("body{}"), 0o644); err != nil {
t.Fatal(err)
}
planned, err := planImplementations(
root,
[]string{"react"},
freePortRange(t, 1),
)
if err != nil {
t.Fatal(err)
}
server, err := startStaticServer(root, planned[0])
if err != nil {
t.Fatal(err)
}
defer server.stop()
status, body := get(
t,
planned[0].Origin()+"/node_modules/todomvc-common/base.css",
)
if status != http.StatusOK || body != "body{}" {
t.Errorf("shared asset: got %d %q", status, body)
}
}
func TestStaticServerWaitReady_ReportsADocumentThatDoesNotAnswer(t *testing.T) {
root := wholeCorpus(t)
document := filepath.Join(root, filepath.FromSlash(documentFor("dojo")))
if err := os.Chmod(document, 0o000); err != nil {
t.Fatal(err)
}
t.Cleanup(func() { os.Chmod(document, 0o644) })
planned, err := planImplementations(
root,
[]string{"dojo"},
freePortRange(t, 1),
)
if err != nil {
t.Fatal(err)
}
server, err := startStaticServer(root, planned[0])
if err != nil {
t.Fatal(err)
}
defer server.stop()
started := time.Now()
err = server.waitReady(context.Background())
if err == nil {
t.Fatal(
"a document that does not answer would be swept as an error page and come back clean",
)
}
if elapsed := time.Since(started); elapsed > readinessTimeout {
t.Errorf(
"waited %s for a status that was never going to change",
elapsed,
)
}
}
func TestStaticServerStop_ReleasesThePort(t *testing.T) {
root := wholeCorpus(t)
planned, err := planImplementations(
root,
[]string{"vue"},
freePortRange(t, 1),
)
if err != nil {
t.Fatal(err)
}
server, err := startStaticServer(root, planned[0])
if err != nil {
t.Fatal(err)
}
if err := server.waitReady(context.Background()); err != nil {
t.Fatal(err)
}
server.stop()
client := &http.Client{Timeout: time.Second}
deadline := time.Now().Add(3 * time.Second)
for time.Now().Before(deadline) {
response, err := client.Get(planned[0].URL())
if err != nil {
return
}
response.Body.Close()
time.Sleep(50 * time.Millisecond)
}
t.Fatalf(
"port %d is still served after stop(), so the next sweep cannot bind it",
planned[0].Port,
)
}
+315
View File
@@ -0,0 +1,315 @@
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strconv"
"sync"
"time"
)
// resolveBinaries turns campaign and sanderling into absolute paths before
// anything is served. Each campaign runs from the sweep's own directory, and a
// binary that is missing altogether has to stop the sweep here rather than fail
// once per implementation and seed. Every one that is missing is named
// together, in flag order: stopping at the first turns that single stop into
// one rerun per missing binary.
func resolveBinaries(configuration *config) error {
var missing []error
for _, binary := range []struct {
name string
value *string
}{
{"--campaign", &configuration.campaignPath},
{"--sanderling", &configuration.sanderlingPath},
} {
resolved, err := exec.LookPath(*binary.value)
if err != nil {
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
continue
}
absolute, err := filepath.Abs(resolved)
if err != nil {
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
continue
}
*binary.value = absolute
}
return errors.Join(missing...)
}
type sweep struct {
configuration config
stdout io.Writer
records io.Writer
servers map[string]*staticServer
mutex sync.Mutex
stalled int
failedRuns int
totalRuns int
}
func runSweep(
ctx context.Context,
configuration config,
stdout io.Writer,
) error {
if _, err := os.Stat(filepath.Join(configuration.outputDirectory, manifestFileName)); err == nil {
return fmt.Errorf(
"%s already exists in %s: pick a fresh --output so two sweeps do not share a directory",
manifestFileName,
configuration.outputDirectory,
)
}
if err := verifyCorpus(configuration.corpusRoot); err != nil {
return err
}
implementations, err := planImplementations(
configuration.corpusRoot,
configuration.implementations,
configuration.basePort,
)
if err != nil {
return err
}
if err := resolveBinaries(&configuration); err != nil {
return err
}
if _, err := os.Stat(configuration.specPath); err != nil {
return fmt.Errorf("--spec: %w", err)
}
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
return fmt.Errorf("create sweep dir: %w", err)
}
host, _ := os.Hostname()
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, implementations, host, time.Now().UTC())); err != nil {
return fmt.Errorf("write %s: %w", manifestFileName, err)
}
recordsFile, err := os.OpenFile(
filepath.Join(configuration.outputDirectory, recordsFileName),
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
0o644,
)
if err != nil {
return fmt.Errorf("open %s: %w", recordsFileName, err)
}
defer recordsFile.Close()
running := &sweep{
configuration: configuration,
stdout: stdout,
records: recordsFile,
}
// Every port is bound before any run starts. A port already taken means
// that implementation cannot have an origin of its own, and continuing
// without one is what the separate origins are there to prevent.
if err := running.serveAll(implementations); err != nil {
return err
}
defer running.stopAll()
fmt.Fprintf(
stdout,
"sweep: %d implementations, %d seeds each, %d at a time, %s\n",
len(
implementations,
),
len(configuration.seeds),
configuration.concurrency,
configuration.outputDirectory,
)
running.work(ctx, implementations)
fmt.Fprintf(
stdout,
"sweep complete: %d of %d implementations never ran, %d of %d campaigns failed\n",
running.stalled,
len(implementations),
running.failedRuns,
running.totalRuns,
)
if running.stalled > 0 || running.failedRuns > 0 {
return fmt.Errorf(
"%d of %d implementations never ran and %d of %d campaigns failed; see %s",
running.stalled,
len(implementations),
running.failedRuns,
running.totalRuns,
recordsFileName,
)
}
return nil
}
func (s *sweep) serveAll(implementations []implementation) error {
s.servers = make(map[string]*staticServer, len(implementations))
for _, target := range implementations {
server, err := startStaticServer(s.configuration.corpusRoot, target)
if err != nil {
s.stopAll()
return err
}
s.servers[target.Name] = server
}
return nil
}
func (s *sweep) stopAll() {
for _, server := range s.servers {
server.stop()
}
}
func (s *sweep) work(ctx context.Context, implementations []implementation) {
queue := make(chan implementation, len(implementations))
for _, target := range implementations {
queue <- target
}
close(queue)
workers := min(s.configuration.concurrency, len(implementations))
var waitGroup sync.WaitGroup
for range workers {
waitGroup.Add(1)
go func() {
defer waitGroup.Done()
for target := range queue {
if ctx.Err() != nil {
return
}
s.report(s.runImplementation(ctx, target))
}
}()
}
waitGroup.Wait()
}
// runImplementation carries one implementation through all of its seeds. A
// document that does not answer is recorded and the sweep moves on: one
// implementation must not cost the other forty-two their runs.
func (s *sweep) runImplementation(
ctx context.Context,
target implementation,
) (record implementationRecord) {
record = implementationRecord{
Name: target.Name,
Document: target.Document,
Port: target.Port,
Origin: target.Origin(),
URL: target.URL(),
StartedAt: time.Now().UTC(),
}
started := time.Now()
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
if err := os.MkdirAll(filepath.Join(s.configuration.outputDirectory, target.Name), 0o755); err != nil {
record.FailedStage = stageServe
record.Error = err.Error()
return record
}
if err := s.servers[target.Name].waitReady(ctx); err != nil {
record.FailedStage = stageServe
record.Error = err.Error()
return record
}
for _, seed := range s.configuration.seeds {
if ctx.Err() != nil {
return record
}
record.Runs = append(record.Runs, s.runSeed(ctx, target, seed))
}
return record
}
func (s *sweep) runSeed(
ctx context.Context,
target implementation,
seed int64,
) (record runRecord) {
seedText := strconv.FormatInt(seed, 10)
directory := campaignDirectory(s.configuration, target, seedText)
record = runRecord{Seed: seed, CampaignDirectory: directory}
started := time.Now()
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
if err := os.MkdirAll(directory, 0o755); err != nil {
record.ExitCode = -1
record.LaunchError = err.Error()
return record
}
exitCode, err := runCommand(
ctx,
s.configuration.campaignPath,
campaignArguments(
s.configuration,
target,
seedText,
),
filepath.Join(directory, "campaign.log"),
)
record.ExitCode = exitCode
if err != nil {
record.LaunchError = err.Error()
}
return record
}
func runCommand(
ctx context.Context,
binary string,
arguments []string,
logPath string,
) (int, error) {
logFile, err := os.Create(logPath)
if err != nil {
return -1, err
}
defer logFile.Close()
command := exec.CommandContext(ctx, binary, arguments...)
command.Stdout = logFile
command.Stderr = logFile
err = command.Run()
if err == nil {
return 0, nil
}
var exitError *exec.ExitError
if errors.As(err, &exitError) {
return exitError.ExitCode(), nil
}
return -1, err
}
func (s *sweep) report(record implementationRecord) {
s.mutex.Lock()
defer s.mutex.Unlock()
if record.FailedStage != "" {
s.stalled++
}
s.totalRuns += len(record.Runs)
failed := 0
for _, run := range record.Runs {
if run.ExitCode != 0 {
failed++
}
}
s.failedRuns += failed
if err := json.NewEncoder(s.records).Encode(record); err != nil {
fmt.Fprintf(s.stdout, "warning: %s record: %v\n", record.Name, err)
}
elapsed := time.Duration(record.MonotonicMillis) * time.Millisecond
if record.FailedStage != "" {
fmt.Fprintf(s.stdout, "%s port=%d failed at %s: %s (%s)\n",
record.Name, record.Port, record.FailedStage, record.Error, elapsed)
return
}
fmt.Fprintf(s.stdout, "%s origin=%s campaigns=%d failed=%d elapsed=%s\n",
record.Name, record.Origin, len(record.Runs), failed, elapsed)
}
@@ -0,0 +1,46 @@
package main
import (
"path/filepath"
"strings"
"testing"
)
// Two binaries missing is one rerun, not two: the operator is told about both
// at once, in flag order, whatever order the check happened to walk.
func TestResolveBinaries_NamesEveryMissingBinaryInFlagOrder(t *testing.T) {
configuration := config{
campaignPath: "campaign-that-is-not-installed",
sanderlingPath: "sanderling-that-is-not-installed",
}
err := resolveBinaries(&configuration)
if err == nil {
t.Fatal("got no error, want both missing binaries named")
}
message := err.Error()
campaign := strings.Index(message, "--campaign")
sanderling := strings.Index(message, "--sanderling")
if campaign < 0 || sanderling < 0 {
t.Fatalf("got %q, want both --campaign and --sanderling named", message)
}
if campaign > sanderling {
t.Errorf("got %q, want --campaign named before --sanderling", message)
}
resolved := config{
campaignPath: writeScript(
t,
filepath.Join(t.TempDir(), "stub-campaign"),
"#!/bin/sh\nexit 0\n",
),
sanderlingPath: "sanderling-that-is-not-installed",
}
err = resolveBinaries(&resolved)
if err == nil {
t.Fatal("got no error, want the missing sanderling named")
}
if strings.Contains(err.Error(), "--campaign") {
t.Errorf("got %q, want the campaign that resolved left out", err.Error())
}
}
@@ -0,0 +1,219 @@
package main
import (
"fmt"
"sort"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// Instance is one defect as the draft identifies it across runs: the property
// that reported, the action attributed as the origin of the failed obligation,
// and the screen the witness observed. Two reports sharing all three are the
// same defect seen twice; a run-level count of violated properties cannot say
// that.
type Instance struct {
Property string `json:"property"`
OriginAction string `json:"origin_action"`
WitnessScreen string `json:"witness_screen"`
Runs []string `json:"runs"`
Seeds []int64 `json:"seeds"`
// Reports counts violations folded into this instance, which exceeds the
// run count only if one run reported the same property twice, and the
// latch says it cannot.
Reports int `json:"reports"`
// RedactedOrigin marks a row whose origin action reached the trace with its
// typed value redacted, so `full` keyed it by selector instead. Two runs
// that typed different values into that field are one row here, which makes
// the count of such rows a floor rather than a total.
RedactedOrigin bool `json:"redacted_origin,omitempty"`
}
// Unattributed is a violation that carries no origin, so the identity rule
// cannot be applied to it. It is reported rather than counted, because
// dropping it understates the defect count and guessing an origin invents one.
type Unattributed struct {
Property string `json:"property"`
Run string `json:"run"`
Step int `json:"step"`
Reason string `json:"reason"`
}
// actionKeyMode selects how much of an action two reports must share to be the
// same origin. The draft names the origin action and does not say which of its
// fields identify it, so the strict reading is available beside the default.
type actionKeyMode string
const (
// bySelector keys an action by what it did and the name of what it did it
// to. Coordinates and generated text differ between two runs that took the
// same action against the same control.
bySelector actionKeyMode = "selector"
// byFullAction adds the text typed and the coordinates dispatched.
byFullAction actionKeyMode = "full"
)
type identityKey struct {
property string
action string
screen string
}
// Corpus is what one invocation read: the instances, the violations it could
// not attribute, and the counts the report needs.
type Corpus struct {
Runs int `json:"runs"`
Instances []Instance `json:"instances"`
Unattributed []Unattributed `json:"unattributed,omitempty"`
UnnamedScreen int `json:"unnamed_witness_screens,omitempty"`
}
// Singletons counts instances that appeared in exactly one run, the number the
// evaluation reports beside the defect count.
func (c Corpus) Singletons() int {
count := 0
for _, instance := range c.Instances {
if len(instance.Runs) == 1 {
count++
}
}
return count
}
// DegradedIdentities counts instances the action key could not be computed for
// in full, which are the rows a reader has to treat as a lower bound.
func (c Corpus) DegradedIdentities() int {
count := 0
for _, instance := range c.Instances {
if instance.RedactedOrigin {
count++
}
}
return count
}
func identify(runs []tracecorpus.Run, mode actionKeyMode) (Corpus, error) {
corpus := Corpus{Runs: len(runs)}
byKey := map[identityKey]*Instance{}
var order []identityKey
for _, run := range runs {
steps := index(run.Steps)
for _, step := range run.Steps {
for _, property := range step.Violations {
witness, ok := step.Witnesses[property]
if !ok || witness.Step == 0 {
corpus.Unattributed = append(corpus.Unattributed, Unattributed{
Property: property,
Run: run.Directory,
Step: step.Index,
Reason: attributionGap(ok),
})
continue
}
origin, ok := steps[witness.Step]
if !ok {
return Corpus{}, fmt.Errorf(
"%s: %s names origin step %d, which the trace does not hold",
run.Directory, property, witness.Step,
)
}
detected, ok := steps[witness.DetectedStep]
if !ok {
return Corpus{}, fmt.Errorf(
"%s: %s names detection step %d, which the trace does not hold",
run.Directory, property, witness.DetectedStep,
)
}
if detected.Screen == "" {
corpus.UnnamedScreen++
}
action, redactedOrigin := actionKey(origin, mode)
key := identityKey{
property: property,
action: action,
screen: detected.Screen,
}
instance, seen := byKey[key]
if !seen {
instance = &Instance{
Property: property,
OriginAction: key.action,
WitnessScreen: key.screen,
RedactedOrigin: redactedOrigin,
}
byKey[key] = instance
order = append(order, key)
}
instance.Reports++
if len(instance.Runs) == 0 ||
instance.Runs[len(instance.Runs)-1] != run.Directory {
instance.Runs = append(instance.Runs, run.Directory)
instance.Seeds = append(instance.Seeds, run.Meta.Seed)
}
}
}
}
for _, key := range order {
corpus.Instances = append(corpus.Instances, *byKey[key])
}
sort.SliceStable(corpus.Instances, func(i, j int) bool {
if corpus.Instances[i].Property != corpus.Instances[j].Property {
return corpus.Instances[i].Property < corpus.Instances[j].Property
}
return corpus.Instances[i].OriginAction < corpus.Instances[j].OriginAction
})
return corpus, nil
}
func attributionGap(hasWitness bool) string {
if !hasWitness {
return "no witness recorded"
}
return "witness records no origin step"
}
func index(steps []trace.Step) map[int]trace.Step {
byIndex := make(map[int]trace.Step, len(steps))
for _, step := range steps {
byIndex[step.Index] = step
}
return byIndex
}
// actionKey renders the action the origin step chose, and reports whether the
// key had to be degraded to the selector. The action recorded on a line is the
// one applied after observing it, which is the alignment that makes an origin
// index name an action at all.
//
// A typed value the record redacted is the same string for every value typed
// into that field, so keying on it would merge distinct actions while reading
// as a whole-action key. The key drops it and says it did, because an identity
// that cannot be computed has to show as an undercount rather than as a count.
func actionKey(origin trace.Step, mode actionKeyMode) (string, bool) {
if origin.NextAction == nil {
return "none", false
}
if origin.ActionSkipped != "" {
return "none (" + origin.ActionSkipped + ")", false
}
action := *origin.NextAction
key := action.Kind
switch {
case action.Selector != "":
key += " " + action.Selector
case action.Key != "":
key += " " + action.Key
case action.X != 0 || action.Y != 0 || action.ToX != 0 || action.ToY != 0:
key += fmt.Sprintf(" (%d,%d)", action.X, action.Y)
}
if mode != byFullAction {
return key, false
}
if action.Text == verifier.RedactedInputText {
return key + " text=redacted", true
}
return fmt.Sprintf("%s text=%q at=(%d,%d)->(%d,%d)",
key, action.Text, action.X, action.Y, action.ToX, action.ToY), false
}
@@ -0,0 +1,242 @@
package main
import (
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// TestOneDefectSeenTwiceIsOneInstance: two runs report the same property from
// the same origin action on the same screen, which is one defect found twice
// and two properties violated.
func TestOneDefectSeenTwiceIsOneInstance(t *testing.T) {
first := run(t, 3,
step(1, "/ledger", tap("id:add-txn")),
violating(2, "/accounts/7", tap("id:save"), "balanceMatches", 2, 2),
)
second := run(t, 5,
step(1, "/ledger", tap("id:add-txn")),
violating(2, "/accounts/7", tap("id:save"), "balanceMatches", 2, 2),
)
corpus := identified(t, bySelector, first, second)
if len(corpus.Instances) != 1 {
t.Fatalf("instances = %d, want 1: %+v", len(corpus.Instances), corpus.Instances)
}
instance := corpus.Instances[0]
if len(instance.Runs) != 2 || instance.Reports != 2 {
t.Fatalf("instance = %+v, want two runs reporting it", instance)
}
if corpus.Singletons() != 0 {
t.Fatalf("singletons = %d, want 0", corpus.Singletons())
}
}
func TestTheSamePropertyFromTwoOriginsIsTwoDefects(t *testing.T) {
first := run(t, 3, violating(1, "/ledger", tap("id:save"), "balanceMatches", 1, 1))
second := run(t, 5, violating(1, "/ledger", tap("id:delete"), "balanceMatches", 1, 1))
corpus := identified(t, bySelector, first, second)
if len(corpus.Instances) != 2 {
t.Fatalf("instances = %d, want 2: %+v", len(corpus.Instances), corpus.Instances)
}
if corpus.Singletons() != 2 {
t.Fatalf("singletons = %d, want 2", corpus.Singletons())
}
}
func TestTheSamePropertyOnTwoScreensIsTwoDefects(t *testing.T) {
first := run(t, 3, violating(1, "/ledger", tap("id:save"), "balanceMatches", 1, 1))
second := run(t, 5, violating(1, "/home", tap("id:save"), "balanceMatches", 1, 1))
corpus := identified(t, bySelector, first, second)
if len(corpus.Instances) != 2 {
t.Fatalf("instances = %d, want 2: %+v", len(corpus.Instances), corpus.Instances)
}
}
// TestADeferredViolationTakesTheActionOnItsOriginLine holds the alignment the
// draft states: the action recorded on line k is the one applied after
// observing k, so an obligation armed at k is attributed to that action and
// witnessed on the screen the detection step observed.
func TestADeferredViolationTakesTheActionOnItsOriginLine(t *testing.T) {
only := run(t, 3,
step(1, "/login", tap("id:login-submit")),
step(2, "/home", tap("id:open-ledger")),
violating(3, "/ledger", tap("id:add-txn"), "landsOnLedger", 2, 3),
)
corpus := identified(t, bySelector, only)
instance := corpus.Instances[0]
if instance.OriginAction != "Tap id:open-ledger" {
t.Fatalf("origin action = %q, want the action on the origin line", instance.OriginAction)
}
if instance.WitnessScreen != "/ledger" {
t.Fatalf("witness screen = %q, want the detection step's screen", instance.WitnessScreen)
}
}
func TestAViolationWithNoWitnessIsReportedNotCounted(t *testing.T) {
unattributed := trace.Step{
Index: 1,
Screen: "/ledger",
Violations: []string{"balanceMatches"},
}
corpus := identified(t, bySelector, run(t, 3, unattributed))
if len(corpus.Instances) != 0 {
t.Fatalf("instances = %d, want 0: %+v", len(corpus.Instances), corpus.Instances)
}
if len(corpus.Unattributed) != 1 ||
corpus.Unattributed[0].Property != "balanceMatches" {
t.Fatalf("unattributed = %+v, want the one violation with no origin", corpus.Unattributed)
}
}
// TestTheStrictActionKeySplitsWhatTheSelectorKeyMerges quantifies the reading
// the draft leaves open: two runs that typed different text into the same
// field are one defect by selector and two by the whole action.
func TestTheStrictActionKeySplitsWhatTheSelectorKeyMerges(t *testing.T) {
first := run(t, 3, violating(1, "/ledger", typing("id:amount", "12"), "balanceMatches", 1, 1))
second := run(t, 5, violating(1, "/ledger", typing("id:amount", "9000"), "balanceMatches", 1, 1))
if got := identified(t, bySelector, first, second); len(got.Instances) != 1 {
t.Fatalf("by selector: instances = %d, want 1", len(got.Instances))
}
if got := identified(t, byFullAction, first, second); len(got.Instances) != 2 {
t.Fatalf("by full action: instances = %d, want 2", len(got.Instances))
}
}
// TestARedactedTypedValueDegradesTheFullKeyVisibly: two runs typed different
// values into one field, both reached the trace redacted, and the whole action
// can no longer tell them apart. The pair is one row, and the report has to say
// so rather than let it read as one defect found twice.
func TestARedactedTypedValueDegradesTheFullKeyVisibly(t *testing.T) {
first := run(t, 3, violating(1, "/login",
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
second := run(t, 5, violating(1, "/login",
typing("id:password", recordedText(t, "correct horse")), "staysSignedIn", 1, 1))
corpus := identified(t, byFullAction, first, second)
if len(corpus.Instances) != 1 {
t.Fatalf("instances = %d, want the redacted pair to be one row: %+v",
len(corpus.Instances), corpus.Instances)
}
report := rendered(corpus)
if !strings.Contains(report, "1 identity") || !strings.Contains(report, "redacted") {
t.Fatalf("report does not say one identity rests on a redacted value:\n%s", report)
}
}
func TestARedactedOriginKeepsTheSelectorApart(t *testing.T) {
first := run(t, 3, violating(1, "/login",
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
second := run(t, 5, violating(1, "/login",
typing("id:pin", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
corpus := identified(t, byFullAction, first, second)
if len(corpus.Instances) != 2 {
t.Fatalf("instances = %d, want two fields to stay two rows: %+v",
len(corpus.Instances), corpus.Instances)
}
if report := rendered(corpus); !strings.Contains(report, "2 identity") {
t.Fatalf("report does not count both degraded identities:\n%s", report)
}
}
func TestARedactedOriginDegradesNothingUnderTheSelectorKey(t *testing.T) {
only := run(t, 3, violating(1, "/login",
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
if report := rendered(identified(t, bySelector, only)); strings.Contains(report, "redacted") {
t.Fatalf("selector key reads no text, so nothing degrades:\n%s", report)
}
}
// recordedText renders a typed value the way the runner records it, so what the
// key sees is redaction as it really happens and not a placeholder the test
// wrote itself.
func recordedText(t *testing.T, typed string) string {
t.Helper()
recorded := verifier.RecordedActionText(verifier.Action{
Kind: verifier.ActionKindInputText,
On: "id:password",
Text: typed,
}, nil)
if recorded == typed {
t.Fatalf("typed value %q reached the record unredacted", typed)
}
return recorded
}
func rendered(corpus Corpus) string {
var report strings.Builder
render(&report, corpus)
return report.String()
}
func identified(t *testing.T, mode actionKeyMode, runs ...tracecorpus.Run) Corpus {
t.Helper()
corpus, err := identify(runs, mode)
if err != nil {
t.Fatal(err)
}
return corpus
}
func run(t *testing.T, seed int64, steps ...trace.Step) tracecorpus.Run {
t.Helper()
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteMeta(trace.Meta{Seed: seed, Platform: "web"}); err != nil {
t.Fatal(err)
}
for _, step := range steps {
if err := writer.WriteStep(step); err != nil {
t.Fatal(err)
}
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
loaded, err := tracecorpus.Load(directory)
if err != nil {
t.Fatal(err)
}
return loaded
}
func step(index int, screen string, action *trace.Action) trace.Step {
return trace.Step{Index: index, Screen: screen, NextAction: action}
}
func violating(
index int,
screen string,
action *trace.Action,
property string,
origin int,
detected int,
) trace.Step {
violated := step(index, screen, action)
violated.Violations = []string{property}
violated.Witnesses = map[string]trace.Witness{
property: {Reason: "predicate false", Step: origin, DetectedStep: detected},
}
return violated
}
func tap(selector string) *trace.Action {
return &trace.Action{Kind: "Tap", Selector: selector, X: 10, Y: 20}
}
func typing(selector string, text string) *trace.Action {
return &trace.Action{Kind: "InputText", Selector: selector, Text: text, X: 10, Y: 20}
}
+120
View File
@@ -0,0 +1,120 @@
// Command defect-identity counts distinct defects across stored runs. A
// property reports at most once per run, so a run-level count is the number of
// properties violated; a defect is identified across runs by the property, the
// action attributed as the origin of the failed obligation, and the screen the
// witness observed.
package main
import (
"encoding/json"
"flag"
"fmt"
"io"
"os"
"text/tabwriter"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
)
func main() {
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
mode := flag.String(
"action-key",
string(bySelector),
"how much of the origin action identifies it: selector or full",
)
flag.Usage = func() {
fmt.Fprintln(
os.Stderr,
"usage: defect-identity [--json] [--action-key selector|full] <run directory> ...",
)
}
flag.Parse()
if flag.NArg() == 0 {
flag.Usage()
os.Exit(2)
}
if *mode != string(bySelector) && *mode != string(byFullAction) {
fmt.Fprintf(os.Stderr, "unknown --action-key %q\n", *mode)
os.Exit(2)
}
runs, err := loadAll(flag.Args())
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if len(runs) == 0 {
fmt.Fprintln(os.Stderr, "no run directory found under the given paths")
os.Exit(1)
}
corpus, err := identify(runs, actionKeyMode(*mode))
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if *jsonOut {
encoder := json.NewEncoder(os.Stdout)
encoder.SetIndent("", " ")
if err := encoder.Encode(corpus); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
return
}
render(os.Stdout, corpus)
}
func loadAll(paths []string) ([]tracecorpus.Run, error) {
var runs []tracecorpus.Run
for _, path := range paths {
directories, err := tracecorpus.Discover(path)
if err != nil {
return nil, err
}
for _, directory := range directories {
run, err := tracecorpus.Load(directory)
if err != nil {
return nil, fmt.Errorf("%s: %w", directory, err)
}
runs = append(runs, run)
}
}
return runs, nil
}
func render(out io.Writer, corpus Corpus) {
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
fmt.Fprintln(writer, "property\torigin action\twitness screen\truns\tseeds")
for _, instance := range corpus.Instances {
screen := instance.WitnessScreen
if screen == "" {
screen = "(unnamed)"
}
fmt.Fprintf(writer, "%s\t%s\t%s\t%d\t%v\n",
instance.Property, instance.OriginAction, screen,
len(instance.Runs), instance.Seeds)
}
writer.Flush()
fmt.Fprintf(out, "\n%d distinct defect(s) over %d run(s); %d seen in exactly one run\n",
len(corpus.Instances), corpus.Runs, corpus.Singletons())
if degraded := corpus.DegradedIdentities(); degraded > 0 {
fmt.Fprintf(out,
"%d identity(ies) rest on the origin selector alone, because the value typed "+
"there is redacted in the record; two runs that typed different values into "+
"that field read as one, so the count above is a floor for those\n",
degraded)
}
if corpus.UnnamedScreen > 0 {
fmt.Fprintf(out,
"%d violation(s) witnessed on a screen the app does not name, "+
"so identity rests on property and origin action alone for those\n",
corpus.UnnamedScreen)
}
for _, gap := range corpus.Unattributed {
fmt.Fprintf(out, "unattributed: %s at step %d of %s (%s)\n",
gap.Property, gap.Step, gap.Run, gap.Reason)
}
}
@@ -0,0 +1,145 @@
// Command exploration-reach counts the distinct structural states a stored
// run visited, and compares two runs by the observation at which their
// hierarchies first differ. Both read the trace alone: no device, no replay.
//
// The state is the settle path's structural hash of the recorded hierarchy,
// the same function the drivers wait on, so a state boundary here is the state
// boundary the harness itself uses.
package main
import (
"encoding/json"
"flag"
"fmt"
"io"
"os"
"text/tabwriter"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
)
func main() {
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
reference := flag.String(
"reference",
"",
"run directory to compare against; reports where each other run's hierarchy first differs",
)
flag.Usage = func() {
fmt.Fprintln(
os.Stderr,
"usage: exploration-reach [--json] [--reference RUN] <run directory> ...",
)
}
flag.Parse()
if flag.NArg() == 0 {
flag.Usage()
os.Exit(2)
}
reaches, err := measureAll(flag.Args())
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if len(reaches) == 0 {
fmt.Fprintln(os.Stderr, "no run directory found under the given paths")
os.Exit(1)
}
report := Report{Runs: reaches, CorpusDistinct: corpusDistinct(reaches)}
if *reference != "" {
base, loadErr := load(*reference)
if loadErr != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", *reference, loadErr)
os.Exit(1)
}
report.Reference = base.Directory
for _, reach := range reaches {
if reach.Directory == base.Directory {
continue
}
report.Divergences = append(report.Divergences, diverge(base, reach))
}
report.MedianDivergence, report.Censored = medianDivergence(report.Divergences)
}
if *jsonOut {
encoder := json.NewEncoder(os.Stdout)
encoder.SetIndent("", " ")
if err := encoder.Encode(report); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
return
}
render(os.Stdout, report)
}
// Report is one invocation's output: reach per run and over the corpus, plus
// the divergence rows when a reference run was named.
type Report struct {
Runs []Reach `json:"runs"`
CorpusDistinct int `json:"corpus_distinct_states"`
Reference string `json:"reference,omitempty"`
Divergences []Divergence `json:"divergences,omitempty"`
MedianDivergence float64 `json:"median_divergence_index,omitempty"`
Censored int `json:"never_diverged,omitempty"`
}
func load(path string) (Reach, error) {
run, err := tracecorpus.Load(path)
if err != nil {
return Reach{}, err
}
return measure(run), nil
}
func measureAll(paths []string) ([]Reach, error) {
var reaches []Reach
for _, path := range paths {
directories, err := tracecorpus.Discover(path)
if err != nil {
return nil, err
}
for _, directory := range directories {
reach, err := load(directory)
if err != nil {
return nil, fmt.Errorf("%s: %w", directory, err)
}
reaches = append(reaches, reach)
}
}
return reaches, nil
}
func render(out io.Writer, report Report) {
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
fmt.Fprintln(writer, "run\tseed\tplatform\tobservations\tdistinct states")
for _, reach := range report.Runs {
fmt.Fprintf(writer, "%s\t%d\t%s\t%d\t%d\n",
reach.Directory, reach.Seed, reach.Platform,
reach.Observations, reach.Distinct)
}
writer.Flush()
fmt.Fprintf(out, "\n%d run(s), %d distinct structural states across the corpus\n",
len(report.Runs), report.CorpusDistinct)
if report.Reference == "" {
return
}
fmt.Fprintf(out, "\nreference: %s\n", report.Reference)
writer = tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
fmt.Fprintln(writer, "run\tseed\tfirst divergence\tobservations compared")
for _, divergence := range report.Divergences {
where := fmt.Sprintf("%d", divergence.Step)
if !divergence.Diverged {
where = fmt.Sprintf("none through %d", divergence.Step)
}
fmt.Fprintf(writer, "%s\t%d\t%s\t%d\n",
divergence.Directory, divergence.Seed, where, divergence.Compared)
}
writer.Flush()
fmt.Fprintf(out, "\nmedian first divergence %.1f over %d run(s), %d never diverged\n",
report.MedianDivergence, len(report.Divergences), report.Censored)
}
@@ -0,0 +1,133 @@
package main
import (
"crypto/sha256"
"encoding/hex"
"sort"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
)
// observation is one hierarchy-bearing step: the index the run gave it and the
// state it observed.
type observation struct {
Step int
State string
}
// Reach is what one run explored. Distinct counts the structural states its
// observations visited, which is the measure; Observations is how many looks
// it took to visit them.
type Reach struct {
Directory string `json:"directory"`
Seed int64 `json:"seed"`
Platform string `json:"platform"`
Arm string `json:"arm,omitempty"`
Observations int `json:"observations"`
Distinct int `json:"distinct_states"`
// Unobserved counts steps carrying no hierarchy, which the run-end
// finalize record is, so an observation count cannot be read as a step
// count by accident.
Unobserved int `json:"steps_without_hierarchy"`
observations []observation
}
// measure keys each observation by a digest of its structural hash. The
// measure is equality between hashes and nothing else, and a corpus holds
// thousands of trees whose hashes run to tens of kilobytes each.
func measure(run tracecorpus.Run) Reach {
reach := Reach{
Directory: run.Directory,
Seed: run.Meta.Seed,
Platform: run.Meta.Platform,
Arm: run.Meta.Arm,
}
distinct := map[string]bool{}
for _, step := range run.Steps {
if step.Hierarchy == nil {
reach.Unobserved++
continue
}
digest := sha256.Sum256([]byte(ioscompanion.StructuralHash(step.Hierarchy)))
key := hex.EncodeToString(digest[:])
reach.observations = append(
reach.observations,
observation{Step: step.Index, State: key},
)
distinct[key] = true
}
reach.Observations = len(reach.observations)
reach.Distinct = len(distinct)
return reach
}
// corpusDistinct counts the structural states the whole corpus reached, which
// is not the sum of the per-run counts: runs of one application revisit the
// same screens.
func corpusDistinct(reaches []Reach) int {
distinct := map[string]bool{}
for _, reach := range reaches {
for _, seen := range reach.observations {
distinct[seen.State] = true
}
}
return len(distinct)
}
// Divergence is where a replay stopped observing what the reference observed.
// Step is the index of the first observation whose structural state differs;
// Diverged is false when the replay matched the reference for every
// observation the two share, in which case Step is that shared length and the
// observation is right-censored.
type Divergence struct {
Directory string `json:"directory"`
Seed int64 `json:"seed"`
Step int `json:"step"`
Diverged bool `json:"diverged"`
Compared int `json:"observations_compared"`
}
func diverge(reference, replay Reach) Divergence {
result := Divergence{Directory: replay.Directory, Seed: replay.Seed}
shared := len(reference.observations)
if len(replay.observations) < shared {
shared = len(replay.observations)
}
result.Compared = shared
for position := 0; position < shared; position++ {
if reference.observations[position].State != replay.observations[position].State {
result.Step = reference.observations[position].Step
result.Diverged = true
return result
}
}
if shared > 0 {
result.Step = reference.observations[shared-1].Step
}
return result
}
// medianDivergence is E6's number: the median observation index at which a
// replay first diverges from the reference. A replay that never diverged
// enters at the last index the two share, which is where the observation is
// censored rather than where it broke, so the count of such replays is
// reported beside the median rather than folded into it.
func medianDivergence(divergences []Divergence) (median float64, censored int) {
if len(divergences) == 0 {
return 0, 0
}
steps := make([]int, 0, len(divergences))
for _, divergence := range divergences {
steps = append(steps, divergence.Step)
if !divergence.Diverged {
censored++
}
}
sort.Ints(steps)
middle := len(steps) / 2
if len(steps)%2 == 1 {
return float64(steps[middle]), censored
}
return float64(steps[middle-1]+steps[middle]) / 2, censored
}
@@ -0,0 +1,177 @@
package main
import (
"testing"
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/tracecorpus"
)
const (
home = `{"attributes": {"text": "Home", "bounds": "[0,0,10,10]"}, "children": [
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []}
]}`
homeScrolled = `{"attributes": {"text": "Home", "bounds": "[0,4,10,14]"}, "children": [
{"attributes": {"text": "row", "bounds": "[0,4,5,9]"}, "children": []}
]}`
ledger = `{"attributes": {"text": "Ledger", "bounds": "[0,0,10,10]"}, "children": [
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []}
]}`
ledgerWithRow = `{"attributes": {"text": "Ledger", "bounds": "[0,0,10,10]"}, "children": [
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []},
{"attributes": {"text": "row"}, "children": []}
]}`
)
// TestReachCountsStructuresNotObservations: five observations of three
// structures, one of them revisited and one differing only in where it sits on
// screen, so the answer is three by construction.
func TestReachCountsStructuresNotObservations(t *testing.T) {
reach := measureRun(t, 7, home, homeScrolled, ledger, home, ledgerWithRow)
if reach.Observations != 5 {
t.Fatalf("observations = %d, want 5", reach.Observations)
}
if reach.Distinct != 3 {
t.Fatalf("distinct states = %d, want 3", reach.Distinct)
}
}
func TestFinalizeRecordIsNoObservation(t *testing.T) {
directory := writeRun(t, 7, home, ledger)
appendFinalize(t, directory, 3)
run, err := tracecorpus.Load(directory)
if err != nil {
t.Fatal(err)
}
reach := measure(run)
if reach.Observations != 2 || reach.Unobserved != 1 {
t.Fatalf("observations = %d, unobserved = %d, want 2 and 1",
reach.Observations, reach.Unobserved)
}
}
// TestStoredTreeHashesAsTheLiveTreeDid is the equivalence the whole measure
// rests on: the hash the settle path computed from the live tree and the hash
// this tool computes from the stored one are the same string, so a reach
// number counts the same state boundaries the drivers wait on.
func TestStoredTreeHashesAsTheLiveTreeDid(t *testing.T) {
live, err := hierarchy.Parse(home)
if err != nil {
t.Fatal(err)
}
liveHash := ioscompanion.StructuralHash(live)
if liveHash == "" {
t.Fatal("live hash is empty, so the test would pass on any stored tree")
}
run, err := tracecorpus.Load(writeRun(t, 7, home))
if err != nil {
t.Fatal(err)
}
stored := ioscompanion.StructuralHash(run.Steps[0].Hierarchy)
if stored != liveHash {
t.Fatalf("stored hash differs from the live one:\n live=%q\n stored=%q",
liveHash, stored)
}
}
func TestFirstDivergenceNamesTheObservationThatDiffers(t *testing.T) {
reference := measureRun(t, 7, home, ledger, home, ledger)
replay := measureRun(t, 7, home, ledger, ledgerWithRow, ledger)
divergence := diverge(reference, replay)
if !divergence.Diverged || divergence.Step != 3 {
t.Fatalf("divergence = %+v, want the third observation", divergence)
}
}
func TestAReplayThatMatchesIsCensoredAtTheSharedLength(t *testing.T) {
reference := measureRun(t, 7, home, ledger, home, ledger)
replay := measureRun(t, 7, home, ledger, home)
divergence := diverge(reference, replay)
if divergence.Diverged {
t.Fatalf("identical observations must not report a divergence: %+v", divergence)
}
if divergence.Step != 3 || divergence.Compared != 3 {
t.Fatalf("divergence = %+v, want censoring at the third observation", divergence)
}
}
func TestMedianDivergenceHoldsCensoredReplaysApart(t *testing.T) {
median, censored := medianDivergence([]Divergence{
{Step: 9, Diverged: true},
{Step: 3, Diverged: true},
{Step: 40, Diverged: false},
})
if median != 9 {
t.Fatalf("median = %v, want 9", median)
}
if censored != 1 {
t.Fatalf("censored = %d, want 1", censored)
}
}
func TestCorpusStatesAreTheUnionNotTheSum(t *testing.T) {
first := measureRun(t, 7, home, ledger)
second := measureRun(t, 11, ledger, ledgerWithRow)
if got := corpusDistinct([]Reach{first, second}); got != 3 {
t.Fatalf("corpus distinct = %d, want 3", got)
}
}
func measureRun(t *testing.T, seed int64, dumps ...string) Reach {
t.Helper()
run, err := tracecorpus.Load(writeRun(t, seed, dumps...))
if err != nil {
t.Fatal(err)
}
return measure(run)
}
func writeRun(t *testing.T, seed int64, dumps ...string) string {
t.Helper()
directory := t.TempDir()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteMeta(trace.Meta{Seed: seed, Platform: "web"}); err != nil {
t.Fatal(err)
}
for index, dump := range dumps {
tree, err := hierarchy.Parse(dump)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteStep(trace.Step{Index: index + 1, Hierarchy: tree}); err != nil {
t.Fatal(err)
}
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
return directory
}
func appendFinalize(t *testing.T, directory string, index int) {
t.Helper()
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatal(err)
}
if err := writer.WriteStep(trace.Step{
Index: index,
Violations: []string{"someTransactionExists"},
}); err != nil {
t.Fatal(err)
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
}
@@ -0,0 +1,482 @@
package main
import (
"bufio"
"bytes"
"encoding/json"
"fmt"
"io"
"math/rand"
"net"
"net/http"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
)
// The test binary doubles as the stub implementation's server and as the
// fetcher the stub campaign uses, so the sweep drives a real preview process
// over a real port and the served URL is answered by a real HTTP server.
func TestMain(m *testing.M) {
switch {
case os.Getenv("SWEEP_TEST_SERVE_PORT") != "":
serveUntilKilled(os.Getenv("SWEEP_TEST_SERVE_PORT"))
case os.Getenv("SWEEP_TEST_FETCH_URL") != "":
recordFetch(
os.Getenv("SWEEP_TEST_FETCH_URL"),
os.Getenv("SWEEP_TEST_FETCH_LOG"),
)
default:
os.Exit(m.Run())
}
}
func serveUntilKilled(port string) {
handler := http.HandlerFunc(
func(writer http.ResponseWriter, request *http.Request) {
fmt.Fprintf(writer, "%s %s", port, request.URL.RequestURI())
},
)
http.ListenAndServe("localhost:"+port, handler)
}
func recordFetch(url, logPath string) {
line := ""
response, err := http.Get(url)
if err != nil {
line = fmt.Sprintf("%s -> error %v\n", url, err)
} else {
body, _ := io.ReadAll(response.Body)
response.Body.Close()
line = fmt.Sprintf("%s -> %d %s\n", url, response.StatusCode, body)
}
logFile, err := os.OpenFile(
logPath,
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
0o644,
)
if err != nil {
return
}
logFile.WriteString(line)
logFile.Close()
}
// stubBun answers install, fails to build impl-02, and serves the preview from
// the test binary on the port it was given.
const stubBun = `#!/bin/sh
echo "$PWD $*" >> "%[1]s"
if [ "$1" = "run" ] && [ "$2" = "build" ]; then
case "$PWD" in *impl-02) echo "TS2322: type error" >&2; exit 1 ;; esac
exit 0
fi
if [ "$1" = "run" ] && [ "$2" = "preview" ]; then
port=""
previous=""
for argument in "$@"; do
if [ "$previous" = "--port" ]; then port="$argument"; fi
previous="$argument"
done
SWEEP_TEST_SERVE_PORT="$port" exec "%[2]s"
fi
exit 0
`
// stubCampaign records the argv it was handed and whether the sweep manifest
// was already on disk when it ran, fetches the URL it was told to drive, writes
// the campaign directory the real tool would write, and fails impl-03 seed 4.
const stubCampaign = `#!/bin/sh
output=""
seed=""
url=""
previous=""
for argument in "$@"; do
case "$previous" in
--output) output="$argument" ;;
--seeds) seed="$argument" ;;
--bundle-id) url="$argument" ;;
esac
previous="$argument"
done
manifest=missing
if [ -f "%[1]s" ]; then manifest=present; fi
echo "manifest=$manifest argv: $*" >> "%[2]s"
SWEEP_TEST_FETCH_URL="$url" SWEEP_TEST_FETCH_LOG="%[3]s" "%[4]s"
mkdir -p "$output"
printf '{"arm":"stub","seeds":[%%s]}\n' "$seed" > "$output/campaign.json"
echo "stub campaign seed=$seed url=$url"
case "$output" in
*impl-03/seed-4) exit 1 ;;
esac
exit 0
`
func TestRun_EndToEndAgainstStubBunAndCampaign(t *testing.T) {
root := t.TempDir()
implementations := filepath.Join(root, "implementations")
for _, name := range []string{"impl-01", "impl-02", "impl-03"} {
if err := os.MkdirAll(filepath.Join(implementations, name), 0o755); err != nil {
t.Fatal(err)
}
}
specPath := filepath.Join(root, "relay.ts")
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
t.Fatal(err)
}
output := filepath.Join(root, "campaigns")
bunLog := filepath.Join(root, "bun.log")
campaignLog := filepath.Join(root, "campaign.log")
fetchLog := filepath.Join(root, "fetch.log")
testBinary, err := filepath.Abs(os.Args[0])
if err != nil {
t.Fatal(err)
}
bunPath := writeScript(
t,
filepath.Join(root, "stub-bun"),
fmt.Sprintf(stubBun, bunLog, testBinary),
)
campaignPath := writeScript(
t,
filepath.Join(root, "stub-campaign"),
fmt.Sprintf(
stubCampaign,
filepath.Join(output, manifestFileName),
campaignLog,
fetchLog,
testBinary,
),
)
sanderlingPath := writeScript(
t,
filepath.Join(root, "stub-sanderling"),
"#!/bin/sh\nexit 0\n",
)
basePort := freePortRange(t, 3)
var stdout bytes.Buffer
err = run([]string{
"--implementations", implementations,
"--spec", specPath,
"--seeds", "4-5",
"--max-steps", "40",
"--duration", "30s",
"--concurrency", "2",
"--base-port", fmt.Sprint(basePort),
"--output", output,
"--bun", bunPath,
"--campaign", campaignPath,
"--sanderling", sanderlingPath,
}, &stdout, io.Discard)
if err == nil ||
!strings.Contains(err.Error(), "1 of 3 implementations never ran") {
t.Fatalf(
"expected the failed build and the failed campaign to be reported, got %v",
err,
)
}
var recorded manifest
manifestBody, err := os.ReadFile(filepath.Join(output, manifestFileName))
if err != nil {
t.Fatal(err)
}
if err := json.Unmarshal(manifestBody, &recorded); err != nil {
t.Fatal(err)
}
if !slices.Equal(recorded.Seeds, []int64{4, 5}) {
t.Errorf("intended seeds: got %v", recorded.Seeds)
}
if len(recorded.Implementations) != 3 {
t.Fatalf("intended implementations: got %v", recorded.Implementations)
}
for index, planned := range recorded.Implementations {
wantPort := basePort + index
if planned.Port != wantPort {
t.Errorf(
"%s port: got %d, want %d",
planned.Name,
planned.Port,
wantPort,
)
}
wantURL := fmt.Sprintf("http://localhost:%d/?seed={seed}", wantPort)
if planned.URLTemplate != wantURL {
t.Errorf(
"%s url template: got %q, want %q",
planned.Name,
planned.URLTemplate,
wantURL,
)
}
}
if recorded.Generator != "seeded" || recorded.MaxSteps != 40 {
t.Errorf(
"manifest generator/budget: got %q/%d",
recorded.Generator,
recorded.MaxSteps,
)
}
// impl-02 fails its build, so the two implementations either side of it
// still have to reach the campaign tool with both seeds.
campaignLines := readLines(t, campaignLog)
if len(campaignLines) != 4 {
t.Fatalf(
"campaign invocations: got %d, want 4:\n%s",
len(campaignLines),
strings.Join(campaignLines, "\n"),
)
}
seen := map[string]bool{}
for _, line := range campaignLines {
if !strings.HasPrefix(line, "manifest=present") {
t.Errorf(
"a campaign ran before the sweep manifest was written: %q",
line,
)
}
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
arm := argumentValue(arguments, "--arm")
seed := argumentValue(arguments, "--seeds")
bundle := argumentValue(arguments, "--bundle-id")
port := basePort + slices.Index(
[]string{"impl-01", "impl-02", "impl-03"},
arm,
)
wantBundle := fmt.Sprintf("http://localhost:%d/?seed=%s", port, seed)
if bundle != wantBundle {
t.Errorf(
"%s seed %s: bundle id %q, want %q",
arm,
seed,
bundle,
wantBundle,
)
}
if got := argumentValue(arguments, "--output"); got != filepath.Join(
output,
arm,
"seed-"+seed,
) {
t.Errorf("%s seed %s: campaign output %q", arm, seed, got)
}
if got := argumentValue(arguments, "--sanderling"); got != sanderlingPath {
t.Errorf("%s seed %s: sanderling path %q", arm, seed, got)
}
seen[arm+"/"+seed] = true
}
for _, want := range []string{"impl-01/4", "impl-01/5", "impl-03/4", "impl-03/5"} {
if !seen[want] {
t.Errorf("%s never reached the campaign tool", want)
}
}
// What the served page actually saw: the right port for the
// implementation, carrying the same seed the campaign was given.
fetched := readLines(t, fetchLog)
for _, want := range []string{
fmt.Sprintf("http://localhost:%d/?seed=4 -> 200 %d /?seed=4", basePort, basePort),
fmt.Sprintf("http://localhost:%d/?seed=5 -> 200 %d /?seed=5", basePort, basePort),
fmt.Sprintf("http://localhost:%d/?seed=4 -> 200 %d /?seed=4", basePort+2, basePort+2),
fmt.Sprintf("http://localhost:%d/?seed=5 -> 200 %d /?seed=5", basePort+2, basePort+2),
} {
if !slices.Contains(fetched, want) {
t.Errorf(
"the served page never saw %q:\n%s",
want,
strings.Join(fetched, "\n"),
)
}
}
records := readRecords(t, filepath.Join(output, recordsFileName))
if len(records) != 3 {
t.Fatalf("implementation records: got %d, want 3", len(records))
}
byName := map[string]implementationRecord{}
for _, record := range records {
byName[record.Name] = record
}
failed := byName["impl-02"]
if failed.FailedStage != stageBuild || len(failed.Runs) != 0 {
t.Errorf(
"impl-02: got stage %q with %d runs, want a build failure and no runs",
failed.FailedStage,
len(failed.Runs),
)
}
if !strings.Contains(failed.Error, "build.log") {
t.Errorf("impl-02 error should point at its log: %q", failed.Error)
}
buildLog, err := os.ReadFile(
filepath.Join(output, "impl-02", stageBuild+".log"),
)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(buildLog), "TS2322") {
t.Errorf("impl-02 build log lost the compiler error: %q", buildLog)
}
for _, name := range []string{"impl-01", "impl-03"} {
record := byName[name]
if record.FailedStage != "" || len(record.Runs) != 2 {
t.Errorf(
"%s: stage %q with %d runs, want no failure and 2 runs",
name,
record.FailedStage,
len(record.Runs),
)
}
if record.MonotonicMillis <= 0 {
t.Errorf(
"%s took %d ms, so nothing timed how long it worked",
name,
record.MonotonicMillis,
)
}
for _, run := range record.Runs {
if run.MonotonicMillis <= 0 {
t.Errorf(
"%s seed %d took %d ms, so nothing timed the campaign",
name,
run.Seed,
run.MonotonicMillis,
)
}
}
}
if exit := byName["impl-03"].Runs[0].ExitCode; exit != 1 {
t.Errorf("impl-03 seed 4 exit code: got %d, want 1", exit)
}
if exit := byName["impl-03"].Runs[1].ExitCode; exit != 0 {
t.Errorf(
"impl-03 seed 5 ran after seed 4 failed and should have exited 0, got %d",
exit,
)
}
for _, name := range []string{"impl-01", "impl-03"} {
for _, seed := range []string{"4", "5"} {
directory := filepath.Join(output, name, "seed-"+seed)
if _, err := os.Stat(filepath.Join(directory, "campaign.json")); err != nil {
t.Errorf(
"%s seed %s: no campaign directory: %v",
name,
seed,
err,
)
}
log, err := os.ReadFile(filepath.Join(directory, "campaign.log"))
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(log), "stub campaign seed="+seed) {
t.Errorf(
"%s seed %s: campaign output was not captured: %q",
name,
seed,
log,
)
}
}
}
installed := readLines(t, bunLog)
for _, name := range []string{"impl-01", "impl-02", "impl-03"} {
if !slices.Contains(
installed,
filepath.Join(implementations, name)+" install",
) {
t.Errorf(
"%s was never installed:\n%s",
name,
strings.Join(installed, "\n"),
)
}
}
// Every preview server the sweep started is gone with it: a leaked one
// holds its port, and the next sweep would be served by the old build.
client := &http.Client{Timeout: 2 * time.Second}
for _, port := range []int{basePort, basePort + 2} {
if response, err := client.Get(readinessURL(port)); err == nil {
response.Body.Close()
t.Errorf("port %d is still served after the sweep finished", port)
}
}
if !strings.Contains(stdout.String(), "failed at build") {
t.Errorf(
"progress output does not name the build failure: %q",
stdout.String(),
)
}
}
func writeScript(t *testing.T, path, body string) string {
t.Helper()
if err := os.WriteFile(path, []byte(body), 0o755); err != nil {
t.Fatal(err)
}
return path
}
func readLines(t *testing.T, path string) []string {
t.Helper()
body, err := os.ReadFile(path)
if err != nil {
t.Fatal(err)
}
var lines []string
scanner := bufio.NewScanner(strings.NewReader(string(body)))
for scanner.Scan() {
if line := strings.TrimSpace(scanner.Text()); line != "" {
lines = append(lines, line)
}
}
return lines
}
func readRecords(t *testing.T, path string) []implementationRecord {
t.Helper()
var records []implementationRecord
for _, line := range readLines(t, path) {
var record implementationRecord
if err := json.Unmarshal([]byte(line), &record); err != nil {
t.Fatalf("%s: %v", line, err)
}
records = append(records, record)
}
return records
}
// freePortRange finds count consecutive free ports, which is what the sweep
// hands out: one port per implementation from --base-port upwards.
func freePortRange(t *testing.T, count int) int {
t.Helper()
for range 100 {
base := 20000 + rand.Intn(20000)
if portsAreFree(base, count) {
return base
}
}
t.Fatalf("no run of %d free ports", count)
return 0
}
func portsAreFree(base, count int) bool {
for offset := range count {
listener, err := net.Listen(
"tcp",
fmt.Sprintf("localhost:%d", base+offset),
)
if err != nil {
return false
}
listener.Close()
}
return true
}
@@ -0,0 +1,296 @@
// Command implementation-sweep runs one identical campaign against every model
// implementation of a single requirement. It installs, builds and serves each
// implementation on its own port, then hands the campaign tool the same seed
// slice, the same step budget and the same generator for all of them, so a
// difference between implementations is not a difference in exploration.
package main
import (
"context"
"errors"
"flag"
"fmt"
"io"
"os"
"os/signal"
"path/filepath"
"strconv"
"syscall"
"time"
"github.com/priyanshujain/sanderling/internal/seedspec"
)
// The generator and the platform are fixed rather than exposed: the
// pre-registration runs the seeded policy against a served web build, and a
// sweep that could quietly run something else records a comparison nobody made.
const (
generator = "seeded"
platform = "web"
)
// defaultConcurrency is how many implementations are built, served and swept at
// once. fleet.md measured eight concurrent web campaigns clean at about 1.1 GB
// resident each, parallel efficiency 0.83 at eight against 0.87 at six, and a
// knee at twelve to sixteen, on a contended laptop it says to re-measure before
// trusting anything above eight. Six sits at the better efficiency, costs about
// 7 GB of a 64 GB host, and leaves the Android emulator farm that shares that
// host its four to six slots. Each worker here also carries a vite server the
// fleet measurement did not include.
const defaultConcurrency = 6
// defaultBasePort is the first port handed out. Vite's own defaults are 5173
// and 4173, so a sweep starting here does not collide with a dev server someone
// left running.
const defaultBasePort = 5300
type config struct {
implementationsDirectory string
specPath string
outputDirectory string
seeds []int64
maxSteps int
duration time.Duration
concurrency int
basePort int
bunPath string
campaignPath string
sanderlingPath string
extraArguments []string
}
const usage = `implementation-sweep runs one seeded campaign per model implementation.
Usage:
implementation-sweep --implementations <dir> --spec <path> --seeds <spec>
--max-steps <n> --output <dir> [flags]
[-- <sanderling test flags>]
Each impl-* directory under --implementations is installed, built and served on
its own port, and every one is swept with the same seeds and the same step
budget. An implementation that fails to install, build or serve is recorded and
the sweep moves on to the next one.
Everything after a bare -- reaches every sanderling test call through the
campaign tool.
`
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
flagSet := flag.NewFlagSet("implementation-sweep", flag.ContinueOnError)
flagSet.SetOutput(stderr)
flagSet.Usage = func() {
fmt.Fprint(stderr, usage)
flagSet.PrintDefaults()
}
var configuration config
var seedSpecification string
flagSet.StringVar(
&configuration.implementationsDirectory,
"implementations",
"",
"directory holding impl-01 to impl-NN (required)",
)
flagSet.StringVar(
&configuration.specPath,
"spec",
"",
"path to the property set every implementation is run against (required)",
)
flagSet.StringVar(
&seedSpecification,
"seeds",
"",
"seeds every implementation runs: ranges and lists, e.g. 1-10,20 (required)",
)
flagSet.IntVar(
&configuration.maxSteps,
"max-steps",
0,
"per-run step budget, identical across implementations (required, must be positive)",
)
flagSet.DurationVar(
&configuration.duration,
"duration",
5*time.Minute,
"per-run wall-clock ceiling passed to each campaign",
)
flagSet.IntVar(
&configuration.concurrency,
"concurrency",
defaultConcurrency,
"implementations built, served and swept at once",
)
flagSet.IntVar(
&configuration.basePort,
"base-port",
defaultBasePort,
"first port served; each implementation takes the next one in name order",
)
flagSet.StringVar(
&configuration.outputDirectory,
"output",
"",
"campaign tree to create (required)",
)
flagSet.StringVar(
&configuration.bunPath,
"bun",
"bun",
"bun binary that installs, builds and serves each implementation",
)
flagSet.StringVar(
&configuration.campaignPath,
"campaign",
"campaign",
"campaign binary to invoke per implementation and seed",
)
flagSet.StringVar(
&configuration.sanderlingPath,
"sanderling",
"sanderling",
"sanderling binary each campaign invokes",
)
if err := flagSet.Parse(arguments); err != nil {
return config{}, err
}
configuration.extraArguments = flagSet.Args()
// Every missing flag is named together, in flag order: stopping at the
// first turns one rerun into one rerun per missing flag.
var missing []error
for _, required := range []struct {
name string
value string
}{
{"--implementations", configuration.implementationsDirectory},
{"--spec", configuration.specPath},
{"--seeds", seedSpecification},
{"--output", configuration.outputDirectory},
} {
if required.value == "" {
missing = append(
missing,
fmt.Errorf("%s is required", required.name),
)
}
}
if err := errors.Join(missing...); err != nil {
return config{}, err
}
if configuration.maxSteps <= 0 {
return config{}, fmt.Errorf(
"--max-steps must be positive: every implementation needs the same step budget",
)
}
if configuration.duration <= 0 {
return config{}, fmt.Errorf(
"--duration must be positive: %s",
configuration.duration,
)
}
if configuration.concurrency <= 0 {
return config{}, fmt.Errorf(
"--concurrency must be positive: %d",
configuration.concurrency,
)
}
if configuration.basePort < 1024 || configuration.basePort > 65535 {
return config{}, fmt.Errorf(
"--base-port %d is outside 1024-65535",
configuration.basePort,
)
}
seeds, err := seedspec.Parse(seedSpecification)
if err != nil {
return config{}, fmt.Errorf("--seeds: %w", err)
}
configuration.seeds = seeds
for name, value := range map[string]*string{
"--implementations": &configuration.implementationsDirectory,
"--spec": &configuration.specPath,
"--output": &configuration.outputDirectory,
} {
absolute, err := filepath.Abs(*value)
if err != nil {
return config{}, fmt.Errorf("%s: %w", name, err)
}
*value = absolute
}
return configuration, nil
}
// servedURL is the one place a seed becomes a URL. The same seed is also handed
// to the campaign as --seeds, which reaches sanderling as --seed and fixes the
// exploration, while the scaffold reads ?seed= and fixes the latency and the
// outcome of every send. A violation replays only when both carry the same
// number, so both come from the seed argument here and never from two flags.
func servedURL(port int, seed string) string {
return fmt.Sprintf("http://localhost:%d/?seed=%s", port, seed)
}
func readinessURL(port int) string {
return fmt.Sprintf("http://localhost:%d/", port)
}
func campaignDirectory(
configuration config,
target implementation,
seed string,
) string {
return filepath.Join(
configuration.outputDirectory,
target.Name,
"seed-"+seed,
)
}
// campaignArguments builds one campaign invocation. The seed is a string
// because it lands in two arguments, --seeds and the ?seed= of --bundle-id,
// and passing it once keeps them from drifting apart.
func campaignArguments(
configuration config,
target implementation,
seed string,
) []string {
arguments := []string{
"--spec", configuration.specPath,
"--bundle-id", servedURL(target.Port, seed),
"--platform", platform,
"--arm", target.Name,
"--generator", generator,
"--max-steps", strconv.Itoa(configuration.maxSteps),
"--duration", configuration.duration.String(),
"--seeds", seed,
"--sanderling", configuration.sanderlingPath,
"--output", campaignDirectory(configuration, target, seed),
}
if len(configuration.extraArguments) > 0 {
arguments = append(arguments, "--")
arguments = append(arguments, configuration.extraArguments...)
}
return arguments
}
func run(arguments []string, stdout, stderr io.Writer) error {
configuration, err := parseArguments(arguments, stderr)
if err != nil {
return err
}
ctx, cancel := signal.NotifyContext(
context.Background(),
os.Interrupt,
syscall.SIGTERM,
)
defer cancel()
return runSweep(ctx, configuration, stdout)
}
func main() {
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
if errors.Is(err, flag.ErrHelp) {
return
}
fmt.Fprintf(os.Stderr, "error: %v\n", err)
os.Exit(1)
}
}
@@ -0,0 +1,241 @@
package main
import (
"io"
"net/url"
"path/filepath"
"slices"
"strconv"
"strings"
"testing"
"time"
)
func baseArguments() []string {
return []string{
"--implementations", "/e4/implementations",
"--spec", "/e4/relay.ts",
"--seeds", "1-3",
"--max-steps", "400",
"--output", "/campaigns/e4",
}
}
func TestParseArguments_DefaultsAndSeeds(t *testing.T) {
configuration, err := parseArguments(baseArguments(), io.Discard)
if err != nil {
t.Fatal(err)
}
if !slices.Equal(configuration.seeds, []int64{1, 2, 3}) {
t.Errorf("seeds: got %v", configuration.seeds)
}
if configuration.concurrency != defaultConcurrency {
t.Errorf(
"concurrency default: got %d, want %d",
configuration.concurrency,
defaultConcurrency,
)
}
if configuration.basePort != defaultBasePort {
t.Errorf(
"base port default: got %d, want %d",
configuration.basePort,
defaultBasePort,
)
}
if configuration.duration != 5*time.Minute {
t.Errorf("duration default: got %s", configuration.duration)
}
for name, got := range map[string]string{
"bun": configuration.bunPath,
"campaign": configuration.campaignPath,
"sanderling": configuration.sanderlingPath,
} {
if got != name {
t.Errorf("%s path default: got %q", name, got)
}
}
}
func TestParseArguments_Rejections(t *testing.T) {
cases := []struct {
name string
arguments []string
want string
}{
{
"missing implementations",
[]string{
"--spec",
"s",
"--seeds",
"1",
"--max-steps",
"10",
"--output",
"o",
},
"--implementations is required",
},
{
"missing spec",
[]string{
"--implementations",
"i",
"--seeds",
"1",
"--max-steps",
"10",
"--output",
"o",
},
"--spec is required",
},
{
"missing output",
[]string{
"--implementations",
"i",
"--spec",
"s",
"--seeds",
"1",
"--max-steps",
"10",
},
"--output is required",
},
{
"zero max steps",
append(baseArguments(), "--max-steps", "0"),
"--max-steps must be positive",
},
{
"zero concurrency",
append(baseArguments(), "--concurrency", "0"),
"--concurrency must be positive",
},
{
"privileged base port",
append(baseArguments(), "--base-port", "80"),
"outside 1024-65535",
},
{
"seed zero",
append(baseArguments(), "--seeds", "0,1"),
"not reproducible",
},
}
for _, testCase := range cases {
_, err := parseArguments(testCase.arguments, io.Discard)
if err == nil {
t.Errorf("%s: expected error", testCase.name)
continue
}
if !strings.Contains(err.Error(), testCase.want) {
t.Errorf(
"%s: got %q, want it to contain %q",
testCase.name,
err,
testCase.want,
)
}
}
}
// Three flags missing is one rerun, not three: the operator is told about all
// of them at once, in flag order, whatever order the check happened to walk.
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
_, err := parseArguments(
[]string{"--spec", "s", "--max-steps", "10"},
io.Discard,
)
if err == nil {
t.Fatal("got no error, want every missing flag named")
}
message := err.Error()
previous := -1
for _, name := range []string{"--implementations", "--seeds", "--output"} {
at := strings.Index(message, name)
if at < 0 {
t.Fatalf("got %q, want %s named", message, name)
}
if at < previous {
t.Errorf("got %q, want the flags named in flag order", message)
}
previous = at
}
if strings.Contains(message, "--spec") {
t.Errorf("got %q, want the supplied --spec left out", message)
}
}
// The seed reaches two independent things, the campaign's own seed and the
// scaffold's failure stream, and a replay reproduces neither unless they carry
// the same number.
func TestCampaignArguments_OneSeedReachesBothTheCampaignAndTheURL(
t *testing.T,
) {
configuration, err := parseArguments(
append(baseArguments(), "--", "--clear-data=false"),
io.Discard,
)
if err != nil {
t.Fatal(err)
}
target := implementation{
Name: "impl-07",
Directory: "/e4/implementations/impl-07",
Port: 5306,
}
for _, seed := range []string{"1", "42"} {
arguments := campaignArguments(configuration, target, seed)
if got := argumentValue(arguments, "--seeds"); got != seed {
t.Errorf("--seeds: got %q, want %q", got, seed)
}
bundle := argumentValue(arguments, "--bundle-id")
parsed, err := url.Parse(bundle)
if err != nil {
t.Fatalf("--bundle-id %q: %v", bundle, err)
}
if got := parsed.Query().Get("seed"); got != seed {
t.Errorf(
"served URL seed: got %q, want %q (from %q)",
got,
seed,
bundle,
)
}
if parsed.Host != "localhost:"+strconv.Itoa(target.Port) {
t.Errorf(
"served host: got %q, want the implementation's own port %d",
parsed.Host,
target.Port,
)
}
for flagName, want := range map[string]string{
"--arm": "impl-07",
"--platform": "web",
"--generator": "seeded",
"--max-steps": "400",
"--spec": "/e4/relay.ts",
"--output": filepath.Join("/campaigns/e4", "impl-07", "seed-"+seed),
} {
if got := argumentValue(arguments, flagName); got != want {
t.Errorf("%s: got %q, want %q", flagName, got, want)
}
}
if arguments[len(arguments)-2] != "--" ||
arguments[len(arguments)-1] != "--clear-data=false" {
t.Errorf("passthrough flags lost: %v", arguments)
}
}
}
func argumentValue(arguments []string, name string) string {
index := slices.Index(arguments, name)
if index < 0 || index+1 >= len(arguments) {
return ""
}
return arguments[index+1]
}
@@ -0,0 +1,117 @@
package main
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
"time"
)
const (
manifestFileName = "sweep.json"
recordsFileName = "implementations.jsonl"
seedPlaceholder = "{seed}"
)
// plannedImplementation is one implementation the sweep intends to run, with
// the port it is served on and the URL every seed is served at.
type plannedImplementation struct {
Name string `json:"name"`
Directory string `json:"directory"`
Port int `json:"port"`
URLTemplate string `json:"url_template"`
}
// manifest is sweep.json: what the sweep INTENDED to run, written before the
// first install so a host that dropped an implementation or a seed shows up as
// a missing run rather than as a smaller sample.
type manifest struct {
Generator string `json:"generator"`
Platform string `json:"platform"`
SpecPath string `json:"spec_path"`
MaxSteps int `json:"max_steps"`
DurationMillis int64 `json:"duration_millis"`
Seeds []int64 `json:"seeds"`
Implementations []plannedImplementation `json:"implementations"`
Concurrency int `json:"concurrency"`
Host string `json:"host"`
BunPath string `json:"bun_path"`
CampaignPath string `json:"campaign_path"`
SanderlingPath string `json:"sanderling_path"`
StartedAt time.Time `json:"started_at"`
}
func buildManifest(
configuration config,
implementations []implementation,
host string,
startedAt time.Time,
) manifest {
planned := make([]plannedImplementation, 0, len(implementations))
for _, target := range implementations {
planned = append(planned, plannedImplementation{
Name: target.Name,
Directory: target.Directory,
Port: target.Port,
URLTemplate: servedURL(target.Port, seedPlaceholder),
})
}
return manifest{
Generator: generator,
Platform: platform,
SpecPath: configuration.specPath,
MaxSteps: configuration.maxSteps,
DurationMillis: configuration.duration.Milliseconds(),
Seeds: configuration.seeds,
Implementations: planned,
Concurrency: configuration.concurrency,
Host: host,
BunPath: configuration.bunPath,
CampaignPath: configuration.campaignPath,
SanderlingPath: configuration.sanderlingPath,
StartedAt: startedAt,
}
}
func writeManifest(directory string, value manifest) error {
body, err := json.MarshalIndent(value, "", " ")
if err != nil {
return fmt.Errorf("marshal manifest: %w", err)
}
return os.WriteFile(
filepath.Join(directory, manifestFileName),
append(body, '\n'),
0o644,
)
}
// runRecord is one campaign, which is one implementation at one seed.
type runRecord struct {
Seed int64 `json:"seed"`
URL string `json:"url"`
ExitCode int `json:"exit_code"`
LaunchError string `json:"launch_error,omitempty"`
CampaignDirectory string `json:"campaign_directory"`
MonotonicMillis int64 `json:"monotonic_millis"`
}
// implementationRecord is one line of implementations.jsonl. FailedStage names
// the step that stopped this implementation, and an implementation that never
// got past install, build or serve carries no runs at all.
type implementationRecord struct {
Name string `json:"implementation"`
Directory string `json:"directory"`
Port int `json:"port"`
FailedStage string `json:"failed_stage,omitempty"`
Error string `json:"error,omitempty"`
StartedAt time.Time `json:"started_at"`
MonotonicMillis int64 `json:"monotonic_millis"`
Runs []runRecord `json:"runs"`
}
const (
stageInstall = "install"
stageBuild = "build"
stageServe = "serve"
)
@@ -0,0 +1,109 @@
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"os/exec"
"syscall"
"time"
)
const (
// serverStartTimeout covers a cold vite start on a host already running
// five other implementations.
serverStartTimeout = 90 * time.Second
serverPollInterval = 250 * time.Millisecond
serverShutdownGrace = 10 * time.Second
)
// server is one implementation's preview server. It runs in its own process
// group so that stopping it takes the whole vite tree with it: a leaked server
// holds its port, and the next sweep against that implementation would be
// served by the previous build.
type server struct {
command *exec.Cmd
logFile *os.File
exited chan struct{}
}
func startServer(
ctx context.Context,
configuration config,
target implementation,
logPath string,
) (*server, error) {
logFile, err := os.Create(logPath)
if err != nil {
return nil, err
}
command := exec.CommandContext(ctx, configuration.bunPath,
"run", "preview", "--port", fmt.Sprint(target.Port), "--strictPort")
command.Dir = target.Directory
command.Stdout = logFile
command.Stderr = logFile
command.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
if err := command.Start(); err != nil {
logFile.Close()
return nil, err
}
running := &server{
command: command,
logFile: logFile,
exited: make(chan struct{}),
}
go func() {
command.Wait()
close(running.exited)
}()
return running, nil
}
// waitReady polls the served page until it answers. A server that exits first
// is reported as such rather than waited on for the full timeout, because the
// usual cause is a port already taken and that answer is in the log.
func (s *server) waitReady(ctx context.Context, url string) error {
client := &http.Client{Timeout: 5 * time.Second}
deadline := time.Now().Add(serverStartTimeout)
for {
select {
case <-s.exited:
return fmt.Errorf("server exited before it answered %s", url)
case <-ctx.Done():
return ctx.Err()
default:
}
response, err := client.Get(url)
if err == nil {
io.Copy(io.Discard, response.Body)
response.Body.Close()
if response.StatusCode == http.StatusOK {
return nil
}
}
if time.Now().After(deadline) {
return fmt.Errorf(
"server did not answer %s within %s",
url,
serverStartTimeout,
)
}
time.Sleep(serverPollInterval)
}
}
func (s *server) stop() {
if s.command.Process != nil {
group := -s.command.Process.Pid
syscall.Kill(group, syscall.SIGTERM)
select {
case <-s.exited:
case <-time.After(serverShutdownGrace):
syscall.Kill(group, syscall.SIGKILL)
<-s.exited
}
}
s.logFile.Close()
}
@@ -0,0 +1,119 @@
package main
import (
"context"
"fmt"
"net/http"
"os"
"path/filepath"
"syscall"
"testing"
"time"
)
// bunSpawningAServer serves from a child process and then waits, which is the
// shape of `bun run preview`: the port belongs to something below the process
// the sweep started.
const bunSpawningAServer = `#!/bin/sh
port=""
previous=""
for argument in "$@"; do
if [ "$previous" = "--port" ]; then port="$argument"; fi
previous="$argument"
done
SWEEP_TEST_SERVE_PORT="$port" "%[1]s" &
wait
`
func TestServerStop_TakesTheProcessBelowItWithTheServer(t *testing.T) {
directory := t.TempDir()
testBinary, err := filepath.Abs(os.Args[0])
if err != nil {
t.Fatal(err)
}
bunPath := writeScript(
t,
filepath.Join(directory, "stub-bun"),
fmt.Sprintf(bunSpawningAServer, testBinary),
)
port := freePortRange(t, 1)
target := implementation{Name: "impl-01", Directory: directory, Port: port}
// Background rather than the test context: only stop() may end this
// server, or a leak would be hidden by the context being cancelled.
running, err := startServer(
context.Background(),
config{bunPath: bunPath},
target,
filepath.Join(directory, "serve.log"),
)
if err != nil {
t.Fatal(err)
}
t.Cleanup(
func() { syscall.Kill(-running.command.Process.Pid, syscall.SIGKILL) },
)
if err := running.waitReady(context.Background(), readinessURL(port)); err != nil {
t.Fatal(err)
}
running.stop()
client := &http.Client{Timeout: time.Second}
deadline := time.Now().Add(5 * time.Second)
for time.Now().Before(deadline) {
response, err := client.Get(readinessURL(port))
if err != nil {
return
}
response.Body.Close()
time.Sleep(100 * time.Millisecond)
}
t.Fatalf(
"port %d is still served after stop(): the server below bun outlived the sweep and holds the port",
port,
)
}
func TestServerWaitReady_ReportsAServerThatExited(t *testing.T) {
directory := t.TempDir()
bunPath := writeScript(
t,
filepath.Join(directory, "stub-bun"),
"#!/bin/sh\necho 'port is already in use' >&2\nexit 1\n",
)
port := freePortRange(t, 1)
running, err := startServer(
context.Background(),
config{bunPath: bunPath},
implementation{
Name: "impl-01",
Directory: directory,
Port: port,
},
filepath.Join(directory, "serve.log"),
)
if err != nil {
t.Fatal(err)
}
defer running.stop()
started := time.Now()
err = running.waitReady(context.Background(), readinessURL(port))
if err == nil {
t.Fatal(
"a server that exited should not be waited on until the start timeout",
)
}
if elapsed := time.Since(started); elapsed > 30*time.Second {
t.Errorf("waited %s for a server that had already exited", elapsed)
}
log, err := os.ReadFile(filepath.Join(directory, "serve.log"))
if err != nil {
t.Fatal(err)
}
if string(log) == "" {
t.Error("serve.log did not capture why the server exited")
}
}
@@ -0,0 +1,399 @@
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"time"
)
// implementation is one directory under --implementations, with the port it
// owns for the whole sweep. The port comes from the implementation's position
// in name order rather than from a pool, so the manifest can name the URL every
// run was served from before anything has been served.
type implementation struct {
Name string
Directory string
Port int
}
const implementationPrefix = "impl-"
func discoverImplementations(
directory string,
basePort int,
) ([]implementation, error) {
entries, err := os.ReadDir(directory)
if err != nil {
return nil, err
}
var found []implementation
for _, entry := range entries {
if !entry.IsDir() ||
!strings.HasPrefix(entry.Name(), implementationPrefix) {
continue
}
found = append(found, implementation{
Name: entry.Name(),
Directory: filepath.Join(directory, entry.Name()),
})
}
if len(found) == 0 {
return nil, fmt.Errorf(
"no %s* directories in %s",
implementationPrefix,
directory,
)
}
sort.Slice(
found,
func(i, j int) bool { return found[i].Name < found[j].Name },
)
if basePort+len(found)-1 > 65535 {
return nil, fmt.Errorf(
"--base-port %d leaves no room for %d implementations",
basePort,
len(found),
)
}
for index := range found {
found[index].Port = basePort + index
}
return found, nil
}
// resolveBinaries turns bun, campaign and sanderling into absolute paths before
// anything is installed. Each campaign runs from the sweep's own directory
// rather than the implementation's, so a relative --sanderling would otherwise
// resolve against the wrong one, and a binary that is missing altogether has to
// stop the sweep here rather than fail once per implementation and seed. Every
// one that is missing is named together, in flag order: stopping at the first
// turns that single stop into one rerun per missing binary.
func resolveBinaries(configuration *config) error {
var missing []error
for _, binary := range []struct {
name string
value *string
}{
{"--bun", &configuration.bunPath},
{"--campaign", &configuration.campaignPath},
{"--sanderling", &configuration.sanderlingPath},
} {
resolved, err := exec.LookPath(*binary.value)
if err != nil {
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
continue
}
absolute, err := filepath.Abs(resolved)
if err != nil {
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
continue
}
*binary.value = absolute
}
return errors.Join(missing...)
}
type sweep struct {
configuration config
stdout io.Writer
records io.Writer
mutex sync.Mutex
stalled int
failedRuns int
totalRuns int
}
func runSweep(
ctx context.Context,
configuration config,
stdout io.Writer,
) error {
if _, err := os.Stat(filepath.Join(configuration.outputDirectory, manifestFileName)); err == nil {
return fmt.Errorf(
"%s already exists in %s: pick a fresh --output so two sweeps do not share a directory",
manifestFileName,
configuration.outputDirectory,
)
}
implementations, err := discoverImplementations(
configuration.implementationsDirectory,
configuration.basePort,
)
if err != nil {
return err
}
if err := resolveBinaries(&configuration); err != nil {
return err
}
if _, err := os.Stat(configuration.specPath); err != nil {
return fmt.Errorf("--spec: %w", err)
}
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
return fmt.Errorf("create sweep dir: %w", err)
}
host, _ := os.Hostname()
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, implementations, host, time.Now().UTC())); err != nil {
return fmt.Errorf("write %s: %w", manifestFileName, err)
}
recordsFile, err := os.OpenFile(
filepath.Join(configuration.outputDirectory, recordsFileName),
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
0o644,
)
if err != nil {
return fmt.Errorf("open %s: %w", recordsFileName, err)
}
defer recordsFile.Close()
running := &sweep{
configuration: configuration,
stdout: stdout,
records: recordsFile,
}
fmt.Fprintf(
stdout,
"sweep: %d implementations, %d seeds each, %d at a time, %s\n",
len(
implementations,
),
len(configuration.seeds),
configuration.concurrency,
configuration.outputDirectory,
)
running.work(ctx, implementations)
fmt.Fprintf(
stdout,
"sweep complete: %d of %d implementations never ran, %d of %d campaigns failed\n",
running.stalled,
len(implementations),
running.failedRuns,
running.totalRuns,
)
if running.stalled > 0 || running.failedRuns > 0 {
return fmt.Errorf(
"%d of %d implementations never ran and %d of %d campaigns failed; see %s",
running.stalled,
len(implementations),
running.failedRuns,
running.totalRuns,
recordsFileName,
)
}
return nil
}
func (s *sweep) work(ctx context.Context, implementations []implementation) {
queue := make(chan implementation, len(implementations))
for _, target := range implementations {
queue <- target
}
close(queue)
workers := min(s.configuration.concurrency, len(implementations))
var waitGroup sync.WaitGroup
for range workers {
waitGroup.Add(1)
go func() {
defer waitGroup.Done()
for target := range queue {
if ctx.Err() != nil {
return
}
s.report(s.runImplementation(ctx, target))
}
}()
}
waitGroup.Wait()
}
// runImplementation carries one implementation from install to its last seed.
// Every failure it can meet is returned in the record: one implementation that
// cannot install, build or serve must not cost the other twenty-three their
// runs.
func (s *sweep) runImplementation(
ctx context.Context,
target implementation,
) (record implementationRecord) {
record = implementationRecord{
Name: target.Name,
Directory: target.Directory,
Port: target.Port,
StartedAt: time.Now().UTC(),
}
started := time.Now()
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
directory := filepath.Join(s.configuration.outputDirectory, target.Name)
if err := os.MkdirAll(directory, 0o755); err != nil {
record.FailedStage = stageInstall
record.Error = err.Error()
return record
}
for _, step := range []struct {
stage string
arguments []string
}{
{stageInstall, []string{"install"}},
{stageBuild, []string{"run", "build"}},
} {
logPath := filepath.Join(directory, step.stage+".log")
exitCode, err := runCommand(
ctx,
target.Directory,
s.configuration.bunPath,
step.arguments,
logPath,
)
if err != nil {
record.FailedStage = step.stage
record.Error = err.Error()
return record
}
if exitCode != 0 {
record.FailedStage = step.stage
record.Error = fmt.Sprintf(
"bun %s exited %d, see %s",
strings.Join(step.arguments, " "),
exitCode,
logPath,
)
return record
}
}
running, err := startServer(
ctx,
s.configuration,
target,
filepath.Join(directory, "serve.log"),
)
if err != nil {
record.FailedStage = stageServe
record.Error = err.Error()
return record
}
defer running.stop()
if err := running.waitReady(ctx, readinessURL(target.Port)); err != nil {
record.FailedStage = stageServe
record.Error = err.Error()
return record
}
for _, seed := range s.configuration.seeds {
if ctx.Err() != nil {
return record
}
record.Runs = append(record.Runs, s.runSeed(ctx, target, seed))
}
return record
}
func (s *sweep) runSeed(
ctx context.Context,
target implementation,
seed int64,
) (record runRecord) {
seedText := strconv.FormatInt(seed, 10)
directory := campaignDirectory(s.configuration, target, seedText)
record = runRecord{
Seed: seed,
URL: servedURL(target.Port, seedText),
CampaignDirectory: directory,
}
started := time.Now()
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
if err := os.MkdirAll(directory, 0o755); err != nil {
record.ExitCode = -1
record.LaunchError = err.Error()
return record
}
exitCode, err := runCommand(
ctx,
"",
s.configuration.campaignPath,
campaignArguments(
s.configuration,
target,
seedText,
),
filepath.Join(directory, "campaign.log"),
)
record.ExitCode = exitCode
if err != nil {
record.LaunchError = err.Error()
}
return record
}
// runCommand runs one step of the pipeline with its output in logPath. An
// empty directory keeps the sweep's own working directory, which is what the
// campaign tool gets: it has no reason to run inside an implementation.
func runCommand(
ctx context.Context,
directory, binary string,
arguments []string,
logPath string,
) (int, error) {
logFile, err := os.Create(logPath)
if err != nil {
return -1, err
}
defer logFile.Close()
command := exec.CommandContext(ctx, binary, arguments...)
command.Dir = directory
command.Stdout = logFile
command.Stderr = logFile
err = command.Run()
if err == nil {
return 0, nil
}
var exitError *exec.ExitError
if errors.As(err, &exitError) {
return exitError.ExitCode(), nil
}
return -1, err
}
func (s *sweep) report(record implementationRecord) {
s.mutex.Lock()
defer s.mutex.Unlock()
if record.FailedStage != "" {
s.stalled++
}
s.totalRuns += len(record.Runs)
for _, run := range record.Runs {
if run.ExitCode != 0 {
s.failedRuns++
}
}
if err := json.NewEncoder(s.records).Encode(record); err != nil {
fmt.Fprintf(s.stdout, "warning: %s record: %v\n", record.Name, err)
}
elapsed := time.Duration(record.MonotonicMillis) * time.Millisecond
if record.FailedStage != "" {
fmt.Fprintf(s.stdout, "%s port=%d failed at %s: %s (%s)\n",
record.Name, record.Port, record.FailedStage, record.Error, elapsed)
return
}
failed := 0
for _, run := range record.Runs {
if run.ExitCode != 0 {
failed++
}
}
fmt.Fprintf(s.stdout, "%s port=%d campaigns=%d failed=%d elapsed=%s\n",
record.Name, record.Port, len(record.Runs), failed, elapsed)
}
@@ -0,0 +1,154 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestDiscoverImplementations_NameOrderAndOnePortEach(t *testing.T) {
directory := t.TempDir()
for _, name := range []string{"impl-03", "impl-01", "impl-10", "impl-02", "scaffold", ".DS_Store"} {
if err := os.MkdirAll(filepath.Join(directory, name), 0o755); err != nil {
t.Fatal(err)
}
}
if err := os.WriteFile(filepath.Join(directory, "impl-notes.md"), []byte("not a directory"), 0o644); err != nil {
t.Fatal(err)
}
found, err := discoverImplementations(directory, 5300)
if err != nil {
t.Fatal(err)
}
want := []implementation{
{Name: "impl-01", Port: 5300},
{Name: "impl-02", Port: 5301},
{Name: "impl-03", Port: 5302},
{Name: "impl-10", Port: 5303},
}
if len(found) != len(want) {
t.Fatalf(
"got %d implementations, want %d: %v",
len(found),
len(want),
found,
)
}
for index, target := range found {
if target.Name != want[index].Name || target.Port != want[index].Port {
t.Errorf(
"position %d: got %s on %d, want %s on %d",
index,
target.Name,
target.Port,
want[index].Name,
want[index].Port,
)
}
if target.Directory != filepath.Join(directory, want[index].Name) {
t.Errorf("%s directory: got %q", target.Name, target.Directory)
}
}
}
func TestDiscoverImplementations_EmptyDirectoryIsRefused(t *testing.T) {
_, err := discoverImplementations(t.TempDir(), 5300)
if err == nil || !strings.Contains(err.Error(), "no impl-* directories") {
t.Fatalf("got %v, want a refusal naming impl-*", err)
}
}
// A binary that is not there fails once, before anything is installed, rather
// than twenty-four times after the sweep has spent its build time.
func TestRunSweep_StopsBeforeItInstallsAnythingWhenABinaryIsMissing(
t *testing.T,
) {
root := t.TempDir()
implementations := filepath.Join(root, "implementations")
if err := os.MkdirAll(filepath.Join(implementations, "impl-01"), 0o755); err != nil {
t.Fatal(err)
}
output := filepath.Join(root, "campaigns")
configuration := config{
implementationsDirectory: implementations,
outputDirectory: output,
basePort: 5300,
concurrency: 1,
bunPath: writeScript(
t,
filepath.Join(root, "stub-bun"),
"#!/bin/sh\nexit 0\n",
),
campaignPath: "campaign-that-is-not-installed",
sanderlingPath: writeScript(
t,
filepath.Join(root, "stub-sanderling"),
"#!/bin/sh\nexit 0\n",
),
}
err := runSweep(t.Context(), configuration, os.Stdout)
if err == nil || !strings.Contains(err.Error(), "--campaign") {
t.Fatalf("got %v, want the missing campaign binary named", err)
}
if _, err := os.Stat(output); !os.IsNotExist(err) {
t.Errorf(
"the sweep created %s before it checked it could run: %v",
output,
err,
)
}
}
// Two binaries missing is one rerun, not two: the operator is told about both
// at once, in flag order, whatever order the check happened to walk.
func TestResolveBinaries_NamesEveryMissingBinaryInFlagOrder(t *testing.T) {
configuration := config{
bunPath: writeScript(
t,
filepath.Join(t.TempDir(), "stub-bun"),
"#!/bin/sh\nexit 0\n",
),
campaignPath: "campaign-that-is-not-installed",
sanderlingPath: "sanderling-that-is-not-installed",
}
err := resolveBinaries(&configuration)
if err == nil {
t.Fatal("got no error, want both missing binaries named")
}
message := err.Error()
campaign := strings.Index(message, "--campaign")
sanderling := strings.Index(message, "--sanderling")
if campaign < 0 || sanderling < 0 {
t.Fatalf("got %q, want both --campaign and --sanderling named", message)
}
if campaign > sanderling {
t.Errorf("got %q, want --campaign named before --sanderling", message)
}
if strings.Contains(message, "--bun") {
t.Errorf("got %q, want the bun that resolved left out", message)
}
}
func TestRunSweep_RefusesADirectoryThatAlreadyHoldsASweep(t *testing.T) {
implementations := t.TempDir()
if err := os.MkdirAll(filepath.Join(implementations, "impl-01"), 0o755); err != nil {
t.Fatal(err)
}
output := t.TempDir()
if err := os.WriteFile(filepath.Join(output, manifestFileName), []byte("{}"), 0o644); err != nil {
t.Fatal(err)
}
configuration := config{
implementationsDirectory: implementations,
outputDirectory: output,
basePort: 5300,
concurrency: 1,
}
if err := runSweep(t.Context(), configuration, os.Stdout); err == nil ||
!strings.Contains(err.Error(), "already exists") {
t.Fatalf("got %v, want a refusal to reuse the directory", err)
}
}
+549
View File
@@ -0,0 +1,549 @@
// Command label-coverage reports how much of an app's interactive surface a
// spec can address, from the hierarchies a run already recorded.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"flag"
"fmt"
"io"
"io/fs"
"os"
"path/filepath"
"sort"
"strings"
"unicode"
)
type element struct {
ResourceID string `json:"resourceId"`
Text string `json:"text"`
Description string `json:"description"`
Class string `json:"class"`
Package string `json:"package"`
Clickable bool `json:"clickable"`
Editable bool `json:"editable"`
Enabled bool `json:"enabled"`
}
type step struct {
Index int `json:"step"`
Screen string `json:"screen"`
Hierarchy *struct {
Elements []element `json:"elements"`
} `json:"hierarchy"`
}
// Counts splits a screen's interactive elements by the strongest selector that
// can reach them. Text is separated from identifier and description because a
// row labelled only by the customer name it displays is addressable in one run
// and gone in the next, which is not the same thing as being addressable. A
// description that carries the data with it is separated for the same reason.
type Counts struct {
Screen string `json:"screen"`
Observations int `json:"observations"`
Elements int `json:"elements"`
Interactive int `json:"interactive"`
ByIdentifier int `json:"by_identifier"`
ByDataID int `json:"by_data_carrying_identifier"`
ByDescription int `json:"by_description"`
ByVolatile int `json:"by_volatile_description"`
ByTextOnly int `json:"by_text_only"`
Unaddressable int `json:"unaddressable"`
AmbiguousIDs int `json:"ambiguous_identifiers"`
// Needing lists the interactive elements no durable selector reaches, which
// is the work list for a label pass rather than a statistic about it.
Needing []string `json:"needing_labels,omitempty"`
}
func (c Counts) stableShare() float64 {
if c.Interactive == 0 {
return 0
}
return float64(
c.ByIdentifier+c.ByDescription,
) / float64(
c.Interactive,
) * 100
}
func main() {
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
show := flag.Int(
"show",
0,
"list up to this many controls per screen that no durable selector reaches",
)
pkg := flag.String(
"package",
"",
"count only elements belonging to this package; system UI is dropped either way",
)
flag.Usage = func() {
fmt.Fprintln(
os.Stderr,
"usage: label-coverage [--json] [--show N] <trace.jsonl | run directory> ...",
)
}
flag.Parse()
if flag.NArg() == 0 {
flag.Usage()
os.Exit(2)
}
traces, err := collect(flag.Args())
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if len(traces) == 0 {
fmt.Fprintln(os.Stderr, "no trace.jsonl found under the given paths")
os.Exit(1)
}
byScreen := map[string]*Counts{}
for _, path := range traces {
if err := accumulate(path, byScreen, *pkg); err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
os.Exit(1)
}
}
screens := make([]Counts, 0, len(byScreen))
for _, counts := range byScreen {
screens = append(screens, *counts)
}
sort.Slice(
screens,
func(i, j int) bool { return screens[i].Screen < screens[j].Screen },
)
if *jsonOut {
encoder := json.NewEncoder(os.Stdout)
encoder.SetIndent("", " ")
if err := encoder.Encode(screens); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
return
}
render(os.Stdout, screens, len(traces), *show)
}
func collect(paths []string) ([]string, error) {
var traces []string
for _, path := range paths {
info, err := os.Stat(path)
if err != nil {
return nil, err
}
if !info.IsDir() {
traces = append(traces, path)
continue
}
err = filepath.WalkDir(
path,
func(candidate string, entry fs.DirEntry, err error) error {
if err != nil {
return err
}
if !entry.IsDir() && entry.Name() == "trace.jsonl" {
traces = append(traces, candidate)
}
return nil
},
)
if err != nil {
return nil, err
}
}
sort.Strings(traces)
return traces, nil
}
// accumulate folds one trace into byScreen, counting each distinct hierarchy
// once. A run that idles on a screen observes it many times, and summing those
// observations would report the screen the explorer sat on rather than the
// screen with the most unlabelled controls.
// systemPackages own nodes that share the screen with the app under test. A
// status bar contributes three well-labelled controls to every capture, and
// counting them lifts an app with no identifiers at all off the floor.
var systemPackages = map[string]bool{
"com.android.systemui": true,
"android": true,
"com.google.android.inputmethod.latin": true,
"com.android.inputmethod.latin": true,
"com.google.android.apps.nexuslauncher": true,
"com.google.android.googlequicksearchbox": true,
}
// scoped drops the windows the app under test does not own. Most of an app's
// own nodes carry no package attribute at all, since only the window roots are
// stamped with one, so an empty package is treated as belonging to the app
// rather than filtered out. Filtering on an exact match alone discards the
// entire application and reports a clean zero.
func scoped(elements []element, pkg string) []element {
kept := make([]element, 0, len(elements))
for _, item := range elements {
if item.Package == "" {
kept = append(kept, item)
continue
}
if pkg != "" {
if item.Package == pkg {
kept = append(kept, item)
}
continue
}
if !systemPackages[item.Package] {
kept = append(kept, item)
}
}
return kept
}
func accumulate(path string, byScreen map[string]*Counts, pkg string) error {
file, err := os.Open(path)
if err != nil {
return err
}
defer file.Close()
seen := map[string]bool{}
decoder := json.NewDecoder(file)
for {
var recorded step
if err := decoder.Decode(&recorded); err != nil {
if errors.Is(err, io.EOF) {
break
}
return err
}
if recorded.Hierarchy == nil || len(recorded.Hierarchy.Elements) == 0 {
continue
}
elements := scoped(recorded.Hierarchy.Elements, pkg)
if len(elements) == 0 {
continue
}
screen := recorded.Screen
if screen == "" {
screen = shape(elements)
}
key := screen + "\x00" + signature(elements)
if seen[key] {
continue
}
seen[key] = true
observed := measure(elements)
observed.Screen = screen
counts, ok := byScreen[screen]
if !ok {
observed.Observations = 1
byScreen[screen] = &observed
continue
}
if observed.Interactive > counts.Interactive {
observed.Observations = counts.Observations
*counts = observed
}
counts.Observations++
}
return nil
}
// shape names a screen the app did not name itself, by the set of distinct
// controls it shows. Repetition is dropped deliberately: a ledger holding three
// rows and the same ledger holding five is one screen, not two.
func shape(elements []element) string {
distinct := map[string]bool{}
for _, item := range elements {
key := item.ResourceID
if key == "" {
key = item.Class + "/" + item.Description
}
distinct[key] = true
}
keys := make([]string, 0, len(distinct))
for key := range distinct {
keys = append(keys, key)
}
sort.Strings(keys)
sum := sha256.Sum256([]byte(strings.Join(keys, "\x00")))
return "shape:" + hex.EncodeToString(sum[:])[:8]
}
// volatile reports whether a description carries the data it labels. Compose
// merges a row's children into one description, so a ledger row arrives
// labelled with the customer name, the balance and a relative age that reprices
// itself every month. Such a description exists, which is why a presence check
// scores it as a selector, and it is not one: the run that recorded it is the
// only run it matches. The test is a heuristic and deliberately blunt, since
// the alternative is to call every one of them stable.
func volatile(description string) bool {
trimmed := strings.TrimSpace(description)
// A one-character label is an avatar initial, and it is the first letter of
// a name the run happened to observe. Renaming the record changes it. The
// digit test below catches an initial drawn from a numeric name and would
// miss every alphabetic one, which is how four of them were counted as
// durable before this was noticed on a real ledger.
if len([]rune(trimmed)) == 1 &&
(unicode.IsLetter([]rune(trimmed)[0]) || unicode.IsDigit([]rune(trimmed)[0])) {
return true
}
for _, character := range trimmed {
if unicode.IsDigit(character) {
return true
}
}
lowered := strings.ToLower(trimmed)
for _, marker := range []string{" ago", "since ", "yesterday", "today", "tomorrow", "last ", "due ", "minute", "hour", "day", "week", "month", "year"} {
if strings.Contains(lowered, marker) {
return true
}
}
return false
}
// dataCarryingID reports whether an identifier embeds the record it names. A
// list row tagged `customer_row_<uuid>` names its role durably and its instance
// not at all, so an exact match on it survives exactly one run. The prefix is
// still worth something, which is what `idPrefix:` is for, but a measure that
// counts the whole string as a durable selector overstates what a spec can say.
// Scoped to the local name so an Android package prefix cannot trip it, and
// tuned to leave ordinary names like `button2` alone.
func dataCarryingID(identifier string) bool {
local := identifier
if index := strings.LastIndex(local, "/"); index >= 0 {
local = local[index+1:]
}
digits := 0
for _, character := range local {
if unicode.IsDigit(character) {
digits++
if digits >= 4 {
return true
}
continue
}
digits = 0
}
hex, groups := 0, 0
for _, character := range local + "-" {
if isHexDigit(character) {
hex++
continue
}
if hex >= 4 {
groups++
}
hex = 0
}
return groups >= 2
}
func isHexDigit(character rune) bool {
return unicode.IsDigit(character) ||
(character >= 'a' && character <= 'f') ||
(character >= 'A' && character <= 'F')
}
// keypadDigits reports whether this screen shows a numeric keypad, in which case
// its single-character digit labels name fixed keys rather than the first
// character of somebody's name. Without the screen for context the two are
// indistinguishable: an avatar initial and a calculator key are both one
// character, and only one of them survives a data change.
func keypadDigits(elements []element) bool {
seen := map[rune]bool{}
for _, item := range elements {
label := strings.TrimSpace(item.Description)
if label == "" {
label = strings.TrimSpace(item.Text)
}
runes := []rune(label)
if len(runes) == 1 && unicode.IsDigit(runes[0]) {
seen[runes[0]] = true
}
}
return len(seen) >= 6
}
func isSingleDigit(label string) bool {
runes := []rune(strings.TrimSpace(label))
return len(runes) == 1 && unicode.IsDigit(runes[0])
}
func measure(elements []element) Counts {
counts := Counts{Elements: len(elements)}
keypad := keypadDigits(elements)
identifiers := map[string]int{}
for _, item := range elements {
if !item.Clickable && !item.Editable {
continue
}
counts.Interactive++
switch {
case item.ResourceID != "" && !dataCarryingID(item.ResourceID):
counts.ByIdentifier++
identifiers[item.ResourceID]++
case item.ResourceID != "":
counts.ByDataID++
counts.Needing = append(counts.Needing, describe(item))
case item.Description != "" && keypad && isSingleDigit(item.Description):
counts.ByDescription++
case item.Description != "" && !volatile(item.Description):
counts.ByDescription++
case item.Description != "":
counts.ByVolatile++
counts.Needing = append(counts.Needing, describe(item))
case item.Text != "":
counts.ByTextOnly++
counts.Needing = append(counts.Needing, describe(item))
default:
counts.Unaddressable++
counts.Needing = append(counts.Needing, describe(item))
}
}
for _, repeats := range identifiers {
if repeats > 1 {
counts.AmbiguousIDs += repeats
}
}
return counts
}
func describe(item element) string {
class := item.Class
if class == "" {
class = "(no class)"
}
switch {
case item.ResourceID != "" && dataCarryingID(item.ResourceID):
return fmt.Sprintf(
"%s id=%q (names its record, not its role)",
class,
item.ResourceID,
)
case item.Description != "":
return fmt.Sprintf(
"%s desc=%q (carries its own data)",
class,
item.Description,
)
case item.Text != "":
return fmt.Sprintf("%s text=%q", class, item.Text)
default:
return class + " (no label at all)"
}
}
func signature(elements []element) string {
var builder strings.Builder
for _, item := range elements {
builder.WriteString(item.Class)
builder.WriteByte('|')
builder.WriteString(item.ResourceID)
builder.WriteByte('|')
builder.WriteString(item.Description)
builder.WriteByte(';')
}
return builder.String()
}
func render(out io.Writer, screens []Counts, traces int, show int) {
fmt.Fprintf(out, "%d trace(s), %d screen(s)\n\n", traces, len(screens))
fmt.Fprintf(
out,
"%-28s %5s %5s %5s %4s %5s %5s %5s %5s %7s\n",
"screen",
"obs",
"inter",
"id",
"id~",
"desc",
"vol",
"text",
"none",
"stable%",
)
total := Counts{Screen: "TOTAL"}
for _, screen := range screens {
fmt.Fprintf(
out,
"%-28s %5d %5d %5d %4d %5d %5d %5d %5d %6.1f%%\n",
truncate(
screen.Screen,
28,
),
screen.Observations,
screen.Interactive,
screen.ByIdentifier,
screen.ByDataID,
screen.ByDescription,
screen.ByVolatile,
screen.ByTextOnly,
screen.Unaddressable,
screen.stableShare(),
)
total.Observations += screen.Observations
total.Elements += screen.Elements
total.Interactive += screen.Interactive
total.ByIdentifier += screen.ByIdentifier
total.ByDataID += screen.ByDataID
total.ByDescription += screen.ByDescription
total.ByVolatile += screen.ByVolatile
total.ByTextOnly += screen.ByTextOnly
total.Unaddressable += screen.Unaddressable
total.AmbiguousIDs += screen.AmbiguousIDs
}
fmt.Fprintf(out, "%-28s %5d %5d %5d %4d %5d %5d %5d %5d %6.1f%%\n",
total.Screen, total.Observations, total.Interactive, total.ByIdentifier,
total.ByDataID, total.ByDescription, total.ByVolatile, total.ByTextOnly,
total.Unaddressable, total.stableShare())
fmt.Fprintf(
out,
"\n%d interactive element(s) carry an identifier, which is the number that survives a data change\n",
total.ByIdentifier,
)
fmt.Fprintf(
out,
"%d share a resource id with another on the same screen\n",
total.AmbiguousIDs,
)
if show <= 0 {
return
}
for _, screen := range screens {
if len(screen.Needing) == 0 {
continue
}
fmt.Fprintf(
out,
"\n%s, %d control(s) no durable selector reaches:\n",
screen.Screen,
len(screen.Needing),
)
for index, item := range screen.Needing {
if index == show {
fmt.Fprintf(
out,
" ... and %d more\n",
len(screen.Needing)-show,
)
break
}
fmt.Fprintf(out, " %s\n", item)
}
}
}
func truncate(value string, width int) string {
if len(value) <= width {
return value
}
return value[:width-1] + "~"
}
@@ -0,0 +1,446 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestMeasureRanksSelectorsByDurability(t *testing.T) {
counts := measure([]element{
{
Class: "Button",
ResourceID: "app:id/save",
Text: "Save",
Clickable: true,
},
{
Class: "Button",
Description: "Filter",
Text: "Filter",
Clickable: true,
},
{Class: "View", Text: "Ramesh Kumar", Clickable: true},
{Class: "View", Clickable: true},
{Class: "EditText", ResourceID: "app:id/amount", Editable: true},
{Class: "TextView", Text: "Balance", Clickable: false},
})
if counts.Elements != 6 {
t.Fatalf("elements: want 6, got %d", counts.Elements)
}
if counts.Interactive != 5 {
t.Fatalf("interactive: want 5, got %d", counts.Interactive)
}
if counts.ByIdentifier != 2 {
t.Errorf("by identifier: want 2, got %d", counts.ByIdentifier)
}
if counts.ByDescription != 1 {
t.Errorf("by description: want 1, got %d", counts.ByDescription)
}
if counts.ByTextOnly != 1 {
t.Errorf("by text only: want 1, got %d", counts.ByTextOnly)
}
if counts.Unaddressable != 1 {
t.Errorf("unaddressable: want 1, got %d", counts.Unaddressable)
}
}
// The row description a merged Compose semantics node produces carries the
// balance and a relative age that reprices itself every month. It is present,
// so a presence check scores it as a selector; it matches only the run that
// recorded it.
func TestVolatileDescriptionsAreNotCountedAsStable(t *testing.T) {
row := "Ramesh ji, 95, Pending Collection Since 15 months, Due"
if !volatile(row) {
t.Fatalf("expected %q to be treated as data-carrying", row)
}
counts := measure([]element{
{Class: "Button", Description: row, Clickable: true},
{Class: "Button", Description: "Filter", Clickable: true},
})
if counts.ByVolatile != 1 {
t.Errorf("volatile: want 1, got %d", counts.ByVolatile)
}
if counts.ByDescription != 1 {
t.Errorf("stable description: want 1, got %d", counts.ByDescription)
}
if got := counts.stableShare(); got != 50 {
t.Fatalf("stable share: want 50, got %.1f", got)
}
if len(counts.Needing) != 1 ||
!strings.Contains(counts.Needing[0], "carries its own data") {
t.Fatalf(
"the volatile row belongs on the work list, got %v",
counts.Needing,
)
}
}
func TestVolatileLeavesPlainLabelsAlone(t *testing.T) {
for _, label := range []string{"Filter", "Search", "Add Relationship", "Share", "Skip"} {
if volatile(label) {
t.Errorf("%q should count as a durable label", label)
}
}
for _, label := range []string{"₹95", "2 days ago", "Due today", "Pending Since 15 months", "Edited on 11 Jun 2026"} {
if !volatile(label) {
t.Errorf("%q carries data and should not count as durable", label)
}
}
}
func TestStableShareExcludesDataDependentText(t *testing.T) {
counts := measure([]element{
{ResourceID: "app:id/add", Clickable: true},
{Text: "Ramesh Kumar", Clickable: true},
{Text: "Suresh Patel", Clickable: true},
{Clickable: true},
})
if got := counts.stableShare(); got != 25 {
t.Fatalf("stable share: want 25, got %.1f", got)
}
}
func TestMeasureCountsRepeatedIdentifiersAsAmbiguous(t *testing.T) {
counts := measure([]element{
{ResourceID: "app:id/row", Text: "Ramesh", Clickable: true},
{ResourceID: "app:id/row", Text: "Suresh", Clickable: true},
{ResourceID: "app:id/add", Clickable: true},
})
if counts.AmbiguousIDs != 2 {
t.Fatalf("ambiguous: want 2, got %d", counts.AmbiguousIDs)
}
}
func TestAccumulateKeepsRichestObservationPerScreen(t *testing.T) {
trace := writeTrace(
t,
`{"step":0,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true}]}}
{"step":1,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true},{"text":"Ramesh","clickable":true},{"class":"View","clickable":true}]}}
{"step":2,"screen":"home","hierarchy":{"elements":[{"description":"Filter","clickable":true}]}}
`,
)
byScreen := map[string]*Counts{}
if err := accumulate(trace, byScreen, ""); err != nil {
t.Fatal(err)
}
ledger := byScreen["ledger"]
if ledger.Interactive != 3 {
t.Fatalf("ledger interactive: want 3, got %d", ledger.Interactive)
}
if ledger.Observations != 2 {
t.Fatalf("ledger observations: want 2, got %d", ledger.Observations)
}
if ledger.Unaddressable != 1 {
t.Errorf("ledger unaddressable: want 1, got %d", ledger.Unaddressable)
}
if byScreen["home"].ByDescription != 1 {
t.Errorf(
"home by description: want 1, got %d",
byScreen["home"].ByDescription,
)
}
}
// An idling probe observes one screen many times. Counting each observation
// would report how long the explorer sat there rather than what it could reach.
func TestAccumulateCountsIdenticalHierarchiesOnce(t *testing.T) {
line := `{"step":0,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true}]}}` + "\n"
byScreen := map[string]*Counts{}
if err := accumulate(writeTrace(t, strings.Repeat(line, 5)), byScreen, ""); err != nil {
t.Fatal(err)
}
if got := byScreen["ledger"].Observations; got != 1 {
t.Fatalf("observations: want 1, got %d", got)
}
}
func TestAccumulateSkipsStepsWithoutHierarchy(t *testing.T) {
byScreen := map[string]*Counts{}
err := accumulate(writeTrace(t, `{"step":0,"screen":"ledger"}
{"step":1,"screen":"ledger","hierarchy":{"elements":[]}}
`), byScreen, "")
if err != nil {
t.Fatal(err)
}
if len(byScreen) != 0 {
t.Fatalf("want no screens, got %d", len(byScreen))
}
}
func TestShapeIgnoresRepetitionButNotComposition(t *testing.T) {
threeRows := shape([]element{
{ResourceID: "app:id/list"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
})
fiveRows := shape([]element{
{ResourceID: "app:id/list"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/row"},
})
if threeRows != fiveRows {
t.Fatalf(
"row count changed the screen identity: %s vs %s",
threeRows,
fiveRows,
)
}
withDialog := shape([]element{
{ResourceID: "app:id/list"},
{ResourceID: "app:id/row"},
{ResourceID: "app:id/confirm_dialog"},
})
if withDialog == threeRows {
t.Fatal(
"a screen showing a dialog must not collapse into the screen behind it",
)
}
}
func TestAccumulateSeparatesUnnamedScreensByShape(t *testing.T) {
byScreen := map[string]*Counts{}
err := accumulate(
writeTrace(
t,
`{"step":0,"hierarchy":{"elements":[{"resourceId":"app:id/ledger","clickable":true}]}}
{"step":1,"hierarchy":{"elements":[{"resourceId":"app:id/add_transaction","clickable":true}]}}
`,
),
byScreen,
"",
)
if err != nil {
t.Fatal(err)
}
if len(byScreen) != 2 {
t.Fatalf("want 2 screens, got %d", len(byScreen))
}
for name := range byScreen {
if !strings.HasPrefix(name, "shape:") {
t.Errorf("unnamed screen should be keyed by shape, got %q", name)
}
}
}
func TestCollectFindsTracesUnderRunDirectories(t *testing.T) {
root := t.TempDir()
for _, run := range []string{"run-1", "run-2"} {
directory := filepath.Join(root, run)
if err := os.MkdirAll(directory, 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte("{}\n"), 0o644); err != nil {
t.Fatal(err)
}
}
traces, err := collect([]string{root})
if err != nil {
t.Fatal(err)
}
if len(traces) != 2 {
t.Fatalf("want 2 traces, got %d: %v", len(traces), traces)
}
}
func writeTrace(t *testing.T, content string) string {
t.Helper()
path := filepath.Join(t.TempDir(), "trace.jsonl")
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
t.Fatal(err)
}
return path
}
// A status bar contributes three well-labelled controls to every capture. An
// app with no identifiers of its own scores 100 percent if they are counted.
func TestAccumulateDropsSystemUIByDefault(t *testing.T) {
trace := writeTrace(
t,
`{"step":0,"screen":"ledger","hierarchy":{"elements":[
{"resourceId":"com.android.systemui:id/clock","package":"com.android.systemui","clickable":true},
{"resourceId":"com.android.systemui:id/battery","package":"com.android.systemui","clickable":true},
{"class":"android.view.View","package":"in.okcredit.merchant.debug","clickable":true}]}}
`,
)
byScreen := map[string]*Counts{}
if err := accumulate(trace, byScreen, ""); err != nil {
t.Fatal(err)
}
counts := byScreen["ledger"]
if counts.Interactive != 1 {
t.Fatalf(
"interactive: want 1 after dropping system UI, got %d",
counts.Interactive,
)
}
if counts.ByIdentifier != 0 {
t.Fatalf(
"system identifiers leaked into the app's score: %d",
counts.ByIdentifier,
)
}
if got := counts.stableShare(); got != 0 {
t.Fatalf("stable share: want 0, got %.1f", got)
}
}
func TestAccumulatePackageFlagPinsExactly(t *testing.T) {
trace := writeTrace(
t,
`{"step":0,"screen":"ledger","hierarchy":{"elements":[
{"resourceId":"other:id/x","package":"com.other.app","clickable":true},
{"class":"android.view.View","package":"in.okcredit.merchant.debug","clickable":true}]}}
`,
)
byScreen := map[string]*Counts{}
if err := accumulate(trace, byScreen, "in.okcredit.merchant.debug"); err != nil {
t.Fatal(err)
}
if got := byScreen["ledger"].Interactive; got != 1 {
t.Fatalf("interactive: want 1, got %d", got)
}
}
// Only window roots carry a package attribute in a hierarchy dump. Treating an
// unstamped node as foreign discards the application and reports a clean zero.
func TestScopedKeepsUnstampedNodes(t *testing.T) {
elements := []element{
{
Package: "com.android.systemui",
ResourceID: "sysui:id/clock",
Clickable: true,
},
{Package: "in.okcredit.merchant.debug", ResourceID: "app:id/root"},
{Class: "android.view.View", Clickable: true},
}
if got := len(scoped(elements, "")); got != 2 {
t.Fatalf("denylist mode: want 2 kept, got %d", got)
}
if got := len(scoped(elements, "in.okcredit.merchant.debug")); got != 2 {
t.Fatalf("pinned mode: want 2 kept, got %d", got)
}
if got := len(scoped(elements, "com.other.app")); got != 1 {
t.Fatalf(
"pinned to a foreign package: want 1 unstamped node kept, got %d",
got,
)
}
}
// An avatar renders the first letter of the name beside it, so a one-character
// description is the name in disguise. Four of these were scored as durable on
// a real ledger before the rule existed.
func TestSingleCharacterDescriptionsAreAvatarInitials(t *testing.T) {
for _, initial := range []string{"R", "C", "T", "A", "9", " R "} {
if !volatile(initial) {
t.Errorf(
"%q is an avatar initial and changes when the record is renamed",
initial,
)
}
}
for _, keypad := range []string{"+", "=", "AC"} {
if volatile(keypad) {
t.Errorf(
"%q is a fixed control label and should count as durable",
keypad,
)
}
}
}
// A row tagged with its record's uuid names its role durably and its instance
// not at all, so an exact match on it survives exactly one run.
func TestDataCarryingIdentifiers(t *testing.T) {
for _, identifier := range []string{
"customer_row_5f338c10-feef-411c-a070-8999b4890a62",
"in.okcredit.merchant.debug:id/customer_row_5f338c10-feef-411c-a070-8999b4890a62",
"txn_20260812",
} {
if !dataCarryingID(identifier) {
t.Errorf("%q embeds the record it names", identifier)
}
}
for _, identifier := range []string{
"customer_supplier_list",
"summary_card",
"in.okcredit.merchant.debug:id/buttonLogin",
"button2",
"add_relationship",
} {
if dataCarryingID(identifier) {
t.Errorf("%q names a role and should count as durable", identifier)
}
}
}
func TestMeasureSeparatesDataCarryingIdentifiers(t *testing.T) {
counts := measure([]element{
{Class: "View", ResourceID: "summary_card", Clickable: true},
{
Class: "View",
ResourceID: "customer_row_5f338c10-feef-411c-a070-8999b4890a62",
Clickable: true,
},
{
Class: "View",
ResourceID: "customer_row_8ea26690-f809-4d23-8560-9cca6d1a5bcd",
Clickable: true,
},
})
if counts.ByIdentifier != 1 {
t.Errorf("durable identifiers: want 1, got %d", counts.ByIdentifier)
}
if counts.ByDataID != 2 {
t.Errorf("data-carrying identifiers: want 2, got %d", counts.ByDataID)
}
if got := counts.stableShare(); got < 33 || got > 34 {
t.Errorf("stable share: want about 33.3, got %.1f", got)
}
if len(counts.Needing) != 2 {
t.Fatalf(
"both row identifiers belong on the work list, got %v",
counts.Needing,
)
}
}
// A calculator key and an avatar initial are both one character. Only the
// screen they sit on tells them apart, so the keypad rule needs that context.
func TestKeypadDigitsAreDurableButAvatarInitialsAreNot(t *testing.T) {
keypad := make([]element, 0, 10)
for _, key := range []string{"1", "2", "3", "4", "5", "6", "7", "8", "9", "0"} {
keypad = append(
keypad,
element{Class: "Button", Description: key, Clickable: true},
)
}
counts := measure(keypad)
if counts.ByDescription != 10 {
t.Fatalf(
"keypad keys are fixed labels: want 10 durable, got %d",
counts.ByDescription,
)
}
counts = measure([]element{
{Class: "Button", Description: "R", Clickable: true},
{Class: "Button", Description: "9", Clickable: true},
{Class: "Button", Description: "Filter", Clickable: true},
})
if counts.ByVolatile != 2 {
t.Fatalf(
"avatar initials carry data: want 2 volatile, got %d",
counts.ByVolatile,
)
}
if counts.ByDescription != 1 {
t.Fatalf("durable descriptions: want 1, got %d", counts.ByDescription)
}
}
+223
View File
@@ -0,0 +1,223 @@
package main
import (
"bufio"
"encoding/json"
"fmt"
"os"
"path/filepath"
"sort"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
const maxStepBytes = 64 * 1024 * 1024
// loadedRun is one run directory: its meta, the steps in file order, and the
// synthetic end-of-run record the runner writes when Finalize convicted a
// liveness obligation.
type loadedRun struct {
Directory string
Meta trace.Meta
Steps []trace.Step
Finalize *trace.Step
}
// loadRun reads a run directory and refuses anything E3 cannot replay. A step
// written before the format change carries no element depths, so its hierarchy
// decodes with a nil root and every selector resolves to nothing: the refusal
// has to name the version rather than let the replay report an empty screen.
func loadRun(directory string) (loadedRun, error) {
metaBody, err := os.ReadFile(filepath.Join(directory, "meta.json"))
if err != nil {
return loadedRun{}, fmt.Errorf("read meta: %w", err)
}
var meta trace.Meta
if err := json.Unmarshal(metaBody, &meta); err != nil {
return loadedRun{}, fmt.Errorf("decode meta: %w", err)
}
file, err := os.Open(filepath.Join(directory, "trace.jsonl"))
if err != nil {
return loadedRun{}, fmt.Errorf("open trace: %w", err)
}
defer file.Close()
run := loadedRun{Directory: directory, Meta: meta}
scanner := bufio.NewScanner(file)
scanner.Buffer(make([]byte, 0, 1024*1024), maxStepBytes)
line := 0
for scanner.Scan() {
line++
if len(scanner.Bytes()) == 0 {
continue
}
var step trace.Step
if err := json.Unmarshal(scanner.Bytes(), &step); err != nil {
return loadedRun{}, fmt.Errorf(
"decode step on line %d: %w",
line,
err,
)
}
if step.TraceVersion != trace.TraceVersion {
return loadedRun{}, fmt.Errorf(
"step %d is trace_version %d, and E3 replays version %d only: "+
"an older step stores no element depths, so its hierarchy decodes "+
"with a nil root and every selector resolves to nothing on replay",
step.Index, step.TraceVersion, trace.TraceVersion)
}
run.Steps = append(run.Steps, step)
}
if err := scanner.Err(); err != nil {
return loadedRun{}, fmt.Errorf("read trace: %w", err)
}
if len(run.Steps) == 0 {
return loadedRun{}, fmt.Errorf("trace has no steps")
}
if last := run.Steps[len(run.Steps)-1]; len(run.Steps) > 1 &&
last.Hierarchy == nil &&
len(last.Violations) > 0 {
run.Finalize = &last
run.Steps = run.Steps[:len(run.Steps)-1]
}
if meta.Seed == 0 {
return loadedRun{}, fmt.Errorf(
"meta records no seed, so the run's bundle cannot be reproduced",
)
}
return run, nil
}
// finalizeIndex is the step index the runner gives its end-of-run record: one
// past the last step it wrote.
func (r loadedRun) finalizeIndex() int {
return r.Steps[len(r.Steps)-1].Index + 1
}
// extractorFold reconstructs every extractor's value at each step from the
// recorded per-step diffs. A run's trace stores what changed, so the value an
// extractor held at a step is the last change at or before it; an extractor
// that never changed held JSON null throughout, which is what an unwritten
// diff means.
func extractorFold(
steps []trace.Step,
names []string,
) ([]map[int]json.RawMessage, error) {
index := make(map[string]int, len(names))
for position, name := range names {
index[name] = position
}
current := make(map[int]json.RawMessage, len(names))
for position := range names {
current[position] = json.RawMessage("null")
}
folded := make([]map[int]json.RawMessage, len(steps))
for step := range steps {
for name, change := range steps[step].ExtractorChanges {
position, ok := index[name]
if !ok {
return nil, fmt.Errorf(
"step %d records extractor %q, which the spec does not register; "+
"the trace and the spec are not the same bundle",
steps[step].Index,
name,
)
}
current[position] = change.Curr
}
snapshot := make(map[int]json.RawMessage, len(current))
for position, value := range current {
snapshot[position] = value
}
folded[step] = snapshot
}
return folded, nil
}
// lastActionFor rebuilds the action the runner had applied before the next
// step observed. An action the runner chose but never dispatched left
// state.lastAction null, and the recorded skip reason is what says so.
func lastActionFor(step trace.Step) *verifier.Action {
if step.NextAction == nil || step.ActionSkipped != "" {
return nil
}
recorded := *step.NextAction
action := verifier.Action{
Kind: verifier.ActionKind(recorded.Kind),
On: recorded.Selector,
Text: recorded.Text,
X: recorded.X,
Y: recorded.Y,
FromX: recorded.FromX,
FromY: recorded.FromY,
ToX: recorded.ToX,
ToY: recorded.ToY,
Key: recorded.Key,
DurationMillis: recorded.DurationMillis,
}
return &action
}
func traceLogs(entries []trace.LogEntry) []verifier.LogEntry {
if len(entries) == 0 {
return nil
}
logs := make([]verifier.LogEntry, 0, len(entries))
for _, entry := range entries {
logs = append(logs, verifier.LogEntry{
UnixMillis: entry.UnixMillis,
Level: entry.Level,
Tag: entry.Tag,
Message: entry.Message,
})
}
return logs
}
func traceExceptions(entries []trace.Exception) []verifier.Exception {
if len(entries) == 0 {
return nil
}
exceptions := make([]verifier.Exception, 0, len(entries))
for _, entry := range entries {
exceptions = append(exceptions, verifier.Exception{
Class: entry.Class,
Message: entry.Message,
StackTrace: entry.StackTrace,
UnixMillis: entry.UnixMillis,
})
}
return exceptions
}
// discoverRuns finds every run directory at or below root, a run directory
// being one holding both meta.json and trace.jsonl.
func discoverRuns(root string) ([]string, error) {
var directories []string
err := filepath.Walk(
root,
func(path string, info os.FileInfo, err error) error {
if err != nil {
return err
}
if !info.IsDir() {
return nil
}
if _, statErr := os.Stat(filepath.Join(path, "trace.jsonl")); statErr != nil {
return nil
}
if _, statErr := os.Stat(filepath.Join(path, "meta.json")); statErr != nil {
return nil
}
directories = append(directories, path)
return nil
},
)
if err != nil {
return nil, err
}
sort.Strings(directories)
return directories, nil
}
@@ -0,0 +1,162 @@
package main
import (
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
func writeRunFiles(t *testing.T, directory, meta string, steps ...string) {
t.Helper()
if err := os.WriteFile(filepath.Join(directory, "meta.json"), []byte(meta), 0o644); err != nil {
t.Fatal(err)
}
body := strings.Join(steps, "\n") + "\n"
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte(body), 0o644); err != nil {
t.Fatal(err)
}
}
func TestLoadRunRefusesAStepFromBeforeTheFormatChange(t *testing.T) {
directory := t.TempDir()
writeRunFiles(
t,
directory,
`{"seed": 3}`,
`{"step":1,"timestamp":"2026-08-15T12:00:00Z"}`,
)
_, err := loadRun(directory)
if err == nil {
t.Fatal("a version-0 step must be refused")
}
if !strings.Contains(err.Error(), "trace_version 0") ||
!strings.Contains(err.Error(), "depths") {
t.Errorf(
"the refusal must name the version and why it cannot replay: %v",
err,
)
}
}
func TestLoadRunRefusesARunWithNoSeed(t *testing.T) {
directory := t.TempDir()
writeRunFiles(
t,
directory,
`{}`,
`{"step":1,"trace_version":1,"timestamp":"2026-08-15T12:00:00Z"}`,
)
_, err := loadRun(directory)
if err == nil || !strings.Contains(err.Error(), "seed") {
t.Errorf("a run without a seed cannot be bundled as it was: %v", err)
}
}
func TestLoadRunSplitsOffTheEndOfRunRecord(t *testing.T) {
directory := t.TempDir()
writeRunFiles(
t,
directory,
`{"seed": 3}`,
`{"step":1,"trace_version":1,"timestamp":"2026-08-15T12:00:00Z","hierarchy":{"elements":[],"depths":[]}}`,
`{"step":2,"trace_version":1,"timestamp":"2026-08-15T12:00:01Z","violations":["reachable"]}`,
)
run, err := loadRun(directory)
if err != nil {
t.Fatal(err)
}
if len(run.Steps) != 1 {
t.Fatalf("observation steps: got %d, want 1", len(run.Steps))
}
if run.Finalize == nil || run.Finalize.Index != 2 {
t.Fatalf("finalize record: %+v", run.Finalize)
}
if run.finalizeIndex() != 2 {
t.Errorf("finalize index: got %d, want 2", run.finalizeIndex())
}
}
func TestExtractorFoldCarriesAValueForwardUntilItChanges(t *testing.T) {
steps := []trace.Step{
{Index: 1, ExtractorChanges: map[string]trace.ExtractorChange{
"count": {
Prev: json.RawMessage("null"),
Curr: json.RawMessage("1"),
},
}},
{Index: 2},
{Index: 3, ExtractorChanges: map[string]trace.ExtractorChange{
"count": {Prev: json.RawMessage("1"), Curr: json.RawMessage("4")},
}},
}
folded, err := extractorFold(steps, []string{"count", "unseen"})
if err != nil {
t.Fatal(err)
}
want := []string{"1", "1", "4"}
for position, expected := range want {
if got := string(folded[position][0]); got != expected {
t.Errorf(
"count at step %d: got %s, want %s",
position+1,
got,
expected,
)
}
if got := string(folded[position][1]); got != "null" {
t.Errorf(
"an extractor that never changed must fold to null, got %s",
got,
)
}
}
}
func TestExtractorFoldRefusesAValueTheSpecCannotPlace(t *testing.T) {
steps := []trace.Step{
{Index: 1, ExtractorChanges: map[string]trace.ExtractorChange{
"gone": {Curr: json.RawMessage("1")},
}},
}
_, err := extractorFold(steps, []string{"count"})
if err == nil || !strings.Contains(err.Error(), "gone") {
t.Errorf("an unplaceable extractor value must be refused: %v", err)
}
}
func TestLastActionIsEmptyWhenTheRunnerNeverDispatchedIt(t *testing.T) {
dispatched := trace.Step{
NextAction: &trace.Action{Kind: "Tap", Selector: "id:save", X: 4, Y: 5},
}
skipped := trace.Step{
NextAction: &trace.Action{Kind: "Tap", Selector: "id:save"},
ActionSkipped: foregroundLossReason,
}
action := lastActionFor(dispatched)
if action == nil || action.Kind != verifier.ActionKindTap ||
action.On != "id:save" ||
action.X != 4 {
t.Fatalf("dispatched action: %+v", action)
}
if lastActionFor(skipped) != nil {
t.Error(
"an action the runner threw away never reached state.lastAction",
)
}
}
+312
View File
@@ -0,0 +1,312 @@
// Command oracle-reduction re-evaluates stored traces offline under four
// oracles and reports what each one refutes: the full engine, a crash-only
// detector, a single-state check, and a single-step property triple. The
// oracles vary while the traces stay fixed, which is what separates a defect an
// oracle cannot express from one an explorer never reached.
//
// The offline engine has to reproduce the verdicts each run recorded. A
// disagreement is a bug here or a gap in the trace, so it is reported as a
// mismatch and exits nonzero rather than being counted as a finding.
package main
import (
"encoding/json"
"flag"
"fmt"
"io"
"os"
"sort"
"strings"
"github.com/priyanshujain/sanderling/internal/testrun"
)
const usage = `oracle-reduction replays stored traces under reduced oracles.
Usage:
oracle-reduction --runs <dir> [--output <path>] [--spec <path>]
--runs is scanned recursively; every directory holding meta.json and
trace.jsonl is one trace. Each trace's spec is bundled from the path its
meta.json records unless --spec overrides it.
Exit status is 2 when the offline engine disagreed with any run's recorded
verdicts, which blocks the experiment rather than reporting the difference as
noise.
A reduced oracle whose rewrite no longer states a property reports
cannot_express for it rather than a verdict, and the property is temporal-only
when the other reductions also fail to state or refute it. A property whose
window is longer than the two observations a triple spans is reported with
single_step_truncates_window, which is the form that decides it.
`
type config struct {
runsRoot string
outputPath string
specPath string
// hostExtractors re-runs the spec's extractor getters over each stored
// hierarchy instead of replaying the values a web run's page computed. It
// asks a different question of the same trace: whether the stored tree
// alone carries what the properties read.
hostExtractors bool
}
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
flagSet := flag.NewFlagSet("oracle-reduction", flag.ContinueOnError)
flagSet.SetOutput(stderr)
flagSet.Usage = func() {
fmt.Fprint(stderr, usage)
flagSet.PrintDefaults()
}
var configuration config
flagSet.StringVar(
&configuration.runsRoot,
"runs",
"",
"directory tree holding the run directories to replay (required)",
)
flagSet.StringVar(
&configuration.outputPath,
"output",
"",
"file to write the per-trace JSONL to (default stdout)",
)
flagSet.StringVar(
&configuration.specPath,
"spec",
"",
"spec to bundle instead of the one each meta.json records",
)
flagSet.BoolVar(
&configuration.hostExtractors,
"host-extractors",
false,
"re-run the spec's extractors over each stored hierarchy instead of replaying the values a web run's page computed",
)
if err := flagSet.Parse(arguments); err != nil {
return config{}, err
}
if configuration.runsRoot == "" {
return config{}, fmt.Errorf("--runs is required")
}
return configuration, nil
}
func main() {
configuration, err := parseArguments(os.Args[1:], os.Stderr)
if err != nil {
fmt.Fprintf(os.Stderr, "oracle-reduction: %v\n", err)
os.Exit(1)
}
code, err := run(configuration, os.Stdout, os.Stderr)
if err != nil {
fmt.Fprintf(os.Stderr, "oracle-reduction: %v\n", err)
os.Exit(1)
}
os.Exit(code)
}
func run(configuration config, stdout, stderr io.Writer) (int, error) {
directories, err := discoverRuns(configuration.runsRoot)
if err != nil {
return 1, err
}
if len(directories) == 0 {
return 1, fmt.Errorf(
"no run directories under %s",
configuration.runsRoot,
)
}
output := stdout
if configuration.outputPath != "" {
file, createErr := os.Create(configuration.outputPath)
if createErr != nil {
return 1, createErr
}
defer file.Close()
output = file
}
encoder := json.NewEncoder(output)
var reports []runReport
rejected := 0
for _, directory := range directories {
loaded, loadErr := loadRun(directory)
if loadErr != nil {
rejected++
fmt.Fprintf(stderr, "skipped %s: %v\n", directory, loadErr)
continue
}
specPath := loaded.Meta.SpecPath
if configuration.specPath != "" {
specPath = configuration.specPath
}
bundle, bundleErr := testrun.BundleSpec(specPath, loaded.Meta.Seed)
if bundleErr != nil {
return 1, fmt.Errorf(
"%s: bundle %s: %w",
directory,
specPath,
bundleErr,
)
}
if loaded.Meta.BundleSHA256 != "" &&
bundle.SHA256 != loaded.Meta.BundleSHA256 {
// The bundler writes each module's path into the output relative to
// the working directory, so this differs whenever the replay is
// invoked from somewhere else than the run was. What a changed spec
// would actually break is caught by the property set, the extractor
// names and the residual comparison below.
fmt.Fprintf(stderr,
"note: %s bundles to %s here and the run recorded %s\n",
directory, bundle.SHA256[:12], loaded.Meta.BundleSHA256[:12])
}
report, replayErr := replay(
loaded,
string(bundle.JavaScript),
configuration.hostExtractors,
)
if replayErr != nil {
return 1, fmt.Errorf("%s: %w", directory, replayErr)
}
if err := encoder.Encode(report); err != nil {
return 1, err
}
reports = append(reports, report)
}
summarize(reports, rejected, stderr)
for _, report := range reports {
if !report.Valid {
return 2, nil
}
}
if len(reports) == 0 {
return 1, fmt.Errorf(
"every run directory was rejected; nothing was replayed",
)
}
return 0, nil
}
func summarize(reports []runReport, rejected int, out io.Writer) {
invalid := 0
crashed := 0
weakest := map[string]int{}
byClass := map[string]map[string]int{}
inexpressible := map[string]map[string]bool{}
unmatched := 0
engineRefutations := 0
for _, report := range reports {
if !report.Valid {
invalid++
}
if report.CrashOnly.Fired {
crashed++
}
for _, property := range report.Properties {
recordInexpressible(
inexpressible,
"single-state",
property.SingleState,
property.Property,
)
recordInexpressible(
inexpressible,
"single-step",
property.SingleStep,
property.Property,
)
if !property.Engine.Refuted {
if property.SingleState.Refuted || property.SingleStep.Refuted {
unmatched++
}
continue
}
engineRefutations++
weakest[property.Weakest]++
if byClass[property.Class] == nil {
byClass[property.Class] = map[string]int{}
}
byClass[property.Class][property.Weakest]++
}
}
fmt.Fprintf(
out,
"\ntraces replayed: %d (rejected: %d)\n",
len(reports),
rejected,
)
fmt.Fprintf(
out,
"validity: %d of %d reproduced the recorded verdicts exactly\n",
len(reports)-invalid,
len(reports),
)
fmt.Fprintf(out, "traces where crash-only fired: %d\n", crashed)
fmt.Fprintf(out, "engine refutations: %d\n", engineRefutations)
if engineRefutations > 0 {
fmt.Fprintf(out, "weakest refuting oracle: %s\n", counts(weakest))
for _, class := range sortedKeys(byClass) {
fmt.Fprintf(out, " %s: %s\n", class, counts(byClass[class]))
}
fmt.Fprintf(out, "temporal-only fraction: %.3f\n",
float64(weakest["temporal-only"])/float64(engineRefutations))
}
fmt.Fprintf(
out,
"properties a reduced oracle cannot express: %s\n",
counts(distinct(inexpressible)),
)
fmt.Fprintf(
out,
"reduced-oracle refutations the engine did not make: %d\n",
unmatched,
)
}
// recordInexpressible counts a property once per oracle however many traces it
// appears on, because whether an oracle can state a property is a fact about
// the property and not about the run.
func recordInexpressible(
seen map[string]map[string]bool,
oracle string,
finding refutation,
property string,
) {
if !finding.CannotExpress {
return
}
if seen[oracle] == nil {
seen[oracle] = map[string]bool{}
}
seen[oracle][property] = true
}
func distinct(seen map[string]map[string]bool) map[string]int {
sizes := map[string]int{}
for oracle, properties := range seen {
sizes[oracle] = len(properties)
}
return sizes
}
func counts(values map[string]int) string {
parts := make([]string, 0, len(values))
for _, key := range sortedKeys(values) {
parts = append(parts, fmt.Sprintf("%s=%d", key, values[key]))
}
return strings.Join(parts, " ")
}
func sortedKeys[V any](values map[string]V) []string {
keys := make([]string, 0, len(values))
for key := range values {
keys = append(keys, key)
}
sort.Strings(keys)
return keys
}
@@ -0,0 +1,549 @@
package main
import (
"bytes"
"encoding/json"
"fmt"
"sort"
"github.com/priyanshujain/sanderling/internal/ltl"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// foregroundLossReason is the skip reason internal/runner records when the app
// was no longer the foreground process at action time. It is the only
// foreground signal a stored trace carries, and it is what the crash-only
// oracle reads for "the application is no longer the foreground process".
const foregroundLossReason = "app_left_foreground"
// refutation is one oracle's finding for one property on one trace. A reduced
// oracle has three of them and not two: CannotExpress says its rewrite no
// longer states the property, which is neither a refutation nor a clean bill.
type refutation struct {
Refuted bool `json:"refuted"`
CannotExpress bool `json:"cannot_express,omitempty"`
// Step is the observation whose evaluation produced the violation;
// OriginStep is the observation whose obligation failed, which for a
// deferred one is earlier.
Step int `json:"step,omitempty"`
OriginStep int `json:"origin_step,omitempty"`
Reason string `json:"reason,omitempty"`
IsError bool `json:"is_error,omitempty"`
}
type propertyReport struct {
Property string `json:"property"`
Class string `json:"class"`
TopLevel string `json:"top_level"`
Engine refutation `json:"engine"`
SingleState refutation `json:"single_state"`
SingleStep refutation `json:"single_step"`
// SingleStepTruncatesWindow marks a property whose window the triple had to
// shorten to two observations, so what its single-step column refutes is a
// stronger property than the one the author wrote.
SingleStepTruncatesWindow bool `json:"single_step_truncates_window,omitempty"`
// Weakest names the weakest oracle that refutes this property on this
// trace, and is empty unless the engine refuted it. "temporal-only" means
// no reduced oracle did.
Weakest string `json:"weakest_refuting_oracle,omitempty"`
}
type crashReport struct {
Fired bool `json:"fired"`
FirstStep int `json:"first_step,omitempty"`
ExceptionSteps []int `json:"exception_steps,omitempty"`
ForegroundLossSteps []int `json:"foreground_loss_steps,omitempty"`
ErrorLogSteps []int `json:"error_log_steps,omitempty"`
}
// mismatch is one disagreement between the offline engine and the verdicts the
// run recorded. Any of these blocks the trace's result.
type mismatch struct {
Property string `json:"property"`
Field string `json:"field"`
Online string `json:"online"`
Offline string `json:"offline"`
}
type runReport struct {
Run string `json:"run"`
Seed int64 `json:"seed"`
Platform string `json:"platform"`
Arm string `json:"arm,omitempty"`
Generator string `json:"generator,omitempty"`
// ReplayMode says where the extractor values under replay came from:
// "page-extractor-values" reuses what the page computed in V8 and the run
// evaluated against, "host-extractors" re-runs the spec's getters over the
// stored hierarchy.
ReplayMode string `json:"replay_mode"`
StepsObserved int `json:"steps_observed"`
StepsSkipped int `json:"steps_skipped"`
Valid bool `json:"valid"`
Mismatches []mismatch `json:"mismatches,omitempty"`
// ResidualMismatches counts every step whose replayed residual formula
// differed from the recorded one, of which Mismatches carries the first
// few. The recorded residual is the engine's whole pending state, so
// agreement on it is a stronger claim than agreement on the verdicts.
ResidualMismatches int `json:"residual_mismatches"`
CrashOnly crashReport `json:"crash_only"`
Properties []propertyReport `json:"properties"`
ActionsUnbounded int `json:"actions_without_resolved_bounds"`
ScrollActions int `json:"scroll_last_actions"`
}
// witnessRecord is one violation as either side reports it: the step it was
// recorded at, plus the witness the engine attached.
type witnessRecord struct {
RecordedStep int
OriginStep int
DetectedStep int
Reason string
IsError bool
}
// replay re-evaluates one run offline under all four oracles. The engine's
// offline verdicts are compared against the ones the run recorded, and the
// report is marked invalid on any disagreement: a mismatch is a bug here or a
// gap in the trace, not a finding.
func replay(
run loadedRun,
bundleJavaScript string,
hostExtractors bool,
) (runReport, error) {
engine, err := verifier.New(
verifier.WithSeed(uint64(run.Meta.Seed)),
verifier.WithPlatform(run.Meta.Platform),
verifier.WithAppPackage(run.Meta.BundleID),
)
if err != nil {
return runReport{}, fmt.Errorf("verifier: %w", err)
}
if err := engine.Load(bundleJavaScript); err != nil {
return runReport{}, fmt.Errorf("load spec: %w", err)
}
formulas, err := engine.PropertyFormulas()
if err != nil {
return runReport{}, err
}
singleState := map[string]*ltl.Evaluator{}
singleStep := map[string]*ltl.Evaluator{}
for name, formula := range formulas {
singleState[name] = ltl.NewEvaluator(singleStateFormula(formula))
singleStep[name] = ltl.NewEvaluator(singleStepFormula(formula))
}
usePageValues := run.Meta.Platform == "web" && !hostExtractors
report := runReport{
Run: run.Directory,
Seed: run.Meta.Seed,
Platform: run.Meta.Platform,
Arm: run.Meta.Arm,
Generator: run.Meta.Generator,
ReplayMode: "host-extractors",
}
var folded []map[int]json.RawMessage
if usePageValues {
report.ReplayMode = "page-extractor-values"
folded, err = extractorFold(run.Steps, engine.ExtractorNames())
if err != nil {
return runReport{}, err
}
}
offline := map[string]witnessRecord{}
stateFired := map[string]int{}
stepFired := map[string]int{}
var residualMismatches []mismatch
var lastAction *verifier.Action
for position, step := range run.Steps {
if step.NextAction != nil && step.NextAction.Selector != "" &&
step.NextAction.ResolvedBounds == nil {
report.ActionsUnbounded++
}
if step.SkippedVerification {
report.StepsSkipped++
residualMismatches = append(
residualMismatches,
compareResiduals(step, engine.Residuals())...)
lastAction = lastActionFor(step)
continue
}
if lastAction != nil && lastAction.Kind == verifier.ActionKindScroll {
report.ScrollActions++
}
if err := engine.PushSnapshot(verifier.SnapshotInput{
Tree: step.Hierarchy,
LastAction: lastAction,
StepTime: step.Timestamp,
StepIndex: step.Index,
RunStart: run.Meta.StartedAt,
Logs: traceLogs(step.Logs),
Exceptions: traceExceptions(step.Exceptions),
}); err != nil {
return runReport{}, fmt.Errorf("step %d push: %w", step.Index, err)
}
if usePageValues {
skipped, overrideErr := engine.OverrideExtractorValues(
folded[position],
)
if overrideErr != nil {
return runReport{}, fmt.Errorf(
"step %d override: %w",
step.Index,
overrideErr,
)
}
if skipped > 0 {
return runReport{}, fmt.Errorf(
"step %d: %d recorded extractor values fell outside the spec's extractor list",
step.Index,
skipped,
)
}
}
engine.EvaluateProperties()
for _, name := range engine.NewlyViolatedProperties() {
offline[name] = witnessFrom(engine.Witness(name), step.Index)
}
residualMismatches = append(
residualMismatches,
compareResiduals(step, engine.Residuals())...)
for name := range formulas {
recordFiring(
stateFired,
name,
singleState[name].ObserveAtStep(step.Timestamp, step.Index),
step.Index,
)
recordFiring(
stepFired,
name,
singleStep[name].ObserveAtStep(step.Timestamp, step.Index),
step.Index,
)
}
report.StepsObserved++
lastAction = lastActionFor(step)
}
finalizeIndex := run.finalizeIndex()
for _, name := range engine.Finalize() {
offline[name] = witnessFrom(engine.Witness(name), finalizeIndex)
}
for name := range formulas {
recordFiring(
stateFired,
name,
singleState[name].Finalize(),
finalizeIndex,
)
recordFiring(
stepFired,
name,
singleStep[name].Finalize(),
finalizeIndex,
)
}
report.CrashOnly = crashOnly(run)
report.Mismatches = compareVerdicts(onlineVerdicts(run), offline)
report.ResidualMismatches = len(residualMismatches)
if len(residualMismatches) > reportedResiduals {
residualMismatches = residualMismatches[:reportedResiduals]
}
report.Mismatches = append(report.Mismatches, residualMismatches...)
report.Valid = len(report.Mismatches) == 0
names := make([]string, 0, len(formulas))
for name := range formulas {
names = append(names, name)
}
sort.Strings(names)
for _, name := range names {
property := propertyReport{
Property: name,
Class: propertyClass(formulas[name]),
TopLevel: topLevelForm(formulas[name]),
Engine: engineRefutation(offline, name),
SingleState: reducedRefutation(
singleStateExpresses(formulas[name]),
singleState[name],
stateFired[name],
),
SingleStep: reducedRefutation(
singleStepExpresses(formulas[name]),
singleStep[name],
stepFired[name],
),
SingleStepTruncatesWindow: truncatesWindow(formulas[name]),
}
property.Weakest = weakestOracle(property, report.CrashOnly)
report.Properties = append(report.Properties, property)
}
return report, nil
}
// reportedResiduals caps how many residual differences one trace lists. The
// count of all of them is reported alongside; the first few are what says
// where the two engines parted.
const reportedResiduals = 5
// compareResiduals checks the replayed pending state against the one the step
// recorded. A step the verifier skipped still recorded the residual it was
// holding, so those steps assert that a skipped step advanced nothing.
func compareResiduals(
step trace.Step,
replayed map[string]ltl.Formula,
) []mismatch {
if len(step.Residuals) == 0 {
return nil
}
var mismatches []mismatch
names := make([]string, 0, len(step.Residuals))
for name := range step.Residuals {
names = append(names, name)
}
sort.Strings(names)
for _, name := range names {
formula, ok := replayed[name]
if !ok {
mismatches = append(mismatches, mismatch{
Property: name,
Field: fmt.Sprintf("residual at step %d", step.Index),
Online: string(step.Residuals[name]),
Offline: "property not registered",
})
continue
}
encoded, err := json.Marshal(formula)
if err != nil {
encoded = []byte(fmt.Sprintf("%q", err.Error()))
}
if !bytes.Equal(encoded, step.Residuals[name]) {
mismatches = append(mismatches, mismatch{
Property: name,
Field: fmt.Sprintf("residual at step %d", step.Index),
Online: string(step.Residuals[name]),
Offline: string(encoded),
})
}
}
return mismatches
}
func witnessFrom(witness *verifier.Witness, recordedStep int) witnessRecord {
record := witnessRecord{RecordedStep: recordedStep}
if witness == nil {
return record
}
record.OriginStep = witness.Step
record.DetectedStep = witness.DetectedStep
record.Reason = witness.Reason
record.IsError = witness.IsError
return record
}
// onlineVerdicts reads the violations the run recorded, including the
// end-of-run record a finalized liveness obligation is written to.
func onlineVerdicts(run loadedRun) map[string]witnessRecord {
recorded := map[string]witnessRecord{}
steps := run.Steps
if run.Finalize != nil {
steps = append(append([]trace.Step(nil), steps...), *run.Finalize)
}
for _, step := range steps {
for _, name := range step.Violations {
record := witnessRecord{RecordedStep: step.Index}
if witness, ok := step.Witnesses[name]; ok {
record.OriginStep = witness.Step
record.DetectedStep = witness.DetectedStep
record.Reason = witness.Reason
record.IsError = witness.IsError
}
recorded[name] = record
}
}
return recorded
}
func compareVerdicts(online, offline map[string]witnessRecord) []mismatch {
var mismatches []mismatch
names := map[string]bool{}
for name := range online {
names[name] = true
}
for name := range offline {
names[name] = true
}
ordered := make([]string, 0, len(names))
for name := range names {
ordered = append(ordered, name)
}
sort.Strings(ordered)
for _, name := range ordered {
recorded, wasRecorded := online[name]
replayed, wasReplayed := offline[name]
switch {
case wasRecorded && !wasReplayed:
mismatches = append(mismatches, mismatch{
Property: name, Field: "violated",
Online: fmt.Sprintf(
"violated at step %d",
recorded.RecordedStep,
),
Offline: "not violated",
})
continue
case !wasRecorded && wasReplayed:
mismatches = append(mismatches, mismatch{
Property: name, Field: "violated",
Online: "not violated",
Offline: fmt.Sprintf(
"violated at step %d",
replayed.RecordedStep,
),
})
continue
case !wasRecorded:
continue
}
for _, field := range []struct {
name string
online string
offline string
}{
{"step", fmt.Sprint(recorded.RecordedStep), fmt.Sprint(replayed.RecordedStep)},
{"origin_step", fmt.Sprint(recorded.OriginStep), fmt.Sprint(replayed.OriginStep)},
{"detected_step", fmt.Sprint(recorded.DetectedStep), fmt.Sprint(replayed.DetectedStep)},
{"reason", recorded.Reason, replayed.Reason},
} {
if field.online != field.offline {
mismatches = append(mismatches, mismatch{
Property: name, Field: field.name,
Online: field.online, Offline: field.offline,
})
}
}
}
return mismatches
}
// crashOnly fires where the application left the foreground or an error
// surface was recorded. Error-level log lines are reported alongside rather
// than folded in: a console error is not a crash, and an analysis that wants
// the looser detector can read the steps from here.
func crashOnly(run loadedRun) crashReport {
report := crashReport{}
for _, step := range run.Steps {
if len(step.Exceptions) > 0 {
report.ExceptionSteps = append(report.ExceptionSteps, step.Index)
}
if step.ActionSkipped == foregroundLossReason {
report.ForegroundLossSteps = append(
report.ForegroundLossSteps,
step.Index,
)
}
for _, entry := range step.Logs {
if entry.Level == "E" || entry.Level == "F" {
report.ErrorLogSteps = append(report.ErrorLogSteps, step.Index)
break
}
}
}
first := 0
for _, step := range append(append([]int(nil), report.ExceptionSteps...), report.ForegroundLossSteps...) {
if first == 0 || step < first {
first = step
}
}
report.Fired = first != 0
report.FirstStep = first
return report
}
func engineRefutation(
offline map[string]witnessRecord,
name string,
) refutation {
record, ok := offline[name]
if !ok {
return refutation{}
}
return refutation{
Refuted: true,
Step: record.RecordedStep,
OriginStep: record.OriginStep,
Reason: record.Reason,
IsError: record.IsError,
}
}
// reducedRefutation reports a reduced oracle's finding, or its inability to
// state the property at all. The evaluator is driven either way so that the
// verdict a reduction would have reached is never what decides whether it was
// entitled to reach one.
func reducedRefutation(
expresses bool,
evaluator *ltl.Evaluator,
firedAt int,
) refutation {
if !expresses {
return refutation{CannotExpress: true}
}
return evaluatorRefutation(evaluator, firedAt)
}
func evaluatorRefutation(evaluator *ltl.Evaluator, firedAt int) refutation {
violation := evaluator.Violation()
if violation == nil {
return refutation{}
}
return refutation{
Refuted: true,
Step: firedAt,
OriginStep: violation.Step,
Reason: violation.Reason,
IsError: violation.IsError,
}
}
// recordFiring keeps the first observation at which a reduced oracle latched,
// which the evaluator itself does not carry.
func recordFiring(
fired map[string]int,
name string,
verdict ltl.Verdict,
step int,
) {
if verdict != ltl.VerdictViolated {
return
}
if _, ok := fired[name]; ok {
return
}
fired[name] = step
}
// weakestOracle names the weakest oracle that refutes a defect the engine
// refuted, in the order a crash detector, a single-state check and a
// single-step triple are weaker than the engine.
func weakestOracle(property propertyReport, crash crashReport) string {
if !property.Engine.Refuted {
return ""
}
switch {
case crash.Fired:
return "crash-only"
case property.SingleState.Refuted:
return "single-state"
case property.SingleStep.Refuted:
return "single-step"
default:
return "temporal-only"
}
}
@@ -0,0 +1,433 @@
package main
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/hierarchy"
"github.com/priyanshujain/sanderling/internal/trace"
"github.com/priyanshujain/sanderling/internal/verifier"
)
// fixtureSpec counts rows on screen and asserts three things about them: a
// bound a single observation can check, a reachability goal, and a growth step
// that only two consecutive observations can see.
const fixtureSpec = `
const rows = __sanderling__.extract((s) => s.ax.findAll({ text: "row" }).length).named("rows");
globalThis.properties = {
fewRows: __sanderling__.always(() => rows.current < 3),
rowsAppear: __sanderling__.eventually(() => rows.current > 0).within(500, "steps"),
rowsKeepGrowing: __sanderling__.always(
__sanderling__.now(() => rows.current === 1).implies(
__sanderling__.next(() => rows.current === 2))),
};
`
func rowTree(t *testing.T, rows int) *hierarchy.Tree {
t.Helper()
children := make([]string, 0, rows)
for range rows {
children = append(
children,
`{"attributes": {"text": "row", "bounds": "[0,0,10,10]"}, "children": []}`,
)
}
source := fmt.Sprintf(
`{"attributes": {"resource-id": "root", "bounds": "[0,0,100,100]"}, "children": [%s]}`,
strings.Join(children, ","),
)
tree, err := hierarchy.Parse(source)
if err != nil {
t.Fatalf("parse tree: %v", err)
}
return tree
}
// writeFixtureRun records a run the way internal/runner does: every step's
// hierarchy, the violations that fired at it, their witnesses, and the residual
// each property was left holding.
func writeFixtureRun(t *testing.T, directory string, rowCounts []int) {
t.Helper()
engine, err := verifier.New(verifier.WithPlatform("android"))
if err != nil {
t.Fatalf("verifier: %v", err)
}
if err := engine.Load(fixtureSpec); err != nil {
t.Fatalf("load spec: %v", err)
}
writer, err := trace.NewWriter(directory)
if err != nil {
t.Fatalf("trace writer: %v", err)
}
defer writer.Close()
runStart := time.Date(2026, 8, 15, 12, 0, 0, 0, time.UTC)
if err := writer.WriteMeta(trace.Meta{
Seed: 11,
SpecPath: "fixture.ts",
Platform: "android",
BundleID: "com.example.fixture",
StartedAt: runStart,
}); err != nil {
t.Fatalf("meta: %v", err)
}
lastIndex := 0
for position, rows := range rowCounts {
index := position + 1
lastIndex = index
stepTime := runStart.Add(time.Duration(index) * time.Second)
tree := rowTree(t, rows)
if err := engine.PushSnapshot(verifier.SnapshotInput{
Tree: tree,
StepTime: stepTime,
StepIndex: index,
RunStart: runStart,
}); err != nil {
t.Fatalf("push step %d: %v", index, err)
}
engine.EvaluateProperties()
violations := engine.NewlyViolatedProperties()
step := trace.Step{
Index: index,
Timestamp: stepTime,
Hierarchy: tree,
Violations: violations,
Witnesses: fixtureWitnesses(engine, violations, index),
Residuals: fixtureResiduals(t, engine),
}
if err := writer.WriteStep(step); err != nil {
t.Fatalf("write step %d: %v", index, err)
}
}
if ended := engine.Finalize(); len(ended) > 0 {
if err := writer.WriteStep(trace.Step{
Index: lastIndex + 1,
Timestamp: runStart.Add(time.Duration(lastIndex+1) * time.Second),
Violations: ended,
Witnesses: fixtureWitnesses(engine, ended, lastIndex+1),
}); err != nil {
t.Fatalf("write finalize step: %v", err)
}
}
}
func fixtureWitnesses(
engine *verifier.Verifier,
properties []string,
index int,
) map[string]trace.Witness {
if len(properties) == 0 {
return nil
}
witnesses := map[string]trace.Witness{}
for _, name := range properties {
witness := engine.Witness(name)
if witness == nil {
continue
}
detected := witness.DetectedStep
if detected == 0 {
detected = index
}
witnesses[name] = trace.Witness{
Reason: witness.Reason,
IsError: witness.IsError,
Step: witness.Step,
DetectedStep: detected,
Extractors: witness.Extractors,
}
}
return witnesses
}
func fixtureResiduals(
t *testing.T,
engine *verifier.Verifier,
) map[string]json.RawMessage {
t.Helper()
residuals := map[string]json.RawMessage{}
for name, formula := range engine.Residuals() {
body, err := json.Marshal(formula)
if err != nil {
t.Fatalf("marshal residual %q: %v", name, err)
}
residuals[name] = body
}
return residuals
}
func replayFixture(t *testing.T, directory string) runReport {
t.Helper()
loaded, err := loadRun(directory)
if err != nil {
t.Fatalf("load run: %v", err)
}
report, err := replay(loaded, fixtureSpec, false)
if err != nil {
t.Fatalf("replay: %v", err)
}
return report
}
func propertyByName(
t *testing.T,
report runReport,
name string,
) propertyReport {
t.Helper()
for _, property := range report.Properties {
if property.Property == name {
return property
}
}
t.Fatalf("property %q missing from the report", name)
return propertyReport{}
}
func TestReplayReproducesTheRecordedVerdicts(t *testing.T) {
directory := t.TempDir()
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
report := replayFixture(t, directory)
if !report.Valid {
t.Fatalf("replay disagreed with the run: %+v", report.Mismatches)
}
if report.ResidualMismatches != 0 {
t.Errorf("residual mismatches: %d", report.ResidualMismatches)
}
if report.StepsObserved != 4 {
t.Errorf("steps observed: got %d, want 4", report.StepsObserved)
}
growth := propertyByName(t, report, "rowsKeepGrowing")
if !growth.Engine.Refuted || growth.Engine.Step != 3 {
t.Errorf("engine on rowsKeepGrowing: %+v", growth.Engine)
}
if growth.SingleState.Refuted {
t.Errorf(
"single-state refuted a growth step it cannot see: %+v",
growth.SingleState,
)
}
if !growth.SingleStep.Refuted {
t.Error("single-step did not refute a one-step obligation")
}
if growth.Weakest != "single-step" {
t.Errorf("weakest oracle for rowsKeepGrowing: got %q", growth.Weakest)
}
bound := propertyByName(t, report, "fewRows")
if !bound.SingleState.Refuted || bound.Weakest != "single-state" {
t.Errorf(
"fewRows: single-state %+v weakest %q",
bound.SingleState,
bound.Weakest,
)
}
}
func TestReplayReportsAViolationTheRunNeverRecorded(t *testing.T) {
directory := t.TempDir()
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
dropRecordedViolation(t, directory, "rowsKeepGrowing")
report := replayFixture(t, directory)
if report.Valid {
t.Fatal("a trace missing a recorded violation must not replay as valid")
}
var found bool
for _, entry := range report.Mismatches {
if entry.Property == "rowsKeepGrowing" && entry.Field == "violated" {
found = true
}
}
if !found {
t.Errorf(
"no mismatch named the dropped violation: %+v",
report.Mismatches,
)
}
}
func TestReplayReportsAResidualTheRunNeverHeld(t *testing.T) {
directory := t.TempDir()
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
rewriteResidual(t, directory, 1, "fewRows", `{"op":"false"}`)
report := replayFixture(t, directory)
if report.Valid {
t.Fatal(
"a trace whose recorded residual differs must not replay as valid",
)
}
if report.ResidualMismatches != 1 {
t.Errorf(
"residual mismatches: got %d, want 1",
report.ResidualMismatches,
)
}
}
// dropRecordedViolation removes one property from every step's violations,
// which is what a trace that failed to record a verdict looks like.
func dropRecordedViolation(t *testing.T, directory, property string) {
t.Helper()
rewriteSteps(t, directory, func(step *trace.Step) {
kept := step.Violations[:0]
for _, name := range step.Violations {
if name != property {
kept = append(kept, name)
}
}
step.Violations = kept
delete(step.Witnesses, property)
})
}
func rewriteResidual(
t *testing.T,
directory string,
index int,
property, residual string,
) {
t.Helper()
rewriteSteps(t, directory, func(step *trace.Step) {
if step.Index == index {
step.Residuals[property] = json.RawMessage(residual)
}
})
}
func rewriteSteps(t *testing.T, directory string, edit func(*trace.Step)) {
t.Helper()
path := filepath.Join(directory, "trace.jsonl")
body, err := os.ReadFile(path)
if err != nil {
t.Fatalf("read trace: %v", err)
}
var rewritten strings.Builder
for _, line := range strings.Split(strings.TrimSpace(string(body)), "\n") {
var step trace.Step
if err := json.Unmarshal([]byte(line), &step); err != nil {
t.Fatalf("decode step: %v", err)
}
edit(&step)
encoded, err := json.Marshal(step)
if err != nil {
t.Fatalf("encode step: %v", err)
}
rewritten.Write(encoded)
rewritten.WriteByte('\n')
}
if err := os.WriteFile(path, []byte(rewritten.String()), 0o644); err != nil {
t.Fatalf("write trace: %v", err)
}
}
func TestCrashOnlyFiresOnAnErrorSurfaceAndOnLeavingTheForeground(t *testing.T) {
run := loadedRun{Steps: []trace.Step{
{Index: 1, Logs: []trace.LogEntry{{Level: "E", Message: "noisy"}}},
{Index: 2, ActionSkipped: foregroundLossReason},
{Index: 3, Exceptions: []trace.Exception{{Class: "TypeError"}}},
}}
report := crashOnly(run)
if !report.Fired || report.FirstStep != 2 {
t.Errorf("crash-only: %+v", report)
}
if len(report.ExceptionSteps) != 1 || report.ExceptionSteps[0] != 3 {
t.Errorf("exception steps: %v", report.ExceptionSteps)
}
if len(report.ErrorLogSteps) != 1 || report.ErrorLogSteps[0] != 1 {
t.Errorf("error log steps: %v", report.ErrorLogSteps)
}
}
func TestCrashOnlyIgnoresAnErrorLogOnItsOwn(t *testing.T) {
run := loadedRun{
Steps: []trace.Step{{Index: 1, Logs: []trace.LogEntry{{Level: "E"}}}},
}
if crashOnly(run).Fired {
t.Error("an error-level log line is not a crash")
}
}
// A reachability goal the run reaches at the fourth observation is clean under
// the window its author wrote and refuted by the same window shortened to two
// observations. Reporting that shortened refutation would convict every clean
// trace, so the single-step column has to admit it cannot state the property.
func TestSingleStepDoesNotConvictACleanTraceOfATruncatedWindow(t *testing.T) {
directory := t.TempDir()
writeFixtureRun(t, directory, []int{0, 0, 0, 1})
report := replayFixture(t, directory)
if !report.Valid {
t.Fatalf("replay disagreed with the run: %+v", report.Mismatches)
}
appear := propertyByName(t, report, "rowsAppear")
if appear.Engine.Refuted {
t.Fatalf(
"the trace reaches the goal inside the 500-step window: %+v",
appear.Engine,
)
}
if appear.SingleStep.Refuted {
t.Errorf(
"single-step refuted a clean trace by shortening the window: %+v",
appear.SingleStep,
)
}
if !appear.SingleStep.CannotExpress {
t.Error(
"single-step must record a window it cannot state as inexpressible",
)
}
if !appear.SingleStepTruncatesWindow {
t.Error("the truncated-window marker must stay visible on the property")
}
}
func TestSingleStateAdmitsWhatOneObservationCannotState(t *testing.T) {
directory := t.TempDir()
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
report := replayFixture(t, directory)
growth := propertyByName(t, report, "rowsKeepGrowing")
if !growth.SingleState.CannotExpress {
t.Errorf(
"single-state kept a next obligation it erases to nothing: %+v",
growth.SingleState,
)
}
appear := propertyByName(t, report, "rowsAppear")
if !appear.SingleState.CannotExpress {
t.Errorf(
"single-state kept a reachability goal it erases to nothing: %+v",
appear.SingleState,
)
}
bound := propertyByName(t, report, "fewRows")
if bound.SingleState.CannotExpress {
t.Error("single-state states a bound read at one observation")
}
if !bound.SingleState.Refuted || bound.Weakest != "single-state" {
t.Errorf(
"fewRows: single-state %+v weakest %q",
bound.SingleState,
bound.Weakest,
)
}
}
@@ -0,0 +1,281 @@
package main
import (
"fmt"
"time"
"github.com/priyanshujain/sanderling/internal/ltl"
)
// tripleWindow is the horizon a property triple has: an obligation armed at one
// observation must discharge at the next, and nothing outlives that.
const tripleWindow = 2
// singleStateFormula is what a checker holding one observation can refute.
// Every operator that defers or repeats an obligation is replaced by the
// constant that makes the formula around it trivially satisfied, so the only
// refutation left is a predicate read at a single observation. The outermost
// always survives because it is what says "check this at each observation",
// which costs no history.
func singleStateFormula(formula ltl.Formula) ltl.Formula {
if always, ok := formula.(ltl.AlwaysFormula); ok {
always.Inner = stateless(always.Inner, true)
return always
}
return stateless(formula, true)
}
// singleStepFormula is the property triple: one step of history, and no
// obligation surviving past the next observation. A next keeps its one-step
// deferral, and an eventually of any window shrinks to the two observations a
// triple spans.
func singleStepFormula(formula ltl.Formula) ltl.Formula {
if always, ok := formula.(ltl.AlwaysFormula); ok {
always.Inner = oneStep(always.Inner, true)
return always
}
return oneStep(formula, true)
}
// singleStateExpresses and singleStepExpresses decide, from the property's form
// alone, whether the reduced oracle can still state the property after its
// rewrite. An oracle that cannot state a property does not get to refute it,
// and is reported as silent on it rather than as either verdict.
//
// Two shapes defeat a reduction. The rewrite shortens an obligation window the
// oracle's horizon cannot hold, leaving it to refute a property strictly
// stronger than the one the author wrote. Or the rewrite leaves a formula whose
// verdict no longer depends on the trace, leaving it to report the same answer
// everywhere. Either way the column would carry a verdict about a property
// nobody wrote, which is worth less than an admission that the oracle is out of
// its depth.
func singleStateExpresses(formula ltl.Formula) bool {
return dependsOnTrace(singleStateFormula(formula))
}
func singleStepExpresses(formula ltl.Formula) bool {
return !truncatesWindow(formula) &&
dependsOnTrace(singleStepFormula(formula))
}
// truncatesWindow reports whether the single-step rewrite had to shorten a
// window to fit a triple. Where it did, the reduced oracle is checking a
// stronger property than the author wrote, so a refutation of it is not the
// same event as a refutation of the property.
func truncatesWindow(formula ltl.Formula) bool {
if always, ok := formula.(ltl.AlwaysFormula); ok {
return shortensWindow(always.Inner, true)
}
return shortensWindow(formula, true)
}
// shortensWindow walks the same shape oneStep rewrites, and reports only the
// shortening that strengthens the formula. Polarity is what separates the two:
// a shorter window under an even number of negations demands the same thing
// sooner, so a refutation of it need not be a refutation of the property, while
// under an odd number it asks for less and its refutations stay sound. The
// operators oneStep erases rather than shortens are erased to the constant its
// position is satisfied by, which also only asks for less.
func shortensWindow(formula ltl.Formula, positive bool) bool {
switch concrete := formula.(type) {
case ltl.EventuallyFormula:
return positive && windowOutlastsTriple(concrete)
case ltl.NowFormula:
return shortensWindow(concrete.Inner, positive)
case ltl.NotFormula:
return shortensWindow(concrete.Inner, !positive)
case ltl.AndFormula:
return shortensWindow(concrete.Left, positive) || shortensWindow(concrete.Right, positive)
case ltl.OrFormula:
return shortensWindow(concrete.Left, positive) || shortensWindow(concrete.Right, positive)
case ltl.ImpliesFormula:
return shortensWindow(concrete.Antecedent, !positive) || shortensWindow(concrete.Consequent, positive)
default:
return false
}
}
// windowOutlastsTriple asks whether an obligation can still be open after the
// two observations a triple spans. A window counted in time is outside what a
// triple can state whatever its length, because a triple has no clock: it can
// say "at the next observation" and nothing about when that arrives.
func windowOutlastsTriple(formula ltl.EventuallyFormula) bool {
if formula.HasStepBound {
return formula.StepBound > tripleWindow
}
return true
}
// dependsOnTrace reports whether a rewritten formula can still read the trace.
// Constants fold through the connectives, and one that folds away to a constant
// answers the same on every trace: silence that reads as "did not refute", or a
// refutation of everything. Neither is a verdict about the run.
func dependsOnTrace(formula ltl.Formula) bool {
_, constant := fold(formula).(ltl.PureFormula)
return !constant
}
// fold propagates the constants the rewrites substituted for erased operators.
// An obligation whose inner formula folded to false keeps its shape, because a
// deferred false is a refutation still owed; only the side that can no longer
// fail folds away.
func fold(formula ltl.Formula) ltl.Formula {
switch concrete := formula.(type) {
case ltl.AlwaysFormula:
concrete.Inner = fold(concrete.Inner)
if pure, ok := concrete.Inner.(ltl.PureFormula); ok {
return pure
}
return concrete
case ltl.EventuallyFormula:
concrete.Inner = fold(concrete.Inner)
if isPure(concrete.Inner, true) {
return ltl.Pure(true)
}
return concrete
case ltl.NextFormula:
inner := fold(concrete.Inner)
if isPure(inner, true) {
return ltl.Pure(true)
}
return ltl.Next(inner)
case ltl.NowFormula:
inner := fold(concrete.Inner)
if _, ok := inner.(ltl.PureFormula); ok {
return inner
}
return ltl.Now(inner)
case ltl.NotFormula:
inner := fold(concrete.Inner)
if pure, ok := inner.(ltl.PureFormula); ok {
return ltl.Pure(!pure.Value)
}
return ltl.Not(inner)
case ltl.AndFormula:
left, right := fold(concrete.Left), fold(concrete.Right)
switch {
case isPure(left, false) || isPure(right, false):
return ltl.Pure(false)
case isPure(left, true):
return right
case isPure(right, true):
return left
}
return ltl.And(left, right)
case ltl.OrFormula:
left, right := fold(concrete.Left), fold(concrete.Right)
switch {
case isPure(left, true) || isPure(right, true):
return ltl.Pure(true)
case isPure(left, false):
return right
case isPure(right, false):
return left
}
return ltl.Or(left, right)
case ltl.ImpliesFormula:
antecedent, consequent := fold(concrete.Antecedent), fold(concrete.Consequent)
switch {
case isPure(antecedent, false) || isPure(consequent, true):
return ltl.Pure(true)
case isPure(antecedent, true):
return consequent
}
return ltl.Implies(antecedent, consequent)
default:
return formula
}
}
func isPure(formula ltl.Formula, value bool) bool {
pure, ok := formula.(ltl.PureFormula)
return ok && pure.Value == value
}
// stateless erases every temporal operator. The replacement constant follows
// the position's polarity: under an even number of negations a temporal
// sub-formula is dropped as satisfied, and under an odd number as failed, so
// that in both cases its negation cannot refute anything either.
func stateless(formula ltl.Formula, positive bool) ltl.Formula {
switch concrete := formula.(type) {
case ltl.AlwaysFormula, ltl.NextFormula, ltl.EventuallyFormula:
return ltl.Pure(positive)
case ltl.NowFormula:
return ltl.Now(stateless(concrete.Inner, positive))
case ltl.NotFormula:
return ltl.Not(stateless(concrete.Inner, !positive))
case ltl.AndFormula:
return ltl.And(stateless(concrete.Left, positive), stateless(concrete.Right, positive))
case ltl.OrFormula:
return ltl.Or(stateless(concrete.Left, positive), stateless(concrete.Right, positive))
case ltl.ImpliesFormula:
return ltl.Implies(
stateless(concrete.Antecedent, !positive),
stateless(concrete.Consequent, positive),
)
default:
return formula
}
}
func oneStep(formula ltl.Formula, positive bool) ltl.Formula {
switch concrete := formula.(type) {
case ltl.AlwaysFormula:
return ltl.Pure(positive)
case ltl.NextFormula:
return ltl.Next(stateless(concrete.Inner, positive))
case ltl.EventuallyFormula:
return ltl.EventuallyWithinSteps(stateless(concrete.Inner, positive), tripleWindow)
case ltl.NowFormula:
return ltl.Now(oneStep(concrete.Inner, positive))
case ltl.NotFormula:
return ltl.Not(oneStep(concrete.Inner, !positive))
case ltl.AndFormula:
return ltl.And(oneStep(concrete.Left, positive), oneStep(concrete.Right, positive))
case ltl.OrFormula:
return ltl.Or(oneStep(concrete.Left, positive), oneStep(concrete.Right, positive))
case ltl.ImpliesFormula:
return ltl.Implies(
oneStep(concrete.Antecedent, !positive),
oneStep(concrete.Consequent, positive),
)
default:
return formula
}
}
// propertyClass splits safety from liveness by the property's top-level form:
// a reachability goal is liveness, everything else is a safety obligation
// re-asserted at each observation.
func propertyClass(formula ltl.Formula) string {
if _, ok := formula.(ltl.EventuallyFormula); ok {
return "liveness"
}
return "safety"
}
func topLevelForm(formula ltl.Formula) string {
switch concrete := formula.(type) {
case ltl.AlwaysFormula:
return "always" + boundSuffix(concrete.HasStepBound, concrete.StepBound, concrete.Duration)
case ltl.EventuallyFormula:
return "eventually" + boundSuffix(concrete.HasStepBound, concrete.StepBound, concrete.Duration)
default:
return "predicate"
}
}
func boundSuffix(
hasStepBound bool,
stepBound int,
duration time.Duration,
) string {
switch {
case hasStepBound:
return fmt.Sprintf(" within %d steps", stepBound)
case duration > 0:
return fmt.Sprintf(" within %s", duration)
default:
return ""
}
}
@@ -0,0 +1,222 @@
package main
import (
"testing"
"time"
"github.com/priyanshujain/sanderling/internal/ltl"
)
// observe drives an evaluator over a fixed sequence of predicate readings and
// returns the observation at which it latched, or 0 if it never did.
func observe(formula ltl.Formula, steps int) int {
evaluator := ltl.NewEvaluator(formula)
base := time.Unix(0, 0)
for step := 1; step <= steps; step++ {
if evaluator.ObserveAtStep(
base.Add(time.Duration(step)*time.Second),
step,
) == ltl.VerdictViolated {
return step
}
}
if evaluator.Finalize() == ltl.VerdictViolated {
return steps + 1
}
return 0
}
func constant(value bool) ltl.Formula {
return ltl.ThunkNamed("p", func() (bool, error) { return value, nil })
}
func TestSingleStateHoldsWhereRefutationNeedsTheNextObservation(t *testing.T) {
property := ltl.Always(
ltl.Implies(ltl.Now(constant(true)), ltl.Next(constant(false))),
)
if engine := observe(property, 4); engine == 0 {
t.Fatal(
"the engine was expected to refute a next obligation that never held",
)
}
if reduced := observe(singleStateFormula(property), 4); reduced != 0 {
t.Errorf(
"single-state refuted at observation %d, and one observation cannot see a next",
reduced,
)
}
if reduced := observe(singleStepFormula(property), 4); reduced == 0 {
t.Error("single-step was expected to refute a one-step obligation")
}
}
func TestSingleStateRefutesAPredicateReadAtOneObservation(t *testing.T) {
property := ltl.Always(constant(false))
if reduced := observe(singleStateFormula(property), 3); reduced != 1 {
t.Errorf(
"single-state latched at %d, want the first observation",
reduced,
)
}
}
func TestSingleStateKeepsANegatedTemporalHarmless(t *testing.T) {
property := ltl.Always(ltl.Not(ltl.Eventually(constant(false))))
if reduced := observe(singleStateFormula(property), 3); reduced != 0 {
t.Errorf(
"single-state refuted at observation %d; erasing a negated eventually must not manufacture a violation",
reduced,
)
}
}
func TestSingleStepShrinksAReachabilityGoalToTwoObservations(t *testing.T) {
property := ltl.EventuallyWithinSteps(constant(false), 500)
if engine := observe(property, 4); engine != 5 {
t.Fatalf(
"a 500-step window closes only at run end, latched at %d",
engine,
)
}
if reduced := observe(singleStepFormula(property), 4); reduced != 2 {
t.Errorf(
"single-step latched at %d, want the second observation",
reduced,
)
}
if !truncatesWindow(property) {
t.Error(
"a 500-step window shortened to two observations must be reported as truncated",
)
}
}
// sequence returns a predicate that reads the given values, one per call, so
// each evaluator gets its own reading counter.
func sequence(readings []bool) ltl.Formula {
position := 0
return ltl.ThunkNamed("p", func() (bool, error) {
value := readings[position]
position++
return value, nil
})
}
func TestSingleStepConvictsAGoalTheEngineSeesReached(t *testing.T) {
readings := []bool{false, false, true, true}
if engine := observe(ltl.EventuallyWithinSteps(sequence(readings), 4), 4); engine != 0 {
t.Fatalf(
"the engine latched at %d, and the goal is reached inside its window",
engine,
)
}
if reduced := observe(singleStepFormula(ltl.EventuallyWithinSteps(sequence(readings), 4)), 4); reduced != 2 {
t.Errorf(
"single-step latched at %d; an obligation armed at the first observation must not reach the third",
reduced,
)
}
}
func TestTruncatesWindowIgnoresAnObligationThatAlreadyFits(t *testing.T) {
property := ltl.Always(
ltl.Implies(ltl.Now(constant(true)), ltl.Next(constant(true))),
)
if truncatesWindow(property) {
t.Error("a next spans two observations already and is not truncated")
}
}
func TestPropertyClassAndFormComeFromTheTopLevel(t *testing.T) {
safety := ltl.Always(constant(true))
liveness := ltl.EventuallyWithinSteps(constant(true), 575)
if got := propertyClass(safety); got != "safety" {
t.Errorf("class of an always: got %q", got)
}
if got := propertyClass(liveness); got != "liveness" {
t.Errorf("class of an eventually: got %q", got)
}
if got := topLevelForm(liveness); got != "eventually within 575 steps" {
t.Errorf("form: got %q", got)
}
if got := topLevelForm(ltl.EventuallyWithin(constant(true), 3*time.Second)); got != "eventually within 3s" {
t.Errorf("form: got %q", got)
}
}
func TestSingleStepCannotExpressAWindowLongerThanATriple(t *testing.T) {
for _, testCase := range []struct {
name string
property ltl.Formula
expresses bool
}{
{"reachability goal in steps", ltl.EventuallyWithinSteps(constant(true), 575), false},
{"reachability goal in time", ltl.EventuallyWithin(constant(true), 3*time.Second), false},
{"unbounded reachability goal", ltl.Eventually(constant(true)), false},
{"window a triple spans exactly", ltl.EventuallyWithinSteps(constant(true), tripleWindow), true},
{"deadline nested under an always", ltl.Always(ltl.Implies(
ltl.Now(constant(true)), ltl.EventuallyWithin(constant(true), 3*time.Second))), false},
{"one-step obligation under an always", ltl.Always(ltl.Implies(
ltl.Now(constant(true)), ltl.Next(constant(true)))), true},
{"predicate under an always", ltl.Always(constant(true)), true},
} {
t.Run(testCase.name, func(t *testing.T) {
if singleStepExpresses(testCase.property) != testCase.expresses {
t.Errorf("single-step expresses %s: got %t, want %t",
testCase.name, !testCase.expresses, testCase.expresses)
}
})
}
}
// A shorter window under a negation asks for less rather than more, so the
// triple's refutations of it stay sound and the property stays expressible.
// Reporting it as inexpressible would push a defect the single-step oracle can
// genuinely catch into the temporal-only column.
func TestSingleStepStillExpressesANegatedWindow(t *testing.T) {
property := ltl.Always(
ltl.Not(ltl.EventuallyWithinSteps(constant(false), 575)),
)
if !singleStepExpresses(property) {
t.Error(
"a window shortened under a negation is weakened, not strengthened",
)
}
if truncatesWindow(property) {
t.Error(
"the truncation marker must not fire where shortening only weakens",
)
}
}
func TestSingleStateCannotExpressWhatItsRewriteEmpties(t *testing.T) {
for _, testCase := range []struct {
name string
property ltl.Formula
expresses bool
}{
{"predicate at one observation", ltl.Always(constant(true)), true},
{"reachability goal", ltl.EventuallyWithinSteps(constant(true), 575), false},
{"next obligation under an always", ltl.Always(ltl.Implies(
ltl.Now(constant(true)), ltl.Next(constant(true)))), false},
{"deadline under an always", ltl.Always(ltl.Implies(
ltl.Now(constant(true)), ltl.EventuallyWithin(constant(true), 3*time.Second))), false},
{"negated eventually under an always", ltl.Always(ltl.Not(ltl.Eventually(constant(false)))), false},
{"predicate conjoined with a next", ltl.Always(ltl.And(constant(true), ltl.Next(constant(true)))), true},
} {
t.Run(testCase.name, func(t *testing.T) {
if singleStateExpresses(testCase.property) != testCase.expresses {
t.Errorf("single-state expresses %s: got %t, want %t",
testCase.name, !testCase.expresses, testCase.expresses)
}
})
}
}
+30 -16
View File
@@ -20,22 +20,25 @@ import (
var Version = "dev"
type testOptions struct {
spec string
bundleID string
platform string
avd string
device string
iosDevice string
iosAppPath string
androidAppPath string
duration time.Duration
maxSteps int
arm string
seed int64
output string
clearData bool
generator string
exitOnViolation bool
spec string
bundleID string
platform string
avd string
device string
iosDevice string
iosAppPath string
androidAppPath string
duration time.Duration
maxSteps int
arm string
seed int64
output string
clearData bool
generator string
labelSource string
exitOnViolation bool
allowNoProperties bool
allowNoGeneratorActions bool
}
const topUsage = `sanderling is a property-based UI fuzzer for mobile apps.
@@ -71,7 +74,10 @@ func parseTestArgs(args []string, stderr io.Writer) (testOptions, error) {
flagSet.BoolVar(&options.clearData, "clear-data", true, "clear app data before launching so each run starts from a fresh install; pass --clear-data=false to resume prior state")
flagSet.StringVar(&options.arm, "arm", "", "experiment cell label, recorded in meta.json so a directory of runs can be attributed to a cell")
flagSet.StringVar(&options.generator, "generator", "seeded", "action generator: seeded (weighted random) or llm (model picks from the same candidate set; requires generator = llm() in the spec)")
flagSet.StringVar(&options.labelSource, "label-source", "visible-text", "how candidates are named to the llm generator: visible-text (what a user reads) or resource-id (the identifier the app assigned). The seeded generator picks by index and ignores this")
flagSet.BoolVar(&options.exitOnViolation, "exit-on-violation", false, "stop the run at the first property violation and exit 2, so CI can tell a found bug (2) from a broken harness (1)")
flagSet.BoolVar(&options.allowNoProperties, "allow-no-properties", false, "run a spec that registers no properties. Such a run judges nothing and can only report no violations, so it is refused by default; pass this when the run measures what the spec extracts")
flagSet.BoolVar(&options.allowNoGeneratorActions, "allow-no-generator-actions", false, "finish a run the action generator never drove. Such a run judged whatever screen the spec's setup left it on and explored nothing, so it is refused by default; pass this when the run measures where the generator reaches and reaching nothing is the measurement")
if err := flagSet.Parse(args); err != nil {
return testOptions{}, err
}
@@ -94,6 +100,14 @@ func parseTestArgs(args []string, stderr io.Writer) (testOptions, error) {
default:
return testOptions{}, fmt.Errorf("unsupported generator: %q (seeded, llm)", options.generator)
}
// Rejected here rather than defaulted, because a campaign that finishes with
// the wrong labelling and a plausible output directory is worse than one
// that never starts.
switch options.labelSource {
case "visible-text", "resource-id":
default:
return testOptions{}, fmt.Errorf("unsupported label source: %q (visible-text, resource-id)", options.labelSource)
}
return options, nil
}
+115
View File
@@ -10,6 +10,7 @@ import (
"time"
"github.com/priyanshujain/sanderling/internal/testrun"
"github.com/priyanshujain/sanderling/internal/verifier"
)
func TestParseTestArgs_Defaults(t *testing.T) {
@@ -132,6 +133,63 @@ func TestParseTestArgs_RejectsUnknownGenerator(t *testing.T) {
}
}
// The label-source cases compare against the verifier's own constants, because
// the flag and the code that reads it are the two halves of one contract: a
// rename on either side would otherwise leave every run silently labelled by
// the default channel while meta.json claimed the other one.
func TestParseTestArgs_LabelSourceDefaultsToVisibleText(t *testing.T) {
options, err := parseTestArgs([]string{"--spec", "s.ts", "--bundle-id", "com.example"}, io.Discard)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if options.labelSource != verifier.LabelSourceVisibleText {
t.Fatalf("label source default: got %q, want %q", options.labelSource, verifier.LabelSourceVisibleText)
}
}
func TestParseTestArgs_AcceptsResourceIDLabelSource(t *testing.T) {
options, err := parseTestArgs([]string{
"--spec", "s.ts",
"--bundle-id", "com.example",
"--label-source", verifier.LabelSourceResourceID,
}, io.Discard)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if options.labelSource != verifier.LabelSourceResourceID {
t.Fatalf("label source: got %q, want %q", options.labelSource, verifier.LabelSourceResourceID)
}
}
func TestParseTestArgs_RejectsUnknownLabelSource(t *testing.T) {
_, err := parseTestArgs([]string{
"--spec", "s.ts",
"--bundle-id", "com.example",
"--label-source", "resource_id",
}, io.Discard)
if err == nil || !strings.Contains(err.Error(), "unsupported label source") {
t.Fatalf("expected unsupported-label-source error, got %v", err)
}
}
func TestPipelineOptionsCarriesTheExperimentCell(t *testing.T) {
options, err := parseTestArgs([]string{
"--spec", "s.ts",
"--bundle-id", "com.example",
"--generator", "llm",
"--label-source", verifier.LabelSourceResourceID,
"--arm", "llm-resource-id",
}, io.Discard)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
pipeline := pipelineOptions(options)
if pipeline.Generator != "llm" || pipeline.LabelSource != verifier.LabelSourceResourceID || pipeline.Arm != "llm-resource-id" {
t.Errorf("cell lost between the flags and the pipeline: generator=%q labelSource=%q arm=%q",
pipeline.Generator, pipeline.LabelSource, pipeline.Arm)
}
}
func TestParseTestArgs_RejectsUnknownPlatform(t *testing.T) {
_, err := parseTestArgs([]string{
"--spec", "s.ts",
@@ -343,6 +401,62 @@ func TestParseTestArgs_ExitOnViolation(t *testing.T) {
}
}
// TestParseTestArgs_AllowNoPropertiesReachesThePipeline pins the opt-out the
// extraction and portability sweeps pass. Dropped here, the guard is either
// unreachable or permanent: a sweep that deliberately judges nothing cannot ask
// for it, and every other run keeps the false green the guard exists to stop.
func TestParseTestArgs_AllowNoPropertiesReachesThePipeline(t *testing.T) {
base := []string{"--spec", "s.ts", "--bundle-id", "com.example"}
options, err := parseTestArgs(base, io.Discard)
if err != nil {
t.Fatal(err)
}
if pipelineOptions(options).AllowNoProperties {
t.Error("allowNoProperties default: got true, want false")
}
options, err = parseTestArgs(append(base, "--allow-no-properties"), io.Discard)
if err != nil {
t.Fatal(err)
}
if !pipelineOptions(options).AllowNoProperties {
t.Error("--allow-no-properties never reached the pipeline options")
}
}
// The exploration sweeps ask for a run the generator never drove, and they ask
// for that alone. Riding on --allow-no-properties made one flag name two
// unrelated waivers, so a sweep that wanted the property-free one silently lost
// the dead-run detector as well.
func TestParseTestArgs_AllowNoGeneratorActionsIsItsOwnFlag(t *testing.T) {
base := []string{"--spec", "s.ts", "--bundle-id", "com.example"}
options, err := parseTestArgs(base, io.Discard)
if err != nil {
t.Fatal(err)
}
if pipelineOptions(options).AllowNoGeneratorActions {
t.Error("allowNoGeneratorActions default: got true, want false")
}
options, err = parseTestArgs(append(base, "--allow-no-generator-actions"), io.Discard)
if err != nil {
t.Fatal(err)
}
if !pipelineOptions(options).AllowNoGeneratorActions {
t.Error("--allow-no-generator-actions never reached the pipeline options")
}
options, err = parseTestArgs(append(base, "--allow-no-properties"), io.Discard)
if err != nil {
t.Fatal(err)
}
if pipelineOptions(options).AllowNoGeneratorActions {
t.Error("--allow-no-properties still waives the dead-run refusal it does not name")
}
}
// TestExitCode_SeparatesFoundBugsFromBrokenHarnesses pins the three statuses CI
// reads: 0 clean, 2 the run found violations, 1 everything else. A workflow
// that asserts "the known bug is still found" is only meaningful while 2 and 1
@@ -358,6 +472,7 @@ func TestExitCode_SeparatesFoundBugsFromBrokenHarnesses(t *testing.T) {
{"help", flag.ErrHelp, 0, ""},
{"violations found", testrun.ViolationsError{Count: 2}, 2, "violations: 2"},
{"broken harness", errors.New("launch app: no device"), 1, "error: launch app"},
{"spec judges nothing", testrun.NoPropertiesError{Spec: "s.ts"}, 1, "registers no properties"},
} {
t.Run(testCase.name, func(t *testing.T) {
var stderr bytes.Buffer
+28 -18
View File
@@ -8,22 +8,32 @@ import (
)
func runTestPipeline(ctx context.Context, options testOptions, stdout io.Writer) error {
return testrun.Execute(ctx, testrun.Options{
Spec: options.spec,
BundleID: options.bundleID,
Platform: options.platform,
AVD: options.avd,
Device: options.device,
IosDevice: options.iosDevice,
IosAppPath: options.iosAppPath,
AndroidAppPath: options.androidAppPath,
Duration: options.duration,
MaxSteps: options.maxSteps,
Seed: options.seed,
Output: options.output,
ClearData: options.clearData,
Generator: options.generator,
Arm: options.arm,
ExitOnViolation: options.exitOnViolation,
}, stdout)
return testrun.Execute(ctx, pipelineOptions(options), stdout)
}
// pipelineOptions maps the parsed flags onto the pipeline's options. A field
// dropped on the way through here is a run that executes one experiment cell
// and records another, which is worth being able to test on its own.
func pipelineOptions(options testOptions) testrun.Options {
return testrun.Options{
Spec: options.spec,
AllowNoProperties: options.allowNoProperties,
AllowNoGeneratorActions: options.allowNoGeneratorActions,
BundleID: options.bundleID,
Platform: options.platform,
AVD: options.avd,
Device: options.device,
IosDevice: options.iosDevice,
IosAppPath: options.iosAppPath,
AndroidAppPath: options.androidAppPath,
Duration: options.duration,
MaxSteps: options.maxSteps,
Seed: options.seed,
Output: options.output,
ClearData: options.clearData,
Generator: options.generator,
LabelSource: options.labelSource,
Arm: options.arm,
ExitOnViolation: options.exitOnViolation,
}
}