mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 11:07:10 +00:00
record what the model picker did, and make both policies see the same actions (#74)
* feat(llmclient): parse usage and the served model An LLM-in-the-loop evaluation has to report tokens per action and cost per defect, and the client discarded both counters. Served model is recorded separately from the requested one because a router can substitute a differently-priced variant. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record one typed outcome per model-driven step llm-calls.jsonl carries the prompts as sent, the candidate list as the model saw it, the screenshot reference, the raw response, tokens, latency and how the step ended. It sits beside trace.jsonl rather than inside it because every trace line already carries a full hierarchy and both the replay server and the campaign summarizer scan all of them; folding prompts in would grow the lines those readers parse for data neither reads. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): expose the step a snapshot was observed at It lags the runner's current step whenever a transitional tree caused an observation to be skipped, which is exactly when the model is shown an older screen than the step it is choosing for. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): a guard-skipped step is no longer a silent log line The strict echo-skip left only a logger.Warn, so a step the guard discarded was indistinguishable in the trace from a picker that legitimately declined. Any yield or actions-per-hour figure computed from model traces mixed the two. Every path that ends a step without a model-chosen action now records its own outcome. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): record when a chosen action was never dispatched A step could carry a next_action that the foreground guard or an apply error stopped from running, and nothing said so. An executed-action count read off trace.jsonl included actions that acted on nothing. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): document llm-calls.jsonl Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(analyze): survival analysis over campaign directories Steps to first violation with clean runs right-censored at the budget, since per-run yield is a binary at 11 to 45 percent and separating two arms on it would need roughly 80 runs per arm. Kaplan-Meier, log-rank, Wilcoxon rank-sum with Vargha-Delaney A12, Holm within each family. A hand-rolled log-rank that is subtly wrong is a silent-wrong-number generator and would be believed, so every statistic is validated against a published worked example with the source named in the test: R survdiff on aml, Freireich 6-MP, Hollander and Wolfe 1973 for the rank sum, printed p.adjust output for Holm. Two could not be: the k>2 log-rank, guarded by calibration instead, and the tie-corrected variance, checked against an exact permutation variance. Failed and timed-out runs are excluded as missing data and counted by reason, never treated as censored observations, which would bias the result. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): select the candidate label source Candidates takes the label source as an argument rather than storing it, which is what keeps the asymmetry structural: the seeded picker selects by index and never calls Candidates, so the mode cannot reach it. That asymmetry is load-bearing, because it makes the two seeded cells of the factorial a manipulation check with identical draw streams. The identifier ladder deliberately has no text rung. A fallback that reached for text would silently turn one arm back into the other. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(runner): thread the label source to the model picker Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record the label source as arm membership Recorded for seeded runs too, unlike model and instructions. Without it the two seeded cells are indistinguishable in the artifact and the manipulation check cannot be grouped. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(cli): add --label-source Unknown values are rejected at parse time rather than falling back to the default, matching the generator check: a campaign that completes with the wrong arm and a correct-looking output directory is worse than one that fails. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): dedup candidates by what they execute, not how they read The dedup key was the rendered description, which embeds the label, so two distinct controls sharing a visible label collapsed to one entry and the survivor carried the first one's action. The second control was not mislabelled, it was absent from the candidate list, so no policy could reach it. Two scrollable containers collapsed the same way, leaving the second unscrollable. The key is now the executable Action struct itself plus whether the model supplies the typed text, so a new Action field cannot silently fall out of it. Descriptions may now repeat; the numbering disambiguates and the echo guard is index-anchored, not description-anchored. This also makes the label source a pure observation-channel change. It was not one before: the label fed the dedup key, so the two arms of the labelling factor enumerated different-sized candidate lists, in both directions depending on the screen. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): report every action that was chosen and never dispatched applyAction could return nil without calling the driver, so the trace showed an action that looked executed and acted on nothing. Six paths did it: a tap, double-tap or long-press whose coordinates do not resolve and which carries no selector, a long-press whose selector is stale, an empty key press, and a zero-duration wait. It now reports whether it dispatched, and the runner records the reason and clears lastAction so the verifier never attributes the next state to an action that did not run. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): the echo guard admits a repeated description Descriptions can now repeat after candidates dedup by what they execute. The guard is index-anchored, so this pins that a repeated string cannot make it misfire in either direction. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): count dispatched actions, not steps A step where the policy declined has no action, and a step whose action was never dispatched did nothing. Both were being counted as actions by everything downstream. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(analyze): divide by actions that ran Defects per thousand actions counted every step, including steps that chose nothing and steps whose action was never dispatched. The inflation is policy-dependent, so it does not cancel between arms: on the fixture campaign the model arm's yield was reported at 60.3 per thousand against a true 120.7, because half its steps did nothing. A runs.jsonl without the count is refused by name and line rather than read as zero actions, which would report every per-action rate wrongly. The report also carries steps beside actions now, so the gap is visible rather than folded into a denominator. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): lower authored actions the way the seeded arm does The authored descriptor path had no parity guard and diverged from the wire format on almost every verb. A Wait lost its duration and was skipped as a zero-duration wait. A Scroll lost its endpoints and its 250ms. A target that resolved to nothing became a tap at the origin, a phantom focus tap, or a swipe to (0,0) instead of being dropped. An authored target object with no x property panicked the whole run at candidate enumeration: ToInteger was called on a nil goja.Value. A target on the screen origin is still kept, so the drop rule cannot swallow it. Builtins were never affected. They serialize through the same path the seeded arm uses, which the existing policy parity test covers. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(verifier): decode a container-only scroll Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): carry the container on an authored scroll serializeAction sent the container's own point as both endpoints, so an authored Scroll({in, direction}) reached the driver as a drag from a point to itself and did nothing, on the seeded arm. The wire now carries the selector and leaves the drag to the runner, which sizes it from the container's bounds and has always had tested support for it that nothing could produce. No rng runs in the serializer, which lowers an already-drawn action, so the draw stream does not move. Builtin scrolls compute both endpoints and their bytes are unchanged. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): both policies must dispatch the same authored action Compares the recorded driver calls across 13 authored shapes. The builtin path had a parity guard and the authored path had none, which is why it drifted on almost every verb. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(hierarchy): match identifiers by role prefix idPrefix: is id: with starts-with in place of equality, so a list whose rows are named <role>_<record id> is reachable by the durable half. The Android package prefix is skipped the same way id: skips it. Routing both prefix kinds through matchAttr also makes the object form work: {descPrefix: ...} matched nothing on the native side while the web runtime honoured it. * feat(chrome): translate idPrefix to a starts-with id match * feat(spec): match idPrefix in the web runtime The DOM has no package prefix, so the native rule reduces to [id^=]. Both prefix kinds now go through the one key table, which drops the separate descPrefix branch that string and object selectors each carried. * feat(sidecar): match idPrefix in the tap-by-selector path * docs(manual): document the idPrefix selector * feat(replay-ui): render idPrefix targets as a prefix tag * fix(spec): read the injected seed per call Binding it at module scope bound it to whenever the module was first imported, so a test file that imported the runtime before setting SANDERLING_SEED froze the seed at zero for every file after it. The bundler still replaces the expression with a literal. * test(chrome): compare both selector matchers over one live page Selector matching is written once per runtime: internal/hierarchy over the dump, web-runtime.ts over the DOM. Nothing made the two agree, and a selector that resolves on one and not the other is silent, since an empty match yields no action and the run still passes. * fix(hierarchy): give id and desc one meaning in both selector forms The object form fell through to the raw attribute map, which carries no id or desc key on any platform, so {id: "save"} matched nothing while "id:save" matched. The repo's own web spec uses the object form thirty times. Both forms now resolve through one switch. Adds the accepted-key list and UnknownSelectorKeys with it, since the same silence hides any mistyped key. A key some element carries is always accepted, so raw driver attributes stay reachable. * test(hierarchy): pin both selector forms and the unknown-key report * feat(verifier): fail the spec on a selector key that cannot match An empty match is indistinguishable from a screen with no such element, so a mistyped key generates no action for the whole run and the campaign finishes clean having explored nothing. The goja boundary now throws, naming the key and the accepted list. * feat(spec): reject an unknown object-selector key in the web runtime Same rule and the same message as the native side: a key no element can carry throws instead of matching nothing. The accepted list is one list, committed as a fixture both suites assert, so a spec cannot be accepted by one runtime and rejected by the other. * test(spec): pin the unknown-key diagnostic to one text The two runtimes each claimed to raise the other's message and nothing checked it. Both now render the committed text for the committed key. * fix(spec): match a merged label by its leading name on web too The native desc rule accepts the label or the label at the head of an iOS merged label; both web translators compared the whole string, so the same selector matched natively and missed on web. The live-page parity test caught it. * test(chrome): drive the live-page parity test through both selector forms * docs(manual): document object-selector key rules * feat(spec): refuse a multi-item authored sampler while enumerating from().generate() draws from the picker's rng, which exists only inside walkActions. The model policy enumerates authored leaves outside that walk, so the sampler silently yielded its first item on every step: measured over 30 draws the seeded arm reached three targets in roughly equal proportion and the model was offered only the first. The two policies had different action spaces and nothing said so. A single-item sampler short-circuits before the rng, so both policies get the same value and it is not refused. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): stop the run on a sampler the model cannot draw, and offer disabled targets Candidates returns an error now. The refusal is thrown at the draw and wrapped with the source of the leaf that made it, since generate() cannot know which leaf it is inside. Only that marked refusal is fatal: this walk calls every leaf on every step, so promoting the rest would kill model runs the seeded arm survives. Authored actions on a disabled target are no longer dropped from the model's candidate list. The seeded picker executes whatever the leaf authored, and a control the application forgot to re-enable is exactly where boundary defects live, so a policy that cannot attempt it cannot find them. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): abort on a candidate enumeration that refused Recorded as candidates_failed before the run stops, so the trace says why. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(spec): refuse a multi-value generator while enumerating integers, strings, emails and edgeCaseText read the same rng from() does, so under the model policy an authored InputText typed the same value on every step while the seeded arm varied it. That is a silently different experiment, not just a silently different action space. Single-valued spans are exempt, because both policies then get the same value: between(7,7), a zero-length string, and a one-entry corpus. length(4,4) is still refused, since the length is pinned but each character is drawn from 62. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(verifier): setup still draws, and the seeded stream is unmoved Setup runs through the picker with the rng under both policies, so a generator there is legitimate and must keep working. Interleaving enumeration and setup catches the flag leaking out of the model's walk. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): value generators are refused under the model policy too Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio): enumerate authored targets and values instead of sampling Sampling inside an authored leaf is refused under the model policy now, because the draw collapses to its first item there. Each sampled leaf offers one action per value instead. Lists are short, three rather than five, because the two form leaves also carry their submit and the seeded picker splits a leaf's probability across the actions it returns. The doubleTaps path that reaches the planted defect is unchanged at 5.88 percent, since no root or defaults weight moved. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio-web): enumerate authored targets and values, declare the llm generator The two edge-case typing leaves become the typing builtin at their combined weight: that text is deliberately not domain-specific, so naming the field and leaving the text to the policy is the designed path, and it keeps the seeded arm on the corpus while the model writes its own. Total weight is unchanged at 165, so every surviving branch keeps its share and submitTxn stays at 9.70 percent. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs: minimal changes, self-documenting code, tests as first-class Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): key web attrs by the names the markup writes attrs was spread from element.dataset, whose DOMStringMap keys are camelCase, so a spec reading attrs["data-cents"] the way every native host reports it read undefined. In folio-web that left ledgerTxnCount and ledgerBalance permanently zero: someTransactionExists could never be satisfied, balanceMatchesTransaction Delta could never fire, and totalBalanceMatchesAccounts compared 0 to 0 and passed vacuously. Three properties reported nothing because the harness was blind, not because the application was correct. The handle also fills hintText and editable now, so an authored InputText on web names its field the way the same action names it on Android instead of rendering as Type "12.34" into "". Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): name a web handle by the same ladder as a tree element The handle fallback read only text, which is textContent and therefore always empty for an input, so the model could not tell the amount field from the note field. It now mirrors visibleLabel's ladder rather than introducing a second naming scheme. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): attrs carries raw attribute names on web too Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): confirm focus moved before typing InputText tapped its target, slept, then typed. Android and web both inject into whatever holds focus, so a tap that missed sent the whole string somewhere else and nothing reported it. On an emulator with a floating keyboard panel parked over the password field, the tap pressed the keyboard's emoji key and every step appended the password to the email instead, forever, because the setup leaf is guarded on the password being empty. The hierarchy is re-read after the tap and the target, or something in its subtree, must hold focus. Platforms whose hierarchy carries no focused attribute skip the read, so they pay nothing. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): signal a timed-out run so it reaps its sidecar CommandContext kills outright, so a run stopped by --run-timeout never ran its own shutdown and left a sidecar holding a port and a quarter gigabyte, reparented to init and deaf to SIGTERM. The timeout exists for unattended hosts, which is exactly where nobody is watching to reap what it leaves. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * perf(runner): confirm focus only when another element holds it Measured over 717 InputText steps: nothing was focused before the tap 23.8 percent of the time, the target already held focus 60.4 percent, and a different element held it 15.8 percent. Silent corruption is only reachable from that third class, and all four real rejections observed came from it. Gating there keeps every rejection, skips 84.2 percent of the extra hierarchy reads, and recovers about 8 percent of Android run time. The pre-tap and post-tap conditions are now the same predicate stated once. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): record both clocks a run was measured on Duration came from the monotonic clock, which does not advance while a host sleeps: one calibration run under-reported by about 15 minutes. A run now carries monotonic_millis for how long it worked and wall_clock_millis for how much time passed, which is what makes a sleep visible at all. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(analyze): divide per-hour rates by time actually worked A host asleep mid-run tested nothing, and charging that sleep to an arm reports it slower for a reason unrelated to the arm. The legend also claimed wall clock while the number was monotonic. Campaigns written before the split are still read through the old field name so their run hours do not silently zero. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(campaign): wait for the trap instead of racing it The reaping test gave the wedged script one second to install its TERM trap, so a loaded machine signalled it first and the test failed for a reason it does not test. It now waits for the script to say the trap exists, then cancels. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(ltl): keep the authored window on a step-bounded obligation reduce decremented StepBound into the residual, so the trace reported the remaining window rather than the authored one: a within(1915, "steps") showed up as 1875 after 40 steps, and the replay UI renders that string verbatim. The duration case was fixed when bounded windows were made to serialize their resolved deadline; the step case was not, and withinFor's comment claimed otherwise. The window is now immutable and the closing observation is resolved once, which mirrors Deadline exactly. A step counts observations the evaluator reduced, not steps the runner executed, because a skipped step gave the property no chance to discharge and transitional-step rate is itself policy-dependent. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(ltl): pin that a slow policy does not fail on time alone Same 300-observation trace at two cadences: a 300 second bound holds for the seeded arm and violates for the model arm eight observations before the predicate fires, while a step bound holds for both. Green before and after, because the step unit already worked; this pins the property rather than fixing it. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(spec): guard the step unit on the authoring surface Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(folio-web): bound the reachability properties by steps At one model call per step the model arm takes 359 seconds where the seeded arm takes 47, so a second-based deadline reported violations that were the arm's speed rather than the application's behaviour. The three cross-arm reachability properties now bound by steps, derived at the measured 6.383 steps per second. The two auth-transition properties keep seconds: a user waits through those regardless of which policy is driving. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs(manual): a step bound counts observations, not runner steps Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): make the label source a cell dimension A 2x2 of policy against labelling needs the runner to express both factors. It could only express the policy, so half the factorial had to go through --extra, where the manifest would not record what was actually run. Rejected at parse rather than on dispatch: a sweep that finds the bad value on run 1 of 40 has already spent a cell's worth of device time. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): record the label source in the manifest A finished sweep should say which cell it ran without anyone having to remember the invocation. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): name a web field by its hint, not its CSS class visibleLabel reads hintText first for an editable element. The dump never emitted it, so an empty web input fell through text, description and descendant text to its class name, and the model was shown an identifier no user can read on exactly the fields a labelling experiment varies. Same ladder as fieldHint in web-runtime.ts, so one field is named one way on both hosts. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(web-runtime): answer clickable for an element reached through ax The handle hardcoded true, so every text node and container a spec reached through state.ax claimed to be a tap target while the enumeration and the hierarchy dump both resolved it through the tappable selector. The parity test now compares the handle against the enumeration element by element in a real browser, which is where the three answers have to agree. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * docs: every target runs on this machine, so start one rather than skip it Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): a selector the tree cannot resolve is not a focus failure otherElementHoldsFocus answered true when FindNode returned nothing, so an unresolvable target read as "another element holds focus". confirmFocus then re-dumped, resolved nothing again, and errored unconditionally. Three of those in a row abort the run. Not knowing where the target is says nothing about where the text would land. The guard's real case, a resolved target with focus outside its subtree, still errors exactly as before. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): emit data-testid so both resolvers name the same element The V8 host names a web target by data-testid and TapSelector translates that selector into a CSS attribute match, but the dump carried no such attribute and no alias could supply one, since an alias only redirects to a key that already holds the value. tree.Find was therefore always nil for exactly the selectors examples/folio-web tags with. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): name an element only when the selector names it alone ax.findAll stamped every result with the query selector, and resolveCoordinates prefers the tree lookup over the element's own coordinates, so N sibling candidates all executed on the first match. On folio's Home screen the fuzzer could never open any account but the first. The gate tests identity rather than cardinality: no node other than this one answers to the rendered string, checked with the same lookup the runner runs. A rendered object selector can resolve somewhere the query never matched, so counting the query would call that unique. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(web-runtime): hold the V8 host to the same naming gate elementHandle stamped the query selector on every result the same way, so the merge carried the sibling collision onto web for authored ax targets. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(runner): sibling taps reach the driver at their own coordinates Drives 40 real draws from a spec that taps each card, through the picker, the serializer and DecodeAction, and asserts on the points the driver saw. Against the shared-selector bug all 40 landed on the first card. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): an ambiguous name loses to the coordinates it was built from Attribute values match by substring, so a selector that named one element where the candidate was built can name several in the tree it resolves against, and the lookup sent every one of them to the first match. The host gates blank an ambiguous tag at enumeration time; this closes the gap between that moment and the action. A bare-string target carries no coordinates, so the first match stays the answer there rather than dropping an authored action. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(verifier): record an element-valued extractor instead of dropping it An element carries find/findAll host functions, so json.Marshal refused the whole value and the encoder answered nil. ChangedExtractors then emitted no entry: no error, no warning, no value. Project the value the way the web host already does (functions dropped, cycles and over-deep branches null, non-finite numbers null) and turn whatever is still beyond JSON into an error the author sees, rather than a missing extractor. * test(verifier): an unrecordable extractor value is reported, not dropped * test(runner): element-valued extractors reach the trace * docs(spec-language): say what a trace records for an element-valued extractor * docs(claude): add delegation and record-keeping sections delegation says to do installs, builds, test runs and greps in subagents and keep the main context for decisions. record-keeping says a finished task updates the files that describe its subject, writes down what was found, corrects old assumptions in place and verifies against the repository. * feat(driver): declare undelivered-action errors and three optional capabilities ErrGestureUndelivered marks a coordinate gesture that reached no element and ErrSelectorMatchedNothing a selector that named nothing, so the runner can tell them apart from a device fault. Scroller lets a driver whose scroll is not a finger drag take Scroll separately from Swipe. ExceptionReporter and NavigationReporter carry an app's uncaught errors and document-replacing navigations to the runner. * feat(driver): add escape to the pressKey surface escape is a key a spec has real use for and no platform could send it. android maps it to KEYCODE_ESCAPE, the ios companion to HID usage 41 and the in-simulator runner to XCUIKeyboardKey.escape, and the Key union accepts it so it can be written at all. * fix(ios): refuse a gesture the screen has no surface under the hierarchy reaches past the screen wherever a scroll container holds content below the fold, so an action derived from it can name a point no touch lands on. tap, double tap, long press and swipe now report ErrGestureUndelivered for such a point, the far edge exclusive because a touch at x == screenWidth arrives at screenWidth-1. resolveSelectorCenter reports ErrSelectorMatchedNothing rather than a bare error. * feat(ios): derive scrollable from the snapshot's tree depth the companion now emits each node's depth, so the hierarchy mapper can find the containers that clip content reaching past their own frame and mark them scrollable:true, the same fact android reads off uiautomator and the web driver derives from overflow. a dump without depth makes every element a root and roots are never marked, so the legacy bridge reports no scroll rather than a guessed one. * fix(sidecar): stop dropping gestures, selectors and keys in silence a point outside the screen is refused with OUT_OF_RANGE, a selector that matches nothing with NOT_FOUND, and a key with no device-driver equivalent throws instead of pressing nothing. parseBounds also reads uiautomator's [left,top][right,bottom] form, which is what a device actually reports and which left every by-selector tap on a device resolving to nothing. * fix(sidecar): map the driver's refusals onto the gesture errors OUT_OF_RANGE becomes ErrGestureUndelivered on tap, long press, double tap, swipe and the selector fallback; NOT_FOUND on TapSelector becomes ErrSelectorMatchedNothing. without this the runner reads either as a plain apply failure and counts it toward the failure streak. * fix(selectors): resolve text to the innermost match and scan the root in both forms an element's text is its whole subtree's text on web and on ios, so every ancestor of a matching element matched too, up to the root. a match a descendant also makes is now dropped, in internal/hierarchy, in the chrome xpath translation and in the page-side web runtime, so all three resolvers name the same element. a raw attribute now matches on a substring (exact for true/false) the way the docs describe, and tree-level FindBySelector considers the root, so ax.find("id:page") and ax.find({id: "page"}) agree. * feat(hierarchy): store the tree shape and tolerate an unreadable boolean flag a Tree marshalled to json kept only the flat element array, so a stored tree decoded with a nil Root and resolved no selector. it now stores each element's pre-order depth and rebuilds Root from it, re-seating elements so Tree.Elements and &node.Element stay one pointer. a stored tree without depths keeps the old shape. a boolean field the producer sent as something other than a boolean now leaves the flag unset and increments UnreadableFlags rather than failing the whole dump. * fix(chrome): emit every markup attribute and read checked and selected off the property the dump emitted a fixed standard attribute set, so a spec reading data-cents or data-account-id saw undefined on the goja host and nothing at all in the trace. it now keys every attribute by the name the markup writes, derived keys overwriting. checked and selected come from the dom property rather than whatever a component left on the object, which is also what the page-side element handle now reports, so a ticked box reads as ticked instead of reporting its starting state forever. * fix(chrome): scroll a gesture point into view and dispatch trusted input getBoundingClientRect keeps reporting elements the growing document pushed below the emulated viewport, and input coordinates are viewport-relative, so a click below the fold was hit-tested to the document root and the step read as an action that landed. every gesture now scrolls the point back in and reports ErrGestureUndelivered when nothing is under it; a selector that names no node reports ErrSelectorMatchedNothing rather than waiting. swipe dispatches a real touch stream instead of page-synthesized pointer events, scroll is a wheel so its distance is exact rather than a fling, and the second tap of a double tap carries click count 2 so dblclick actually fires. * feat(chrome): read the page's exceptions and navigations, and hold the picker state across them a page navigation replaces the runtime, so the seeded picker restarted the seed's stream at its first draw on every reload and a trace could not tell a reload from a generator repeating itself. the driver now drains the main-frame navigations it saw, reports the page's buffered uncaught errors so state.exceptions is the page's list on the goja host too, and carries the picker's draw position out of v8 and back in around each decision. * feat(trace): version each step and record its logs, exceptions and navigations a step now carries trace_version, the platform log lines and uncaught errors behind state.logs and state.exceptions, the document-replacing navigations seen since the previous step, and observation_error naming why a device read produced no tree. version 0 is a step written before those fields existed, which is what separates a trace that cannot answer the question from a step that had nothing to report. * feat(runner): bound every device call and record the actions that never reached the app observation and apply now run under a timeout, so a driver that stops answering ends the step rather than the run. an undelivered gesture and a selector that matched nothing are recorded as their own skip reasons instead of counting toward the apply-failure streak, a failed observation is counted apart from a screen with nothing on it, and the summary names both. resolveCoordinates hands a point outside the viewport to the driver rather than dropping it: only the driver knows whether it can scroll that point back into reach. exceptions and navigations are collected per step and a Scroll goes to a driver's Scroller when it has one. * feat(verifier): expose extractor names and rebuilt property formulas an offline replay of a trace needs the name-to-index mapping the spec fixed at load, because a trace records extractor values by name, and needs each property's formula built over this verifier's own predicates so a rewritten formula observes exactly what the engine's evaluator does. * feat(testrun): expose the seeded bundle a run loaded BundleSpec produces the goja bundle a run of a spec loaded, seeded as that run was. an offline replay has to load the same javascript, and the seed is one of the bundle's defines, so it is part of the bundle's identity. * feat(tracecorpus): load recorded runs for offline measures reads a run directory's meta and every step, and refuses a step whose trace_version is not the current one: an older step stores no element depths, so its hierarchy decodes with a nil root and a structural hash over it is the empty string for every screen. Discover walks a tree for the directories holding both meta.json and trace.jsonl. * refactor(seedspec): move seed spec parsing out of the campaign command the campaign tool and the sweep tools that drive it have to read a seed specification the same way, or a sweep records an intent that differs from what ran. parseSeeds becomes seedspec.Parse with no behaviour change. * feat(analyze): time an event at the step it was detected and report the quartiles an obligation that never discharges is reported when the run ends, and timing it by the step that armed it recorded a liveness failure flushed at the budget as a violation found on step 1. the survival analysis now measures the detected step, falling back to the origin for campaigns written before the field existed, and says how many events that moved. the report gains the first and third quartiles beside the median. * feat(analyze): add the seed-paired signed-rank comparison and record the holm family --paired contrasts two arms running the same seeds seed by seed with the wilcoxon signed-rank test rather than treating them as two independent samples, reporting the per-seed differences, the sign, a12 within pairs and the seeds usable in one arm only. --question names the family holm corrected within, and the family size is recorded next to the p-values rather than left to the reader to reconstruct. * test(analyze): recover planted effects through the tool's own entry point a pipeline exercised only on data whose answer nobody knows reports that it runs, not that it is right. these plant effects whose value follows from the generating model and require the tool to recover them from campaign directories it reads off disk. * feat(label-coverage): report the addressable share of an app's interactive surface reads the hierarchies a run already recorded and splits each screen's interactive elements by the strongest selector that can name them, so a spec's reach over an app is a number rather than an impression. * feat(exploration-reach): count the distinct structural states a stored run visited the state is the settle path's structural hash of the recorded hierarchy, the same function the drivers wait on, so a state boundary here is the one the harness itself uses. --reference reports the observation at which two runs' hierarchies first differ. trace only: no device, no replay. * feat(defect-identity): count distinct defects across stored runs a property reports at most once per run, so a run-level count is just the number of properties violated. a defect is identified across runs by the property, the action attributed as the origin of the failed obligation and the screen the witness observed. * feat(oracle-reduction): replay stored traces under four reduced oracles re-evaluates each trace offline under the full engine, a crash-only detector, a single-state check and a single-step property triple, and reports what each refutes: the oracles vary while the traces stay fixed, which separates a defect an oracle cannot express from one an explorer never reached. a disagreement with the verdicts a run recorded exits nonzero rather than being counted as a finding. * feat(implementation-sweep): run one campaign against every implementation of a requirement installs, builds and serves each implementation on its own port, then hands the campaign tool the same seed slice, step budget and generator for all of them, so a difference between implementations is not a difference in exploration. the generator and platform are fixed rather than exposed. * feat(corpus-sweep): run one specification against a served corpus of implementations same fixed campaign as implementation-sweep, over a corpus that needs no build. each implementation gets its own port: the corpus holds pairs that write the same localStorage key, and one shared origin is one stored record shared between them. * docs(manual): document innermost text matching, escape and the web scroll verb text: names the innermost match and both selector forms scan the same set, root included. escape joins the key list, with a per-platform note and the rule that a key the platform cannot send fails the action. scroll and swipe are one gesture on a touch device and two different ones in a browser, so say which reaches what. * test(browser): assert an uncaught page exception reaches the trace the page buffered its uncaught errors in v8 and nothing carried them out, so state.exceptions was empty on the host and no trace held one, leaving an offline crash oracle nothing to read. asserts the recorded trace steps rather than the summary. * feat(trace): a step can name the precondition it could not meet A step that never had the app under test in front of it observed something else, and nothing in the trace said so. Index 0 carries the startup gate's verdict, so a run that never started is a trace holding that record and nothing else rather than a run that explored and found nothing. * fix(runner): budget the foreground gate in time, not in polls Eight polls is not a budget. Each poll costs whatever the driver's idle wait happens to take, so the same launch cleared the gate on one device and exhausted it on another: across 80 runs of one app, the gate reported "app never reached foreground" on 38 of 40 Android 14 runs and 0 of 40 Android 16 runs, and it was wrong every time. On API 34 settleForForeground returned in ~100ms, so the eight polls gave up 1.2s into a launch whose window drew at ~1.9s; on API 36 the same eight polls spanned 3s and covered it. The Android 14 runs then spent their first step on the launch animation instead of the app, which is the one-step offset that came out of that campaign looking like a platform difference. The gate now polls for a fixed 15s at a 250ms floor, so its verdict is the same duration on every device, and a verdict of "not in front" ends the run instead of warning and carrying on: a run that never got its app on screen holds no evidence about the app, and the trace records why at step 0. * test(runner): the gate keeps looking until its budget runs out Locks the three facts the campaign was missing: a window that draws after more polls than the old count allowed still clears the gate, an app that never comes forward ends the run with a typed error, and both the startup verdict and every mid-run step the guard could not recover are readable off trace.jsonl. * feat(campaign): count the runs that were never in the app A run that failed its precondition has zero steps and no violations, which is what a short clean run looks like too. The summary now counts the trace records naming an unmet precondition, so a campaign directory answers "how many of these were never in the app" without grepping any log. * docs(triage): name the trace field a run that never started leaves * fix(selectors): tag names the whole tag, not a substring of it matchSelectorKind had no case for tag, so it fell through to the raw attribute path and matched by substring. web-runtime.ts compiles tag to a CSS type selector, so tag:li resolved to <todo-list> on the Go side and to nothing on the web side. * test(chrome): both resolvers agree on tag where a container's name contains its child's * fix(make): build the binary instead of matching the build directory build/ exists at the repo root, so make build was satisfied by the directory and left a stale bin/sanderling in place. * feat(verifier): expose the property names a loaded spec registered * feat(testrun): refuse a run against a spec that registers no properties A spec with no properties drove the app and reported no violations, which is indistinguishable from a spec that judged something and found nothing. Execute now aborts after loading the spec unless the run asks for the opt-out by name. * feat(cli): --allow-no-properties opts a run out of the refusal * docs(cli): document --allow-no-properties * feat(bundle-check): fail a spec that bundles but registers no properties * test(bundle-check): cover the zero-property refusal and pin the reported bundle * feat(folio-web): predicates for counting commits against submit actions * feat(folio-web): judge one commit per submit over a home-card window Replaces totalBalanceMatchesAccounts and balanceMatchesTransactionDelta, which compared two consecutive steps on one screen and so could not see a double submission that lands across a navigation. * fix(folio-web): keep submit live for 400ms after saving Defers the navigation back so the button is tappable while the label reads Saved, widening the double-submit window the counting property is there to catch. * feat(confusion-matrix): score the checker against a blind reviewer Cross-tabulates the properties that fired against the human verdict, one cell per implementation, over a sweep whose implementations all passed their own generated tests. An implementation that failed to build, has no usable run, or carries no filed verdict is listed as missing data rather than counted as a clean cell. Landing the package in one commit because the intermediate splits would not link. * test(confusion-matrix): reject malformed inputs and keep missing data out of the cells * test(confusion-matrix): cover cell assignment, precision and recall * fix(chrome): focus descends into the shadow root document.activeElement names the host, not the node focused inside it, so a Compose-for-wasm app that mounts its tree in a shadow root reported focus on div#app forever. confirmFocus could never be satisfied and every InputText step aborted the run after three tries. selectAllScript already descends the boundary; the tree builder did not. * test(implementation-sweep): supply the binaries the missing-binary test does not test resolveBinaries ranges a map, so with more than one binary absent the error named whichever it reached first. The test passed locally only because bun and sanderling were on PATH; on CI it was a three-way coin flip. * fix(replay-ui): read data-* attributes by their markup names The web runtime now publishes raw markup attribute names, so attrs["step"] read nothing where the markup writes data-step. Three properties went vacuous and exactlyOneStepIsSelected reported false against a UI that was fine. The test also fails if a dataOf key gains no matching attribute, or if an attribute it derives is rendered nowhere. * fix(web-runtime): focus descends into the shadow root here too The Go driver already descends the boundary; the V8 host did not, so the two enumerations disagreed about focus on any shadow-mounted app. The harness now answers activeElement the way a real root does: a root names a node of its own tree, so only the shadow root itself names the field. * fix(implementation-sweep): name every missing binary, in flag order Ranging a map returned at the first failure, so an operator missing three binaries was told about one, fixed it, reran, and was told about the next. The function exists to stop the sweep once rather than fail per implementation and seed. Two identical runs also printed different errors, which is why this reached master as a flake instead of a clean red. * fix(chrome): focus follows the caret to the field it types into Compose for wasm never focuses the semantics node carrying the testTag. It proxies keystrokes through a hidden 1px backing input that is a sibling of the a11y tree, so the node the runner tapped never held focus and confirmFocus refused to type into every Compose text field. Focus is re-attributed to the smallest editable whose box holds the caret's centre. Centre-point rather than full containment because the caret's height comes from the text style and the field's from its layout box, so a taller font would silently drop back to refusing. * fix(corpus-sweep): name every missing binary, in flag order Same map-ranging bug as the sibling tool, and this copy had no test on the missing-binary path at all. * fix(web-runtime): a handle answers editable for itself, not its container isContentEditable is inherited, so every span inside a contenteditable div called itself typeable. collectTargets and the chrome dump both require the element itself to match; the handle was the one that did not. * test(chrome): a hinted field is not named by its css class The fixture inputs carried no class at all, so the test could not fail the way the bug did. They now carry folio-web-shaped classes, and the test asserts the editable gate the hint is read behind. * test(chrome): the handle and the enumeration agree on editable too The helper compared clickable alone, so the inherited-contenteditable bug was caught by unit test only and never in a real browser. * fix(web-runtime): focus follows the caret to the field it types into Mirrors the driver, so the two hosts agree about focus on a Compose page. The harness inherits custom properties down the parent chain the way CSS does, so an implementation matching the inline style attribute fails. * fix(campaign): name every missing required flag, in flag order Five required flags ranged as a map, so omitting three told the operator about one, chosen at random. * fix(corpus-sweep): name every missing required flag, in flag order * fix(implementation-sweep): name every missing required flag, in flag order * fix(confusion-matrix): name every missing required flag, in flag order * ci: pin the idb-companion tap to the formula the companion is staged from The tap moved to 1.5.0, whose bundle has no top-level Frameworks/, and prepare.sh stages bin/ and Frameworks/ as siblings because the binary resolves through @rpath. Floating on it also made the hard-coded companion-1.1.8 output name a lie. The ios-assets cache does not cover this: it restores and make rebuilds anyway, because checkout stamps prepare.sh newer than the archived tarball. Master was green only because its last run predated the bump. * fix(campaign): refuse to start on a device that is not there A sweep launched at six serials, three of which had been deleted from the host. 19 of 20 runs were lost, and not because half the devices were wrong: a worker on a dead serial fails in about 31 seconds and immediately pulls another seed, so three bad workers drained sixteen seeds while the three good workers were still inside their first run. Fast failure is more dangerous than slow failure, because the fast failure consumes the resource the slow one would have left alone. Preflight names every missing serial before the first seed is dispatched. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(campaign): quarantine a device that keeps failing fast Preflight cannot catch a device that disappears mid-sweep, which is what happened: the serials were alive the previous day. Three consecutive failures under two minutes, with no run that worked in between, is a property of the device and not a coincidence. The manifest records which device was quarantined and which seeds have no result, so an aborted sweep says so in its own artefact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record the device a run executed on meta.json carried the host but not the device, so a trace could not say what hardware produced it without the campaign manifest beside it. An experiment splitting cells across api levels could only join them through that manifest. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * ci: let a restored ios bundle survive make's mtime check The cache restored and the build ran anyway: a restored tarball keeps the mtime it was archived with while checkout stamps the sources, so make read every bundle as stale. Both logged Cache hit and rebuilt regardless. Dating the bundles after their sources fixes the lie where it is told. Order-only prerequisites would have fixed it in make, but a laptop has no cache key, so editing prepare.sh would silently embed the previous tarball. The formula version joins the key because a hit now decides what gets embedded, and the key was blind to the brew install: the 1.1.8 and 1.5.0.b2 runs shared a key. * fix(confusion-matrix): a campaign that died is missing data, not a true negative The sweep-level loop excluded a run on launch_error alone, while excludedBecause already checked the campaign process's exit code. An interrupted campaign wrote exit_code -1 with an empty launch_error, so its one completed seed scored the implementation as a clean cell on a tenth of the planned evidence. The fixture builder wrote one exit code into both the sweep record and the campaign run record, which is why no test could tell the two levels apart. * fix(analyze): censor a clean run at the steps it ran, and refuse mismatched budgets A run stops at whichever comes first, the step budget or --duration, so a clean run that reached the wall clock exited with fewer steps than the budget and was still credited with the whole of it. The model arm pays a network call and a screenshot per step, so it reaches the wall sooner and was handed exposure it never had. Nothing checked that two arms shared a budget either. Thirty identical clean runs under budgets of 400 and 100 read a12 0.000 and p 1.685e-14 from the rank-sum while the log-rank in the same report read p 1.0000. groupArms already refused this within one arm. The claims the old convention left in comments and report lines are corrected rather than left standing beside the new behaviour. * fix(runner): a source that was asked and handed nothing says so NextAction returning ErrNoAction left the step with no skip reason, so a run whose every model call failed on transport, a non-2xx, an empty choices array or an echo mismatch printed no violations and exited 0. Only llm-calls.jsonl knew it had never touched the app. The reason now travels the path the other five already take, so it reaches the trace, the summary, and the campaign's dispatched-action exclusion. A held step never asks and keeps carrying nothing. * feat(testrun): refuse a run that dispatched none of its actions Same argument as the zero-property refusal: an instrument that drove nothing must not report a clean result. A first-screen violation still wins under --exit-on-violation, --allow-no-properties exempts the extraction sweeps that measure reach rather than judge, and one dispatched action is enough, so a generator quiet on some screens is untouched. * docs(cli): document --label-source * docs(spec-language): name the hintText selector's host divergence The line said the key matches placeholder alone, which is true of the web runtime and not of the tree, where it resolves against the derived attribute. A spec author reading it wrote a selector that matched on one host and not the other. * feat(bundle-check): --allow-no-properties opts out of the refusal The run path grew the opt-out and the freeze gate did not, so a spec the extraction and portability sweeps register nothing for on purpose could be run but never frozen. The refusal now names the flag the way the runner's does. * test(verifier): an unreadable committed fixture fails, it does not skip The comment said the round trip always runs. A skip on a fixture that is committed turns a missing or truncated file into a green. * fix(testrun): the refusal asks whether the generator drove, not whether anything did A dead provider against folio exited 0 on a real emulator: the login setup dispatched three actions before the generator was consulted, so DispatchedActions was 3 and the gate never fired while the generator drove the app zero times across 83 steps. Any spec with a login setup was immune, which is the normal case. Summary counts generator actions separately and the refusal reads that. NoActionsDispatchedError becomes NoGeneratorActionsError, because a run that dispatched three login taps was lying in the old name. * feat(runner): the summary says how many steps the generator drove A green llm run carried no evidence of how much the generator actually drove: the count was inferable only from llm-calls.jsonl outcomes, and the number the refusal turns on was invisible in the run's own output. * fix(testrun): an ios run records the simulator it executed on Device was read from --device, which only an android run sets, so every ios meta.json left the field empty and the trace could not say what hardware produced it. * fix(campaign): the action count leaves the setup's login out on a model run Defects per thousand actions divided by every dispatched step, so a spec whose setup logs in inflated the denominator by however many steps that took. It is the same error the run gate had, and it does not cancel between arms. A model run is separable because only an llm-selected action stamps next_action.source. A seeded run is not: its setup returns through the same entry with no marker, and 11261 dispatched steps across the 169 recorded runs carry no source at all, so excluding on it blind would report every seeded run as having explored nothing. The seeded arm counts as before and a test pins that. * feat(hierarchy): an element reports whether it masks what is typed into it ios reads it off SecureTextField, which the companion already sent and nothing read; web reads input[type=password]. Android cannot: the native tree mapper drops the password attribute before the sidecar sees it, so the fact is three-valued and null there rather than a false that would read as "not secure". * fix(verifier): a secure field's typed value never reaches the record A folio login run wrote the account email and password in cleartext into llm-calls.jsonl, 166 times in one run, beside screenshots of the same screens. Three sites rendered it: the recent-action memory, the candidate list, and the trace. One helper now covers all three so a fourth cannot bypass it, and the driver still receives the real text. Android redacts every typed value because it cannot tell a secure field from any other. That asymmetry is deliberate and documented: safe by default on the target that cannot tell. * fix(runner): a secure field's value does not reach state.lastAction either folio extracts lastAction, and extractor values are persisted as extractor_changes, so the password still reached the run directory through the spec after the three render sites were closed. The wrap sits in the runner rather than in lastActionFields because the hosts hold the next step's tree, not the one the action was chosen against: a field that stops being secure between the two would publish what the trace withheld. Live and replay now agree byte for byte. * fix(trace): an action names the generator that produced it The setup exclusion landed for the model arm only, because only a model pick stamped a source. A seeded run returned setup's action through the same entry with no marker, so its denominator still counted the login while the model arm's did not, and the two are compared. serializeAction names setup and seeded on the wire, so both arms are counted by one rule. An already-recorded trace names nothing and keeps exactly the count it was reported with; unattributed_actions counts those steps so the old denominator cannot pass as the new one. TraceVersion is deliberately unbumped: oracle-reduction refuses a differing version, and a bump would make all 169 recorded runs unreplayable. * fix(defect-identity): degrade a redacted origin action to its selector The full action key read the typed value straight from the trace, where redaction renders every value typed into one field as the same string, so two runs that typed different values there collapsed into one identity and the report said nothing about it. The key now drops a redacted value, falls back to the selector for that action, and counts the rows it did that to, so the undercount reads as an undercount. * fix(campaign): a record always says how many actions named no producer An omitted count reads the same as a run recorded before actions carried a source, so the two cannot be told apart by anything downstream. * fix(analyze): read how much of a record's action count names no producer A runs.jsonl written before actions named one has no field, and its whole count is of unknown provenance rather than none of it. * fix(analyze): refuse to compare attributed and unattributed denominators One arm's actions may include the login the spec's setup drove and the other's cannot, so a per-action rate over the two divides by different things and the tests rank the bookkeeping. * fix(analyze): mark an action count of unknown provenance in the report * docs(manual): what an action count with no producer means for a rate * fix(folio): install through adb so a remote adb server works Gradle's install task talks to adb through ddmlib, which reads only ANDROID_ADB_SERVER_PORT and dials the loopback address, so ADB_SERVER_SOCKET never reaches it and `just test` could not touch a remote emulator. Gradle now only assembles the APK and adb does the install, which picks up the same server every other call in the run talks to. * docs(folio): say how to point just test at a remote adb server * test(conformance): the g4 fixture holds what a redacted android trace holds Android reports no secure fact for any field, so every InputText it records writes the redaction placeholder rather than the typed value. The fixture still carried the real value, which is the only reason the gate reported itself as catching the doubling. Two more fixtures come with it: a repeated-character corpus value that reads as its own doubling and must not fail, and a backend that does record the typed value. Red at this commit: G4 reports PASS on a doubled field it cannot see. * fix(testrun): a recorded violation outranks the dead-run refusal A campaign never passes --exit-on-violation, so the refusal was discarding runs that had found something: exit_code 1 in the record and the analysis drops them as missing data. A run that recorded a violation holds a verdict, which is the whole reason the refusal exists. * fix(testrun): the dead-run refusal gets its own opt-out --allow-no-properties was waiving two unrelated refusals, so a sweep passing it for the property-free reason silently lost a detector it never asked to disable, and a run with properties could only get the dead-run exemption by claiming one it did not want. * feat(cli): --allow-no-generator-actions The flag the dead-run refusal names, wired through to the pipeline. The property-free flag goes back to meaning what it says. * refactor(analyze): open the log-rank up to a weight on the risk set The log-rank is one member of a family that differs only in how much each event time counts. Nothing else changes: the counts it reports stay counts whatever the weight, and the published-dataset results are unmoved. * feat(analyze): add the gehan generalized wilcoxon test The rank-sum carried over to right-censored samples: every pair of runs is scored by which one outlived the other, and a pair censoring cannot order counts as half rather than as a difference neither run supports. The effect size and the p-value are the same statistic, and with nothing censored both are exactly what the rank-sum reports. * fix(analyze): compare arms on censored runs, not on flattened step counts stepTimes threw the censoring flag away and handed the rank-sum a plain number per run, so a run the wall clock stopped at step 12 was ranked as one that violated at step 12. That was defensible while every clean run sat at the budget, the largest value any run could take, and it stopped being defensible when a clean run started being censored where it stopped. Twenty runs clean at step 12 against twenty violations at step 100 read a12 0.000 and p 4.683e-10 from the rank-sum, in the same report as a log-rank reading p 1.0000. The pairwise comparison is now the Gehan test over the observations themselves, and the report says how many run pairs censoring left with no order between them, which is how much of the effect size is the null value rather than an observation. * fix(conformance): g4 reads a doubling off the observed field value The typed value stopped reaching the trace on any target that reports no secure fact for the field, which on android is every field, so the gate was comparing the redaction placeholder against itself and passing whatever the driver did. The observed value is not redacted, and a field holding one string twice over is the doubling itself. A value that is a single character repeated stays exempt: the corpus types "a" 4096 times and a pair of spaces, and neither can be told apart from its own doubling. The recorded-value check stays for the targets that do record it, where it also catches a doubling appended to content the field already held. * fix(spec): a secure selector names the password field on web secure is derived from the field type, not written by the markup, so matching it as a raw attribute reached nothing: the key is accepted, no unknown-key error fires, and find answered undefined on web for the field it answers with on ios. false is every editable field that is not a password entry, since an element that is no field reports null and answers to neither value. * test(chrome): resolve the secure selector on both matchers the fixture covers the password entry, the three shapes of editable field that are not one, and a checkbox that is no field at all. * test(chrome): compare the secure fact across both producers it is the fourth fact the dump and the web runtime derive independently, and the one that decides whether a typed value is written into the shared record. three-valued, so the fixture guard requires all three states rather than both polarities. * docs(manual): state what a secure selector matches * test(conformance): g4 keeps checking past an input typed at coordinates An InputText that names no field aborts the analyzer, so the gate reports the whole run as failed and checks none of the steps after it. 129 of the 485 recorded traces hold such a step. Red at this commit: jq stops on a null selector and the gate reports FAIL. * fix(conformance): g4 skips an input that names no field jq splits an empty string into no segments, so reading the last one off an action typed at coordinates threw and took the rest of the run's steps with it. Such a step names nothing to check; the gate now passes over it and keeps checking the ones that do. * test(browser): the exit code a dead run and a violated one actually leave Drives the built binary against a page with nothing to tap and reads the process status, then the same run through campaign to pin what lands in runs.jsonl: exit_code 1 there is a detection the analysis drops as missing data. * fix(spec): keep a secure selector valid beside another key a multi-key object selector concatenates its parts into one compound, and a type selector is valid only at the head of one, so {id, secure} built '[id="pwd"]input[type="password"]' and querySelectorAll threw. * fix(analyze): score a seed pair by which run outlived the other The paired path had the same defect as the unpaired one: it subtracted two step counts and handed the differences to the signed-rank test, so a pair holding a run the wall clock stopped at step 12 entered as a difference neither run supports. Twenty seeds where the first arm was still clean at step 12 and the second violated at step 5 in six of them read sign -1 and p 0.0011, pointing at the arm that never violated. A pair is now scored the way the unpaired comparison scores one and tested by the exact sign test over the pairs whose order censoring determines, which is what the log-rank stratified by seed reduces to here. The signed-rank goes with the differences it needed: a magnitude-based paired test wants a difference from every pair, and the arms censor on different clocks. The median difference stays, over the pairs where both runs violated, and says so. * docs(analyze): name the tests the tool actually runs The --paired flag advertised the signed-rank, two comments and a test message still said rank-sum, and nothing said what rankSum is doing in the tree now that no campaign reaches it. * docs(manual): exit 1 also means a run that holds no verdict And the flag the dead-run refusal now names, which --allow-no-properties used to double as. * docs(skills): quote the summary line the runner prints now The setup skill's empty-page claim was the stale one that mattered: that run records no_action_produced on every step and exits 1, it does not sit at exit 0 with no violations. Numbers remeasured against the counter and throwing fixtures. * test(conformance): g4 sees a doubling appended to what the field held Redaction cost the gate this shape on android: the driver typed the value twice onto existing content, so the whole value is not its own doubling and the typed value is not in the trace to compare against. The recorded-value check still catches it on the backends that record one. Red at this commit: G4 reports PASS on a field that grew by one string twice. * refactor(analyze): hoist the sign test's loop bound * fix(conformance): g4 reads a doubling out of what the field grew by The whole-value check misses a driver that typed the value twice onto content the field already held, which is the append-vs-replace shape the recorded value used to catch before it was redacted. What the field grew by over the snapshot the action was chosen against is the same signal and needs no typed value. Checked against every recorded trace under conformance/runs: 485 traces, 299 of them carrying an InputText, none newly failing. * fix(analyze): write an undefined paired p-value as null, not as NaN A paired contrast where censoring orders no pair has no p-value, and JSON has no NaN, so --json failed with 'marshal summary: json: unsupported value: NaN' and wrote no summary at all after printing a complete report. The two fields join the medians and the rates already carried as pointers, undefined reading as null in the summary and n/a in the report. Reachable since a clean run started being censored where it stopped: an arm the wall clock stops before its partner ever violates orders nothing. * fix(spec): a boolean state selector names what the live element reports clickable, enabled, focused, checked and selected are derived from the element rather than written by the markup, so matching them as raw attributes built [clickable="true"] and reached nothing: the keys are accepted, no unknown-key error fires, and the worked example in docs/manual/spec-language.md found no element on web and passed having checked nothing. Each key is answered by the same function elementHandle derives the fact with, since no CSS says what any of them says: :focus names the shadow host of a focused field as well, :checked misses a checked custom element and answers for a selected option besides, and [checked] is the state the page loaded with rather than the one the user left it in. * fix(spec): keep a tag selector valid beside another key a multi-key object selector concatenates its parts into one compound, and a type selector is valid only at the head of one, so {id, tag} built '[id="amount"]input' and querySelectorAll threw. whether a spec got an exception or an element depended on the order its author wrote the keys in. * fix(chrome): state every boolean flag the dump can state internal/hierarchy writes the attribute a selector matches on only where the producer stated the flag, so a state emitted as null is one no selector can ask about: {clickable: false} and {enabled: false} matched nothing at all against a web dump while matching on android, which states every flag both ways. only secure stays three-valued. * test(chrome): resolve the five state selectors on both matchers the fixture differs one state at a time: a disabled button and an aria-disabled role control, a box ticked by script with no checked attribute beside one cleared by script that has it, and a select whose first option is selected without the markup saying so anywhere. half the states are asked inside one container, because a state the whole page has an opinion about answers with most of the document and a want list nobody can check by reading. * test(chrome): compare checked, selected and focused across both producers the target enumeration carries none of the three, so they reach a spec through the ax handle alone, and a selector naming one of them resolves against that same reading. the shadow fixture holds the focused control inside its shadow root, where document.activeElement names the mount element and only a producer that descends finds the field. * fix(spec): keep a selector out of the head subtree the head renders nothing, so the hierarchy dump drops it and so does the enumeration the picker walks, but a selector still resolved into it: a whole-page findAll answered with <head> and <title> here and with neither on the goja host, which is a divergence the moment a state selector asks a question every element has an answer to. * docs(manual): state what the other boolean state selectors match * fix(folio): refuse to install and fuzz a device nobody named adb falls through to the local server when ADB_SERVER_SOCKET is unset, and claims the only device attached there. That could be a personal handset, and a run installs the app, clears its state and fuzzes it. Every recipe that touches a device now resolves the target through _require-device, which only picks on its own when a single local emulator is all adb sees. * docs(folio): state that android recipes need ANDROID_DEVICE * fix(spec): and text with the keys written beside it a compound object selector dropped text and matched on the other keys alone, so {testTag: "Row", text: "Alice"} selected every row carrying the tag where internal/hierarchy selects the one row the author named. matching more than the spec said is silent: the find lands on a row nobody wrote and every property over it still passes. text is answered against the element the way the boolean states are, since css cannot ask what an element's text says and the xpath that can cannot ask about the rest, and the innermost rule now holds over what the whole selector matched, where internal/hierarchy holds it. a text-only selector still compiles to the same innermost xpath. * test(spec): pin text against the key beside it in either order object keys iterate in insertion order, so the order the author wrote them in decided what a compound selector meant. the innermost rule is pinned over the whole selector's matches: a row whose badge carries the class and the text both is dropped, one whose badge carries the text alone is kept, and a state key is anded before either. * test(chrome): compare a compound text selector across both matchers one page, both resolvers, text written before and after the key beside it. the object form now encodes its keys in the order the filters state them rather than the order a map iterates, so both orders are asked. the row and the badge under it share a class so the innermost rule has something to drop, and {text, clickable} pins that text is anded before that rule runs: the innermost element carrying "January" is the option, and the select is the only element that is both. * docs(manual): state how text combines with the key beside it the object selector section said every pair must match without saying where the innermost rule then lands. * fix(hierarchy): reach the class attribute through className className is an accepted selector key that no producer writes: android reports the view class, ios the element type and the chrome dump el.className, all of them under `class`. With no alias onto that key the selector matched NOTHING here on every platform while the web runtime resolved it against the live DOM, so {className: "status"} named the row and the badge on one host and no element at all on the other. The failure is silent: the key is accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. * test(chrome): compare className across both matchers one page, both resolvers, the two names for the one attribute. class is asked beside className so the pair is pinned to the same elements rather than each to itself: the row and the badge under it both carry it. * test(spec): pin className and class on the same elements this host answers both names against the live DOM and internal/hierarchy now aliases the second onto the first, so a name dropped from the table here would match nothing on web while the dump still answers it. * docs(manual): list className among the cross-platform aliases the key was already typed on the spec surface and already resolved on web, and the alias table said nothing about which attribute it reads. * fix(hierarchy): reach the accessible label through every name for it label and accessibilityLabel aliased onto accessibilityText alone, which only the ios sidecar writes, and alias expansion is ONE level: the hop from accessibilityText to content-desc was never taken, so both keys matched nothing on android and on the chrome dump, which write the fact under content-desc. ariaLabel and contentDescription aliased onto nothing at all and matched nothing anywhere. The web runtime resolves all four against the live DOM, so a selector naming a field this way found it on one host and no element at all on the other. The keys are accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. Each name lists both keys rather than chaining through accessibilityText: transitive expansion would silently widen every existing key at once. * fix(hierarchy): reach a web test tag through testTag and testID Compose for Web writes a test tag as data-testid, which is what the web runtime resolves both names against. testTag aliased onto the three identifier keys and not that one, and testID aliased onto nothing at all, so a tag the web runtime found on every row of a list named no element here and every property over it passed vacuously. * fix(spec): resolve the identifier, label and class aliases against the DOM identifier, accessibilityIdentifier, accessibilityText and elementType are the names ios writes four facts under, and internal/hierarchy aliases each onto the key the other producers write. This table listed none of them, so each fell through to a raw attribute lookup and built [accessibilityIdentifier="summary_card"], which no element carries. Every one of them resolved against the dump on the goja host and named nothing here. The keys are accepted, so no unknown-key error fires, and a property over the element that was never found passes having checked nothing. * fix(spec): an editable or scrollable selector names what this host derives Both facts are derived from the live element rather than written by the markup, and matching them as attributes built [editable="true"], which no page carries. Both resolve against the dump on the goja host, so a spec naming a field or a scroll container that way found it there and no element at all here, with no unknown-key error to say so. Each reads the same function the fact is derived with, so a selector cannot name an element this host calls something else: the handle, the picker's target list and the editable selector all go through isEditable, and scrollable reads the overflow test collectTargets reads. scrollable false names nothing rather than every element that does not scroll: both producers state the fact only where it holds, the way an element that is no field at all answers to neither value of secure. * test(chrome): compare the alias keys and the two derived facts one page, both resolvers, the ten names that resolved on one host only. each alias is asked beside the key it resolves through, so the pair is pinned to the same elements rather than each to itself. the page grows a container that overflows its box and a neighbour that does not, because scrollable is derived from the box: without one the only scrolling element on the page is the document root, whose answer moves with the window. * fix(hierarchy): bounds is a raw attribute, not a cross-platform key Every native dump writes the rectangle out as a string under bounds, and no DOM element carries an attribute of that name, so the key resolved against the dump and matched nothing on web on every page there is. It is accepted, so no unknown-key error said so, and no mapping can be invented for it: there is no DOM fact to map it to. Off the accepted list the web runtime raises the unknown-key error instead of matching nothing in silence, and the key still resolves wherever a producer writes it, through the escape hatch every other raw attribute already uses: a key some element carries is a key that can match, on both sides. * docs(manual): state which attribute each alias reads, and what bounds is the table listed neither name for the accessible label that a web page writes, nor the key a web test tag lands on, and said nothing about elementType. editable and scrollable are boolean states like the rest, and scrollable is the one of them the platforms state only where it holds. bounds is a raw driver attribute rather than an accepted key. * docs(hierarchy): the package doc names every key an alias reaches it described the alias table as it stood before the label and test-tag names reached the keys android and web write, and said nothing about expansion being one level deep, which is why each name has to list every key rather than hop through another alias. * fix(spec): a hint selector names the ladder both producers derive hintText and placeholderValue are the accessible-name ladder, derived from the live element, and compiling them to [placeholder="..."] made them name the wrong field or none at all. A field labelled by an aria-label or a bound <label> carries no placeholder, so it resolved against the dump on the goja host and reached nothing here; one carrying both answered to its placeholder here where the dump answers to its aria-label, which lands a find on an element nobody named. Both keys read the same fieldHint elementHandle and the hierarchy dump (internal/driver/chrome/driver.go) derive the fact with, so a selector cannot name a field this host calls something else. An empty hint names nothing rather than everything that is no field: both producers write the fact only where the ladder answered. placeholder stays the attribute the markup writes, which is what the dump carries under that name too, so a field whose hint is something else still answers to it on both hosts. * fix(chrome): a hint target is not tapped by the placeholder attribute TapSelector is a third resolver, and it built [placeholder="..."] for hintText and placeholderValue too. Now that both matchers read the accessible-name ladder, that CSS names a field whose hint is its aria-label and whose placeholder happens to carry the value, which is an element neither matcher named. No CSS says what the ladder says, so both keys fall through to a match that reaches nothing and the step fails naming the selector, the way every other derived key in this file already does. A selector reaches here only where the dump resolved it to no coordinates at all. * test(chrome): compare the hint keys and placeholder across both matchers The page gains four fields that differ one rung at a time: a bound label, a placeholder, a placeholder an aria-label outranks, and the name the form gives the field. Only the placeholder rung was reachable before, so hintText and placeholderValue named a field on the goja host and no element at all on web for the other three, and named the field here and nothing there for the rung the ladder passed over. placeholder was measured empty on both hosts because nothing on the page carried the attribute, which said nothing about it. It now names the field the markup wrote it on and not the field whose hint is its aria-label. The third resolver reads the same selectors: what TranslateStringSelector builds for a hint key has to match nothing over CDP rather than the field carrying the value as a placeholder. * docs(manual): both hosts read the hint ladder, placeholder is the attribute The web section said the hintText key does not read the ladder on both hosts and told authors to select such a field by attrs.hintText instead. Both hosts read it now, so that instruction is gone rather than left standing beside a newer sentence. placeholder is stated as the attribute the markup writes and nothing more, the tap path is stated as failing by name where no CSS says what the ladder says, and the alias table gains the row it was missing.
This commit is contained in:
243 files changed
+33333
-923
No files matched your search
@@ -39,23 +39,55 @@ runs:
|
||||
|
||||
# idb-companion is not in homebrew-core, only in facebook/homebrew-fb, so
|
||||
# it has to be named by its full tap path. xcodegen and just are core.
|
||||
#
|
||||
# The tap is checked out at the 1.1.8 formula because companionassets stages
|
||||
# bin/ and Frameworks/ as siblings and names its output companion-1.1.8. The
|
||||
# tap moved to 1.5.0, which drops the top-level Frameworks/, so floating on
|
||||
# it broke the staging and made the embedded version string a lie. Moving to
|
||||
# 1.5.0 is a companion change, not a CI one.
|
||||
- name: Install idb-companion, xcodegen and just
|
||||
id: idb
|
||||
if: inputs.platform == 'ios'
|
||||
shell: bash
|
||||
run: brew install facebook/fb/idb-companion xcodegen just
|
||||
env:
|
||||
HOMEBREW_NO_AUTO_UPDATE: 1
|
||||
run: |
|
||||
brew install xcodegen just
|
||||
brew tap facebook/fb
|
||||
git -C "$(brew --repository facebook/fb)" checkout --quiet c0386793f59da10c619787f2aa18d938ef1d69c9
|
||||
brew install facebook/fb/idb-companion
|
||||
echo "companion-version=$(brew list --versions idb-companion | awk '{print $2}')" >> "$GITHUB_OUTPUT"
|
||||
|
||||
# Both asset tarballs are built by the prepare scripts, and the runner
|
||||
# bundle is an xcodebuild of companion/Sources. Keyed on the scripts and
|
||||
# the versions the Makefile embeds, so a later run reuses them. This has to
|
||||
# land before `make sanderling-ios`, which is what consumes them.
|
||||
#
|
||||
# The formula version is in the key because it is the companion tarball's
|
||||
# largest input and it lives outside the repository: prepare.sh copies
|
||||
# whatever brew installed. Without it, a tap pin bumped on its own would hit
|
||||
# a cache filled from the old formula and embed it under the new pin.
|
||||
- name: Cache the companion and runner bundles
|
||||
id: ios-assets
|
||||
if: inputs.platform == 'ios'
|
||||
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
|
||||
with:
|
||||
path: |
|
||||
internal/driver/ioscompanion/companionassets/assets
|
||||
internal/driver/ioscompanion/runnerassets/assets
|
||||
key: ios-assets-${{ runner.os }}-${{ hashFiles('internal/driver/ioscompanion/companionassets/prepare.sh', 'companion/prepare.sh', 'companion/project.yml', 'companion/Sources/**') }}
|
||||
key: ios-assets-${{ runner.os }}-idb${{ steps.idb.outputs.companion-version }}-${{ hashFiles('internal/driver/ioscompanion/companionassets/prepare.sh', 'companion/prepare.sh', 'companion/project.yml', 'companion/Sources/**') }}
|
||||
|
||||
# A restored tarball keeps the mtime it was archived with, while the
|
||||
# checkout stamps the prepare scripts and companion sources with checkout
|
||||
# time, so make reads every restored bundle as older than its sources and
|
||||
# rebuilds it. The key covers all of those sources exactly, so a hit means
|
||||
# the bundles match them and the timestamps are the only thing lying.
|
||||
- name: Date the restored bundles after their sources
|
||||
if: inputs.platform == 'ios' && steps.ios-assets.outputs.cache-hit == 'true'
|
||||
shell: bash
|
||||
run: |
|
||||
touch -c internal/driver/ioscompanion/companionassets/assets/*.tar.gz \
|
||||
internal/driver/ioscompanion/runnerassets/assets/*.tar.gz
|
||||
|
||||
# Without this the emulator falls back to software rendering and every
|
||||
# step costs several seconds.
|
||||
|
||||
@@ -30,6 +30,15 @@
|
||||
order to accommodate a signature change. Its original subject must survive.
|
||||
- Carry intent through test names and assertions rather than through prose comments.
|
||||
|
||||
## Verification Targets
|
||||
|
||||
- All three targets run on this machine when it is macOS: the Android emulator, the iOS
|
||||
simulator, and Chrome. A change to a driver, the runner, the verifier or the spec surface
|
||||
is verified on every target it can affect, not on whichever one is already booted.
|
||||
- If a simulator is not running, start it. "Nothing was booted" is not a reason to skip a
|
||||
target, and neither is a missing tool on PATH: fix discovery or the install flow rather
|
||||
than prefixing the run with environment variables.
|
||||
|
||||
## Git Branch Rules
|
||||
|
||||
- No slashes in branch names (e.g., use `fix-something` not `fix/something`).
|
||||
@@ -58,3 +67,30 @@
|
||||
- Keep commits small: aim for under 20 lines changed per commit.
|
||||
- Don't batch multiple unrelated changes into one commit.
|
||||
- Commit early and often. A working 5-line change is better than a pending 200-line change.
|
||||
|
||||
## Delegation
|
||||
|
||||
- Do the work in subagents, not in the main context. Installs, builds, test runs,
|
||||
file-by-file writing, greps across the tree and any multi-step verification go to an
|
||||
Agent that reports back a short result.
|
||||
- The main context is for deciding what to do, reviewing what comes back, and talking to
|
||||
the user. Keep it clean. Never paste build output, test output or file listings into it.
|
||||
- One specific task per subagent, with the context it needs. Launch independent tasks in
|
||||
parallel in a single message.
|
||||
- Give each agent its own scratchpad subdirectory, named for its task, and tell it to
|
||||
delete nothing it did not create. Agents run concurrently and a shared scratch directory
|
||||
means one deletes another's work mid-run.
|
||||
|
||||
## Keeping the record
|
||||
|
||||
- Finishing a task includes updating the files that describe its subject. Status lines,
|
||||
schedules, gates and readiness notes go stale the moment work lands, and a plan that
|
||||
says "not started" about something that ran is worse than no plan.
|
||||
- Write down what was found, not just what was changed. A measurement, a number, a thing
|
||||
that turned out not to work: it goes in the file where someone would look for it, with
|
||||
the path, commit or number behind it.
|
||||
- Correct old assumptions explicitly. When something turns out to be wrong, fix the
|
||||
sentence that said it rather than adding a newer sentence beside it. Say what it used
|
||||
to claim if the change matters.
|
||||
- Verify against the repository rather than recalling. A file that says it was checked
|
||||
against HEAD and was not is the failure this project exists to catch.
|
||||
@@ -29,7 +29,7 @@ WEB_DIST := replay-ui/dist
|
||||
|
||||
GOLINES := $(shell $(GO) env GOPATH)/bin/golines
|
||||
|
||||
.PHONY: bootstrap proto sidecar sidecar-embed sanderling sanderling-web sanderling-android sanderling-ios install test test-go test-browser test-companion test-kotlin test-folio test-spec-api test-ci-scripts spec-typecheck web-test web-typecheck web-build web-dev replay-dev docs clean release-cli release-npm-dry fmt fmt-go fmt-kotlin fmt-ts fmt-swift
|
||||
.PHONY: bootstrap proto sidecar sidecar-embed sanderling build sanderling-web sanderling-android sanderling-ios install test test-go test-browser test-companion test-kotlin test-folio test-spec-api test-ci-scripts spec-typecheck web-test web-typecheck web-build web-dev replay-dev docs clean release-cli release-npm-dry fmt fmt-go fmt-kotlin fmt-ts fmt-swift
|
||||
|
||||
bootstrap:
|
||||
$(GO) mod download
|
||||
@@ -48,6 +48,12 @@ sidecar-embed: $(SIDECAR_EMBED)
|
||||
|
||||
sanderling: $(SANDERLING_BIN)
|
||||
|
||||
# `build` is a directory at the repository root, so `make build` matched it and
|
||||
# reported nothing to be done while leaving a stale binary in bin/ for the next
|
||||
# command to use. Aliasing it is cheaper than expecting everyone to remember
|
||||
# that the target is named after the binary.
|
||||
build: sanderling
|
||||
|
||||
$(SANDERLING_BIN): $(SIDECAR_EMBED) $(COMPANION_EMBED) $(RUNNER_EMBED) web-build
|
||||
mkdir -p bin
|
||||
$(GO) build -tags "withsidecar withcompanion" -o $(SANDERLING_BIN) ./cmd/sanderling
|
||||
|
||||
@@ -0,0 +1,318 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"slices"
|
||||
"time"
|
||||
)
|
||||
|
||||
type armSummary struct {
|
||||
Arm string `json:"arm"`
|
||||
Generator string `json:"generator,omitempty"`
|
||||
Platform string `json:"platform,omitempty"`
|
||||
StepBudget int `json:"step_budget"`
|
||||
Directories []string `json:"directories"`
|
||||
Recorded int `json:"recorded_runs"`
|
||||
Usable int `json:"usable_runs"`
|
||||
Violated int `json:"violated_runs"`
|
||||
Censored int `json:"censored_runs"`
|
||||
Excluded int `json:"excluded_runs"`
|
||||
ExcludedByReason map[string]int `json:"excluded_by_reason,omitempty"`
|
||||
MissingSeeds []int64 `json:"missing_seeds,omitempty"`
|
||||
EventsHeldAtBudget int `json:"events_held_at_budget"`
|
||||
EventsDetectedAfterOrigin int `json:"events_detected_after_origin"`
|
||||
MedianStepsToFirstViolation *float64 `json:"median_steps_to_first_violation"`
|
||||
FirstQuartileSteps *float64 `json:"first_quartile_steps_to_first_violation"`
|
||||
ThirdQuartileSteps *float64 `json:"third_quartile_steps_to_first_violation"`
|
||||
SurvivalCurve []survivalPoint `json:"survival_curve,omitempty"`
|
||||
ViolationRate *float64 `json:"violation_rate"`
|
||||
TotalSteps int `json:"total_steps"`
|
||||
TotalActions int `json:"total_actions"`
|
||||
UnattributedActions int `json:"unattributed_actions"`
|
||||
TotalRunHours float64 `json:"total_run_hours"`
|
||||
Detections int `json:"detections"`
|
||||
DefectsPerThousandActions *float64 `json:"defects_per_thousand_actions"`
|
||||
DefectsPerHour *float64 `json:"defects_per_hour"`
|
||||
DistinctDefects int `json:"distinct_defects"`
|
||||
SingletonDefects int `json:"singleton_defects"`
|
||||
SingletonFraction *float64 `json:"singleton_fraction"`
|
||||
DefectRunCounts map[string]int `json:"defect_run_counts,omitempty"`
|
||||
}
|
||||
|
||||
type pairwiseResult struct {
|
||||
First string `json:"first"`
|
||||
Second string `json:"second"`
|
||||
FirstSize int `json:"first_size"`
|
||||
SecondSize int `json:"second_size"`
|
||||
Statistic float64 `json:"u"`
|
||||
A12 float64 `json:"a12"`
|
||||
// Unordered is how many of the run pairs behind U and A12 have no order
|
||||
// between them, because they were tied or because censoring stopped one run
|
||||
// before the other violated. Each of them counts as half, so it is also how
|
||||
// much of the effect size is the null value rather than an observation.
|
||||
Unordered int `json:"unordered_pairs"`
|
||||
PValue float64 `json:"p_value"`
|
||||
HolmPValue float64 `json:"holm_p_value"`
|
||||
}
|
||||
|
||||
type analysis struct {
|
||||
GeneratedAt time.Time `json:"generated_at"`
|
||||
Outcome string `json:"outcome"`
|
||||
// Question names the family Holm corrects within. The correction is applied
|
||||
// across the comparisons of one research question and never across the
|
||||
// paper, so the family a p-value was adjusted in has to be recorded next to
|
||||
// it rather than left to the reader to reconstruct.
|
||||
Question string `json:"question,omitempty"`
|
||||
HolmFamilySize int `json:"holm_family_size"`
|
||||
Arms []armSummary `json:"arms"`
|
||||
LogRank *logRankResult `json:"log_rank"`
|
||||
Pairwise []pairwiseResult `json:"pairwise"`
|
||||
Paired *pairedComparison `json:"paired,omitempty"`
|
||||
Notes []string `json:"notes,omitempty"`
|
||||
}
|
||||
|
||||
const outcomeDescription = "steps to first violation, right-censored at the last step a clean run reached"
|
||||
|
||||
func analyse(arms []arm, now time.Time) (analysis, error) {
|
||||
result, testable, err := baseAnalysis(arms, now)
|
||||
if err != nil {
|
||||
return analysis{}, err
|
||||
}
|
||||
if len(testable) >= 2 {
|
||||
result.Pairwise = comparePairs(testable)
|
||||
result.HolmFamilySize = countCorrected(result.Pairwise)
|
||||
}
|
||||
return result, nil
|
||||
}
|
||||
|
||||
// analysePaired is the seed-matched design of the actuation ablation: two arms
|
||||
// running the same seeds, contrasted seed by seed rather than as two
|
||||
// independent samples.
|
||||
func analysePaired(arms []arm, now time.Time) (analysis, error) {
|
||||
result, testable, err := baseAnalysis(arms, now)
|
||||
if err != nil {
|
||||
return analysis{}, err
|
||||
}
|
||||
if len(testable) != 2 {
|
||||
return analysis{}, fmt.Errorf("a paired comparison needs exactly two arms with usable runs, found %d", len(testable))
|
||||
}
|
||||
comparison, err := pairArms(testable[0], testable[1])
|
||||
if err != nil {
|
||||
return analysis{}, err
|
||||
}
|
||||
if comparison.Pairs == 0 {
|
||||
return analysis{}, fmt.Errorf("arms %q and %q share no seed with a usable run in both",
|
||||
testable[0].Name, testable[1].Name)
|
||||
}
|
||||
if comparison.PValue != nil {
|
||||
adjusted := holm([]float64{*comparison.PValue})[0]
|
||||
comparison.HolmPValue = &adjusted
|
||||
result.HolmFamilySize = 1
|
||||
}
|
||||
result.Paired = &comparison
|
||||
return result, nil
|
||||
}
|
||||
|
||||
func baseAnalysis(arms []arm, now time.Time) (analysis, []arm, error) {
|
||||
result := analysis{GeneratedAt: now, Outcome: outcomeDescription}
|
||||
for _, current := range arms {
|
||||
result.Arms = append(result.Arms, summarize(current))
|
||||
}
|
||||
|
||||
var testable []arm
|
||||
for _, current := range arms {
|
||||
if len(current.observations()) > 0 {
|
||||
testable = append(testable, current)
|
||||
}
|
||||
}
|
||||
if len(testable) < len(arms) {
|
||||
result.Notes = append(result.Notes,
|
||||
"arms with no usable runs are reported but left out of the log-rank test and the pairwise comparisons")
|
||||
}
|
||||
if err := sameBudget(testable); err != nil {
|
||||
return analysis{}, nil, err
|
||||
}
|
||||
if err := sameAttribution(testable); err != nil {
|
||||
return analysis{}, nil, err
|
||||
}
|
||||
if len(testable) >= 2 {
|
||||
names := make([]string, len(testable))
|
||||
groups := make([][]observation, len(testable))
|
||||
for index, current := range testable {
|
||||
names[index] = current.Name
|
||||
groups[index] = current.observations()
|
||||
}
|
||||
test := logRank(names, groups)
|
||||
result.LogRank = &test
|
||||
}
|
||||
return result, testable, nil
|
||||
}
|
||||
|
||||
// sameBudget refuses arms that were given different exposure. A clean run is
|
||||
// censored somewhere at or below its arm's budget, so the arm with the larger
|
||||
// budget carries censored runs the smaller arm could not have produced, and
|
||||
// every test that ranks the two against each other reads that as the arm
|
||||
// surviving longer. It is the cross-arm form of what groupArms already refuses
|
||||
// within one arm.
|
||||
func sameBudget(arms []arm) error {
|
||||
for index := 1; index < len(arms); index++ {
|
||||
if arms[index].Budget != arms[0].Budget {
|
||||
return fmt.Errorf("arm %q has step budget %d and arm %q has %d: "+
|
||||
"runs censored at different budgets cannot be compared",
|
||||
arms[0].Name, arms[0].Budget, arms[index].Name, arms[index].Budget)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// sameAttribution refuses arms whose actions were counted against different
|
||||
// denominators. An arm recorded before an action named its producer counts
|
||||
// whatever the spec's setup dispatched among its actions, and an arm recorded
|
||||
// after leaves the login out, so the same rate over the two divides by
|
||||
// different things and the tests rank a bookkeeping difference. Two arms of the
|
||||
// same unknown provenance are diluted alike and compare; one of each does not.
|
||||
func sameAttribution(arms []arm) error {
|
||||
for index := 1; index < len(arms); index++ {
|
||||
unknown, attributed := arms[0], arms[index]
|
||||
if (unknown.unattributedActions() == 0) == (attributed.unattributedActions() == 0) {
|
||||
continue
|
||||
}
|
||||
if unknown.unattributedActions() == 0 {
|
||||
unknown, attributed = attributed, unknown
|
||||
}
|
||||
return fmt.Errorf("arm %q counts %d action(s) of unknown provenance and arm %q counts none: "+
|
||||
"a denominator that may include the spec's setup cannot be compared against one that excludes it",
|
||||
unknown.Name, unknown.unattributedActions(), attributed.Name)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func countCorrected(pairs []pairwiseResult) int {
|
||||
corrected := 0
|
||||
for _, pair := range pairs {
|
||||
if !math.IsNaN(pair.PValue) {
|
||||
corrected++
|
||||
}
|
||||
}
|
||||
return corrected
|
||||
}
|
||||
|
||||
func comparePairs(arms []arm) []pairwiseResult {
|
||||
var pairs []pairwiseResult
|
||||
for first := 0; first < len(arms); first++ {
|
||||
for second := first + 1; second < len(arms); second++ {
|
||||
test := gehanTest(arms[first].observations(), arms[second].observations())
|
||||
pairs = append(pairs, pairwiseResult{
|
||||
First: arms[first].Name,
|
||||
Second: arms[second].Name,
|
||||
FirstSize: test.FirstSize,
|
||||
SecondSize: test.SecondSize,
|
||||
Statistic: test.Statistic,
|
||||
A12: test.A12,
|
||||
Unordered: test.Unordered,
|
||||
PValue: test.PValue,
|
||||
HolmPValue: math.NaN(),
|
||||
})
|
||||
}
|
||||
}
|
||||
// Holm runs over this one family of comparisons. A comparison whose p-value
|
||||
// could not be computed is not part of the family and does not shrink the
|
||||
// correction the others receive.
|
||||
var family []int
|
||||
var raw []float64
|
||||
for index, pair := range pairs {
|
||||
if math.IsNaN(pair.PValue) {
|
||||
continue
|
||||
}
|
||||
family = append(family, index)
|
||||
raw = append(raw, pair.PValue)
|
||||
}
|
||||
for position, adjusted := range holm(raw) {
|
||||
pairs[family[position]].HolmPValue = adjusted
|
||||
}
|
||||
return pairs
|
||||
}
|
||||
|
||||
func summarize(current arm) armSummary {
|
||||
summary := armSummary{
|
||||
Arm: current.Name,
|
||||
Generator: current.Generator,
|
||||
Platform: current.Platform,
|
||||
StepBudget: current.Budget,
|
||||
Directories: current.Directories,
|
||||
Recorded: len(current.Runs),
|
||||
MissingSeeds: current.MissingSeeds,
|
||||
UnattributedActions: current.unattributedActions(),
|
||||
}
|
||||
runsPerDefect := map[string]int{}
|
||||
for _, item := range current.Runs {
|
||||
if item.ExcludedBecause != "" {
|
||||
summary.Excluded++
|
||||
if summary.ExcludedByReason == nil {
|
||||
summary.ExcludedByReason = map[string]int{}
|
||||
}
|
||||
summary.ExcludedByReason[item.ExcludedBecause]++
|
||||
continue
|
||||
}
|
||||
summary.Usable++
|
||||
summary.TotalSteps += item.Steps
|
||||
// Steps and actions differ by the steps that chose no action, the steps
|
||||
// whose action was never dispatched, and the steps the spec's setup
|
||||
// drove into position. Only what the action generator dispatched
|
||||
// explored the app, so only that belongs in a per-action rate.
|
||||
summary.TotalActions += item.Actions
|
||||
summary.TotalRunHours += float64(item.MonotonicMillis) / float64(time.Hour/time.Millisecond)
|
||||
if item.ClampedToBudget {
|
||||
summary.EventsHeldAtBudget++
|
||||
}
|
||||
if item.Violated && item.EventStep > item.OriginStep {
|
||||
summary.EventsDetectedAfterOrigin++
|
||||
}
|
||||
if item.Violated {
|
||||
summary.Violated++
|
||||
} else {
|
||||
summary.Censored++
|
||||
}
|
||||
distinct := slices.Compact(slices.Sorted(slices.Values(item.ViolatedProperties)))
|
||||
summary.Detections += len(distinct)
|
||||
for _, property := range distinct {
|
||||
runsPerDefect[property]++
|
||||
}
|
||||
}
|
||||
|
||||
summary.SurvivalCurve = kaplanMeier(current.observations())
|
||||
if median, ok := medianSurvival(summary.SurvivalCurve); ok {
|
||||
summary.MedianStepsToFirstViolation = &median
|
||||
}
|
||||
if lower, ok := quantileSurvival(summary.SurvivalCurve, 0.25); ok {
|
||||
summary.FirstQuartileSteps = &lower
|
||||
}
|
||||
if upper, ok := quantileSurvival(summary.SurvivalCurve, 0.75); ok {
|
||||
summary.ThirdQuartileSteps = &upper
|
||||
}
|
||||
if summary.Usable > 0 {
|
||||
rate := float64(summary.Violated) / float64(summary.Usable)
|
||||
summary.ViolationRate = &rate
|
||||
}
|
||||
if summary.TotalActions > 0 {
|
||||
perThousand := 1000 * float64(summary.Detections) / float64(summary.TotalActions)
|
||||
summary.DefectsPerThousandActions = &perThousand
|
||||
}
|
||||
if summary.TotalRunHours > 0 {
|
||||
perHour := float64(summary.Detections) / summary.TotalRunHours
|
||||
summary.DefectsPerHour = &perHour
|
||||
}
|
||||
if len(runsPerDefect) > 0 {
|
||||
summary.DefectRunCounts = runsPerDefect
|
||||
summary.DistinctDefects = len(runsPerDefect)
|
||||
for _, count := range runsPerDefect {
|
||||
if count == 1 {
|
||||
summary.SingletonDefects++
|
||||
}
|
||||
}
|
||||
fraction := float64(summary.SingletonDefects) / float64(summary.DistinctDefects)
|
||||
summary.SingletonFraction = &fraction
|
||||
}
|
||||
return summary
|
||||
}
|
||||
@@ -0,0 +1,422 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"io"
|
||||
"math"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func violatingRun(seed int64, steps, origin int, properties ...string) classifiedRun {
|
||||
return classifiedRun{
|
||||
Seed: seed,
|
||||
Steps: steps,
|
||||
Actions: steps,
|
||||
MonotonicMillis: 60_000,
|
||||
OriginStep: origin,
|
||||
EventStep: origin,
|
||||
Violated: true,
|
||||
ViolatedProperties: properties,
|
||||
}
|
||||
}
|
||||
|
||||
func cleanRun(seed int64, steps int) classifiedRun {
|
||||
return classifiedRun{Seed: seed, Steps: steps, Actions: steps, MonotonicMillis: 60_000}
|
||||
}
|
||||
|
||||
func TestSummarize_ArmWhereNoRunViolated(t *testing.T) {
|
||||
summary := summarize(arm{
|
||||
Name: "quiet",
|
||||
Budget: 40,
|
||||
Runs: []classifiedRun{cleanRun(1, 40), cleanRun(2, 40), cleanRun(3, 40)},
|
||||
})
|
||||
if summary.Usable != 3 || summary.Censored != 3 || summary.Violated != 0 {
|
||||
t.Errorf("summary %+v, want three censored runs", summary)
|
||||
}
|
||||
if summary.MedianStepsToFirstViolation != nil {
|
||||
t.Errorf("median %v, want undefined", *summary.MedianStepsToFirstViolation)
|
||||
}
|
||||
if summary.ViolationRate == nil || *summary.ViolationRate != 0 {
|
||||
t.Errorf("violation rate %v, want 0", summary.ViolationRate)
|
||||
}
|
||||
if summary.DistinctDefects != 0 || summary.SingletonFraction != nil {
|
||||
t.Errorf("defects %d singleton fraction %v, want none", summary.DistinctDefects, summary.SingletonFraction)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarize_ArmWhereEveryRunViolated(t *testing.T) {
|
||||
summary := summarize(arm{
|
||||
Name: "loud",
|
||||
Budget: 40,
|
||||
Runs: []classifiedRun{
|
||||
violatingRun(1, 5, 5, "cartTotal"),
|
||||
violatingRun(2, 9, 9, "cartTotal"),
|
||||
violatingRun(3, 11, 11, "cartTotal", "backNavigation"),
|
||||
},
|
||||
})
|
||||
if summary.Violated != 3 || summary.Censored != 0 {
|
||||
t.Errorf("summary %+v, want three events", summary)
|
||||
}
|
||||
if summary.MedianStepsToFirstViolation == nil || *summary.MedianStepsToFirstViolation != 9 {
|
||||
t.Errorf("median %v, want 9", summary.MedianStepsToFirstViolation)
|
||||
}
|
||||
if *summary.ViolationRate != 1 {
|
||||
t.Errorf("violation rate %v, want 1", *summary.ViolationRate)
|
||||
}
|
||||
if summary.Detections != 4 || summary.DistinctDefects != 2 {
|
||||
t.Errorf("detections %d distinct %d, want 4 and 2", summary.Detections, summary.DistinctDefects)
|
||||
}
|
||||
// backNavigation appears in one run of three, cartTotal in all three.
|
||||
if summary.SingletonDefects != 1 || math.Abs(*summary.SingletonFraction-0.5) > 1e-12 {
|
||||
t.Errorf("singletons %d fraction %v, want 1 and 0.5", summary.SingletonDefects, summary.SingletonFraction)
|
||||
}
|
||||
// 25 actions over 3 minutes.
|
||||
if math.Abs(*summary.DefectsPerThousandActions-160) > 1e-9 {
|
||||
t.Errorf("defects per thousand actions %v, want 160", *summary.DefectsPerThousandActions)
|
||||
}
|
||||
if math.Abs(*summary.DefectsPerHour-80) > 1e-9 {
|
||||
t.Errorf("defects per hour %v, want 80", *summary.DefectsPerHour)
|
||||
}
|
||||
}
|
||||
|
||||
// A step that chose no action, and a step whose action was never dispatched,
|
||||
// left the app untouched. Counting them would inflate the denominator of every
|
||||
// per-action rate, and the inflation differs by arm so it does not cancel.
|
||||
func TestSummarize_CountsDispatchedActionsNotSteps(t *testing.T) {
|
||||
summary := summarize(arm{
|
||||
Name: "declines",
|
||||
Budget: 40,
|
||||
Runs: []classifiedRun{
|
||||
{Seed: 1, Steps: 40, Actions: 10, MonotonicMillis: 3_600_000,
|
||||
Violated: true, OriginStep: 12, ViolatedProperties: []string{"cartTotal"}},
|
||||
{Seed: 2, Steps: 40, Actions: 6, MonotonicMillis: 3_600_000},
|
||||
},
|
||||
})
|
||||
if summary.TotalSteps != 80 {
|
||||
t.Errorf("total steps %d, want 80", summary.TotalSteps)
|
||||
}
|
||||
if summary.TotalActions != 16 {
|
||||
t.Errorf("total actions %d, want 16 dispatched of 80 steps", summary.TotalActions)
|
||||
}
|
||||
if summary.DefectsPerThousandActions == nil {
|
||||
t.Fatal("no defects per thousand actions")
|
||||
}
|
||||
expected := 1000.0 / 16.0
|
||||
if math.Abs(*summary.DefectsPerThousandActions-expected) > 1e-9 {
|
||||
t.Errorf("defects per thousand actions %v, want %v", *summary.DefectsPerThousandActions, expected)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarize_ArmThatDispatchedNothingHasNoPerActionRate(t *testing.T) {
|
||||
summary := summarize(arm{
|
||||
Name: "inert",
|
||||
Budget: 20,
|
||||
Runs: []classifiedRun{
|
||||
{Seed: 1, Steps: 20, Actions: 0, MonotonicMillis: 3_600_000,
|
||||
Violated: true, OriginStep: 3, ViolatedProperties: []string{"cartTotal"}},
|
||||
},
|
||||
})
|
||||
if summary.TotalActions != 0 || summary.TotalSteps != 20 {
|
||||
t.Errorf("steps %d actions %d, want 20 and 0", summary.TotalSteps, summary.TotalActions)
|
||||
}
|
||||
if summary.DefectsPerThousandActions != nil {
|
||||
t.Errorf("defects per thousand actions %v, want none with nothing dispatched", *summary.DefectsPerThousandActions)
|
||||
}
|
||||
if summary.DefectsPerHour == nil || *summary.DefectsPerHour != 1 {
|
||||
t.Errorf("defects per hour %v, want 1: the run still consumed an hour", summary.DefectsPerHour)
|
||||
}
|
||||
}
|
||||
|
||||
func runHoursFor(t *testing.T, record map[string]any) float64 {
|
||||
t.Helper()
|
||||
directory := filepath.Join(t.TempDir(), "campaign")
|
||||
record["seed"] = 1
|
||||
writeCampaign(t, directory,
|
||||
map[string]any{"arm": "seeded", "max_steps": 50, "seeds": []int{1}},
|
||||
[]map[string]any{record})
|
||||
arms, err := groupArms([]string{directory})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return summarize(arms[0]).TotalRunHours
|
||||
}
|
||||
|
||||
// A host asleep mid-run advanced the wall clock while testing nothing, so the
|
||||
// sleep has no place in the denominator of a per-hour rate.
|
||||
func TestSummarize_RunHoursCountTheTimeWorkedNotTheTimeThatPassed(t *testing.T) {
|
||||
hours := runHoursFor(t, map[string]any{
|
||||
"exit_code": 0, "steps": 50, "actions": 50,
|
||||
"monotonic_millis": 3_600_000, "wall_clock_millis": 5_400_000,
|
||||
})
|
||||
if hours != 1 {
|
||||
t.Errorf("run hours %v, want the 1 hour worked rather than the 1.5 hours that passed", hours)
|
||||
}
|
||||
}
|
||||
|
||||
// Campaigns recorded before the two clocks were split gave the same monotonic
|
||||
// reading the name duration_millis, and their hours still have to count.
|
||||
func TestSummarize_RunHoursReadCampaignsWrittenBeforeTheClocksWereSplit(t *testing.T) {
|
||||
hours := runHoursFor(t, map[string]any{
|
||||
"exit_code": 0, "steps": 50, "actions": 50, "duration_millis": 3_600_000,
|
||||
})
|
||||
if hours != 1 {
|
||||
t.Errorf("run hours %v, want 1 from the older duration_millis field", hours)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarize_ArmWithNoUsableRunsAfterExclusions(t *testing.T) {
|
||||
summary := summarize(arm{
|
||||
Name: "broken",
|
||||
Budget: 40,
|
||||
Runs: []classifiedRun{
|
||||
{Seed: 1, ExcludedBecause: reasonTimedOut},
|
||||
{Seed: 2, ExcludedBecause: reasonNonzeroExit},
|
||||
{Seed: 3, ExcludedBecause: reasonNonzeroExit},
|
||||
},
|
||||
})
|
||||
if summary.Usable != 0 || summary.Excluded != 3 {
|
||||
t.Errorf("summary %+v, want no usable runs and three exclusions", summary)
|
||||
}
|
||||
if summary.ExcludedByReason[reasonNonzeroExit] != 2 || summary.ExcludedByReason[reasonTimedOut] != 1 {
|
||||
t.Errorf("exclusions %v", summary.ExcludedByReason)
|
||||
}
|
||||
if summary.ViolationRate != nil || summary.MedianStepsToFirstViolation != nil {
|
||||
t.Error("reported a rate or a median for an arm with nothing in it")
|
||||
}
|
||||
if summary.DefectsPerThousandActions != nil || summary.DefectsPerHour != nil {
|
||||
t.Error("reported a yield rate with no actions and no time")
|
||||
}
|
||||
if len(summary.SurvivalCurve) != 0 {
|
||||
t.Errorf("survival curve %v, want empty", summary.SurvivalCurve)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarize_SingleRunArm(t *testing.T) {
|
||||
summary := summarize(arm{Name: "one", Budget: 40, Runs: []classifiedRun{violatingRun(1, 6, 6, "cartTotal")}})
|
||||
if summary.Usable != 1 || summary.Violated != 1 {
|
||||
t.Errorf("summary %+v", summary)
|
||||
}
|
||||
if summary.MedianStepsToFirstViolation == nil || *summary.MedianStepsToFirstViolation != 6 {
|
||||
t.Errorf("median %v, want 6", summary.MedianStepsToFirstViolation)
|
||||
}
|
||||
if summary.SingletonDefects != 1 || *summary.SingletonFraction != 1 {
|
||||
t.Errorf("singletons %d fraction %v, want 1 and 1", summary.SingletonDefects, summary.SingletonFraction)
|
||||
}
|
||||
}
|
||||
|
||||
// Excluded runs must not reach the survival data at all, and the counts must
|
||||
// keep them visible.
|
||||
func TestAnalyse_ExcludedRunsNeverBecomeObservations(t *testing.T) {
|
||||
current := arm{
|
||||
Name: "mixed",
|
||||
Budget: 30,
|
||||
Runs: []classifiedRun{
|
||||
violatingRun(1, 8, 8, "cartTotal"),
|
||||
cleanRun(2, 30),
|
||||
{Seed: 3, ExcludedBecause: reasonTimedOut},
|
||||
},
|
||||
}
|
||||
observations := current.observations()
|
||||
if len(observations) != 2 {
|
||||
t.Fatalf("%d observations, want 2", len(observations))
|
||||
}
|
||||
summary := summarize(current)
|
||||
if summary.Usable != 2 || summary.Excluded != 1 || summary.Violated != 1 || summary.Censored != 1 {
|
||||
t.Errorf("summary %+v", summary)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAnalyse_ArmWithNoUsableRunsIsReportedButNotTested(t *testing.T) {
|
||||
result, err := analyse([]arm{
|
||||
{Name: "a", Budget: 30, Runs: []classifiedRun{violatingRun(1, 4, 4), violatingRun(2, 6, 6)}},
|
||||
{Name: "b", Budget: 30, Runs: []classifiedRun{cleanRun(1, 30), cleanRun(2, 30)}},
|
||||
{Name: "c", Budget: 30, Runs: []classifiedRun{{Seed: 1, ExcludedBecause: reasonNonzeroExit}}},
|
||||
}, time.Unix(0, 0).UTC())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if len(result.Arms) != 3 {
|
||||
t.Fatalf("%d arms reported, want all 3", len(result.Arms))
|
||||
}
|
||||
if result.LogRank == nil || len(result.LogRank.Groups) != 2 {
|
||||
t.Fatalf("log-rank %+v, want the two testable arms", result.LogRank)
|
||||
}
|
||||
if len(result.Pairwise) != 1 {
|
||||
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
|
||||
}
|
||||
if len(result.Notes) == 0 {
|
||||
t.Error("no note explaining the dropped arm")
|
||||
}
|
||||
}
|
||||
|
||||
// With a single testable arm there is nothing to compare against, and the tool
|
||||
// must say so instead of producing a statistic.
|
||||
func TestAnalyse_SingleArmHasNoTests(t *testing.T) {
|
||||
result, err := analyse([]arm{
|
||||
{Name: "a", Budget: 30, Runs: []classifiedRun{violatingRun(1, 4, 4)}},
|
||||
}, time.Unix(0, 0).UTC())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if result.LogRank != nil || len(result.Pairwise) != 0 {
|
||||
t.Errorf("log-rank %+v pairwise %v, want neither", result.LogRank, result.Pairwise)
|
||||
}
|
||||
}
|
||||
|
||||
// Holm is applied within the family of pairwise comparisons, so with three arms
|
||||
// the smallest raw p-value is multiplied by three.
|
||||
func TestComparePairs_AppliesHolmWithinTheFamily(t *testing.T) {
|
||||
arms := []arm{
|
||||
{Name: "a", Budget: 40, Runs: manyRuns(12, 4)},
|
||||
{Name: "b", Budget: 40, Runs: manyRuns(12, 20)},
|
||||
{Name: "c", Budget: 40, Runs: manyRuns(12, 36)},
|
||||
}
|
||||
pairs := comparePairs(arms)
|
||||
if len(pairs) != 3 {
|
||||
t.Fatalf("%d comparisons, want 3", len(pairs))
|
||||
}
|
||||
raw := make([]float64, len(pairs))
|
||||
for index, pair := range pairs {
|
||||
raw[index] = pair.PValue
|
||||
if pair.HolmPValue < pair.PValue-1e-12 {
|
||||
t.Errorf("%s vs %s: holm p %v below raw p %v", pair.First, pair.Second, pair.HolmPValue, pair.PValue)
|
||||
}
|
||||
}
|
||||
expected := holm(raw)
|
||||
for index, pair := range pairs {
|
||||
if math.Abs(pair.HolmPValue-expected[index]) > 1e-12 {
|
||||
t.Errorf("comparison %d holm p %v, want %v", index, pair.HolmPValue, expected[index])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// a12 above one half means the first arm needed more steps before its first
|
||||
// violation, so the arm that finds defects sooner sits below one half.
|
||||
func TestComparePairs_A12DirectionFollowsStepCounts(t *testing.T) {
|
||||
pairs := comparePairs([]arm{
|
||||
{Name: "slow", Budget: 40, Runs: manyRuns(6, 30)},
|
||||
{Name: "fast", Budget: 40, Runs: manyRuns(6, 4)},
|
||||
})
|
||||
if pairs[0].A12 <= 0.5 {
|
||||
t.Errorf("a12 %v for the slower arm listed first, want above 0.5", pairs[0].A12)
|
||||
}
|
||||
}
|
||||
|
||||
func writeCleanCampaign(t *testing.T, directory, name string, budget, steps, runs int) {
|
||||
t.Helper()
|
||||
seeds := make([]int, 0, runs)
|
||||
records := make([]map[string]any, 0, runs)
|
||||
for seed := 1; seed <= runs; seed++ {
|
||||
seeds = append(seeds, seed)
|
||||
records = append(records, map[string]any{
|
||||
"seed": seed, "exit_code": 0, "steps": steps, "actions": steps, "monotonic_millis": 60_000,
|
||||
})
|
||||
}
|
||||
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": budget, "seeds": seeds}, records)
|
||||
}
|
||||
|
||||
// Arms censored at different budgets are not on the same clock: every clean run
|
||||
// of the wider arm outranks every clean run of the narrower one whatever the
|
||||
// app did, so the pairwise and the paired test reach a foregone conclusion the
|
||||
// log-rank in the same report contradicts. groupArms already refuses this
|
||||
// within one arm, and comparing across arms is the same hazard.
|
||||
func TestRun_RefusesToCompareArmsCensoredAtDifferentBudgets(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
wideSteps int
|
||||
arguments []string
|
||||
}{
|
||||
{name: "identical runs under different budgets", wideSteps: 100},
|
||||
{name: "each arm run to its own budget", wideSteps: 400},
|
||||
{name: "paired", wideSteps: 400, arguments: []string{"--paired"}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
root := t.TempDir()
|
||||
wide := filepath.Join(root, "wide")
|
||||
narrow := filepath.Join(root, "narrow")
|
||||
writeCleanCampaign(t, wide, "wide", 400, test.wideSteps, 30)
|
||||
writeCleanCampaign(t, narrow, "narrow", 100, 100, 30)
|
||||
|
||||
err := run(append(test.arguments, wide, narrow), io.Discard, io.Discard)
|
||||
if err == nil {
|
||||
t.Fatalf("%s: arms censored at 400 and at 100 steps were compared without complaint", test.name)
|
||||
}
|
||||
for _, fragment := range []string{"wide", "400", "narrow", "100", "different budgets"} {
|
||||
if !strings.Contains(err.Error(), fragment) {
|
||||
t.Errorf("%s: error %q is missing %q", test.name, err, fragment)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func writeSourcedCampaign(t *testing.T, directory, name string, steps, unattributed int) {
|
||||
t.Helper()
|
||||
const budget, runs, actions = 40, 6, 20
|
||||
seeds := make([]int, 0, runs)
|
||||
records := make([]map[string]any, 0, runs)
|
||||
for seed := 1; seed <= runs; seed++ {
|
||||
seeds = append(seeds, seed)
|
||||
records = append(records, map[string]any{
|
||||
"seed": seed, "exit_code": 0, "steps": steps, "actions": actions,
|
||||
"monotonic_millis": 60_000, "unattributed_actions": unattributed,
|
||||
})
|
||||
}
|
||||
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": budget, "seeds": seeds}, records)
|
||||
}
|
||||
|
||||
// An arm whose actions name no producer counts whatever the spec's setup
|
||||
// dispatched in the denominator of every per-action rate, and an arm whose
|
||||
// actions name one leaves the login out of it. The two denominators measure
|
||||
// different things, so a test that ranks one arm against the other reads a
|
||||
// difference in what was counted as a difference in what the arms found.
|
||||
func TestRun_RefusesToCompareArmsWhoseActionsWereCountedDifferently(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
arguments []string
|
||||
}{
|
||||
{name: "independent samples"},
|
||||
{name: "paired", arguments: []string{"--paired"}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
root := t.TempDir()
|
||||
attributed := filepath.Join(root, "attributed")
|
||||
legacy := filepath.Join(root, "legacy")
|
||||
writeSourcedCampaign(t, attributed, "attributed", 40, 0)
|
||||
writeSourcedCampaign(t, legacy, "legacy", 30, 20)
|
||||
|
||||
err := run(append(test.arguments, attributed, legacy), io.Discard, io.Discard)
|
||||
if err == nil {
|
||||
t.Fatalf("%s: an arm of unknown provenance was tested against an attributed one without complaint", test.name)
|
||||
}
|
||||
for _, fragment := range []string{"attributed", "legacy", "120", "unknown provenance"} {
|
||||
if !strings.Contains(err.Error(), fragment) {
|
||||
t.Errorf("%s: error %q is missing %q", test.name, err, fragment)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Two arms recorded before actions named their producer are on the same
|
||||
// denominator as each other, diluted the same way, so they compare.
|
||||
func TestRun_ComparesTwoArmsThatBothNameNoProducer(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
first := filepath.Join(root, "first")
|
||||
second := filepath.Join(root, "second")
|
||||
writeSourcedCampaign(t, first, "first", 40, 20)
|
||||
writeSourcedCampaign(t, second, "second", 30, 20)
|
||||
|
||||
if err := run([]string{first, second}, io.Discard, io.Discard); err != nil {
|
||||
t.Fatalf("two arms of the same unknown provenance were refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func manyRuns(count, originStep int) []classifiedRun {
|
||||
runs := make([]classifiedRun, 0, count)
|
||||
for index := 0; index < count; index++ {
|
||||
runs = append(runs, violatingRun(int64(index), originStep+index, originStep+index, "cartTotal"))
|
||||
}
|
||||
return runs
|
||||
}
|
||||
@@ -0,0 +1,148 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"math"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// A run stops at whichever comes first, the step budget or the campaign's wall
|
||||
// clock, so an arm that spends more wall clock per step leaves runs censored far
|
||||
// below the budget. Those runs are not observations of a violation at that step,
|
||||
// and the comparison between arms has to read them as the bounds they are.
|
||||
|
||||
type wallClockRun struct {
|
||||
steps int
|
||||
violated bool
|
||||
}
|
||||
|
||||
func stoppedShort(count, steps int) []wallClockRun {
|
||||
runs := make([]wallClockRun, 0, count)
|
||||
for index := 0; index < count; index++ {
|
||||
runs = append(runs, wallClockRun{steps: steps})
|
||||
}
|
||||
return runs
|
||||
}
|
||||
|
||||
func violatedAt(count, steps int) []wallClockRun {
|
||||
runs := make([]wallClockRun, 0, count)
|
||||
for index := 0; index < count; index++ {
|
||||
runs = append(runs, wallClockRun{steps: steps, violated: true})
|
||||
}
|
||||
return runs
|
||||
}
|
||||
|
||||
func writeWallClockCampaign(t *testing.T, directory, name string, budget int, runs []wallClockRun) string {
|
||||
t.Helper()
|
||||
seeds := make([]int, 0, len(runs))
|
||||
records := make([]map[string]any, 0, len(runs))
|
||||
for index, run := range runs {
|
||||
seed := index + 1
|
||||
seeds = append(seeds, seed)
|
||||
record := map[string]any{
|
||||
"seed": seed, "exit_code": 0, "steps": run.steps, "actions": run.steps,
|
||||
"monotonic_millis": 60_000,
|
||||
}
|
||||
if run.violated {
|
||||
record["first_violation_origin_step"] = run.steps
|
||||
record["violated_properties"] = []string{"plantedProperty"}
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
writeCampaign(t, directory, map[string]any{
|
||||
"arm": name, "generator": "seeded", "platform": "android",
|
||||
"max_steps": budget, "seeds": seeds,
|
||||
}, records)
|
||||
return directory
|
||||
}
|
||||
|
||||
func wallClockArms(t *testing.T, early, late []wallClockRun) (string, string) {
|
||||
t.Helper()
|
||||
root := t.TempDir()
|
||||
return writeWallClockCampaign(t, filepath.Join(root, "early"), "early", 400, early),
|
||||
writeWallClockCampaign(t, filepath.Join(root, "late"), "late", 400, late)
|
||||
}
|
||||
|
||||
// Twenty runs stopped clean at step 12 against twenty violations at step 100.
|
||||
// Nothing in the first arm was observed past step 12, so no pair of runs across
|
||||
// the arms has a determined order and there is no difference to report.
|
||||
func TestRun_RunsStoppedBeforeEveryEventCarryNoComparison(t *testing.T) {
|
||||
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), violatedAt(20, 100))
|
||||
pair := analyseCampaigns(t, earlyDirectory, lateDirectory).Pairwise[0]
|
||||
|
||||
if pair.First != "early" || pair.Second != "late" {
|
||||
t.Fatalf("comparison %s vs %s, want early vs late", pair.First, pair.Second)
|
||||
}
|
||||
if math.Abs(pair.A12-0.5) > 1e-12 {
|
||||
t.Errorf("a12 %.4f between an arm censored at 12 and one violating at 100, want 0.5: "+
|
||||
"a run that stopped at step 12 never reached step 100", pair.A12)
|
||||
}
|
||||
if pair.PValue < 0.05 {
|
||||
t.Errorf("p %.3e, want no significant difference: the arms were never observed over the same steps",
|
||||
pair.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// Where the two arms were observed together, over the first twelve steps, the
|
||||
// arm the wall clock stopped is the one that did not violate. The effect size
|
||||
// has to follow that and not the step counts the flattening reads.
|
||||
func TestRun_EffectSizeFollowsWhatCensoringDetermines(t *testing.T) {
|
||||
late := append(violatedAt(6, 5), violatedAt(14, 100)...)
|
||||
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), late)
|
||||
pair := analyseCampaigns(t, earlyDirectory, lateDirectory).Pairwise[0]
|
||||
|
||||
if pair.A12 <= 0.5 {
|
||||
t.Errorf("a12 %.4f, want above 0.5: six of the late arm's runs violated by step 5 "+
|
||||
"and none of the early arm's twenty had violated by step 12", pair.A12)
|
||||
}
|
||||
}
|
||||
|
||||
// The seed-matched contrast reads the same censored runs and reaches the same
|
||||
// conclusion or it is not measuring the same thing.
|
||||
func TestRun_PairedContrastFollowsWhatCensoringDetermines(t *testing.T) {
|
||||
late := append(violatedAt(6, 5), violatedAt(14, 100)...)
|
||||
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), late)
|
||||
result := analyseCampaigns(t, "--paired", earlyDirectory, lateDirectory)
|
||||
|
||||
paired := *result.Paired
|
||||
if paired.First != "early" || paired.Second != "late" {
|
||||
t.Fatalf("paired %s minus %s, want early minus late", paired.First, paired.Second)
|
||||
}
|
||||
if paired.Sign != 1 {
|
||||
t.Errorf("sign %+d, want +1: the late arm is the one seen to violate first, in the six pairs "+
|
||||
"where the order is determined at all", paired.Sign)
|
||||
}
|
||||
if paired.A12 <= 0.5 {
|
||||
t.Errorf("a12 within pairs %.4f, want above 0.5", paired.A12)
|
||||
}
|
||||
}
|
||||
|
||||
// Nothing orders any pair here, so there is no test to report. The summary has
|
||||
// to say that rather than failing to write a number that does not exist: a
|
||||
// NaN p-value is not JSON and the whole summary went unwritten behind it.
|
||||
func TestRun_PairedContrastWithNoOrderedPairSaysSo(t *testing.T) {
|
||||
earlyDirectory, lateDirectory := wallClockArms(t, stoppedShort(20, 12), violatedAt(20, 100))
|
||||
result := analyseCampaigns(t, "--paired", earlyDirectory, lateDirectory)
|
||||
|
||||
paired := *result.Paired
|
||||
if paired.Pairs != 20 || paired.Unordered != 20 {
|
||||
t.Fatalf("paired %+v, want twenty pairs and all of them unordered", paired)
|
||||
}
|
||||
if paired.PValue != nil || paired.HolmPValue != nil {
|
||||
t.Errorf("p %v and holm p %v, want both undefined", paired.PValue, paired.HolmPValue)
|
||||
}
|
||||
if result.HolmFamilySize != 0 {
|
||||
t.Errorf("holm family of %d, want none where nothing was tested", result.HolmFamilySize)
|
||||
}
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{"--paired", earlyDirectory, lateDirectory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "sign test over the 0 ordered pair(s), p n/a, holm p n/a") {
|
||||
t.Errorf("report does not say the test was not run:\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,109 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// These records are the shape a real folio-web campaign produced: an
|
||||
// `eventually` obligation armed on step 1, never satisfied, and reported when
|
||||
// the run ended at the step budget. Timing that event by the step that armed it
|
||||
// puts the first violation on step 1 and reports a median of one step to first
|
||||
// violation for an arm that spent its whole budget before it could know.
|
||||
func liveness(seed int64, budget int, properties ...string) map[string]any {
|
||||
return map[string]any{
|
||||
"seed": seed, "exit_code": 0, "steps": budget, "actions": budget - 1,
|
||||
"monotonic_millis": 15000,
|
||||
"first_violation_origin_step": 1,
|
||||
"first_violation_detected_step": budget,
|
||||
"first_violation_reason": "eventually never satisfied",
|
||||
"violated_properties": properties,
|
||||
}
|
||||
}
|
||||
|
||||
func TestClassify_ObligationReportedAtTheRunEndIsTimedAtItsDetection(t *testing.T) {
|
||||
detected := 40
|
||||
item := classify(runRecord{
|
||||
Seed: 2, Steps: 40, Actions: stepPointer(39),
|
||||
FirstViolationOriginStep: stepPointer(1),
|
||||
FirstViolationDetectedStep: &detected,
|
||||
ViolatedProperties: []string{"someTransactionExists"},
|
||||
}, 40)
|
||||
|
||||
if !item.Violated {
|
||||
t.Fatal("run not marked as violated")
|
||||
}
|
||||
if item.OriginStep != 1 {
|
||||
t.Errorf("origin step %d, want the step that armed the obligation", item.OriginStep)
|
||||
}
|
||||
if item.EventStep != 40 {
|
||||
t.Errorf("event step %d, want the step the run could know, 40", item.EventStep)
|
||||
}
|
||||
current := arm{Budget: 40, Runs: []classifiedRun{item}}
|
||||
observations := current.observations()
|
||||
if len(observations) != 1 || !observations[0].Event || observations[0].Steps != 40 {
|
||||
t.Errorf("observations %+v, want one event at 40", observations)
|
||||
}
|
||||
}
|
||||
|
||||
// A safety property that trips under its own action is detected on the step
|
||||
// that armed it, so nothing about the existing outcome moves.
|
||||
func TestClassify_SafetyViolationKeepsItsOriginStep(t *testing.T) {
|
||||
detected := 12
|
||||
item := classify(runRecord{
|
||||
Steps: 12, FirstViolationOriginStep: stepPointer(12), FirstViolationDetectedStep: &detected,
|
||||
}, 400)
|
||||
if item.EventStep != 12 || item.OriginStep != 12 {
|
||||
t.Errorf("run %+v, want an event at 12", item)
|
||||
}
|
||||
}
|
||||
|
||||
// A campaign written before the detected step was recorded still reads, and
|
||||
// keeps timing its events at the origin.
|
||||
func TestClassify_MissingDetectedStepKeepsTheOrigin(t *testing.T) {
|
||||
item := classify(runRecord{Steps: 30, FirstViolationOriginStep: stepPointer(18)}, 400)
|
||||
if item.EventStep != 18 {
|
||||
t.Errorf("event step %d, want the origin step 18", item.EventStep)
|
||||
}
|
||||
}
|
||||
|
||||
// The whole-pipeline form of the same thing, on the record shape a real web
|
||||
// campaign wrote. Before the outcome was timed at detection this reported a
|
||||
// median of 1 step to first violation.
|
||||
func TestRun_LivenessFlushedAtTheBudgetDoesNotReportAOneStepMedian(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "seeded-web")
|
||||
writeCampaign(t, directory, map[string]any{
|
||||
"arm": "seeded-web", "generator": "seeded", "platform": "web",
|
||||
"max_steps": 40, "seeds": []int{1, 2, 3, 4, 5},
|
||||
}, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 38, "monotonic_millis": 13512},
|
||||
liveness(2, 40, "accountCreationReachable", "someTransactionExists"),
|
||||
liveness(3, 40, "someTransactionExists"),
|
||||
liveness(4, 40, "accountCreationReachable", "someTransactionExists"),
|
||||
{"seed": 5, "exit_code": 0, "steps": 40, "actions": 40, "monotonic_millis": 15433},
|
||||
})
|
||||
|
||||
summary := armByName(t, analyseCampaigns(t, directory), "seeded-web")
|
||||
if summary.MedianStepsToFirstViolation == nil {
|
||||
t.Fatal("median undefined, want it at the budget")
|
||||
}
|
||||
if *summary.MedianStepsToFirstViolation != 40 {
|
||||
t.Errorf("median %v steps to first violation, want 40: no run could know before the budget",
|
||||
*summary.MedianStepsToFirstViolation)
|
||||
}
|
||||
if summary.EventsDetectedAfterOrigin != 3 {
|
||||
t.Errorf("%d events detected after their origin, want 3", summary.EventsDetectedAfterOrigin)
|
||||
}
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{directory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "timed 3 violation(s) at the step they were detected") {
|
||||
t.Errorf("the report does not say the events were timed at detection\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,83 @@
|
||||
package main
|
||||
|
||||
import "math"
|
||||
|
||||
// standardNormalUpperTail is P(Z > z) for a standard normal Z.
|
||||
func standardNormalUpperTail(z float64) float64 {
|
||||
return 0.5 * math.Erfc(z/math.Sqrt2)
|
||||
}
|
||||
|
||||
// chiSquareUpperTail is P(X > x) for a chi-square variate with the given
|
||||
// degrees of freedom, which is the regularized upper incomplete gamma
|
||||
// Q(degreesOfFreedom/2, x/2).
|
||||
func chiSquareUpperTail(x float64, degreesOfFreedom int) float64 {
|
||||
if degreesOfFreedom <= 0 || math.IsNaN(x) {
|
||||
return math.NaN()
|
||||
}
|
||||
if x <= 0 {
|
||||
return 1
|
||||
}
|
||||
return regularizedUpperGamma(float64(degreesOfFreedom)/2, x/2)
|
||||
}
|
||||
|
||||
const (
|
||||
gammaIterationLimit = 2000
|
||||
gammaTolerance = 1e-15
|
||||
gammaTiny = 1e-300
|
||||
)
|
||||
|
||||
// regularizedUpperGamma is Q(shape, x). The series is used below the crossover
|
||||
// and the continued fraction above it, as in Numerical Recipes in C, 2nd ed.,
|
||||
// section 6.2 (gammp/gammq).
|
||||
func regularizedUpperGamma(shape, x float64) float64 {
|
||||
if x < shape+1 {
|
||||
return 1 - lowerGammaSeries(shape, x)
|
||||
}
|
||||
return upperGammaContinuedFraction(shape, x)
|
||||
}
|
||||
|
||||
func lowerGammaSeries(shape, x float64) float64 {
|
||||
term := 1 / shape
|
||||
sum := term
|
||||
for iteration := 1; iteration < gammaIterationLimit; iteration++ {
|
||||
term *= x / (shape + float64(iteration))
|
||||
sum += term
|
||||
if math.Abs(term) < math.Abs(sum)*gammaTolerance {
|
||||
break
|
||||
}
|
||||
}
|
||||
return sum * math.Exp(-x+shape*math.Log(x)-logGamma(shape))
|
||||
}
|
||||
|
||||
// upperGammaContinuedFraction evaluates Q(shape, x) with the modified Lentz
|
||||
// algorithm, Numerical Recipes in C, 2nd ed., section 5.2.
|
||||
func upperGammaContinuedFraction(shape, x float64) float64 {
|
||||
b := x + 1 - shape
|
||||
c := 1 / gammaTiny
|
||||
d := 1 / b
|
||||
h := d
|
||||
for iteration := 1; iteration < gammaIterationLimit; iteration++ {
|
||||
numerator := -float64(iteration) * (float64(iteration) - shape)
|
||||
b += 2
|
||||
d = numerator*d + b
|
||||
if math.Abs(d) < gammaTiny {
|
||||
d = gammaTiny
|
||||
}
|
||||
c = b + numerator/c
|
||||
if math.Abs(c) < gammaTiny {
|
||||
c = gammaTiny
|
||||
}
|
||||
d = 1 / d
|
||||
delta := d * c
|
||||
h *= delta
|
||||
if math.Abs(delta-1) < gammaTolerance {
|
||||
break
|
||||
}
|
||||
}
|
||||
return h * math.Exp(-x+shape*math.Log(x)-logGamma(shape))
|
||||
}
|
||||
|
||||
func logGamma(x float64) float64 {
|
||||
value, _ := math.Lgamma(x)
|
||||
return value
|
||||
}
|
||||
@@ -0,0 +1,74 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Chi-square critical values are the standard published table entries: the
|
||||
// upper-tail probability of each of these statistics is the stated alpha in any
|
||||
// chi-square table, for example Pearson and Hartley, Biometrika Tables for
|
||||
// Statisticians, Table 8.
|
||||
func TestChiSquareUpperTail_MatchesPublishedCriticalValues(t *testing.T) {
|
||||
cases := []struct {
|
||||
statistic float64
|
||||
degreesOfFreedom int
|
||||
expected float64
|
||||
}{
|
||||
{3.841459, 1, 0.05},
|
||||
{6.634897, 1, 0.01},
|
||||
{10.827566, 1, 0.001},
|
||||
{5.991465, 2, 0.05},
|
||||
{9.210340, 2, 0.01},
|
||||
{7.814728, 3, 0.05},
|
||||
{11.344867, 3, 0.01},
|
||||
{9.487729, 4, 0.05},
|
||||
{18.307038, 10, 0.05},
|
||||
}
|
||||
for _, test := range cases {
|
||||
got := chiSquareUpperTail(test.statistic, test.degreesOfFreedom)
|
||||
if math.Abs(got-test.expected) > 1e-6 {
|
||||
t.Errorf("chiSquareUpperTail(%v, %d) = %v, want %v", test.statistic, test.degreesOfFreedom, got, test.expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// For one degree of freedom the upper tail has the closed form erfc(sqrt(x/2)),
|
||||
// which is an independent check on the incomplete gamma routine.
|
||||
func TestChiSquareUpperTail_AgreesWithClosedFormAtOneDegreeOfFreedom(t *testing.T) {
|
||||
for _, statistic := range []float64{0.1, 1, 3.4, 16.79, 40, 120} {
|
||||
expected := math.Erfc(math.Sqrt(statistic / 2))
|
||||
got := chiSquareUpperTail(statistic, 1)
|
||||
if math.Abs(got-expected) > 1e-12*math.Max(1, expected) {
|
||||
t.Errorf("chiSquareUpperTail(%v, 1) = %v, want %v", statistic, got, expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestChiSquareUpperTail_ZeroStatisticIsCertain(t *testing.T) {
|
||||
if got := chiSquareUpperTail(0, 1); got != 1 {
|
||||
t.Errorf("chiSquareUpperTail(0, 1) = %v, want 1", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Standard normal quantiles from any published normal table.
|
||||
func TestStandardNormalUpperTail_MatchesPublishedQuantiles(t *testing.T) {
|
||||
cases := []struct {
|
||||
z float64
|
||||
expected float64
|
||||
}{
|
||||
{1.281552, 0.10},
|
||||
{1.644854, 0.05},
|
||||
{1.959964, 0.025},
|
||||
{2.326348, 0.01},
|
||||
{2.575829, 0.005},
|
||||
{3.090232, 0.001},
|
||||
{0, 0.5},
|
||||
}
|
||||
for _, test := range cases {
|
||||
got := standardNormalUpperTail(test.z)
|
||||
if math.Abs(got-test.expected) > 1e-6 {
|
||||
t.Errorf("standardNormalUpperTail(%v) = %v, want %v", test.z, got, test.expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,314 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"math"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// buildFixtureCampaign writes a campaign directory shaped exactly like the one
|
||||
// the campaign tool emits: campaign.json plus one runs.jsonl line per seed.
|
||||
func buildFixtureCampaign(t *testing.T, directory, armName string, budget int, records []map[string]any) {
|
||||
t.Helper()
|
||||
seeds := make([]int, 0, len(records))
|
||||
for _, record := range records {
|
||||
seeds = append(seeds, record["seed"].(int))
|
||||
}
|
||||
writeCampaign(t, directory, map[string]any{
|
||||
"arm": armName,
|
||||
"generator": "seeded",
|
||||
"platform": "web",
|
||||
"spec_path": "/specs/folio.ts",
|
||||
"bundle_id": "app.folio",
|
||||
"max_steps": budget,
|
||||
"seeds": seeds,
|
||||
"host": "experiment-host",
|
||||
"started_at": "2026-08-12T00:00:00Z",
|
||||
"argument_temp": nil,
|
||||
}, records)
|
||||
}
|
||||
|
||||
func seededArmRecords() []map[string]any {
|
||||
// Ten runs: two violate early, one violates late, six run the budget clean,
|
||||
// one times out and is missing data rather than a censored observation.
|
||||
return []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 2, "exit_code": 0, "steps": 18, "actions": 16, "duration_millis": 120000,
|
||||
"first_violation_origin_step": 14, "violated_properties": []string{"cartTotalMatches"}},
|
||||
{"seed": 3, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 4, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 5, "exit_code": 0, "steps": 44, "actions": 42, "duration_millis": 240000,
|
||||
"first_violation_origin_step": 41, "violated_properties": []string{"cartTotalMatches", "backLeavesApp"}},
|
||||
{"seed": 6, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 7, "exit_code": -1, "timed_out": true, "actions": 0, "duration_millis": 900000},
|
||||
{"seed": 8, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 9, "exit_code": 0, "steps": 60, "actions": 58, "duration_millis": 300000, "first_violation_origin_step": nil},
|
||||
{"seed": 10, "exit_code": 0, "steps": 21, "actions": 19, "duration_millis": 130000,
|
||||
"first_violation_origin_step": 19, "violated_properties": []string{"cartTotalMatches"}},
|
||||
}
|
||||
}
|
||||
|
||||
func llmArmRecords() []map[string]any {
|
||||
// Eight runs: six violate, one clean, one failed to launch. This arm
|
||||
// dispatches an action on about half its steps, which is the asymmetry the
|
||||
// per-action denominator has to survive.
|
||||
return []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 7, "actions": 3, "duration_millis": 400000,
|
||||
"first_violation_origin_step": 5, "violated_properties": []string{"cartTotalMatches"}},
|
||||
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 5, "duration_millis": 420000,
|
||||
"first_violation_origin_step": 8, "violated_properties": []string{"backLeavesApp"}},
|
||||
{"seed": 3, "exit_code": 0, "steps": 60, "actions": 30, "duration_millis": 1800000, "first_violation_origin_step": nil},
|
||||
{"seed": 4, "exit_code": 0, "steps": 5, "actions": 2, "duration_millis": 380000,
|
||||
"first_violation_origin_step": 3, "violated_properties": []string{"cartTotalMatches"}},
|
||||
{"seed": 5, "exit_code": 0, "steps": 13, "actions": 7, "duration_millis": 500000,
|
||||
"first_violation_origin_step": 11, "violated_properties": []string{"cartTotalMatches", "priceNeverNegative"}},
|
||||
{"seed": 6, "exit_code": -1, "actions": 0, "launch_error": "fork/exec sanderling: no such file or directory"},
|
||||
{"seed": 7, "exit_code": 0, "steps": 6, "actions": 3, "duration_millis": 390000,
|
||||
"first_violation_origin_step": 6, "violated_properties": []string{"cartTotalMatches"}},
|
||||
{"seed": 8, "exit_code": 0, "steps": 16, "actions": 8, "duration_millis": 520000,
|
||||
"first_violation_origin_step": 15, "violated_properties": []string{"backLeavesApp"}},
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_EndToEndOverFixtureCampaignDirectories(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
seededDirectory := filepath.Join(root, "seeded-web")
|
||||
llmDirectory := filepath.Join(root, "llm-web")
|
||||
buildFixtureCampaign(t, seededDirectory, "seeded", 60, seededArmRecords())
|
||||
buildFixtureCampaign(t, llmDirectory, "llm", 60, llmArmRecords())
|
||||
summaryPath := filepath.Join(root, "analysis.json")
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{"--json", summaryPath, seededDirectory, llmDirectory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
text := stdout.String()
|
||||
for _, fragment := range []string{
|
||||
"steps to first violation, right-censored at the last step a clean run reached",
|
||||
"log-rank across 2 arms",
|
||||
"pairwise gehan generalized wilcoxon",
|
||||
"llm vs seeded",
|
||||
"excluded 1 run(s) as missing data",
|
||||
} {
|
||||
if !strings.Contains(text, fragment) {
|
||||
t.Errorf("stdout is missing %q\n%s", fragment, text)
|
||||
}
|
||||
}
|
||||
|
||||
body, err := os.ReadFile(summaryPath)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var result analysis
|
||||
if err := json.Unmarshal(body, &result); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if len(result.Arms) != 2 {
|
||||
t.Fatalf("%d arms, want 2", len(result.Arms))
|
||||
}
|
||||
byName := map[string]armSummary{}
|
||||
for _, summary := range result.Arms {
|
||||
byName[summary.Arm] = summary
|
||||
}
|
||||
|
||||
seeded := byName["seeded"]
|
||||
if seeded.Usable != 9 || seeded.Violated != 3 || seeded.Censored != 6 || seeded.Excluded != 1 {
|
||||
t.Errorf("seeded arm %+v, want 9 usable, 3 violated, 6 censored, 1 excluded", seeded)
|
||||
}
|
||||
if seeded.ExcludedByReason[reasonTimedOut] != 1 {
|
||||
t.Errorf("seeded exclusions %v, want one timeout", seeded.ExcludedByReason)
|
||||
}
|
||||
if seeded.MedianStepsToFirstViolation != nil {
|
||||
t.Errorf("seeded median %v, want undefined with 3 of 9 violating",
|
||||
*seeded.MedianStepsToFirstViolation)
|
||||
}
|
||||
if math.Abs(*seeded.ViolationRate-3.0/9.0) > 1e-12 {
|
||||
t.Errorf("seeded violation rate %v, want 1/3", *seeded.ViolationRate)
|
||||
}
|
||||
// cartTotalMatches in 3 runs, backLeavesApp in 1 of 2 distinct defects.
|
||||
if seeded.DistinctDefects != 2 || seeded.SingletonDefects != 1 {
|
||||
t.Errorf("seeded defects %d singletons %d, want 2 and 1", seeded.DistinctDefects, seeded.SingletonDefects)
|
||||
}
|
||||
if seeded.TotalSteps != 443 || seeded.TotalActions != 425 {
|
||||
t.Errorf("seeded steps %d actions %d, want 443 and 425", seeded.TotalSteps, seeded.TotalActions)
|
||||
}
|
||||
|
||||
llm := byName["llm"]
|
||||
if llm.Usable != 7 || llm.Violated != 6 || llm.Censored != 1 || llm.Excluded != 1 {
|
||||
t.Errorf("llm arm %+v, want 7 usable, 6 violated, 1 censored, 1 excluded", llm)
|
||||
}
|
||||
if llm.ExcludedByReason[reasonLaunchError] != 1 {
|
||||
t.Errorf("llm exclusions %v, want one launch error", llm.ExcludedByReason)
|
||||
}
|
||||
if llm.MedianStepsToFirstViolation == nil || *llm.MedianStepsToFirstViolation != 8 {
|
||||
t.Errorf("llm median %v, want 8", llm.MedianStepsToFirstViolation)
|
||||
}
|
||||
// The arm dispatches an action on half its steps, so counting steps would
|
||||
// halve its yield per thousand actions and flatter it against seeded.
|
||||
if llm.TotalSteps != 116 || llm.TotalActions != 58 {
|
||||
t.Errorf("llm steps %d actions %d, want 116 and 58", llm.TotalSteps, llm.TotalActions)
|
||||
}
|
||||
if llm.Detections != 7 {
|
||||
t.Fatalf("llm detections %d, want 7", llm.Detections)
|
||||
}
|
||||
expected := 7000.0 / 58.0
|
||||
if math.Abs(*llm.DefectsPerThousandActions-expected) > 1e-9 {
|
||||
t.Errorf("llm defects per thousand actions %v, want %v", *llm.DefectsPerThousandActions, expected)
|
||||
}
|
||||
|
||||
if result.LogRank == nil {
|
||||
t.Fatal("no log-rank result")
|
||||
}
|
||||
if result.LogRank.DegreesOfFreedom != 1 {
|
||||
t.Errorf("log-rank df %d, want 1", result.LogRank.DegreesOfFreedom)
|
||||
}
|
||||
if result.LogRank.PValue > 0.05 {
|
||||
t.Errorf("log-rank p %v, want the two clearly different arms to separate", result.LogRank.PValue)
|
||||
}
|
||||
if len(result.Pairwise) != 1 {
|
||||
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
|
||||
}
|
||||
pair := result.Pairwise[0]
|
||||
if pair.First != "llm" || pair.Second != "seeded" {
|
||||
t.Errorf("comparison %s vs %s, want arms in sorted order", pair.First, pair.Second)
|
||||
}
|
||||
if pair.A12 >= 0.5 {
|
||||
t.Errorf("a12 %v, want the arm that violates sooner below 0.5", pair.A12)
|
||||
}
|
||||
if pair.HolmPValue != pair.PValue {
|
||||
t.Errorf("holm p %v differs from raw p %v in a family of one", pair.HolmPValue, pair.PValue)
|
||||
}
|
||||
// Six seeded runs and one llm run ran the budget clean, and two censored
|
||||
// runs have no order between them whatever step either stopped on.
|
||||
if pair.Unordered < 6 {
|
||||
t.Errorf("%d unordered pair(s), want at least the six pairs of censored runs", pair.Unordered)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_ReportsBothArmsWhenOneHasNothingUsable(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
good := filepath.Join(root, "good")
|
||||
broken := filepath.Join(root, "broken")
|
||||
buildFixtureCampaign(t, good, "good", 30, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28},
|
||||
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 1000,
|
||||
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
|
||||
})
|
||||
buildFixtureCampaign(t, broken, "broken", 30, []map[string]any{
|
||||
{"seed": 1, "exit_code": 3, "actions": 0},
|
||||
{"seed": 2, "timed_out": true, "exit_code": -1, "actions": 0},
|
||||
})
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{good, broken}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
text := stdout.String()
|
||||
if !strings.Contains(text, "broken") {
|
||||
t.Errorf("the arm with nothing usable is not reported\n%s", text)
|
||||
}
|
||||
if strings.Contains(text, "log-rank across") {
|
||||
t.Errorf("ran a log-rank with only one testable arm\n%s", text)
|
||||
}
|
||||
if !strings.Contains(text, "arms with no usable runs are reported but left out") {
|
||||
t.Errorf("no note about the dropped arm\n%s", text)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_JsonToStdout(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "only")
|
||||
buildFixtureCampaign(t, directory, "only", 20, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 20, "actions": 17},
|
||||
})
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{"--json", "-", "--campaign", directory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
start := strings.Index(stdout.String(), "{")
|
||||
if start < 0 {
|
||||
t.Fatalf("no json in stdout\n%s", stdout.String())
|
||||
}
|
||||
var result analysis
|
||||
if err := json.Unmarshal([]byte(stdout.String()[start:]), &result); err != nil {
|
||||
t.Fatalf("json: %v", err)
|
||||
}
|
||||
if len(result.Arms) != 1 || result.Arms[0].Arm != "only" {
|
||||
t.Errorf("arms %+v", result.Arms)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_RejectsTheSameDirectoryTwice(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "one")
|
||||
buildFixtureCampaign(t, directory, "one", 20, []map[string]any{{"seed": 1, "exit_code": 0, "steps": 20, "actions": 20}})
|
||||
err := run([]string{directory, directory}, io.Discard, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), "twice") {
|
||||
t.Fatalf("error %v, want a refusal to double count", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_RequiresACampaignDirectory(t *testing.T) {
|
||||
if err := run(nil, io.Discard, io.Discard); err == nil {
|
||||
t.Fatal("expected an error with no campaign directories")
|
||||
}
|
||||
}
|
||||
|
||||
// A campaign recorded before actions named their producer cannot say whether
|
||||
// the login taps of the spec's setup are inside its per-action denominator, and
|
||||
// the report has to say so where that denominator is read rather than leave the
|
||||
// reader to date the file.
|
||||
func TestRun_MarksAnArmWhoseActionsNameNoProducer(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "before-source")
|
||||
buildFixtureCampaign(t, directory, "before-source", 30, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28, "duration_millis": 60000},
|
||||
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 60000,
|
||||
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
|
||||
})
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{"--json", "-", directory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
text := stdout.String()
|
||||
if !strings.Contains(text, "37 (37 unattributed)") {
|
||||
t.Errorf("the actions cell does not carry the unattributed count\n%s", text)
|
||||
}
|
||||
if !strings.Contains(text, "unknown provenance") {
|
||||
t.Errorf("the report does not say the denominator's provenance is unknown\n%s", text)
|
||||
}
|
||||
|
||||
var result analysis
|
||||
if err := json.Unmarshal([]byte(text[strings.Index(text, "{"):]), &result); err != nil {
|
||||
t.Fatalf("json: %v", err)
|
||||
}
|
||||
if len(result.Arms) != 1 || result.Arms[0].UnattributedActions != 37 {
|
||||
t.Errorf("arms %+v, want 37 unattributed actions", result.Arms)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_LeavesAnArmWhoseActionsAllNameAProducerUnmarked(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "sourced")
|
||||
buildFixtureCampaign(t, directory, "sourced", 30, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 30, "actions": 28, "duration_millis": 60000, "unattributed_actions": 0},
|
||||
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 9, "duration_millis": 60000, "unattributed_actions": 0,
|
||||
"first_violation_origin_step": 9, "violated_properties": []string{"cartTotalMatches"}},
|
||||
})
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{directory}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if strings.Contains(stdout.String(), "unattributed") || strings.Contains(stdout.String(), "unknown provenance") {
|
||||
t.Errorf("an arm whose every action names a producer was marked\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,55 @@
|
||||
package main
|
||||
|
||||
// Published right-censored datasets whose log-rank and Kaplan-Meier results are
|
||||
// reported in the survival-analysis literature and in R's survival package, so
|
||||
// every expected number in these tests can be checked against a source rather
|
||||
// than against this tool's own output.
|
||||
|
||||
// gehanSixMercaptopurine and gehanPlacebo are remission times in weeks from the
|
||||
// 6-MP versus placebo trial in acute leukaemia, Freireich et al. (1963). This is
|
||||
// the dataset R's survival literature calls gehan. A trailing plus in the
|
||||
// published listing marks a censored time.
|
||||
//
|
||||
// 6-MP: 6, 6, 6, 6+, 7, 9+, 10, 10+, 11+, 13, 16, 17+, 19+, 20+, 22, 23, 25+, 32+, 32+, 34+, 35+
|
||||
// placebo: 1, 1, 2, 2, 3, 4, 4, 5, 5, 8, 8, 8, 8, 11, 11, 12, 12, 15, 17, 22, 23
|
||||
var (
|
||||
gehanSixMercaptopurine = []observation{
|
||||
{6, true}, {6, true}, {6, true}, {6, false},
|
||||
{7, true}, {9, false}, {10, true}, {10, false},
|
||||
{11, false}, {13, true}, {16, true}, {17, false},
|
||||
{19, false}, {20, false}, {22, true}, {23, true},
|
||||
{25, false}, {32, false}, {32, false}, {34, false}, {35, false},
|
||||
}
|
||||
gehanPlacebo = []observation{
|
||||
{1, true}, {1, true}, {2, true}, {2, true}, {3, true},
|
||||
{4, true}, {4, true}, {5, true}, {5, true}, {8, true},
|
||||
{8, true}, {8, true}, {8, true}, {11, true}, {11, true},
|
||||
{12, true}, {12, true}, {15, true}, {17, true}, {22, true}, {23, true},
|
||||
}
|
||||
)
|
||||
|
||||
// amlMaintained and amlNonmaintained are the acute myelogenous leukaemia
|
||||
// survival times in weeks from Miller (1997), shipped as the aml dataset in R's
|
||||
// survival package. Five subjects are censored, at 13, 16, 28, 45 and 161 weeks.
|
||||
//
|
||||
// maintained: 9, 13, 13+, 18, 23, 28+, 31, 34, 45+, 48, 161+
|
||||
// nonmaintained: 5, 5, 8, 8, 12, 16+, 23, 27, 30, 33, 43, 45
|
||||
var (
|
||||
amlMaintained = []observation{
|
||||
{9, true}, {13, true}, {13, false}, {18, true}, {23, true},
|
||||
{28, false}, {31, true}, {34, true}, {45, false}, {48, true}, {161, false},
|
||||
}
|
||||
amlNonmaintained = []observation{
|
||||
{5, true}, {5, true}, {8, true}, {8, true}, {12, true}, {16, false},
|
||||
{23, true}, {27, true}, {30, true}, {33, true}, {43, true}, {45, true},
|
||||
}
|
||||
)
|
||||
|
||||
// chorioamnionTerm and chorioamnionEarly are permeability constants of the human
|
||||
// chorioamnion at term and between 12 and 26 weeks gestational age, Hollander
|
||||
// and Wolfe (1973), 69f. R's wilcox.test help page uses exactly these vectors as
|
||||
// its two-sample example.
|
||||
var (
|
||||
chorioamnionTerm = []float64{0.80, 0.83, 1.89, 1.04, 1.45, 1.38, 1.91, 1.64, 0.73, 1.46}
|
||||
chorioamnionEarly = []float64{1.15, 0.88, 0.90, 0.74, 1.21}
|
||||
)
|
||||
@@ -0,0 +1,97 @@
|
||||
package main
|
||||
|
||||
import "math"
|
||||
|
||||
// gehanResult is one pairwise comparison of two arms of right-censored runs: an
|
||||
// effect size, the run pairs that have no order between them, and the test of
|
||||
// the same statistic against the null of equal hazards.
|
||||
type gehanResult struct {
|
||||
FirstSize int
|
||||
SecondSize int
|
||||
Statistic float64
|
||||
A12 float64
|
||||
Unordered int
|
||||
PValue float64
|
||||
}
|
||||
|
||||
// outlives orders two runs the only way right-censoring allows. A run censored
|
||||
// at step t violated at no step up to t and stopped for a reason of its own, so
|
||||
// it outlives a violation at or before t and nothing orders it against a
|
||||
// violation after t or against another censored run. Comparing the two step
|
||||
// counts as plain numbers instead reads a run the wall clock stopped at step 12
|
||||
// as one that violated at step 12.
|
||||
func outlives(left, right observation) int {
|
||||
switch {
|
||||
case left.Event && right.Event:
|
||||
switch {
|
||||
case left.Steps > right.Steps:
|
||||
return 1
|
||||
case left.Steps < right.Steps:
|
||||
return -1
|
||||
}
|
||||
case left.Event:
|
||||
if right.Steps >= left.Steps {
|
||||
return -1
|
||||
}
|
||||
case right.Event:
|
||||
if left.Steps >= right.Steps {
|
||||
return 1
|
||||
}
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// atRiskWeight is Gehan's weight: an event counts for as many runs as were still
|
||||
// at risk when it happened. It is what makes the weighted log-rank statistic the
|
||||
// same quantity as the pairwise count below, so the effect size and the p-value
|
||||
// are one statistic rather than two that can disagree.
|
||||
func atRiskWeight(atRisk float64) float64 { return atRisk }
|
||||
|
||||
// gehanTest is the Gehan-Breslow generalized Wilcoxon test: the rank-sum
|
||||
// carried over to right-censored samples by scoring every pair of runs by which
|
||||
// one outlived the other and leaving the pairs censoring cannot order out of the
|
||||
// count. Gehan (1965), "A Generalized Wilcoxon Test for Comparing Arbitrarily
|
||||
// Singly-Censored Samples", Biometrika 52(1-2), 203-223; Breslow (1970).
|
||||
//
|
||||
// Statistic is that count, U, and A12 is it over the number of pairs: the share
|
||||
// of run pairs in which the first arm survived longer, an unordered pair
|
||||
// counting as half. With nothing censored the two are exactly the Mann-Whitney U
|
||||
// and the Vargha-Delaney A12 the uncensored rank-sum reports. Where censoring
|
||||
// leaves a pair unordered, the half it contributes is the null value, so an
|
||||
// unordered pair can only pull the effect size toward 0.5 and can never
|
||||
// manufacture a direction.
|
||||
//
|
||||
// The p-value is the same statistic standardized: the weighted log-rank with
|
||||
// Gehan's weight has this U for its statistic, and its variance is the
|
||||
// conditional hypergeometric one summed over event times, which is what keeps
|
||||
// the test honest when the arms censor on different schedules. The permutation
|
||||
// variance Gehan originally paired with the statistic does not.
|
||||
func gehanTest(first, second []observation) gehanResult {
|
||||
result := gehanResult{
|
||||
FirstSize: len(first),
|
||||
SecondSize: len(second),
|
||||
Statistic: math.NaN(),
|
||||
A12: math.NaN(),
|
||||
PValue: math.NaN(),
|
||||
}
|
||||
if len(first) == 0 || len(second) == 0 {
|
||||
return result
|
||||
}
|
||||
outlived := 0.0
|
||||
for _, left := range first {
|
||||
for _, right := range second {
|
||||
switch outlives(left, right) {
|
||||
case 1:
|
||||
outlived++
|
||||
case 0:
|
||||
outlived += 0.5
|
||||
result.Unordered++
|
||||
}
|
||||
}
|
||||
}
|
||||
result.Statistic = outlived
|
||||
result.A12 = outlived / float64(len(first)*len(second))
|
||||
test := weightedLogRank([]string{"first", "second"}, [][]observation{first, second}, atRiskWeight)
|
||||
result.PValue = test.PValue
|
||||
return result
|
||||
}
|
||||
@@ -0,0 +1,256 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func tiedPairs(first, second []float64) int {
|
||||
tied := 0
|
||||
for _, left := range first {
|
||||
for _, right := range second {
|
||||
if left == right {
|
||||
tied++
|
||||
}
|
||||
}
|
||||
}
|
||||
return tied
|
||||
}
|
||||
|
||||
func events(values []float64) []observation {
|
||||
items := make([]observation, 0, len(values))
|
||||
for _, value := range values {
|
||||
items = append(items, observation{Steps: value, Event: true})
|
||||
}
|
||||
return items
|
||||
}
|
||||
|
||||
// Every ordering censoring supports and every ordering it does not.
|
||||
func TestOutlives_OrdersOnlyWhatTheCensoringSupports(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
left, right observation
|
||||
want int
|
||||
}{
|
||||
{"two violations", observation{30, true}, observation{10, true}, 1},
|
||||
{"two violations the other way", observation{10, true}, observation{30, true}, -1},
|
||||
{"two violations at the same step", observation{10, true}, observation{10, true}, 0},
|
||||
{"censored after the violation", observation{30, false}, observation{10, true}, 1},
|
||||
{"censored on the violation's own step", observation{10, false}, observation{10, true}, 1},
|
||||
{"censored before the violation", observation{10, false}, observation{30, true}, 0},
|
||||
{"violation before the censoring", observation{10, true}, observation{30, false}, -1},
|
||||
{"violation after the censoring", observation{30, true}, observation{10, false}, 0},
|
||||
{"both censored", observation{10, false}, observation{30, false}, 0},
|
||||
}
|
||||
for _, test := range cases {
|
||||
if got := outlives(test.left, test.right); got != test.want {
|
||||
t.Errorf("%s: %v against %v ordered %+d, want %+d", test.name, test.left, test.right, got, test.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// With nothing censored the test is the rank-sum, so its statistic and effect
|
||||
// size have to be the ones the rank-sum reports on the same numbers, ties
|
||||
// included.
|
||||
func TestGehanTest_ReducesToTheRankSumWhenNothingIsCensored(t *testing.T) {
|
||||
cases := [][2][]float64{
|
||||
{chorioamnionTerm, chorioamnionEarly},
|
||||
{{1, 2, 3, 4}, {3, 4, 5, 6}},
|
||||
{{40, 40, 40}, {40, 40, 40, 40}},
|
||||
{{5, 6, 7}, {1, 2}},
|
||||
{{3}, {9}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
result := gehanTest(events(test[0]), events(test[1]))
|
||||
reference := rankSum(test[0], test[1])
|
||||
if result.Statistic != reference.Statistic {
|
||||
t.Errorf("u %v over %v and %v, want the rank-sum's %v",
|
||||
result.Statistic, test[0], test[1], reference.Statistic)
|
||||
}
|
||||
if result.A12 != reference.A12 {
|
||||
t.Errorf("a12 %v over %v and %v, want the rank-sum's %v",
|
||||
result.A12, test[0], test[1], reference.A12)
|
||||
}
|
||||
if want := tiedPairs(test[0], test[1]); result.Unordered != want {
|
||||
t.Errorf("%d unordered pair(s) over %v and %v, want the %d tied ones and no others",
|
||||
result.Unordered, test[0], test[1], want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The failure the flattening produced: an arm the wall clock stopped at step 12
|
||||
// says nothing about step 100, so there is no difference to find and no
|
||||
// direction to report.
|
||||
func TestGehanTest_RunsStoppedBeforeEveryViolationOrderNothing(t *testing.T) {
|
||||
stopped := []observation{{12, false}, {12, false}, {12, false}, {12, false}}
|
||||
violated := []observation{{100, true}, {100, true}, {100, true}}
|
||||
result := gehanTest(stopped, violated)
|
||||
|
||||
if result.Unordered != 12 || result.A12 != 0.5 {
|
||||
t.Errorf("%d of 12 pairs unordered, a12 %v, want all of them and 0.5", result.Unordered, result.A12)
|
||||
}
|
||||
if result.PValue < 0.05 {
|
||||
t.Errorf("p %v, want no difference between arms never observed over the same steps", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// A censored run outliving a violation is evidence, and it is the only kind the
|
||||
// wall-clock case leaves: four runs still clean at step 12 against three
|
||||
// violations by step 5.
|
||||
func TestGehanTest_CensoringLeavesTheEvidenceItDoesSupport(t *testing.T) {
|
||||
stopped := []observation{{12, false}, {12, false}, {12, false}, {12, false}}
|
||||
violated := []observation{{5, true}, {5, true}, {5, true}}
|
||||
result := gehanTest(stopped, violated)
|
||||
|
||||
if result.Unordered != 0 || result.Statistic != 12 || result.A12 != 1 {
|
||||
t.Errorf("result %+v, want every pair ordered for the arm that had not violated", result)
|
||||
}
|
||||
if result.PValue > 0.05 {
|
||||
t.Errorf("p %v, want the arms to separate", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
func riskAndDeaths(group []observation, steps float64) (float64, float64) {
|
||||
atRisk, deaths := 0.0, 0.0
|
||||
for _, item := range group {
|
||||
if item.Steps >= steps {
|
||||
atRisk++
|
||||
}
|
||||
if item.Steps == steps && item.Event {
|
||||
deaths++
|
||||
}
|
||||
}
|
||||
return atRisk, deaths
|
||||
}
|
||||
|
||||
// gehanReference is Gehan's statistic and its conditional variance written
|
||||
// straight from the definitions,
|
||||
//
|
||||
// S = sum over event times of (Y2*d1 - Y1*d2)
|
||||
// V = sum over event times of d(Y-d)/(Y-1) * Y1*Y2
|
||||
//
|
||||
// which is an independent calculation rather than a second call into the code
|
||||
// under test.
|
||||
func gehanReference(first, second []observation) (float64, float64) {
|
||||
pooled := append(append([]observation{}, first...), second...)
|
||||
statistic, variance := 0.0, 0.0
|
||||
for _, steps := range distinctSteps(pooled) {
|
||||
firstAtRisk, firstDeaths := riskAndDeaths(first, steps)
|
||||
secondAtRisk, secondDeaths := riskAndDeaths(second, steps)
|
||||
deaths := firstDeaths + secondDeaths
|
||||
if deaths == 0 {
|
||||
continue
|
||||
}
|
||||
atRisk := firstAtRisk + secondAtRisk
|
||||
statistic += secondAtRisk*firstDeaths - firstAtRisk*secondDeaths
|
||||
if atRisk > 1 {
|
||||
variance += deaths * (atRisk - deaths) / (atRisk - 1) * firstAtRisk * secondAtRisk
|
||||
}
|
||||
}
|
||||
return statistic, variance
|
||||
}
|
||||
|
||||
// The effect size and the p-value have to be the same statistic seen twice, or
|
||||
// the report can carry a direction its p-value does not support. Counting run
|
||||
// pairs and accumulating over risk sets are two routes to Gehan's statistic, and
|
||||
// they are tied by S = mn - 2U.
|
||||
func TestGehanTest_PairCountAndRiskSetAgreeOnOneStatistic(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
first, second []observation
|
||||
}{
|
||||
{"6-mp against placebo", gehanSixMercaptopurine, gehanPlacebo},
|
||||
{"maintained against nonmaintained", amlMaintained, amlNonmaintained},
|
||||
{"stopped short against violating late", []observation{{12, false}, {12, false}, {14, false}},
|
||||
[]observation{{5, true}, {100, true}, {100, true}}},
|
||||
{"censoring tied with an event", []observation{{20, false}, {20, true}, {35, true}},
|
||||
[]observation{{20, true}, {20, false}, {9, true}}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
result := gehanTest(test.first, test.second)
|
||||
statistic, variance := gehanReference(test.first, test.second)
|
||||
pairs := float64(len(test.first) * len(test.second))
|
||||
|
||||
if got := pairs - 2*result.Statistic; math.Abs(got-statistic) > 1e-9 {
|
||||
t.Errorf("%s: pair count gives a statistic of %v, the risk sets give %v", test.name, got, statistic)
|
||||
}
|
||||
expected := chiSquareUpperTail(statistic*statistic/variance, 1)
|
||||
if math.Abs(result.PValue-expected) > 1e-12 {
|
||||
t.Errorf("%s: p %v, want %v from statistic %v over variance %v",
|
||||
test.name, result.PValue, expected, statistic, variance)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The 6-MP trial is the dataset the test is named for. Its log-rank result is
|
||||
// checked elsewhere against the published one; here the generalized Wilcoxon
|
||||
// has to reach the same conclusion, with the maintained arm outliving the
|
||||
// placebo arm on both routes.
|
||||
func TestGehanTest_SeparatesThePublishedLeukaemiaTrial(t *testing.T) {
|
||||
result := gehanTest(gehanSixMercaptopurine, gehanPlacebo)
|
||||
if result.A12 <= 0.5 {
|
||||
t.Errorf("a12 %v, want the 6-mp arm to outlive the placebo arm", result.A12)
|
||||
}
|
||||
if result.PValue > 0.001 {
|
||||
t.Errorf("p %v, want the arms to separate as the log-rank has them separate", result.PValue)
|
||||
}
|
||||
logRankResult := logRank([]string{"6-mp", "placebo"},
|
||||
[][]observation{gehanSixMercaptopurine, gehanPlacebo})
|
||||
if logRankResult.PValue > 0.001 {
|
||||
t.Fatalf("log-rank p %v: the comparison being made is not the one this test assumes", logRankResult.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
func censoredAt(count int, steps float64) []observation {
|
||||
items := make([]observation, 0, count)
|
||||
for index := 0; index < count; index++ {
|
||||
items = append(items, observation{Steps: steps})
|
||||
}
|
||||
return items
|
||||
}
|
||||
|
||||
// A specification whose violations are all obligations reported when the run
|
||||
// ends puts every event on one step, and the comparison collapses to a single
|
||||
// two-by-two table of violated against clean. The statistic there is the
|
||||
// Mantel-Haenszel chi-square of that table,
|
||||
//
|
||||
// (N-1)(ad-bc)^2 / ((a+b)(c+d)(a+c)(b+d))
|
||||
//
|
||||
// and the (Y-d)/(Y-1) term in the variance is what carries the (N-1)/N that
|
||||
// separates it from the Pearson chi-square. Nothing else in the pipeline
|
||||
// exercises that term hard, because tied events are otherwise rare.
|
||||
func TestGehanTest_EveryViolationOnOneStepIsTheMantelHaenszelTable(t *testing.T) {
|
||||
cases := [][4]int{{30, 20, 10, 40}, {25, 25, 15, 35}, {20, 20, 12, 28}}
|
||||
for _, test := range cases {
|
||||
firstEvents, firstCensored, secondEvents, secondCensored := test[0], test[1], test[2], test[3]
|
||||
first := append(events(repeated(firstEvents, 400)), censoredAt(firstCensored, 400)...)
|
||||
second := append(events(repeated(secondEvents, 400)), censoredAt(secondCensored, 400)...)
|
||||
|
||||
total := float64(firstEvents + firstCensored + secondEvents + secondCensored)
|
||||
crossProduct := float64(firstEvents*secondCensored - firstCensored*secondEvents)
|
||||
chiSquare := (total - 1) * crossProduct * crossProduct /
|
||||
(float64(firstEvents+firstCensored) * float64(secondEvents+secondCensored) *
|
||||
float64(firstEvents+secondEvents) * float64(firstCensored+secondCensored))
|
||||
|
||||
got := gehanTest(first, second).PValue
|
||||
want := chiSquareUpperTail(chiSquare, 1)
|
||||
if math.Abs(got-want) > 1e-12 {
|
||||
t.Errorf("%v: p %v, want the table's %v from chi-square %v", test, got, want, chiSquare)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func repeated(count int, value float64) []float64 {
|
||||
values := make([]float64, 0, count)
|
||||
for index := 0; index < count; index++ {
|
||||
values = append(values, value)
|
||||
}
|
||||
return values
|
||||
}
|
||||
|
||||
func TestGehanTest_EmptyArmHasNoComparison(t *testing.T) {
|
||||
result := gehanTest(nil, []observation{{5, true}})
|
||||
if !math.IsNaN(result.PValue) || !math.IsNaN(result.A12) || !math.IsNaN(result.Statistic) {
|
||||
t.Errorf("result %+v, want everything undefined", result)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,36 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"slices"
|
||||
)
|
||||
|
||||
// holm applies the Holm (1979) step-down correction within one family of
|
||||
// comparisons, enforcing monotonicity across the sorted p-values the way R's
|
||||
// p.adjust does. Holm, "A Simple Sequentially Rejective Multiple Test
|
||||
// Procedure", Scandinavian Journal of Statistics 6(2), 65-70.
|
||||
func holm(pValues []float64) []float64 {
|
||||
count := len(pValues)
|
||||
adjusted := make([]float64, count)
|
||||
order := make([]int, count)
|
||||
for index := range order {
|
||||
order[index] = index
|
||||
}
|
||||
slices.SortStableFunc(order, func(left, right int) int {
|
||||
switch {
|
||||
case pValues[left] < pValues[right]:
|
||||
return -1
|
||||
case pValues[left] > pValues[right]:
|
||||
return 1
|
||||
default:
|
||||
return 0
|
||||
}
|
||||
})
|
||||
running := 0.0
|
||||
for position, index := range order {
|
||||
scaled := float64(count-position) * pValues[index]
|
||||
running = math.Max(running, scaled)
|
||||
adjusted[index] = math.Min(running, 1)
|
||||
}
|
||||
return adjusted
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Both cases are printed R output in Eve Slavich, "Four strategies for dealing
|
||||
// with multiple comparisons", UNSW Stats Central, slides 9 and 10:
|
||||
//
|
||||
// pValues = c(0.01, 0.2, 0.08, 0.03)
|
||||
// p.adjust(pValues, method = "holm")
|
||||
// ## [1] 0.04 0.20 0.16 0.09
|
||||
//
|
||||
// pValues = c(0.01, 0.2, 0.08, 0.03, 0.02, 0.01)
|
||||
// p.adjust(pValues, method = "holm")
|
||||
// ## [1] 0.06 0.20 0.16 0.09 0.08 0.06
|
||||
//
|
||||
// The second case exercises the monotonicity step: sorted p are
|
||||
// .01 .01 .02 .03 .08 .20, scaled by 6 5 4 3 2 1 to .06 .05 .08 .09 .16 .20,
|
||||
// and the running maximum lifts the second back to .06.
|
||||
func TestHolm_MatchesPublishedAdjustment(t *testing.T) {
|
||||
cases := []struct {
|
||||
raw []float64
|
||||
expected []float64
|
||||
}{
|
||||
{[]float64{0.01, 0.2, 0.08, 0.03}, []float64{0.04, 0.20, 0.16, 0.09}},
|
||||
{[]float64{0.01, 0.2, 0.08, 0.03, 0.02, 0.01}, []float64{0.06, 0.20, 0.16, 0.09, 0.08, 0.06}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
adjusted := holm(test.raw)
|
||||
for index, want := range test.expected {
|
||||
if math.Abs(adjusted[index]-want) > 1e-12 {
|
||||
t.Errorf("holm(%v)[%d] = %v, want %v", test.raw, index, adjusted[index], want)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestHolm_CapsAtOneAndKeepsOrder(t *testing.T) {
|
||||
adjusted := holm([]float64{0.4, 0.5, 0.9})
|
||||
for index, value := range adjusted {
|
||||
if value != 1 {
|
||||
t.Errorf("adjusted[%d] = %v, want 1", index, value)
|
||||
}
|
||||
}
|
||||
single := holm([]float64{0.03})
|
||||
if len(single) != 1 || single[0] != 0.03 {
|
||||
t.Errorf("single comparison adjusted to %v, want 0.03 unchanged", single)
|
||||
}
|
||||
if got := holm(nil); len(got) != 0 {
|
||||
t.Errorf("holm(nil) = %v, want empty", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,155 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// A campaign that produced nothing still has a manifest, and the manifest is
|
||||
// what makes the difference between an arm that ran nothing and an arm that was
|
||||
// never scheduled. Reading one must report every intended seed as missing rather
|
||||
// than an arm with a small clean sample.
|
||||
func TestRun_EmptyCampaignIsReportedAsEverySeedMissing(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
empty := filepath.Join(root, "empty")
|
||||
writeCampaign(t, empty, map[string]any{
|
||||
"arm": "seeded", "max_steps": 400, "seeds": []int{1, 2, 3, 4, 5},
|
||||
}, nil)
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{empty}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
result := analyseCampaigns(t, empty)
|
||||
summary := armByName(t, result, "seeded")
|
||||
if summary.Recorded != 0 || summary.Usable != 0 {
|
||||
t.Errorf("recorded %d usable %d, want none", summary.Recorded, summary.Usable)
|
||||
}
|
||||
if len(summary.MissingSeeds) != 5 {
|
||||
t.Errorf("missing seeds %v, want all five", summary.MissingSeeds)
|
||||
}
|
||||
if summary.MedianStepsToFirstViolation != nil || summary.ViolationRate != nil {
|
||||
t.Errorf("median %v violation rate %v, want neither from no runs",
|
||||
summary.MedianStepsToFirstViolation, summary.ViolationRate)
|
||||
}
|
||||
if result.LogRank != nil || len(result.Pairwise) != 0 {
|
||||
t.Errorf("log-rank %+v pairwise %v, want no tests", result.LogRank, result.Pairwise)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "seeded") {
|
||||
t.Errorf("the empty arm is not in the report\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
// A host that stopped part way through leaves a directory that looks complete.
|
||||
// The seeds it never reached are the difference between a partial campaign and
|
||||
// a smaller one, and the report has to carry that count.
|
||||
func TestRun_PartialCampaignCountsTheSeedsTheHostNeverReached(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
partial := filepath.Join(root, "partial")
|
||||
var records []map[string]any
|
||||
for seed := 1; seed <= 6; seed++ {
|
||||
records = append(records, map[string]any{
|
||||
"seed": seed, "exit_code": 0, "steps": 400, "actions": 380, "monotonic_millis": 360000,
|
||||
})
|
||||
}
|
||||
writeCampaign(t, partial, map[string]any{
|
||||
"arm": "seeded", "max_steps": 400,
|
||||
"seeds": []int{1, 2, 3, 4, 5, 6, 7, 8, 9, 10},
|
||||
}, records)
|
||||
|
||||
summary := armByName(t, analyseCampaigns(t, partial), "seeded")
|
||||
if summary.Usable != 6 {
|
||||
t.Errorf("%d usable runs, want the 6 that landed", summary.Usable)
|
||||
}
|
||||
if len(summary.MissingSeeds) != 4 {
|
||||
t.Errorf("missing seeds %v, want the 4 the host never reached", summary.MissingSeeds)
|
||||
}
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{partial}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "missing") {
|
||||
t.Errorf("no missing-seed column in the report\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
// The ablation is seed-matched, so a run lost on one arm removes its partner
|
||||
// from the comparison too. Those seeds are named because a paired sample that
|
||||
// silently shrinks is how a campaign reports a difference between two arms that
|
||||
// were not in fact matched.
|
||||
func TestRun_PairedComparisonNamesSeedsLostOnOneArm(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
declared := []int{1, 2, 3, 4, 5, 6}
|
||||
pre := filepath.Join(root, "pre")
|
||||
post := filepath.Join(root, "post")
|
||||
writeCampaign(t, pre, map[string]any{"arm": "pre", "max_steps": 400, "seeds": declared}, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 300, "actions": 290, "monotonic_millis": 1000, "first_violation_origin_step": 300, "violated_properties": []string{"p"}},
|
||||
{"seed": 2, "exit_code": 0, "steps": 400, "actions": 390, "monotonic_millis": 1000},
|
||||
{"seed": 3, "exit_code": 0, "steps": 250, "actions": 240, "monotonic_millis": 1000, "first_violation_origin_step": 250, "violated_properties": []string{"p"}},
|
||||
{"seed": 4, "exit_code": -1, "timed_out": true, "actions": 0},
|
||||
})
|
||||
writeCampaign(t, post, map[string]any{"arm": "post", "max_steps": 400, "seeds": declared}, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 38, "monotonic_millis": 1000, "first_violation_origin_step": 40, "violated_properties": []string{"p"}},
|
||||
{"seed": 2, "exit_code": 0, "steps": 60, "actions": 55, "monotonic_millis": 1000, "first_violation_origin_step": 60, "violated_properties": []string{"p"}},
|
||||
{"seed": 3, "exit_code": 0, "steps": 50, "actions": 47, "monotonic_millis": 1000, "first_violation_origin_step": 50, "violated_properties": []string{"p"}},
|
||||
{"seed": 4, "exit_code": 0, "steps": 70, "actions": 66, "monotonic_millis": 1000, "first_violation_origin_step": 70, "violated_properties": []string{"p"}},
|
||||
})
|
||||
|
||||
result := analyseCampaigns(t, "--paired", pre, post)
|
||||
if result.Paired == nil {
|
||||
t.Fatal("no paired comparison")
|
||||
}
|
||||
if result.Paired.Pairs != 3 {
|
||||
t.Errorf("%d pairs, want the 3 seeds usable on both arms", result.Paired.Pairs)
|
||||
}
|
||||
if len(result.Paired.UnpairedSeeds) != 1 || result.Paired.UnpairedSeeds[0] != 4 {
|
||||
t.Errorf("unpaired seeds %v, want [4]", result.Paired.UnpairedSeeds)
|
||||
}
|
||||
for _, summary := range result.Arms {
|
||||
if len(summary.MissingSeeds) != 2 {
|
||||
t.Errorf("arm %s missing seeds %v, want seeds 5 and 6", summary.Arm, summary.MissingSeeds)
|
||||
}
|
||||
}
|
||||
|
||||
var stdout bytes.Buffer
|
||||
if err := run([]string{"--paired", pre, post}, &stdout, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "usable in one arm only") {
|
||||
t.Errorf("the report does not name the lost seed\n%s", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_PairedRefusesAnythingOtherThanTwoArms(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
var directories []string
|
||||
for _, name := range []string{"a", "b", "c"} {
|
||||
directory := filepath.Join(root, name)
|
||||
writeCampaign(t, directory, map[string]any{"arm": name, "max_steps": 40, "seeds": []int{1}},
|
||||
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
|
||||
directories = append(directories, directory)
|
||||
}
|
||||
err := run(append([]string{"--paired"}, directories...), io.Discard, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), "exactly two arms") {
|
||||
t.Fatalf("error %v, want a refusal to pair three arms", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_PairedRefusesArmsThatShareNoSeed(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
north := filepath.Join(root, "north")
|
||||
south := filepath.Join(root, "south")
|
||||
writeCampaign(t, north, map[string]any{"arm": "north", "max_steps": 40, "seeds": []int{1, 2}},
|
||||
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
|
||||
writeCampaign(t, south, map[string]any{"arm": "south", "max_steps": 40, "seeds": []int{3, 4}},
|
||||
[]map[string]any{{"seed": 3, "exit_code": 0, "steps": 40, "actions": 40}})
|
||||
|
||||
err := run([]string{"--paired", north, south}, io.Discard, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), "share no seed") {
|
||||
t.Fatalf("error %v, want a refusal to pair arms that ran different seeds", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,330 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
const (
|
||||
manifestFileName = "campaign.json"
|
||||
recordsFileName = "runs.jsonl"
|
||||
maxRecordBytes = 4 * 1024 * 1024
|
||||
)
|
||||
|
||||
// manifest mirrors the fields analyze reads from campaign.json. The step budget
|
||||
// lives here rather than in any run, because it is the exposure every run in an
|
||||
// arm was given and the ceiling a clean run can be censored at.
|
||||
type manifest struct {
|
||||
Arm string `json:"arm"`
|
||||
Generator string `json:"generator"`
|
||||
Platform string `json:"platform"`
|
||||
MaxSteps int `json:"max_steps"`
|
||||
Seeds []int64 `json:"seeds"`
|
||||
Host string `json:"host"`
|
||||
}
|
||||
|
||||
// runRecord mirrors the fields analyze reads from one line of runs.jsonl.
|
||||
type runRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
ExitCode int `json:"exit_code"`
|
||||
LaunchError string `json:"launch_error"`
|
||||
TimedOut bool `json:"timed_out"`
|
||||
// MonotonicMillis is how long the run worked, and it is what every
|
||||
// per-hour rate here divides by: a host asleep mid-run tested nothing, so
|
||||
// charging that time to the arm would report it as slower for a reason
|
||||
// that has nothing to do with the arm. The wall clock the campaign also
|
||||
// records answers the other question, how much time passed.
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
// DurationMillis is the name campaigns written before the two clocks were
|
||||
// split gave the same monotonic reading, so those files still read.
|
||||
DurationMillis int64 `json:"duration_millis"`
|
||||
TraceError string `json:"trace_error"`
|
||||
Steps int `json:"steps"`
|
||||
FirstViolationOriginStep *int `json:"first_violation_origin_step"`
|
||||
// FirstViolationDetectedStep is the step the violation was reported on,
|
||||
// which is the origin step for a safety property tripping under its own
|
||||
// action and the end of the budget for an obligation that never discharged.
|
||||
// It is what the survival analysis times the event by; see eventStep.
|
||||
FirstViolationDetectedStep *int `json:"first_violation_detected_step"`
|
||||
ViolatedProperties []string `json:"violated_properties"`
|
||||
// Actions is the count of steps on which the action generator dispatched an
|
||||
// action, which excludes the steps the spec's setup drove before the
|
||||
// generator was consulted wherever the trace names them. It is a pointer so
|
||||
// that a runs.jsonl written before the campaign tool counted them is refused
|
||||
// rather than read as an arm that acted zero times. The campaign tool always
|
||||
// emits the field, so its absence dates the file.
|
||||
Actions *int `json:"actions"`
|
||||
// UnattributedActions is how many of the run's actions name no producer, so
|
||||
// nothing can say whether the spec's setup drove them. It is a pointer
|
||||
// because a runs.jsonl written before actions named a producer at all has no
|
||||
// field, and reading that silence as none would let a denominator of unknown
|
||||
// provenance pass for one that excludes the setup's login.
|
||||
UnattributedActions *int `json:"unattributed_actions"`
|
||||
}
|
||||
|
||||
// Exclusion reasons. A run that failed or timed out is missing data, not a
|
||||
// censored observation: it broke off, so its step count is not exposure the app
|
||||
// survived and counting it as one would bias the survival estimate downward.
|
||||
const (
|
||||
reasonLaunchError = "launch error"
|
||||
reasonTimedOut = "timed out"
|
||||
reasonNonzeroExit = "nonzero exit"
|
||||
reasonTraceError = "unreadable trace"
|
||||
reasonMalformedStep = "violation step outside the budget"
|
||||
)
|
||||
|
||||
type classifiedRun struct {
|
||||
Seed int64
|
||||
Steps int
|
||||
Actions int
|
||||
UnattributedActions int
|
||||
MonotonicMillis int64
|
||||
OriginStep int
|
||||
// EventStep is when the run could know, and it is what the survival
|
||||
// analysis measures. It is the origin step whenever the two agree.
|
||||
EventStep int
|
||||
Violated bool
|
||||
ClampedToBudget bool
|
||||
ViolatedProperties []string
|
||||
ExcludedBecause string
|
||||
}
|
||||
|
||||
type arm struct {
|
||||
Name string
|
||||
Budget int
|
||||
Generator string
|
||||
Platform string
|
||||
Directories []string
|
||||
Runs []classifiedRun
|
||||
MissingSeeds []int64
|
||||
}
|
||||
|
||||
func loadCampaign(directory string) (manifest, []runRecord, error) {
|
||||
body, err := os.ReadFile(filepath.Join(directory, manifestFileName))
|
||||
if err != nil {
|
||||
return manifest{}, nil, fmt.Errorf("read %s: %w", manifestFileName, err)
|
||||
}
|
||||
var declared manifest
|
||||
if err := json.Unmarshal(body, &declared); err != nil {
|
||||
return manifest{}, nil, fmt.Errorf("parse %s in %s: %w", manifestFileName, directory, err)
|
||||
}
|
||||
if declared.Arm == "" {
|
||||
return manifest{}, nil, fmt.Errorf("%s in %s has no arm", manifestFileName, directory)
|
||||
}
|
||||
if declared.MaxSteps <= 0 {
|
||||
return manifest{}, nil, fmt.Errorf("%s in %s has max_steps %d: clean runs have nothing to be censored at",
|
||||
manifestFileName, directory, declared.MaxSteps)
|
||||
}
|
||||
|
||||
file, err := os.Open(filepath.Join(directory, recordsFileName))
|
||||
if err != nil {
|
||||
return manifest{}, nil, fmt.Errorf("read %s: %w", recordsFileName, err)
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
var records []runRecord
|
||||
scanner := bufio.NewScanner(file)
|
||||
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
|
||||
lineNumber := 0
|
||||
for scanner.Scan() {
|
||||
lineNumber++
|
||||
raw := strings.TrimSpace(scanner.Text())
|
||||
if raw == "" {
|
||||
continue
|
||||
}
|
||||
var record runRecord
|
||||
if err := json.Unmarshal([]byte(raw), &record); err != nil {
|
||||
return manifest{}, nil, fmt.Errorf("%s line %d in %s: %w", recordsFileName, lineNumber, directory, err)
|
||||
}
|
||||
if record.Actions == nil {
|
||||
return manifest{}, nil, fmt.Errorf("%s line %d in %s has no actions count: it was written before "+
|
||||
"dispatched actions were counted, and reading the missing count as zero would report every "+
|
||||
"per-action rate wrongly; re-run the campaign to produce it",
|
||||
recordsFileName, lineNumber, directory)
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
if err := scanner.Err(); err != nil {
|
||||
return manifest{}, nil, fmt.Errorf("read %s in %s: %w", recordsFileName, directory, err)
|
||||
}
|
||||
return declared, records, nil
|
||||
}
|
||||
|
||||
// unattributedActions is how many of the record's actions carry no producer. A
|
||||
// campaign written before the field existed carries none at all, so its whole
|
||||
// action count is of unknown provenance rather than none of it.
|
||||
func unattributedActions(record runRecord) int {
|
||||
if record.UnattributedActions != nil {
|
||||
return *record.UnattributedActions
|
||||
}
|
||||
if record.Actions == nil {
|
||||
return 0
|
||||
}
|
||||
return *record.Actions
|
||||
}
|
||||
|
||||
func (r runRecord) workingMillis() int64 {
|
||||
if r.MonotonicMillis != 0 {
|
||||
return r.MonotonicMillis
|
||||
}
|
||||
return r.DurationMillis
|
||||
}
|
||||
|
||||
// classify turns one record into the run the analysis works with, deciding
|
||||
// whether it is usable and, if it is, whether it is an event or censored.
|
||||
func classify(record runRecord, budget int) classifiedRun {
|
||||
item := classifiedRun{
|
||||
Seed: record.Seed,
|
||||
Steps: record.Steps,
|
||||
MonotonicMillis: record.workingMillis(),
|
||||
ViolatedProperties: slices.Clone(record.ViolatedProperties),
|
||||
}
|
||||
if record.Actions != nil {
|
||||
item.Actions = *record.Actions
|
||||
}
|
||||
item.UnattributedActions = unattributedActions(record)
|
||||
switch {
|
||||
case record.LaunchError != "":
|
||||
item.ExcludedBecause = reasonLaunchError
|
||||
return item
|
||||
case record.TimedOut:
|
||||
item.ExcludedBecause = reasonTimedOut
|
||||
return item
|
||||
case record.ExitCode != 0:
|
||||
item.ExcludedBecause = reasonNonzeroExit
|
||||
return item
|
||||
case record.TraceError != "":
|
||||
item.ExcludedBecause = reasonTraceError
|
||||
return item
|
||||
}
|
||||
if record.FirstViolationOriginStep == nil {
|
||||
if len(record.ViolatedProperties) > 0 {
|
||||
item.ExcludedBecause = reasonMalformedStep
|
||||
}
|
||||
return item
|
||||
}
|
||||
origin := *record.FirstViolationOriginStep
|
||||
if origin < 1 {
|
||||
item.ExcludedBecause = reasonMalformedStep
|
||||
return item
|
||||
}
|
||||
item.Violated = true
|
||||
item.OriginStep = origin
|
||||
item.EventStep = eventStep(record, origin)
|
||||
if item.EventStep > budget {
|
||||
// The run-end finalize line reports obligations that never discharged
|
||||
// at an index one past the last executed step. That is a real detection
|
||||
// but not a real step, so it is held at the budget and counted.
|
||||
item.EventStep = budget
|
||||
item.ClampedToBudget = true
|
||||
}
|
||||
return item
|
||||
}
|
||||
|
||||
// eventStep is the step at which the run could know it had violated. A safety
|
||||
// property tripping under its own action is detected on the step that armed it
|
||||
// and the two agree. An obligation that never discharges is reported when the
|
||||
// run ends, and timing that at the step that armed it would record a liveness
|
||||
// failure flushed at the budget as a violation found on the first step, which
|
||||
// is a number the run cannot support and which no censored run can be compared
|
||||
// against: the end of the run is the clock the clean runs are censored on, so
|
||||
// the events have to be on it too. A campaign written before the field existed carries no
|
||||
// detected step and keeps the origin.
|
||||
func eventStep(record runRecord, origin int) int {
|
||||
if record.FirstViolationDetectedStep == nil {
|
||||
return origin
|
||||
}
|
||||
if detected := *record.FirstViolationDetectedStep; detected > origin {
|
||||
return detected
|
||||
}
|
||||
return origin
|
||||
}
|
||||
|
||||
// groupArms folds every campaign directory into its arm. Two directories with
|
||||
// the same arm label are pooled, which is how a campaign split across hosts is
|
||||
// analysed, but they must agree on the step budget.
|
||||
func groupArms(directories []string) ([]arm, error) {
|
||||
byName := map[string]*arm{}
|
||||
var order []string
|
||||
for _, directory := range directories {
|
||||
declared, records, err := loadCampaign(directory)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
current, seen := byName[declared.Arm]
|
||||
if !seen {
|
||||
current = &arm{
|
||||
Name: declared.Arm,
|
||||
Budget: declared.MaxSteps,
|
||||
Generator: declared.Generator,
|
||||
Platform: declared.Platform,
|
||||
}
|
||||
byName[declared.Arm] = current
|
||||
order = append(order, declared.Arm)
|
||||
}
|
||||
if current.Budget != declared.MaxSteps {
|
||||
return nil, fmt.Errorf("arm %q has step budget %d in an earlier campaign and %d in %s: "+
|
||||
"runs censored at different budgets cannot be pooled",
|
||||
declared.Arm, current.Budget, declared.MaxSteps, directory)
|
||||
}
|
||||
current.Directories = append(current.Directories, directory)
|
||||
|
||||
present := map[int64]bool{}
|
||||
for _, record := range records {
|
||||
present[record.Seed] = true
|
||||
current.Runs = append(current.Runs, classify(record, declared.MaxSteps))
|
||||
}
|
||||
for _, seed := range declared.Seeds {
|
||||
if !present[seed] {
|
||||
current.MissingSeeds = append(current.MissingSeeds, seed)
|
||||
}
|
||||
}
|
||||
}
|
||||
slices.Sort(order)
|
||||
arms := make([]arm, 0, len(order))
|
||||
for _, name := range order {
|
||||
arms = append(arms, *byName[name])
|
||||
}
|
||||
return arms, nil
|
||||
}
|
||||
|
||||
// unattributedActions is how much of the arm's per-action denominator has no
|
||||
// producer behind it, counted over the runs the denominator is built from.
|
||||
func (a arm) unattributedActions() int {
|
||||
total := 0
|
||||
for _, item := range a.Runs {
|
||||
if item.ExcludedBecause != "" {
|
||||
continue
|
||||
}
|
||||
total += item.UnattributedActions
|
||||
}
|
||||
return total
|
||||
}
|
||||
|
||||
// observations returns the usable runs as survival data: an event at the step
|
||||
// that armed the first violation, or a censored observation at the last step
|
||||
// the run reached.
|
||||
func (a arm) observations() []observation {
|
||||
var result []observation
|
||||
for _, item := range a.Runs {
|
||||
if item.ExcludedBecause != "" {
|
||||
continue
|
||||
}
|
||||
result = append(result, observationOf(item, a.Budget))
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func observationOf(item classifiedRun, budget int) observation {
|
||||
if item.Violated {
|
||||
return observation{Steps: float64(item.EventStep), Event: true}
|
||||
}
|
||||
// A run that hit the campaign's wall clock stopped short of the budget, and
|
||||
// the steps it never ran are not exposure it survived.
|
||||
return observation{Steps: float64(min(item.Steps, budget)), Event: false}
|
||||
}
|
||||
@@ -0,0 +1,218 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func stepPointer(value int) *int { return &value }
|
||||
|
||||
func TestClassify_FailedAndTimedOutRunsAreMissingDataNotCensored(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
record runRecord
|
||||
reason string
|
||||
}{
|
||||
{"launch error", runRecord{LaunchError: "fork/exec: no such file"}, reasonLaunchError},
|
||||
{"timed out", runRecord{TimedOut: true, ExitCode: -1}, reasonTimedOut},
|
||||
{"nonzero exit", runRecord{ExitCode: 3}, reasonNonzeroExit},
|
||||
{"unreadable trace", runRecord{TraceError: "no run directory with meta.json"}, reasonTraceError},
|
||||
{"violation at step zero", runRecord{FirstViolationOriginStep: stepPointer(0)}, reasonMalformedStep},
|
||||
{"violation without a step", runRecord{ViolatedProperties: []string{"cartTotal"}}, reasonMalformedStep},
|
||||
}
|
||||
for _, test := range cases {
|
||||
item := classify(test.record, 50)
|
||||
if item.ExcludedBecause != test.reason {
|
||||
t.Errorf("%s: excluded because %q, want %q", test.name, item.ExcludedBecause, test.reason)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A run also ends when the campaign's wall clock does, so a clean run can stop
|
||||
// well short of the budget. Censoring it at the budget would credit it with
|
||||
// steps it never ran, and a slower arm loses fewer steps in the same wall clock
|
||||
// than a fast one, so the credit does not cancel between arms.
|
||||
func TestClassify_CleanRunIsCensoredAtTheStepsItRan(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
steps int
|
||||
budget int
|
||||
censored float64
|
||||
}{
|
||||
{"stopped by the wall clock short of the budget", 12, 400, 12},
|
||||
{"ran the whole budget", 400, 400, 400},
|
||||
{"recorded more steps than the manifest budget", 420, 400, 400},
|
||||
}
|
||||
for _, test := range cases {
|
||||
item := classify(runRecord{Seed: 4, Steps: test.steps, DurationMillis: 1000}, test.budget)
|
||||
if item.ExcludedBecause != "" || item.Violated {
|
||||
t.Fatalf("%s: run %+v, want a usable clean run", test.name, item)
|
||||
}
|
||||
current := arm{Budget: test.budget, Runs: []classifiedRun{item}}
|
||||
observations := current.observations()
|
||||
if len(observations) != 1 || observations[0].Event || observations[0].Steps != test.censored {
|
||||
t.Errorf("%s: observations %+v, want one censored observation at %v",
|
||||
test.name, observations, test.censored)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestClassify_ViolationIsAnEventAtTheOriginStep(t *testing.T) {
|
||||
item := classify(runRecord{Seed: 5, Steps: 12, FirstViolationOriginStep: stepPointer(7)}, 50)
|
||||
if !item.Violated || item.OriginStep != 7 || item.ClampedToBudget {
|
||||
t.Fatalf("run %+v, want an unclamped event at step 7", item)
|
||||
}
|
||||
current := arm{Budget: 50, Runs: []classifiedRun{item}}
|
||||
observations := current.observations()
|
||||
if len(observations) != 1 || !observations[0].Event || observations[0].Steps != 7 {
|
||||
t.Errorf("observations %+v, want one event at 7", observations)
|
||||
}
|
||||
}
|
||||
|
||||
// The run-end finalize line reports at an index one past the last executed step,
|
||||
// so an origin past the budget is held at the budget and counted rather than
|
||||
// silently turned into a censored run.
|
||||
func TestClassify_ViolationPastTheBudgetIsHeldAtTheBudget(t *testing.T) {
|
||||
item := classify(runRecord{FirstViolationOriginStep: stepPointer(51)}, 50)
|
||||
if !item.Violated || item.EventStep != 50 || !item.ClampedToBudget {
|
||||
t.Errorf("run %+v, want a clamped event at 50", item)
|
||||
}
|
||||
}
|
||||
|
||||
func writeCampaign(t *testing.T, directory string, declared map[string]any, records []map[string]any) {
|
||||
t.Helper()
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body, err := json.MarshalIndent(declared, "", " ")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(directory, manifestFileName), append(body, '\n'), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var lines strings.Builder
|
||||
for _, record := range records {
|
||||
line, err := json.Marshal(record)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
lines.Write(line)
|
||||
lines.WriteByte('\n')
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(directory, recordsFileName), []byte(lines.String()), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestClassify_CarriesTheDispatchedActionCount(t *testing.T) {
|
||||
actions := 7
|
||||
item := classify(runRecord{Seed: 4, Steps: 30, Actions: &actions}, 50)
|
||||
if item.Steps != 30 || item.Actions != 7 {
|
||||
t.Errorf("run %+v, want 30 steps and 7 actions", item)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGroupArms_PoolsDirectoriesSharingAnArmAndReportsMissingSeeds(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeCampaign(t, filepath.Join(root, "north"), map[string]any{
|
||||
"arm": "seeded", "max_steps": 40, "seeds": []int{1, 2, 3},
|
||||
}, []map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 33},
|
||||
{"seed": 2, "exit_code": 0, "steps": 9, "actions": 8, "first_violation_origin_step": 9},
|
||||
})
|
||||
writeCampaign(t, filepath.Join(root, "south"), map[string]any{
|
||||
"arm": "seeded", "max_steps": 40, "seeds": []int{4},
|
||||
}, []map[string]any{
|
||||
{"seed": 4, "exit_code": 0, "steps": 40, "actions": 40},
|
||||
})
|
||||
|
||||
arms, err := groupArms([]string{filepath.Join(root, "north"), filepath.Join(root, "south")})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(arms) != 1 {
|
||||
t.Fatalf("%d arms, want 1", len(arms))
|
||||
}
|
||||
if len(arms[0].Runs) != 3 {
|
||||
t.Errorf("%d runs, want 3", len(arms[0].Runs))
|
||||
}
|
||||
if len(arms[0].MissingSeeds) != 1 || arms[0].MissingSeeds[0] != 3 {
|
||||
t.Errorf("missing seeds %v, want [3]", arms[0].MissingSeeds)
|
||||
}
|
||||
if len(arms[0].Directories) != 2 {
|
||||
t.Errorf("directories %v, want both", arms[0].Directories)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGroupArms_RejectsDisagreeingStepBudgets(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeCampaign(t, filepath.Join(root, "a"), map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}},
|
||||
[]map[string]any{{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40}})
|
||||
writeCampaign(t, filepath.Join(root, "b"), map[string]any{"arm": "seeded", "max_steps": 80, "seeds": []int{2}},
|
||||
[]map[string]any{{"seed": 2, "exit_code": 0, "steps": 80, "actions": 80}})
|
||||
|
||||
_, err := groupArms([]string{filepath.Join(root, "a"), filepath.Join(root, "b")})
|
||||
if err == nil || !strings.Contains(err.Error(), "different budgets") {
|
||||
t.Fatalf("error %v, want a refusal to pool different budgets", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGroupArms_RejectsAMissingStepBudget(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeCampaign(t, filepath.Join(root, "a"), map[string]any{"arm": "seeded", "seeds": []int{1}},
|
||||
[]map[string]any{{"seed": 1, "exit_code": 0}})
|
||||
_, err := groupArms([]string{filepath.Join(root, "a")})
|
||||
if err == nil || !strings.Contains(err.Error(), "censored at") {
|
||||
t.Fatalf("error %v, want a complaint about max_steps", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGroupArms_ReportsBadRecordLines(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "a")
|
||||
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}}, nil)
|
||||
if err := os.WriteFile(filepath.Join(directory, recordsFileName),
|
||||
[]byte("{\"seed\":1,\"actions\":0}\nnot json\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
_, err := groupArms([]string{directory})
|
||||
if err == nil || !strings.Contains(err.Error(), "line 2") {
|
||||
t.Fatalf("error %v, want the offending line number", err)
|
||||
}
|
||||
}
|
||||
|
||||
// A runs.jsonl written before the campaign counted dispatched actions has no
|
||||
// such field. Reading the absence as zero would divide by zero, so the whole
|
||||
// campaign is refused instead.
|
||||
func TestGroupArms_RefusesRecordsWithoutADispatchedActionCount(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "old-format")
|
||||
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1, 2}},
|
||||
[]map[string]any{
|
||||
{"seed": 1, "exit_code": 0, "steps": 40, "actions": 40},
|
||||
{"seed": 2, "exit_code": 0, "steps": 40},
|
||||
})
|
||||
_, err := groupArms([]string{directory})
|
||||
if err == nil {
|
||||
t.Fatal("a runs.jsonl without an action count was accepted")
|
||||
}
|
||||
for _, fragment := range []string{"line 2", "actions", directory} {
|
||||
if !strings.Contains(err.Error(), fragment) {
|
||||
t.Errorf("error %q is missing %q", err, fragment)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestGroupArms_RefusesAnExcludedRecordWithoutAnActionCount(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
directory := filepath.Join(root, "old-format")
|
||||
writeCampaign(t, directory, map[string]any{"arm": "seeded", "max_steps": 40, "seeds": []int{1}},
|
||||
[]map[string]any{{"seed": 1, "exit_code": 3}})
|
||||
if _, err := groupArms([]string{directory}); err == nil {
|
||||
t.Fatal("an old-format record was accepted because the run was excluded anyway")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,121 @@
|
||||
// Command analyze reduces campaign directories to the statistics the
|
||||
// evaluation reports. The primary outcome is steps to first violation, with
|
||||
// clean runs right-censored where they stopped rather than discarded: defect
|
||||
// yield per run is a binary that would need on the order of eighty runs an arm
|
||||
// to separate, while survival analysis uses every run, including the clean ones.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
const usage = `analyze reports the statistics of a sanderling evaluation from campaign directories.
|
||||
|
||||
Usage:
|
||||
analyze [--json <path>] <campaign-dir> [<campaign-dir> ...]
|
||||
|
||||
Each directory is one produced by the campaign tool and must hold campaign.json
|
||||
and runs.jsonl. Directories sharing an arm label are pooled and must agree on
|
||||
the step budget, and arms compared against each other must agree on it too.
|
||||
|
||||
One invocation is one research question: Holm corrects across the comparisons it
|
||||
produces and across nothing else.
|
||||
`
|
||||
|
||||
type stringList []string
|
||||
|
||||
func (list *stringList) String() string { return strings.Join(*list, ",") }
|
||||
|
||||
func (list *stringList) Set(value string) error {
|
||||
if strings.TrimSpace(value) == "" {
|
||||
return errors.New("empty campaign directory")
|
||||
}
|
||||
*list = append(*list, value)
|
||||
return nil
|
||||
}
|
||||
|
||||
func run(arguments []string, stdout, stderr io.Writer) error {
|
||||
flagSet := flag.NewFlagSet("analyze", flag.ContinueOnError)
|
||||
flagSet.SetOutput(stderr)
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprint(stderr, usage)
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
var directories stringList
|
||||
var jsonPath string
|
||||
var question string
|
||||
var paired bool
|
||||
flagSet.Var(&directories, "campaign", "campaign directory to read; repeat for more, or pass them as arguments")
|
||||
flagSet.StringVar(&jsonPath, "json", "", "write the machine-readable summary here, or - for stdout")
|
||||
flagSet.StringVar(&question, "question", "", "the research question these campaigns answer; Holm corrects within one invocation, and this records which family that was")
|
||||
flagSet.BoolVar(&paired, "paired", false, "the two arms ran the same seeds: contrast them seed by seed with the sign test instead of pooling them into two independent samples")
|
||||
if err := flagSet.Parse(arguments); err != nil {
|
||||
return err
|
||||
}
|
||||
directories = append(directories, flagSet.Args()...)
|
||||
if len(directories) == 0 {
|
||||
return errors.New("no campaign directories given")
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
for _, directory := range directories {
|
||||
resolved, err := filepath.Abs(directory)
|
||||
if err != nil {
|
||||
return fmt.Errorf("resolve %s: %w", directory, err)
|
||||
}
|
||||
if seen[resolved] {
|
||||
return fmt.Errorf("campaign directory %s given twice: its runs would be counted twice", directory)
|
||||
}
|
||||
seen[resolved] = true
|
||||
}
|
||||
|
||||
arms, err := groupArms(directories)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
var result analysis
|
||||
if paired {
|
||||
result, err = analysePaired(arms, time.Now().UTC())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
} else {
|
||||
result, err = analyse(arms, time.Now().UTC())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
result.Question = question
|
||||
writeReport(result, stdout)
|
||||
|
||||
if jsonPath == "" {
|
||||
return nil
|
||||
}
|
||||
body, err := json.MarshalIndent(result, "", " ")
|
||||
if err != nil {
|
||||
return fmt.Errorf("marshal summary: %w", err)
|
||||
}
|
||||
body = append(body, '\n')
|
||||
if jsonPath == "-" {
|
||||
_, err = stdout.Write(body)
|
||||
return err
|
||||
}
|
||||
return os.WriteFile(jsonPath, body, 0o644)
|
||||
}
|
||||
|
||||
func main() {
|
||||
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(os.Stderr, "error: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,174 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"slices"
|
||||
)
|
||||
|
||||
// signTest is the exact two-sided sign test over matched pairs. Under the null
|
||||
// that neither arm reaches its first violation sooner, a pair whose order the
|
||||
// censoring determines falls either way with probability one half, so the count
|
||||
// is binomial and the two-sided p-value doubles the smaller tail. Pairs left
|
||||
// with no order carry no information and are not trials.
|
||||
//
|
||||
// It is the seed-matched form of the comparison the unpaired test makes, and
|
||||
// it is what the log-rank stratified by seed reduces to with one run per arm in
|
||||
// each stratum. The magnitude-based alternatives are not available: a
|
||||
// difference in steps needs both runs to have violated, and a paired test built
|
||||
// on scores of censored times, the paired Prentice-Wilcoxon among them, is
|
||||
// centred at zero under the null only when the two arms censor alike, which is
|
||||
// exactly what the wall clock stops them from doing.
|
||||
func signTest(favouringFirst, favouringSecond int) float64 {
|
||||
trials := favouringFirst + favouringSecond
|
||||
if trials == 0 {
|
||||
return math.NaN()
|
||||
}
|
||||
smaller := min(favouringFirst, favouringSecond)
|
||||
tail := 0.0
|
||||
for count := 0; count <= smaller; count++ {
|
||||
tail += math.Exp(logBinomialCoefficient(trials, count) - float64(trials)*math.Ln2)
|
||||
}
|
||||
return math.Min(2*tail, 1)
|
||||
}
|
||||
|
||||
func logBinomialCoefficient(trials, chosen int) float64 {
|
||||
all, _ := math.Lgamma(float64(trials + 1))
|
||||
picked, _ := math.Lgamma(float64(chosen + 1))
|
||||
rest, _ := math.Lgamma(float64(trials-chosen) + 1)
|
||||
return all - picked - rest
|
||||
}
|
||||
|
||||
// pairedComparison is the seed-matched contrast the actuation ablation reports.
|
||||
// A pair is scored the way the unpaired comparison scores one, by which run
|
||||
// outlived the other, so Sign is +1 when the second arm is the one seen to
|
||||
// violate sooner across the pairs whose order censoring determines.
|
||||
type pairedComparison struct {
|
||||
First string `json:"first"`
|
||||
Second string `json:"second"`
|
||||
Pairs int `json:"pairs"`
|
||||
UnpairedSeeds []int64 `json:"unpaired_seeds,omitempty"`
|
||||
Sign int `json:"sign"`
|
||||
FirstSooner int `json:"first_sooner"`
|
||||
SecondSooner int `json:"second_sooner"`
|
||||
// Unordered is the pairs the censoring leaves in no order, either because
|
||||
// both runs ended clean or because the run that stopped first stopped before
|
||||
// the other violated. They are not evidence either way and are not trials.
|
||||
Unordered int `json:"unordered_pairs"`
|
||||
// MedianDifference is in steps and is undefined unless some pair has both
|
||||
// runs violating, which is the only shape a difference in steps can be read
|
||||
// off. BothViolated says how many pairs it summarizes, because it describes
|
||||
// those pairs and not the sample.
|
||||
MedianDifference *float64 `json:"median_step_difference"`
|
||||
BothViolated int `json:"both_violated_pairs"`
|
||||
// A12 is the within-pair form of the Vargha-Delaney effect size, the share
|
||||
// of matched seeds on which the first arm took more steps, an unordered pair
|
||||
// counting as half. A matched design has no reason to compare the two arms
|
||||
// as pooled bags of runs when each seed has a partner.
|
||||
A12 float64 `json:"a12_within_pairs"`
|
||||
// PValue is undefined, and null in the summary, when censoring left no pair
|
||||
// ordered: there is nothing for the test to be a test of, and JSON has no
|
||||
// way to write the number that is not there.
|
||||
PValue *float64 `json:"p_value"`
|
||||
HolmPValue *float64 `json:"holm_p_value"`
|
||||
}
|
||||
|
||||
// pairArms matches the two arms by seed and contrasts them pair by pair. A seed
|
||||
// usable in one arm and not the other is named rather than dropped silently,
|
||||
// because that is a host that lost a run and it is what the campaign manifest
|
||||
// exists to make visible.
|
||||
func pairArms(first, second arm) (pairedComparison, error) {
|
||||
firstBySeed, err := usableBySeed(first)
|
||||
if err != nil {
|
||||
return pairedComparison{}, err
|
||||
}
|
||||
secondBySeed, err := usableBySeed(second)
|
||||
if err != nil {
|
||||
return pairedComparison{}, err
|
||||
}
|
||||
|
||||
comparison := pairedComparison{First: first.Name, Second: second.Name, A12: math.NaN()}
|
||||
var differences []float64
|
||||
for _, seed := range sortedSeeds(firstBySeed, secondBySeed) {
|
||||
left, inFirst := firstBySeed[seed]
|
||||
right, inSecond := secondBySeed[seed]
|
||||
if !inFirst || !inSecond {
|
||||
comparison.UnpairedSeeds = append(comparison.UnpairedSeeds, seed)
|
||||
continue
|
||||
}
|
||||
leftRun := observationOf(left, first.Budget)
|
||||
rightRun := observationOf(right, second.Budget)
|
||||
comparison.Pairs++
|
||||
switch outlives(leftRun, rightRun) {
|
||||
case 1:
|
||||
comparison.SecondSooner++
|
||||
case -1:
|
||||
comparison.FirstSooner++
|
||||
default:
|
||||
comparison.Unordered++
|
||||
}
|
||||
if leftRun.Event && rightRun.Event {
|
||||
comparison.BothViolated++
|
||||
differences = append(differences, leftRun.Steps-rightRun.Steps)
|
||||
}
|
||||
}
|
||||
if comparison.Pairs == 0 {
|
||||
return comparison, nil
|
||||
}
|
||||
|
||||
if len(differences) > 0 {
|
||||
median := medianOf(differences)
|
||||
comparison.MedianDifference = &median
|
||||
}
|
||||
switch {
|
||||
case comparison.SecondSooner > comparison.FirstSooner:
|
||||
comparison.Sign = 1
|
||||
case comparison.FirstSooner > comparison.SecondSooner:
|
||||
comparison.Sign = -1
|
||||
}
|
||||
comparison.A12 = (float64(comparison.SecondSooner) + 0.5*float64(comparison.Unordered)) / float64(comparison.Pairs)
|
||||
if tested := signTest(comparison.FirstSooner, comparison.SecondSooner); !math.IsNaN(tested) {
|
||||
comparison.PValue = &tested
|
||||
}
|
||||
return comparison, nil
|
||||
}
|
||||
|
||||
func usableBySeed(current arm) (map[int64]classifiedRun, error) {
|
||||
bySeed := map[int64]classifiedRun{}
|
||||
for _, item := range current.Runs {
|
||||
if item.ExcludedBecause != "" {
|
||||
continue
|
||||
}
|
||||
if _, repeated := bySeed[item.Seed]; repeated {
|
||||
return nil, fmt.Errorf("arm %q has more than one usable run for seed %d: a seed-matched "+
|
||||
"comparison cannot choose between them", current.Name, item.Seed)
|
||||
}
|
||||
bySeed[item.Seed] = item
|
||||
}
|
||||
return bySeed, nil
|
||||
}
|
||||
|
||||
func sortedSeeds(sets ...map[int64]classifiedRun) []int64 {
|
||||
var seeds []int64
|
||||
seen := map[int64]bool{}
|
||||
for _, set := range sets {
|
||||
for seed := range set {
|
||||
if seen[seed] {
|
||||
continue
|
||||
}
|
||||
seen[seed] = true
|
||||
seeds = append(seeds, seed)
|
||||
}
|
||||
}
|
||||
slices.Sort(seeds)
|
||||
return seeds
|
||||
}
|
||||
|
||||
func medianOf(values []float64) float64 {
|
||||
sorted := slices.Sorted(slices.Values(values))
|
||||
middle := len(sorted) / 2
|
||||
if len(sorted)%2 == 1 {
|
||||
return sorted[middle]
|
||||
}
|
||||
return (sorted[middle-1] + sorted[middle]) / 2
|
||||
}
|
||||
@@ -0,0 +1,190 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The two-sided sign test is R's binom.test(k, n) at p = 0.5, which is the
|
||||
// doubled tail of a symmetric binomial and can be worked out by hand from the
|
||||
// coefficients: 2 * sum(C(n, i), i <= min(k, n-k)) / 2^n.
|
||||
func TestSignTest_MatchesTheBinomialTail(t *testing.T) {
|
||||
cases := []struct {
|
||||
first, second int
|
||||
want float64
|
||||
}{
|
||||
{0, 10, 2.0 / 1024},
|
||||
{1, 9, 2 * 11.0 / 1024},
|
||||
{3, 7, 2 * 176.0 / 1024},
|
||||
{5, 5, 1},
|
||||
{0, 1, 1},
|
||||
{2, 0, 0.5},
|
||||
}
|
||||
for _, test := range cases {
|
||||
got := signTest(test.first, test.second)
|
||||
if math.Abs(got-test.want) > 1e-12 {
|
||||
t.Errorf("sign test on %d against %d gives %v, want %v", test.first, test.second, got, test.want)
|
||||
}
|
||||
if reversed := signTest(test.second, test.first); math.Abs(reversed-got) > 1e-12 {
|
||||
t.Errorf("sign test on %d against %d gives %v reversed and %v forward",
|
||||
test.first, test.second, reversed, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A campaign runs tens of seeds, not tens of thousands, but the tail is summed
|
||||
// through log-gamma rather than through factorials so that a lopsided family
|
||||
// stays a number rather than becoming an overflow.
|
||||
func TestSignTest_LargeCountsStayFinite(t *testing.T) {
|
||||
if got := signTest(0, 200); got <= 0 || got > 1e-59 {
|
||||
t.Errorf("sign test on 0 against 200 gives %v, want a positive value around 2^-199", got)
|
||||
}
|
||||
if got := signTest(100, 100); math.Abs(got-1) > 1e-12 {
|
||||
t.Errorf("sign test on an even split gives %v, want 1", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSignTest_NoOrderedPairHasNoTest(t *testing.T) {
|
||||
if got := signTest(0, 0); !math.IsNaN(got) {
|
||||
t.Errorf("sign test with nothing to test gives %v, want undefined", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPairArms_ScoresEachPairByWhichRunOutlivedTheOther(t *testing.T) {
|
||||
pre := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
|
||||
violatingRun(1, 30, 30, "doubleTapCharges"),
|
||||
cleanRun(2, 40),
|
||||
violatingRun(3, 25, 25, "doubleTapCharges"),
|
||||
}}
|
||||
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{
|
||||
violatingRun(1, 10, 10, "doubleTapCharges"),
|
||||
violatingRun(2, 12, 12, "doubleTapCharges"),
|
||||
violatingRun(3, 25, 25, "doubleTapCharges"),
|
||||
}}
|
||||
|
||||
comparison, err := pairArms(pre, post)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if comparison.Pairs != 3 {
|
||||
t.Fatalf("%d pairs, want 3", comparison.Pairs)
|
||||
}
|
||||
// Seed 1 violated at 30 against 10 and seed 2 was still clean at 40 when its
|
||||
// partner violated at 12, so both go to the second arm; seed 3 violated on
|
||||
// the same step in both and has no order.
|
||||
if comparison.SecondSooner != 2 || comparison.FirstSooner != 0 || comparison.Unordered != 1 {
|
||||
t.Errorf("counts %+v, want two favouring the second arm and one unordered", comparison)
|
||||
}
|
||||
if comparison.Sign != 1 {
|
||||
t.Errorf("sign %d, want +1 for the arm that violated later", comparison.Sign)
|
||||
}
|
||||
if math.Abs(comparison.A12-2.5/3) > 1e-12 {
|
||||
t.Errorf("a12 within pairs %v, want %v", comparison.A12, 2.5/3)
|
||||
}
|
||||
// Only seeds 1 and 3 have a difference in steps to take a median of, 20 and
|
||||
// 0: the pair holding a clean run has no difference either arm supports.
|
||||
if comparison.BothViolated != 2 || comparison.MedianDifference == nil || *comparison.MedianDifference != 10 {
|
||||
t.Errorf("median difference %v over %d pair(s), want 10 over 2",
|
||||
comparison.MedianDifference, comparison.BothViolated)
|
||||
}
|
||||
if want := signTest(0, 2); comparison.PValue == nil || *comparison.PValue != want {
|
||||
t.Errorf("p %v, want the sign test's %v over the two ordered pairs", comparison.PValue, want)
|
||||
}
|
||||
}
|
||||
|
||||
// Two clean runs are two runs that were still going when they stopped, whatever
|
||||
// step each stopped on, so the pair says nothing and is not a trial.
|
||||
func TestPairArms_PairsOfCleanRunsAreNotEvidence(t *testing.T) {
|
||||
early := arm{Name: "early", Budget: 400, Runs: []classifiedRun{cleanRun(1, 12), cleanRun(2, 14)}}
|
||||
late := arm{Name: "late", Budget: 400, Runs: []classifiedRun{cleanRun(1, 400), cleanRun(2, 380)}}
|
||||
|
||||
comparison, err := pairArms(early, late)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if comparison.Unordered != 2 || comparison.Sign != 0 {
|
||||
t.Errorf("comparison %+v, want both pairs unordered and no direction", comparison)
|
||||
}
|
||||
if comparison.PValue != nil {
|
||||
t.Errorf("p %v, want undefined with no ordered pair", *comparison.PValue)
|
||||
}
|
||||
if comparison.MedianDifference != nil {
|
||||
t.Errorf("median difference %v, want undefined where no pair has two violations",
|
||||
*comparison.MedianDifference)
|
||||
}
|
||||
if comparison.A12 != 0.5 {
|
||||
t.Errorf("a12 within pairs %v, want 0.5", comparison.A12)
|
||||
}
|
||||
}
|
||||
|
||||
// A run excluded as missing data cannot be paired against anything, and the
|
||||
// seed it came from has to be named rather than silently shrinking the sample.
|
||||
func TestPairArms_NamesSeedsUsableInOneArmOnly(t *testing.T) {
|
||||
pre := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
|
||||
violatingRun(1, 30, 30, "p"),
|
||||
{Seed: 2, ExcludedBecause: reasonTimedOut},
|
||||
violatingRun(3, 20, 20, "p"),
|
||||
}}
|
||||
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{
|
||||
violatingRun(1, 10, 10, "p"),
|
||||
violatingRun(2, 11, 11, "p"),
|
||||
}}
|
||||
|
||||
comparison, err := pairArms(pre, post)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if comparison.Pairs != 1 {
|
||||
t.Fatalf("%d pairs, want 1", comparison.Pairs)
|
||||
}
|
||||
if len(comparison.UnpairedSeeds) != 2 || comparison.UnpairedSeeds[0] != 2 || comparison.UnpairedSeeds[1] != 3 {
|
||||
t.Errorf("unpaired seeds %v, want [2 3]", comparison.UnpairedSeeds)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPairArms_RefusesTwoUsableRunsForOneSeed(t *testing.T) {
|
||||
pooled := arm{Name: "pre", Budget: 40, Runs: []classifiedRun{
|
||||
violatingRun(1, 30, 30, "p"),
|
||||
violatingRun(1, 12, 12, "p"),
|
||||
}}
|
||||
post := arm{Name: "post", Budget: 40, Runs: []classifiedRun{violatingRun(1, 10, 10, "p")}}
|
||||
|
||||
if _, err := pairArms(pooled, post); err == nil {
|
||||
t.Fatal("paired two arms where one seed ran twice")
|
||||
}
|
||||
}
|
||||
|
||||
// The paired comparison is the ablation's decision rule, so the direction it
|
||||
// reports has to survive the arms being passed the other way round.
|
||||
func TestPairArms_DirectionReversesWithTheArms(t *testing.T) {
|
||||
pre := arm{Name: "pre", Budget: 400, Runs: []classifiedRun{
|
||||
cleanRun(1, 400), cleanRun(2, 400), violatingRun(3, 380, 380, "p"),
|
||||
cleanRun(4, 400), violatingRun(5, 350, 350, "p"),
|
||||
}}
|
||||
post := arm{Name: "post", Budget: 400, Runs: []classifiedRun{
|
||||
violatingRun(1, 40, 40, "p"), violatingRun(2, 90, 90, "p"), violatingRun(3, 60, 60, "p"),
|
||||
violatingRun(4, 120, 120, "p"), violatingRun(5, 30, 30, "p"),
|
||||
}}
|
||||
|
||||
forward, err := pairArms(pre, post)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
reversed, err := pairArms(post, pre)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if forward.Sign != 1 || reversed.Sign != -1 {
|
||||
t.Errorf("signs %+d and %+d, want +1 then -1", forward.Sign, reversed.Sign)
|
||||
}
|
||||
if *forward.MedianDifference != -*reversed.MedianDifference {
|
||||
t.Errorf("median differences %v and %v, want opposites",
|
||||
*forward.MedianDifference, *reversed.MedianDifference)
|
||||
}
|
||||
if math.Abs(*forward.PValue-*reversed.PValue) > 1e-12 {
|
||||
t.Errorf("p-values %v and %v, want the same two-sided value", *forward.PValue, *reversed.PValue)
|
||||
}
|
||||
if math.Abs(forward.A12+reversed.A12-1) > 1e-12 {
|
||||
t.Errorf("a12 %v and %v, want them to sum to 1", forward.A12, reversed.A12)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,676 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"math"
|
||||
"math/rand"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// A pipeline exercised only on data whose answer nobody knows reports that it
|
||||
// runs, not that it is right. These tests plant effects whose value is known in
|
||||
// advance from the model that generated the data, and require the pipeline to
|
||||
// recover them from campaign directories it reads off disk through its own
|
||||
// entry point.
|
||||
//
|
||||
// plantedModel is the generator: a run that has not yet violated does so at
|
||||
// every step with probability Hazard, and a run reaching Budget without
|
||||
// violating is right-censored there. Steps to first violation is therefore
|
||||
// geometric, truncated at the budget, and every quantity the pipeline reports
|
||||
// about it has a closed form below.
|
||||
type plantedModel struct {
|
||||
Hazard float64
|
||||
Budget int
|
||||
}
|
||||
|
||||
// recordedValues is the distribution of what one run contributes to the
|
||||
// analysis: the violation step when it violates, and the budget when it does
|
||||
// not, since a censored run is held at the budget. Index t carries P(value = t).
|
||||
func (model plantedModel) recordedValues() []float64 {
|
||||
probabilities := make([]float64, model.Budget+1)
|
||||
survival := 1.0
|
||||
for step := 1; step < model.Budget; step++ {
|
||||
probabilities[step] = survival * model.Hazard
|
||||
survival *= 1 - model.Hazard
|
||||
}
|
||||
probabilities[model.Budget] = survival
|
||||
return probabilities
|
||||
}
|
||||
|
||||
func (model plantedModel) survivalAt(step int) float64 {
|
||||
return math.Pow(1-model.Hazard, float64(step))
|
||||
}
|
||||
|
||||
func (model plantedModel) censoredShare() float64 {
|
||||
return model.survivalAt(model.Budget)
|
||||
}
|
||||
|
||||
// medianSteps is the smallest step at which the true survival function falls to
|
||||
// or below one half, which is what the Kaplan-Meier median estimates.
|
||||
func (model plantedModel) medianSteps() (int, bool) {
|
||||
for step := 1; step <= model.Budget; step++ {
|
||||
if model.survivalAt(step) <= 0.5 {
|
||||
return step, true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// populationA12 is the Vargha-Delaney effect size between two planted models,
|
||||
// computed from the models rather than from any sample: the probability that a
|
||||
// run of the first takes more steps than a run of the second, counting a tie as
|
||||
// half. Every run held at the budget ties with every other, which is the term a
|
||||
// pipeline mishandling censoring gets wrong.
|
||||
func populationA12(first, second plantedModel) float64 {
|
||||
left, right := first.recordedValues(), second.recordedValues()
|
||||
total := 0.0
|
||||
for value, probability := range left {
|
||||
if probability == 0 {
|
||||
continue
|
||||
}
|
||||
for other, otherProbability := range right {
|
||||
switch {
|
||||
case value > other:
|
||||
total += probability * otherProbability
|
||||
case value == other:
|
||||
total += 0.5 * probability * otherProbability
|
||||
}
|
||||
}
|
||||
}
|
||||
return total
|
||||
}
|
||||
|
||||
type plantedRun struct {
|
||||
seed int64
|
||||
steps int
|
||||
violated bool
|
||||
}
|
||||
|
||||
func (model plantedModel) draw(seed int64, source *rand.Rand) plantedRun {
|
||||
for step := 1; step <= model.Budget; step++ {
|
||||
if source.Float64() < model.Hazard {
|
||||
return plantedRun{seed: seed, steps: step, violated: true}
|
||||
}
|
||||
}
|
||||
return plantedRun{seed: seed, steps: model.Budget}
|
||||
}
|
||||
|
||||
func (model plantedModel) drawCampaign(source *rand.Rand, seeds int) []plantedRun {
|
||||
runs := make([]plantedRun, 0, seeds)
|
||||
for seed := int64(1); seed <= int64(seeds); seed++ {
|
||||
runs = append(runs, model.draw(seed, source))
|
||||
}
|
||||
return runs
|
||||
}
|
||||
|
||||
// writePlantedCampaign emits the campaign directory shape the campaign tool
|
||||
// writes, so the planted data reaches the statistics through the same loader,
|
||||
// classifier and censoring rules as a real sweep.
|
||||
func writePlantedCampaign(t *testing.T, directory, armName string, budget int, runs []plantedRun) string {
|
||||
t.Helper()
|
||||
seeds := make([]int64, 0, len(runs))
|
||||
records := make([]map[string]any, 0, len(runs))
|
||||
for _, run := range runs {
|
||||
seeds = append(seeds, run.seed)
|
||||
record := map[string]any{
|
||||
"seed": run.seed,
|
||||
"exit_code": 0,
|
||||
"steps": run.steps,
|
||||
"actions": run.steps,
|
||||
"monotonic_millis": int64(run.steps) * 900,
|
||||
}
|
||||
if run.violated {
|
||||
record["first_violation_origin_step"] = run.steps
|
||||
record["violated_properties"] = []string{"plantedProperty"}
|
||||
} else {
|
||||
record["first_violation_origin_step"] = nil
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
writeCampaign(t, directory, map[string]any{
|
||||
"arm": armName,
|
||||
"generator": "seeded",
|
||||
"platform": "android",
|
||||
"max_steps": budget,
|
||||
"seeds": seeds,
|
||||
"host": "planted",
|
||||
}, records)
|
||||
return directory
|
||||
}
|
||||
|
||||
// analyseCampaigns runs the tool exactly as the command line does and returns
|
||||
// the machine-readable summary it wrote.
|
||||
func analyseCampaigns(t *testing.T, arguments ...string) analysis {
|
||||
t.Helper()
|
||||
summaryPath := filepath.Join(t.TempDir(), "analysis.json")
|
||||
if err := run(append([]string{"--json", summaryPath}, arguments...), io.Discard, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body, err := os.ReadFile(summaryPath)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var result analysis
|
||||
if err := json.Unmarshal(body, &result); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func armByName(t *testing.T, result analysis, name string) armSummary {
|
||||
t.Helper()
|
||||
for _, summary := range result.Arms {
|
||||
if summary.Arm == name {
|
||||
return summary
|
||||
}
|
||||
}
|
||||
t.Fatalf("no arm %q in %v", name, result.Arms)
|
||||
return armSummary{}
|
||||
}
|
||||
|
||||
// plantTwoArms writes two campaign directories drawn from the given models and
|
||||
// returns them in the order they were written.
|
||||
func plantTwoArms(t *testing.T, sourceSeed int64, seeds int, first, second plantedModel) (string, string) {
|
||||
t.Helper()
|
||||
root := t.TempDir()
|
||||
source := rand.New(rand.NewSource(sourceSeed))
|
||||
firstDirectory := writePlantedCampaign(t, filepath.Join(root, "first"), "first", first.Budget,
|
||||
first.drawCampaign(source, seeds))
|
||||
secondDirectory := writePlantedCampaign(t, filepath.Join(root, "second"), "second", second.Budget,
|
||||
second.drawCampaign(source, seeds))
|
||||
return firstDirectory, secondDirectory
|
||||
}
|
||||
|
||||
// The effect size is the number the paper reports as the size of a difference,
|
||||
// so it is checked against the value the generating models define rather than
|
||||
// against anything this tool produced. Four independent draws are used because
|
||||
// a tolerance that holds for one lucky sample is not a check.
|
||||
func TestPlanted_RecoversTheKnownEffectSize(t *testing.T) {
|
||||
slow := plantedModel{Hazard: 0.005, Budget: 400}
|
||||
fast := plantedModel{Hazard: 0.02, Budget: 400}
|
||||
expected := populationA12(slow, fast)
|
||||
if expected < 0.6 {
|
||||
t.Fatalf("planted effect %v is too small to be worth checking", expected)
|
||||
}
|
||||
|
||||
for _, sourceSeed := range []int64{1, 2, 3, 4} {
|
||||
slowDirectory, fastDirectory := plantTwoArms(t, sourceSeed, 400, slow, fast)
|
||||
result := analyseCampaigns(t, slowDirectory, fastDirectory)
|
||||
|
||||
if len(result.Pairwise) != 1 {
|
||||
t.Fatalf("%d comparisons, want 1", len(result.Pairwise))
|
||||
}
|
||||
pair := result.Pairwise[0]
|
||||
if pair.First != "first" || pair.Second != "second" {
|
||||
t.Fatalf("comparison %s vs %s, want first vs second", pair.First, pair.Second)
|
||||
}
|
||||
if math.Abs(pair.A12-expected) > 0.03 {
|
||||
t.Errorf("seed %d: a12 %.4f, want the planted %.4f", sourceSeed, pair.A12, expected)
|
||||
}
|
||||
if pair.PValue > 1e-6 {
|
||||
t.Errorf("seed %d: p-value %v for an effect this large", sourceSeed, pair.PValue)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A12 above one half has to mean the first arm took longer, and a sign flip
|
||||
// anywhere in the pipeline is the failure that produces a plausible wrong
|
||||
// answer rather than an obvious one.
|
||||
func TestPlanted_EffectSizeDirectionFollowsTheArmOrder(t *testing.T) {
|
||||
slow := plantedModel{Hazard: 0.005, Budget: 400}
|
||||
fast := plantedModel{Hazard: 0.02, Budget: 400}
|
||||
source := rand.New(rand.NewSource(7))
|
||||
slowRuns := slow.drawCampaign(source, 300)
|
||||
fastRuns := fast.drawCampaign(source, 300)
|
||||
|
||||
// Comparisons are ordered by arm label, so the same two samples are
|
||||
// analysed twice with the labels swapped.
|
||||
root := t.TempDir()
|
||||
forward := analyseCampaigns(t,
|
||||
writePlantedCampaign(t, filepath.Join(root, "slow-first"), "a", slow.Budget, slowRuns),
|
||||
writePlantedCampaign(t, filepath.Join(root, "fast-second"), "b", fast.Budget, fastRuns)).Pairwise[0]
|
||||
reversed := analyseCampaigns(t,
|
||||
writePlantedCampaign(t, filepath.Join(root, "slow-second"), "b", slow.Budget, slowRuns),
|
||||
writePlantedCampaign(t, filepath.Join(root, "fast-first"), "a", fast.Budget, fastRuns)).Pairwise[0]
|
||||
if forward.A12 <= 0.5 {
|
||||
t.Errorf("a12 %.4f with the slower arm first, want above 0.5", forward.A12)
|
||||
}
|
||||
if math.Abs(forward.A12+reversed.A12-1) > 1e-12 {
|
||||
t.Errorf("a12 %.4f and %.4f, want them to sum to 1", forward.A12, reversed.A12)
|
||||
}
|
||||
if math.Abs(forward.PValue-reversed.PValue) > 1e-12 {
|
||||
t.Errorf("p-values %v and %v, want the same two-sided value", forward.PValue, reversed.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// The median and the quartiles are read off the Kaplan-Meier curve, and the
|
||||
// curve itself is checked point by point against the survival function that
|
||||
// generated the data. A pipeline counting a censored run as a violation drives
|
||||
// this estimate down hard.
|
||||
func TestPlanted_SurvivalCurveTracksTheGeneratingHazard(t *testing.T) {
|
||||
model := plantedModel{Hazard: 0.005, Budget: 400}
|
||||
if model.censoredShare() < 0.1 {
|
||||
t.Fatalf("planted censoring is %.3f, too little to exercise the estimator", model.censoredShare())
|
||||
}
|
||||
directory, _ := plantTwoArms(t, 11, 600, model, model)
|
||||
summary := armByName(t, analyseCampaigns(t, directory), "first")
|
||||
|
||||
for _, step := range []int{20, 50, 100, 200, 300} {
|
||||
estimated, ok := survivalAtStep(summary.SurvivalCurve, float64(step))
|
||||
if !ok {
|
||||
t.Fatalf("curve has no point at or before step %d", step)
|
||||
}
|
||||
if want := model.survivalAt(step); math.Abs(estimated-want) > 0.05 {
|
||||
t.Errorf("survival at %d is %.4f, want the planted %.4f", step, estimated, want)
|
||||
}
|
||||
}
|
||||
|
||||
median, ok := model.medianSteps()
|
||||
if !ok {
|
||||
t.Fatal("the planted model has no median to recover")
|
||||
}
|
||||
if summary.MedianStepsToFirstViolation == nil {
|
||||
t.Fatal("median undefined, want the planted median")
|
||||
}
|
||||
if math.Abs(*summary.MedianStepsToFirstViolation-float64(median)) > 10 {
|
||||
t.Errorf("median %v, want the planted %d", *summary.MedianStepsToFirstViolation, median)
|
||||
}
|
||||
assertQuartile(t, summary.FirstQuartileSteps, model, 0.25)
|
||||
assertQuartile(t, summary.ThirdQuartileSteps, model, 0.75)
|
||||
}
|
||||
|
||||
func assertQuartile(t *testing.T, got *float64, model plantedModel, fraction float64) {
|
||||
t.Helper()
|
||||
if got == nil {
|
||||
t.Fatalf("quantile %v undefined, want a value", fraction)
|
||||
}
|
||||
if want := model.survivalAt(int(*got)); want > 1-fraction+0.05 {
|
||||
t.Errorf("quantile %v at step %v, where the planted survival is %.4f", fraction, *got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func survivalAtStep(curve []survivalPoint, step float64) (float64, bool) {
|
||||
value, found := 0.0, false
|
||||
for _, point := range curve {
|
||||
if point.Steps > step {
|
||||
break
|
||||
}
|
||||
value, found = point.Survival, true
|
||||
}
|
||||
return value, found
|
||||
}
|
||||
|
||||
// A true null must not be called significant more often than the test's own
|
||||
// level allows. This is the check that catches a wrong variance term, which
|
||||
// leaves every point estimate looking reasonable and only shows up in how often
|
||||
// the pipeline claims a difference that is not there.
|
||||
func TestPlanted_TrueNullIsNotCalledSignificantAboveItsLevel(t *testing.T) {
|
||||
const replicates = 300
|
||||
model := plantedModel{Hazard: 0.01, Budget: 400}
|
||||
rejectedByGehan, rejectedByLogRank := 0, 0
|
||||
for replicate := 0; replicate < replicates; replicate++ {
|
||||
first, second := plantTwoArms(t, int64(1000+replicate), 30, model, model)
|
||||
result := analyseCampaigns(t, first, second)
|
||||
if result.Pairwise[0].PValue < 0.05 {
|
||||
rejectedByGehan++
|
||||
}
|
||||
if result.LogRank.PValue < 0.05 {
|
||||
rejectedByLogRank++
|
||||
}
|
||||
}
|
||||
// Three standard errors around 0.05 at 300 replicates is 0.05 +/- 0.038.
|
||||
// The lower bound is asserted too: a test that never rejects has bought its
|
||||
// level by losing the power the experiment is sized for.
|
||||
for name, rejected := range map[string]int{"gehan": rejectedByGehan, "log-rank": rejectedByLogRank} {
|
||||
rate := float64(rejected) / replicates
|
||||
if rate > 0.09 || rate < 0.015 {
|
||||
t.Errorf("%s called a true null significant in %.1f%% of %d replicates, want about 5%%",
|
||||
name, 100*rate, replicates)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The same null with most runs censored, where nearly every observation is tied
|
||||
// at the budget and the tie correction is what keeps the level honest.
|
||||
func TestPlanted_TrueNullUnderHeavyCensoringKeepsItsLevel(t *testing.T) {
|
||||
const replicates = 300
|
||||
model := plantedModel{Hazard: 0.0004, Budget: 400}
|
||||
if model.censoredShare() < 0.8 {
|
||||
t.Fatalf("planted censoring is %.3f, want most runs censored", model.censoredShare())
|
||||
}
|
||||
rejected := 0
|
||||
for replicate := 0; replicate < replicates; replicate++ {
|
||||
first, second := plantTwoArms(t, int64(9000+replicate), 40, model, model)
|
||||
if analyseCampaigns(t, first, second).Pairwise[0].PValue < 0.05 {
|
||||
rejected++
|
||||
}
|
||||
}
|
||||
if rate := float64(rejected) / replicates; rate > 0.09 {
|
||||
t.Errorf("gehan called a true null significant in %.1f%% of %d replicates under heavy censoring, want at most about 5%%",
|
||||
100*rate, replicates)
|
||||
}
|
||||
}
|
||||
|
||||
// Most runs never violating is the expected shape of an arm at a budget chosen
|
||||
// for a harder subject, and it is the case where an implementation that drops
|
||||
// censored runs instead of holding them at the budget still produces a number.
|
||||
func TestPlanted_HeavyCensoringLeavesTheMedianUndefinedAndKeepsTheEffect(t *testing.T) {
|
||||
quiet := plantedModel{Hazard: 0.00025, Budget: 400}
|
||||
loud := plantedModel{Hazard: 0.002, Budget: 400}
|
||||
if quiet.censoredShare() < 0.85 {
|
||||
t.Fatalf("planted censoring is %.3f, want most runs censored", quiet.censoredShare())
|
||||
}
|
||||
if _, ok := quiet.medianSteps(); ok {
|
||||
t.Fatal("the heavily censored model has a median, so the test is not testing what it says")
|
||||
}
|
||||
|
||||
quietDirectory, loudDirectory := plantTwoArms(t, 21, 400, quiet, loud)
|
||||
result := analyseCampaigns(t, quietDirectory, loudDirectory)
|
||||
summary := armByName(t, result, "first")
|
||||
|
||||
if summary.MedianStepsToFirstViolation != nil {
|
||||
t.Errorf("median %v, want undefined where fewer than half the runs violate",
|
||||
*summary.MedianStepsToFirstViolation)
|
||||
}
|
||||
if summary.ThirdQuartileSteps != nil {
|
||||
t.Errorf("third quartile %v, want undefined", *summary.ThirdQuartileSteps)
|
||||
}
|
||||
if summary.Usable != 400 {
|
||||
t.Errorf("%d usable runs, want all 400: a censored run is data", summary.Usable)
|
||||
}
|
||||
censoredShare := float64(summary.Censored) / float64(summary.Usable)
|
||||
if math.Abs(censoredShare-quiet.censoredShare()) > 0.04 {
|
||||
t.Errorf("censored share %.3f, want the planted %.3f", censoredShare, quiet.censoredShare())
|
||||
}
|
||||
|
||||
expected := populationA12(quiet, loud)
|
||||
pair := result.Pairwise[0]
|
||||
if math.Abs(pair.A12-expected) > 0.03 {
|
||||
t.Errorf("a12 %.4f under heavy censoring, want the planted %.4f", pair.A12, expected)
|
||||
}
|
||||
if pair.PValue > 0.01 {
|
||||
t.Errorf("p-value %v, want the effect to survive the censoring", pair.PValue)
|
||||
}
|
||||
if result.LogRank.Observed[0] >= result.LogRank.Expected[0] {
|
||||
t.Errorf("log-rank observed %v against expected %v for the quiet arm, want fewer than expected",
|
||||
result.LogRank.Observed[0], result.LogRank.Expected[0])
|
||||
}
|
||||
}
|
||||
|
||||
// An arm at a real budget is mostly one enormous group of runs censored
|
||||
// together at it. No exact null distribution covers that, so the conditional
|
||||
// variance carries the p-value, and it is checked against a permutation p-value
|
||||
// over the pipeline's own two samples.
|
||||
func TestPlanted_CensoringAtTheBudgetTracksThePermutationPValue(t *testing.T) {
|
||||
cases := []struct {
|
||||
sourceSeed int64
|
||||
quiet, loud float64
|
||||
seeds int
|
||||
}{
|
||||
{33, 0.0008, 0.0025, 60},
|
||||
{51, 0.0008, 0.0014, 50},
|
||||
{52, 0.0006, 0.0012, 50},
|
||||
{53, 0.0010, 0.0016, 50},
|
||||
}
|
||||
discriminating := 0
|
||||
for _, test := range cases {
|
||||
quiet := plantedModel{Hazard: test.quiet, Budget: 400}
|
||||
loud := plantedModel{Hazard: test.loud, Budget: 400}
|
||||
source := rand.New(rand.NewSource(test.sourceSeed))
|
||||
quietRuns := quiet.drawCampaign(source, test.seeds)
|
||||
loudRuns := loud.drawCampaign(source, test.seeds)
|
||||
|
||||
root := t.TempDir()
|
||||
quietDirectory := writePlantedCampaign(t, filepath.Join(root, "quiet"), "quiet", quiet.Budget, quietRuns)
|
||||
loudDirectory := writePlantedCampaign(t, filepath.Join(root, "loud"), "loud", loud.Budget, loudRuns)
|
||||
pair := analyseCampaigns(t, loudDirectory, quietDirectory).Pairwise[0]
|
||||
|
||||
permuted := permutationGehanTwoSided(plantedObservations(loudRuns), plantedObservations(quietRuns))
|
||||
if permuted > 0.01 {
|
||||
discriminating++
|
||||
}
|
||||
// The conditional variance and the permutation variance are two
|
||||
// estimators and not one, so they are near rather than equal. Over these
|
||||
// four samples the gap is at most 0.011, at a p-value of 0.49 where
|
||||
// nothing is decided; where a decision is made it is under 0.005. A
|
||||
// variance wrong by a factor moves the p-value across orders of
|
||||
// magnitude, which is what this catches.
|
||||
if math.Abs(pair.PValue-permuted) > 0.015 {
|
||||
t.Errorf("hazards %v against %v: p-value %.4f, want near the permutation p-value %.4f",
|
||||
test.quiet, test.loud, pair.PValue, permuted)
|
||||
}
|
||||
}
|
||||
// A permutation p-value already at the floor cannot show a variance moving
|
||||
// by a fraction, so the check has to include cases that are not overwhelming.
|
||||
if discriminating < 3 {
|
||||
t.Fatalf("%d of %d cases carry a p-value large enough to discriminate", discriminating, len(cases))
|
||||
}
|
||||
}
|
||||
|
||||
func plantedObservations(runs []plantedRun) []observation {
|
||||
items := make([]observation, 0, len(runs))
|
||||
for _, run := range runs {
|
||||
items = append(items, observation{Steps: float64(run.steps), Event: run.violated})
|
||||
}
|
||||
return items
|
||||
}
|
||||
|
||||
// permutationGehanTwoSided is the randomization p-value of Gehan's statistic:
|
||||
// relabel the pooled runs many times, censoring and all, and count how often the
|
||||
// statistic lands at least as far from zero as the observed one. Both arms here
|
||||
// censor on the same schedule, which is what makes relabelling them a null.
|
||||
func permutationGehanTwoSided(first, second []observation) float64 {
|
||||
pooled := append(append([]observation{}, first...), second...)
|
||||
observed, _ := gehanReference(first, second)
|
||||
source := rand.New(rand.NewSource(99))
|
||||
const shuffles = 20000
|
||||
extreme := 0
|
||||
for shuffle := 0; shuffle < shuffles; shuffle++ {
|
||||
source.Shuffle(len(pooled), func(left, right int) {
|
||||
pooled[left], pooled[right] = pooled[right], pooled[left]
|
||||
})
|
||||
statistic, _ := gehanReference(pooled[:len(first)], pooled[len(first):])
|
||||
if math.Abs(statistic) >= math.Abs(observed)-1e-9 {
|
||||
extreme++
|
||||
}
|
||||
}
|
||||
return float64(extreme) / shuffles
|
||||
}
|
||||
|
||||
// Holm has to change an answer somewhere, or its presence in the pipeline is
|
||||
// decoration. Six arms drawn from hazards close together produce a family where
|
||||
// uncorrected p-values call several comparisons significant and the correction
|
||||
// withdraws some of them.
|
||||
func TestPlanted_HolmWithdrawsAConclusionUncorrectedPValuesReach(t *testing.T) {
|
||||
budget := 400
|
||||
hazards := []float64{0.0100, 0.0125, 0.0150, 0.0175, 0.0200, 0.0225}
|
||||
root := t.TempDir()
|
||||
source := rand.New(rand.NewSource(4242))
|
||||
var directories []string
|
||||
for index, hazard := range hazards {
|
||||
model := plantedModel{Hazard: hazard, Budget: budget}
|
||||
name := fmt.Sprintf("arm%d", index)
|
||||
directories = append(directories, writePlantedCampaign(t,
|
||||
filepath.Join(root, name), name, budget, model.drawCampaign(source, 40)))
|
||||
}
|
||||
|
||||
result := analyseCampaigns(t, directories...)
|
||||
if len(result.Pairwise) != 15 {
|
||||
t.Fatalf("%d comparisons, want 15 over six arms", len(result.Pairwise))
|
||||
}
|
||||
if result.HolmFamilySize != 15 {
|
||||
t.Errorf("holm family size %d, want 15", result.HolmFamilySize)
|
||||
}
|
||||
|
||||
uncorrected, corrected, withdrawn := 0, 0, 0
|
||||
for _, pair := range result.Pairwise {
|
||||
if pair.PValue < 0.05 {
|
||||
uncorrected++
|
||||
}
|
||||
if pair.HolmPValue < 0.05 {
|
||||
corrected++
|
||||
}
|
||||
if pair.PValue < 0.05 && pair.HolmPValue >= 0.05 {
|
||||
withdrawn++
|
||||
}
|
||||
}
|
||||
if withdrawn == 0 {
|
||||
t.Fatalf("holm withdrew no conclusion: %d of 15 comparisons significant before and after", uncorrected)
|
||||
}
|
||||
if corrected >= uncorrected {
|
||||
t.Errorf("%d comparisons significant uncorrected and %d after holm, want fewer", uncorrected, corrected)
|
||||
}
|
||||
}
|
||||
|
||||
// Holm is applied within one research question and not across the paper, so the
|
||||
// same two campaigns analysed on their own must keep their raw p-value while the
|
||||
// same comparison inside a larger family is corrected against that family.
|
||||
func TestPlanted_HolmFamilyIsTheInvocationAndNotEveryComparisonEverMade(t *testing.T) {
|
||||
budget := 400
|
||||
root := t.TempDir()
|
||||
source := rand.New(rand.NewSource(777))
|
||||
var directories []string
|
||||
for index, hazard := range []float64{0.010, 0.014, 0.018, 0.022} {
|
||||
model := plantedModel{Hazard: hazard, Budget: budget}
|
||||
name := fmt.Sprintf("arm%d", index)
|
||||
directories = append(directories, writePlantedCampaign(t,
|
||||
filepath.Join(root, name), name, budget, model.drawCampaign(source, 40)))
|
||||
}
|
||||
|
||||
whole := analyseCampaigns(t, append([]string{"--question", "RQ4"}, directories...)...)
|
||||
alone := analyseCampaigns(t, "--question", "RQ4a", directories[0], directories[1])
|
||||
|
||||
if whole.Question != "RQ4" || alone.Question != "RQ4a" {
|
||||
t.Errorf("questions %q and %q, want them recorded next to the corrected values", whole.Question, alone.Question)
|
||||
}
|
||||
if whole.HolmFamilySize != 6 || alone.HolmFamilySize != 1 {
|
||||
t.Fatalf("family sizes %d and %d, want 6 and 1", whole.HolmFamilySize, alone.HolmFamilySize)
|
||||
}
|
||||
if alone.Pairwise[0].HolmPValue != alone.Pairwise[0].PValue {
|
||||
t.Errorf("holm p %v against raw p %v in a family of one, want them equal",
|
||||
alone.Pairwise[0].HolmPValue, alone.Pairwise[0].PValue)
|
||||
}
|
||||
|
||||
var inFamily pairwiseResult
|
||||
for _, pair := range whole.Pairwise {
|
||||
if pair.First == alone.Pairwise[0].First && pair.Second == alone.Pairwise[0].Second {
|
||||
inFamily = pair
|
||||
}
|
||||
}
|
||||
if math.Abs(inFamily.PValue-alone.Pairwise[0].PValue) > 1e-12 {
|
||||
t.Fatalf("the same comparison has raw p %v in one family and %v in the other",
|
||||
inFamily.PValue, alone.Pairwise[0].PValue)
|
||||
}
|
||||
if inFamily.HolmPValue <= alone.Pairwise[0].HolmPValue {
|
||||
t.Errorf("holm p %v inside a family of six, want it above the %v it carries alone",
|
||||
inFamily.HolmPValue, alone.Pairwise[0].HolmPValue)
|
||||
}
|
||||
}
|
||||
|
||||
// The ablation's outcome is a per-seed difference with a sign, so the planted
|
||||
// shift is applied seed by seed and the pipeline has to return that shift with
|
||||
// that sign from the campaign directories alone.
|
||||
func TestPlanted_PairedComparisonRecoversTheShiftAndItsSign(t *testing.T) {
|
||||
budget := 400
|
||||
base := plantedModel{Hazard: 0.01, Budget: budget}
|
||||
const shift = 60
|
||||
source := rand.New(rand.NewSource(5150))
|
||||
|
||||
var post, pre []plantedRun
|
||||
// Arms are contrasted in the order their labels sort, so the plant is stated
|
||||
// the same way: post-repair against pre-repair. A seed where the unshifted
|
||||
// run violated is a pair the shifted run is known to have outlived, whether
|
||||
// it violated later or ran on clean; a seed where neither violated is a pair
|
||||
// with no order.
|
||||
firstSooner, bothViolated, unordered := 0, 0, 0
|
||||
for seed := int64(1); seed <= 30; seed++ {
|
||||
fast := base.draw(seed, source)
|
||||
slow := plantedRun{seed: seed, steps: fast.steps + shift, violated: fast.violated}
|
||||
if !fast.violated || slow.steps > budget {
|
||||
slow = plantedRun{seed: seed, steps: budget}
|
||||
}
|
||||
post = append(post, fast)
|
||||
pre = append(pre, slow)
|
||||
switch {
|
||||
case !fast.violated:
|
||||
unordered++
|
||||
case slow.violated:
|
||||
bothViolated++
|
||||
firstSooner++
|
||||
default:
|
||||
firstSooner++
|
||||
}
|
||||
}
|
||||
if bothViolated == 0 {
|
||||
t.Fatal("no pair has two violations, so the planted shift is nowhere the analysis can read it")
|
||||
}
|
||||
|
||||
root := t.TempDir()
|
||||
preDirectory := writePlantedCampaign(t, filepath.Join(root, "pre"), "pre-repair", budget, pre)
|
||||
postDirectory := writePlantedCampaign(t, filepath.Join(root, "post"), "post-repair", budget, post)
|
||||
result := analyseCampaigns(t, "--paired", "--question", "RQ4 ablation", preDirectory, postDirectory)
|
||||
|
||||
if result.Paired == nil {
|
||||
t.Fatal("no paired comparison, want one")
|
||||
}
|
||||
if len(result.Pairwise) != 0 {
|
||||
t.Errorf("%d unpaired comparisons alongside the paired one, want none in a seed-matched design",
|
||||
len(result.Pairwise))
|
||||
}
|
||||
paired := *result.Paired
|
||||
if paired.First != "post-repair" || paired.Second != "pre-repair" {
|
||||
t.Fatalf("paired %s minus %s, want post-repair minus pre-repair", paired.First, paired.Second)
|
||||
}
|
||||
if paired.Pairs != 30 {
|
||||
t.Errorf("%d pairs, want 30", paired.Pairs)
|
||||
}
|
||||
// Every pair where both runs violated was shifted by exactly the plant, so
|
||||
// the median over them is the plant itself rather than a mixture of it with
|
||||
// the step counts censored runs never reached.
|
||||
if paired.BothViolated != bothViolated {
|
||||
t.Errorf("%d pair(s) with two violations, want %d", paired.BothViolated, bothViolated)
|
||||
}
|
||||
if paired.MedianDifference == nil || *paired.MedianDifference != -shift {
|
||||
t.Errorf("median difference %v, want the planted %d", paired.MedianDifference, -shift)
|
||||
}
|
||||
if paired.Sign != -1 {
|
||||
t.Errorf("sign %+d, want -1 for the arm that violated sooner", paired.Sign)
|
||||
}
|
||||
if paired.FirstSooner != firstSooner || paired.SecondSooner != 0 || paired.Unordered != unordered {
|
||||
t.Errorf("counts %+v, want %d favouring the shifted arm, none the other way and %d unordered",
|
||||
paired, firstSooner, unordered)
|
||||
}
|
||||
if want := 0.5 * float64(unordered) / 30; paired.A12 != want {
|
||||
t.Errorf("a12 within pairs %v, want %v where no pair favours the first arm", paired.A12, want)
|
||||
}
|
||||
if paired.PValue == nil || *paired.PValue > 0.001 {
|
||||
t.Errorf("p-value %v for a shift planted in every pair", paired.PValue)
|
||||
}
|
||||
if *paired.HolmPValue != *paired.PValue {
|
||||
t.Errorf("holm p %v in a family of one, want the raw %v", *paired.HolmPValue, *paired.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// The paired test is the ablation's decision rule, so a null there has to stay a
|
||||
// null: two arms drawn from the same model, matched by seed, must not report a
|
||||
// difference more often than the level allows.
|
||||
func TestPlanted_PairedNullIsNotCalledSignificantAboveItsLevel(t *testing.T) {
|
||||
model := plantedModel{Hazard: 0.01, Budget: 400}
|
||||
const replicates = 200
|
||||
rejected := 0
|
||||
for replicate := 0; replicate < replicates; replicate++ {
|
||||
first, second := plantTwoArms(t, int64(5000+replicate), 30, model, model)
|
||||
result := analyseCampaigns(t, "--paired", first, second)
|
||||
if *result.Paired.PValue < 0.05 {
|
||||
rejected++
|
||||
}
|
||||
}
|
||||
if rate := float64(rejected) / replicates; rate > 0.10 {
|
||||
t.Errorf("the paired test called a true null significant in %.1f%% of %d replicates, want about 5%%",
|
||||
100*rate, replicates)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,17 @@
|
||||
package main
|
||||
|
||||
// quantileSurvival is the smallest step count at which the product-limit
|
||||
// estimate falls to or below 1-fraction, which is the fraction-th quantile of
|
||||
// steps to first violation. It is undefined whenever the curve never falls that
|
||||
// far, which is what an arm where most runs exhaust the budget produces, and the
|
||||
// second return value says so rather than substituting a number the data does
|
||||
// not contain.
|
||||
func quantileSurvival(curve []survivalPoint, fraction float64) (float64, bool) {
|
||||
threshold := 1 - fraction
|
||||
for _, point := range curve {
|
||||
if point.Survival <= threshold {
|
||||
return point.Steps, true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
@@ -0,0 +1,225 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"slices"
|
||||
)
|
||||
|
||||
type rankSumResult struct {
|
||||
FirstSize int `json:"first_size"`
|
||||
SecondSize int `json:"second_size"`
|
||||
Statistic float64 `json:"mann_whitney_u"`
|
||||
A12 float64 `json:"a12"`
|
||||
PValue float64 `json:"p_value"`
|
||||
Exact bool `json:"exact"`
|
||||
}
|
||||
|
||||
// exactRankSumLimit matches R's wilcox.test: the exact null distribution is
|
||||
// used only when both samples are below this size and nothing is tied.
|
||||
const exactRankSumLimit = 50
|
||||
|
||||
// vargaDelaneyA12 is the probability that a value drawn from first exceeds one
|
||||
// drawn from second, counting a tie as half:
|
||||
//
|
||||
// A = P(X > Y) + 0.5 * P(X = Y)
|
||||
//
|
||||
// Vargha and Delaney (2000), "A Critique and Improvement of the CL Common
|
||||
// Language Effect Size Statistics of McGraw and Wong", Journal of Educational
|
||||
// and Behavioral Statistics 25(2), 101-132.
|
||||
func vargaDelaneyA12(first, second []float64) float64 {
|
||||
if len(first) == 0 || len(second) == 0 {
|
||||
return math.NaN()
|
||||
}
|
||||
total := 0.0
|
||||
for _, left := range first {
|
||||
for _, right := range second {
|
||||
switch {
|
||||
case left > right:
|
||||
total++
|
||||
case left == right:
|
||||
total += 0.5
|
||||
}
|
||||
}
|
||||
}
|
||||
return total / float64(len(first)*len(second))
|
||||
}
|
||||
|
||||
// rankSum is the two-sided Wilcoxon rank-sum (Mann-Whitney) test. The reported
|
||||
// statistic is U for the first sample, the same quantity R's wilcox.test calls
|
||||
// W. The exact null distribution is used when there are no ties and both
|
||||
// samples are small; otherwise the normal approximation is used with the
|
||||
// continuity correction and the tie correction to the variance.
|
||||
//
|
||||
// It takes plain numbers, so it cannot be given campaign runs, where a clean one
|
||||
// carries a bound and not a value. What it is here for is the uncensored case
|
||||
// the pipeline's comparison has to reproduce: this implementation is checked
|
||||
// against R on a published sample, and gehanTest is checked against this one.
|
||||
func rankSum(first, second []float64) rankSumResult {
|
||||
firstSize, secondSize := len(first), len(second)
|
||||
result := rankSumResult{
|
||||
FirstSize: firstSize,
|
||||
SecondSize: secondSize,
|
||||
Statistic: math.NaN(),
|
||||
A12: math.NaN(),
|
||||
PValue: math.NaN(),
|
||||
}
|
||||
if firstSize == 0 || secondSize == 0 {
|
||||
return result
|
||||
}
|
||||
pooled := make([]float64, 0, firstSize+secondSize)
|
||||
pooled = append(pooled, first...)
|
||||
pooled = append(pooled, second...)
|
||||
ranks, tieGroups := midRanks(pooled)
|
||||
|
||||
rankTotal := 0.0
|
||||
for index := 0; index < firstSize; index++ {
|
||||
rankTotal += ranks[index]
|
||||
}
|
||||
statistic := rankTotal - float64(firstSize)*float64(firstSize+1)/2
|
||||
result.Statistic = statistic
|
||||
result.A12 = vargaDelaneyA12(first, second)
|
||||
|
||||
if len(tieGroups) == 0 && firstSize < exactRankSumLimit && secondSize < exactRankSumLimit {
|
||||
result.Exact = true
|
||||
result.PValue = exactRankSumTwoSided(statistic, firstSize, secondSize)
|
||||
return result
|
||||
}
|
||||
result.PValue = normalRankSumTwoSided(statistic, firstSize, secondSize, tieGroups)
|
||||
return result
|
||||
}
|
||||
|
||||
// normalRankSumTwoSided follows the large-sample branch of R's wilcox.test:
|
||||
//
|
||||
// sigma^2 = (m*n/12) * ((N+1) - sum(t^3 - t) / (N*(N-1)))
|
||||
//
|
||||
// where t runs over the sizes of the tied groups. The 0.5 shift toward the null
|
||||
// mean is the continuity correction.
|
||||
func normalRankSumTwoSided(statistic float64, firstSize, secondSize int, tieGroups []int) float64 {
|
||||
sizeProduct := float64(firstSize) * float64(secondSize)
|
||||
variance := rankSumVariance(firstSize, secondSize, tieGroups)
|
||||
if variance <= 0 {
|
||||
return 1
|
||||
}
|
||||
centered := statistic - sizeProduct/2
|
||||
correction := 0.0
|
||||
switch {
|
||||
case centered > 0:
|
||||
correction = 0.5
|
||||
case centered < 0:
|
||||
correction = -0.5
|
||||
}
|
||||
z := (centered - correction) / math.Sqrt(variance)
|
||||
tail := math.Min(standardNormalUpperTail(z), standardNormalUpperTail(-z))
|
||||
return math.Min(2*tail, 1)
|
||||
}
|
||||
|
||||
func rankSumVariance(firstSize, secondSize int, tieGroups []int) float64 {
|
||||
sizeProduct := float64(firstSize) * float64(secondSize)
|
||||
total := float64(firstSize + secondSize)
|
||||
tieAdjustment := 0.0
|
||||
for _, size := range tieGroups {
|
||||
count := float64(size)
|
||||
tieAdjustment += count*count*count - count
|
||||
}
|
||||
return (sizeProduct / 12) * ((total + 1) - tieAdjustment/(total*(total-1)))
|
||||
}
|
||||
|
||||
// exactRankSumTwoSided doubles the smaller exact tail, as R's wilcox.test does.
|
||||
func exactRankSumTwoSided(statistic float64, firstSize, secondSize int) float64 {
|
||||
counts := exactRankSumCounts(firstSize, secondSize)
|
||||
total := 0.0
|
||||
for _, count := range counts {
|
||||
total += count
|
||||
}
|
||||
value := int(math.Round(statistic))
|
||||
tail := 0.0
|
||||
if statistic > float64(firstSize*secondSize)/2 {
|
||||
for u := value; u < len(counts); u++ {
|
||||
tail += counts[u]
|
||||
}
|
||||
} else {
|
||||
for u := 0; u <= value && u < len(counts); u++ {
|
||||
tail += counts[u]
|
||||
}
|
||||
}
|
||||
return math.Min(2*tail/total, 1)
|
||||
}
|
||||
|
||||
// exactRankSumUpperTail is P(U >= statistic) under the null with no ties.
|
||||
func exactRankSumUpperTail(statistic float64, firstSize, secondSize int) float64 {
|
||||
counts := exactRankSumCounts(firstSize, secondSize)
|
||||
total, tail := 0.0, 0.0
|
||||
for u, count := range counts {
|
||||
total += count
|
||||
if float64(u) >= statistic {
|
||||
tail += count
|
||||
}
|
||||
}
|
||||
return tail / total
|
||||
}
|
||||
|
||||
// exactRankSumCounts returns the number of untied assignments producing each
|
||||
// value of U from 0 to firstSize*secondSize. U equals the sum of the zero-based
|
||||
// pooled positions held by the first sample, less firstSize*(firstSize-1)/2, so
|
||||
// the count is a subset-sum tally over those positions.
|
||||
func exactRankSumCounts(firstSize, secondSize int) []float64 {
|
||||
maximum := firstSize * secondSize
|
||||
offset := firstSize * (firstSize - 1) / 2
|
||||
high := maximum + offset
|
||||
table := make([][]float64, firstSize+1)
|
||||
for index := range table {
|
||||
table[index] = make([]float64, high+1)
|
||||
}
|
||||
table[0][0] = 1
|
||||
for position := 0; position < firstSize+secondSize; position++ {
|
||||
for chosen := min(position+1, firstSize); chosen >= 1; chosen-- {
|
||||
row, previous := table[chosen], table[chosen-1]
|
||||
for sum := high; sum >= position; sum-- {
|
||||
if previous[sum-position] != 0 {
|
||||
row[sum] += previous[sum-position]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
counts := make([]float64, maximum+1)
|
||||
for u := range counts {
|
||||
counts[u] = table[firstSize][u+offset]
|
||||
}
|
||||
return counts
|
||||
}
|
||||
|
||||
// midRanks ranks values from 1, averaging the ranks within a tied group, and
|
||||
// also returns the size of every group of size two or more.
|
||||
func midRanks(values []float64) ([]float64, []int) {
|
||||
order := make([]int, len(values))
|
||||
for index := range order {
|
||||
order[index] = index
|
||||
}
|
||||
slices.SortStableFunc(order, func(left, right int) int {
|
||||
switch {
|
||||
case values[left] < values[right]:
|
||||
return -1
|
||||
case values[left] > values[right]:
|
||||
return 1
|
||||
default:
|
||||
return 0
|
||||
}
|
||||
})
|
||||
ranks := make([]float64, len(values))
|
||||
var tieGroups []int
|
||||
for start := 0; start < len(order); {
|
||||
end := start + 1
|
||||
for end < len(order) && values[order[end]] == values[order[start]] {
|
||||
end++
|
||||
}
|
||||
shared := float64(start+1+end) / 2
|
||||
for index := start; index < end; index++ {
|
||||
ranks[order[index]] = shared
|
||||
}
|
||||
if end-start > 1 {
|
||||
tieGroups = append(tieGroups, end-start)
|
||||
}
|
||||
start = end
|
||||
}
|
||||
return ranks, tieGroups
|
||||
}
|
||||
@@ -0,0 +1,211 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// R's wilcox.test on the Hollander and Wolfe (1973), 69f chorioamnion data
|
||||
// reports W = 35 with an exact two-sided p-value of 0.2544; the one-sided
|
||||
// greater alternative that the help page uses reports the same W with
|
||||
// p-value = 0.1272.
|
||||
func TestRankSum_MatchesPublishedChorioamnionResult(t *testing.T) {
|
||||
result := rankSum(chorioamnionTerm, chorioamnionEarly)
|
||||
|
||||
if result.Statistic != 35 {
|
||||
t.Errorf("statistic %v, want 35", result.Statistic)
|
||||
}
|
||||
if !result.Exact {
|
||||
t.Error("expected the exact null distribution for untied samples this small")
|
||||
}
|
||||
if math.Abs(result.PValue-0.2544) > 5e-5 {
|
||||
t.Errorf("two-sided p-value %.6f, want 0.2544", result.PValue)
|
||||
}
|
||||
upper := exactRankSumUpperTail(35, len(chorioamnionTerm), len(chorioamnionEarly))
|
||||
if math.Abs(upper-0.1272) > 5e-5 {
|
||||
t.Errorf("one-sided p-value %.6f, want 0.1272", upper)
|
||||
}
|
||||
}
|
||||
|
||||
// A12 is P(X > Y) + 0.5 P(X = Y), which is U/(mn). With the published W = 35
|
||||
// and sample sizes 10 and 5 the effect size is 35/50 = 0.70. The expected value
|
||||
// is therefore the published Mann-Whitney statistic combined with the published
|
||||
// definition in Vargha and Delaney (2000), not a number this tool produced.
|
||||
func TestVargaDelaneyA12_MatchesPublishedChorioamnionStatistic(t *testing.T) {
|
||||
got := vargaDelaneyA12(chorioamnionTerm, chorioamnionEarly)
|
||||
if math.Abs(got-0.70) > 1e-12 {
|
||||
t.Errorf("A12 = %v, want 0.70", got)
|
||||
}
|
||||
if reversed := vargaDelaneyA12(chorioamnionEarly, chorioamnionTerm); math.Abs(reversed-0.30) > 1e-12 {
|
||||
t.Errorf("reversed A12 = %v, want 0.30", reversed)
|
||||
}
|
||||
}
|
||||
|
||||
// Identical samples are stochastically equal, which Vargha and Delaney define
|
||||
// as A = 0.5, and complete separation gives 1 and 0.
|
||||
func TestVargaDelaneyA12_BoundaryCases(t *testing.T) {
|
||||
same := []float64{1, 2, 3, 4}
|
||||
if got := vargaDelaneyA12(same, same); got != 0.5 {
|
||||
t.Errorf("A12 of a sample against itself = %v, want 0.5", got)
|
||||
}
|
||||
if got := vargaDelaneyA12([]float64{5, 6, 7}, []float64{1, 2}); got != 1 {
|
||||
t.Errorf("A12 with complete dominance = %v, want 1", got)
|
||||
}
|
||||
if got := vargaDelaneyA12([]float64{1, 2}, []float64{5, 6, 7}); got != 0 {
|
||||
t.Errorf("A12 with complete subordination = %v, want 0", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The counting definition and the rank-sum route must agree, including when the
|
||||
// samples are tied against each other, which is the case the evaluation data is
|
||||
// usually in because censored runs pile up on the step they stopped at.
|
||||
func TestVargaDelaneyA12_AgreesWithRankSumStatistic(t *testing.T) {
|
||||
cases := [][2][]float64{
|
||||
{{1, 2, 3}, {2, 3, 4}},
|
||||
{{40, 40, 40, 12}, {40, 7, 3}},
|
||||
{{5}, {5, 5, 5}},
|
||||
{{9, 9, 9}, {9, 9, 9}},
|
||||
}
|
||||
for _, test := range cases {
|
||||
result := rankSum(test[0], test[1])
|
||||
expected := result.Statistic / float64(len(test[0])*len(test[1]))
|
||||
if math.Abs(result.A12-expected) > 1e-12 {
|
||||
t.Errorf("A12 %v for %v vs %v, want U/(mn) = %v", result.A12, test[0], test[1], expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// The tie-corrected variance is checked against the exact permutation variance
|
||||
// of the statistic, computed here by enumerating every way to split the pooled
|
||||
// midranks. That is an independent calculation, not a second call into the
|
||||
// implementation under test.
|
||||
func TestRankSumVariance_MatchesExactPermutationVariance(t *testing.T) {
|
||||
cases := [][]float64{
|
||||
{1, 2, 3, 4, 5, 6, 7, 8},
|
||||
{40, 40, 40, 40, 12, 7, 3, 3},
|
||||
{5, 5, 5, 5, 5, 5, 5, 9},
|
||||
{2, 2, 3, 3, 3, 4, 9, 9, 9},
|
||||
}
|
||||
for _, pooled := range cases {
|
||||
firstSize := len(pooled) / 2
|
||||
ranks, tieGroups := midRanks(pooled)
|
||||
mean, variance := permutationMomentsOfRankSum(ranks, firstSize)
|
||||
expectedMean := float64(firstSize*(len(pooled)-firstSize)) / 2
|
||||
if math.Abs(mean-expectedMean) > 1e-9 {
|
||||
t.Errorf("%v: permutation mean %v, want %v", pooled, mean, expectedMean)
|
||||
}
|
||||
got := rankSumVariance(firstSize, len(pooled)-firstSize, tieGroups)
|
||||
if math.Abs(got-variance) > 1e-9 {
|
||||
t.Errorf("%v: tie-corrected variance %v, want the permutation variance %v", pooled, got, variance)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// permutationMomentsOfRankSum enumerates every subset of the given size and
|
||||
// returns the mean and variance of the Mann-Whitney statistic over them.
|
||||
func permutationMomentsOfRankSum(ranks []float64, firstSize int) (float64, float64) {
|
||||
offset := float64(firstSize) * float64(firstSize+1) / 2
|
||||
var values []float64
|
||||
chosen := make([]int, 0, firstSize)
|
||||
var walk func(start int)
|
||||
walk = func(start int) {
|
||||
if len(chosen) == firstSize {
|
||||
total := 0.0
|
||||
for _, index := range chosen {
|
||||
total += ranks[index]
|
||||
}
|
||||
values = append(values, total-offset)
|
||||
return
|
||||
}
|
||||
for index := start; index < len(ranks); index++ {
|
||||
chosen = append(chosen, index)
|
||||
walk(index + 1)
|
||||
chosen = chosen[:len(chosen)-1]
|
||||
}
|
||||
}
|
||||
walk(0)
|
||||
mean := 0.0
|
||||
for _, value := range values {
|
||||
mean += value
|
||||
}
|
||||
mean /= float64(len(values))
|
||||
variance := 0.0
|
||||
for _, value := range values {
|
||||
variance += (value - mean) * (value - mean)
|
||||
}
|
||||
return mean, variance / float64(len(values))
|
||||
}
|
||||
|
||||
func TestMidRanks_AveragesTiedGroups(t *testing.T) {
|
||||
ranks, tieGroups := midRanks([]float64{3, 1, 3, 2, 3})
|
||||
expected := []float64{4, 1, 4, 2, 4}
|
||||
for index, want := range expected {
|
||||
if ranks[index] != want {
|
||||
t.Errorf("rank %d = %v, want %v", index, ranks[index], want)
|
||||
}
|
||||
}
|
||||
if len(tieGroups) != 1 || tieGroups[0] != 3 {
|
||||
t.Errorf("tie groups %v, want [3]", tieGroups)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRankSum_TiedSamplesUseTheNormalApproximation(t *testing.T) {
|
||||
result := rankSum([]float64{1, 2, 3, 4}, []float64{3, 4, 5, 6})
|
||||
if result.Exact {
|
||||
t.Error("used the exact null distribution despite ties")
|
||||
}
|
||||
if math.IsNaN(result.PValue) || result.PValue < 0 || result.PValue > 1 {
|
||||
t.Errorf("p-value %v", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// Every observation identical carries no information, and the test must say so
|
||||
// rather than dividing by a zero variance.
|
||||
func TestRankSum_AllValuesIdentical(t *testing.T) {
|
||||
result := rankSum([]float64{40, 40, 40}, []float64{40, 40, 40, 40})
|
||||
if result.PValue != 1 {
|
||||
t.Errorf("p-value %v, want 1", result.PValue)
|
||||
}
|
||||
if result.A12 != 0.5 {
|
||||
t.Errorf("A12 %v, want 0.5", result.A12)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRankSum_SingleObservationPerSample(t *testing.T) {
|
||||
result := rankSum([]float64{3}, []float64{9})
|
||||
if result.Statistic != 0 {
|
||||
t.Errorf("statistic %v, want 0", result.Statistic)
|
||||
}
|
||||
if result.A12 != 0 {
|
||||
t.Errorf("A12 %v, want 0", result.A12)
|
||||
}
|
||||
if math.IsNaN(result.PValue) || result.PValue > 1 {
|
||||
t.Errorf("p-value %v", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRankSum_EmptySampleHasNoStatistic(t *testing.T) {
|
||||
result := rankSum(nil, []float64{1, 2, 3})
|
||||
if !math.IsNaN(result.PValue) || !math.IsNaN(result.A12) {
|
||||
t.Errorf("result %+v, want everything undefined", result)
|
||||
}
|
||||
}
|
||||
|
||||
// The exact null distribution must be a proper distribution: the counts sum to
|
||||
// the binomial coefficient and the distribution is symmetric about mn/2.
|
||||
func TestExactRankSumCounts_FormAProperSymmetricDistribution(t *testing.T) {
|
||||
counts := exactRankSumCounts(4, 6)
|
||||
total := 0.0
|
||||
for _, count := range counts {
|
||||
total += count
|
||||
}
|
||||
if total != 210 {
|
||||
t.Errorf("counts sum to %v, want C(10,4) = 210", total)
|
||||
}
|
||||
for index := range counts {
|
||||
mirrored := counts[len(counts)-1-index]
|
||||
if counts[index] != mirrored {
|
||||
t.Errorf("count at %d is %v but %v at the mirrored point", index, counts[index], mirrored)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,228 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"io"
|
||||
"maps"
|
||||
"math"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"text/tabwriter"
|
||||
)
|
||||
|
||||
func writeReport(result analysis, out io.Writer) {
|
||||
fmt.Fprintf(out, "primary outcome: %s\n\n", result.Outcome)
|
||||
|
||||
writeTable(out, []string{"arm", "runs", "violated", "censored", "excluded", "missing", "median steps", "iqr steps", "violation rate"},
|
||||
func(add func(...string)) {
|
||||
for _, summary := range result.Arms {
|
||||
add(
|
||||
summary.Arm,
|
||||
strconv.Itoa(summary.Usable),
|
||||
strconv.Itoa(summary.Violated),
|
||||
strconv.Itoa(summary.Censored),
|
||||
strconv.Itoa(summary.Excluded),
|
||||
strconv.Itoa(len(summary.MissingSeeds)),
|
||||
formatMedian(summary.MedianStepsToFirstViolation),
|
||||
formatMedian(summary.FirstQuartileSteps)+" to "+formatMedian(summary.ThirdQuartileSteps),
|
||||
formatRatio(summary.ViolationRate, 3),
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
fmt.Fprintln(out)
|
||||
fmt.Fprintln(out, "a detection is one distinct property violated in one run; run hours sum the time the runs worked,")
|
||||
fmt.Fprintln(out, "on the monotonic clock, so a host that slept mid-run is not charged for the sleep")
|
||||
fmt.Fprintln(out, "actions count the steps the generator dispatched one on; the rest chose nothing, had the choice")
|
||||
fmt.Fprintln(out, "thrown away, or were the spec's setup driving the app into position before the generator ran")
|
||||
writeTable(out, []string{"arm", "steps", "actions", "run hours", "detections", "defects/1k actions", "defects/hour", "distinct defects", "found in one run"},
|
||||
func(add func(...string)) {
|
||||
for _, summary := range result.Arms {
|
||||
add(
|
||||
summary.Arm,
|
||||
strconv.Itoa(summary.TotalSteps),
|
||||
formatActions(summary),
|
||||
fmt.Sprintf("%.2f", summary.TotalRunHours),
|
||||
strconv.Itoa(summary.Detections),
|
||||
formatRatio(summary.DefectsPerThousandActions, 2),
|
||||
formatRatio(summary.DefectsPerHour, 2),
|
||||
strconv.Itoa(summary.DistinctDefects),
|
||||
formatSingletons(summary),
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
for _, summary := range result.Arms {
|
||||
if len(summary.ExcludedByReason) == 0 {
|
||||
continue
|
||||
}
|
||||
var parts []string
|
||||
for _, reason := range sortedKeys(summary.ExcludedByReason) {
|
||||
parts = append(parts, fmt.Sprintf("%s=%d", reason, summary.ExcludedByReason[reason]))
|
||||
}
|
||||
fmt.Fprintf(out, "\n%s excluded %d run(s) as missing data, not as censored observations: %s",
|
||||
summary.Arm, summary.Excluded, strings.Join(parts, ", "))
|
||||
}
|
||||
for _, summary := range result.Arms {
|
||||
if summary.UnattributedActions > 0 {
|
||||
fmt.Fprintf(out, "\n%s counts %d action(s) of unknown provenance, recorded before an action named its producer: "+
|
||||
"its per-action denominator may include the login the spec's setup drove",
|
||||
summary.Arm, summary.UnattributedActions)
|
||||
}
|
||||
}
|
||||
for _, summary := range result.Arms {
|
||||
if summary.EventsHeldAtBudget > 0 {
|
||||
fmt.Fprintf(out, "\n%s held %d violation(s) reported past the budget at %d steps",
|
||||
summary.Arm, summary.EventsHeldAtBudget, summary.StepBudget)
|
||||
}
|
||||
}
|
||||
for _, summary := range result.Arms {
|
||||
if summary.EventsDetectedAfterOrigin > 0 {
|
||||
fmt.Fprintf(out, "\n%s timed %d violation(s) at the step they were detected rather than the step that armed them, "+
|
||||
"which is what an obligation reported only when the run ended looks like",
|
||||
summary.Arm, summary.EventsDetectedAfterOrigin)
|
||||
}
|
||||
}
|
||||
if len(result.Arms) > 0 {
|
||||
fmt.Fprintln(out)
|
||||
}
|
||||
|
||||
if result.LogRank != nil {
|
||||
fmt.Fprintf(out, "\nlog-rank across %d arms: chi-square %.4f on %d df, p %s\n",
|
||||
len(result.LogRank.Groups), result.LogRank.ChiSquare, result.LogRank.DegreesOfFreedom,
|
||||
formatPValue(result.LogRank.PValue))
|
||||
writeTable(out, []string{"arm", "n", "observed", "expected"}, func(add func(...string)) {
|
||||
for index, name := range result.LogRank.Groups {
|
||||
add(name,
|
||||
strconv.Itoa(result.LogRank.Sizes[index]),
|
||||
fmt.Sprintf("%.0f", result.LogRank.Observed[index]),
|
||||
fmt.Sprintf("%.2f", result.LogRank.Expected[index]))
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
if result.Paired != nil {
|
||||
writePaired(out, *result.Paired)
|
||||
}
|
||||
|
||||
if len(result.Pairwise) > 0 {
|
||||
fmt.Fprintln(out, "\npairwise gehan generalized wilcoxon, which reads a censored run as the bound it is")
|
||||
fmt.Fprintln(out, "a12 above 0.5 means the first arm takes more steps to its first violation; u counts the run")
|
||||
fmt.Fprintln(out, "pairs the first arm outlived, and the unordered pairs count as half in both u and a12")
|
||||
writeTable(out, []string{"comparison", "n1", "n2", "u", "a12", "unordered", "p", "holm p"}, func(add func(...string)) {
|
||||
for _, pair := range result.Pairwise {
|
||||
add(
|
||||
pair.First+" vs "+pair.Second,
|
||||
strconv.Itoa(pair.FirstSize),
|
||||
strconv.Itoa(pair.SecondSize),
|
||||
fmt.Sprintf("%.1f", pair.Statistic),
|
||||
fmt.Sprintf("%.3f", pair.A12),
|
||||
fmt.Sprintf("%d of %d", pair.Unordered, pair.FirstSize*pair.SecondSize),
|
||||
formatPValue(pair.PValue),
|
||||
formatPValue(pair.HolmPValue),
|
||||
)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
if result.HolmFamilySize > 0 {
|
||||
family := "this invocation"
|
||||
if result.Question != "" {
|
||||
family = result.Question
|
||||
}
|
||||
fmt.Fprintf(out, "\nholm correction applied within %s, over %d comparison(s)\n", family, result.HolmFamilySize)
|
||||
}
|
||||
|
||||
for _, note := range result.Notes {
|
||||
fmt.Fprintf(out, "\nnote: %s\n", note)
|
||||
}
|
||||
}
|
||||
|
||||
func writePaired(out io.Writer, comparison pairedComparison) {
|
||||
fmt.Fprintf(out, "\npaired per-seed contrast, %s against %s, each pair scored by which run outlived the other\n",
|
||||
comparison.First, comparison.Second)
|
||||
fmt.Fprintf(out, "%d seed pair(s): %s sooner in %d, %s sooner in %d, left in no order by censoring in %d\n",
|
||||
comparison.Pairs, comparison.First, comparison.FirstSooner,
|
||||
comparison.Second, comparison.SecondSooner, comparison.Unordered)
|
||||
fmt.Fprintf(out, "median difference %s over the %d pair(s) where both runs violated, sign %+d, a12 within pairs %.3f\n",
|
||||
formatStepDifference(comparison.MedianDifference), comparison.BothViolated, comparison.Sign, comparison.A12)
|
||||
fmt.Fprintf(out, "sign test over the %d ordered pair(s), p %s, holm p %s\n",
|
||||
comparison.FirstSooner+comparison.SecondSooner,
|
||||
formatOptionalPValue(comparison.PValue), formatOptionalPValue(comparison.HolmPValue))
|
||||
if len(comparison.UnpairedSeeds) > 0 {
|
||||
fmt.Fprintf(out, "%d seed(s) usable in one arm only and left out of the pairing: %v\n",
|
||||
len(comparison.UnpairedSeeds), comparison.UnpairedSeeds)
|
||||
}
|
||||
}
|
||||
|
||||
func sortedKeys(counts map[string]int) []string {
|
||||
return slices.Sorted(maps.Keys(counts))
|
||||
}
|
||||
|
||||
func writeTable(out io.Writer, header []string, rows func(add func(...string))) {
|
||||
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
|
||||
fmt.Fprintln(writer, strings.Join(header, "\t"))
|
||||
rows(func(cells ...string) {
|
||||
fmt.Fprintln(writer, strings.Join(cells, "\t"))
|
||||
})
|
||||
writer.Flush()
|
||||
}
|
||||
|
||||
// formatMedian says undefined rather than substituting a mean, because a curve
|
||||
// that never reaches one half has no median to report.
|
||||
func formatMedian(value *float64) string {
|
||||
if value == nil {
|
||||
return "undefined"
|
||||
}
|
||||
return strconv.FormatFloat(*value, 'f', -1, 64)
|
||||
}
|
||||
|
||||
func formatStepDifference(value *float64) string {
|
||||
if value == nil {
|
||||
return "undefined"
|
||||
}
|
||||
return fmt.Sprintf("%+.1f steps", *value)
|
||||
}
|
||||
|
||||
// formatActions marks a denominator with actions whose producer nothing names,
|
||||
// because the rate beside it then divides by a count that may include the
|
||||
// login the spec's setup drove.
|
||||
func formatActions(summary armSummary) string {
|
||||
if summary.UnattributedActions == 0 {
|
||||
return strconv.Itoa(summary.TotalActions)
|
||||
}
|
||||
return fmt.Sprintf("%d (%d unattributed)", summary.TotalActions, summary.UnattributedActions)
|
||||
}
|
||||
|
||||
func formatRatio(value *float64, digits int) string {
|
||||
if value == nil {
|
||||
return "n/a"
|
||||
}
|
||||
return strconv.FormatFloat(*value, 'f', digits, 64)
|
||||
}
|
||||
|
||||
func formatSingletons(summary armSummary) string {
|
||||
if summary.SingletonFraction == nil {
|
||||
return "n/a"
|
||||
}
|
||||
return fmt.Sprintf("%d/%d (%.3f)", summary.SingletonDefects, summary.DistinctDefects, *summary.SingletonFraction)
|
||||
}
|
||||
|
||||
func formatOptionalPValue(value *float64) string {
|
||||
if value == nil {
|
||||
return "n/a"
|
||||
}
|
||||
return formatPValue(*value)
|
||||
}
|
||||
|
||||
func formatPValue(value float64) string {
|
||||
switch {
|
||||
case math.IsNaN(value):
|
||||
return "n/a"
|
||||
case value < 1e-4:
|
||||
return fmt.Sprintf("%.3e", value)
|
||||
default:
|
||||
return fmt.Sprintf("%.4f", value)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,247 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"slices"
|
||||
)
|
||||
|
||||
// observation is one run reduced to what the survival analysis needs: the step
|
||||
// count at which it left the risk set, and whether it left because a violation
|
||||
// was found (an event) or because the run ended without one (right-censored). A run that failed or timed out is neither and never
|
||||
// reaches this type.
|
||||
type observation struct {
|
||||
Steps float64
|
||||
Event bool
|
||||
}
|
||||
|
||||
type survivalPoint struct {
|
||||
Steps float64 `json:"steps"`
|
||||
AtRisk int `json:"at_risk"`
|
||||
Events int `json:"events"`
|
||||
Censored int `json:"censored"`
|
||||
Survival float64 `json:"survival"`
|
||||
}
|
||||
|
||||
// kaplanMeier is the product-limit estimate of Kaplan and Meier (1958), one row
|
||||
// per distinct observed step count. Runs censored at a step count tied with an
|
||||
// event are counted in the risk set for that event, which is the standard
|
||||
// convention.
|
||||
func kaplanMeier(observations []observation) []survivalPoint {
|
||||
if len(observations) == 0 {
|
||||
return nil
|
||||
}
|
||||
remaining := len(observations)
|
||||
survival := 1.0
|
||||
var curve []survivalPoint
|
||||
for _, steps := range distinctSteps(observations) {
|
||||
events, censored := 0, 0
|
||||
for _, item := range observations {
|
||||
if item.Steps != steps {
|
||||
continue
|
||||
}
|
||||
if item.Event {
|
||||
events++
|
||||
} else {
|
||||
censored++
|
||||
}
|
||||
}
|
||||
atRisk := remaining
|
||||
if events > 0 {
|
||||
survival *= 1 - float64(events)/float64(atRisk)
|
||||
}
|
||||
curve = append(curve, survivalPoint{
|
||||
Steps: steps,
|
||||
AtRisk: atRisk,
|
||||
Events: events,
|
||||
Censored: censored,
|
||||
Survival: survival,
|
||||
})
|
||||
remaining -= events + censored
|
||||
}
|
||||
return curve
|
||||
}
|
||||
|
||||
// medianSurvival is the smallest step count at which the estimate falls to or
|
||||
// below one half. It is undefined whenever fewer than half the runs violate,
|
||||
// and the second return value says so: substituting a mean there would report a
|
||||
// number the data does not contain.
|
||||
func medianSurvival(curve []survivalPoint) (float64, bool) {
|
||||
return quantileSurvival(curve, 0.5)
|
||||
}
|
||||
|
||||
func distinctSteps(observations []observation) []float64 {
|
||||
steps := make([]float64, 0, len(observations))
|
||||
for _, item := range observations {
|
||||
steps = append(steps, item.Steps)
|
||||
}
|
||||
slices.Sort(steps)
|
||||
return slices.Compact(steps)
|
||||
}
|
||||
|
||||
type logRankResult struct {
|
||||
Groups []string `json:"groups"`
|
||||
Sizes []int `json:"sizes"`
|
||||
Observed []float64 `json:"observed"`
|
||||
Expected []float64 `json:"expected"`
|
||||
ChiSquare float64 `json:"chi_square"`
|
||||
DegreesOfFreedom int `json:"degrees_of_freedom"`
|
||||
PValue float64 `json:"p_value"`
|
||||
}
|
||||
|
||||
// logRank is the k-sample Mantel-Haenszel log-rank test, the member of the
|
||||
// weighted family that counts every event time alike. Mantel (1966); Peto and
|
||||
// Peto (1972).
|
||||
func logRank(names []string, groups [][]observation) logRankResult {
|
||||
return weightedLogRank(names, groups, func(atRisk float64) float64 { return 1 })
|
||||
}
|
||||
|
||||
// weightedLogRank is the family the log-rank belongs to. At every distinct event
|
||||
// time it contrasts observed with expected events under the null of equal
|
||||
// hazards, weights that difference by weight(atRisk), and combines the k-1
|
||||
// independent weighted differences through their covariance matrix:
|
||||
// chi-square = U' V^-1 U on k-1 degrees of freedom. Observed and Expected stay
|
||||
// event counts whatever the weight, because a weighted count is not one.
|
||||
// Klein and Moeschberger, Survival Analysis, 2nd ed., section 7.3.
|
||||
func weightedLogRank(names []string, groups [][]observation, weight func(atRisk float64) float64) logRankResult {
|
||||
var keptNames []string
|
||||
var kept [][]observation
|
||||
for index, group := range groups {
|
||||
if len(group) == 0 {
|
||||
continue
|
||||
}
|
||||
keptNames = append(keptNames, names[index])
|
||||
kept = append(kept, group)
|
||||
}
|
||||
names, groups = keptNames, kept
|
||||
|
||||
count := len(groups)
|
||||
result := logRankResult{
|
||||
Groups: names,
|
||||
Sizes: make([]int, count),
|
||||
Observed: make([]float64, count),
|
||||
Expected: make([]float64, count),
|
||||
DegreesOfFreedom: count - 1,
|
||||
PValue: math.NaN(),
|
||||
}
|
||||
if count < 2 {
|
||||
return result
|
||||
}
|
||||
var pooled []observation
|
||||
for index, group := range groups {
|
||||
result.Sizes[index] = len(group)
|
||||
pooled = append(pooled, group...)
|
||||
}
|
||||
|
||||
covariance := make([][]float64, count)
|
||||
for index := range covariance {
|
||||
covariance[index] = make([]float64, count)
|
||||
}
|
||||
atRisk := make([]float64, count)
|
||||
deaths := make([]float64, count)
|
||||
weightedDifference := make([]float64, count)
|
||||
|
||||
for _, steps := range distinctSteps(pooled) {
|
||||
totalAtRisk, totalDeaths := 0.0, 0.0
|
||||
for index, group := range groups {
|
||||
atRisk[index], deaths[index] = 0, 0
|
||||
for _, item := range group {
|
||||
if item.Steps >= steps {
|
||||
atRisk[index]++
|
||||
}
|
||||
if item.Steps == steps && item.Event {
|
||||
deaths[index]++
|
||||
}
|
||||
}
|
||||
totalAtRisk += atRisk[index]
|
||||
totalDeaths += deaths[index]
|
||||
}
|
||||
if totalDeaths == 0 {
|
||||
continue
|
||||
}
|
||||
weightAtStep := weight(totalAtRisk)
|
||||
for index := range groups {
|
||||
expected := totalDeaths * atRisk[index] / totalAtRisk
|
||||
result.Observed[index] += deaths[index]
|
||||
result.Expected[index] += expected
|
||||
weightedDifference[index] += weightAtStep * (deaths[index] - expected)
|
||||
}
|
||||
if totalAtRisk <= 1 {
|
||||
continue
|
||||
}
|
||||
scale := weightAtStep * weightAtStep * totalDeaths * (totalAtRisk - totalDeaths) / (totalAtRisk - 1)
|
||||
for row := range groups {
|
||||
share := atRisk[row] / totalAtRisk
|
||||
covariance[row][row] += scale * share * (1 - share)
|
||||
for column := range groups {
|
||||
if column == row {
|
||||
continue
|
||||
}
|
||||
covariance[row][column] -= scale * share * atRisk[column] / totalAtRisk
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
reduced := make([][]float64, count-1)
|
||||
difference := make([]float64, count-1)
|
||||
for row := 0; row < count-1; row++ {
|
||||
reduced[row] = make([]float64, count-1)
|
||||
copy(reduced[row], covariance[row][:count-1])
|
||||
difference[row] = weightedDifference[row]
|
||||
}
|
||||
solution, ok := solveLinearSystem(reduced, difference)
|
||||
if !ok {
|
||||
result.ChiSquare = 0
|
||||
result.PValue = 1
|
||||
return result
|
||||
}
|
||||
statistic := 0.0
|
||||
for index := range difference {
|
||||
statistic += difference[index] * solution[index]
|
||||
}
|
||||
if statistic < 0 || math.IsNaN(statistic) {
|
||||
statistic = 0
|
||||
}
|
||||
result.ChiSquare = statistic
|
||||
result.PValue = chiSquareUpperTail(statistic, result.DegreesOfFreedom)
|
||||
return result
|
||||
}
|
||||
|
||||
// solveLinearSystem solves matrix*x = vector by Gaussian elimination with
|
||||
// partial pivoting, reporting failure rather than a value when the matrix is
|
||||
// singular, which is what a group with no events at all produces.
|
||||
func solveLinearSystem(matrix [][]float64, vector []float64) ([]float64, bool) {
|
||||
size := len(vector)
|
||||
work := make([][]float64, size)
|
||||
for row := range work {
|
||||
work[row] = make([]float64, size+1)
|
||||
copy(work[row], matrix[row])
|
||||
work[row][size] = vector[row]
|
||||
}
|
||||
for column := 0; column < size; column++ {
|
||||
pivot := column
|
||||
for row := column + 1; row < size; row++ {
|
||||
if math.Abs(work[row][column]) > math.Abs(work[pivot][column]) {
|
||||
pivot = row
|
||||
}
|
||||
}
|
||||
if math.Abs(work[pivot][column]) < 1e-12 {
|
||||
return nil, false
|
||||
}
|
||||
work[column], work[pivot] = work[pivot], work[column]
|
||||
for row := column + 1; row < size; row++ {
|
||||
factor := work[row][column] / work[column][column]
|
||||
for next := column; next <= size; next++ {
|
||||
work[row][next] -= factor * work[column][next]
|
||||
}
|
||||
}
|
||||
}
|
||||
solution := make([]float64, size)
|
||||
for row := size - 1; row >= 0; row-- {
|
||||
total := work[row][size]
|
||||
for column := row + 1; column < size; column++ {
|
||||
total -= work[row][column] * solution[column]
|
||||
}
|
||||
solution[row] = total / work[row][row]
|
||||
}
|
||||
return solution, true
|
||||
}
|
||||
@@ -0,0 +1,323 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"math"
|
||||
"math/rand"
|
||||
"slices"
|
||||
"strconv"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// The product-limit estimates for the 6-MP arm of Freireich et al. (1963) are
|
||||
// the worked example reproduced in Collett, Modelling Survival Data in Medical
|
||||
// Research, and in the standard course treatments of the gehan data:
|
||||
//
|
||||
// t: 6 7 10 13 16 22 23
|
||||
// S(t): 0.857 0.807 0.753 0.690 0.627 0.538 0.448
|
||||
func TestKaplanMeier_MatchesPublishedGehanEstimates(t *testing.T) {
|
||||
curve := kaplanMeier(gehanSixMercaptopurine)
|
||||
expected := map[float64]float64{
|
||||
6: 0.857, 7: 0.807, 10: 0.753, 13: 0.690, 16: 0.627, 22: 0.538, 23: 0.448,
|
||||
}
|
||||
seen := 0
|
||||
for _, point := range curve {
|
||||
want, ok := expected[point.Steps]
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
seen++
|
||||
if math.Abs(point.Survival-want) > 5e-4 {
|
||||
t.Errorf("S(%v) = %.4f, want %v", point.Steps, point.Survival, want)
|
||||
}
|
||||
}
|
||||
if seen != len(expected) {
|
||||
t.Fatalf("matched %d of %d published times", seen, len(expected))
|
||||
}
|
||||
}
|
||||
|
||||
// The risk set at each time is the count of runs still under observation, with
|
||||
// runs censored at a tied time counted as at risk for that event.
|
||||
func TestKaplanMeier_RiskSetHandlesTiesAndCensoring(t *testing.T) {
|
||||
curve := kaplanMeier(gehanSixMercaptopurine)
|
||||
expected := map[float64]struct {
|
||||
atRisk int
|
||||
events int
|
||||
censored int
|
||||
}{
|
||||
6: {21, 3, 1},
|
||||
7: {17, 1, 0},
|
||||
9: {16, 0, 1},
|
||||
10: {15, 1, 1},
|
||||
13: {12, 1, 0},
|
||||
23: {6, 1, 0},
|
||||
}
|
||||
for _, point := range curve {
|
||||
want, ok := expected[point.Steps]
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if point.AtRisk != want.atRisk || point.Events != want.events || point.Censored != want.censored {
|
||||
t.Errorf("at %v: risk=%d events=%d censored=%d, want risk=%d events=%d censored=%d",
|
||||
point.Steps, point.AtRisk, point.Events, point.Censored, want.atRisk, want.events, want.censored)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Published medians: 23 weeks for 6-MP against 8 weeks for placebo (Gehan and
|
||||
// Freireich, "The 6-MP versus placebo clinical trial in acute leukemia",
|
||||
// Clinical Trials 8(3), 2011), and 31 against 23 weeks for the aml arms as
|
||||
// reported by survfit in R's survival package.
|
||||
func TestMedianSurvival_MatchesPublishedMedians(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
observations []observation
|
||||
expected float64
|
||||
}{
|
||||
{"gehan 6-MP", gehanSixMercaptopurine, 23},
|
||||
{"gehan placebo", gehanPlacebo, 8},
|
||||
{"aml maintained", amlMaintained, 31},
|
||||
{"aml nonmaintained", amlNonmaintained, 23},
|
||||
}
|
||||
for _, test := range cases {
|
||||
median, ok := medianSurvival(kaplanMeier(test.observations))
|
||||
if !ok {
|
||||
t.Errorf("%s: median undefined, want %v", test.name, test.expected)
|
||||
continue
|
||||
}
|
||||
if median != test.expected {
|
||||
t.Errorf("%s: median %v, want %v", test.name, median, test.expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// R's survival package documents this log-rank on the aml data:
|
||||
//
|
||||
// N Observed Expected (O-E)^2/E (O-E)^2/V
|
||||
// x=Maintained 11 7 10.69 1.27 3.4
|
||||
// x=Nonmaintained 12 11 7.31 1.86 3.4
|
||||
// Chisq= 3.4 on 1 degrees of freedom, p= 0.0653
|
||||
func TestLogRank_MatchesPublishedAmlResult(t *testing.T) {
|
||||
result := logRank([]string{"maintained", "nonmaintained"}, [][]observation{amlMaintained, amlNonmaintained})
|
||||
|
||||
if result.Observed[0] != 7 || result.Observed[1] != 11 {
|
||||
t.Errorf("observed %v, want [7 11]", result.Observed)
|
||||
}
|
||||
if math.Abs(result.Expected[0]-10.69) > 5e-3 || math.Abs(result.Expected[1]-7.31) > 5e-3 {
|
||||
t.Errorf("expected %v, want [10.69 7.31]", result.Expected)
|
||||
}
|
||||
for index, want := range []float64{1.27, 1.86} {
|
||||
difference := result.Observed[index] - result.Expected[index]
|
||||
got := difference * difference / result.Expected[index]
|
||||
if math.Abs(got-want) > 5e-3 {
|
||||
t.Errorf("(O-E)^2/E for group %d = %.4f, want %v", index, got, want)
|
||||
}
|
||||
}
|
||||
if math.Abs(result.ChiSquare-3.4) > 5e-2 {
|
||||
t.Errorf("chi-square %.4f, want 3.4", result.ChiSquare)
|
||||
}
|
||||
if math.Abs(result.PValue-0.0653) > 5e-4 {
|
||||
t.Errorf("p-value %.6f, want 0.0653", result.PValue)
|
||||
}
|
||||
if result.DegreesOfFreedom != 1 {
|
||||
t.Errorf("degrees of freedom %d, want 1", result.DegreesOfFreedom)
|
||||
}
|
||||
}
|
||||
|
||||
// The log-rank on the Freireich 6-MP trial is the textbook worked example:
|
||||
// observed 9 against 19.25 expected in the treated arm and 21 against 10.75 in
|
||||
// the control arm, Mantel-Haenszel chi-square 16.79 on 1 degree of freedom,
|
||||
// p = 4.17e-05. Reported for instance in Rodriguez, Kaplan-Meier and
|
||||
// Mantel-Haenszel, https://grodri.github.io/survival/gehan, and in Collett.
|
||||
func TestLogRank_MatchesPublishedGehanResult(t *testing.T) {
|
||||
result := logRank([]string{"6-MP", "placebo"}, [][]observation{gehanSixMercaptopurine, gehanPlacebo})
|
||||
|
||||
if result.Observed[0] != 9 || result.Observed[1] != 21 {
|
||||
t.Errorf("observed %v, want [9 21]", result.Observed)
|
||||
}
|
||||
if math.Abs(result.Expected[0]-19.25) > 5e-3 || math.Abs(result.Expected[1]-10.75) > 5e-3 {
|
||||
t.Errorf("expected %v, want [19.25 10.75]", result.Expected)
|
||||
}
|
||||
if math.Abs(result.ChiSquare-16.79) > 5e-3 {
|
||||
t.Errorf("chi-square %.4f, want 16.79", result.ChiSquare)
|
||||
}
|
||||
if math.Abs(result.PValue-4.17e-5) > 5e-8 {
|
||||
t.Errorf("p-value %.3e, want 4.17e-05", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// An arm with no usable runs contributes nothing and must not consume a degree
|
||||
// of freedom or make the covariance matrix singular.
|
||||
func TestLogRank_EmptyGroupIsDropped(t *testing.T) {
|
||||
two := logRank([]string{"a", "b"}, [][]observation{amlMaintained, amlNonmaintained})
|
||||
three := logRank([]string{"a", "b", "c"}, [][]observation{amlMaintained, amlNonmaintained, nil})
|
||||
if math.Abs(two.ChiSquare-three.ChiSquare) > 1e-12 {
|
||||
t.Errorf("chi-square %.10f with an empty third group, want %.10f", three.ChiSquare, two.ChiSquare)
|
||||
}
|
||||
if three.DegreesOfFreedom != 1 {
|
||||
t.Errorf("degrees of freedom %d, want 1", three.DegreesOfFreedom)
|
||||
}
|
||||
if len(three.Groups) != 2 {
|
||||
t.Errorf("groups %v, want the empty arm dropped", three.Groups)
|
||||
}
|
||||
}
|
||||
|
||||
// No published multi-arm dataset with a printed log-rank chi-square was found
|
||||
// small enough to embed, so the k-group covariance algebra is checked against
|
||||
// its own null distribution instead: under the null of equal hazards the
|
||||
// statistic is asymptotically chi-square on k-1 degrees of freedom, so its mean
|
||||
// over random relabellings has to sit near k-1. A wrong variance term or a wrong
|
||||
// degrees-of-freedom count moves this badly.
|
||||
func TestLogRank_NullMeanTracksDegreesOfFreedom(t *testing.T) {
|
||||
pooled := make([]observation, 0, 36)
|
||||
for index := 0; index < 36; index++ {
|
||||
pooled = append(pooled, observation{Steps: float64(index%17 + 1), Event: index%5 != 0})
|
||||
}
|
||||
for _, groupCount := range []int{2, 3, 4} {
|
||||
generator := rand.New(rand.NewSource(20260812))
|
||||
total := 0.0
|
||||
const replicates = 4000
|
||||
for replicate := 0; replicate < replicates; replicate++ {
|
||||
shuffled := slices.Clone(pooled)
|
||||
generator.Shuffle(len(shuffled), func(left, right int) {
|
||||
shuffled[left], shuffled[right] = shuffled[right], shuffled[left]
|
||||
})
|
||||
names := make([]string, groupCount)
|
||||
groups := make([][]observation, groupCount)
|
||||
for index, item := range shuffled {
|
||||
groups[index%groupCount] = append(groups[index%groupCount], item)
|
||||
}
|
||||
for index := range names {
|
||||
names[index] = strconv.Itoa(index)
|
||||
}
|
||||
total += logRank(names, groups).ChiSquare
|
||||
}
|
||||
mean := total / replicates
|
||||
expected := float64(groupCount - 1)
|
||||
if math.Abs(mean-expected) > 0.15*expected {
|
||||
t.Errorf("%d groups: null mean chi-square %.3f, want near %v", groupCount, mean, expected)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A three-group split of one homogeneous sample must not look significant, and
|
||||
// the statistic must be finite on 2 degrees of freedom.
|
||||
func TestLogRank_ThreeIdenticalGroupsAreNotSignificant(t *testing.T) {
|
||||
group := []observation{{4, true}, {7, true}, {9, false}, {12, true}, {20, false}}
|
||||
result := logRank([]string{"a", "b", "c"}, [][]observation{group, group, group})
|
||||
if result.DegreesOfFreedom != 2 {
|
||||
t.Fatalf("degrees of freedom %d, want 2", result.DegreesOfFreedom)
|
||||
}
|
||||
if result.ChiSquare > 1e-9 {
|
||||
t.Errorf("chi-square %.10f for three identical groups, want 0", result.ChiSquare)
|
||||
}
|
||||
if math.Abs(result.PValue-1) > 1e-9 {
|
||||
t.Errorf("p-value %v, want 1", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaplanMeier_EveryObservationCensored(t *testing.T) {
|
||||
observations := []observation{{40, false}, {40, false}, {40, false}}
|
||||
curve := kaplanMeier(observations)
|
||||
if len(curve) != 1 {
|
||||
t.Fatalf("curve has %d points, want 1", len(curve))
|
||||
}
|
||||
if curve[0].Survival != 1 || curve[0].Events != 0 || curve[0].Censored != 3 {
|
||||
t.Errorf("point %+v, want survival 1 with 3 censored", curve[0])
|
||||
}
|
||||
if _, ok := medianSurvival(curve); ok {
|
||||
t.Error("median defined for an arm where nothing violated")
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaplanMeier_EveryObservationAnEvent(t *testing.T) {
|
||||
curve := kaplanMeier([]observation{{2, true}, {4, true}, {6, true}, {8, true}})
|
||||
last := curve[len(curve)-1]
|
||||
if last.Survival != 0 {
|
||||
t.Errorf("final survival %v, want 0", last.Survival)
|
||||
}
|
||||
median, ok := medianSurvival(curve)
|
||||
if !ok || median != 4 {
|
||||
t.Errorf("median %v ok=%v, want 4", median, ok)
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaplanMeier_TiedEventTimesDropOnce(t *testing.T) {
|
||||
curve := kaplanMeier([]observation{{5, true}, {5, true}, {5, true}, {9, true}})
|
||||
if len(curve) != 2 {
|
||||
t.Fatalf("curve has %d points, want 2", len(curve))
|
||||
}
|
||||
if curve[0].Events != 3 || math.Abs(curve[0].Survival-0.25) > 1e-12 {
|
||||
t.Errorf("first point %+v, want 3 events and survival 0.25", curve[0])
|
||||
}
|
||||
if curve[1].Survival != 0 {
|
||||
t.Errorf("second point %+v, want survival 0", curve[1])
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaplanMeier_SingleObservation(t *testing.T) {
|
||||
event := kaplanMeier([]observation{{11, true}})
|
||||
if len(event) != 1 || event[0].Survival != 0 {
|
||||
t.Fatalf("single event curve %+v", event)
|
||||
}
|
||||
median, ok := medianSurvival(event)
|
||||
if !ok || median != 11 {
|
||||
t.Errorf("median %v ok=%v, want 11", median, ok)
|
||||
}
|
||||
censored := kaplanMeier([]observation{{11, false}})
|
||||
if len(censored) != 1 || censored[0].Survival != 1 {
|
||||
t.Fatalf("single censored curve %+v", censored)
|
||||
}
|
||||
if _, ok := medianSurvival(censored); ok {
|
||||
t.Error("median defined for a single censored observation")
|
||||
}
|
||||
}
|
||||
|
||||
func TestKaplanMeier_NoObservations(t *testing.T) {
|
||||
if curve := kaplanMeier(nil); curve != nil {
|
||||
t.Errorf("curve %v for no observations, want nil", curve)
|
||||
}
|
||||
if _, ok := medianSurvival(nil); ok {
|
||||
t.Error("median defined for an empty curve")
|
||||
}
|
||||
}
|
||||
|
||||
func TestLogRank_SingleObservationPerGroup(t *testing.T) {
|
||||
result := logRank([]string{"a", "b"}, [][]observation{{{3, true}}, {{9, true}}})
|
||||
if math.IsNaN(result.ChiSquare) || result.ChiSquare < 0 {
|
||||
t.Errorf("chi-square %v", result.ChiSquare)
|
||||
}
|
||||
if math.IsNaN(result.PValue) || result.PValue > 1 || result.PValue < 0 {
|
||||
t.Errorf("p-value %v", result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// With no events anywhere the covariance matrix is singular and there is
|
||||
// nothing to test, which must report no difference rather than a divide by zero.
|
||||
func TestLogRank_NoEventsAnywhere(t *testing.T) {
|
||||
result := logRank([]string{"a", "b"}, [][]observation{
|
||||
{{40, false}, {40, false}},
|
||||
{{40, false}, {40, false}, {40, false}},
|
||||
})
|
||||
if result.ChiSquare != 0 || result.PValue != 1 {
|
||||
t.Errorf("chi-square %v p-value %v, want 0 and 1", result.ChiSquare, result.PValue)
|
||||
}
|
||||
}
|
||||
|
||||
// The weighted family reports event counts, not weighted ones: the report
|
||||
// prints observed against expected as counts of violations, and a weight that
|
||||
// changes the statistic must leave those alone.
|
||||
func TestWeightedLogRank_WeightsTheStatisticAndNotTheCounts(t *testing.T) {
|
||||
names := []string{"6-mp", "placebo"}
|
||||
groups := [][]observation{gehanSixMercaptopurine, gehanPlacebo}
|
||||
plain := logRank(names, groups)
|
||||
weighted := weightedLogRank(names, groups, atRiskWeight)
|
||||
|
||||
if !slices.Equal(plain.Observed, weighted.Observed) || !slices.Equal(plain.Expected, weighted.Expected) {
|
||||
t.Errorf("weighted observed %v expected %v, want the counts %v and %v",
|
||||
weighted.Observed, weighted.Expected, plain.Observed, plain.Expected)
|
||||
}
|
||||
if math.Abs(weighted.ChiSquare-plain.ChiSquare) < 1e-9 {
|
||||
t.Errorf("chi-square %v under Gehan's weight and %v unweighted, want the weight to reach the statistic",
|
||||
weighted.ChiSquare, plain.ChiSquare)
|
||||
}
|
||||
}
|
||||
@@ -1,14 +1,25 @@
|
||||
// Command bundle-check is a developer tool that bundles a spec file to confirm it compiles.
|
||||
// Command bundle-check is a developer tool that bundles a spec file and loads
|
||||
// it into the evaluator to confirm it compiles and registers properties.
|
||||
package main
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/bundler"
|
||||
"github.com/priyanshujain/sanderling/internal/testrun"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
// checkSeed keeps the load deterministic. The bundle it seeds is only loaded,
|
||||
// never hashed or reported, so the value is arbitrary.
|
||||
const checkSeed = 1
|
||||
|
||||
func bundleSpec(specSrc, entryFile string) (bundler.Result, error) {
|
||||
return bundler.Bundle(bundler.Options{
|
||||
EntryFile: entryFile,
|
||||
@@ -20,13 +31,67 @@ func bundleSpec(specSrc, entryFile string) (bundler.Result, error) {
|
||||
})
|
||||
}
|
||||
|
||||
// registeredProperties bundles the spec the way a run bundles it, with the
|
||||
// runtime entry that assigns globalThis.properties, and loads it into the real
|
||||
// evaluator. Bundling alone proves nothing about registration: a spec that
|
||||
// registers no property compiles perfectly and then judges nothing.
|
||||
func registeredProperties(entryFile string) ([]string, error) {
|
||||
bundle, err := testrun.BundleSpec(entryFile, checkSeed)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("bundle with runtime entry: %w", err)
|
||||
}
|
||||
evaluator, err := verifier.New()
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("evaluator: %w", err)
|
||||
}
|
||||
if err := evaluator.Load(string(bundle.JavaScript)); err != nil {
|
||||
return nil, fmt.Errorf("load spec: %w", err)
|
||||
}
|
||||
return evaluator.PropertyNames(), nil
|
||||
}
|
||||
|
||||
func check(specSrc, entryFile string, stdout io.Writer) error {
|
||||
return checkWithOptions(specSrc, entryFile, false, stdout)
|
||||
}
|
||||
|
||||
func checkWithOptions(specSrc, entryFile string, allowNoProperties bool, stdout io.Writer) error {
|
||||
result, err := bundleSpec(specSrc, entryFile)
|
||||
if err != nil {
|
||||
return fmt.Errorf("bundle: %w", err)
|
||||
}
|
||||
fmt.Fprintf(stdout, "bundled: %d bytes, sha256=%s\n", len(result.JavaScript), result.SHA256)
|
||||
|
||||
names, err := registeredProperties(entryFile)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if len(names) == 0 && !allowNoProperties {
|
||||
return errors.New("the spec bundles and loads cleanly but registers no properties: " +
|
||||
"nothing is wrong with the source, and a run against it would check nothing " +
|
||||
"and report no violations. Pass --allow-no-properties for a spec that measures " +
|
||||
"what it extracts or where the generator reaches")
|
||||
}
|
||||
fmt.Fprintf(stdout, "properties registered: %d (%s)\n", len(names), strings.Join(names, ", "))
|
||||
return nil
|
||||
}
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
fmt.Fprintln(os.Stderr, "usage: bundle-check <spec.ts>")
|
||||
flagSet := flag.NewFlagSet("bundle-check", flag.ExitOnError)
|
||||
allowNoProperties := flagSet.Bool("allow-no-properties", false,
|
||||
"accept a spec that registers no properties, for a pre-registration that measures what the spec extracts or where the generator reaches")
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprintln(flagSet.Output(), "usage: bundle-check [--allow-no-properties] <spec.ts>")
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
if err := flagSet.Parse(os.Args[1:]); err != nil {
|
||||
os.Exit(1)
|
||||
}
|
||||
if flagSet.NArg() != 1 {
|
||||
flagSet.Usage()
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
entryFile, err := filepath.Abs(os.Args[1])
|
||||
entryFile, err := filepath.Abs(flagSet.Arg(0))
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "resolve spec path: %v\n", err)
|
||||
os.Exit(1)
|
||||
@@ -38,10 +103,8 @@ func main() {
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
result, err := bundleSpec(filepath.Join(repoRoot, "pkg/spec/src"), entryFile)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "bundle: %v\n", err)
|
||||
if err := checkWithOptions(filepath.Join(repoRoot, "pkg/spec/src"), entryFile, *allowNoProperties, os.Stdout); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "%v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
fmt.Printf("bundled: %d bytes, sha256=%s\n", len(result.JavaScript), result.SHA256)
|
||||
}
|
||||
@@ -1,8 +1,10 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
@@ -55,3 +57,72 @@ func TestBundleSpec_ResolvesSpecAliases(t *testing.T) {
|
||||
t.Errorf("unstable bundle hash: %s vs %s", result.SHA256, second.SHA256)
|
||||
}
|
||||
}
|
||||
|
||||
func testdataSpec(t *testing.T, name string) string {
|
||||
t.Helper()
|
||||
path, err := filepath.Abs(filepath.Join("testdata", name))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
func TestCheck_RejectsSpecThatRegistersNoProperties(t *testing.T) {
|
||||
var stdout bytes.Buffer
|
||||
err := check(repoSpecSrc(t), testdataSpec(t, "no-properties.ts"), &stdout)
|
||||
if err == nil {
|
||||
t.Fatal("a spec registering zero properties passed the gate: a run against it reports no violations while checking nothing")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "registers no properties") {
|
||||
t.Errorf("failure must name the missing registration, got: %v", err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "bundled: ") {
|
||||
t.Errorf("bundle size and hash must still be reported for a spec awaiting its properties, got: %q", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
// The extraction and portability sweeps freeze a spec that registers nothing
|
||||
// on purpose, and run it with --allow-no-properties. Without the same opt-out
|
||||
// here the gate that is supposed to freeze those pre-registrations is the one
|
||||
// thing that cannot accept them.
|
||||
func TestCheck_RunsTheZeroPropertySpecUnderTheOptOut(t *testing.T) {
|
||||
var stdout bytes.Buffer
|
||||
if err := checkWithOptions(repoSpecSrc(t), testdataSpec(t, "no-properties.ts"), true, &stdout); err != nil {
|
||||
t.Fatalf("the opt-out did not admit a spec that registers nothing: %v", err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "properties registered: 0") {
|
||||
t.Errorf("the report must still say nothing was registered, got: %q", stdout.String())
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "bundled: ") {
|
||||
t.Errorf("bundle size and hash must still be reported, got: %q", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestCheck_ReportsRegisteredPropertyCountAndNames(t *testing.T) {
|
||||
var stdout bytes.Buffer
|
||||
if err := check(repoSpecSrc(t), testdataSpec(t, "spec.ts"), &stdout); err != nil {
|
||||
t.Fatalf("check: %v", err)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "properties registered: 1 (noUncaughtExceptions)") {
|
||||
t.Errorf("expected the registered count and name, got: %q", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
// TestCheck_ReportsUnchangedBundleHash pins the reported bundle of a fixture
|
||||
// that imports nothing, so the hash moves only when the bundler itself does.
|
||||
// Frozen pre-registrations record these hashes, so bundle-check reporting a
|
||||
// different bundle (the runtime-entry one, say) silently invalidates them.
|
||||
func TestCheck_ReportsUnchangedBundleHash(t *testing.T) {
|
||||
const (
|
||||
plainBytes = 57
|
||||
plainSHA256 = "bd757084b3f29c04c68ba2aa4c7b93e63b6dc6dd08ea3f726e0d1878be845888"
|
||||
)
|
||||
result, err := bundleSpec(repoSpecSrc(t), testdataSpec(t, "plain.ts"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(result.JavaScript) != plainBytes || result.SHA256 != plainSHA256 {
|
||||
t.Errorf("reported bundle changed: %d bytes sha256=%s, want %d bytes sha256=%s",
|
||||
len(result.JavaScript), result.SHA256, plainBytes, plainSHA256)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
import { Tap, actions } from "@sanderling/spec";
|
||||
|
||||
export const properties = {};
|
||||
export const actionsRoot = actions((state) => {
|
||||
const button = state.ax.find("desc:primary");
|
||||
return button ? [Tap({ on: button })] : [];
|
||||
});
|
||||
@@ -0,0 +1 @@
|
||||
export const answer = 42;
|
||||
@@ -10,10 +10,14 @@ import (
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/android"
|
||||
)
|
||||
|
||||
// commandExecutor runs one sanderling invocation and returns its exit code.
|
||||
@@ -25,6 +29,12 @@ func executeCommand(ctx context.Context, binary string, arguments []string, outp
|
||||
command := exec.CommandContext(ctx, binary, arguments...)
|
||||
command.Stdout = output
|
||||
command.Stderr = output
|
||||
// SIGTERM rather than the default kill: a run killed outright never runs its
|
||||
// own shutdown, and its sidecar survives holding a port and a quarter
|
||||
// gigabyte. The run timeout exists for unattended hosts, which is exactly
|
||||
// where nobody is watching to reap what it leaves.
|
||||
command.Cancel = func() error { return command.Process.Signal(syscall.SIGTERM) }
|
||||
command.WaitDelay = runShutdownGrace
|
||||
err := command.Run()
|
||||
if err == nil {
|
||||
return 0, nil
|
||||
@@ -33,10 +43,38 @@ func executeCommand(ctx context.Context, binary string, arguments []string, outp
|
||||
if errors.As(err, &exitError) {
|
||||
return exitError.ExitCode(), nil
|
||||
}
|
||||
if ctx.Err() != nil {
|
||||
return -1, nil
|
||||
}
|
||||
return -1, err
|
||||
}
|
||||
|
||||
// runRecord is one line of runs.jsonl.
|
||||
// connectedDevices lists the devices the host currently has. A variable so a
|
||||
// preflight test runs without a device farm attached.
|
||||
var connectedDevices = android.ConnectedDevices
|
||||
|
||||
// A failure that came back in less than fastFailureThreshold never did the
|
||||
// work the run was asked to do: a step-budgeted run takes tens of minutes,
|
||||
// while a worker pointed at a device that is gone gives up in about half a
|
||||
// minute. fastFailuresBeforeQuarantine of those in a row, with no run that
|
||||
// worked in between to reset the count, is a property of the device rather
|
||||
// than a flake, and it is where the cost of being wrong (one worker's
|
||||
// throughput, since the seeds stay on the shared queue) is still smaller than
|
||||
// the cost of being right one seed later.
|
||||
const (
|
||||
fastFailureThreshold = 2 * time.Minute
|
||||
fastFailuresBeforeQuarantine = 3
|
||||
)
|
||||
|
||||
// runShutdownGrace bounds how long a signalled run gets to stop its sidecar
|
||||
// before it is killed. It exceeds the sidecar's own 15s shutdown grace, or the
|
||||
// escalation would land while the run was still doing what it was asked.
|
||||
const runShutdownGrace = 30 * time.Second
|
||||
|
||||
// runRecord is one line of runs.jsonl. MonotonicMillis is how long the run
|
||||
// worked and WallClockMillis is how much time passed; they answer different
|
||||
// questions and differ by however long the host slept mid-run, which an
|
||||
// unattended overnight sweep is exactly where to expect.
|
||||
type runRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
Device string `json:"device,omitempty"`
|
||||
@@ -44,20 +82,96 @@ type runRecord struct {
|
||||
LaunchError string `json:"launch_error,omitempty"`
|
||||
TimedOut bool `json:"timed_out,omitempty"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
DurationMillis int64 `json:"duration_millis"`
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
WallClockMillis int64 `json:"wall_clock_millis"`
|
||||
RunDirectory string `json:"run_directory,omitempty"`
|
||||
TraceError string `json:"trace_error,omitempty"`
|
||||
traceSummary
|
||||
}
|
||||
|
||||
// clocks reads the two measures a run is timed on. The monotonic clock does
|
||||
// not advance while the host is asleep, so on its own it reports a run that
|
||||
// slept through a quarter of an hour as a quarter of an hour shorter than it
|
||||
// was; the wall clock advances but can be stepped by the host.
|
||||
type clocks struct {
|
||||
monotonicNow func() time.Time
|
||||
wallClockNow func() time.Time
|
||||
}
|
||||
|
||||
func systemClocks() clocks {
|
||||
return clocks{
|
||||
monotonicNow: time.Now,
|
||||
// Round(0) drops the monotonic reading time.Now carries, so
|
||||
// subtracting two of these readings uses the wall clock.
|
||||
wallClockNow: func() time.Time { return time.Now().Round(0) },
|
||||
}
|
||||
}
|
||||
|
||||
type campaign struct {
|
||||
configuration config
|
||||
executor commandExecutor
|
||||
stdout io.Writer
|
||||
records io.Writer
|
||||
clocks clocks
|
||||
mutex sync.Mutex
|
||||
failures int
|
||||
unreadable int
|
||||
quarantined []quarantinedDevice
|
||||
// unrunSeeds are the seeds left without a trustworthy result: those a
|
||||
// quarantined device consumed on its way out, and those still queued when
|
||||
// the last worker stopped.
|
||||
unrunSeeds []int64
|
||||
}
|
||||
|
||||
// failureStreak is one worker's run of fast failures, and the seeds they cost.
|
||||
type failureStreak struct {
|
||||
fastFailures int
|
||||
consumedSeeds []int64
|
||||
}
|
||||
|
||||
// failedFast reports a run that came back non-zero too quickly to have done the
|
||||
// work it was given. A killed run is excluded: it outlived the run timeout,
|
||||
// which is the opposite of failing fast.
|
||||
func failedFast(record runRecord) bool {
|
||||
return record.ExitCode != 0 && !record.TimedOut &&
|
||||
time.Duration(record.MonotonicMillis)*time.Millisecond < fastFailureThreshold
|
||||
}
|
||||
|
||||
// preflightDevices refuses to start until every serial in --devices is present.
|
||||
// A worker aimed at a serial that is gone fails in seconds and pulls the next
|
||||
// seed, so a handful of dead serials drain the queue while the healthy workers
|
||||
// are still inside their first run. Discovering that on run 1 of 20 is already
|
||||
// too late: the sweep is spent, and its output does not say why.
|
||||
func preflightDevices(ctx context.Context, configuration config) error {
|
||||
// Android is the platform whose worker names a device this can enumerate: a
|
||||
// web worker is a label with nothing behind it, and an --ios-device is
|
||||
// resolved by simctl or by devicectl depending on whether it names a
|
||||
// simulator or a paired phone.
|
||||
if configuration.platform != "android" || len(configuration.devices) == 0 {
|
||||
return nil
|
||||
}
|
||||
present, err := connectedDevices(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("list android devices: %w", err)
|
||||
}
|
||||
var missing []string
|
||||
for _, device := range configuration.devices {
|
||||
if !slices.Contains(present, device) {
|
||||
missing = append(missing, device)
|
||||
}
|
||||
}
|
||||
if len(missing) == 0 {
|
||||
return nil
|
||||
}
|
||||
return fmt.Errorf("not starting: %d of %d --devices not connected: %s (adb reports %s)",
|
||||
len(missing), len(configuration.devices), strings.Join(missing, ", "), presentDevices(present))
|
||||
}
|
||||
|
||||
func presentDevices(present []string) string {
|
||||
if len(present) == 0 {
|
||||
return "no devices"
|
||||
}
|
||||
return strings.Join(present, ", ")
|
||||
}
|
||||
|
||||
func runCampaign(ctx context.Context, configuration config, executor commandExecutor, stdout io.Writer) error {
|
||||
@@ -65,6 +179,9 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
|
||||
return fmt.Errorf("%s already exists in %s: pick a fresh --output so two campaigns do not share a directory",
|
||||
manifestFileName, configuration.outputDirectory)
|
||||
}
|
||||
if err := preflightDevices(ctx, configuration); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
|
||||
return fmt.Errorf("create campaign dir: %w", err)
|
||||
}
|
||||
@@ -79,7 +196,8 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
|
||||
return err
|
||||
}
|
||||
host, _ := os.Hostname()
|
||||
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, host, binaryPath, version, time.Now().UTC())); err != nil {
|
||||
intended := buildManifest(configuration, host, binaryPath, version, time.Now().UTC())
|
||||
if err := writeManifest(configuration.outputDirectory, intended); err != nil {
|
||||
return fmt.Errorf("write %s: %w", manifestFileName, err)
|
||||
}
|
||||
|
||||
@@ -90,13 +208,26 @@ func runCampaign(ctx context.Context, configuration config, executor commandExec
|
||||
}
|
||||
defer recordsFile.Close()
|
||||
|
||||
sweep := &campaign{configuration: configuration, executor: executor, stdout: stdout, records: recordsFile}
|
||||
sweep := &campaign{configuration: configuration, executor: executor, stdout: stdout, records: recordsFile, clocks: systemClocks()}
|
||||
fmt.Fprintf(stdout, "campaign %s: %d seeds, %d worker(s), %s\n",
|
||||
configuration.arm, len(configuration.seeds), len(workerDevices(configuration.devices)), configuration.outputDirectory)
|
||||
sweep.sweep(ctx)
|
||||
|
||||
fmt.Fprintf(stdout, "campaign complete: %d of %d runs failed, %d produced an unreadable trace\n",
|
||||
sweep.failures, len(configuration.seeds), sweep.unreadable)
|
||||
if len(sweep.quarantined) > 0 || len(sweep.unrunSeeds) > 0 {
|
||||
intended.Quarantined = sweep.quarantined
|
||||
intended.UnrunSeeds = sweep.unrunSeeds
|
||||
if err := writeManifest(configuration.outputDirectory, intended); err != nil {
|
||||
fmt.Fprintf(stdout, "warning: rewrite %s: %v\n", manifestFileName, err)
|
||||
}
|
||||
fmt.Fprintf(stdout, "%d of %d seeds have no result: %v\n",
|
||||
len(sweep.unrunSeeds), len(configuration.seeds), sweep.unrunSeeds)
|
||||
}
|
||||
if len(sweep.quarantined) > 0 && len(sweep.quarantined) == len(workerDevices(configuration.devices)) {
|
||||
return fmt.Errorf("every device was quarantined (%s); %d of %d seeds have no result",
|
||||
strings.Join(quarantinedNames(sweep.quarantined), ", "), len(sweep.unrunSeeds), len(configuration.seeds))
|
||||
}
|
||||
if sweep.failures > 0 {
|
||||
return fmt.Errorf("%d of %d runs failed", sweep.failures, len(configuration.seeds))
|
||||
}
|
||||
@@ -142,15 +273,68 @@ func (c *campaign) sweep(ctx context.Context) {
|
||||
waitGroup.Add(1)
|
||||
go func(device string) {
|
||||
defer waitGroup.Done()
|
||||
c.work(ctx, device, queue)
|
||||
}(device)
|
||||
}
|
||||
waitGroup.Wait()
|
||||
|
||||
for seed := range queue {
|
||||
c.unrunSeeds = append(c.unrunSeeds, seed)
|
||||
}
|
||||
slices.Sort(c.unrunSeeds)
|
||||
}
|
||||
|
||||
// work runs seeds on one device until the queue is empty or the device has
|
||||
// failed fast often enough in a row to be quarantined. The streak is worker
|
||||
// local: one worker drives one device, and any run that did its work clears it.
|
||||
func (c *campaign) work(ctx context.Context, device string, queue <-chan int64) {
|
||||
var streak failureStreak
|
||||
for seed := range queue {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
c.report(c.runSeed(ctx, seed, device))
|
||||
record := c.runSeed(ctx, seed, device)
|
||||
c.report(record)
|
||||
// A cancelled campaign fails every run in flight in seconds, which is
|
||||
// the shutdown and not the device.
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
}(device)
|
||||
if !failedFast(record) {
|
||||
streak = failureStreak{}
|
||||
continue
|
||||
}
|
||||
waitGroup.Wait()
|
||||
streak.fastFailures++
|
||||
streak.consumedSeeds = append(streak.consumedSeeds, seed)
|
||||
if streak.fastFailures >= fastFailuresBeforeQuarantine {
|
||||
c.quarantine(device, streak)
|
||||
return
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// quarantine takes a device out of the sweep. The seeds it burned are reported
|
||||
// unrun rather than requeued: each already holds that attempt's log, and a
|
||||
// second record for the same seed would make one seed count as two runs.
|
||||
func (c *campaign) quarantine(device string, streak failureStreak) {
|
||||
c.mutex.Lock()
|
||||
defer c.mutex.Unlock()
|
||||
c.quarantined = append(c.quarantined, quarantinedDevice{
|
||||
Device: device,
|
||||
FastFailures: streak.fastFailures,
|
||||
ConsumedSeeds: streak.consumedSeeds,
|
||||
})
|
||||
c.unrunSeeds = append(c.unrunSeeds, streak.consumedSeeds...)
|
||||
fmt.Fprintf(c.stdout, "quarantined device %q: %d runs in a row failed in under %s; seeds %v have no result\n",
|
||||
device, streak.fastFailures, fastFailureThreshold, streak.consumedSeeds)
|
||||
}
|
||||
|
||||
func quarantinedNames(devices []quarantinedDevice) []string {
|
||||
names := make([]string, 0, len(devices))
|
||||
for _, device := range devices {
|
||||
names = append(names, device.Device)
|
||||
}
|
||||
return names
|
||||
}
|
||||
|
||||
func (c *campaign) runSeed(ctx context.Context, seed int64, device string) runRecord {
|
||||
@@ -173,12 +357,14 @@ func (c *campaign) runSeed(ctx context.Context, seed int64, device string) runRe
|
||||
|
||||
runCtx, cancelRun := context.WithTimeout(ctx, c.configuration.runTimeout)
|
||||
defer cancelRun()
|
||||
start := time.Now()
|
||||
monotonicStart := c.clocks.monotonicNow()
|
||||
wallClockStart := c.clocks.wallClockNow()
|
||||
exitCode, runErr := c.executor(runCtx, c.configuration.sanderlingPath, runArguments(c.configuration, seedText, device), logFile)
|
||||
if runCtx.Err() != nil && ctx.Err() == nil {
|
||||
record.TimedOut = true
|
||||
}
|
||||
record.DurationMillis = time.Since(start).Milliseconds()
|
||||
record.MonotonicMillis = c.clocks.monotonicNow().Sub(monotonicStart).Milliseconds()
|
||||
record.WallClockMillis = c.clocks.wallClockNow().Sub(wallClockStart).Milliseconds()
|
||||
record.ExitCode = exitCode
|
||||
if runErr != nil {
|
||||
record.LaunchError = runErr.Error()
|
||||
@@ -207,9 +393,10 @@ func (c *campaign) report(record runRecord) {
|
||||
if err := json.NewEncoder(c.records).Encode(record); err != nil {
|
||||
fmt.Fprintf(c.stdout, "warning: seed %d record: %v\n", record.Seed, err)
|
||||
}
|
||||
fmt.Fprintf(c.stdout, "seed=%d device=%q outcome=%s steps=%d exit=%d duration=%s\n",
|
||||
fmt.Fprintf(c.stdout, "seed=%d device=%q outcome=%s steps=%d exit=%d monotonic=%s wall_clock=%s\n",
|
||||
record.Seed, record.Device, outcome(record), record.Steps, record.ExitCode,
|
||||
time.Duration(record.DurationMillis)*time.Millisecond)
|
||||
time.Duration(record.MonotonicMillis)*time.Millisecond,
|
||||
time.Duration(record.WallClockMillis)*time.Millisecond)
|
||||
}
|
||||
|
||||
func outcome(record runRecord) string {
|
||||
|
||||
@@ -69,6 +69,91 @@ func writeFakeRun(t *testing.T, arguments []string, steps []trace.Step) {
|
||||
writeRunDirectory(t, argumentValue(arguments, "--output"), "20260101-000000", steps)
|
||||
}
|
||||
|
||||
// readings hands out the given instants in turn, so a test can script a clock
|
||||
// that jumps across a host sleep independently of one that stops through it.
|
||||
func readings(instants ...time.Time) func() time.Time {
|
||||
var index int
|
||||
return func() time.Time {
|
||||
instant := instants[min(index, len(instants)-1)]
|
||||
index++
|
||||
return instant
|
||||
}
|
||||
}
|
||||
|
||||
// A sleeping host stops the monotonic clock and not the wall clock, so a run
|
||||
// timed on the monotonic clock alone reports the sleep as time that never
|
||||
// passed. The record carries both, named for the clock each came from.
|
||||
func TestRunSeed_RecordsTheTimeWorkedAndTheTimeThatPassed(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
startedAt := time.Date(2026, 8, 14, 2, 0, 0, 0, time.UTC)
|
||||
var records bytes.Buffer
|
||||
sweep := &campaign{
|
||||
configuration: testConfiguration(t, directory, "--seeds", "1"),
|
||||
executor: func(context.Context, string, []string, io.Writer) (int, error) {
|
||||
return 0, nil
|
||||
},
|
||||
stdout: io.Discard,
|
||||
records: &records,
|
||||
clocks: clocks{
|
||||
monotonicNow: readings(startedAt, startedAt.Add(2*time.Minute)),
|
||||
wallClockNow: readings(startedAt, startedAt.Add(17*time.Minute)),
|
||||
},
|
||||
}
|
||||
|
||||
sweep.report(sweep.runSeed(context.Background(), 1, ""))
|
||||
|
||||
var written map[string]any
|
||||
if err := json.Unmarshal(records.Bytes(), &written); err != nil {
|
||||
t.Fatalf("decode %q: %v", records.String(), err)
|
||||
}
|
||||
if written["monotonic_millis"] != float64((2 * time.Minute).Milliseconds()) {
|
||||
t.Errorf("monotonic_millis %v, want the two minutes of work", written["monotonic_millis"])
|
||||
}
|
||||
if written["wall_clock_millis"] != float64((17 * time.Minute).Milliseconds()) {
|
||||
t.Errorf("wall_clock_millis %v, want the seventeen minutes that passed", written["wall_clock_millis"])
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunCampaign_RecordsDispatchedActionsNotSteps(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1")
|
||||
|
||||
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
|
||||
writeFakeRun(t, arguments, []trace.Step{
|
||||
actingStep(1), observedStep(2), skippedActionStep(3, "unresolved_selector"), actingStep(4),
|
||||
})
|
||||
return 0, nil
|
||||
})
|
||||
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
records := readRecords(t, directory)
|
||||
if len(records) != 1 {
|
||||
t.Fatalf("records: got %d, want 1", len(records))
|
||||
}
|
||||
if records[0].Steps != 4 || records[0].Actions != 2 {
|
||||
t.Errorf("steps %d actions %d, want 4 and 2", records[0].Steps, records[0].Actions)
|
||||
}
|
||||
}
|
||||
|
||||
// A run that never produced a readable trace still has to carry the field, so
|
||||
// analysis can tell a zero-action run from a file written before the count. The
|
||||
// same holds for the unattributed count: a run whose every action named its
|
||||
// producer says so with a zero, and a file that says nothing is one recorded
|
||||
// before actions named one at all.
|
||||
func TestRunRecord_AlwaysCarriesTheActionCounts(t *testing.T) {
|
||||
body, err := json.Marshal(runRecord{Seed: 7, TraceError: "no run directory with meta.json"})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, field := range []string{`"actions":0`, `"unattributed_actions":0`} {
|
||||
if !strings.Contains(string(body), field) {
|
||||
t.Errorf("record %s omits %s", body, field)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunCampaign_WritesManifestBeforeAnyRun(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory)
|
||||
@@ -172,6 +257,7 @@ func TestRunCampaign_RecordsPerRunSummary(t *testing.T) {
|
||||
func TestRunCampaign_DistributesSeedsAcrossDeviceWorkers(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1-9", "--devices", "device-a,device-b,device-c")
|
||||
devicesPresent(t, "device-a", "device-b", "device-c")
|
||||
|
||||
var mutex sync.Mutex
|
||||
assignments := map[int64]string{}
|
||||
@@ -263,6 +349,28 @@ func TestRunCampaign_ContinuesAfterFailingRun(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunCampaign_GivesEveryRunTheCellsLabelSource(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1-3", "--label-source", "resource-id")
|
||||
|
||||
var mutex sync.Mutex
|
||||
var dispatched []string
|
||||
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
|
||||
mutex.Lock()
|
||||
dispatched = append(dispatched, argumentValue(arguments, "--label-source"))
|
||||
mutex.Unlock()
|
||||
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
|
||||
return 0, nil
|
||||
})
|
||||
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if !slices.Equal(dispatched, []string{"resource-id", "resource-id", "resource-id"}) {
|
||||
t.Errorf("--label-source reaching sanderling: got %v, want resource-id on every run", dispatched)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunCampaign_RefusesToReuseACampaignDirectory(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory)
|
||||
@@ -360,3 +468,208 @@ func TestParseArguments_RunTimeoutDefaultsToThreeTimesDuration(t *testing.T) {
|
||||
t.Errorf("run timeout default: got %s, want 12m", configuration.runTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
// devicesPresent points the preflight at a fixed set of serials, so a campaign
|
||||
// can be preflighted on a host with no device farm attached.
|
||||
func devicesPresent(t *testing.T, present ...string) {
|
||||
t.Helper()
|
||||
original := connectedDevices
|
||||
connectedDevices = func(context.Context) ([]string, error) { return present, nil }
|
||||
t.Cleanup(func() { connectedDevices = original })
|
||||
}
|
||||
|
||||
// A worker aimed at a serial that no longer exists fails in seconds and pulls
|
||||
// the next seed, so a few dead serials drain the queue while the healthy
|
||||
// workers are still inside their first run. That has to be caught before the
|
||||
// first seed is dispatched: a sweep that discovers it on run 1 of 20 has
|
||||
// already been destroyed, and its output does not say so.
|
||||
func TestRunCampaign_RefusesToStartWhenADeviceIsMissing(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory,
|
||||
"--seeds", "1-20", "--devices", "emulator-5554,emulator-5564,emulator-5556")
|
||||
devicesPresent(t, "emulator-5554", "emulator-5556")
|
||||
|
||||
executor := func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
|
||||
t.Errorf("the campaign dispatched %v despite a missing device", arguments)
|
||||
return 0, nil
|
||||
}
|
||||
err := runCampaign(context.Background(), configuration, executor, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), "not connected: emulator-5564") {
|
||||
t.Fatalf("the error must name the missing serial, got %v", err)
|
||||
}
|
||||
if _, statErr := os.Stat(filepath.Join(directory, manifestFileName)); statErr == nil {
|
||||
t.Error("a campaign that cannot run wrote a manifest")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunCampaign_RunsWhenEveryDeviceIsPresent(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1-2", "--devices", "emulator-5554,emulator-5556")
|
||||
devicesPresent(t, "emulator-5556", "emulator-5580", "emulator-5554")
|
||||
|
||||
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
|
||||
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
|
||||
return 0, nil
|
||||
})
|
||||
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got := len(readRecords(t, directory)); got != 2 {
|
||||
t.Errorf("records: got %d, want 2", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Preflight only means something where the worker names a device. A campaign
|
||||
// without --devices has one worker and no serial, and a web worker is a label
|
||||
// with no device behind it; neither may be blocked by a device check.
|
||||
func TestRunCampaign_PreflightsOnlyWhereWorkersNameADevice(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
name string
|
||||
extra []string
|
||||
}{
|
||||
{"no devices", []string{"--seeds", "1"}},
|
||||
{"web workers", []string{"--seeds", "1", "--platform", "web", "--devices", "worker-a,worker-b"}},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, testCase.extra...)
|
||||
original := connectedDevices
|
||||
connectedDevices = func(context.Context) ([]string, error) {
|
||||
t.Error("preflight looked for devices where the workers name none")
|
||||
return nil, fmt.Errorf("no adb server")
|
||||
}
|
||||
t.Cleanup(func() { connectedDevices = original })
|
||||
|
||||
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, _ io.Writer) (int, error) {
|
||||
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
|
||||
return 0, nil
|
||||
})
|
||||
if err := runCampaign(context.Background(), configuration, executor, io.Discard); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Preflight cannot catch a device that disappears mid-campaign. A device that
|
||||
// keeps failing in seconds is drained of seeds by exactly the speed of its
|
||||
// failure, so it has to stop being given any.
|
||||
func TestRunCampaign_QuarantinesADeviceThatKeepsFailingFast(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1-8", "--devices", "device-a,device-b")
|
||||
devicesPresent(t, "device-a", "device-b")
|
||||
|
||||
failedEnough := make(chan struct{})
|
||||
var closeOnce sync.Once
|
||||
var timedOut atomic.Bool
|
||||
var mutex sync.Mutex
|
||||
dispatched := map[string][]int64{}
|
||||
|
||||
executor := versionAnswering(func(_ context.Context, _ string, arguments []string, output io.Writer) (int, error) {
|
||||
device := argumentValue(arguments, "--device")
|
||||
seed, err := strconv.ParseInt(argumentValue(arguments, "--seed"), 10, 64)
|
||||
if err != nil {
|
||||
t.Errorf("seed argument: %v", err)
|
||||
}
|
||||
mutex.Lock()
|
||||
dispatched[device] = append(dispatched[device], seed)
|
||||
failures := len(dispatched["device-a"])
|
||||
mutex.Unlock()
|
||||
if device == "device-a" {
|
||||
fmt.Fprintln(output, "device 'device-a' not found")
|
||||
if failures >= fastFailuresBeforeQuarantine {
|
||||
closeOnce.Do(func() { close(failedEnough) })
|
||||
}
|
||||
return 1, nil
|
||||
}
|
||||
// The healthy worker holds its seed until the sick one has failed its
|
||||
// way to quarantine, so the split of the queue is the scheduler's
|
||||
// decision and not a race between two equally fast fakes.
|
||||
select {
|
||||
case <-failedEnough:
|
||||
case <-time.After(5 * time.Second):
|
||||
timedOut.Store(true)
|
||||
}
|
||||
writeFakeRun(t, arguments, []trace.Step{observedStep(1)})
|
||||
return 0, nil
|
||||
})
|
||||
|
||||
var stdout bytes.Buffer
|
||||
err := runCampaign(context.Background(), configuration, executor, &stdout)
|
||||
if timedOut.Load() {
|
||||
t.Fatal("device-a never reached quarantine")
|
||||
}
|
||||
if err == nil || !strings.Contains(err.Error(), "3 of 8 runs failed") {
|
||||
t.Fatalf("expected the three fast failures to be reported, got %v", err)
|
||||
}
|
||||
if got := len(dispatched["device-a"]); got != fastFailuresBeforeQuarantine {
|
||||
t.Errorf("seeds sent to the failing device: got %d, want %d", got, fastFailuresBeforeQuarantine)
|
||||
}
|
||||
if got := len(dispatched["device-b"]); got != 8-fastFailuresBeforeQuarantine {
|
||||
t.Errorf("seeds sent to the healthy device: got %d, want %d", got, 8-fastFailuresBeforeQuarantine)
|
||||
}
|
||||
ran := append(append([]int64{}, dispatched["device-a"]...), dispatched["device-b"]...)
|
||||
slices.Sort(ran)
|
||||
if !slices.Equal(ran, []int64{1, 2, 3, 4, 5, 6, 7, 8}) {
|
||||
t.Errorf("every seed must be dispatched exactly once: got %v", ran)
|
||||
}
|
||||
if !strings.Contains(stdout.String(), `quarantined device "device-a"`) {
|
||||
t.Errorf("the quarantine must be reported in the campaign output:\n%s", stdout.String())
|
||||
}
|
||||
|
||||
recorded := readManifest(t, directory)
|
||||
if len(recorded.Quarantined) != 1 || recorded.Quarantined[0].Device != "device-a" {
|
||||
t.Fatalf("the manifest must record the quarantine: %+v", recorded.Quarantined)
|
||||
}
|
||||
burned := slices.Clone(dispatched["device-a"])
|
||||
slices.Sort(burned)
|
||||
if !slices.Equal(recorded.Quarantined[0].ConsumedSeeds, burned) {
|
||||
t.Errorf("consumed seeds: got %v, want %v", recorded.Quarantined[0].ConsumedSeeds, burned)
|
||||
}
|
||||
if !slices.Equal(recorded.UnrunSeeds, burned) {
|
||||
t.Errorf("the seeds the quarantined device burned have no result and must be reported unrun: got %v, want %v",
|
||||
recorded.UnrunSeeds, burned)
|
||||
}
|
||||
}
|
||||
|
||||
// Every device gone is not a campaign that should keep pulling seeds: the rest
|
||||
// of the queue would be spent producing the same failure.
|
||||
func TestRunCampaign_AbortsWhenEveryDeviceIsQuarantined(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
configuration := testConfiguration(t, directory, "--seeds", "1-10", "--devices", "device-a,device-b")
|
||||
devicesPresent(t, "device-a", "device-b")
|
||||
|
||||
var dispatches atomic.Int32
|
||||
executor := versionAnswering(func(_ context.Context, _ string, _ []string, output io.Writer) (int, error) {
|
||||
dispatches.Add(1)
|
||||
fmt.Fprintln(output, "sidecar health check: context deadline exceeded")
|
||||
return 1, nil
|
||||
})
|
||||
|
||||
var stdout bytes.Buffer
|
||||
err := runCampaign(context.Background(), configuration, executor, &stdout)
|
||||
if err == nil || !strings.Contains(err.Error(), "every device was quarantined") {
|
||||
t.Fatalf("a campaign with no device left must abort with a clear error, got %v", err)
|
||||
}
|
||||
if got := dispatches.Load(); got != int32(2*fastFailuresBeforeQuarantine) {
|
||||
t.Errorf("runs dispatched: got %d, want %d: the sweep must stop rather than spin through the queue",
|
||||
got, 2*fastFailuresBeforeQuarantine)
|
||||
}
|
||||
recorded := readManifest(t, directory)
|
||||
if !slices.Equal(recorded.UnrunSeeds, []int64{1, 2, 3, 4, 5, 6, 7, 8, 9, 10}) {
|
||||
t.Errorf("every seed is either burned by a quarantined device or never dispatched: got %v", recorded.UnrunSeeds)
|
||||
}
|
||||
}
|
||||
|
||||
func readManifest(t *testing.T, campaignDirectory string) manifest {
|
||||
t.Helper()
|
||||
body, err := os.ReadFile(filepath.Join(campaignDirectory, manifestFileName))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var recorded manifest
|
||||
if err := json.Unmarshal(body, &recorded); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return recorded
|
||||
}
|
||||
@@ -2,6 +2,7 @@ package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"io"
|
||||
"os"
|
||||
@@ -9,6 +10,7 @@ import (
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// stubSanderling answers `version`, writes a run directory shaped like the one
|
||||
@@ -142,3 +144,53 @@ func TestRun_EndToEndAgainstStubBinary(t *testing.T) {
|
||||
t.Errorf("progress output: %q", stdout.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestExecuteCommand_RunTimeoutSignalsSoTheRunReapsItsChildren(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
marker := filepath.Join(directory, "child-reaped")
|
||||
trapped := filepath.Join(directory, "trap-installed")
|
||||
script := filepath.Join(directory, "wedged")
|
||||
body := "#!/bin/sh\n" +
|
||||
"sleep 300 &\n" +
|
||||
"child=$!\n" +
|
||||
"trap 'kill $child; echo reaped > " + marker + "; exit 143' TERM\n" +
|
||||
"echo installed > " + trapped + "\n" +
|
||||
"wait $child\n"
|
||||
if err := os.WriteFile(script, []byte(body), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// Cancel only once the script has installed its trap. A fixed deadline
|
||||
// races the shell under a loaded machine, and a run signalled before its
|
||||
// trap exists fails this test for a reason it does not test.
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
defer cancel()
|
||||
finished := make(chan error, 1)
|
||||
go func() {
|
||||
_, err := executeCommand(ctx, script, nil, io.Discard)
|
||||
finished <- err
|
||||
}()
|
||||
waitForFile(t, trapped)
|
||||
cancel()
|
||||
|
||||
if err := <-finished; err != nil {
|
||||
t.Fatalf("execute: %v", err)
|
||||
}
|
||||
|
||||
if _, err := os.Stat(marker); err != nil {
|
||||
t.Fatal("the run timeout killed the run outright, so it never reaped its own children: " +
|
||||
"an unattended sweep leaks one sidecar per wedged run")
|
||||
}
|
||||
}
|
||||
|
||||
func waitForFile(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(30 * time.Second)
|
||||
for time.Now().Before(deadline) {
|
||||
if _, err := os.Stat(path); err == nil {
|
||||
return
|
||||
}
|
||||
time.Sleep(5 * time.Millisecond)
|
||||
}
|
||||
t.Fatalf("%s never appeared", path)
|
||||
}
|
||||
@@ -15,6 +15,8 @@ import (
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/seedspec"
|
||||
)
|
||||
|
||||
type config struct {
|
||||
@@ -23,6 +25,7 @@ type config struct {
|
||||
platform string
|
||||
arm string
|
||||
generator string
|
||||
labelSource string
|
||||
maxSteps int
|
||||
duration time.Duration
|
||||
seeds []int64
|
||||
@@ -58,6 +61,7 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
flagSet.StringVar(&configuration.platform, "platform", "android", "target platform: android, ios, web")
|
||||
flagSet.StringVar(&configuration.arm, "arm", "", "experiment cell label recorded on every run (required)")
|
||||
flagSet.StringVar(&configuration.generator, "generator", "seeded", "action generator: seeded or llm")
|
||||
flagSet.StringVar(&configuration.labelSource, "label-source", "visible-text", "how candidates are named to the llm generator: visible-text (what a user reads) or resource-id (the identifier the app assigned). The seeded generator picks by index and ignores this")
|
||||
flagSet.IntVar(&configuration.maxSteps, "max-steps", 0, "per-run step budget (required, must be positive)")
|
||||
flagSet.DurationVar(&configuration.duration, "duration", 5*time.Minute, "per-run wall-clock ceiling")
|
||||
flagSet.StringVar(&seedSpecification, "seeds", "", "seeds to run: ranges and lists, e.g. 1-10,20,30-32 (required)")
|
||||
@@ -70,17 +74,26 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
}
|
||||
configuration.extraArguments = flagSet.Args()
|
||||
|
||||
for name, value := range map[string]string{
|
||||
"--spec": configuration.specPath,
|
||||
"--bundle-id": configuration.bundleID,
|
||||
"--arm": configuration.arm,
|
||||
"--seeds": seedSpecification,
|
||||
"--output": configuration.outputDirectory,
|
||||
// Every missing flag is named together, in flag order: stopping at the
|
||||
// first turns one rerun into one rerun per missing flag.
|
||||
var missing []error
|
||||
for _, required := range []struct {
|
||||
name string
|
||||
value string
|
||||
}{
|
||||
{"--spec", configuration.specPath},
|
||||
{"--bundle-id", configuration.bundleID},
|
||||
{"--arm", configuration.arm},
|
||||
{"--seeds", seedSpecification},
|
||||
{"--output", configuration.outputDirectory},
|
||||
} {
|
||||
if value == "" {
|
||||
return config{}, fmt.Errorf("%s is required", name)
|
||||
if required.value == "" {
|
||||
missing = append(missing, fmt.Errorf("%s is required", required.name))
|
||||
}
|
||||
}
|
||||
if err := errors.Join(missing...); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
switch configuration.platform {
|
||||
case "android", "ios", "web":
|
||||
default:
|
||||
@@ -91,6 +104,13 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
default:
|
||||
return config{}, fmt.Errorf("unsupported generator: %q (seeded, llm)", configuration.generator)
|
||||
}
|
||||
// Rejected before the first run, because a sweep that discovers the bad
|
||||
// value on run 1 of 40 has already spent the device time of a whole cell.
|
||||
switch configuration.labelSource {
|
||||
case "visible-text", "resource-id":
|
||||
default:
|
||||
return config{}, fmt.Errorf("unsupported label source: %q (visible-text, resource-id)", configuration.labelSource)
|
||||
}
|
||||
if configuration.maxSteps <= 0 {
|
||||
// Steps to first violation is right-censored at the budget, so a
|
||||
// campaign without one has nothing to censor its clean runs at.
|
||||
@@ -109,7 +129,7 @@ func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
return config{}, fmt.Errorf("--run-timeout %s must exceed --duration %s, or every run is killed before it finishes",
|
||||
configuration.runTimeout, configuration.duration)
|
||||
}
|
||||
seeds, err := parseSeeds(seedSpecification)
|
||||
seeds, err := seedspec.Parse(seedSpecification)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("--seeds: %w", err)
|
||||
}
|
||||
@@ -170,6 +190,7 @@ func runArguments(configuration config, seed, device string) []string {
|
||||
"--platform", configuration.platform,
|
||||
"--arm", configuration.arm,
|
||||
"--generator", configuration.generator,
|
||||
"--label-source", configuration.labelSource,
|
||||
"--max-steps", strconv.Itoa(configuration.maxSteps),
|
||||
"--duration", configuration.duration.String(),
|
||||
"--seed", seed,
|
||||
|
||||
@@ -54,6 +54,7 @@ func TestParseArguments_Rejections(t *testing.T) {
|
||||
{"missing output", []string{"--spec", "s", "--bundle-id", "a", "--arm", "b", "--seeds", "1", "--max-steps", "10"}, "--output is required"},
|
||||
{"bad platform", append(baseArguments(), "--platform", "windows"), "unsupported platform"},
|
||||
{"bad generator", append(baseArguments(), "--generator", "vibes"), "unsupported generator"},
|
||||
{"bad label source", append(baseArguments(), "--label-source", "resource_id"), `unsupported label source: "resource_id"`},
|
||||
{"zero max steps", append(baseArguments(), "--max-steps", "0"), "--max-steps must be positive"},
|
||||
{"seed zero", append(baseArguments(), "--seeds", "0-2"), "not reproducible"},
|
||||
{"duplicate device", append(baseArguments(), "--devices", "a,a"), "duplicate device"},
|
||||
@@ -70,6 +71,35 @@ func TestParseArguments_Rejections(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// Three flags missing is one rerun, not three: the operator is told about all
|
||||
// of them at once, in flag order, whatever order the check happened to walk.
|
||||
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
|
||||
_, err := parseArguments(
|
||||
[]string{"--bundle-id", "a", "--seeds", "1", "--max-steps", "10"},
|
||||
io.Discard,
|
||||
)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want every missing flag named")
|
||||
}
|
||||
message := err.Error()
|
||||
previous := -1
|
||||
for _, name := range []string{"--spec", "--arm", "--output"} {
|
||||
at := strings.Index(message, name)
|
||||
if at < 0 {
|
||||
t.Fatalf("got %q, want %s named", message, name)
|
||||
}
|
||||
if at < previous {
|
||||
t.Errorf("got %q, want the flags named in flag order", message)
|
||||
}
|
||||
previous = at
|
||||
}
|
||||
for _, supplied := range []string{"--bundle-id", "--seeds"} {
|
||||
if strings.Contains(message, supplied) {
|
||||
t.Errorf("got %q, want the supplied %s left out", message, supplied)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunArguments_PlatformDeviceFlagAndPassthrough(t *testing.T) {
|
||||
cases := []struct {
|
||||
platform string
|
||||
@@ -113,6 +143,16 @@ func TestRunArguments_PlatformDeviceFlagAndPassthrough(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunArguments_LabelSourceDefaultsToVisibleText(t *testing.T) {
|
||||
configuration, err := parseArguments(baseArguments(), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got := argumentValue(runArguments(configuration, "7", ""), "--label-source"); got != "visible-text" {
|
||||
t.Errorf("--label-source = %q, want visible-text", got)
|
||||
}
|
||||
}
|
||||
|
||||
func argumentValue(arguments []string, name string) string {
|
||||
index := slices.Index(arguments, name)
|
||||
if index < 0 || index+1 >= len(arguments) {
|
||||
|
||||
@@ -21,6 +21,7 @@ const (
|
||||
type manifest struct {
|
||||
Arm string `json:"arm"`
|
||||
Generator string `json:"generator"`
|
||||
LabelSource string `json:"label_source"`
|
||||
Platform string `json:"platform"`
|
||||
SpecPath string `json:"spec_path"`
|
||||
BundleID string `json:"bundle_id"`
|
||||
@@ -34,6 +35,22 @@ type manifest struct {
|
||||
SanderlingVersion string `json:"sanderling_version"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
ArgumentTemplate []string `json:"argument_template"`
|
||||
|
||||
// Quarantined and UnrunSeeds are filled in when the sweep ends, so the
|
||||
// artifact says a device dropped out and which seeds have no result rather
|
||||
// than leaving both to be inferred from a thin runs.jsonl.
|
||||
Quarantined []quarantinedDevice `json:"quarantined,omitempty"`
|
||||
UnrunSeeds []int64 `json:"unrun_seeds,omitempty"`
|
||||
}
|
||||
|
||||
// quarantinedDevice is a device the sweep stopped assigning seeds to, with the
|
||||
// seeds it consumed on the way out. Those seeds are reported unrun rather than
|
||||
// requeued: their directories already hold the failed attempt's log, and a
|
||||
// second record for the same seed would make a seed count two runs.
|
||||
type quarantinedDevice struct {
|
||||
Device string `json:"device"`
|
||||
FastFailures int `json:"fast_failures"`
|
||||
ConsumedSeeds []int64 `json:"consumed_seeds"`
|
||||
}
|
||||
|
||||
func buildManifest(configuration config, host, binaryPath, version string, startedAt time.Time) manifest {
|
||||
@@ -44,6 +61,7 @@ func buildManifest(configuration config, host, binaryPath, version string, start
|
||||
return manifest{
|
||||
Arm: configuration.arm,
|
||||
Generator: configuration.generator,
|
||||
LabelSource: configuration.labelSource,
|
||||
Platform: configuration.platform,
|
||||
SpecPath: configuration.specPath,
|
||||
BundleID: configuration.bundleID,
|
||||
|
||||
@@ -41,6 +41,30 @@ func TestBuildManifest_RecordsIntendedRunsAndTemplate(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteManifest_RecordsTheLabelSourceCell(t *testing.T) {
|
||||
configuration, err := parseArguments(append(baseArguments(), "--label-source", "resource-id"), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
directory := t.TempDir()
|
||||
if err := writeManifest(directory, buildManifest(configuration, "host", "sanderling", "dev", time.Now())); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body, err := os.ReadFile(filepath.Join(directory, manifestFileName))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var decoded struct {
|
||||
LabelSource string `json:"label_source"`
|
||||
}
|
||||
if err := json.Unmarshal(body, &decoded); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if decoded.LabelSource != "resource-id" {
|
||||
t.Errorf("label_source: got %q, want resource-id", decoded.LabelSource)
|
||||
}
|
||||
}
|
||||
|
||||
func TestBuildManifest_EmptyDeviceListSerializesAsArray(t *testing.T) {
|
||||
configuration, err := parseArguments(baseArguments(), io.Discard)
|
||||
if err != nil {
|
||||
|
||||
@@ -20,12 +20,34 @@ const maxTraceLineBytes = 16 * 1024 * 1024
|
||||
// to open trace.jsonl again.
|
||||
type traceSummary struct {
|
||||
Steps int `json:"steps"`
|
||||
// Actions counts only the steps on which the action generator both chose an
|
||||
// action and dispatched it. A step whose policy declined to act, a step
|
||||
// whose chosen action the runner threw away, and a step the spec's setup
|
||||
// drove before the generator was ever consulted all explored nothing, so
|
||||
// counting any of them as an action inflates the denominator of every
|
||||
// per-action rate. The inflation is policy-dependent, so it does not cancel
|
||||
// between arms.
|
||||
Actions int `json:"actions"`
|
||||
// UnattributedActions counts the dispatched steps whose action names no
|
||||
// producer, which only a trace recorded before actions carried one can do.
|
||||
// Such a run's Actions is the count it was already reported with rather than
|
||||
// a setup-excluding one, and this is what says so. It is written even when
|
||||
// it is zero, because a record that omits it is one the analysis has to read
|
||||
// as unattributable and a run where every action named a producer is the
|
||||
// opposite of that.
|
||||
UnattributedActions int `json:"unattributed_actions"`
|
||||
FirstViolationOriginStep *int `json:"first_violation_origin_step"`
|
||||
FirstViolationDetectedStep *int `json:"first_violation_detected_step"`
|
||||
FirstViolationProperties []string `json:"first_violation_properties,omitempty"`
|
||||
FirstViolationReason string `json:"first_violation_reason,omitempty"`
|
||||
FirstViolationIsError bool `json:"first_violation_is_error,omitempty"`
|
||||
ViolatedProperties []string `json:"violated_properties,omitempty"`
|
||||
// PreconditionFailures counts the trace records naming a precondition the
|
||||
// run could not meet: the startup gate's verdict at step 0, and every later
|
||||
// step the scope guard could not bring the app back for. A run with one of
|
||||
// these and no steps never started, and counting it as a run that explored
|
||||
// and found nothing puts a harness failure in the same column as evidence.
|
||||
PreconditionFailures int `json:"precondition_failures,omitempty"`
|
||||
}
|
||||
|
||||
type traceLine struct {
|
||||
@@ -33,10 +55,52 @@ type traceLine struct {
|
||||
// Hierarchy is read only for its presence: the run-end finalize line is the
|
||||
// one line carrying violations without an observed hierarchy.
|
||||
Hierarchy json.RawMessage `json:"hierarchy"`
|
||||
// NextAction is nil on a step that chose no action. Its Source names the
|
||||
// backend that chose the action, which is the only thing in the trace
|
||||
// separating a step the spec's setup drove from one the generator drove.
|
||||
NextAction *actionLine `json:"next_action"`
|
||||
ActionSkipped string `json:"action_skipped"`
|
||||
PreconditionFailure string `json:"precondition_failure"`
|
||||
Violations []string `json:"violations"`
|
||||
Witnesses map[string]trace.Witness `json:"witnesses"`
|
||||
}
|
||||
|
||||
type actionLine struct {
|
||||
Source string `json:"source"`
|
||||
}
|
||||
|
||||
// generatorDispatched reports whether this step drove the app on the action
|
||||
// generator's behalf, which is the exposure a per-action rate divides by. A
|
||||
// spec's setup puts the app into its starting position before the generator is
|
||||
// consulted, so its login taps are dispatched actions that explored nothing.
|
||||
// Both arms are read the same way, off the source the action names, because a
|
||||
// denominator that excludes setup on one arm and includes it on the other makes
|
||||
// the two rates incomparable.
|
||||
//
|
||||
// An action naming no source at all is one recorded before the distinction
|
||||
// existed and cannot be attributed now, so each arm keeps the count it was
|
||||
// already reported with: everything the seeded picker dispatched, and only what
|
||||
// the model stamped. summarizeTrace counts those steps separately so a
|
||||
// pre-source run cannot pass its denominator off as a setup-excluding one.
|
||||
func generatorDispatched(line traceLine, generator string) bool {
|
||||
if line.ActionSkipped != "" || line.NextAction == nil {
|
||||
return false
|
||||
}
|
||||
switch line.NextAction.Source {
|
||||
case "":
|
||||
return generator != trace.ActionSourceModel
|
||||
case trace.ActionSourceSetup:
|
||||
return false
|
||||
default:
|
||||
return true
|
||||
}
|
||||
}
|
||||
|
||||
// actionUnattributed reports a dispatched step whose action names no producer.
|
||||
func actionUnattributed(line traceLine) bool {
|
||||
return line.ActionSkipped == "" && line.NextAction != nil && line.NextAction.Source == ""
|
||||
}
|
||||
|
||||
// findRunDirectory returns the run directory `sanderling test` created inside
|
||||
// seedDirectory. Names are UTC timestamps, so the last one sorted is the newest.
|
||||
func findRunDirectory(seedDirectory string) (string, error) {
|
||||
@@ -68,14 +132,35 @@ func summarizeRun(seedDirectory string) (string, traceSummary, error) {
|
||||
if err != nil {
|
||||
return "", traceSummary{}, err
|
||||
}
|
||||
summary, err := summarizeTrace(filepath.Join(seedDirectory, name, "trace.jsonl"))
|
||||
directory := filepath.Join(seedDirectory, name)
|
||||
generator, err := runGenerator(directory)
|
||||
if err != nil {
|
||||
return name, traceSummary{}, err
|
||||
}
|
||||
summary, err := summarizeTrace(filepath.Join(directory, "trace.jsonl"), generator)
|
||||
if err != nil {
|
||||
return name, traceSummary{}, err
|
||||
}
|
||||
return name, summary, nil
|
||||
}
|
||||
|
||||
func summarizeTrace(tracePath string) (traceSummary, error) {
|
||||
// runGenerator reads which picker drove the run. A trace recorded before
|
||||
// actions named their source cannot say whether an unstamped action came from
|
||||
// the seeded picker or from the spec's setup under the model picker, and the
|
||||
// two count differently.
|
||||
func runGenerator(runDirectory string) (string, error) {
|
||||
body, err := os.ReadFile(filepath.Join(runDirectory, "meta.json"))
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("read meta: %w", err)
|
||||
}
|
||||
var meta trace.Meta
|
||||
if err := json.Unmarshal(body, &meta); err != nil {
|
||||
return "", fmt.Errorf("decode meta: %w", err)
|
||||
}
|
||||
return meta.Generator, nil
|
||||
}
|
||||
|
||||
func summarizeTrace(tracePath, generator string) (traceSummary, error) {
|
||||
file, err := os.Open(tracePath)
|
||||
if err != nil {
|
||||
return traceSummary{}, fmt.Errorf("open trace: %w", err)
|
||||
@@ -103,6 +188,15 @@ func summarizeTrace(tracePath string) (traceSummary, error) {
|
||||
if !synthetic && line.Index > summary.Steps {
|
||||
summary.Steps = line.Index
|
||||
}
|
||||
if !synthetic && generatorDispatched(line, generator) {
|
||||
summary.Actions++
|
||||
}
|
||||
if !synthetic && actionUnattributed(line) {
|
||||
summary.UnattributedActions++
|
||||
}
|
||||
if line.PreconditionFailure != "" {
|
||||
summary.PreconditionFailures++
|
||||
}
|
||||
for _, property := range line.Violations {
|
||||
violated[property] = true
|
||||
recordViolation(&summary, line, property)
|
||||
|
||||
@@ -22,13 +22,70 @@ func observedStep(index int) trace.Step {
|
||||
}
|
||||
}
|
||||
|
||||
func actingStep(index int) trace.Step {
|
||||
step := observedStep(index)
|
||||
step.NextAction = &trace.Action{Kind: "tap", X: 12, Y: 34}
|
||||
return step
|
||||
}
|
||||
|
||||
func skippedActionStep(index int, reason string) trace.Step {
|
||||
step := actingStep(index)
|
||||
step.ActionSkipped = reason
|
||||
return step
|
||||
}
|
||||
|
||||
func setupStep(index int) trace.Step {
|
||||
step := observedStep(index)
|
||||
step.NextAction = &trace.Action{
|
||||
Kind: "InputText",
|
||||
Selector: "testTag:LoginScreen > testTag:LoginEmail",
|
||||
Text: "[email protected]",
|
||||
Source: trace.ActionSourceSetup,
|
||||
}
|
||||
return step
|
||||
}
|
||||
|
||||
func seededStep(index int) trace.Step {
|
||||
step := actingStep(index)
|
||||
step.NextAction.Source = trace.ActionSourceSeeded
|
||||
return step
|
||||
}
|
||||
|
||||
func modelStep(index int) trace.Step {
|
||||
step := actingStep(index)
|
||||
step.NextAction.Source = trace.ActionSourceModel
|
||||
return step
|
||||
}
|
||||
|
||||
func skippedModelStep(index int, reason string) trace.Step {
|
||||
step := modelStep(index)
|
||||
step.ActionSkipped = reason
|
||||
return step
|
||||
}
|
||||
|
||||
func writeRunDirectory(t *testing.T, seedDirectory, name string, steps []trace.Step) string {
|
||||
t.Helper()
|
||||
return writeRunDirectoryWithMeta(
|
||||
t,
|
||||
seedDirectory,
|
||||
name,
|
||||
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline"},
|
||||
steps,
|
||||
)
|
||||
}
|
||||
|
||||
func writeRunDirectoryWithMeta(
|
||||
t *testing.T,
|
||||
seedDirectory, name string,
|
||||
declared trace.Meta,
|
||||
steps []trace.Step,
|
||||
) string {
|
||||
t.Helper()
|
||||
directory := filepath.Join(seedDirectory, name)
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
meta, err := json.Marshal(trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline"})
|
||||
meta, err := json.Marshal(declared)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
@@ -72,6 +129,206 @@ func TestSummarizeRun_CleanRunIsCensored(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_CountsOnlyStepsThatDispatchedAnAction(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
|
||||
actingStep(1),
|
||||
observedStep(2),
|
||||
actingStep(3),
|
||||
skippedActionStep(4, "no_target"),
|
||||
skippedActionStep(5, "unresolved_selector"),
|
||||
skippedActionStep(6, "missing_key"),
|
||||
skippedActionStep(7, "zero_duration_wait"),
|
||||
skippedActionStep(8, "app_left_foreground"),
|
||||
skippedActionStep(9, "apply_error"),
|
||||
actingStep(10),
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Steps != 10 {
|
||||
t.Errorf("steps: got %d, want 10", summary.Steps)
|
||||
}
|
||||
if summary.Actions != 3 {
|
||||
t.Errorf("actions: got %d, want 3 (one step chose nothing and six were never dispatched)", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_ModelRunLeavesTheSetupsLoginOutOfTheActionCount(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
|
||||
trace.Meta{Seed: 11, Platform: "web", Arm: "llm-identifier", Generator: "llm"},
|
||||
[]trace.Step{
|
||||
setupStep(1),
|
||||
setupStep(2),
|
||||
setupStep(3),
|
||||
modelStep(4),
|
||||
skippedModelStep(5, "unresolved_selector"),
|
||||
modelStep(6),
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Steps != 6 {
|
||||
t.Errorf("steps: got %d, want 6", summary.Steps)
|
||||
}
|
||||
if summary.Actions != 2 {
|
||||
t.Errorf("actions: got %d, want 2 (three login steps were setup's and one generator action was thrown away)", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_ModelRunWithoutSetupCountsEveryDispatchedStep(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
|
||||
trace.Meta{Seed: 11, Platform: "web", Arm: "llm-identifier", Generator: "llm"},
|
||||
[]trace.Step{modelStep(1), modelStep(2), modelStep(3), modelStep(4)})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 4 {
|
||||
t.Errorf("actions: got %d, want 4", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_SeededRunLeavesTheSetupsLoginOutOfTheActionCount(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
|
||||
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline", Generator: "seeded"},
|
||||
[]trace.Step{
|
||||
setupStep(1),
|
||||
setupStep(2),
|
||||
setupStep(3),
|
||||
seededStep(4),
|
||||
seededStep(5),
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 2 {
|
||||
t.Errorf("actions: got %d, want 2 (three login steps were setup's), which is the same "+
|
||||
"rule the model arm is counted by", summary.Actions)
|
||||
}
|
||||
if summary.UnattributedActions != 0 {
|
||||
t.Errorf("unattributed actions: got %d, want 0: every step named its source",
|
||||
summary.UnattributedActions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_SeededRunWithoutSetupCountsEveryDispatchedStep(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
|
||||
trace.Meta{Seed: 11, Platform: "web", Arm: "seeded-baseline", Generator: "seeded"},
|
||||
[]trace.Step{seededStep(1), seededStep(2), seededStep(3), seededStep(4)})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 4 {
|
||||
t.Errorf("actions: got %d, want 4", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
// Traces recorded before actions named their source cannot be re-attributed
|
||||
// after the fact, so each arm keeps the count it was already reported with: the
|
||||
// seeded arm counts every dispatched step, the model arm counts only what the
|
||||
// model stamped. UnattributedActions is how such a run says so rather than
|
||||
// passing its old denominator off as a setup-excluding one.
|
||||
func TestSummarizeRun_TraceWithoutSourcesKeepsTheCountItWasReportedWith(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
generator string
|
||||
actions int
|
||||
unattributedActions int
|
||||
}{
|
||||
{generator: "seeded", actions: 5, unattributedActions: 5},
|
||||
{generator: "llm", actions: 0, unattributedActions: 5},
|
||||
} {
|
||||
t.Run(testCase.generator, func(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectoryWithMeta(t, seedDirectory, "20260812-090000",
|
||||
trace.Meta{Seed: 11, Platform: "web", Generator: testCase.generator},
|
||||
[]trace.Step{
|
||||
actingStep(1), actingStep(2), actingStep(3), actingStep(4), actingStep(5),
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != testCase.actions {
|
||||
t.Errorf("actions: got %d, want %d", summary.Actions, testCase.actions)
|
||||
}
|
||||
if summary.UnattributedActions != testCase.unattributedActions {
|
||||
t.Errorf("unattributed actions: got %d, want %d",
|
||||
summary.UnattributedActions, testCase.unattributedActions)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_SkipReasonsEachSuppressTheAction(t *testing.T) {
|
||||
for _, reason := range []string{
|
||||
"no_target", "unresolved_selector", "missing_key",
|
||||
"zero_duration_wait", "app_left_foreground", "apply_error",
|
||||
} {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
|
||||
actingStep(1), skippedActionStep(2, reason),
|
||||
})
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 1 {
|
||||
t.Errorf("%s: actions got %d, want 1", reason, summary.Actions)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeTrace_NullActionIsNoAction(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
directory := writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{observedStep(1)})
|
||||
lines := "{\"step\":1,\"hierarchy\":{},\"next_action\":null}\n" +
|
||||
"{\"step\":2,\"hierarchy\":{},\"next_action\":{\"kind\":\"tap\"},\"action_skipped\":\"\"}\n"
|
||||
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte(lines), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 1 {
|
||||
t.Errorf("actions: got %d, want 1", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_FinalizeLineIsNotAnAction(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
finalize := trace.Step{
|
||||
Index: 3,
|
||||
Timestamp: time.Now().UTC(),
|
||||
NextAction: &trace.Action{Kind: "tap"},
|
||||
Violations: []string{"eventuallySettles"},
|
||||
}
|
||||
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{actingStep(1), actingStep(2), finalize})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.Actions != 2 {
|
||||
t.Errorf("actions: got %d, want 2", summary.Actions)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSummarizeRun_UsesWitnessOriginNotDetectionStep(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
violating := observedStep(9)
|
||||
@@ -198,3 +455,49 @@ func TestSummarizeTrace_MalformedLine(t *testing.T) {
|
||||
t.Fatal("expected an error for a malformed trace line")
|
||||
}
|
||||
}
|
||||
|
||||
// A run whose app never came to the foreground has to be countable off the
|
||||
// summary. Without it the campaign row for a run that never started is a row of
|
||||
// zero steps and no violations, which is what a clean short run looks like too.
|
||||
func TestSummarizeRun_CountsPreconditionFailures(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
|
||||
{Index: 0, PreconditionFailure: "app_not_in_foreground"},
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.PreconditionFailures != 1 {
|
||||
t.Errorf("precondition failures: got %d, want 1", summary.PreconditionFailures)
|
||||
}
|
||||
if summary.Steps != 0 {
|
||||
t.Errorf("steps: got %d, want 0; the run never observed anything", summary.Steps)
|
||||
}
|
||||
}
|
||||
|
||||
// The same fact mid-run: steps the scope guard could not bring the app back for
|
||||
// are still steps, and they are counted separately from the ones that explored.
|
||||
func TestSummarizeRun_CountsMidRunPreconditionFailures(t *testing.T) {
|
||||
seedDirectory := t.TempDir()
|
||||
outsideApp := func(index int) trace.Step {
|
||||
step := actingStep(index)
|
||||
step.PreconditionFailure = "app_not_in_foreground"
|
||||
return step
|
||||
}
|
||||
writeRunDirectory(t, seedDirectory, "20260812-090000", []trace.Step{
|
||||
actingStep(1), outsideApp(2), outsideApp(3), actingStep(4),
|
||||
})
|
||||
|
||||
_, summary, err := summarizeRun(seedDirectory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if summary.PreconditionFailures != 2 {
|
||||
t.Errorf("precondition failures: got %d, want 2", summary.PreconditionFailures)
|
||||
}
|
||||
if summary.Steps != 4 {
|
||||
t.Errorf("steps: got %d, want 4", summary.Steps)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,55 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"regexp"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// The three models the sample is drawn from, in the capability order
|
||||
// model-implementations.md fixed before any implementation was generated.
|
||||
var capabilityOrder = []string{"Sonnet 5", "Opus 5", "Fable 5"}
|
||||
|
||||
var (
|
||||
implementationName = regexp.MustCompile(`^impl-\d+$`)
|
||||
nonAlphanumeric = regexp.MustCompile(`[^a-z0-9]`)
|
||||
)
|
||||
|
||||
func loadAssignment(path string) (map[string]string, error) {
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read assignment: %w", err)
|
||||
}
|
||||
assignments := map[string]string{}
|
||||
for _, row := range parseTableRows(string(body)) {
|
||||
name := row.cell(0)
|
||||
if !implementationName.MatchString(name) {
|
||||
continue
|
||||
}
|
||||
model, ok := canonicalModel(row.cell(1))
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("%s line %d: %s is assigned model %q, which is none of %s",
|
||||
path, row.Line, name, row.cell(1), strings.Join(capabilityOrder, ", "))
|
||||
}
|
||||
if existing, seen := assignments[name]; seen && existing != model {
|
||||
return nil, fmt.Errorf("%s line %d: %s is assigned to both %s and %s",
|
||||
path, row.Line, name, existing, model)
|
||||
}
|
||||
assignments[name] = model
|
||||
}
|
||||
if len(assignments) == 0 {
|
||||
return nil, fmt.Errorf("%s maps no implementation to a model", path)
|
||||
}
|
||||
return assignments, nil
|
||||
}
|
||||
|
||||
func canonicalModel(value string) (string, bool) {
|
||||
key := nonAlphanumeric.ReplaceAllString(strings.ToLower(value), "")
|
||||
for _, model := range capabilityOrder {
|
||||
if key == nonAlphanumeric.ReplaceAllString(strings.ToLower(model), "") {
|
||||
return model, true
|
||||
}
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
@@ -0,0 +1,345 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"maps"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
const (
|
||||
sweepManifestFileName = "sweep.json"
|
||||
sweepRecordsFileName = "implementations.jsonl"
|
||||
campaignManifestFileName = "campaign.json"
|
||||
campaignRecordsFileName = "runs.jsonl"
|
||||
traceFileName = "trace.jsonl"
|
||||
surfacesExtractor = "locatableSurfaces"
|
||||
maxRecordBytes = 4 * 1024 * 1024
|
||||
maxTraceLineBytes = 16 * 1024 * 1024
|
||||
)
|
||||
|
||||
// Run exclusion reasons, kept in the vocabulary analyze already uses so the two
|
||||
// tools describe the same run the same way.
|
||||
const (
|
||||
reasonLaunchError = "launch error"
|
||||
reasonTimedOut = "timed out"
|
||||
reasonNonzeroExit = "nonzero exit"
|
||||
reasonTraceError = "unreadable trace"
|
||||
)
|
||||
|
||||
type sweepManifest struct {
|
||||
SpecPath string `json:"spec_path"`
|
||||
Implementations []struct {
|
||||
Name string `json:"name"`
|
||||
} `json:"implementations"`
|
||||
}
|
||||
|
||||
type sweepRunRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
ExitCode int `json:"exit_code"`
|
||||
LaunchError string `json:"launch_error"`
|
||||
CampaignDirectory string `json:"campaign_directory"`
|
||||
}
|
||||
|
||||
type sweepImplementationRecord struct {
|
||||
Name string `json:"implementation"`
|
||||
FailedStage string `json:"failed_stage"`
|
||||
Error string `json:"error"`
|
||||
Runs []sweepRunRecord `json:"runs"`
|
||||
}
|
||||
|
||||
type campaignRunRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
ExitCode int `json:"exit_code"`
|
||||
LaunchError string `json:"launch_error"`
|
||||
TimedOut bool `json:"timed_out"`
|
||||
TraceError string `json:"trace_error"`
|
||||
RunDirectory string `json:"run_directory"`
|
||||
ViolatedProperties []string `json:"violated_properties"`
|
||||
}
|
||||
|
||||
// checkerVerdict is one implementation's whole checker side, pooled across the
|
||||
// seeds it was swept at.
|
||||
type checkerVerdict struct {
|
||||
Implementation string
|
||||
FailedStage string
|
||||
FailedError string
|
||||
RunsRecorded int
|
||||
RunsUsable int
|
||||
ExcludedByReason map[string]int
|
||||
FiredProperties []string
|
||||
// SurfacesObserved holds every locatable surface seen true on at least one
|
||||
// step of at least one usable run. A surface missing from it was never
|
||||
// located across the whole sweep of this implementation.
|
||||
SurfacesObserved map[string]bool
|
||||
SurfacesKnown bool
|
||||
TraceErrors []string
|
||||
}
|
||||
|
||||
func (v checkerVerdict) fired() bool { return len(v.FiredProperties) > 0 }
|
||||
|
||||
type checkerSide struct {
|
||||
Directory string
|
||||
SpecPath string
|
||||
Planned []string
|
||||
Verdicts []checkerVerdict
|
||||
}
|
||||
|
||||
func loadChecker(directory string) (checkerSide, error) {
|
||||
body, err := os.ReadFile(filepath.Join(directory, sweepManifestFileName))
|
||||
if err != nil {
|
||||
return checkerSide{}, fmt.Errorf("read %s: %w", sweepManifestFileName, err)
|
||||
}
|
||||
var declared sweepManifest
|
||||
if err := json.Unmarshal(body, &declared); err != nil {
|
||||
return checkerSide{}, fmt.Errorf("parse %s in %s: %w", sweepManifestFileName, directory, err)
|
||||
}
|
||||
side := checkerSide{Directory: directory, SpecPath: declared.SpecPath}
|
||||
for _, planned := range declared.Implementations {
|
||||
side.Planned = append(side.Planned, planned.Name)
|
||||
}
|
||||
|
||||
records, err := readSweepRecords(filepath.Join(directory, sweepRecordsFileName))
|
||||
if err != nil {
|
||||
return checkerSide{}, err
|
||||
}
|
||||
for _, record := range records {
|
||||
verdict, err := readImplementation(directory, record)
|
||||
if err != nil {
|
||||
return checkerSide{}, err
|
||||
}
|
||||
side.Verdicts = append(side.Verdicts, verdict)
|
||||
}
|
||||
slices.SortFunc(side.Verdicts, func(a, b checkerVerdict) int {
|
||||
return strings.Compare(a.Implementation, b.Implementation)
|
||||
})
|
||||
return side, nil
|
||||
}
|
||||
|
||||
func readSweepRecords(path string) ([]sweepImplementationRecord, error) {
|
||||
file, err := os.Open(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", sweepRecordsFileName, err)
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
var records []sweepImplementationRecord
|
||||
scanner := bufio.NewScanner(file)
|
||||
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
|
||||
lineNumber := 0
|
||||
seen := map[string]bool{}
|
||||
for scanner.Scan() {
|
||||
lineNumber++
|
||||
raw := strings.TrimSpace(scanner.Text())
|
||||
if raw == "" {
|
||||
continue
|
||||
}
|
||||
var record sweepImplementationRecord
|
||||
if err := json.Unmarshal([]byte(raw), &record); err != nil {
|
||||
return nil, fmt.Errorf("%s line %d: %w", sweepRecordsFileName, lineNumber, err)
|
||||
}
|
||||
if record.Name == "" {
|
||||
return nil, fmt.Errorf("%s line %d names no implementation", sweepRecordsFileName, lineNumber)
|
||||
}
|
||||
if seen[record.Name] {
|
||||
return nil, fmt.Errorf("%s line %d records %s a second time: its runs would be pooled twice",
|
||||
sweepRecordsFileName, lineNumber, record.Name)
|
||||
}
|
||||
seen[record.Name] = true
|
||||
records = append(records, record)
|
||||
}
|
||||
if err := scanner.Err(); err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", sweepRecordsFileName, err)
|
||||
}
|
||||
return records, nil
|
||||
}
|
||||
|
||||
func readImplementation(sweepDirectory string, record sweepImplementationRecord) (checkerVerdict, error) {
|
||||
verdict := checkerVerdict{
|
||||
Implementation: record.Name,
|
||||
FailedStage: record.FailedStage,
|
||||
FailedError: record.Error,
|
||||
ExcludedByReason: map[string]int{},
|
||||
SurfacesObserved: map[string]bool{},
|
||||
}
|
||||
fired := map[string]bool{}
|
||||
for _, run := range record.Runs {
|
||||
verdict.RunsRecorded++
|
||||
if reason := sweepRunExcludedBecause(run); reason != "" {
|
||||
verdict.ExcludedByReason[reason]++
|
||||
continue
|
||||
}
|
||||
directory := resolveCampaignDirectory(sweepDirectory, record.Name, run)
|
||||
campaignRuns, err := readCampaignRuns(directory)
|
||||
if err != nil {
|
||||
return checkerVerdict{}, err
|
||||
}
|
||||
for _, campaignRun := range campaignRuns {
|
||||
if reason := excludedBecause(campaignRun); reason != "" {
|
||||
verdict.ExcludedByReason[reason]++
|
||||
continue
|
||||
}
|
||||
verdict.RunsUsable++
|
||||
for _, property := range campaignRun.ViolatedProperties {
|
||||
fired[property] = true
|
||||
}
|
||||
observed, err := readObservedSurfaces(filepath.Join(directory, campaignRun.RunDirectory))
|
||||
if err != nil {
|
||||
verdict.TraceErrors = append(verdict.TraceErrors, err.Error())
|
||||
continue
|
||||
}
|
||||
verdict.SurfacesKnown = true
|
||||
for surface, seen := range observed {
|
||||
if seen {
|
||||
verdict.SurfacesObserved[surface] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(fired) > 0 {
|
||||
verdict.FiredProperties = slices.Sorted(maps.Keys(fired))
|
||||
}
|
||||
if len(verdict.ExcludedByReason) == 0 {
|
||||
verdict.ExcludedByReason = nil
|
||||
}
|
||||
return verdict, nil
|
||||
}
|
||||
|
||||
// resolveCampaignDirectory prefers the path the sweep recorded and falls back to
|
||||
// the layout it names, so a sweep directory read on another machine than the one
|
||||
// that wrote its absolute paths still resolves.
|
||||
func resolveCampaignDirectory(sweepDirectory, name string, run sweepRunRecord) string {
|
||||
if run.CampaignDirectory != "" {
|
||||
if _, err := os.Stat(run.CampaignDirectory); err == nil {
|
||||
return run.CampaignDirectory
|
||||
}
|
||||
}
|
||||
return filepath.Join(sweepDirectory, name, "seed-"+strconv.FormatInt(run.Seed, 10))
|
||||
}
|
||||
|
||||
func readCampaignRuns(directory string) ([]campaignRunRecord, error) {
|
||||
if _, err := os.Stat(filepath.Join(directory, campaignManifestFileName)); err != nil {
|
||||
return nil, fmt.Errorf("read %s in %s: %w", campaignManifestFileName, directory, err)
|
||||
}
|
||||
file, err := os.Open(filepath.Join(directory, campaignRecordsFileName))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", campaignRecordsFileName, err)
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
var records []campaignRunRecord
|
||||
scanner := bufio.NewScanner(file)
|
||||
scanner.Buffer(make([]byte, 0, 64*1024), maxRecordBytes)
|
||||
lineNumber := 0
|
||||
for scanner.Scan() {
|
||||
lineNumber++
|
||||
raw := strings.TrimSpace(scanner.Text())
|
||||
if raw == "" {
|
||||
continue
|
||||
}
|
||||
var record campaignRunRecord
|
||||
if err := json.Unmarshal([]byte(raw), &record); err != nil {
|
||||
return nil, fmt.Errorf("%s line %d in %s: %w", campaignRecordsFileName, lineNumber, directory, err)
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
if err := scanner.Err(); err != nil {
|
||||
return nil, fmt.Errorf("read %s in %s: %w", campaignRecordsFileName, directory, err)
|
||||
}
|
||||
return records, nil
|
||||
}
|
||||
|
||||
// sweepRunExcludedBecause reads the outcome of the campaign process itself,
|
||||
// which the records inside its directory cannot report. One sweep run is one
|
||||
// campaign of one seed, so a campaign that died left a runs.jsonl that is
|
||||
// partial or empty, and scoring the runs it did write reads the seeds it never
|
||||
// reached as agreement.
|
||||
func sweepRunExcludedBecause(record sweepRunRecord) string {
|
||||
switch {
|
||||
case record.LaunchError != "":
|
||||
return reasonLaunchError
|
||||
case record.ExitCode != 0:
|
||||
return reasonNonzeroExit
|
||||
default:
|
||||
return ""
|
||||
}
|
||||
}
|
||||
|
||||
func excludedBecause(record campaignRunRecord) string {
|
||||
switch {
|
||||
case record.LaunchError != "":
|
||||
return reasonLaunchError
|
||||
case record.TimedOut:
|
||||
return reasonTimedOut
|
||||
case record.ExitCode != 0:
|
||||
return reasonNonzeroExit
|
||||
case record.TraceError != "":
|
||||
return reasonTraceError
|
||||
default:
|
||||
return ""
|
||||
}
|
||||
}
|
||||
|
||||
type traceLine struct {
|
||||
ExtractorChanges map[string]struct {
|
||||
Curr json.RawMessage `json:"curr"`
|
||||
} `json:"extractor_changes"`
|
||||
}
|
||||
|
||||
// readObservedSurfaces replays one run's locatableSurfaces readings. The
|
||||
// verifier emits every extractor as a change on the first snapshot and only on
|
||||
// a difference afterwards, so a surface true on any recorded change was located
|
||||
// at least once, and one absent from every change was never located at all.
|
||||
func readObservedSurfaces(runDirectory string) (map[string]bool, error) {
|
||||
path := filepath.Join(runDirectory, traceFileName)
|
||||
file, err := os.Open(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", path, err)
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
observed := map[string]bool{}
|
||||
found := false
|
||||
scanner := bufio.NewScanner(file)
|
||||
scanner.Buffer(make([]byte, 0, 64*1024), maxTraceLineBytes)
|
||||
lineNumber := 0
|
||||
for scanner.Scan() {
|
||||
lineNumber++
|
||||
raw := strings.TrimSpace(scanner.Text())
|
||||
if raw == "" {
|
||||
continue
|
||||
}
|
||||
var line traceLine
|
||||
if err := json.Unmarshal([]byte(raw), &line); err != nil {
|
||||
return nil, fmt.Errorf("%s line %d: %w", path, lineNumber, err)
|
||||
}
|
||||
change, present := line.ExtractorChanges[surfacesExtractor]
|
||||
if !present || len(change.Curr) == 0 {
|
||||
continue
|
||||
}
|
||||
var reading map[string]bool
|
||||
if err := json.Unmarshal(change.Curr, &reading); err != nil {
|
||||
return nil, fmt.Errorf("%s line %d: %s is not an object of booleans: %w",
|
||||
path, lineNumber, surfacesExtractor, err)
|
||||
}
|
||||
found = true
|
||||
for surface, located := range reading {
|
||||
if located {
|
||||
observed[surface] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
if err := scanner.Err(); err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", path, err)
|
||||
}
|
||||
if !found {
|
||||
return nil, fmt.Errorf("%s records no %s reading: this run cannot say whether a surface was located",
|
||||
path, surfacesExtractor)
|
||||
}
|
||||
return observed, nil
|
||||
}
|
||||
@@ -0,0 +1,308 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// fixtureRun is one campaign of one implementation at one seed, in the shape
|
||||
// implementation-sweep and campaign write together.
|
||||
type fixtureRun struct {
|
||||
Seed int64
|
||||
ExitCode int
|
||||
// CampaignExitCode is the campaign process's own exit status, which the
|
||||
// sweep records beside the campaign directory. It is not the exit status of
|
||||
// the run inside that campaign.
|
||||
CampaignExitCode int
|
||||
TimedOut bool
|
||||
Violated []string
|
||||
// Surfaces is the locatableSurfaces reading the trace records. A nil map
|
||||
// with NoTrace false still writes a reading of every surface false.
|
||||
Surfaces map[string]bool
|
||||
NoTrace bool
|
||||
}
|
||||
|
||||
type fixtureReview struct {
|
||||
Overall string
|
||||
Clauses map[string]string
|
||||
Minutes int
|
||||
}
|
||||
|
||||
type fixtureImplementation struct {
|
||||
Name string
|
||||
Model string
|
||||
FailedStage string
|
||||
Runs []fixtureRun
|
||||
Review *fixtureReview
|
||||
RawReview string
|
||||
Adjudicated map[string]string
|
||||
}
|
||||
|
||||
type fixture struct {
|
||||
Sweep string
|
||||
Reviews string
|
||||
Assignment string
|
||||
Mapping string
|
||||
}
|
||||
|
||||
const defaultMapping = `
|
||||
| property | clauses | surfaces |
|
||||
| --- | --- | --- |
|
||||
| sentOnlyAfterConfirmation | R5 R15 | stateWords |
|
||||
| serverHoldsEachMessageOnce | R15 | none |
|
||||
| unsentReachesZero | R14 | unsentCount |
|
||||
|
||||
| surface | never observed | note |
|
||||
| --- | --- | --- |
|
||||
| composer | unlocatable | R1 obliges one |
|
||||
| stateWords | unlocatable | R4 obliges one on every composed message |
|
||||
| unsentCount | inconclusive | R6 hides the count at zero |
|
||||
`
|
||||
|
||||
func writeFixture(t *testing.T, implementations []fixtureImplementation, mappingBody string) fixture {
|
||||
t.Helper()
|
||||
root := t.TempDir()
|
||||
built := fixture{
|
||||
Sweep: filepath.Join(root, "sweep"),
|
||||
Reviews: filepath.Join(root, "reviews"),
|
||||
Assignment: filepath.Join(root, "assignment.md"),
|
||||
Mapping: filepath.Join(root, "property-clauses.md"),
|
||||
}
|
||||
mustMkdir(t, built.Sweep)
|
||||
mustMkdir(t, built.Reviews)
|
||||
mustWrite(t, built.Mapping, mappingBody)
|
||||
|
||||
var planned []map[string]any
|
||||
var assignmentRows []string
|
||||
assignmentRows = append(assignmentRows, "| implementation | model |", "| --- | --- |")
|
||||
var records []string
|
||||
for _, implementation := range implementations {
|
||||
planned = append(planned, map[string]any{"name": implementation.Name})
|
||||
if implementation.Model != "" {
|
||||
assignmentRows = append(assignmentRows,
|
||||
fmt.Sprintf("| %s | %s |", implementation.Name, implementation.Model))
|
||||
}
|
||||
records = append(records, writeImplementation(t, built, implementation))
|
||||
writeReviewFiles(t, built.Reviews, implementation)
|
||||
}
|
||||
mustWrite(t, built.Assignment, strings.Join(assignmentRows, "\n")+"\n")
|
||||
mustWriteJSON(t, filepath.Join(built.Sweep, sweepManifestFileName), map[string]any{
|
||||
"generator": "seeded",
|
||||
"platform": "web",
|
||||
"spec_path": "paper/experiments/e4/spec/spec.ts",
|
||||
"max_steps": 400,
|
||||
"seeds": []int{1, 2},
|
||||
"host": "anton",
|
||||
"implementations": planned,
|
||||
})
|
||||
mustWrite(t, filepath.Join(built.Sweep, sweepRecordsFileName), strings.Join(records, "\n")+"\n")
|
||||
return built
|
||||
}
|
||||
|
||||
func writeImplementation(t *testing.T, built fixture, implementation fixtureImplementation) string {
|
||||
t.Helper()
|
||||
record := map[string]any{"implementation": implementation.Name}
|
||||
if implementation.FailedStage != "" {
|
||||
record["failed_stage"] = implementation.FailedStage
|
||||
record["error"] = "bun run build exited 1"
|
||||
return encode(t, record)
|
||||
}
|
||||
var runs []map[string]any
|
||||
for _, run := range implementation.Runs {
|
||||
seedText := strconv.FormatInt(run.Seed, 10)
|
||||
campaignDirectory := filepath.Join(built.Sweep, implementation.Name, "seed-"+seedText)
|
||||
mustMkdir(t, campaignDirectory)
|
||||
mustWriteJSON(t, filepath.Join(campaignDirectory, campaignManifestFileName), map[string]any{
|
||||
"arm": implementation.Name,
|
||||
"generator": "seeded",
|
||||
"platform": "web",
|
||||
"max_steps": 400,
|
||||
"seeds": []int64{run.Seed},
|
||||
})
|
||||
runDirectory := filepath.Join("seed-"+seedText, "20260820T110000Z")
|
||||
campaignRun := map[string]any{
|
||||
"seed": run.Seed,
|
||||
"exit_code": run.ExitCode,
|
||||
"steps": 400,
|
||||
"actions": 380,
|
||||
"run_directory": runDirectory,
|
||||
}
|
||||
if run.TimedOut {
|
||||
campaignRun["timed_out"] = true
|
||||
}
|
||||
if len(run.Violated) > 0 {
|
||||
campaignRun["violated_properties"] = run.Violated
|
||||
}
|
||||
mustWrite(t, filepath.Join(campaignDirectory, campaignRecordsFileName), encode(t, campaignRun)+"\n")
|
||||
if !run.NoTrace {
|
||||
writeTrace(t, filepath.Join(campaignDirectory, runDirectory), run.Surfaces)
|
||||
}
|
||||
runs = append(runs, map[string]any{
|
||||
"seed": run.Seed,
|
||||
"exit_code": run.CampaignExitCode,
|
||||
"campaign_directory": campaignDirectory,
|
||||
})
|
||||
}
|
||||
record["runs"] = runs
|
||||
return encode(t, record)
|
||||
}
|
||||
|
||||
// writeTrace writes the one line the verifier emits on the first snapshot, when
|
||||
// every extractor is reported as a change from null.
|
||||
func writeTrace(t *testing.T, directory string, surfaces map[string]bool) {
|
||||
t.Helper()
|
||||
mustMkdir(t, directory)
|
||||
reading := map[string]bool{}
|
||||
for _, surface := range []string{"appRoot", "composer", "submit", "stateWords",
|
||||
"unsentCount", "pendingIndicator", "offlineIndicator", "retryControl"} {
|
||||
reading[surface] = surfaces[surface]
|
||||
}
|
||||
line := map[string]any{
|
||||
"step": 1,
|
||||
"extractor_changes": map[string]any{
|
||||
surfacesExtractor: map[string]any{"prev": nil, "curr": reading},
|
||||
},
|
||||
}
|
||||
mustWrite(t, filepath.Join(directory, traceFileName), encode(t, line)+"\n")
|
||||
}
|
||||
|
||||
func writeReviewFiles(t *testing.T, directory string, implementation fixtureImplementation) {
|
||||
t.Helper()
|
||||
if implementation.RawReview != "" {
|
||||
mustWrite(t, filepath.Join(directory, implementation.Name+".md"), implementation.RawReview)
|
||||
return
|
||||
}
|
||||
if implementation.Review == nil {
|
||||
return
|
||||
}
|
||||
mustWrite(t, filepath.Join(directory, implementation.Name+".md"), renderReview(*implementation.Review))
|
||||
if len(implementation.Adjudicated) == 0 {
|
||||
return
|
||||
}
|
||||
rows := []string{"| clause | verdict | why |", "| --- | --- | --- |"}
|
||||
for _, clause := range allClauses() {
|
||||
label, resolved := implementation.Adjudicated[clause]
|
||||
if !resolved {
|
||||
continue
|
||||
}
|
||||
rows = append(rows, fmt.Sprintf("| %s | %s | joint reread |", clause, label))
|
||||
}
|
||||
mustWrite(t, filepath.Join(directory, implementation.Name+"-adjudication.md"),
|
||||
strings.Join(rows, "\n")+"\n")
|
||||
}
|
||||
|
||||
func renderReview(review fixtureReview) string {
|
||||
minutes := review.Minutes
|
||||
if minutes == 0 {
|
||||
minutes = 45
|
||||
}
|
||||
lines := []string{
|
||||
"# review",
|
||||
"",
|
||||
"reviewer: Jane",
|
||||
"date: 2026-09-02",
|
||||
fmt.Sprintf("minutes: %d", minutes),
|
||||
"",
|
||||
"| clause | verdict | justification | steps |",
|
||||
"| --- | --- | --- | --- |",
|
||||
}
|
||||
for _, clause := range allClauses() {
|
||||
label := review.Clauses[clause]
|
||||
if label == "" {
|
||||
label = clauseMeets
|
||||
}
|
||||
lines = append(lines, fmt.Sprintf("| %s | %s | seen by hand | offline, compose, online |", clause, label))
|
||||
}
|
||||
lines = append(lines, "", "overall: "+review.Overall, "")
|
||||
return strings.Join(lines, "\n")
|
||||
}
|
||||
|
||||
func runTool(t *testing.T, built fixture) (result, string) {
|
||||
t.Helper()
|
||||
var stdout, stderr bytes.Buffer
|
||||
jsonPath := filepath.Join(t.TempDir(), "matrix.json")
|
||||
err := run([]string{
|
||||
"--sweep", built.Sweep,
|
||||
"--reviews", built.Reviews,
|
||||
"--assignment", built.Assignment,
|
||||
"--property-clauses", built.Mapping,
|
||||
"--json", jsonPath,
|
||||
}, &stdout, &stderr)
|
||||
if err != nil {
|
||||
t.Fatalf("run: %v\nstderr: %s", err, stderr.String())
|
||||
}
|
||||
body, err := os.ReadFile(jsonPath)
|
||||
if err != nil {
|
||||
t.Fatalf("read emitted summary: %v", err)
|
||||
}
|
||||
var emitted result
|
||||
if err := json.Unmarshal(body, &emitted); err != nil {
|
||||
t.Fatalf("emitted summary is not valid JSON: %v", err)
|
||||
}
|
||||
return emitted, stdout.String()
|
||||
}
|
||||
|
||||
func cleanRun(seed int64, violated ...string) fixtureRun {
|
||||
return fixtureRun{
|
||||
Seed: seed,
|
||||
Violated: violated,
|
||||
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": true},
|
||||
}
|
||||
}
|
||||
|
||||
func mustMkdir(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
if err := os.MkdirAll(path, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
func mustWrite(t *testing.T, path, body string) {
|
||||
t.Helper()
|
||||
mustMkdir(t, filepath.Dir(path))
|
||||
if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
func mustWriteJSON(t *testing.T, path string, value any) {
|
||||
t.Helper()
|
||||
mustWrite(t, path, encode(t, value)+"\n")
|
||||
}
|
||||
|
||||
func encode(t *testing.T, value any) string {
|
||||
t.Helper()
|
||||
body, err := json.Marshal(value)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return string(body)
|
||||
}
|
||||
|
||||
func outcomeFor(t *testing.T, emitted result, name string) implementationOutcome {
|
||||
t.Helper()
|
||||
for _, row := range emitted.Outcomes {
|
||||
if row.Implementation == name {
|
||||
return row
|
||||
}
|
||||
}
|
||||
t.Fatalf("%s carries no cell; excluded as %v", name, emitted.Excluded)
|
||||
return implementationOutcome{}
|
||||
}
|
||||
|
||||
func exclusionFor(t *testing.T, emitted result, name string) exclusion {
|
||||
t.Helper()
|
||||
for _, entry := range emitted.Excluded {
|
||||
if entry.Implementation == name {
|
||||
return entry
|
||||
}
|
||||
}
|
||||
t.Fatalf("%s was not excluded; it scored %v", name, emitted.Outcomes)
|
||||
return exclusion{}
|
||||
}
|
||||
@@ -0,0 +1,120 @@
|
||||
// Command confusion-matrix cross-tabulates e4's checker verdicts against the
|
||||
// blind human review, which is the measure model-implementations.md
|
||||
// pre-registers: an implementation whose own suite passed, scored on whether a
|
||||
// property fired and on whether the reviewer found a defect.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"time"
|
||||
)
|
||||
|
||||
const usage = `confusion-matrix cross-tabulates the e4 checker against the blind human review.
|
||||
|
||||
Usage:
|
||||
confusion-matrix --sweep <dir> --reviews <dir> --assignment <path> --property-clauses <path> [--json <path>]
|
||||
|
||||
--sweep is the directory implementation-sweep wrote: sweep.json, implementations.jsonl
|
||||
and the per-implementation campaign directories under it.
|
||||
|
||||
--reviews holds one impl-NN.md per implementation in the shape review-protocol.md
|
||||
fixes, plus any impl-NN-adjudication.md whose resolved labels replace the first
|
||||
rater's for the clauses it names.
|
||||
|
||||
--assignment is implementations/assignment.md, the blinded implementation-to-model
|
||||
mapping, opened only after the last verdict is filed.
|
||||
|
||||
--property-clauses declares which requirement clauses each property covers and which
|
||||
locatable surfaces it reads. Without it a fired property cannot be scored against a
|
||||
clause and a portability miss cannot be told from a clean run.
|
||||
`
|
||||
|
||||
func run(arguments []string, stdout, stderr io.Writer) error {
|
||||
flagSet := flag.NewFlagSet("confusion-matrix", flag.ContinueOnError)
|
||||
flagSet.SetOutput(stderr)
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprint(stderr, usage)
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
var sweepDirectory string
|
||||
var reviewsDirectory string
|
||||
var assignmentPath string
|
||||
var mappingPath string
|
||||
var jsonPath string
|
||||
flagSet.StringVar(&sweepDirectory, "sweep", "", "directory implementation-sweep wrote")
|
||||
flagSet.StringVar(&reviewsDirectory, "reviews", "", "directory holding impl-NN.md verdict forms")
|
||||
flagSet.StringVar(&assignmentPath, "assignment", "", "implementations/assignment.md, the implementation-to-model mapping")
|
||||
flagSet.StringVar(&mappingPath, "property-clauses", "", "the declared property-to-clause and surface mapping")
|
||||
flagSet.StringVar(&jsonPath, "json", "", "write the machine-readable summary here, or - for stdout")
|
||||
if err := flagSet.Parse(arguments); err != nil {
|
||||
return err
|
||||
}
|
||||
// Every missing flag is named together, in flag order: stopping at the
|
||||
// first turns one rerun into one rerun per missing flag.
|
||||
var missing []error
|
||||
for _, required := range []struct {
|
||||
name string
|
||||
value string
|
||||
}{
|
||||
{"--sweep", sweepDirectory},
|
||||
{"--reviews", reviewsDirectory},
|
||||
{"--assignment", assignmentPath},
|
||||
{"--property-clauses", mappingPath},
|
||||
} {
|
||||
if required.value == "" {
|
||||
missing = append(missing, fmt.Errorf("%s is required", required.name))
|
||||
}
|
||||
}
|
||||
if err := errors.Join(missing...); err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
mapping, err := loadMapping(mappingPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
assignments, err := loadAssignment(assignmentPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
checker, err := loadChecker(sweepDirectory)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
reviews, err := loadReviews(reviewsDirectory)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
result := crossTabulate(checker, reviews, assignments, mapping, time.Now().UTC())
|
||||
writeReport(result, stdout)
|
||||
|
||||
if jsonPath == "" {
|
||||
return nil
|
||||
}
|
||||
body, err := json.MarshalIndent(result, "", " ")
|
||||
if err != nil {
|
||||
return fmt.Errorf("marshal summary: %w", err)
|
||||
}
|
||||
body = append(body, '\n')
|
||||
if jsonPath == "-" {
|
||||
_, err = stdout.Write(body)
|
||||
return err
|
||||
}
|
||||
return os.WriteFile(jsonPath, body, 0o644)
|
||||
}
|
||||
|
||||
func main() {
|
||||
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(os.Stderr, "error: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,232 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"regexp"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// clauseCount is R1 to R20, the twenty clauses requirement.md numbers and the
|
||||
// twenty rows review-protocol.md requires on every verdict form.
|
||||
const clauseCount = 20
|
||||
|
||||
const (
|
||||
surfaceUnlocatable = "unlocatable"
|
||||
surfaceInconclusive = "inconclusive"
|
||||
)
|
||||
|
||||
const todoMarker = "todo"
|
||||
|
||||
// propertyMapping is one row of the property table: the clauses a property is
|
||||
// the oracle for, and the locatable surfaces it reads.
|
||||
type propertyMapping struct {
|
||||
Property string
|
||||
Clauses []string
|
||||
Surfaces []string
|
||||
}
|
||||
|
||||
// surfaceMapping says what a surface never observed across an implementation's
|
||||
// whole sweep means. Only a surface the requirement obliges every
|
||||
// implementation to show at all times can be read as unlocatable; a surface
|
||||
// that is legitimately absent when nothing is in that state is inconclusive
|
||||
// and can never mark a property unevaluated.
|
||||
type surfaceMapping struct {
|
||||
Surface string
|
||||
NeverObserved string
|
||||
Note string
|
||||
}
|
||||
|
||||
type mapping struct {
|
||||
Path string
|
||||
Properties []propertyMapping
|
||||
Surfaces []surfaceMapping
|
||||
PropertyTodoRows int
|
||||
byProperty map[string]propertyMapping
|
||||
coveringProperty map[string][]string
|
||||
unlocatableSurface map[string]bool
|
||||
}
|
||||
|
||||
var clausePattern = regexp.MustCompile(`^[Rr]([0-9]{1,2})$`)
|
||||
|
||||
func canonicalClause(value string) (string, bool) {
|
||||
match := clausePattern.FindStringSubmatch(strings.TrimSpace(value))
|
||||
if match == nil {
|
||||
return "", false
|
||||
}
|
||||
number := match[1]
|
||||
trimmed := strings.TrimLeft(number, "0")
|
||||
if trimmed == "" {
|
||||
return "", false
|
||||
}
|
||||
clause := "R" + trimmed
|
||||
if !slices.Contains(allClauses(), clause) {
|
||||
return "", false
|
||||
}
|
||||
return clause, true
|
||||
}
|
||||
|
||||
func allClauses() []string {
|
||||
clauses := make([]string, 0, clauseCount)
|
||||
for index := 1; index <= clauseCount; index++ {
|
||||
clauses = append(clauses, fmt.Sprintf("R%d", index))
|
||||
}
|
||||
return clauses
|
||||
}
|
||||
|
||||
func loadMapping(path string) (mapping, error) {
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return mapping{}, fmt.Errorf("read property-clause mapping: %w", err)
|
||||
}
|
||||
result := mapping{
|
||||
Path: path,
|
||||
byProperty: map[string]propertyMapping{},
|
||||
coveringProperty: map[string][]string{},
|
||||
unlocatableSurface: map[string]bool{},
|
||||
}
|
||||
section := ""
|
||||
for _, row := range parseTableRows(string(body)) {
|
||||
head := strings.ToLower(row.cell(0))
|
||||
switch head {
|
||||
case "property":
|
||||
section = "property"
|
||||
continue
|
||||
case "surface":
|
||||
section = "surface"
|
||||
continue
|
||||
}
|
||||
switch section {
|
||||
case "property":
|
||||
if err := result.addProperty(row); err != nil {
|
||||
return mapping{}, fmt.Errorf("%s line %d: %w", path, row.Line, err)
|
||||
}
|
||||
case "surface":
|
||||
if err := result.addSurface(row); err != nil {
|
||||
return mapping{}, fmt.Errorf("%s line %d: %w", path, row.Line, err)
|
||||
}
|
||||
default:
|
||||
return mapping{}, fmt.Errorf("%s line %d: table row before any header naming property or surface", path, row.Line)
|
||||
}
|
||||
}
|
||||
if len(result.Surfaces) == 0 {
|
||||
return mapping{}, fmt.Errorf("%s declares no surfaces: a portability miss cannot be told from a clean run without them", path)
|
||||
}
|
||||
for _, property := range result.Properties {
|
||||
for _, surface := range property.Surfaces {
|
||||
if _, declared := result.surfaceByName(surface); !declared {
|
||||
return mapping{}, fmt.Errorf("%s: property %q reads surface %q, which the surface table does not declare",
|
||||
path, property.Property, surface)
|
||||
}
|
||||
}
|
||||
}
|
||||
return result, nil
|
||||
}
|
||||
|
||||
func (m *mapping) addProperty(row tableRow) error {
|
||||
name := row.cell(0)
|
||||
if name == "" {
|
||||
return nil
|
||||
}
|
||||
if strings.EqualFold(name, todoMarker) {
|
||||
m.PropertyTodoRows++
|
||||
return nil
|
||||
}
|
||||
if _, seen := m.byProperty[name]; seen {
|
||||
return fmt.Errorf("property %q is mapped twice", name)
|
||||
}
|
||||
entry := propertyMapping{Property: name}
|
||||
for _, item := range splitList(row.cell(1)) {
|
||||
if strings.EqualFold(item, todoMarker) || item == "-" || strings.EqualFold(item, "none") {
|
||||
continue
|
||||
}
|
||||
clause, ok := canonicalClause(item)
|
||||
if !ok {
|
||||
return fmt.Errorf("property %q names clause %q, which is not one of R1 to R%d", name, item, clauseCount)
|
||||
}
|
||||
if slices.Contains(entry.Clauses, clause) {
|
||||
continue
|
||||
}
|
||||
entry.Clauses = append(entry.Clauses, clause)
|
||||
}
|
||||
for _, item := range splitList(row.cell(2)) {
|
||||
if strings.EqualFold(item, todoMarker) || item == "-" || strings.EqualFold(item, "none") {
|
||||
continue
|
||||
}
|
||||
if !slices.Contains(entry.Surfaces, item) {
|
||||
entry.Surfaces = append(entry.Surfaces, item)
|
||||
}
|
||||
}
|
||||
m.Properties = append(m.Properties, entry)
|
||||
m.byProperty[name] = entry
|
||||
for _, clause := range entry.Clauses {
|
||||
m.coveringProperty[clause] = append(m.coveringProperty[clause], name)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (m *mapping) addSurface(row tableRow) error {
|
||||
name := row.cell(0)
|
||||
if name == "" || strings.EqualFold(name, todoMarker) {
|
||||
return nil
|
||||
}
|
||||
meaning := strings.ToLower(row.cell(1))
|
||||
if meaning != surfaceUnlocatable && meaning != surfaceInconclusive {
|
||||
return fmt.Errorf("surface %q says %q for never observed, want %s or %s",
|
||||
name, row.cell(1), surfaceUnlocatable, surfaceInconclusive)
|
||||
}
|
||||
if _, seen := m.surfaceByName(name); seen {
|
||||
return fmt.Errorf("surface %q is declared twice", name)
|
||||
}
|
||||
m.Surfaces = append(m.Surfaces, surfaceMapping{Surface: name, NeverObserved: meaning, Note: row.cell(2)})
|
||||
if meaning == surfaceUnlocatable {
|
||||
m.unlocatableSurface[name] = true
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (m mapping) surfaceByName(name string) (surfaceMapping, bool) {
|
||||
for _, surface := range m.Surfaces {
|
||||
if surface.Surface == name {
|
||||
return surface, true
|
||||
}
|
||||
}
|
||||
return surfaceMapping{}, false
|
||||
}
|
||||
|
||||
// unevaluable reports whether a property could not be evaluated against this
|
||||
// implementation because a surface it reads was never located. A surface the
|
||||
// mapping calls inconclusive never makes a property unevaluable, however many
|
||||
// steps failed to observe it.
|
||||
func (m mapping) unevaluable(property string, observed map[string]bool) bool {
|
||||
entry, known := m.byProperty[property]
|
||||
if !known {
|
||||
return false
|
||||
}
|
||||
for _, surface := range entry.Surfaces {
|
||||
if m.unlocatableSurface[surface] && !observed[surface] {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func (m mapping) covering(clause string) []string {
|
||||
return m.coveringProperty[clause]
|
||||
}
|
||||
|
||||
func (m mapping) knows(property string) bool {
|
||||
_, known := m.byProperty[property]
|
||||
return known
|
||||
}
|
||||
|
||||
func (m mapping) unlocatableSurfaces() []string {
|
||||
var names []string
|
||||
for _, surface := range m.Surfaces {
|
||||
if surface.NeverObserved == surfaceUnlocatable {
|
||||
names = append(names, surface.Surface)
|
||||
}
|
||||
}
|
||||
return names
|
||||
}
|
||||
@@ -0,0 +1,510 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"maps"
|
||||
"slices"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// The four cells of the pre-registered measure, named as
|
||||
// model-implementations.md describes them.
|
||||
const (
|
||||
cellTruePositive = "checker fired, review confirmed"
|
||||
cellFalsePositive = "checker fired, review found nothing"
|
||||
cellFalseNegative = "checker silent, review found a defect"
|
||||
cellTrueNegative = "checker silent, review found nothing"
|
||||
)
|
||||
|
||||
// Reasons an implementation is missing data rather than a cell of the matrix.
|
||||
const (
|
||||
missingSweepStage = "sweep stopped before any run"
|
||||
missingNoUsableRun = "no usable run"
|
||||
missingNoSweepRecord = "no sweep record"
|
||||
missingNoReview = "no verdict filed"
|
||||
missingMalformed = "malformed verdict form"
|
||||
missingSurfaces = "surface locatability unknown"
|
||||
missingNoModel = "not in the assignment mapping"
|
||||
)
|
||||
|
||||
type matrix struct {
|
||||
Unit string `json:"unit"`
|
||||
Scored int `json:"scored"`
|
||||
TruePositive int `json:"true_positive"`
|
||||
FalsePositive int `json:"false_positive"`
|
||||
FalseNegative int `json:"false_negative"`
|
||||
TrueNegative int `json:"true_negative"`
|
||||
Precision *float64 `json:"precision"`
|
||||
Recall *float64 `json:"recall"`
|
||||
}
|
||||
|
||||
func (m *matrix) add(checkerPositive, humanPositive bool) {
|
||||
m.Scored++
|
||||
switch {
|
||||
case checkerPositive && humanPositive:
|
||||
m.TruePositive++
|
||||
case checkerPositive:
|
||||
m.FalsePositive++
|
||||
case humanPositive:
|
||||
m.FalseNegative++
|
||||
default:
|
||||
m.TrueNegative++
|
||||
}
|
||||
}
|
||||
|
||||
func (m *matrix) finish() {
|
||||
m.Precision = ratio(m.TruePositive, m.TruePositive+m.FalsePositive)
|
||||
m.Recall = ratio(m.TruePositive, m.TruePositive+m.FalseNegative)
|
||||
}
|
||||
|
||||
func ratio(numerator, denominator int) *float64 {
|
||||
if denominator == 0 {
|
||||
return nil
|
||||
}
|
||||
value := float64(numerator) / float64(denominator)
|
||||
return &value
|
||||
}
|
||||
|
||||
// clauseMatrix scores one implementation-clause pair. Its three side buckets
|
||||
// hold the pairs that are not evidence either way: a clause the reviewer could
|
||||
// not judge, a clause whose every covering property was unevaluable because a
|
||||
// surface was never located, and, tracked but still scored, a clause no
|
||||
// property covers at all.
|
||||
type clauseMatrix struct {
|
||||
matrix
|
||||
CannotTell int `json:"cannot_tell"`
|
||||
UnevaluatedSurfaceMissed int `json:"unevaluated_surface_missed"`
|
||||
UncoveredScored int `json:"uncovered_clause_pairs_scored"`
|
||||
UncoveredFalseNegative int `json:"uncovered_clause_false_negatives"`
|
||||
}
|
||||
|
||||
type implementationOutcome struct {
|
||||
Implementation string `json:"implementation"`
|
||||
Model string `json:"model"`
|
||||
Cell string `json:"cell"`
|
||||
FiredProperties []string `json:"fired_properties,omitempty"`
|
||||
ViolatedClauses []string `json:"violated_clauses,omitempty"`
|
||||
CannotTellClauses []string `json:"cannot_tell_clauses,omitempty"`
|
||||
UnlocatableSurfaces []string `json:"unlocatable_surfaces,omitempty"`
|
||||
UnevaluatedProperties []string `json:"unevaluated_properties,omitempty"`
|
||||
// DefectOnlyOnUnevaluatedClauses marks a false negative the checker was
|
||||
// never in a position to catch: every clause the reviewer faulted is
|
||||
// covered only by properties a missing surface left unevaluable. It is a
|
||||
// portability miss reported beside the matrix, never inside it.
|
||||
DefectOnlyOnUnevaluatedClauses bool `json:"defect_only_on_unevaluated_clauses,omitempty"`
|
||||
RunsUsable int `json:"runs_usable"`
|
||||
ReviewMinutes int `json:"review_minutes,omitempty"`
|
||||
}
|
||||
|
||||
type exclusion struct {
|
||||
Implementation string `json:"implementation"`
|
||||
Model string `json:"model,omitempty"`
|
||||
Reason string `json:"reason"`
|
||||
Detail string `json:"detail,omitempty"`
|
||||
}
|
||||
|
||||
type modelBreakdown struct {
|
||||
Model string `json:"model"`
|
||||
Implementations matrix `json:"implementation_matrix"`
|
||||
Clauses clauseMatrix `json:"clause_matrix"`
|
||||
Excluded int `json:"excluded"`
|
||||
DefectOnlyOnUnevaluatedClauses int `json:"defect_only_on_unevaluated_clauses"`
|
||||
}
|
||||
|
||||
type portability struct {
|
||||
Scored int `json:"implementations_scored"`
|
||||
WithUnlocatableSurface int `json:"implementations_with_an_unlocatable_surface"`
|
||||
BySurface map[string]int `json:"implementations_by_unlocatable_surface,omitempty"`
|
||||
SurfacesReadAsInconclusive []string `json:"surfaces_a_miss_cannot_be_read_from,omitempty"`
|
||||
}
|
||||
|
||||
type coverage struct {
|
||||
MappedProperties int `json:"mapped_properties"`
|
||||
MappingTodoRows int `json:"mapping_todo_rows"`
|
||||
ClausesCovered []string `json:"clauses_covered,omitempty"`
|
||||
ClausesUncovered []string `json:"clauses_no_property_covers,omitempty"`
|
||||
FiredPropertiesNotMapped []string `json:"fired_properties_not_in_the_mapping,omitempty"`
|
||||
}
|
||||
|
||||
type result struct {
|
||||
GeneratedAt time.Time `json:"generated_at"`
|
||||
SweepDirectory string `json:"sweep_directory"`
|
||||
ReviewsDirectory string `json:"reviews_directory"`
|
||||
MappingPath string `json:"property_clause_mapping"`
|
||||
SpecPath string `json:"spec_path,omitempty"`
|
||||
Implementations matrix `json:"implementation_matrix"`
|
||||
Clauses clauseMatrix `json:"clause_matrix"`
|
||||
ByModel []modelBreakdown `json:"by_model"`
|
||||
Outcomes []implementationOutcome `json:"outcomes"`
|
||||
Excluded []exclusion `json:"excluded,omitempty"`
|
||||
Portability portability `json:"portability"`
|
||||
ReviewMinutes int `json:"review_minutes_over_scored_implementations"`
|
||||
Coverage coverage `json:"clause_coverage"`
|
||||
Notes []string `json:"notes,omitempty"`
|
||||
}
|
||||
|
||||
func crossTabulate(
|
||||
checker checkerSide,
|
||||
reviews reviewSide,
|
||||
assignments map[string]string,
|
||||
declared mapping,
|
||||
now time.Time,
|
||||
) result {
|
||||
outcome := result{
|
||||
GeneratedAt: now,
|
||||
SweepDirectory: checker.Directory,
|
||||
ReviewsDirectory: reviews.Directory,
|
||||
MappingPath: declared.Path,
|
||||
SpecPath: checker.SpecPath,
|
||||
Implementations: matrix{Unit: "implementation"},
|
||||
Clauses: clauseMatrix{matrix: matrix{Unit: "implementation-clause pair"}},
|
||||
Portability: portability{BySurface: map[string]int{}},
|
||||
}
|
||||
|
||||
reviewByName := map[string]reviewVerdict{}
|
||||
for _, verdict := range reviews.Verdicts {
|
||||
reviewByName[verdict.Implementation] = verdict
|
||||
}
|
||||
malformedByName := map[string]malformedReview{}
|
||||
for _, entry := range reviews.Malformed {
|
||||
malformedByName[entry.Implementation] = entry
|
||||
}
|
||||
checkerByName := map[string]checkerVerdict{}
|
||||
for _, verdict := range checker.Verdicts {
|
||||
checkerByName[verdict.Implementation] = verdict
|
||||
}
|
||||
|
||||
byModel := map[string]*modelBreakdown{}
|
||||
modelOf := func(name string) string { return assignments[name] }
|
||||
breakdown := func(model string) *modelBreakdown {
|
||||
current, seen := byModel[model]
|
||||
if !seen {
|
||||
current = &modelBreakdown{
|
||||
Model: model,
|
||||
Implementations: matrix{Unit: "implementation"},
|
||||
Clauses: clauseMatrix{matrix: matrix{Unit: "implementation-clause pair"}},
|
||||
}
|
||||
byModel[model] = current
|
||||
}
|
||||
return current
|
||||
}
|
||||
|
||||
unmappedFired := map[string]bool{}
|
||||
for _, name := range implementationNames(checker, reviews, assignments) {
|
||||
model := modelOf(name)
|
||||
verdict, swept := checkerByName[name]
|
||||
review, reviewed := reviewByName[name]
|
||||
|
||||
if reason, detail := missingData(name, model, verdict, swept, reviewed, malformedByName); reason != "" {
|
||||
outcome.Excluded = append(outcome.Excluded, exclusion{
|
||||
Implementation: name, Model: model, Reason: reason, Detail: detail,
|
||||
})
|
||||
if model != "" {
|
||||
breakdown(model).Excluded++
|
||||
}
|
||||
continue
|
||||
}
|
||||
|
||||
for _, property := range verdict.FiredProperties {
|
||||
if !declared.knows(property) {
|
||||
unmappedFired[property] = true
|
||||
}
|
||||
}
|
||||
row := scoreImplementation(verdict, review, declared)
|
||||
row.Model = model
|
||||
outcome.Outcomes = append(outcome.Outcomes, row)
|
||||
|
||||
modelRow := breakdown(model)
|
||||
outcome.Implementations.add(verdict.fired(), review.defective())
|
||||
modelRow.Implementations.add(verdict.fired(), review.defective())
|
||||
scoreClauses(&outcome.Clauses, &modelRow.Clauses, verdict, review, declared)
|
||||
|
||||
if row.DefectOnlyOnUnevaluatedClauses {
|
||||
modelRow.DefectOnlyOnUnevaluatedClauses++
|
||||
}
|
||||
outcome.Portability.Scored++
|
||||
outcome.ReviewMinutes += review.Minutes
|
||||
if len(row.UnlocatableSurfaces) > 0 {
|
||||
outcome.Portability.WithUnlocatableSurface++
|
||||
for _, surface := range row.UnlocatableSurfaces {
|
||||
outcome.Portability.BySurface[surface]++
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
outcome.Implementations.finish()
|
||||
outcome.Clauses.finish()
|
||||
for _, model := range capabilityOrder {
|
||||
current, seen := byModel[model]
|
||||
if !seen {
|
||||
continue
|
||||
}
|
||||
current.Implementations.finish()
|
||||
current.Clauses.finish()
|
||||
outcome.ByModel = append(outcome.ByModel, *current)
|
||||
}
|
||||
|
||||
outcome.Coverage = describeCoverage(declared, unmappedFired)
|
||||
outcome.Portability.SurfacesReadAsInconclusive = inconclusiveSurfaces(declared)
|
||||
outcome.Notes = buildNotes(outcome, declared, reviews)
|
||||
return outcome
|
||||
}
|
||||
|
||||
func implementationNames(checker checkerSide, reviews reviewSide, assignments map[string]string) []string {
|
||||
names := map[string]bool{}
|
||||
for _, planned := range checker.Planned {
|
||||
names[planned] = true
|
||||
}
|
||||
for _, verdict := range checker.Verdicts {
|
||||
names[verdict.Implementation] = true
|
||||
}
|
||||
for _, verdict := range reviews.Verdicts {
|
||||
names[verdict.Implementation] = true
|
||||
}
|
||||
for _, entry := range reviews.Malformed {
|
||||
names[entry.Implementation] = true
|
||||
}
|
||||
for name := range assignments {
|
||||
names[name] = true
|
||||
}
|
||||
return slices.Sorted(maps.Keys(names))
|
||||
}
|
||||
|
||||
// missingData names why an implementation carries no cell. A build that never
|
||||
// finished, a run that never produced a usable campaign and a verdict that was
|
||||
// never filed are all absent evidence: scoring any of them as a clean run would
|
||||
// read the gap as agreement.
|
||||
func missingData(
|
||||
name string,
|
||||
model string,
|
||||
verdict checkerVerdict,
|
||||
swept bool,
|
||||
reviewed bool,
|
||||
malformed map[string]malformedReview,
|
||||
) (string, string) {
|
||||
if model == "" {
|
||||
return missingNoModel, ""
|
||||
}
|
||||
if !swept {
|
||||
return missingNoSweepRecord, ""
|
||||
}
|
||||
if verdict.FailedStage != "" {
|
||||
return missingSweepStage, fmt.Sprintf("%s: %s", verdict.FailedStage, verdict.FailedError)
|
||||
}
|
||||
if verdict.RunsUsable == 0 {
|
||||
return missingNoUsableRun, excludedSummary(verdict)
|
||||
}
|
||||
if entry, broken := malformed[name]; broken {
|
||||
return missingMalformed, entry.Reason
|
||||
}
|
||||
if !reviewed {
|
||||
return missingNoReview, ""
|
||||
}
|
||||
if !verdict.SurfacesKnown {
|
||||
return missingSurfaces, strings.Join(verdict.TraceErrors, "; ")
|
||||
}
|
||||
return "", ""
|
||||
}
|
||||
|
||||
func excludedSummary(verdict checkerVerdict) string {
|
||||
if len(verdict.ExcludedByReason) == 0 {
|
||||
return ""
|
||||
}
|
||||
var parts []string
|
||||
for _, reason := range slices.Sorted(maps.Keys(verdict.ExcludedByReason)) {
|
||||
parts = append(parts, fmt.Sprintf("%s=%d", reason, verdict.ExcludedByReason[reason]))
|
||||
}
|
||||
return strings.Join(parts, ", ")
|
||||
}
|
||||
|
||||
func scoreImplementation(verdict checkerVerdict, review reviewVerdict, declared mapping) implementationOutcome {
|
||||
row := implementationOutcome{
|
||||
Implementation: verdict.Implementation,
|
||||
Cell: cellOf(verdict.fired(), review.defective()),
|
||||
FiredProperties: verdict.FiredProperties,
|
||||
ViolatedClauses: review.violatedClauses(),
|
||||
RunsUsable: verdict.RunsUsable,
|
||||
ReviewMinutes: review.Minutes,
|
||||
}
|
||||
for _, clause := range allClauses() {
|
||||
if review.Clauses[clause] == clauseCannotTell {
|
||||
row.CannotTellClauses = append(row.CannotTellClauses, clause)
|
||||
}
|
||||
}
|
||||
for _, surface := range declared.unlocatableSurfaces() {
|
||||
if !verdict.SurfacesObserved[surface] {
|
||||
row.UnlocatableSurfaces = append(row.UnlocatableSurfaces, surface)
|
||||
}
|
||||
}
|
||||
for _, property := range declared.Properties {
|
||||
if declared.unevaluable(property.Property, verdict.SurfacesObserved) {
|
||||
row.UnevaluatedProperties = append(row.UnevaluatedProperties, property.Property)
|
||||
}
|
||||
}
|
||||
row.DefectOnlyOnUnevaluatedClauses = attributableToUnevaluated(verdict, review, declared)
|
||||
return row
|
||||
}
|
||||
|
||||
// attributableToUnevaluated reports a false negative the checker could not have
|
||||
// caught: it fired nothing, the reviewer faulted at least one clause, and every
|
||||
// clause the reviewer faulted is covered only by properties a missing surface
|
||||
// left unevaluable.
|
||||
func attributableToUnevaluated(verdict checkerVerdict, review reviewVerdict, declared mapping) bool {
|
||||
if verdict.fired() || !review.defective() {
|
||||
return false
|
||||
}
|
||||
violated := review.violatedClauses()
|
||||
if len(violated) == 0 {
|
||||
return false
|
||||
}
|
||||
for _, clause := range violated {
|
||||
covering := declared.covering(clause)
|
||||
if len(covering) == 0 {
|
||||
return false
|
||||
}
|
||||
for _, property := range covering {
|
||||
if !declared.unevaluable(property, verdict.SurfacesObserved) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func cellOf(checkerPositive, humanPositive bool) string {
|
||||
switch {
|
||||
case checkerPositive && humanPositive:
|
||||
return cellTruePositive
|
||||
case checkerPositive:
|
||||
return cellFalsePositive
|
||||
case humanPositive:
|
||||
return cellFalseNegative
|
||||
default:
|
||||
return cellTrueNegative
|
||||
}
|
||||
}
|
||||
|
||||
func scoreClauses(overall, model *clauseMatrix, verdict checkerVerdict, review reviewVerdict, declared mapping) {
|
||||
fired := map[string]bool{}
|
||||
for _, property := range verdict.FiredProperties {
|
||||
fired[property] = true
|
||||
}
|
||||
for _, clause := range allClauses() {
|
||||
label := review.Clauses[clause]
|
||||
if label == clauseCannotTell {
|
||||
overall.CannotTell++
|
||||
model.CannotTell++
|
||||
continue
|
||||
}
|
||||
covering := declared.covering(clause)
|
||||
checkerPositive := false
|
||||
for _, property := range covering {
|
||||
if fired[property] {
|
||||
checkerPositive = true
|
||||
break
|
||||
}
|
||||
}
|
||||
if !checkerPositive && len(covering) > 0 && allUnevaluable(covering, verdict, declared) {
|
||||
overall.UnevaluatedSurfaceMissed++
|
||||
model.UnevaluatedSurfaceMissed++
|
||||
continue
|
||||
}
|
||||
humanPositive := label == clauseViolates
|
||||
overall.add(checkerPositive, humanPositive)
|
||||
model.add(checkerPositive, humanPositive)
|
||||
if len(covering) == 0 {
|
||||
overall.UncoveredScored++
|
||||
model.UncoveredScored++
|
||||
if humanPositive {
|
||||
overall.UncoveredFalseNegative++
|
||||
model.UncoveredFalseNegative++
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func allUnevaluable(covering []string, verdict checkerVerdict, declared mapping) bool {
|
||||
for _, property := range covering {
|
||||
if !declared.unevaluable(property, verdict.SurfacesObserved) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func describeCoverage(declared mapping, unmappedFired map[string]bool) coverage {
|
||||
result := coverage{
|
||||
MappedProperties: len(declared.Properties),
|
||||
MappingTodoRows: declared.PropertyTodoRows,
|
||||
}
|
||||
for _, clause := range allClauses() {
|
||||
if len(declared.covering(clause)) > 0 {
|
||||
result.ClausesCovered = append(result.ClausesCovered, clause)
|
||||
continue
|
||||
}
|
||||
result.ClausesUncovered = append(result.ClausesUncovered, clause)
|
||||
}
|
||||
if len(unmappedFired) > 0 {
|
||||
result.FiredPropertiesNotMapped = slices.Sorted(maps.Keys(unmappedFired))
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func inconclusiveSurfaces(declared mapping) []string {
|
||||
var names []string
|
||||
for _, surface := range declared.Surfaces {
|
||||
if surface.NeverObserved == surfaceInconclusive {
|
||||
names = append(names, surface.Surface)
|
||||
}
|
||||
}
|
||||
return names
|
||||
}
|
||||
|
||||
func buildNotes(outcome result, declared mapping, reviews reviewSide) []string {
|
||||
var notes []string
|
||||
if len(declared.Properties) == 0 {
|
||||
notes = append(notes, fmt.Sprintf(
|
||||
"%s maps no property to a clause, so the clause matrix is empty and no portability miss can be detected; "+
|
||||
"the implementation matrix below stands on its own", declared.Path))
|
||||
}
|
||||
if declared.PropertyTodoRows > 0 {
|
||||
notes = append(notes, fmt.Sprintf("%s still carries %d TODO row(s) in its property table",
|
||||
declared.Path, declared.PropertyTodoRows))
|
||||
}
|
||||
if len(outcome.Coverage.FiredPropertiesNotMapped) > 0 {
|
||||
notes = append(notes, fmt.Sprintf(
|
||||
"%d fired propert(ies) are absent from the mapping and could not be attributed to a clause: %s",
|
||||
len(outcome.Coverage.FiredPropertiesNotMapped),
|
||||
strings.Join(outcome.Coverage.FiredPropertiesNotMapped, ", ")))
|
||||
}
|
||||
for _, verdict := range reviews.Verdicts {
|
||||
if !verdict.defective() && len(verdict.violatedClauses()) > 0 {
|
||||
notes = append(notes, fmt.Sprintf(
|
||||
"%s files %d violating clause(s) under an overall verdict of %s; the overall verdict is what the matrix scores",
|
||||
verdict.Implementation, len(verdict.violatedClauses()), overallNotDefective))
|
||||
}
|
||||
if len(verdict.Adjudicated) > 0 {
|
||||
notes = append(notes, fmt.Sprintf("%s uses the adjudicated label for %s",
|
||||
verdict.Implementation, strings.Join(verdict.Adjudicated, ", ")))
|
||||
}
|
||||
}
|
||||
if count := attributedFalseNegatives(outcome); count > 0 {
|
||||
notes = append(notes, fmt.Sprintf(
|
||||
"%d false negative(s) fault only clauses whose every property a missing surface left unevaluable: "+
|
||||
"those are portability misses, not blind spots", count))
|
||||
}
|
||||
notes = append(notes, "second-rater agreement and Cohen's kappa are not computed here; "+
|
||||
"an impl-NN-adjudication.md is read and its resolved labels replace the first rater's")
|
||||
return notes
|
||||
}
|
||||
|
||||
func attributedFalseNegatives(outcome result) int {
|
||||
count := 0
|
||||
for _, row := range outcome.Outcomes {
|
||||
if row.DefectOnlyOnUnevaluatedClauses {
|
||||
count++
|
||||
}
|
||||
}
|
||||
return count
|
||||
}
|
||||
@@ -0,0 +1,461 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestCrossTabulateScoresEachImplementationIntoOneCell(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
implementation fixtureImplementation
|
||||
wantCell string
|
||||
wantExcluded string
|
||||
}{
|
||||
{
|
||||
name: "checker and reviewer agree",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce"), cleanRun(2)},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
|
||||
},
|
||||
wantCell: cellTruePositive,
|
||||
},
|
||||
{
|
||||
name: "checker fired and the reviewer found nothing",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-02", Model: "Sonnet 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
wantCell: cellFalsePositive,
|
||||
},
|
||||
{
|
||||
name: "reviewer found a defect no property fired on",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-03", Model: "Fable 5",
|
||||
Runs: []fixtureRun{cleanRun(1), cleanRun(2)},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
|
||||
},
|
||||
wantCell: cellFalseNegative,
|
||||
},
|
||||
{
|
||||
name: "both clean",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-04", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1), cleanRun(2)},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
wantCell: cellTrueNegative,
|
||||
},
|
||||
{
|
||||
name: "the build never finished",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-05", Model: "Sonnet 5", FailedStage: "build",
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R11": clauseViolates}},
|
||||
},
|
||||
wantExcluded: missingSweepStage,
|
||||
},
|
||||
{
|
||||
name: "every run timed out",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-06", Model: "Fable 5",
|
||||
Runs: []fixtureRun{{Seed: 1, ExitCode: -1, TimedOut: true}},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
wantExcluded: missingNoUsableRun,
|
||||
},
|
||||
{
|
||||
name: "no verdict was filed",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-07", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
},
|
||||
wantExcluded: missingNoReview,
|
||||
},
|
||||
{
|
||||
name: "the verdict form is missing a clause",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-08", Model: "Sonnet 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
RawReview: "reviewer: Jane\n\n| clause | verdict |\n| --- | --- |\n" +
|
||||
"| R1 | meets |\n| R2 | meets |\n\noverall: not defective\n",
|
||||
},
|
||||
wantExcluded: missingMalformed,
|
||||
},
|
||||
{
|
||||
name: "no trace says whether a surface was located",
|
||||
implementation: fixtureImplementation{
|
||||
Name: "impl-09", Model: "Fable 5",
|
||||
Runs: []fixtureRun{{Seed: 1, NoTrace: true}},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
wantExcluded: missingSurfaces,
|
||||
},
|
||||
}
|
||||
|
||||
for _, test := range tests {
|
||||
t.Run(test.name, func(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{test.implementation}, defaultMapping))
|
||||
if test.wantExcluded != "" {
|
||||
entry := exclusionFor(t, emitted, test.implementation.Name)
|
||||
if entry.Reason != test.wantExcluded {
|
||||
t.Fatalf("%s excluded as %q, want %q", test.implementation.Name, entry.Reason, test.wantExcluded)
|
||||
}
|
||||
if emitted.Implementations.Scored != 0 {
|
||||
t.Errorf("missing data scored %d implementation(s), want none: absent evidence is not a clean run",
|
||||
emitted.Implementations.Scored)
|
||||
}
|
||||
if !strings.Contains(stdout, "carry no cell") {
|
||||
t.Errorf("the report never separates the missing data from the matrix:\n%s", stdout)
|
||||
}
|
||||
return
|
||||
}
|
||||
row := outcomeFor(t, emitted, test.implementation.Name)
|
||||
if row.Cell != test.wantCell {
|
||||
t.Fatalf("%s landed in %q, want %q", test.implementation.Name, row.Cell, test.wantCell)
|
||||
}
|
||||
if emitted.Implementations.Scored != 1 {
|
||||
t.Errorf("scored %d implementation(s), want 1", emitted.Implementations.Scored)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestACampaignThatDiedIsMissingDataNotATrueNegative covers the sweep it was
|
||||
// interrupted on: the one seed the campaign got through wrote a clean run
|
||||
// before the process died, and scoring the implementation on it reads the nine
|
||||
// seeds that never ran as agreement between the checker and the reviewer.
|
||||
func TestACampaignThatDiedIsMissingDataNotATrueNegative(t *testing.T) {
|
||||
interrupted := cleanRun(1)
|
||||
interrupted.CampaignExitCode = -1
|
||||
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-07", Model: "Opus 5",
|
||||
Runs: []fixtureRun{interrupted},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
}}, defaultMapping))
|
||||
|
||||
entry := exclusionFor(t, emitted, "impl-07")
|
||||
if entry.Reason != missingNoUsableRun {
|
||||
t.Fatalf("impl-07 excluded as %q, want %q", entry.Reason, missingNoUsableRun)
|
||||
}
|
||||
if !strings.Contains(entry.Detail, reasonNonzeroExit) {
|
||||
t.Errorf("exclusion detail %q does not name %q", entry.Detail, reasonNonzeroExit)
|
||||
}
|
||||
if emitted.Implementations.TrueNegative != 0 || emitted.Implementations.Scored != 0 {
|
||||
t.Errorf("implementation matrix = %+v, want no cell: a dead campaign is absent evidence",
|
||||
emitted.Implementations)
|
||||
}
|
||||
if !strings.Contains(stdout, "carry no cell") {
|
||||
t.Errorf("the report never separates the dead campaign from the matrix:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnlocatableSurfaceIsNeitherAPositiveNorANegative pins the rule the
|
||||
// pre-registration turns on: a clause whose every covering property a
|
||||
// never-located surface left unrunnable is a portability miss. Counting it as a
|
||||
// true positive credits the oracle for a property that never ran, and counting
|
||||
// it as a false negative charges the oracle for a defect it was never in a
|
||||
// position to see.
|
||||
func TestUnlocatableSurfaceIsNeitherAPositiveNorANegative(t *testing.T) {
|
||||
built := writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01",
|
||||
Model: "Opus 5",
|
||||
Runs: []fixtureRun{{
|
||||
Seed: 1,
|
||||
Violated: []string{"serverHoldsEachMessageOnce"},
|
||||
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
|
||||
}},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{
|
||||
"R5": clauseViolates,
|
||||
"R15": clauseViolates,
|
||||
}},
|
||||
}}, defaultMapping)
|
||||
|
||||
emitted, stdout := runTool(t, built)
|
||||
|
||||
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
|
||||
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1 for R5 behind an unlocated stateWords", got)
|
||||
}
|
||||
if got := emitted.Clauses.TruePositive; got != 1 {
|
||||
t.Errorf("clause matrix reports %d true positive(s), want 1: only R15 had a property that ran and fired", got)
|
||||
}
|
||||
if got := emitted.Clauses.FalsePositive; got != 0 {
|
||||
t.Errorf("clause matrix reports %d false positive(s), want 0", got)
|
||||
}
|
||||
if got := emitted.Clauses.FalseNegative; got != 0 {
|
||||
t.Errorf("clause matrix reports %d false negative(s), want 0: R5 was never evaluated, so it was not missed", got)
|
||||
}
|
||||
if got := emitted.Clauses.Scored; got != clauseCount-1 {
|
||||
t.Errorf("clause matrix scored %d pair(s), want %d: the unevaluated pair is excluded, not scored", got, clauseCount-1)
|
||||
}
|
||||
if total := emitted.Clauses.TruePositive + emitted.Clauses.FalsePositive +
|
||||
emitted.Clauses.FalseNegative + emitted.Clauses.TrueNegative; total != emitted.Clauses.Scored {
|
||||
t.Errorf("the four cells sum to %d against %d scored: an excluded pair leaked into a cell", total, emitted.Clauses.Scored)
|
||||
}
|
||||
row := outcomeFor(t, emitted, "impl-01")
|
||||
if want := []string{"stateWords"}; !equalStrings(row.UnlocatableSurfaces, want) {
|
||||
t.Errorf("impl-01 reports unlocatable surfaces %v, want %v", row.UnlocatableSurfaces, want)
|
||||
}
|
||||
if want := []string{"sentOnlyAfterConfirmation"}; !equalStrings(row.UnevaluatedProperties, want) {
|
||||
t.Errorf("impl-01 reports unevaluated properties %v, want %v", row.UnevaluatedProperties, want)
|
||||
}
|
||||
if emitted.Portability.WithUnlocatableSurface != 1 {
|
||||
t.Errorf("portability counts %d implementation(s) needing a locating adaptation, want 1",
|
||||
emitted.Portability.WithUnlocatableSurface)
|
||||
}
|
||||
if !strings.Contains(stdout, "needed a locating adaptation") {
|
||||
t.Errorf("the report never names the portability count:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnlocatableSurfaceOnAMetClauseIsNotATrueNegative is the other half: a
|
||||
// property that never ran is not evidence the implementation met the clause.
|
||||
func TestUnlocatableSurfaceOnAMetClauseIsNotATrueNegative(t *testing.T) {
|
||||
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01",
|
||||
Model: "Sonnet 5",
|
||||
Runs: []fixtureRun{{
|
||||
Seed: 1,
|
||||
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
|
||||
}},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
}}, defaultMapping))
|
||||
|
||||
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
|
||||
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1: R5 is covered only by a property "+
|
||||
"stateWords left unrunnable", got)
|
||||
}
|
||||
if got := emitted.Clauses.TrueNegative; got != clauseCount-1 {
|
||||
t.Errorf("clause matrix reports %d true negative(s), want %d: R5 is excluded rather than credited", got, clauseCount-1)
|
||||
}
|
||||
if got := emitted.Clauses.FalsePositive; got != 0 {
|
||||
t.Errorf("clause matrix reports %d false positive(s), want 0: a property that never ran cannot have fired", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestFiredPropertyOutranksAnUnevaluableSibling keeps the exclusion narrow: a
|
||||
// clause is unevaluated only when every property covering it was unrunnable.
|
||||
func TestFiredPropertyOutranksAnUnevaluableSibling(t *testing.T) {
|
||||
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01",
|
||||
Model: "Fable 5",
|
||||
Runs: []fixtureRun{{
|
||||
Seed: 1,
|
||||
Violated: []string{"serverHoldsEachMessageOnce"},
|
||||
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
|
||||
}},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
|
||||
}}, defaultMapping))
|
||||
|
||||
row := outcomeFor(t, emitted, "impl-01")
|
||||
if row.Cell != cellTruePositive {
|
||||
t.Fatalf("impl-01 landed in %q, want %q", row.Cell, cellTruePositive)
|
||||
}
|
||||
if got := emitted.Clauses.TruePositive; got != 1 {
|
||||
t.Errorf("R15 produced %d true positive(s), want 1: one covering property ran and fired", got)
|
||||
}
|
||||
if got := emitted.Clauses.UnevaluatedSurfaceMissed; got != 1 {
|
||||
t.Errorf("clause matrix reports %d unevaluated pair(s), want 1 for R5 alone", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestFalseNegativeBehindAnUnlocatedSurfaceIsReportedApart keeps a portability
|
||||
// miss out of the blind-spot story it would otherwise be read as.
|
||||
func TestFalseNegativeBehindAnUnlocatedSurfaceIsReportedApart(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01",
|
||||
Model: "Opus 5",
|
||||
Runs: []fixtureRun{{
|
||||
Seed: 1,
|
||||
Surfaces: map[string]bool{"appRoot": true, "composer": true, "submit": true, "stateWords": false},
|
||||
}},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R5": clauseViolates}},
|
||||
}}, defaultMapping))
|
||||
|
||||
row := outcomeFor(t, emitted, "impl-01")
|
||||
if row.Cell != cellFalseNegative {
|
||||
t.Fatalf("impl-01 landed in %q, want %q at the implementation unit", row.Cell, cellFalseNegative)
|
||||
}
|
||||
if !row.DefectOnlyOnUnevaluatedClauses {
|
||||
t.Error("impl-01 faults only R5, whose one property never ran, and is not marked as a portability miss")
|
||||
}
|
||||
if got := emitted.Clauses.FalseNegative; got != 0 {
|
||||
t.Errorf("clause matrix reports %d false negative(s), want 0", got)
|
||||
}
|
||||
if !strings.Contains(stdout, "portability misses, not blind spots") {
|
||||
t.Errorf("the report never separates the portability miss from the blind spot:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPrecisionRecallAndPerModelBreakdown(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{
|
||||
{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
|
||||
},
|
||||
{
|
||||
Name: "impl-02", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
{
|
||||
Name: "impl-03", Model: "Sonnet 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
|
||||
},
|
||||
{
|
||||
Name: "impl-04", Model: "Fable 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
}, defaultMapping))
|
||||
|
||||
if emitted.Implementations.Scored != 4 {
|
||||
t.Fatalf("scored %d implementation(s), want 4", emitted.Implementations.Scored)
|
||||
}
|
||||
wantCells := map[string]int{"tp": 1, "fp": 1, "fn": 1, "tn": 1}
|
||||
got := map[string]int{
|
||||
"tp": emitted.Implementations.TruePositive,
|
||||
"fp": emitted.Implementations.FalsePositive,
|
||||
"fn": emitted.Implementations.FalseNegative,
|
||||
"tn": emitted.Implementations.TrueNegative,
|
||||
}
|
||||
for cell, want := range wantCells {
|
||||
if got[cell] != want {
|
||||
t.Errorf("implementation matrix %s = %d, want %d", cell, got[cell], want)
|
||||
}
|
||||
}
|
||||
if emitted.Implementations.Precision == nil || *emitted.Implementations.Precision != 0.5 {
|
||||
t.Errorf("precision = %v, want 0.5", emitted.Implementations.Precision)
|
||||
}
|
||||
if emitted.Implementations.Recall == nil || *emitted.Implementations.Recall != 0.5 {
|
||||
t.Errorf("recall = %v, want 0.5", emitted.Implementations.Recall)
|
||||
}
|
||||
|
||||
if len(emitted.ByModel) != 3 {
|
||||
t.Fatalf("broke down %d model(s), want 3", len(emitted.ByModel))
|
||||
}
|
||||
if order := []string{emitted.ByModel[0].Model, emitted.ByModel[1].Model, emitted.ByModel[2].Model}; !equalStrings(order, capabilityOrder) {
|
||||
t.Errorf("models reported in %v, want the pre-registered capability order %v", order, capabilityOrder)
|
||||
}
|
||||
for _, model := range emitted.ByModel {
|
||||
switch model.Model {
|
||||
case "Opus 5":
|
||||
if model.Implementations.TruePositive != 1 || model.Implementations.FalsePositive != 1 {
|
||||
t.Errorf("Opus 5 = %+v, want one true positive and one false positive", model.Implementations)
|
||||
}
|
||||
case "Sonnet 5":
|
||||
if model.Implementations.FalseNegative != 1 {
|
||||
t.Errorf("Sonnet 5 = %+v, want one false negative", model.Implementations)
|
||||
}
|
||||
case "Fable 5":
|
||||
if model.Implementations.TrueNegative != 1 {
|
||||
t.Errorf("Fable 5 = %+v, want one true negative", model.Implementations)
|
||||
}
|
||||
}
|
||||
}
|
||||
if emitted.ReviewMinutes != 180 {
|
||||
t.Errorf("review cost = %d minutes, want 180: four forms recording 45 each", emitted.ReviewMinutes)
|
||||
}
|
||||
for _, want := range []string{"precision", "recall", "Sonnet 5", "Opus 5", "Fable 5",
|
||||
"review cost over the scored implementations: 180 minutes"} {
|
||||
if !strings.Contains(stdout, want) {
|
||||
t.Errorf("the report never prints %q:\n%s", want, stdout)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCannotTellIsEvidenceNeitherWay(t *testing.T) {
|
||||
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
Review: &fixtureReview{Overall: overallNotDefective, Clauses: map[string]string{
|
||||
"R17": clauseCannotTell,
|
||||
"R18": clauseCannotTell,
|
||||
}},
|
||||
}}, defaultMapping))
|
||||
|
||||
if got := emitted.Clauses.CannotTell; got != 2 {
|
||||
t.Fatalf("clause matrix reports %d cannot-tell pair(s), want 2", got)
|
||||
}
|
||||
if got := emitted.Clauses.Scored; got != clauseCount-2 {
|
||||
t.Errorf("clause matrix scored %d pair(s), want %d: a specification error is not a true negative", got, clauseCount-2)
|
||||
}
|
||||
row := outcomeFor(t, emitted, "impl-01")
|
||||
if want := []string{"R17", "R18"}; !equalStrings(row.CannotTellClauses, want) {
|
||||
t.Errorf("impl-01 reports cannot-tell clauses %v, want %v", row.CannotTellClauses, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUncoveredClauseMissIsSeparatedFromACoveredOne(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Fable 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
|
||||
}}, defaultMapping))
|
||||
|
||||
if got := emitted.Clauses.FalseNegative; got != 1 {
|
||||
t.Fatalf("clause matrix reports %d false negative(s), want 1 for R17", got)
|
||||
}
|
||||
if got := emitted.Clauses.UncoveredFalseNegative; got != 1 {
|
||||
t.Errorf("clause matrix reports %d false negative(s) on a clause no property covers, want 1", got)
|
||||
}
|
||||
if want := []string{"R5", "R14", "R15"}; !equalStrings(emitted.Coverage.ClausesCovered, want) {
|
||||
t.Errorf("coverage reports %v covered, want %v", emitted.Coverage.ClausesCovered, want)
|
||||
}
|
||||
if !strings.Contains(stdout, "no property covers R1, R2") {
|
||||
t.Errorf("the report never names the clauses no property covers:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAdjudicatedLabelReplacesTheFirstRaters(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseCannotTell}},
|
||||
Adjudicated: map[string]string{"R15": clauseViolates},
|
||||
}}, defaultMapping))
|
||||
|
||||
if got := emitted.Clauses.TruePositive; got != 1 {
|
||||
t.Fatalf("clause matrix reports %d true positive(s), want 1: R15 resolved to %s", got, clauseViolates)
|
||||
}
|
||||
if got := emitted.Clauses.CannotTell; got != 0 {
|
||||
t.Errorf("clause matrix reports %d cannot-tell pair(s), want 0: the adjudicated label replaces it", got)
|
||||
}
|
||||
if !strings.Contains(stdout, "uses the adjudicated label for R15") {
|
||||
t.Errorf("the report never says an adjudicated label was used:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFiredPropertyOutsideTheMappingIsReported(t *testing.T) {
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Sonnet 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "orderingHoldsAtTheServer")},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R17": clauseViolates}},
|
||||
}}, defaultMapping))
|
||||
|
||||
if want := []string{"orderingHoldsAtTheServer"}; !equalStrings(emitted.Coverage.FiredPropertiesNotMapped, want) {
|
||||
t.Fatalf("coverage reports unmapped fired properties %v, want %v", emitted.Coverage.FiredPropertiesNotMapped, want)
|
||||
}
|
||||
if row := outcomeFor(t, emitted, "impl-01"); row.Cell != cellTruePositive {
|
||||
t.Errorf("impl-01 landed in %q, want %q: an unmapped property still fired", row.Cell, cellTruePositive)
|
||||
}
|
||||
if !strings.Contains(stdout, "could not be attributed to a clause") {
|
||||
t.Errorf("the report never flags the unmapped property:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
func equalStrings(got, want []string) bool {
|
||||
if len(got) != len(want) {
|
||||
return false
|
||||
}
|
||||
for index := range got {
|
||||
if got[index] != want[index] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
@@ -0,0 +1,247 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func runExpectingError(t *testing.T, built fixture) string {
|
||||
t.Helper()
|
||||
var stdout, stderr bytes.Buffer
|
||||
err := run([]string{
|
||||
"--sweep", built.Sweep,
|
||||
"--reviews", built.Reviews,
|
||||
"--assignment", built.Assignment,
|
||||
"--property-clauses", built.Mapping,
|
||||
}, &stdout, &stderr)
|
||||
if err == nil {
|
||||
t.Fatalf("the tool reported success on input it must refuse:\n%s", stdout.String())
|
||||
}
|
||||
return err.Error()
|
||||
}
|
||||
|
||||
func scoredFixture(name, model string) fixtureImplementation {
|
||||
return fixtureImplementation{
|
||||
Name: name, Model: model,
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
}
|
||||
}
|
||||
|
||||
// Three flags missing is one rerun, not three: the operator is told about all
|
||||
// of them at once, in flag order, whatever order the check happened to walk.
|
||||
func TestRunNamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
|
||||
var stdout, stderr bytes.Buffer
|
||||
err := run([]string{"--sweep", "s"}, &stdout, &stderr)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want every missing flag named")
|
||||
}
|
||||
message := err.Error()
|
||||
previous := -1
|
||||
for _, name := range []string{"--reviews", "--assignment", "--property-clauses"} {
|
||||
at := strings.Index(message, name)
|
||||
if at < 0 {
|
||||
t.Fatalf("got %q, want %s named", message, name)
|
||||
}
|
||||
if at < previous {
|
||||
t.Errorf("got %q, want the flags named in flag order", message)
|
||||
}
|
||||
previous = at
|
||||
}
|
||||
if strings.Contains(message, "--sweep") {
|
||||
t.Errorf("got %q, want the supplied --sweep left out", message)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMappingRefusesInputTheMatrixCannotBeScoredFrom(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
mapping string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
name: "a property reads a surface the surface table never declares",
|
||||
mapping: "| property | clauses | surfaces |\n| p | R1 | badgeRow |\n" +
|
||||
"| surface | never observed | note |\n| composer | unlocatable | |\n",
|
||||
want: "the surface table does not declare",
|
||||
},
|
||||
{
|
||||
name: "no surface is declared at all",
|
||||
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n",
|
||||
want: "declares no surfaces",
|
||||
},
|
||||
{
|
||||
name: "a property names a clause outside the requirement",
|
||||
mapping: "| property | clauses | surfaces |\n| p | R21 | none |\n" +
|
||||
"| surface | never observed | note |\n| composer | unlocatable | |\n",
|
||||
want: "which is not one of R1 to R20",
|
||||
},
|
||||
{
|
||||
name: "a surface says something other than what a miss means",
|
||||
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n" +
|
||||
"| surface | never observed | note |\n| composer | maybe | |\n",
|
||||
want: "want unlocatable or inconclusive",
|
||||
},
|
||||
{
|
||||
name: "one property is mapped twice",
|
||||
mapping: "| property | clauses | surfaces |\n| p | R1 | none |\n| p | R2 | none |\n" +
|
||||
"| surface | never observed | note |\n| composer | unlocatable | |\n",
|
||||
want: "is mapped twice",
|
||||
},
|
||||
}
|
||||
for _, test := range tests {
|
||||
t.Run(test.name, func(t *testing.T) {
|
||||
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, test.mapping)
|
||||
if got := runExpectingError(t, built); !strings.Contains(got, test.want) {
|
||||
t.Fatalf("error %q does not name the problem %q", got, test.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestAssignmentRefusesAModelTheSampleWasNotDrawnFrom(t *testing.T) {
|
||||
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, defaultMapping)
|
||||
mustWrite(t, built.Assignment, "| implementation | model |\n| impl-01 | Opus 4 |\n")
|
||||
if got := runExpectingError(t, built); !strings.Contains(got, "which is none of Sonnet 5, Opus 5, Fable 5") {
|
||||
t.Fatalf("error %q does not refuse the unknown model", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSweepRecordingAnImplementationTwiceIsRefused(t *testing.T) {
|
||||
built := writeFixture(t, []fixtureImplementation{scoredFixture("impl-01", "Opus 5")}, defaultMapping)
|
||||
path := filepath.Join(built.Sweep, sweepRecordsFileName)
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
mustWrite(t, path, string(body)+string(body))
|
||||
if got := runExpectingError(t, built); !strings.Contains(got, "a second time") {
|
||||
t.Fatalf("error %q does not refuse the repeated implementation", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMalformedVerdictFormsAreExcludedWithTheirReason(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
body string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
name: "a clause carries a word that is not a verdict",
|
||||
body: reviewWithRow("| R7 | probably fine | |"),
|
||||
want: "want meets, violates or cannot tell",
|
||||
},
|
||||
{
|
||||
name: "the form closes with no overall verdict",
|
||||
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
|
||||
"overall: not defective", "", 1),
|
||||
want: "no overall verdict",
|
||||
},
|
||||
{
|
||||
name: "the overall verdict is neither answer",
|
||||
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
|
||||
"overall: not defective", "overall: mostly ok", 1),
|
||||
want: "is neither defective nor not defective",
|
||||
},
|
||||
{
|
||||
name: "a clause is filed twice with two labels",
|
||||
body: renderReview(fixtureReview{Overall: overallNotDefective}) +
|
||||
"\n| R3 | violates | filed again |\n",
|
||||
want: "is filed twice",
|
||||
},
|
||||
{
|
||||
name: "a clause row is missing",
|
||||
body: strings.Replace(renderReview(fixtureReview{Overall: overallNotDefective}),
|
||||
"| R12 | meets | seen by hand | offline, compose, online |\n", "", 1),
|
||||
want: "no row for clause R12",
|
||||
},
|
||||
}
|
||||
for _, test := range tests {
|
||||
t.Run(test.name, func(t *testing.T) {
|
||||
built := writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1)},
|
||||
RawReview: test.body,
|
||||
}}, defaultMapping)
|
||||
emitted, stdout := runTool(t, built)
|
||||
entry := exclusionFor(t, emitted, "impl-01")
|
||||
if entry.Reason != missingMalformed {
|
||||
t.Fatalf("impl-01 excluded as %q, want %q", entry.Reason, missingMalformed)
|
||||
}
|
||||
if !strings.Contains(entry.Detail, test.want) {
|
||||
t.Errorf("exclusion detail %q does not name %q", entry.Detail, test.want)
|
||||
}
|
||||
if !strings.Contains(stdout, missingMalformed) {
|
||||
t.Errorf("the report never prints the malformed form:\n%s", stdout)
|
||||
}
|
||||
if emitted.Implementations.Scored != 0 {
|
||||
t.Errorf("a malformed form scored %d implementation(s), want none", emitted.Implementations.Scored)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestAnUnfilledMappingStillReportsTheImplementationMatrix(t *testing.T) {
|
||||
unfilled := "| property | clauses | surfaces |\n| TODO | TODO | TODO |\n" +
|
||||
"| surface | never observed | note |\n| stateWords | unlocatable | R4 obliges one |\n"
|
||||
emitted, stdout := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Opus 5",
|
||||
Runs: []fixtureRun{cleanRun(1, "serverHoldsEachMessageOnce")},
|
||||
Review: &fixtureReview{Overall: overallDefective, Clauses: map[string]string{"R15": clauseViolates}},
|
||||
}}, unfilled))
|
||||
|
||||
if emitted.Implementations.TruePositive != 1 {
|
||||
t.Fatalf("implementation matrix = %+v, want one true positive", emitted.Implementations)
|
||||
}
|
||||
if emitted.Clauses.TruePositive != 0 || emitted.Clauses.FalseNegative != 1 {
|
||||
t.Errorf("clause matrix = %+v, want every clause uncovered", emitted.Clauses)
|
||||
}
|
||||
if emitted.Coverage.MappingTodoRows != 1 {
|
||||
t.Errorf("coverage reports %d TODO row(s), want 1", emitted.Coverage.MappingTodoRows)
|
||||
}
|
||||
if !strings.Contains(stdout, "maps no property to a clause") {
|
||||
t.Errorf("the report never says the mapping is unfilled:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAnImplementationOutsideTheAssignmentIsMissingData(t *testing.T) {
|
||||
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{
|
||||
scoredFixture("impl-01", "Opus 5"),
|
||||
{
|
||||
Name: "impl-02",
|
||||
Runs: []fixtureRun{cleanRun(1)}, Review: &fixtureReview{Overall: overallNotDefective},
|
||||
},
|
||||
}, defaultMapping))
|
||||
|
||||
if entry := exclusionFor(t, emitted, "impl-02"); entry.Reason != missingNoModel {
|
||||
t.Fatalf("impl-02 excluded as %q, want %q", entry.Reason, missingNoModel)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSurfacesArePooledAcrossTheSeedsOneImplementationWasSweptAt(t *testing.T) {
|
||||
emitted, _ := runTool(t, writeFixture(t, []fixtureImplementation{{
|
||||
Name: "impl-01", Model: "Fable 5",
|
||||
Runs: []fixtureRun{
|
||||
{Seed: 1, Surfaces: map[string]bool{"composer": true}},
|
||||
{Seed: 2, Surfaces: map[string]bool{"stateWords": true}},
|
||||
},
|
||||
Review: &fixtureReview{Overall: overallNotDefective},
|
||||
}}, defaultMapping))
|
||||
|
||||
row := outcomeFor(t, emitted, "impl-01")
|
||||
if len(row.UnlocatableSurfaces) != 0 {
|
||||
t.Fatalf("impl-01 reports %v unlocatable, want none: each surface was located on one seed or the other",
|
||||
row.UnlocatableSurfaces)
|
||||
}
|
||||
if emitted.Clauses.UnevaluatedSurfaceMissed != 0 {
|
||||
t.Errorf("clause matrix reports %d unevaluated pair(s), want none", emitted.Clauses.UnevaluatedSurfaceMissed)
|
||||
}
|
||||
}
|
||||
|
||||
func reviewWithRow(row string) string {
|
||||
body := renderReview(fixtureReview{Overall: overallNotDefective})
|
||||
return strings.Replace(body, "| R7 | meets | seen by hand | offline, compose, online |", row, 1)
|
||||
}
|
||||
@@ -0,0 +1,169 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"io"
|
||||
"maps"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"text/tabwriter"
|
||||
)
|
||||
|
||||
// labelledMatrix is one row of a printed matrix: the whole sample, or one model.
|
||||
type labelledMatrix struct {
|
||||
Label string
|
||||
Value clauseMatrix
|
||||
Excluded int
|
||||
}
|
||||
|
||||
func writeReport(outcome result, out io.Writer) {
|
||||
fmt.Fprintln(out, "unit: implementation whose own suite passed, scored on whether a property fired and on the reviewer's overall verdict")
|
||||
fmt.Fprintln(out, "an implementation with no cell is missing data and is listed separately, never counted as a clean run")
|
||||
fmt.Fprintln(out)
|
||||
writeTable(out, []string{"group", "scored", "true pos", "false pos", "false neg", "true neg", "precision", "recall", "no cell"},
|
||||
func(add func(...string)) {
|
||||
for _, row := range implementationRows(outcome) {
|
||||
add(
|
||||
row.Label,
|
||||
strconv.Itoa(row.Value.Scored),
|
||||
strconv.Itoa(row.Value.TruePositive),
|
||||
strconv.Itoa(row.Value.FalsePositive),
|
||||
strconv.Itoa(row.Value.FalseNegative),
|
||||
strconv.Itoa(row.Value.TrueNegative),
|
||||
formatRatio(row.Value.Precision),
|
||||
formatRatio(row.Value.Recall),
|
||||
strconv.Itoa(row.Excluded),
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
fmt.Fprintln(out)
|
||||
fmt.Fprintln(out, "clause pairs are an implementation against one of R1 to R20, scored through the declared property-to-clause mapping")
|
||||
fmt.Fprintln(out, "cannot tell is the reviewer's specification-error answer and is evidence neither way")
|
||||
fmt.Fprintln(out, "unevaluated is a clause whose every covering property a never-located surface left unrunnable: a portability miss, not a cell")
|
||||
writeTable(out, []string{"group", "scored", "true pos", "false pos", "false neg", "true neg",
|
||||
"precision", "recall", "cannot tell", "unevaluated", "uncovered", "uncovered false neg"},
|
||||
func(add func(...string)) {
|
||||
for _, row := range clauseRows(outcome) {
|
||||
add(
|
||||
row.Label,
|
||||
strconv.Itoa(row.Value.Scored),
|
||||
strconv.Itoa(row.Value.TruePositive),
|
||||
strconv.Itoa(row.Value.FalsePositive),
|
||||
strconv.Itoa(row.Value.FalseNegative),
|
||||
strconv.Itoa(row.Value.TrueNegative),
|
||||
formatRatio(row.Value.Precision),
|
||||
formatRatio(row.Value.Recall),
|
||||
strconv.Itoa(row.Value.CannotTell),
|
||||
strconv.Itoa(row.Value.UnevaluatedSurfaceMissed),
|
||||
strconv.Itoa(row.Value.UncoveredScored),
|
||||
strconv.Itoa(row.Value.UncoveredFalseNegative),
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
fmt.Fprintln(out)
|
||||
writeTable(out, []string{"implementation", "model", "cell", "usable runs", "fired", "violated clauses", "cannot tell"},
|
||||
func(add func(...string)) {
|
||||
for _, row := range outcome.Outcomes {
|
||||
add(
|
||||
row.Implementation,
|
||||
row.Model,
|
||||
row.Cell,
|
||||
strconv.Itoa(row.RunsUsable),
|
||||
joinOrDash(row.FiredProperties),
|
||||
joinOrDash(row.ViolatedClauses),
|
||||
strconv.Itoa(len(row.CannotTellClauses)),
|
||||
)
|
||||
}
|
||||
})
|
||||
|
||||
if len(outcome.Excluded) > 0 {
|
||||
fmt.Fprintf(out, "\n%d implementation(s) carry no cell: missing data, never a clean run\n", len(outcome.Excluded))
|
||||
writeTable(out, []string{"implementation", "model", "reason", "detail"}, func(add func(...string)) {
|
||||
for _, entry := range outcome.Excluded {
|
||||
add(entry.Implementation, orDash(entry.Model), entry.Reason, orDash(entry.Detail))
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
fmt.Fprintf(out, "\nspecification portability over the %d scored implementation(s): %d needed a locating adaptation\n",
|
||||
outcome.Portability.Scored, outcome.Portability.WithUnlocatableSurface)
|
||||
if len(outcome.Portability.BySurface) > 0 {
|
||||
writeTable(out, []string{"surface never located", "implementations"}, func(add func(...string)) {
|
||||
for _, surface := range slices.Sorted(maps.Keys(outcome.Portability.BySurface)) {
|
||||
add(surface, strconv.Itoa(outcome.Portability.BySurface[surface]))
|
||||
}
|
||||
})
|
||||
}
|
||||
if len(outcome.Portability.SurfacesReadAsInconclusive) > 0 {
|
||||
fmt.Fprintf(out, "a miss on %s cannot be told from that surface being legitimately absent, so neither counts against portability\n",
|
||||
strings.Join(outcome.Portability.SurfacesReadAsInconclusive, ", "))
|
||||
}
|
||||
|
||||
fmt.Fprintf(out, "\nclause coverage: %d mapped propert(ies) cover %d of %d clauses\n",
|
||||
outcome.Coverage.MappedProperties, len(outcome.Coverage.ClausesCovered), clauseCount)
|
||||
if len(outcome.Coverage.ClausesUncovered) > 0 {
|
||||
fmt.Fprintf(out, "no property covers %s\n", strings.Join(outcome.Coverage.ClausesUncovered, ", "))
|
||||
}
|
||||
fmt.Fprintf(out, "review cost over the scored implementations: %d minutes\n", outcome.ReviewMinutes)
|
||||
|
||||
for _, note := range outcome.Notes {
|
||||
fmt.Fprintf(out, "\nnote: %s\n", note)
|
||||
}
|
||||
}
|
||||
|
||||
func implementationRows(outcome result) []labelledMatrix {
|
||||
rows := []labelledMatrix{{
|
||||
Label: "all",
|
||||
Value: clauseMatrix{matrix: outcome.Implementations},
|
||||
Excluded: len(outcome.Excluded),
|
||||
}}
|
||||
for _, model := range outcome.ByModel {
|
||||
rows = append(rows, labelledMatrix{
|
||||
Label: model.Model,
|
||||
Value: clauseMatrix{matrix: model.Implementations},
|
||||
Excluded: model.Excluded,
|
||||
})
|
||||
}
|
||||
return rows
|
||||
}
|
||||
|
||||
func clauseRows(outcome result) []labelledMatrix {
|
||||
rows := []labelledMatrix{{Label: "all", Value: outcome.Clauses}}
|
||||
for _, model := range outcome.ByModel {
|
||||
rows = append(rows, labelledMatrix{Label: model.Model, Value: model.Clauses})
|
||||
}
|
||||
return rows
|
||||
}
|
||||
|
||||
func writeTable(out io.Writer, header []string, rows func(add func(...string))) {
|
||||
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
|
||||
fmt.Fprintln(writer, strings.Join(header, "\t"))
|
||||
rows(func(cells ...string) {
|
||||
fmt.Fprintln(writer, strings.Join(cells, "\t"))
|
||||
})
|
||||
writer.Flush()
|
||||
}
|
||||
|
||||
func formatRatio(value *float64) string {
|
||||
if value == nil {
|
||||
return "n/a"
|
||||
}
|
||||
return strconv.FormatFloat(*value, 'f', 3, 64)
|
||||
}
|
||||
|
||||
func joinOrDash(values []string) string {
|
||||
if len(values) == 0 {
|
||||
return "-"
|
||||
}
|
||||
return strings.Join(values, " ")
|
||||
}
|
||||
|
||||
func orDash(value string) string {
|
||||
if value == "" {
|
||||
return "-"
|
||||
}
|
||||
return value
|
||||
}
|
||||
@@ -0,0 +1,237 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
const (
|
||||
clauseMeets = "meets"
|
||||
clauseViolates = "violates"
|
||||
clauseCannotTell = "cannot tell"
|
||||
)
|
||||
|
||||
const (
|
||||
overallDefective = "defective"
|
||||
overallNotDefective = "not defective"
|
||||
)
|
||||
|
||||
// reviewVerdict is one filed verdict form: the twenty clause rows and the
|
||||
// overall verdict review-protocol.md requires, with any adjudicated label
|
||||
// already substituted for the first rater's.
|
||||
type reviewVerdict struct {
|
||||
Implementation string
|
||||
Path string
|
||||
Reviewer string
|
||||
Date string
|
||||
Minutes int
|
||||
Overall string
|
||||
Clauses map[string]string
|
||||
Adjudicated []string
|
||||
}
|
||||
|
||||
func (v reviewVerdict) defective() bool { return v.Overall == overallDefective }
|
||||
|
||||
func (v reviewVerdict) violatedClauses() []string {
|
||||
var violated []string
|
||||
for _, clause := range allClauses() {
|
||||
if v.Clauses[clause] == clauseViolates {
|
||||
violated = append(violated, clause)
|
||||
}
|
||||
}
|
||||
return violated
|
||||
}
|
||||
|
||||
type malformedReview struct {
|
||||
Implementation string
|
||||
Path string
|
||||
Reason string
|
||||
}
|
||||
|
||||
type reviewSide struct {
|
||||
Directory string
|
||||
Verdicts []reviewVerdict
|
||||
Malformed []malformedReview
|
||||
}
|
||||
|
||||
var (
|
||||
reviewFileName = regexp.MustCompile(`^(impl-\d+)\.md$`)
|
||||
adjudicationName = regexp.MustCompile(`^(impl-\d+)-adjudication\.md$`)
|
||||
secondRaterName = regexp.MustCompile(`^impl-\d+-r2\.md$`)
|
||||
keyValueLine = regexp.MustCompile(`(?i)^\s*[-*]?\s*(reviewer|date|minutes|overall verdict|overall)\s*:\s*(.+?)\s*$`)
|
||||
leadingWholeNumbers = regexp.MustCompile(`^\d+`)
|
||||
)
|
||||
|
||||
func loadReviews(directory string) (reviewSide, error) {
|
||||
entries, err := os.ReadDir(directory)
|
||||
if err != nil {
|
||||
return reviewSide{}, fmt.Errorf("read reviews: %w", err)
|
||||
}
|
||||
side := reviewSide{Directory: directory}
|
||||
adjudications := map[string]map[string]string{}
|
||||
var forms []struct {
|
||||
name string
|
||||
path string
|
||||
}
|
||||
for _, entry := range entries {
|
||||
if entry.IsDir() {
|
||||
continue
|
||||
}
|
||||
path := filepath.Join(directory, entry.Name())
|
||||
switch {
|
||||
case secondRaterName.MatchString(entry.Name()):
|
||||
continue
|
||||
case adjudicationName.MatchString(entry.Name()):
|
||||
name := adjudicationName.FindStringSubmatch(entry.Name())[1]
|
||||
resolved, err := readAdjudication(path)
|
||||
if err != nil {
|
||||
side.Malformed = append(side.Malformed, malformedReview{
|
||||
Implementation: name, Path: path, Reason: err.Error(),
|
||||
})
|
||||
continue
|
||||
}
|
||||
adjudications[name] = resolved
|
||||
case reviewFileName.MatchString(entry.Name()):
|
||||
forms = append(forms, struct {
|
||||
name string
|
||||
path string
|
||||
}{reviewFileName.FindStringSubmatch(entry.Name())[1], path})
|
||||
}
|
||||
}
|
||||
|
||||
for _, form := range forms {
|
||||
verdict, err := readReview(form.name, form.path)
|
||||
if err != nil {
|
||||
side.Malformed = append(side.Malformed, malformedReview{
|
||||
Implementation: form.name, Path: form.path, Reason: err.Error(),
|
||||
})
|
||||
continue
|
||||
}
|
||||
for clause, label := range adjudications[form.name] {
|
||||
if verdict.Clauses[clause] != label {
|
||||
verdict.Adjudicated = append(verdict.Adjudicated, clause)
|
||||
}
|
||||
verdict.Clauses[clause] = label
|
||||
}
|
||||
slices.Sort(verdict.Adjudicated)
|
||||
side.Verdicts = append(side.Verdicts, verdict)
|
||||
}
|
||||
slices.SortFunc(side.Verdicts, func(a, b reviewVerdict) int {
|
||||
return strings.Compare(a.Implementation, b.Implementation)
|
||||
})
|
||||
slices.SortFunc(side.Malformed, func(a, b malformedReview) int {
|
||||
return strings.Compare(a.Path, b.Path)
|
||||
})
|
||||
return side, nil
|
||||
}
|
||||
|
||||
func readReview(name, path string) (reviewVerdict, error) {
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return reviewVerdict{}, err
|
||||
}
|
||||
verdict := reviewVerdict{Implementation: name, Path: path, Clauses: map[string]string{}}
|
||||
for _, line := range strings.Split(string(body), "\n") {
|
||||
match := keyValueLine.FindStringSubmatch(normalizeCell(line))
|
||||
if match == nil {
|
||||
continue
|
||||
}
|
||||
value := strings.TrimSpace(match[2])
|
||||
switch strings.ToLower(match[1]) {
|
||||
case "reviewer":
|
||||
verdict.Reviewer = value
|
||||
case "date":
|
||||
verdict.Date = value
|
||||
case "minutes":
|
||||
if digits := leadingWholeNumbers.FindString(value); digits != "" {
|
||||
verdict.Minutes, _ = strconv.Atoi(digits)
|
||||
}
|
||||
case "overall", "overall verdict":
|
||||
overall, ok := canonicalOverall(value)
|
||||
if !ok {
|
||||
return reviewVerdict{}, fmt.Errorf("overall verdict %q is neither %s nor %s", value, overallDefective, overallNotDefective)
|
||||
}
|
||||
verdict.Overall = overall
|
||||
}
|
||||
}
|
||||
clauses, err := readClauseRows(string(body))
|
||||
if err != nil {
|
||||
return reviewVerdict{}, err
|
||||
}
|
||||
verdict.Clauses = clauses
|
||||
for _, clause := range allClauses() {
|
||||
if _, filed := verdict.Clauses[clause]; !filed {
|
||||
return reviewVerdict{}, fmt.Errorf("no row for clause %s: the form must carry all %d", clause, clauseCount)
|
||||
}
|
||||
}
|
||||
if verdict.Overall == "" {
|
||||
return reviewVerdict{}, fmt.Errorf("no overall verdict: the matrix scores the reviewer's own %s or %s", overallDefective, overallNotDefective)
|
||||
}
|
||||
return verdict, nil
|
||||
}
|
||||
|
||||
func readClauseRows(body string) (map[string]string, error) {
|
||||
clauses := map[string]string{}
|
||||
for _, row := range parseTableRows(body) {
|
||||
clause, ok := canonicalClause(row.cell(0))
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
label, ok := canonicalClauseVerdict(row.cell(1))
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("clause %s has verdict %q, want %s, %s or %s",
|
||||
clause, row.cell(1), clauseMeets, clauseViolates, clauseCannotTell)
|
||||
}
|
||||
if existing, seen := clauses[clause]; seen && existing != label {
|
||||
return nil, fmt.Errorf("clause %s is filed twice, as %q and %q", clause, existing, label)
|
||||
}
|
||||
clauses[clause] = label
|
||||
}
|
||||
return clauses, nil
|
||||
}
|
||||
|
||||
func readAdjudication(path string) (map[string]string, error) {
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
resolved, err := readClauseRows(string(body))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if len(resolved) == 0 {
|
||||
return nil, fmt.Errorf("no resolved clause rows")
|
||||
}
|
||||
return resolved, nil
|
||||
}
|
||||
|
||||
func canonicalClauseVerdict(value string) (string, bool) {
|
||||
switch strings.ToLower(strings.TrimSpace(value)) {
|
||||
case clauseMeets, "meet", "met":
|
||||
return clauseMeets, true
|
||||
case clauseViolates, "violate", "violated":
|
||||
return clauseViolates, true
|
||||
case clauseCannotTell, "cannot-tell", "cannot_tell", "can't tell", "cant tell":
|
||||
return clauseCannotTell, true
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
}
|
||||
|
||||
func canonicalOverall(value string) (string, bool) {
|
||||
cleaned := strings.ToLower(strings.TrimSpace(value))
|
||||
cleaned = strings.TrimSuffix(cleaned, ".")
|
||||
switch cleaned {
|
||||
case overallDefective:
|
||||
return overallDefective, true
|
||||
case overallNotDefective, "not-defective", "no defect", "clean":
|
||||
return overallNotDefective, true
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,79 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"regexp"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// tableRow is one pipe-delimited markdown row with its 1-based line number, so
|
||||
// a malformed cell can name the line the author has to go and fix.
|
||||
type tableRow struct {
|
||||
Line int
|
||||
Cells []string
|
||||
}
|
||||
|
||||
var separatorCell = regexp.MustCompile(`^:?-{2,}:?$`)
|
||||
|
||||
func parseTableRows(body string) []tableRow {
|
||||
var rows []tableRow
|
||||
for index, raw := range strings.Split(body, "\n") {
|
||||
line := strings.TrimSpace(raw)
|
||||
if !strings.HasPrefix(line, "|") {
|
||||
continue
|
||||
}
|
||||
cells := splitCells(line)
|
||||
if len(cells) == 0 || isSeparatorRow(cells) {
|
||||
continue
|
||||
}
|
||||
rows = append(rows, tableRow{Line: index + 1, Cells: cells})
|
||||
}
|
||||
return rows
|
||||
}
|
||||
|
||||
func splitCells(line string) []string {
|
||||
trimmed := strings.Trim(line, "|")
|
||||
parts := strings.Split(trimmed, "|")
|
||||
cells := make([]string, 0, len(parts))
|
||||
for _, part := range parts {
|
||||
cells = append(cells, normalizeCell(part))
|
||||
}
|
||||
return cells
|
||||
}
|
||||
|
||||
func isSeparatorRow(cells []string) bool {
|
||||
for _, cell := range cells {
|
||||
if !separatorCell.MatchString(cell) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func normalizeCell(value string) string {
|
||||
cleaned := strings.ReplaceAll(value, "**", "")
|
||||
cleaned = strings.ReplaceAll(cleaned, "`", "")
|
||||
return strings.Join(strings.Fields(cleaned), " ")
|
||||
}
|
||||
|
||||
func (row tableRow) cell(index int) string {
|
||||
if index >= len(row.Cells) {
|
||||
return ""
|
||||
}
|
||||
return row.Cells[index]
|
||||
}
|
||||
|
||||
var listSeparators = regexp.MustCompile(`[,\s]+`)
|
||||
|
||||
func splitList(value string) []string {
|
||||
trimmed := strings.TrimSpace(value)
|
||||
if trimmed == "" {
|
||||
return nil
|
||||
}
|
||||
var items []string
|
||||
for _, item := range listSeparators.Split(trimmed, -1) {
|
||||
if item != "" {
|
||||
items = append(items, item)
|
||||
}
|
||||
}
|
||||
return items
|
||||
}
|
||||
@@ -0,0 +1,174 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// The corpus is tastejs/todomvc. Its examples/ directory holds 48
|
||||
// implementations of one requirement; five are excluded and the 43 that remain
|
||||
// are this experiment's population.
|
||||
//
|
||||
// The population is named here rather than read from whatever directories
|
||||
// happen to exist, so a corpus at the wrong commit fails verifyCorpus instead
|
||||
// of quietly sweeping a different sample and reporting it as this one.
|
||||
var includedImplementations = []string{
|
||||
"angular-dart", "angular2", "angular2_es2015", "angularjs", "angularjs_require",
|
||||
"aurelia", "backbone", "backbone_marionette", "backbone_require", "binding-scala",
|
||||
"canjs", "canjs_require", "closure", "dijon", "dojo",
|
||||
"duel", "elm", "emberjs", "enyo_backbone", "exoskeleton",
|
||||
"jquery", "js_of_ocaml", "jsblocks", "knockback", "knockoutjs",
|
||||
"knockoutjs_require", "kotlin-react", "lavaca_require", "mithril", "polymer",
|
||||
"ractive", "react", "react-alt", "react-backbone", "reagent",
|
||||
"riotjs", "scalajs-react", "typescript-angular", "typescript-backbone", "typescript-react",
|
||||
"vanilla-es6", "vanillajs", "vue",
|
||||
}
|
||||
|
||||
// excludedImplementations are the five directories under examples/ that the
|
||||
// corpus survey dropped. They are listed rather than merely omitted so
|
||||
// verifyCorpus can insist the corpus holds exactly these 48 names: an example
|
||||
// added or renamed upstream then stops the sweep rather than silently shrinking
|
||||
// or growing the sample.
|
||||
var excludedImplementations = []string{
|
||||
"cujo", "emberjs_require", "firebase-angular", "gwt", "react-hooks",
|
||||
}
|
||||
|
||||
const examplesDirectory = "examples"
|
||||
|
||||
// documentPath names the served document for the implementations that do not
|
||||
// keep an index.html at the root of their example directory.
|
||||
var documentPath = map[string]string{
|
||||
"angular-dart": "examples/angular-dart/web/index.html",
|
||||
"duel": "examples/duel/www/index.html",
|
||||
}
|
||||
|
||||
func documentFor(name string) string {
|
||||
if override, ok := documentPath[name]; ok {
|
||||
return override
|
||||
}
|
||||
return examplesDirectory + "/" + name + "/index.html"
|
||||
}
|
||||
|
||||
// implementation is one member of the population, with the port it owns for the
|
||||
// whole sweep. The port is what keeps implementations apart: every one is
|
||||
// served on its own, so every one is its own web origin and localStorage keeps
|
||||
// their records in separate partitions. Four pairs in this corpus write the
|
||||
// same key, and one origin between them is one record between them.
|
||||
type implementation struct {
|
||||
Name string
|
||||
Document string
|
||||
Port int
|
||||
}
|
||||
|
||||
func (i implementation) Origin() string {
|
||||
return fmt.Sprintf("http://127.0.0.1:%d", i.Port)
|
||||
}
|
||||
|
||||
func (i implementation) URL() string {
|
||||
return i.Origin() + "/" + i.Document
|
||||
}
|
||||
|
||||
// verifyCorpus insists the corpus holds exactly the 48 example directories this
|
||||
// population was drawn from. A sweep against a different checkout would still
|
||||
// run, and its 43 arms would still be labelled, which is why the check is here
|
||||
// and not left to whoever reads the results.
|
||||
func verifyCorpus(corpusRoot string) error {
|
||||
entries, err := os.ReadDir(filepath.Join(corpusRoot, examplesDirectory))
|
||||
if err != nil {
|
||||
return fmt.Errorf("--corpus: %w", err)
|
||||
}
|
||||
var found []string
|
||||
for _, entry := range entries {
|
||||
if entry.IsDir() {
|
||||
found = append(found, entry.Name())
|
||||
}
|
||||
}
|
||||
expected := slices.Concat(includedImplementations, excludedImplementations)
|
||||
slices.Sort(expected)
|
||||
slices.Sort(found)
|
||||
if slices.Equal(expected, found) {
|
||||
return nil
|
||||
}
|
||||
var missing, unexpected []string
|
||||
for _, name := range expected {
|
||||
if !slices.Contains(found, name) {
|
||||
missing = append(missing, name)
|
||||
}
|
||||
}
|
||||
for _, name := range found {
|
||||
if !slices.Contains(expected, name) {
|
||||
unexpected = append(unexpected, name)
|
||||
}
|
||||
}
|
||||
return fmt.Errorf(
|
||||
"%s/%s is not the corpus this population was drawn from: missing %v, unexpected %v",
|
||||
corpusRoot,
|
||||
examplesDirectory,
|
||||
missing,
|
||||
unexpected,
|
||||
)
|
||||
}
|
||||
|
||||
// selectImplementations resolves --implementations against the population. An
|
||||
// empty selection is the whole population, which is what a real sweep runs; a
|
||||
// named subset is for smoke runs and is recorded in the manifest like any other
|
||||
// intent.
|
||||
func selectImplementations(selection string) ([]string, error) {
|
||||
if strings.TrimSpace(selection) == "" {
|
||||
return slices.Clone(includedImplementations), nil
|
||||
}
|
||||
var names []string
|
||||
seen := map[string]bool{}
|
||||
for _, part := range strings.Split(selection, ",") {
|
||||
name := strings.TrimSpace(part)
|
||||
if name == "" {
|
||||
return nil, fmt.Errorf("empty implementation in %q", selection)
|
||||
}
|
||||
if !slices.Contains(includedImplementations, name) {
|
||||
return nil, fmt.Errorf(
|
||||
"%q is not one of the %d implementations in this population",
|
||||
name,
|
||||
len(includedImplementations),
|
||||
)
|
||||
}
|
||||
if seen[name] {
|
||||
return nil, fmt.Errorf("duplicate implementation %q", name)
|
||||
}
|
||||
seen[name] = true
|
||||
names = append(names, name)
|
||||
}
|
||||
return names, nil
|
||||
}
|
||||
|
||||
// planImplementations gives every selected implementation its document and its
|
||||
// own port, in population order so the manifest can name the URL each arm was
|
||||
// served from before anything has been served.
|
||||
func planImplementations(
|
||||
corpusRoot string,
|
||||
names []string,
|
||||
basePort int,
|
||||
) ([]implementation, error) {
|
||||
if basePort+len(names)-1 > 65535 {
|
||||
return nil, fmt.Errorf(
|
||||
"--base-port %d leaves no room for %d implementations",
|
||||
basePort,
|
||||
len(names),
|
||||
)
|
||||
}
|
||||
planned := make([]implementation, 0, len(names))
|
||||
for index, name := range names {
|
||||
document := documentFor(name)
|
||||
if _, err := os.Stat(filepath.Join(corpusRoot, filepath.FromSlash(document))); err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", name, err)
|
||||
}
|
||||
planned = append(planned, implementation{
|
||||
Name: name,
|
||||
Document: document,
|
||||
Port: basePort + index,
|
||||
})
|
||||
}
|
||||
return planned, nil
|
||||
}
|
||||
@@ -0,0 +1,161 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// writeCorpus builds a tree with the shape the sweep expects: every example
|
||||
// directory the population was drawn from, each holding the document that
|
||||
// implementation is served at.
|
||||
func writeCorpus(t *testing.T, names []string) string {
|
||||
t.Helper()
|
||||
root := t.TempDir()
|
||||
for _, name := range names {
|
||||
if err := os.MkdirAll(filepath.Join(root, examplesDirectory, name), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
document := filepath.Join(root, filepath.FromSlash(documentFor(name)))
|
||||
if err := os.MkdirAll(filepath.Dir(document), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body := "<!doctype html><title>" + name + "</title><ul class=\"todo-list\"></ul>"
|
||||
if err := os.WriteFile(document, []byte(body), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return root
|
||||
}
|
||||
|
||||
func wholeCorpus(t *testing.T) string {
|
||||
t.Helper()
|
||||
return writeCorpus(
|
||||
t,
|
||||
slices.Concat(includedImplementations, excludedImplementations),
|
||||
)
|
||||
}
|
||||
|
||||
func TestPopulation_IsFortyThreeNamesDisjointFromTheExclusions(t *testing.T) {
|
||||
if len(includedImplementations) != 43 {
|
||||
t.Errorf(
|
||||
"population size: got %d, want 43",
|
||||
len(includedImplementations),
|
||||
)
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
for _, name := range includedImplementations {
|
||||
if seen[name] {
|
||||
t.Errorf("%q appears twice in the population", name)
|
||||
}
|
||||
seen[name] = true
|
||||
}
|
||||
for _, name := range excludedImplementations {
|
||||
if seen[name] {
|
||||
t.Errorf("%q is both included and excluded", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestVerifyCorpus_RejectsATreeThatIsNotThisCorpus(t *testing.T) {
|
||||
if err := verifyCorpus(wholeCorpus(t)); err != nil {
|
||||
t.Fatalf(
|
||||
"the corpus this population was drawn from should verify: %v",
|
||||
err,
|
||||
)
|
||||
}
|
||||
|
||||
short := writeCorpus(
|
||||
t,
|
||||
slices.Concat(includedImplementations[1:], excludedImplementations),
|
||||
)
|
||||
err := verifyCorpus(short)
|
||||
if err == nil {
|
||||
t.Fatal(
|
||||
"a corpus missing an implementation swept 42 arms and reported them as 43",
|
||||
)
|
||||
}
|
||||
if !strings.Contains(err.Error(), includedImplementations[0]) {
|
||||
t.Errorf("the error should name what is missing: %v", err)
|
||||
}
|
||||
|
||||
extra := writeCorpus(
|
||||
t,
|
||||
slices.Concat(
|
||||
includedImplementations,
|
||||
excludedImplementations,
|
||||
[]string{"svelte"},
|
||||
),
|
||||
)
|
||||
err = verifyCorpus(extra)
|
||||
if err == nil {
|
||||
t.Fatal(
|
||||
"a corpus with an implementation this population never drew from verified",
|
||||
)
|
||||
}
|
||||
if !strings.Contains(err.Error(), "svelte") {
|
||||
t.Errorf("the error should name what is unexpected: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDocumentFor_ServesTheOverriddenPathForImplementationsWithoutARootIndex(
|
||||
t *testing.T,
|
||||
) {
|
||||
for name, want := range map[string]string{
|
||||
"angular-dart": "examples/angular-dart/web/index.html",
|
||||
"duel": "examples/duel/www/index.html",
|
||||
"vanillajs": "examples/vanillajs/index.html",
|
||||
"react": "examples/react/index.html",
|
||||
} {
|
||||
if got := documentFor(name); got != want {
|
||||
t.Errorf("%s document: got %q, want %q", name, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestPlanImplementations_RefusesAnImplementationWhoseDocumentIsMissing(
|
||||
t *testing.T,
|
||||
) {
|
||||
root := wholeCorpus(t)
|
||||
if err := os.Remove(filepath.Join(root, filepath.FromSlash(documentFor("duel")))); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
_, err := planImplementations(root, []string{"duel"}, 5400)
|
||||
if err == nil {
|
||||
t.Fatal(
|
||||
"an implementation whose document is missing would be swept as a 404 page",
|
||||
)
|
||||
}
|
||||
if !strings.Contains(err.Error(), "duel") {
|
||||
t.Errorf("the error should name the implementation: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSelectImplementations_DefaultsToThePopulationAndRejectsNamesOutsideIt(
|
||||
t *testing.T,
|
||||
) {
|
||||
all, err := selectImplementations("")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !slices.Equal(all, includedImplementations) {
|
||||
t.Errorf(
|
||||
"an empty selection should be the whole population, got %d names",
|
||||
len(all),
|
||||
)
|
||||
}
|
||||
subset, err := selectImplementations("react, angular2_es2015")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !slices.Equal(subset, []string{"react", "angular2_es2015"}) {
|
||||
t.Errorf("subset: got %v", subset)
|
||||
}
|
||||
for _, selection := range []string{"cujo", "svelte", "react,react"} {
|
||||
if _, err := selectImplementations(selection); err == nil {
|
||||
t.Errorf("selection %q should have been rejected", selection)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,468 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"math/rand"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// The test binary doubles as the fetcher the stub campaign uses, so the URL a
|
||||
// campaign is handed is really requested while the sweep is serving, and what
|
||||
// came back is on disk for the test to read.
|
||||
func TestMain(m *testing.M) {
|
||||
if url := os.Getenv("CORPUS_SWEEP_TEST_FETCH_URL"); url != "" {
|
||||
recordFetch(url, os.Getenv("CORPUS_SWEEP_TEST_FETCH_LOG"))
|
||||
return
|
||||
}
|
||||
os.Exit(m.Run())
|
||||
}
|
||||
|
||||
func recordFetch(url, logPath string) {
|
||||
line := ""
|
||||
response, err := http.Get(url)
|
||||
if err != nil {
|
||||
line = fmt.Sprintf("%s -> error %v\n", url, err)
|
||||
} else {
|
||||
body, _ := io.ReadAll(response.Body)
|
||||
response.Body.Close()
|
||||
line = fmt.Sprintf("%s -> %d %s\n", url, response.StatusCode, body)
|
||||
}
|
||||
logFile, err := os.OpenFile(
|
||||
logPath,
|
||||
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
|
||||
0o644,
|
||||
)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
logFile.WriteString(line)
|
||||
logFile.Close()
|
||||
}
|
||||
|
||||
// stubCampaign records the argv it was handed and whether the sweep manifest
|
||||
// was already on disk when it ran, fetches the URL it was told to drive, writes
|
||||
// the campaign directory the real tool would write, and fails one seed.
|
||||
const stubCampaign = `#!/bin/sh
|
||||
output=""
|
||||
seed=""
|
||||
url=""
|
||||
previous=""
|
||||
for argument in "$@"; do
|
||||
case "$previous" in
|
||||
--output) output="$argument" ;;
|
||||
--seeds) seed="$argument" ;;
|
||||
--bundle-id) url="$argument" ;;
|
||||
esac
|
||||
previous="$argument"
|
||||
done
|
||||
manifest=missing
|
||||
if [ -f "%[1]s" ]; then manifest=present; fi
|
||||
echo "manifest=$manifest argv: $*" >> "%[2]s"
|
||||
CORPUS_SWEEP_TEST_FETCH_URL="$url" CORPUS_SWEEP_TEST_FETCH_LOG="%[3]s" "%[4]s"
|
||||
mkdir -p "$output"
|
||||
printf '{"seeds":[%%s]}\n' "$seed" > "$output/campaign.json"
|
||||
echo "stub campaign seed=$seed url=$url"
|
||||
case "$output" in
|
||||
*angular2_es2015/seed-5) exit 1 ;;
|
||||
esac
|
||||
exit 0
|
||||
`
|
||||
|
||||
func TestRun_EndToEndAgainstAServedCorpusAndAStubCampaign(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
corpus := wholeCorpus(t)
|
||||
// dojo's document is served but cannot be read, so it stalls at serve and
|
||||
// the two implementations either side of it still have to reach the
|
||||
// campaign tool with both seeds.
|
||||
unreadable := filepath.Join(corpus, filepath.FromSlash(documentFor("dojo")))
|
||||
if err := os.Chmod(unreadable, 0o000); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(func() { os.Chmod(unreadable, 0o644) })
|
||||
|
||||
specPath := filepath.Join(root, "todo.ts")
|
||||
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
output := filepath.Join(root, "campaigns")
|
||||
campaignLog := filepath.Join(root, "campaign.log")
|
||||
fetchLog := filepath.Join(root, "fetch.log")
|
||||
testBinary, err := filepath.Abs(os.Args[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
campaignPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-campaign"),
|
||||
fmt.Sprintf(
|
||||
stubCampaign,
|
||||
filepath.Join(output, manifestFileName),
|
||||
campaignLog,
|
||||
fetchLog,
|
||||
testBinary,
|
||||
),
|
||||
)
|
||||
sanderlingPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-sanderling"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
)
|
||||
selected := []string{"angular2", "angular2_es2015", "dojo"}
|
||||
basePort := freePortRange(t, len(selected))
|
||||
|
||||
var stdout bytes.Buffer
|
||||
err = run([]string{
|
||||
"--corpus", corpus,
|
||||
"--spec", specPath,
|
||||
"--implementations", strings.Join(selected, ","),
|
||||
"--seeds", "4-5",
|
||||
"--max-steps", "40",
|
||||
"--duration", "30s",
|
||||
"--concurrency", "2",
|
||||
"--base-port", fmt.Sprint(basePort),
|
||||
"--output", output,
|
||||
"--campaign", campaignPath,
|
||||
"--sanderling", sanderlingPath,
|
||||
}, &stdout, io.Discard)
|
||||
if err == nil ||
|
||||
!strings.Contains(err.Error(), "1 of 3 implementations never ran") {
|
||||
t.Fatalf(
|
||||
"expected the stalled implementation and the failed campaign to be reported, got %v",
|
||||
err,
|
||||
)
|
||||
}
|
||||
|
||||
recorded := readManifest(t, filepath.Join(output, manifestFileName))
|
||||
if !slices.Equal(recorded.Seeds, []int64{4, 5}) {
|
||||
t.Errorf("intended seeds: got %v", recorded.Seeds)
|
||||
}
|
||||
if len(recorded.Implementations) != len(selected) {
|
||||
t.Fatalf("intended implementations: got %v", recorded.Implementations)
|
||||
}
|
||||
for index, planned := range recorded.Implementations {
|
||||
if planned.Name != selected[index] {
|
||||
t.Errorf(
|
||||
"implementation %d: got %q, want %q",
|
||||
index,
|
||||
planned.Name,
|
||||
selected[index],
|
||||
)
|
||||
}
|
||||
wantPort := basePort + index
|
||||
wantURL := fmt.Sprintf(
|
||||
"http://127.0.0.1:%d/examples/%s/index.html",
|
||||
wantPort,
|
||||
planned.Name,
|
||||
)
|
||||
if planned.Port != wantPort || planned.URL != wantURL {
|
||||
t.Errorf(
|
||||
"%s: port %d url %q, want %d %q",
|
||||
planned.Name,
|
||||
planned.Port,
|
||||
planned.URL,
|
||||
wantPort,
|
||||
wantURL,
|
||||
)
|
||||
}
|
||||
}
|
||||
if recorded.Generator != "seeded" || recorded.Platform != "web" ||
|
||||
recorded.MaxSteps != 40 {
|
||||
t.Errorf(
|
||||
"manifest generator/platform/budget: %q/%q/%d",
|
||||
recorded.Generator,
|
||||
recorded.Platform,
|
||||
recorded.MaxSteps,
|
||||
)
|
||||
}
|
||||
if recorded.CorpusRoot != corpus {
|
||||
t.Errorf(
|
||||
"manifest corpus root: got %q, want %q",
|
||||
recorded.CorpusRoot,
|
||||
corpus,
|
||||
)
|
||||
}
|
||||
|
||||
campaignLines := readLines(t, campaignLog)
|
||||
if len(campaignLines) != 4 {
|
||||
t.Fatalf(
|
||||
"campaign invocations: got %d, want 4:\n%s",
|
||||
len(campaignLines),
|
||||
strings.Join(campaignLines, "\n"),
|
||||
)
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
for _, line := range campaignLines {
|
||||
if !strings.HasPrefix(line, "manifest=present") {
|
||||
t.Errorf(
|
||||
"a campaign ran before the sweep manifest was written: %q",
|
||||
line,
|
||||
)
|
||||
}
|
||||
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
|
||||
arm := argumentValue(arguments, "--arm")
|
||||
seed := argumentValue(arguments, "--seeds")
|
||||
port := basePort + slices.Index(selected, arm)
|
||||
wantBundle := fmt.Sprintf(
|
||||
"http://127.0.0.1:%d/examples/%s/index.html",
|
||||
port,
|
||||
arm,
|
||||
)
|
||||
if got := argumentValue(arguments, "--bundle-id"); got != wantBundle {
|
||||
t.Errorf(
|
||||
"%s seed %s: bundle id %q, want %q",
|
||||
arm,
|
||||
seed,
|
||||
got,
|
||||
wantBundle,
|
||||
)
|
||||
}
|
||||
if got := argumentValue(arguments, "--output"); got != filepath.Join(
|
||||
output,
|
||||
arm,
|
||||
"seed-"+seed,
|
||||
) {
|
||||
t.Errorf("%s seed %s: campaign output %q", arm, seed, got)
|
||||
}
|
||||
if got := argumentValue(arguments, "--sanderling"); got != sanderlingPath {
|
||||
t.Errorf("%s seed %s: sanderling path %q", arm, seed, got)
|
||||
}
|
||||
seen[arm+"/"+seed] = true
|
||||
}
|
||||
for _, want := range []string{"angular2/4", "angular2/5", "angular2_es2015/4", "angular2_es2015/5"} {
|
||||
if !seen[want] {
|
||||
t.Errorf("%s never reached the campaign tool", want)
|
||||
}
|
||||
}
|
||||
|
||||
// What the served page actually returned: each port answered with its own
|
||||
// implementation's document, so no arm was driven against another's.
|
||||
fetched := readLines(t, fetchLog)
|
||||
for index, name := range []string{"angular2", "angular2_es2015"} {
|
||||
want := fmt.Sprintf(
|
||||
"http://127.0.0.1:%d/examples/%s/index.html -> 200 <!doctype html><title>%s</title>",
|
||||
basePort+index,
|
||||
name,
|
||||
name,
|
||||
)
|
||||
matched := 0
|
||||
for _, line := range fetched {
|
||||
if strings.HasPrefix(line, want) {
|
||||
matched++
|
||||
}
|
||||
}
|
||||
if matched != 2 {
|
||||
t.Errorf(
|
||||
"%s was served its own document %d times, want 2:\n%s",
|
||||
name,
|
||||
matched,
|
||||
strings.Join(fetched, "\n"),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
records := readRecords(t, filepath.Join(output, recordsFileName))
|
||||
if len(records) != 3 {
|
||||
t.Fatalf("implementation records: got %d, want 3", len(records))
|
||||
}
|
||||
byName := map[string]implementationRecord{}
|
||||
for _, record := range records {
|
||||
byName[record.Name] = record
|
||||
}
|
||||
stalled := byName["dojo"]
|
||||
if stalled.FailedStage != stageServe || len(stalled.Runs) != 0 {
|
||||
t.Errorf(
|
||||
"dojo: stage %q with %d runs, want a serve failure and no runs",
|
||||
stalled.FailedStage,
|
||||
len(stalled.Runs),
|
||||
)
|
||||
}
|
||||
for _, name := range []string{"angular2", "angular2_es2015"} {
|
||||
record := byName[name]
|
||||
if record.FailedStage != "" || len(record.Runs) != 2 {
|
||||
t.Errorf(
|
||||
"%s: stage %q with %d runs, want no failure and 2 runs",
|
||||
name,
|
||||
record.FailedStage,
|
||||
len(record.Runs),
|
||||
)
|
||||
}
|
||||
if record.MonotonicMillis <= 0 {
|
||||
t.Errorf(
|
||||
"%s took %d ms, so nothing timed how long it worked",
|
||||
name,
|
||||
record.MonotonicMillis,
|
||||
)
|
||||
}
|
||||
}
|
||||
if exit := byName["angular2_es2015"].Runs[1].ExitCode; exit != 1 {
|
||||
t.Errorf("angular2_es2015 seed 5 exit code: got %d, want 1", exit)
|
||||
}
|
||||
if exit := byName["angular2_es2015"].Runs[0].ExitCode; exit != 0 {
|
||||
t.Errorf("angular2_es2015 seed 4 exit code: got %d, want 0", exit)
|
||||
}
|
||||
|
||||
for _, name := range []string{"angular2", "angular2_es2015"} {
|
||||
for _, seed := range []string{"4", "5"} {
|
||||
directory := filepath.Join(output, name, "seed-"+seed)
|
||||
if _, err := os.Stat(filepath.Join(directory, "campaign.json")); err != nil {
|
||||
t.Errorf(
|
||||
"%s seed %s: no campaign directory: %v",
|
||||
name,
|
||||
seed,
|
||||
err,
|
||||
)
|
||||
}
|
||||
log, err := os.ReadFile(filepath.Join(directory, "campaign.log"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(log), "stub campaign seed="+seed) {
|
||||
t.Errorf(
|
||||
"%s seed %s: campaign output was not captured: %q",
|
||||
name,
|
||||
seed,
|
||||
log,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
client := &http.Client{Timeout: 2 * time.Second}
|
||||
for offset := range selected {
|
||||
if response, err := client.Get(fmt.Sprintf("http://127.0.0.1:%d/", basePort+offset)); err == nil {
|
||||
response.Body.Close()
|
||||
t.Errorf(
|
||||
"port %d is still served after the sweep finished",
|
||||
basePort+offset,
|
||||
)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "failed at serve") {
|
||||
t.Errorf(
|
||||
"progress output does not name the stalled implementation: %q",
|
||||
stdout.String(),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRun_RefusesAnOutputDirectoryThatAlreadyHoldsASweep(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
output := filepath.Join(root, "campaigns")
|
||||
if err := os.MkdirAll(output, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(output, manifestFileName), []byte("{}"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
specPath := filepath.Join(root, "todo.ts")
|
||||
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
err := run([]string{
|
||||
"--corpus", wholeCorpus(t), "--spec", specPath, "--seeds", "1",
|
||||
"--max-steps", "10", "--output", output,
|
||||
}, io.Discard, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), manifestFileName) {
|
||||
t.Fatalf(
|
||||
"two sweeps sharing a directory would interleave their records, got %v",
|
||||
err,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func writeScript(t *testing.T, path, body string) string {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(path, []byte(body), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
func readLines(t *testing.T, path string) []string {
|
||||
t.Helper()
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var lines []string
|
||||
scanner := bufio.NewScanner(strings.NewReader(string(body)))
|
||||
for scanner.Scan() {
|
||||
if line := strings.TrimSpace(scanner.Text()); line != "" {
|
||||
lines = append(lines, line)
|
||||
}
|
||||
}
|
||||
return lines
|
||||
}
|
||||
|
||||
func readManifest(t *testing.T, path string) manifest {
|
||||
t.Helper()
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var recorded manifest
|
||||
if err := json.Unmarshal(body, &recorded); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return recorded
|
||||
}
|
||||
|
||||
func readRecords(t *testing.T, path string) []implementationRecord {
|
||||
t.Helper()
|
||||
var records []implementationRecord
|
||||
for _, line := range readLines(t, path) {
|
||||
var record implementationRecord
|
||||
if err := json.Unmarshal([]byte(line), &record); err != nil {
|
||||
t.Fatalf("%s: %v", line, err)
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
return records
|
||||
}
|
||||
|
||||
func argumentValue(arguments []string, name string) string {
|
||||
index := slices.Index(arguments, name)
|
||||
if index < 0 || index+1 >= len(arguments) {
|
||||
return ""
|
||||
}
|
||||
return arguments[index+1]
|
||||
}
|
||||
|
||||
// freePortRange finds count consecutive free ports, which is what the sweep
|
||||
// hands out: one port per implementation from --base-port upwards.
|
||||
func freePortRange(t *testing.T, count int) int {
|
||||
t.Helper()
|
||||
for range 100 {
|
||||
base := 20000 + rand.Intn(20000)
|
||||
if portsAreFree(base, count) {
|
||||
return base
|
||||
}
|
||||
}
|
||||
t.Fatalf("no run of %d free ports", count)
|
||||
return 0
|
||||
}
|
||||
|
||||
func portsAreFree(base, count int) bool {
|
||||
for offset := range count {
|
||||
listener, err := net.Listen(
|
||||
"tcp",
|
||||
fmt.Sprintf("127.0.0.1:%d", base+offset),
|
||||
)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
listener.Close()
|
||||
}
|
||||
return true
|
||||
}
|
||||
@@ -0,0 +1,287 @@
|
||||
// Command corpus-sweep runs one specification against every implementation in a
|
||||
// served corpus of independent implementations of the same requirement. It
|
||||
// serves each one on its own port, which is what keeps them apart: the corpus
|
||||
// holds pairs that write the same localStorage key, and one origin shared
|
||||
// between two of them is one stored record shared between them.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/seedspec"
|
||||
)
|
||||
|
||||
// The generator and the platform are fixed rather than exposed: the
|
||||
// pre-registration runs the seeded policy against the served corpus, and a
|
||||
// sweep that could quietly run something else records a comparison nobody made.
|
||||
const (
|
||||
generator = "seeded"
|
||||
platform = "web"
|
||||
)
|
||||
|
||||
// defaultConcurrency is how many implementations are swept at once. fleet.md
|
||||
// measured eight concurrent web campaigns clean at about 1.1 GB resident each,
|
||||
// parallel efficiency 0.83 at eight against 0.87 at six, and said to re-measure
|
||||
// before trusting anything above eight on a contended host. Six sits at the
|
||||
// better efficiency and leaves the emulator farm that shares the host its
|
||||
// slots. Serving costs nothing here: the corpus needs no build and no separate
|
||||
// server process, so a worker is one browser.
|
||||
const defaultConcurrency = 6
|
||||
|
||||
// defaultBasePort starts above the range the model-implementation sweep hands
|
||||
// out, so the two can run on one host without either being served the other's
|
||||
// pages.
|
||||
const defaultBasePort = 5400
|
||||
|
||||
type config struct {
|
||||
corpusRoot string
|
||||
specPath string
|
||||
outputDirectory string
|
||||
implementations []string
|
||||
seeds []int64
|
||||
maxSteps int
|
||||
duration time.Duration
|
||||
concurrency int
|
||||
basePort int
|
||||
campaignPath string
|
||||
sanderlingPath string
|
||||
extraArguments []string
|
||||
}
|
||||
|
||||
const usage = `corpus-sweep runs one seeded campaign per implementation of a served corpus.
|
||||
|
||||
Usage:
|
||||
corpus-sweep --corpus <dir> --spec <path> --seeds <spec>
|
||||
--max-steps <n> --output <dir> [flags]
|
||||
[-- <sanderling test flags>]
|
||||
|
||||
Every implementation is served on its own port, so every one is its own origin
|
||||
and none can read another's stored state, and every one is swept with the same
|
||||
seeds and the same step budget. Each run's arm is the implementation's name.
|
||||
|
||||
Everything after a bare -- reaches every sanderling test call through the
|
||||
campaign tool.
|
||||
`
|
||||
|
||||
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
flagSet := flag.NewFlagSet("corpus-sweep", flag.ContinueOnError)
|
||||
flagSet.SetOutput(stderr)
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprint(stderr, usage)
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
var configuration config
|
||||
var seedSpecification string
|
||||
var selection string
|
||||
flagSet.StringVar(
|
||||
&configuration.corpusRoot,
|
||||
"corpus",
|
||||
"",
|
||||
"root of the checked-out corpus, holding examples/ (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.specPath,
|
||||
"spec",
|
||||
"",
|
||||
"path to the property set every implementation is run against (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&seedSpecification,
|
||||
"seeds",
|
||||
"",
|
||||
"seeds every implementation runs: ranges and lists, e.g. 1-10,20 (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&selection,
|
||||
"implementations",
|
||||
"",
|
||||
"comma-separated subset of the population to sweep (default: all of it)",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.maxSteps,
|
||||
"max-steps",
|
||||
0,
|
||||
"per-run step budget, identical across implementations (required, must be positive)",
|
||||
)
|
||||
flagSet.DurationVar(
|
||||
&configuration.duration,
|
||||
"duration",
|
||||
5*time.Minute,
|
||||
"per-run wall-clock ceiling passed to each campaign",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.concurrency,
|
||||
"concurrency",
|
||||
defaultConcurrency,
|
||||
"implementations swept at once",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.basePort,
|
||||
"base-port",
|
||||
defaultBasePort,
|
||||
"first port served; each implementation takes the next one in population order",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.outputDirectory,
|
||||
"output",
|
||||
"",
|
||||
"campaign tree to create (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.campaignPath,
|
||||
"campaign",
|
||||
"campaign",
|
||||
"campaign binary to invoke per implementation and seed",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.sanderlingPath,
|
||||
"sanderling",
|
||||
"sanderling",
|
||||
"sanderling binary each campaign invokes",
|
||||
)
|
||||
if err := flagSet.Parse(arguments); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
configuration.extraArguments = flagSet.Args()
|
||||
|
||||
// Every missing flag is named together, in flag order: stopping at the
|
||||
// first turns one rerun into one rerun per missing flag.
|
||||
var missing []error
|
||||
for _, required := range []struct {
|
||||
name string
|
||||
value string
|
||||
}{
|
||||
{"--corpus", configuration.corpusRoot},
|
||||
{"--spec", configuration.specPath},
|
||||
{"--seeds", seedSpecification},
|
||||
{"--output", configuration.outputDirectory},
|
||||
} {
|
||||
if required.value == "" {
|
||||
missing = append(
|
||||
missing,
|
||||
fmt.Errorf("%s is required", required.name),
|
||||
)
|
||||
}
|
||||
}
|
||||
if err := errors.Join(missing...); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
if configuration.maxSteps <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--max-steps must be positive: every implementation needs the same step budget",
|
||||
)
|
||||
}
|
||||
if configuration.duration <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--duration must be positive: %s",
|
||||
configuration.duration,
|
||||
)
|
||||
}
|
||||
if configuration.concurrency <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--concurrency must be positive: %d",
|
||||
configuration.concurrency,
|
||||
)
|
||||
}
|
||||
if configuration.basePort < 1024 || configuration.basePort > 65535 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--base-port %d is outside 1024-65535",
|
||||
configuration.basePort,
|
||||
)
|
||||
}
|
||||
seeds, err := seedspec.Parse(seedSpecification)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("--seeds: %w", err)
|
||||
}
|
||||
configuration.seeds = seeds
|
||||
names, err := selectImplementations(selection)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("--implementations: %w", err)
|
||||
}
|
||||
configuration.implementations = names
|
||||
for name, value := range map[string]*string{
|
||||
"--corpus": &configuration.corpusRoot,
|
||||
"--spec": &configuration.specPath,
|
||||
"--output": &configuration.outputDirectory,
|
||||
} {
|
||||
absolute, err := filepath.Abs(*value)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("%s: %w", name, err)
|
||||
}
|
||||
*value = absolute
|
||||
}
|
||||
return configuration, nil
|
||||
}
|
||||
|
||||
func campaignDirectory(
|
||||
configuration config,
|
||||
target implementation,
|
||||
seed string,
|
||||
) string {
|
||||
return filepath.Join(
|
||||
configuration.outputDirectory,
|
||||
target.Name,
|
||||
"seed-"+seed,
|
||||
)
|
||||
}
|
||||
|
||||
// campaignArguments builds one campaign invocation. The arm is the
|
||||
// implementation's name, so every run in the tree can be attributed to the
|
||||
// implementation it came from without reading back which port served it.
|
||||
func campaignArguments(
|
||||
configuration config,
|
||||
target implementation,
|
||||
seed string,
|
||||
) []string {
|
||||
arguments := []string{
|
||||
"--spec", configuration.specPath,
|
||||
"--bundle-id", target.URL(),
|
||||
"--platform", platform,
|
||||
"--arm", target.Name,
|
||||
"--generator", generator,
|
||||
"--max-steps", strconv.Itoa(configuration.maxSteps),
|
||||
"--duration", configuration.duration.String(),
|
||||
"--seeds", seed,
|
||||
"--sanderling", configuration.sanderlingPath,
|
||||
"--output", campaignDirectory(configuration, target, seed),
|
||||
}
|
||||
if len(configuration.extraArguments) > 0 {
|
||||
arguments = append(arguments, "--")
|
||||
arguments = append(arguments, configuration.extraArguments...)
|
||||
}
|
||||
return arguments
|
||||
}
|
||||
|
||||
func run(arguments []string, stdout, stderr io.Writer) error {
|
||||
configuration, err := parseArguments(arguments, stderr)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ctx, cancel := signal.NotifyContext(
|
||||
context.Background(),
|
||||
os.Interrupt,
|
||||
syscall.SIGTERM,
|
||||
)
|
||||
defer cancel()
|
||||
return runSweep(ctx, configuration, stdout)
|
||||
}
|
||||
|
||||
func main() {
|
||||
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(os.Stderr, "error: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,34 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"io"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Three flags missing is one rerun, not three: the operator is told about all
|
||||
// of them at once, in flag order, whatever order the check happened to walk.
|
||||
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
|
||||
_, err := parseArguments(
|
||||
[]string{"--spec", "s", "--max-steps", "10"},
|
||||
io.Discard,
|
||||
)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want every missing flag named")
|
||||
}
|
||||
message := err.Error()
|
||||
previous := -1
|
||||
for _, name := range []string{"--corpus", "--seeds", "--output"} {
|
||||
at := strings.Index(message, name)
|
||||
if at < 0 {
|
||||
t.Fatalf("got %q, want %s named", message, name)
|
||||
}
|
||||
if at < previous {
|
||||
t.Errorf("got %q, want the flags named in flag order", message)
|
||||
}
|
||||
previous = at
|
||||
}
|
||||
if strings.Contains(message, "--spec") {
|
||||
t.Errorf("got %q, want the supplied --spec left out", message)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,132 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
const (
|
||||
manifestFileName = "sweep.json"
|
||||
recordsFileName = "implementations.jsonl"
|
||||
)
|
||||
|
||||
// plannedImplementation is one implementation the sweep intends to run, with
|
||||
// the origin that keeps its stored state its own and the URL every seed is
|
||||
// driven at.
|
||||
type plannedImplementation struct {
|
||||
Name string `json:"name"`
|
||||
Document string `json:"document"`
|
||||
Port int `json:"port"`
|
||||
Origin string `json:"origin"`
|
||||
URL string `json:"url"`
|
||||
}
|
||||
|
||||
// manifest is sweep.json: what the sweep INTENDED to run, written before the
|
||||
// first run so a host that dropped an implementation or a seed shows up as a
|
||||
// missing run rather than as a smaller sample.
|
||||
type manifest struct {
|
||||
Generator string `json:"generator"`
|
||||
Platform string `json:"platform"`
|
||||
SpecPath string `json:"spec_path"`
|
||||
CorpusRoot string `json:"corpus_root"`
|
||||
CorpusCommit string `json:"corpus_commit"`
|
||||
MaxSteps int `json:"max_steps"`
|
||||
DurationMillis int64 `json:"duration_millis"`
|
||||
Seeds []int64 `json:"seeds"`
|
||||
Implementations []plannedImplementation `json:"implementations"`
|
||||
Concurrency int `json:"concurrency"`
|
||||
Host string `json:"host"`
|
||||
CampaignPath string `json:"campaign_path"`
|
||||
SanderlingPath string `json:"sanderling_path"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
}
|
||||
|
||||
func buildManifest(
|
||||
configuration config,
|
||||
implementations []implementation,
|
||||
host string,
|
||||
startedAt time.Time,
|
||||
) manifest {
|
||||
planned := make([]plannedImplementation, 0, len(implementations))
|
||||
for _, target := range implementations {
|
||||
planned = append(planned, plannedImplementation{
|
||||
Name: target.Name,
|
||||
Document: target.Document,
|
||||
Port: target.Port,
|
||||
Origin: target.Origin(),
|
||||
URL: target.URL(),
|
||||
})
|
||||
}
|
||||
return manifest{
|
||||
Generator: generator,
|
||||
Platform: platform,
|
||||
SpecPath: configuration.specPath,
|
||||
CorpusRoot: configuration.corpusRoot,
|
||||
CorpusCommit: corpusCommit(configuration.corpusRoot),
|
||||
MaxSteps: configuration.maxSteps,
|
||||
DurationMillis: configuration.duration.Milliseconds(),
|
||||
Seeds: configuration.seeds,
|
||||
Implementations: planned,
|
||||
Concurrency: configuration.concurrency,
|
||||
Host: host,
|
||||
CampaignPath: configuration.campaignPath,
|
||||
SanderlingPath: configuration.sanderlingPath,
|
||||
StartedAt: startedAt,
|
||||
}
|
||||
}
|
||||
|
||||
// corpusCommit records which checkout was swept. It is empty rather than fatal
|
||||
// for a corpus that is not a git working tree, because the population check has
|
||||
// already established the corpus holds the right examples.
|
||||
func corpusCommit(corpusRoot string) string {
|
||||
command := exec.Command("git", "-C", corpusRoot, "rev-parse", "HEAD")
|
||||
output, err := command.Output()
|
||||
if err != nil {
|
||||
return ""
|
||||
}
|
||||
return strings.TrimSpace(string(output))
|
||||
}
|
||||
|
||||
func writeManifest(directory string, value manifest) error {
|
||||
body, err := json.MarshalIndent(value, "", " ")
|
||||
if err != nil {
|
||||
return fmt.Errorf("marshal manifest: %w", err)
|
||||
}
|
||||
return os.WriteFile(
|
||||
filepath.Join(directory, manifestFileName),
|
||||
append(body, '\n'),
|
||||
0o644,
|
||||
)
|
||||
}
|
||||
|
||||
// runRecord is one campaign, which is one implementation at one seed.
|
||||
type runRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
ExitCode int `json:"exit_code"`
|
||||
LaunchError string `json:"launch_error,omitempty"`
|
||||
CampaignDirectory string `json:"campaign_directory"`
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
}
|
||||
|
||||
// implementationRecord is one line of implementations.jsonl. FailedStage names
|
||||
// the step that stopped this implementation, and one that never got served
|
||||
// carries no runs at all.
|
||||
type implementationRecord struct {
|
||||
Name string `json:"implementation"`
|
||||
Document string `json:"document"`
|
||||
Port int `json:"port"`
|
||||
Origin string `json:"origin"`
|
||||
URL string `json:"url"`
|
||||
FailedStage string `json:"failed_stage,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
Runs []runRecord `json:"runs"`
|
||||
}
|
||||
|
||||
const stageServe = "serve"
|
||||
@@ -0,0 +1,194 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"io"
|
||||
"net/url"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// collisionPairs are implementations in this corpus that write the same
|
||||
// localStorage key as each other. Served from one origin they share one stored
|
||||
// record: in the corpus survey angular2_es2015 crashed at bootstrap on a record
|
||||
// angular2 had written, which is a violation belonging to no implementation.
|
||||
var collisionPairs = [][2]string{
|
||||
{"angular2", "angular2_es2015"},
|
||||
{"backbone", "backbone_require"},
|
||||
{"canjs", "canjs_require"},
|
||||
{"react", "typescript-react"},
|
||||
}
|
||||
|
||||
func originOf(t *testing.T, rawURL string) string {
|
||||
t.Helper()
|
||||
parsed, err := url.Parse(rawURL)
|
||||
if err != nil {
|
||||
t.Fatalf("parse %q: %v", rawURL, err)
|
||||
}
|
||||
return parsed.Scheme + "://" + parsed.Host
|
||||
}
|
||||
|
||||
// The whole population is swept so the assertion covers every implementation
|
||||
// rather than the four pairs already known to collide: the survey found those
|
||||
// four, and an unexamined fifth would be just as damaging.
|
||||
func TestSweep_GivesEveryImplementationAnOriginNoOtherImplementationShares(
|
||||
t *testing.T,
|
||||
) {
|
||||
root := t.TempDir()
|
||||
corpus := wholeCorpus(t)
|
||||
specPath := filepath.Join(root, "todo.ts")
|
||||
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
output := filepath.Join(root, "campaigns")
|
||||
campaignLog := filepath.Join(root, "campaign.log")
|
||||
fetchLog := filepath.Join(root, "fetch.log")
|
||||
testBinary, err := filepath.Abs(os.Args[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
campaignPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-campaign"),
|
||||
fmt.Sprintf(
|
||||
stubCampaign,
|
||||
filepath.Join(output, manifestFileName),
|
||||
campaignLog,
|
||||
fetchLog,
|
||||
testBinary,
|
||||
),
|
||||
)
|
||||
sanderlingPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-sanderling"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
)
|
||||
|
||||
err = run([]string{
|
||||
"--corpus", corpus,
|
||||
"--spec", specPath,
|
||||
"--seeds", "1",
|
||||
"--max-steps", "10",
|
||||
"--duration", "30s",
|
||||
"--concurrency", "8",
|
||||
"--base-port", fmt.Sprint(freePortRange(t, len(includedImplementations))),
|
||||
"--output", output,
|
||||
"--campaign", campaignPath,
|
||||
"--sanderling", sanderlingPath,
|
||||
}, io.Discard, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatalf("sweep: %v", err)
|
||||
}
|
||||
|
||||
// What the sweep declared it would serve each implementation from.
|
||||
recorded := readManifest(t, filepath.Join(output, manifestFileName))
|
||||
if len(recorded.Implementations) != len(includedImplementations) {
|
||||
t.Fatalf(
|
||||
"intended implementations: got %d, want %d",
|
||||
len(recorded.Implementations),
|
||||
len(includedImplementations),
|
||||
)
|
||||
}
|
||||
plannedOrigin := map[string]string{}
|
||||
ownerOfOrigin := map[string]string{}
|
||||
for _, planned := range recorded.Implementations {
|
||||
if owner, taken := ownerOfOrigin[planned.Origin]; taken {
|
||||
t.Errorf(
|
||||
"%s and %s are both served from %s, so they share one localStorage",
|
||||
owner,
|
||||
planned.Name,
|
||||
planned.Origin,
|
||||
)
|
||||
}
|
||||
ownerOfOrigin[planned.Origin] = planned.Name
|
||||
plannedOrigin[planned.Name] = planned.Origin
|
||||
if got := originOf(t, planned.URL); got != planned.Origin {
|
||||
t.Errorf(
|
||||
"%s: url %q is not under the origin %q the manifest claims",
|
||||
planned.Name,
|
||||
planned.URL,
|
||||
planned.Origin,
|
||||
)
|
||||
}
|
||||
}
|
||||
for _, pair := range collisionPairs {
|
||||
if plannedOrigin[pair[0]] == plannedOrigin[pair[1]] {
|
||||
t.Errorf(
|
||||
"%s and %s write the same localStorage key and are both served from %s",
|
||||
pair[0],
|
||||
pair[1],
|
||||
plannedOrigin[pair[0]],
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// What actually reached the driver. The web driver parses the origin to
|
||||
// clear out of --bundle-id, so two arms sharing a bundle-id origin clear
|
||||
// and repopulate one another's storage however carefully they are labelled.
|
||||
drivenOrigin := map[string]string{}
|
||||
for _, line := range readLines(t, campaignLog) {
|
||||
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
|
||||
arm := argumentValue(arguments, "--arm")
|
||||
if arm == "" {
|
||||
t.Fatalf(
|
||||
"a campaign ran with no arm, so its runs cannot be attributed: %q",
|
||||
line,
|
||||
)
|
||||
}
|
||||
drivenOrigin[arm] = originOf(t, argumentValue(arguments, "--bundle-id"))
|
||||
}
|
||||
if len(drivenOrigin) != len(includedImplementations) {
|
||||
t.Fatalf(
|
||||
"arms that reached the campaign tool: got %d, want %d",
|
||||
len(drivenOrigin),
|
||||
len(includedImplementations),
|
||||
)
|
||||
}
|
||||
armOfOrigin := map[string]string{}
|
||||
for arm, origin := range drivenOrigin {
|
||||
if other, taken := armOfOrigin[origin]; taken {
|
||||
t.Errorf(
|
||||
"arms %s and %s were both driven at %s",
|
||||
other,
|
||||
arm,
|
||||
origin,
|
||||
)
|
||||
}
|
||||
armOfOrigin[origin] = arm
|
||||
if origin != plannedOrigin[arm] {
|
||||
t.Errorf(
|
||||
"%s was driven at %s but the manifest promised %s",
|
||||
arm,
|
||||
origin,
|
||||
plannedOrigin[arm],
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// And what each origin answered with, which is the check that the ports
|
||||
// are not merely distinct but each carries its own implementation.
|
||||
served := map[string]string{}
|
||||
for _, line := range readLines(t, fetchLog) {
|
||||
requested, body, found := strings.Cut(line, " -> 200 ")
|
||||
if !found {
|
||||
t.Errorf("a served page did not answer: %q", line)
|
||||
continue
|
||||
}
|
||||
served[originOf(t, requested)] = body
|
||||
}
|
||||
for arm, origin := range drivenOrigin {
|
||||
if want := "<title>" + arm + "</title>"; !strings.Contains(
|
||||
served[origin],
|
||||
want,
|
||||
) {
|
||||
t.Errorf(
|
||||
"%s at %s was served %q, which is not its own document",
|
||||
arm,
|
||||
origin,
|
||||
served[origin],
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,178 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"io/fs"
|
||||
"net"
|
||||
"net/http"
|
||||
"path"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
const (
|
||||
readinessTimeout = 5 * time.Second
|
||||
readinessInterval = 100 * time.Millisecond
|
||||
)
|
||||
|
||||
// stubScriptPath is served from every implementation's own origin in place of a
|
||||
// dependency that no longer answers.
|
||||
const stubScriptPath = "/__corpus-sweep__/stub.js"
|
||||
|
||||
// deadDependency is polyfill.io, which was shut down after this corpus was
|
||||
// pinned. One implementation loads it from its index.html; the request fails
|
||||
// and the application works regardless, but it fails slowly, once per run, on
|
||||
// every run. The corpus is served from this process, so the reference is
|
||||
// rewritten to a local no-op on the way out.
|
||||
var deadDependency = regexp.MustCompile(`https?://polyfill\.io/[^"'\s>]*`)
|
||||
|
||||
// staticServer serves the whole corpus tree on one implementation's port. The
|
||||
// tree rather than the implementation's own directory, because an example that
|
||||
// asks for a path above itself has to resolve the same way it would in the
|
||||
// upstream repository. Only one implementation is ever visited on this port, so
|
||||
// the origin still belongs to it alone.
|
||||
type staticServer struct {
|
||||
implementation implementation
|
||||
listener net.Listener
|
||||
server *http.Server
|
||||
}
|
||||
|
||||
func startStaticServer(
|
||||
corpusRoot string,
|
||||
target implementation,
|
||||
) (*staticServer, error) {
|
||||
listener, err := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", target.Port))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("serve %s: %w", target.Name, err)
|
||||
}
|
||||
server := &http.Server{Handler: corpusHandler(corpusRoot)}
|
||||
running := &staticServer{
|
||||
implementation: target,
|
||||
listener: listener,
|
||||
server: server,
|
||||
}
|
||||
go server.Serve(listener)
|
||||
return running, nil
|
||||
}
|
||||
|
||||
func (s *staticServer) stop() {
|
||||
s.server.Close()
|
||||
}
|
||||
|
||||
// waitReady confirms the served document answers before any run drives it. A
|
||||
// wrong document path would otherwise reach the driver as a 404 page, and a
|
||||
// sweep of 404 pages produces clean runs for every implementation.
|
||||
//
|
||||
// Only a transport error is retried. The listener is bound before the sweep
|
||||
// starts, so a status that is not 200 is the server's answer about this
|
||||
// document and waiting will not change it.
|
||||
func (s *staticServer) waitReady(ctx context.Context) error {
|
||||
url := s.implementation.URL()
|
||||
client := &http.Client{Timeout: 5 * time.Second}
|
||||
deadline := time.Now().Add(readinessTimeout)
|
||||
var lastErr error
|
||||
for {
|
||||
if ctx.Err() != nil {
|
||||
return ctx.Err()
|
||||
}
|
||||
response, err := client.Get(url)
|
||||
if err == nil {
|
||||
io.Copy(io.Discard, response.Body)
|
||||
response.Body.Close()
|
||||
if response.StatusCode == http.StatusOK {
|
||||
return nil
|
||||
}
|
||||
return fmt.Errorf(
|
||||
"%s answered %d, not the document",
|
||||
url,
|
||||
response.StatusCode,
|
||||
)
|
||||
}
|
||||
lastErr = err
|
||||
if time.Now().After(deadline) {
|
||||
return fmt.Errorf(
|
||||
"%s did not answer within %s: %w",
|
||||
url,
|
||||
readinessTimeout,
|
||||
lastErr,
|
||||
)
|
||||
}
|
||||
time.Sleep(readinessInterval)
|
||||
}
|
||||
}
|
||||
|
||||
func corpusHandler(corpusRoot string) http.Handler {
|
||||
files := http.FileServer(http.Dir(corpusRoot))
|
||||
return http.HandlerFunc(
|
||||
func(writer http.ResponseWriter, request *http.Request) {
|
||||
cleaned := path.Clean(
|
||||
"/" + strings.TrimPrefix(request.URL.Path, "/"),
|
||||
)
|
||||
if cleaned == stubScriptPath {
|
||||
writer.Header().Set("Content-Type", "application/javascript")
|
||||
io.WriteString(
|
||||
writer,
|
||||
"/* corpus-sweep: dependency removed upstream */\n",
|
||||
)
|
||||
return
|
||||
}
|
||||
if !strings.HasSuffix(cleaned, ".html") {
|
||||
files.ServeHTTP(writer, request)
|
||||
return
|
||||
}
|
||||
body, modified, err := readDocument(corpusRoot, cleaned)
|
||||
if err != nil {
|
||||
// Never the file server's fallback: it answers a document it
|
||||
// cannot read with a 200 directory listing, and a run driven at a
|
||||
// listing explores nothing and comes back clean.
|
||||
http.Error(writer, err.Error(), documentStatus(err))
|
||||
return
|
||||
}
|
||||
rewritten := deadDependency.ReplaceAll(body, []byte(stubScriptPath))
|
||||
http.ServeContent(
|
||||
writer,
|
||||
request,
|
||||
path.Base(cleaned),
|
||||
modified,
|
||||
bytes.NewReader(rewritten),
|
||||
)
|
||||
},
|
||||
)
|
||||
}
|
||||
|
||||
func documentStatus(err error) int {
|
||||
switch {
|
||||
case errors.Is(err, fs.ErrNotExist):
|
||||
return http.StatusNotFound
|
||||
case errors.Is(err, fs.ErrPermission):
|
||||
return http.StatusForbidden
|
||||
default:
|
||||
return http.StatusInternalServerError
|
||||
}
|
||||
}
|
||||
|
||||
func readDocument(corpusRoot, cleaned string) ([]byte, time.Time, error) {
|
||||
file, err := http.Dir(corpusRoot).Open(cleaned)
|
||||
if err != nil {
|
||||
return nil, time.Time{}, err
|
||||
}
|
||||
defer file.Close()
|
||||
info, err := file.Stat()
|
||||
if err != nil || info.IsDir() {
|
||||
return nil, time.Time{}, fmt.Errorf(
|
||||
"%s is not a document",
|
||||
filepath.FromSlash(cleaned),
|
||||
)
|
||||
}
|
||||
body, err := io.ReadAll(file)
|
||||
if err != nil {
|
||||
return nil, time.Time{}, err
|
||||
}
|
||||
return body, info.ModTime(), nil
|
||||
}
|
||||
@@ -0,0 +1,214 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func get(t *testing.T, url string) (int, string) {
|
||||
t.Helper()
|
||||
client := &http.Client{Timeout: 5 * time.Second}
|
||||
response, err := client.Get(url)
|
||||
if err != nil {
|
||||
t.Fatalf("GET %s: %v", url, err)
|
||||
}
|
||||
defer response.Body.Close()
|
||||
body, err := io.ReadAll(response.Body)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return response.StatusCode, string(body)
|
||||
}
|
||||
|
||||
func TestStaticServer_ServesEachImplementationsDocumentFromItsOwnPort(
|
||||
t *testing.T,
|
||||
) {
|
||||
root := wholeCorpus(t)
|
||||
planned, err := planImplementations(
|
||||
root,
|
||||
[]string{"angular-dart", "duel", "react"},
|
||||
freePortRange(t, 3),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, target := range planned {
|
||||
server, err := startStaticServer(root, target)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer server.stop()
|
||||
if err := server.waitReady(context.Background()); err != nil {
|
||||
t.Fatalf("%s: %v", target.Name, err)
|
||||
}
|
||||
status, body := get(t, target.URL())
|
||||
if status != http.StatusOK ||
|
||||
!strings.Contains(body, "<title>"+target.Name+"</title>") {
|
||||
t.Errorf(
|
||||
"%s at %s: got %d %q",
|
||||
target.Name,
|
||||
target.URL(),
|
||||
status,
|
||||
body,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestStaticServer_ReplacesTheDependencyThatNoLongerAnswers(t *testing.T) {
|
||||
root := wholeCorpus(t)
|
||||
document := filepath.Join(root, filepath.FromSlash(documentFor("aurelia")))
|
||||
body := `<!doctype html><script src="https://polyfill.io/v3/polyfill.min.js?features=Promise"></script>`
|
||||
if err := os.WriteFile(document, []byte(body), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
planned, err := planImplementations(
|
||||
root,
|
||||
[]string{"aurelia"},
|
||||
freePortRange(t, 1),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
server, err := startStaticServer(root, planned[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer server.stop()
|
||||
|
||||
status, served := get(t, planned[0].URL())
|
||||
if status != http.StatusOK {
|
||||
t.Fatalf("document answered %d", status)
|
||||
}
|
||||
if strings.Contains(served, "polyfill.io") {
|
||||
t.Errorf(
|
||||
"a request to a host that no longer answers reaches the browser once per run: %q",
|
||||
served,
|
||||
)
|
||||
}
|
||||
if !strings.Contains(served, `src="`+stubScriptPath+`"`) {
|
||||
t.Errorf(
|
||||
"the reference was removed rather than pointed at a local no-op: %q",
|
||||
served,
|
||||
)
|
||||
}
|
||||
|
||||
stubStatus, stub := get(t, planned[0].Origin()+stubScriptPath)
|
||||
if stubStatus != http.StatusOK || stub == "" {
|
||||
t.Errorf("stub script: got %d %q", stubStatus, stub)
|
||||
}
|
||||
}
|
||||
|
||||
// An example that asks for a path above its own directory has to resolve the
|
||||
// same way it does upstream, which is why the whole tree is served rather than
|
||||
// the one directory.
|
||||
func TestStaticServer_ServesPathsAboveTheImplementationsOwnDirectory(
|
||||
t *testing.T,
|
||||
) {
|
||||
root := wholeCorpus(t)
|
||||
shared := filepath.Join(root, "node_modules", "todomvc-common")
|
||||
if err := os.MkdirAll(shared, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(shared, "base.css"), []byte("body{}"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
planned, err := planImplementations(
|
||||
root,
|
||||
[]string{"react"},
|
||||
freePortRange(t, 1),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
server, err := startStaticServer(root, planned[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer server.stop()
|
||||
|
||||
status, body := get(
|
||||
t,
|
||||
planned[0].Origin()+"/node_modules/todomvc-common/base.css",
|
||||
)
|
||||
if status != http.StatusOK || body != "body{}" {
|
||||
t.Errorf("shared asset: got %d %q", status, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestStaticServerWaitReady_ReportsADocumentThatDoesNotAnswer(t *testing.T) {
|
||||
root := wholeCorpus(t)
|
||||
document := filepath.Join(root, filepath.FromSlash(documentFor("dojo")))
|
||||
if err := os.Chmod(document, 0o000); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(func() { os.Chmod(document, 0o644) })
|
||||
planned, err := planImplementations(
|
||||
root,
|
||||
[]string{"dojo"},
|
||||
freePortRange(t, 1),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
server, err := startStaticServer(root, planned[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer server.stop()
|
||||
|
||||
started := time.Now()
|
||||
err = server.waitReady(context.Background())
|
||||
if err == nil {
|
||||
t.Fatal(
|
||||
"a document that does not answer would be swept as an error page and come back clean",
|
||||
)
|
||||
}
|
||||
if elapsed := time.Since(started); elapsed > readinessTimeout {
|
||||
t.Errorf(
|
||||
"waited %s for a status that was never going to change",
|
||||
elapsed,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestStaticServerStop_ReleasesThePort(t *testing.T) {
|
||||
root := wholeCorpus(t)
|
||||
planned, err := planImplementations(
|
||||
root,
|
||||
[]string{"vue"},
|
||||
freePortRange(t, 1),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
server, err := startStaticServer(root, planned[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := server.waitReady(context.Background()); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
server.stop()
|
||||
|
||||
client := &http.Client{Timeout: time.Second}
|
||||
deadline := time.Now().Add(3 * time.Second)
|
||||
for time.Now().Before(deadline) {
|
||||
response, err := client.Get(planned[0].URL())
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
response.Body.Close()
|
||||
time.Sleep(50 * time.Millisecond)
|
||||
}
|
||||
t.Fatalf(
|
||||
"port %d is still served after stop(), so the next sweep cannot bind it",
|
||||
planned[0].Port,
|
||||
)
|
||||
}
|
||||
@@ -0,0 +1,315 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
// resolveBinaries turns campaign and sanderling into absolute paths before
|
||||
// anything is served. Each campaign runs from the sweep's own directory, and a
|
||||
// binary that is missing altogether has to stop the sweep here rather than fail
|
||||
// once per implementation and seed. Every one that is missing is named
|
||||
// together, in flag order: stopping at the first turns that single stop into
|
||||
// one rerun per missing binary.
|
||||
func resolveBinaries(configuration *config) error {
|
||||
var missing []error
|
||||
for _, binary := range []struct {
|
||||
name string
|
||||
value *string
|
||||
}{
|
||||
{"--campaign", &configuration.campaignPath},
|
||||
{"--sanderling", &configuration.sanderlingPath},
|
||||
} {
|
||||
resolved, err := exec.LookPath(*binary.value)
|
||||
if err != nil {
|
||||
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
|
||||
continue
|
||||
}
|
||||
absolute, err := filepath.Abs(resolved)
|
||||
if err != nil {
|
||||
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
|
||||
continue
|
||||
}
|
||||
*binary.value = absolute
|
||||
}
|
||||
return errors.Join(missing...)
|
||||
}
|
||||
|
||||
type sweep struct {
|
||||
configuration config
|
||||
stdout io.Writer
|
||||
records io.Writer
|
||||
servers map[string]*staticServer
|
||||
mutex sync.Mutex
|
||||
stalled int
|
||||
failedRuns int
|
||||
totalRuns int
|
||||
}
|
||||
|
||||
func runSweep(
|
||||
ctx context.Context,
|
||||
configuration config,
|
||||
stdout io.Writer,
|
||||
) error {
|
||||
if _, err := os.Stat(filepath.Join(configuration.outputDirectory, manifestFileName)); err == nil {
|
||||
return fmt.Errorf(
|
||||
"%s already exists in %s: pick a fresh --output so two sweeps do not share a directory",
|
||||
manifestFileName,
|
||||
configuration.outputDirectory,
|
||||
)
|
||||
}
|
||||
if err := verifyCorpus(configuration.corpusRoot); err != nil {
|
||||
return err
|
||||
}
|
||||
implementations, err := planImplementations(
|
||||
configuration.corpusRoot,
|
||||
configuration.implementations,
|
||||
configuration.basePort,
|
||||
)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := resolveBinaries(&configuration); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := os.Stat(configuration.specPath); err != nil {
|
||||
return fmt.Errorf("--spec: %w", err)
|
||||
}
|
||||
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
|
||||
return fmt.Errorf("create sweep dir: %w", err)
|
||||
}
|
||||
host, _ := os.Hostname()
|
||||
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, implementations, host, time.Now().UTC())); err != nil {
|
||||
return fmt.Errorf("write %s: %w", manifestFileName, err)
|
||||
}
|
||||
|
||||
recordsFile, err := os.OpenFile(
|
||||
filepath.Join(configuration.outputDirectory, recordsFileName),
|
||||
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
|
||||
0o644,
|
||||
)
|
||||
if err != nil {
|
||||
return fmt.Errorf("open %s: %w", recordsFileName, err)
|
||||
}
|
||||
defer recordsFile.Close()
|
||||
|
||||
running := &sweep{
|
||||
configuration: configuration,
|
||||
stdout: stdout,
|
||||
records: recordsFile,
|
||||
}
|
||||
// Every port is bound before any run starts. A port already taken means
|
||||
// that implementation cannot have an origin of its own, and continuing
|
||||
// without one is what the separate origins are there to prevent.
|
||||
if err := running.serveAll(implementations); err != nil {
|
||||
return err
|
||||
}
|
||||
defer running.stopAll()
|
||||
|
||||
fmt.Fprintf(
|
||||
stdout,
|
||||
"sweep: %d implementations, %d seeds each, %d at a time, %s\n",
|
||||
len(
|
||||
implementations,
|
||||
),
|
||||
len(configuration.seeds),
|
||||
configuration.concurrency,
|
||||
configuration.outputDirectory,
|
||||
)
|
||||
running.work(ctx, implementations)
|
||||
|
||||
fmt.Fprintf(
|
||||
stdout,
|
||||
"sweep complete: %d of %d implementations never ran, %d of %d campaigns failed\n",
|
||||
running.stalled,
|
||||
len(implementations),
|
||||
running.failedRuns,
|
||||
running.totalRuns,
|
||||
)
|
||||
if running.stalled > 0 || running.failedRuns > 0 {
|
||||
return fmt.Errorf(
|
||||
"%d of %d implementations never ran and %d of %d campaigns failed; see %s",
|
||||
running.stalled,
|
||||
len(implementations),
|
||||
running.failedRuns,
|
||||
running.totalRuns,
|
||||
recordsFileName,
|
||||
)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (s *sweep) serveAll(implementations []implementation) error {
|
||||
s.servers = make(map[string]*staticServer, len(implementations))
|
||||
for _, target := range implementations {
|
||||
server, err := startStaticServer(s.configuration.corpusRoot, target)
|
||||
if err != nil {
|
||||
s.stopAll()
|
||||
return err
|
||||
}
|
||||
s.servers[target.Name] = server
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (s *sweep) stopAll() {
|
||||
for _, server := range s.servers {
|
||||
server.stop()
|
||||
}
|
||||
}
|
||||
|
||||
func (s *sweep) work(ctx context.Context, implementations []implementation) {
|
||||
queue := make(chan implementation, len(implementations))
|
||||
for _, target := range implementations {
|
||||
queue <- target
|
||||
}
|
||||
close(queue)
|
||||
|
||||
workers := min(s.configuration.concurrency, len(implementations))
|
||||
var waitGroup sync.WaitGroup
|
||||
for range workers {
|
||||
waitGroup.Add(1)
|
||||
go func() {
|
||||
defer waitGroup.Done()
|
||||
for target := range queue {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
s.report(s.runImplementation(ctx, target))
|
||||
}
|
||||
}()
|
||||
}
|
||||
waitGroup.Wait()
|
||||
}
|
||||
|
||||
// runImplementation carries one implementation through all of its seeds. A
|
||||
// document that does not answer is recorded and the sweep moves on: one
|
||||
// implementation must not cost the other forty-two their runs.
|
||||
func (s *sweep) runImplementation(
|
||||
ctx context.Context,
|
||||
target implementation,
|
||||
) (record implementationRecord) {
|
||||
record = implementationRecord{
|
||||
Name: target.Name,
|
||||
Document: target.Document,
|
||||
Port: target.Port,
|
||||
Origin: target.Origin(),
|
||||
URL: target.URL(),
|
||||
StartedAt: time.Now().UTC(),
|
||||
}
|
||||
started := time.Now()
|
||||
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
|
||||
|
||||
if err := os.MkdirAll(filepath.Join(s.configuration.outputDirectory, target.Name), 0o755); err != nil {
|
||||
record.FailedStage = stageServe
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
if err := s.servers[target.Name].waitReady(ctx); err != nil {
|
||||
record.FailedStage = stageServe
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
|
||||
for _, seed := range s.configuration.seeds {
|
||||
if ctx.Err() != nil {
|
||||
return record
|
||||
}
|
||||
record.Runs = append(record.Runs, s.runSeed(ctx, target, seed))
|
||||
}
|
||||
return record
|
||||
}
|
||||
|
||||
func (s *sweep) runSeed(
|
||||
ctx context.Context,
|
||||
target implementation,
|
||||
seed int64,
|
||||
) (record runRecord) {
|
||||
seedText := strconv.FormatInt(seed, 10)
|
||||
directory := campaignDirectory(s.configuration, target, seedText)
|
||||
record = runRecord{Seed: seed, CampaignDirectory: directory}
|
||||
started := time.Now()
|
||||
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
|
||||
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
record.ExitCode = -1
|
||||
record.LaunchError = err.Error()
|
||||
return record
|
||||
}
|
||||
exitCode, err := runCommand(
|
||||
ctx,
|
||||
s.configuration.campaignPath,
|
||||
campaignArguments(
|
||||
s.configuration,
|
||||
target,
|
||||
seedText,
|
||||
),
|
||||
filepath.Join(directory, "campaign.log"),
|
||||
)
|
||||
record.ExitCode = exitCode
|
||||
if err != nil {
|
||||
record.LaunchError = err.Error()
|
||||
}
|
||||
return record
|
||||
}
|
||||
|
||||
func runCommand(
|
||||
ctx context.Context,
|
||||
binary string,
|
||||
arguments []string,
|
||||
logPath string,
|
||||
) (int, error) {
|
||||
logFile, err := os.Create(logPath)
|
||||
if err != nil {
|
||||
return -1, err
|
||||
}
|
||||
defer logFile.Close()
|
||||
command := exec.CommandContext(ctx, binary, arguments...)
|
||||
command.Stdout = logFile
|
||||
command.Stderr = logFile
|
||||
err = command.Run()
|
||||
if err == nil {
|
||||
return 0, nil
|
||||
}
|
||||
var exitError *exec.ExitError
|
||||
if errors.As(err, &exitError) {
|
||||
return exitError.ExitCode(), nil
|
||||
}
|
||||
return -1, err
|
||||
}
|
||||
|
||||
func (s *sweep) report(record implementationRecord) {
|
||||
s.mutex.Lock()
|
||||
defer s.mutex.Unlock()
|
||||
if record.FailedStage != "" {
|
||||
s.stalled++
|
||||
}
|
||||
s.totalRuns += len(record.Runs)
|
||||
failed := 0
|
||||
for _, run := range record.Runs {
|
||||
if run.ExitCode != 0 {
|
||||
failed++
|
||||
}
|
||||
}
|
||||
s.failedRuns += failed
|
||||
if err := json.NewEncoder(s.records).Encode(record); err != nil {
|
||||
fmt.Fprintf(s.stdout, "warning: %s record: %v\n", record.Name, err)
|
||||
}
|
||||
elapsed := time.Duration(record.MonotonicMillis) * time.Millisecond
|
||||
if record.FailedStage != "" {
|
||||
fmt.Fprintf(s.stdout, "%s port=%d failed at %s: %s (%s)\n",
|
||||
record.Name, record.Port, record.FailedStage, record.Error, elapsed)
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(s.stdout, "%s origin=%s campaigns=%d failed=%d elapsed=%s\n",
|
||||
record.Name, record.Origin, len(record.Runs), failed, elapsed)
|
||||
}
|
||||
@@ -0,0 +1,46 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// Two binaries missing is one rerun, not two: the operator is told about both
|
||||
// at once, in flag order, whatever order the check happened to walk.
|
||||
func TestResolveBinaries_NamesEveryMissingBinaryInFlagOrder(t *testing.T) {
|
||||
configuration := config{
|
||||
campaignPath: "campaign-that-is-not-installed",
|
||||
sanderlingPath: "sanderling-that-is-not-installed",
|
||||
}
|
||||
|
||||
err := resolveBinaries(&configuration)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want both missing binaries named")
|
||||
}
|
||||
message := err.Error()
|
||||
campaign := strings.Index(message, "--campaign")
|
||||
sanderling := strings.Index(message, "--sanderling")
|
||||
if campaign < 0 || sanderling < 0 {
|
||||
t.Fatalf("got %q, want both --campaign and --sanderling named", message)
|
||||
}
|
||||
if campaign > sanderling {
|
||||
t.Errorf("got %q, want --campaign named before --sanderling", message)
|
||||
}
|
||||
|
||||
resolved := config{
|
||||
campaignPath: writeScript(
|
||||
t,
|
||||
filepath.Join(t.TempDir(), "stub-campaign"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
),
|
||||
sanderlingPath: "sanderling-that-is-not-installed",
|
||||
}
|
||||
err = resolveBinaries(&resolved)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want the missing sanderling named")
|
||||
}
|
||||
if strings.Contains(err.Error(), "--campaign") {
|
||||
t.Errorf("got %q, want the campaign that resolved left out", err.Error())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,219 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"sort"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
// Instance is one defect as the draft identifies it across runs: the property
|
||||
// that reported, the action attributed as the origin of the failed obligation,
|
||||
// and the screen the witness observed. Two reports sharing all three are the
|
||||
// same defect seen twice; a run-level count of violated properties cannot say
|
||||
// that.
|
||||
type Instance struct {
|
||||
Property string `json:"property"`
|
||||
OriginAction string `json:"origin_action"`
|
||||
WitnessScreen string `json:"witness_screen"`
|
||||
Runs []string `json:"runs"`
|
||||
Seeds []int64 `json:"seeds"`
|
||||
// Reports counts violations folded into this instance, which exceeds the
|
||||
// run count only if one run reported the same property twice, and the
|
||||
// latch says it cannot.
|
||||
Reports int `json:"reports"`
|
||||
// RedactedOrigin marks a row whose origin action reached the trace with its
|
||||
// typed value redacted, so `full` keyed it by selector instead. Two runs
|
||||
// that typed different values into that field are one row here, which makes
|
||||
// the count of such rows a floor rather than a total.
|
||||
RedactedOrigin bool `json:"redacted_origin,omitempty"`
|
||||
}
|
||||
|
||||
// Unattributed is a violation that carries no origin, so the identity rule
|
||||
// cannot be applied to it. It is reported rather than counted, because
|
||||
// dropping it understates the defect count and guessing an origin invents one.
|
||||
type Unattributed struct {
|
||||
Property string `json:"property"`
|
||||
Run string `json:"run"`
|
||||
Step int `json:"step"`
|
||||
Reason string `json:"reason"`
|
||||
}
|
||||
|
||||
// actionKeyMode selects how much of an action two reports must share to be the
|
||||
// same origin. The draft names the origin action and does not say which of its
|
||||
// fields identify it, so the strict reading is available beside the default.
|
||||
type actionKeyMode string
|
||||
|
||||
const (
|
||||
// bySelector keys an action by what it did and the name of what it did it
|
||||
// to. Coordinates and generated text differ between two runs that took the
|
||||
// same action against the same control.
|
||||
bySelector actionKeyMode = "selector"
|
||||
// byFullAction adds the text typed and the coordinates dispatched.
|
||||
byFullAction actionKeyMode = "full"
|
||||
)
|
||||
|
||||
type identityKey struct {
|
||||
property string
|
||||
action string
|
||||
screen string
|
||||
}
|
||||
|
||||
// Corpus is what one invocation read: the instances, the violations it could
|
||||
// not attribute, and the counts the report needs.
|
||||
type Corpus struct {
|
||||
Runs int `json:"runs"`
|
||||
Instances []Instance `json:"instances"`
|
||||
Unattributed []Unattributed `json:"unattributed,omitempty"`
|
||||
UnnamedScreen int `json:"unnamed_witness_screens,omitempty"`
|
||||
}
|
||||
|
||||
// Singletons counts instances that appeared in exactly one run, the number the
|
||||
// evaluation reports beside the defect count.
|
||||
func (c Corpus) Singletons() int {
|
||||
count := 0
|
||||
for _, instance := range c.Instances {
|
||||
if len(instance.Runs) == 1 {
|
||||
count++
|
||||
}
|
||||
}
|
||||
return count
|
||||
}
|
||||
|
||||
// DegradedIdentities counts instances the action key could not be computed for
|
||||
// in full, which are the rows a reader has to treat as a lower bound.
|
||||
func (c Corpus) DegradedIdentities() int {
|
||||
count := 0
|
||||
for _, instance := range c.Instances {
|
||||
if instance.RedactedOrigin {
|
||||
count++
|
||||
}
|
||||
}
|
||||
return count
|
||||
}
|
||||
|
||||
func identify(runs []tracecorpus.Run, mode actionKeyMode) (Corpus, error) {
|
||||
corpus := Corpus{Runs: len(runs)}
|
||||
byKey := map[identityKey]*Instance{}
|
||||
var order []identityKey
|
||||
for _, run := range runs {
|
||||
steps := index(run.Steps)
|
||||
for _, step := range run.Steps {
|
||||
for _, property := range step.Violations {
|
||||
witness, ok := step.Witnesses[property]
|
||||
if !ok || witness.Step == 0 {
|
||||
corpus.Unattributed = append(corpus.Unattributed, Unattributed{
|
||||
Property: property,
|
||||
Run: run.Directory,
|
||||
Step: step.Index,
|
||||
Reason: attributionGap(ok),
|
||||
})
|
||||
continue
|
||||
}
|
||||
origin, ok := steps[witness.Step]
|
||||
if !ok {
|
||||
return Corpus{}, fmt.Errorf(
|
||||
"%s: %s names origin step %d, which the trace does not hold",
|
||||
run.Directory, property, witness.Step,
|
||||
)
|
||||
}
|
||||
detected, ok := steps[witness.DetectedStep]
|
||||
if !ok {
|
||||
return Corpus{}, fmt.Errorf(
|
||||
"%s: %s names detection step %d, which the trace does not hold",
|
||||
run.Directory, property, witness.DetectedStep,
|
||||
)
|
||||
}
|
||||
if detected.Screen == "" {
|
||||
corpus.UnnamedScreen++
|
||||
}
|
||||
action, redactedOrigin := actionKey(origin, mode)
|
||||
key := identityKey{
|
||||
property: property,
|
||||
action: action,
|
||||
screen: detected.Screen,
|
||||
}
|
||||
instance, seen := byKey[key]
|
||||
if !seen {
|
||||
instance = &Instance{
|
||||
Property: property,
|
||||
OriginAction: key.action,
|
||||
WitnessScreen: key.screen,
|
||||
RedactedOrigin: redactedOrigin,
|
||||
}
|
||||
byKey[key] = instance
|
||||
order = append(order, key)
|
||||
}
|
||||
instance.Reports++
|
||||
if len(instance.Runs) == 0 ||
|
||||
instance.Runs[len(instance.Runs)-1] != run.Directory {
|
||||
instance.Runs = append(instance.Runs, run.Directory)
|
||||
instance.Seeds = append(instance.Seeds, run.Meta.Seed)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for _, key := range order {
|
||||
corpus.Instances = append(corpus.Instances, *byKey[key])
|
||||
}
|
||||
sort.SliceStable(corpus.Instances, func(i, j int) bool {
|
||||
if corpus.Instances[i].Property != corpus.Instances[j].Property {
|
||||
return corpus.Instances[i].Property < corpus.Instances[j].Property
|
||||
}
|
||||
return corpus.Instances[i].OriginAction < corpus.Instances[j].OriginAction
|
||||
})
|
||||
return corpus, nil
|
||||
}
|
||||
|
||||
func attributionGap(hasWitness bool) string {
|
||||
if !hasWitness {
|
||||
return "no witness recorded"
|
||||
}
|
||||
return "witness records no origin step"
|
||||
}
|
||||
|
||||
func index(steps []trace.Step) map[int]trace.Step {
|
||||
byIndex := make(map[int]trace.Step, len(steps))
|
||||
for _, step := range steps {
|
||||
byIndex[step.Index] = step
|
||||
}
|
||||
return byIndex
|
||||
}
|
||||
|
||||
// actionKey renders the action the origin step chose, and reports whether the
|
||||
// key had to be degraded to the selector. The action recorded on a line is the
|
||||
// one applied after observing it, which is the alignment that makes an origin
|
||||
// index name an action at all.
|
||||
//
|
||||
// A typed value the record redacted is the same string for every value typed
|
||||
// into that field, so keying on it would merge distinct actions while reading
|
||||
// as a whole-action key. The key drops it and says it did, because an identity
|
||||
// that cannot be computed has to show as an undercount rather than as a count.
|
||||
func actionKey(origin trace.Step, mode actionKeyMode) (string, bool) {
|
||||
if origin.NextAction == nil {
|
||||
return "none", false
|
||||
}
|
||||
if origin.ActionSkipped != "" {
|
||||
return "none (" + origin.ActionSkipped + ")", false
|
||||
}
|
||||
action := *origin.NextAction
|
||||
key := action.Kind
|
||||
switch {
|
||||
case action.Selector != "":
|
||||
key += " " + action.Selector
|
||||
case action.Key != "":
|
||||
key += " " + action.Key
|
||||
case action.X != 0 || action.Y != 0 || action.ToX != 0 || action.ToY != 0:
|
||||
key += fmt.Sprintf(" (%d,%d)", action.X, action.Y)
|
||||
}
|
||||
if mode != byFullAction {
|
||||
return key, false
|
||||
}
|
||||
if action.Text == verifier.RedactedInputText {
|
||||
return key + " text=redacted", true
|
||||
}
|
||||
return fmt.Sprintf("%s text=%q at=(%d,%d)->(%d,%d)",
|
||||
key, action.Text, action.X, action.Y, action.ToX, action.ToY), false
|
||||
}
|
||||
@@ -0,0 +1,242 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
// TestOneDefectSeenTwiceIsOneInstance: two runs report the same property from
|
||||
// the same origin action on the same screen, which is one defect found twice
|
||||
// and two properties violated.
|
||||
func TestOneDefectSeenTwiceIsOneInstance(t *testing.T) {
|
||||
first := run(t, 3,
|
||||
step(1, "/ledger", tap("id:add-txn")),
|
||||
violating(2, "/accounts/7", tap("id:save"), "balanceMatches", 2, 2),
|
||||
)
|
||||
second := run(t, 5,
|
||||
step(1, "/ledger", tap("id:add-txn")),
|
||||
violating(2, "/accounts/7", tap("id:save"), "balanceMatches", 2, 2),
|
||||
)
|
||||
|
||||
corpus := identified(t, bySelector, first, second)
|
||||
if len(corpus.Instances) != 1 {
|
||||
t.Fatalf("instances = %d, want 1: %+v", len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
instance := corpus.Instances[0]
|
||||
if len(instance.Runs) != 2 || instance.Reports != 2 {
|
||||
t.Fatalf("instance = %+v, want two runs reporting it", instance)
|
||||
}
|
||||
if corpus.Singletons() != 0 {
|
||||
t.Fatalf("singletons = %d, want 0", corpus.Singletons())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTheSamePropertyFromTwoOriginsIsTwoDefects(t *testing.T) {
|
||||
first := run(t, 3, violating(1, "/ledger", tap("id:save"), "balanceMatches", 1, 1))
|
||||
second := run(t, 5, violating(1, "/ledger", tap("id:delete"), "balanceMatches", 1, 1))
|
||||
|
||||
corpus := identified(t, bySelector, first, second)
|
||||
if len(corpus.Instances) != 2 {
|
||||
t.Fatalf("instances = %d, want 2: %+v", len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
if corpus.Singletons() != 2 {
|
||||
t.Fatalf("singletons = %d, want 2", corpus.Singletons())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTheSamePropertyOnTwoScreensIsTwoDefects(t *testing.T) {
|
||||
first := run(t, 3, violating(1, "/ledger", tap("id:save"), "balanceMatches", 1, 1))
|
||||
second := run(t, 5, violating(1, "/home", tap("id:save"), "balanceMatches", 1, 1))
|
||||
|
||||
corpus := identified(t, bySelector, first, second)
|
||||
if len(corpus.Instances) != 2 {
|
||||
t.Fatalf("instances = %d, want 2: %+v", len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
}
|
||||
|
||||
// TestADeferredViolationTakesTheActionOnItsOriginLine holds the alignment the
|
||||
// draft states: the action recorded on line k is the one applied after
|
||||
// observing k, so an obligation armed at k is attributed to that action and
|
||||
// witnessed on the screen the detection step observed.
|
||||
func TestADeferredViolationTakesTheActionOnItsOriginLine(t *testing.T) {
|
||||
only := run(t, 3,
|
||||
step(1, "/login", tap("id:login-submit")),
|
||||
step(2, "/home", tap("id:open-ledger")),
|
||||
violating(3, "/ledger", tap("id:add-txn"), "landsOnLedger", 2, 3),
|
||||
)
|
||||
|
||||
corpus := identified(t, bySelector, only)
|
||||
instance := corpus.Instances[0]
|
||||
if instance.OriginAction != "Tap id:open-ledger" {
|
||||
t.Fatalf("origin action = %q, want the action on the origin line", instance.OriginAction)
|
||||
}
|
||||
if instance.WitnessScreen != "/ledger" {
|
||||
t.Fatalf("witness screen = %q, want the detection step's screen", instance.WitnessScreen)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAViolationWithNoWitnessIsReportedNotCounted(t *testing.T) {
|
||||
unattributed := trace.Step{
|
||||
Index: 1,
|
||||
Screen: "/ledger",
|
||||
Violations: []string{"balanceMatches"},
|
||||
}
|
||||
corpus := identified(t, bySelector, run(t, 3, unattributed))
|
||||
|
||||
if len(corpus.Instances) != 0 {
|
||||
t.Fatalf("instances = %d, want 0: %+v", len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
if len(corpus.Unattributed) != 1 ||
|
||||
corpus.Unattributed[0].Property != "balanceMatches" {
|
||||
t.Fatalf("unattributed = %+v, want the one violation with no origin", corpus.Unattributed)
|
||||
}
|
||||
}
|
||||
|
||||
// TestTheStrictActionKeySplitsWhatTheSelectorKeyMerges quantifies the reading
|
||||
// the draft leaves open: two runs that typed different text into the same
|
||||
// field are one defect by selector and two by the whole action.
|
||||
func TestTheStrictActionKeySplitsWhatTheSelectorKeyMerges(t *testing.T) {
|
||||
first := run(t, 3, violating(1, "/ledger", typing("id:amount", "12"), "balanceMatches", 1, 1))
|
||||
second := run(t, 5, violating(1, "/ledger", typing("id:amount", "9000"), "balanceMatches", 1, 1))
|
||||
|
||||
if got := identified(t, bySelector, first, second); len(got.Instances) != 1 {
|
||||
t.Fatalf("by selector: instances = %d, want 1", len(got.Instances))
|
||||
}
|
||||
if got := identified(t, byFullAction, first, second); len(got.Instances) != 2 {
|
||||
t.Fatalf("by full action: instances = %d, want 2", len(got.Instances))
|
||||
}
|
||||
}
|
||||
|
||||
// TestARedactedTypedValueDegradesTheFullKeyVisibly: two runs typed different
|
||||
// values into one field, both reached the trace redacted, and the whole action
|
||||
// can no longer tell them apart. The pair is one row, and the report has to say
|
||||
// so rather than let it read as one defect found twice.
|
||||
func TestARedactedTypedValueDegradesTheFullKeyVisibly(t *testing.T) {
|
||||
first := run(t, 3, violating(1, "/login",
|
||||
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
|
||||
second := run(t, 5, violating(1, "/login",
|
||||
typing("id:password", recordedText(t, "correct horse")), "staysSignedIn", 1, 1))
|
||||
|
||||
corpus := identified(t, byFullAction, first, second)
|
||||
if len(corpus.Instances) != 1 {
|
||||
t.Fatalf("instances = %d, want the redacted pair to be one row: %+v",
|
||||
len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
report := rendered(corpus)
|
||||
if !strings.Contains(report, "1 identity") || !strings.Contains(report, "redacted") {
|
||||
t.Fatalf("report does not say one identity rests on a redacted value:\n%s", report)
|
||||
}
|
||||
}
|
||||
|
||||
func TestARedactedOriginKeepsTheSelectorApart(t *testing.T) {
|
||||
first := run(t, 3, violating(1, "/login",
|
||||
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
|
||||
second := run(t, 5, violating(1, "/login",
|
||||
typing("id:pin", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
|
||||
|
||||
corpus := identified(t, byFullAction, first, second)
|
||||
if len(corpus.Instances) != 2 {
|
||||
t.Fatalf("instances = %d, want two fields to stay two rows: %+v",
|
||||
len(corpus.Instances), corpus.Instances)
|
||||
}
|
||||
if report := rendered(corpus); !strings.Contains(report, "2 identity") {
|
||||
t.Fatalf("report does not count both degraded identities:\n%s", report)
|
||||
}
|
||||
}
|
||||
|
||||
func TestARedactedOriginDegradesNothingUnderTheSelectorKey(t *testing.T) {
|
||||
only := run(t, 3, violating(1, "/login",
|
||||
typing("id:password", recordedText(t, "hunter2")), "staysSignedIn", 1, 1))
|
||||
|
||||
if report := rendered(identified(t, bySelector, only)); strings.Contains(report, "redacted") {
|
||||
t.Fatalf("selector key reads no text, so nothing degrades:\n%s", report)
|
||||
}
|
||||
}
|
||||
|
||||
// recordedText renders a typed value the way the runner records it, so what the
|
||||
// key sees is redaction as it really happens and not a placeholder the test
|
||||
// wrote itself.
|
||||
func recordedText(t *testing.T, typed string) string {
|
||||
t.Helper()
|
||||
recorded := verifier.RecordedActionText(verifier.Action{
|
||||
Kind: verifier.ActionKindInputText,
|
||||
On: "id:password",
|
||||
Text: typed,
|
||||
}, nil)
|
||||
if recorded == typed {
|
||||
t.Fatalf("typed value %q reached the record unredacted", typed)
|
||||
}
|
||||
return recorded
|
||||
}
|
||||
|
||||
func rendered(corpus Corpus) string {
|
||||
var report strings.Builder
|
||||
render(&report, corpus)
|
||||
return report.String()
|
||||
}
|
||||
|
||||
func identified(t *testing.T, mode actionKeyMode, runs ...tracecorpus.Run) Corpus {
|
||||
t.Helper()
|
||||
corpus, err := identify(runs, mode)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return corpus
|
||||
}
|
||||
|
||||
func run(t *testing.T, seed int64, steps ...trace.Step) tracecorpus.Run {
|
||||
t.Helper()
|
||||
directory := t.TempDir()
|
||||
writer, err := trace.NewWriter(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := writer.WriteMeta(trace.Meta{Seed: seed, Platform: "web"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, step := range steps {
|
||||
if err := writer.WriteStep(step); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := writer.Close(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
loaded, err := tracecorpus.Load(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return loaded
|
||||
}
|
||||
|
||||
func step(index int, screen string, action *trace.Action) trace.Step {
|
||||
return trace.Step{Index: index, Screen: screen, NextAction: action}
|
||||
}
|
||||
|
||||
func violating(
|
||||
index int,
|
||||
screen string,
|
||||
action *trace.Action,
|
||||
property string,
|
||||
origin int,
|
||||
detected int,
|
||||
) trace.Step {
|
||||
violated := step(index, screen, action)
|
||||
violated.Violations = []string{property}
|
||||
violated.Witnesses = map[string]trace.Witness{
|
||||
property: {Reason: "predicate false", Step: origin, DetectedStep: detected},
|
||||
}
|
||||
return violated
|
||||
}
|
||||
|
||||
func tap(selector string) *trace.Action {
|
||||
return &trace.Action{Kind: "Tap", Selector: selector, X: 10, Y: 20}
|
||||
}
|
||||
|
||||
func typing(selector string, text string) *trace.Action {
|
||||
return &trace.Action{Kind: "InputText", Selector: selector, Text: text, X: 10, Y: 20}
|
||||
}
|
||||
@@ -0,0 +1,120 @@
|
||||
// Command defect-identity counts distinct defects across stored runs. A
|
||||
// property reports at most once per run, so a run-level count is the number of
|
||||
// properties violated; a defect is identified across runs by the property, the
|
||||
// action attributed as the origin of the failed obligation, and the screen the
|
||||
// witness observed.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"text/tabwriter"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
)
|
||||
|
||||
func main() {
|
||||
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
|
||||
mode := flag.String(
|
||||
"action-key",
|
||||
string(bySelector),
|
||||
"how much of the origin action identifies it: selector or full",
|
||||
)
|
||||
flag.Usage = func() {
|
||||
fmt.Fprintln(
|
||||
os.Stderr,
|
||||
"usage: defect-identity [--json] [--action-key selector|full] <run directory> ...",
|
||||
)
|
||||
}
|
||||
flag.Parse()
|
||||
if flag.NArg() == 0 {
|
||||
flag.Usage()
|
||||
os.Exit(2)
|
||||
}
|
||||
if *mode != string(bySelector) && *mode != string(byFullAction) {
|
||||
fmt.Fprintf(os.Stderr, "unknown --action-key %q\n", *mode)
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
runs, err := loadAll(flag.Args())
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
if len(runs) == 0 {
|
||||
fmt.Fprintln(os.Stderr, "no run directory found under the given paths")
|
||||
os.Exit(1)
|
||||
}
|
||||
corpus, err := identify(runs, actionKeyMode(*mode))
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
if *jsonOut {
|
||||
encoder := json.NewEncoder(os.Stdout)
|
||||
encoder.SetIndent("", " ")
|
||||
if err := encoder.Encode(corpus); err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
return
|
||||
}
|
||||
render(os.Stdout, corpus)
|
||||
}
|
||||
|
||||
func loadAll(paths []string) ([]tracecorpus.Run, error) {
|
||||
var runs []tracecorpus.Run
|
||||
for _, path := range paths {
|
||||
directories, err := tracecorpus.Discover(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for _, directory := range directories {
|
||||
run, err := tracecorpus.Load(directory)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", directory, err)
|
||||
}
|
||||
runs = append(runs, run)
|
||||
}
|
||||
}
|
||||
return runs, nil
|
||||
}
|
||||
|
||||
func render(out io.Writer, corpus Corpus) {
|
||||
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
|
||||
fmt.Fprintln(writer, "property\torigin action\twitness screen\truns\tseeds")
|
||||
for _, instance := range corpus.Instances {
|
||||
screen := instance.WitnessScreen
|
||||
if screen == "" {
|
||||
screen = "(unnamed)"
|
||||
}
|
||||
fmt.Fprintf(writer, "%s\t%s\t%s\t%d\t%v\n",
|
||||
instance.Property, instance.OriginAction, screen,
|
||||
len(instance.Runs), instance.Seeds)
|
||||
}
|
||||
writer.Flush()
|
||||
|
||||
fmt.Fprintf(out, "\n%d distinct defect(s) over %d run(s); %d seen in exactly one run\n",
|
||||
len(corpus.Instances), corpus.Runs, corpus.Singletons())
|
||||
if degraded := corpus.DegradedIdentities(); degraded > 0 {
|
||||
fmt.Fprintf(out,
|
||||
"%d identity(ies) rest on the origin selector alone, because the value typed "+
|
||||
"there is redacted in the record; two runs that typed different values into "+
|
||||
"that field read as one, so the count above is a floor for those\n",
|
||||
degraded)
|
||||
}
|
||||
if corpus.UnnamedScreen > 0 {
|
||||
fmt.Fprintf(out,
|
||||
"%d violation(s) witnessed on a screen the app does not name, "+
|
||||
"so identity rests on property and origin action alone for those\n",
|
||||
corpus.UnnamedScreen)
|
||||
}
|
||||
for _, gap := range corpus.Unattributed {
|
||||
fmt.Fprintf(out, "unattributed: %s at step %d of %s (%s)\n",
|
||||
gap.Property, gap.Step, gap.Run, gap.Reason)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,145 @@
|
||||
// Command exploration-reach counts the distinct structural states a stored
|
||||
// run visited, and compares two runs by the observation at which their
|
||||
// hierarchies first differ. Both read the trace alone: no device, no replay.
|
||||
//
|
||||
// The state is the settle path's structural hash of the recorded hierarchy,
|
||||
// the same function the drivers wait on, so a state boundary here is the state
|
||||
// boundary the harness itself uses.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"text/tabwriter"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
)
|
||||
|
||||
func main() {
|
||||
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
|
||||
reference := flag.String(
|
||||
"reference",
|
||||
"",
|
||||
"run directory to compare against; reports where each other run's hierarchy first differs",
|
||||
)
|
||||
flag.Usage = func() {
|
||||
fmt.Fprintln(
|
||||
os.Stderr,
|
||||
"usage: exploration-reach [--json] [--reference RUN] <run directory> ...",
|
||||
)
|
||||
}
|
||||
flag.Parse()
|
||||
if flag.NArg() == 0 {
|
||||
flag.Usage()
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
reaches, err := measureAll(flag.Args())
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
if len(reaches) == 0 {
|
||||
fmt.Fprintln(os.Stderr, "no run directory found under the given paths")
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
report := Report{Runs: reaches, CorpusDistinct: corpusDistinct(reaches)}
|
||||
if *reference != "" {
|
||||
base, loadErr := load(*reference)
|
||||
if loadErr != nil {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", *reference, loadErr)
|
||||
os.Exit(1)
|
||||
}
|
||||
report.Reference = base.Directory
|
||||
for _, reach := range reaches {
|
||||
if reach.Directory == base.Directory {
|
||||
continue
|
||||
}
|
||||
report.Divergences = append(report.Divergences, diverge(base, reach))
|
||||
}
|
||||
report.MedianDivergence, report.Censored = medianDivergence(report.Divergences)
|
||||
}
|
||||
|
||||
if *jsonOut {
|
||||
encoder := json.NewEncoder(os.Stdout)
|
||||
encoder.SetIndent("", " ")
|
||||
if err := encoder.Encode(report); err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
return
|
||||
}
|
||||
render(os.Stdout, report)
|
||||
}
|
||||
|
||||
// Report is one invocation's output: reach per run and over the corpus, plus
|
||||
// the divergence rows when a reference run was named.
|
||||
type Report struct {
|
||||
Runs []Reach `json:"runs"`
|
||||
CorpusDistinct int `json:"corpus_distinct_states"`
|
||||
Reference string `json:"reference,omitempty"`
|
||||
Divergences []Divergence `json:"divergences,omitempty"`
|
||||
MedianDivergence float64 `json:"median_divergence_index,omitempty"`
|
||||
Censored int `json:"never_diverged,omitempty"`
|
||||
}
|
||||
|
||||
func load(path string) (Reach, error) {
|
||||
run, err := tracecorpus.Load(path)
|
||||
if err != nil {
|
||||
return Reach{}, err
|
||||
}
|
||||
return measure(run), nil
|
||||
}
|
||||
|
||||
func measureAll(paths []string) ([]Reach, error) {
|
||||
var reaches []Reach
|
||||
for _, path := range paths {
|
||||
directories, err := tracecorpus.Discover(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for _, directory := range directories {
|
||||
reach, err := load(directory)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", directory, err)
|
||||
}
|
||||
reaches = append(reaches, reach)
|
||||
}
|
||||
}
|
||||
return reaches, nil
|
||||
}
|
||||
|
||||
func render(out io.Writer, report Report) {
|
||||
writer := tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
|
||||
fmt.Fprintln(writer, "run\tseed\tplatform\tobservations\tdistinct states")
|
||||
for _, reach := range report.Runs {
|
||||
fmt.Fprintf(writer, "%s\t%d\t%s\t%d\t%d\n",
|
||||
reach.Directory, reach.Seed, reach.Platform,
|
||||
reach.Observations, reach.Distinct)
|
||||
}
|
||||
writer.Flush()
|
||||
fmt.Fprintf(out, "\n%d run(s), %d distinct structural states across the corpus\n",
|
||||
len(report.Runs), report.CorpusDistinct)
|
||||
|
||||
if report.Reference == "" {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(out, "\nreference: %s\n", report.Reference)
|
||||
writer = tabwriter.NewWriter(out, 0, 0, 2, ' ', 0)
|
||||
fmt.Fprintln(writer, "run\tseed\tfirst divergence\tobservations compared")
|
||||
for _, divergence := range report.Divergences {
|
||||
where := fmt.Sprintf("%d", divergence.Step)
|
||||
if !divergence.Diverged {
|
||||
where = fmt.Sprintf("none through %d", divergence.Step)
|
||||
}
|
||||
fmt.Fprintf(writer, "%s\t%d\t%s\t%d\n",
|
||||
divergence.Directory, divergence.Seed, where, divergence.Compared)
|
||||
}
|
||||
writer.Flush()
|
||||
fmt.Fprintf(out, "\nmedian first divergence %.1f over %d run(s), %d never diverged\n",
|
||||
report.MedianDivergence, len(report.Divergences), report.Censored)
|
||||
}
|
||||
@@ -0,0 +1,133 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"sort"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion"
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
)
|
||||
|
||||
// observation is one hierarchy-bearing step: the index the run gave it and the
|
||||
// state it observed.
|
||||
type observation struct {
|
||||
Step int
|
||||
State string
|
||||
}
|
||||
|
||||
// Reach is what one run explored. Distinct counts the structural states its
|
||||
// observations visited, which is the measure; Observations is how many looks
|
||||
// it took to visit them.
|
||||
type Reach struct {
|
||||
Directory string `json:"directory"`
|
||||
Seed int64 `json:"seed"`
|
||||
Platform string `json:"platform"`
|
||||
Arm string `json:"arm,omitempty"`
|
||||
Observations int `json:"observations"`
|
||||
Distinct int `json:"distinct_states"`
|
||||
// Unobserved counts steps carrying no hierarchy, which the run-end
|
||||
// finalize record is, so an observation count cannot be read as a step
|
||||
// count by accident.
|
||||
Unobserved int `json:"steps_without_hierarchy"`
|
||||
observations []observation
|
||||
}
|
||||
|
||||
// measure keys each observation by a digest of its structural hash. The
|
||||
// measure is equality between hashes and nothing else, and a corpus holds
|
||||
// thousands of trees whose hashes run to tens of kilobytes each.
|
||||
func measure(run tracecorpus.Run) Reach {
|
||||
reach := Reach{
|
||||
Directory: run.Directory,
|
||||
Seed: run.Meta.Seed,
|
||||
Platform: run.Meta.Platform,
|
||||
Arm: run.Meta.Arm,
|
||||
}
|
||||
distinct := map[string]bool{}
|
||||
for _, step := range run.Steps {
|
||||
if step.Hierarchy == nil {
|
||||
reach.Unobserved++
|
||||
continue
|
||||
}
|
||||
digest := sha256.Sum256([]byte(ioscompanion.StructuralHash(step.Hierarchy)))
|
||||
key := hex.EncodeToString(digest[:])
|
||||
reach.observations = append(
|
||||
reach.observations,
|
||||
observation{Step: step.Index, State: key},
|
||||
)
|
||||
distinct[key] = true
|
||||
}
|
||||
reach.Observations = len(reach.observations)
|
||||
reach.Distinct = len(distinct)
|
||||
return reach
|
||||
}
|
||||
|
||||
// corpusDistinct counts the structural states the whole corpus reached, which
|
||||
// is not the sum of the per-run counts: runs of one application revisit the
|
||||
// same screens.
|
||||
func corpusDistinct(reaches []Reach) int {
|
||||
distinct := map[string]bool{}
|
||||
for _, reach := range reaches {
|
||||
for _, seen := range reach.observations {
|
||||
distinct[seen.State] = true
|
||||
}
|
||||
}
|
||||
return len(distinct)
|
||||
}
|
||||
|
||||
// Divergence is where a replay stopped observing what the reference observed.
|
||||
// Step is the index of the first observation whose structural state differs;
|
||||
// Diverged is false when the replay matched the reference for every
|
||||
// observation the two share, in which case Step is that shared length and the
|
||||
// observation is right-censored.
|
||||
type Divergence struct {
|
||||
Directory string `json:"directory"`
|
||||
Seed int64 `json:"seed"`
|
||||
Step int `json:"step"`
|
||||
Diverged bool `json:"diverged"`
|
||||
Compared int `json:"observations_compared"`
|
||||
}
|
||||
|
||||
func diverge(reference, replay Reach) Divergence {
|
||||
result := Divergence{Directory: replay.Directory, Seed: replay.Seed}
|
||||
shared := len(reference.observations)
|
||||
if len(replay.observations) < shared {
|
||||
shared = len(replay.observations)
|
||||
}
|
||||
result.Compared = shared
|
||||
for position := 0; position < shared; position++ {
|
||||
if reference.observations[position].State != replay.observations[position].State {
|
||||
result.Step = reference.observations[position].Step
|
||||
result.Diverged = true
|
||||
return result
|
||||
}
|
||||
}
|
||||
if shared > 0 {
|
||||
result.Step = reference.observations[shared-1].Step
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
// medianDivergence is E6's number: the median observation index at which a
|
||||
// replay first diverges from the reference. A replay that never diverged
|
||||
// enters at the last index the two share, which is where the observation is
|
||||
// censored rather than where it broke, so the count of such replays is
|
||||
// reported beside the median rather than folded into it.
|
||||
func medianDivergence(divergences []Divergence) (median float64, censored int) {
|
||||
if len(divergences) == 0 {
|
||||
return 0, 0
|
||||
}
|
||||
steps := make([]int, 0, len(divergences))
|
||||
for _, divergence := range divergences {
|
||||
steps = append(steps, divergence.Step)
|
||||
if !divergence.Diverged {
|
||||
censored++
|
||||
}
|
||||
}
|
||||
sort.Ints(steps)
|
||||
middle := len(steps) / 2
|
||||
if len(steps)%2 == 1 {
|
||||
return float64(steps[middle]), censored
|
||||
}
|
||||
return float64(steps[middle-1]+steps[middle]) / 2, censored
|
||||
}
|
||||
@@ -0,0 +1,177 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/driver/ioscompanion"
|
||||
"github.com/priyanshujain/sanderling/internal/hierarchy"
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/tracecorpus"
|
||||
)
|
||||
|
||||
const (
|
||||
home = `{"attributes": {"text": "Home", "bounds": "[0,0,10,10]"}, "children": [
|
||||
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []}
|
||||
]}`
|
||||
homeScrolled = `{"attributes": {"text": "Home", "bounds": "[0,4,10,14]"}, "children": [
|
||||
{"attributes": {"text": "row", "bounds": "[0,4,5,9]"}, "children": []}
|
||||
]}`
|
||||
ledger = `{"attributes": {"text": "Ledger", "bounds": "[0,0,10,10]"}, "children": [
|
||||
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []}
|
||||
]}`
|
||||
ledgerWithRow = `{"attributes": {"text": "Ledger", "bounds": "[0,0,10,10]"}, "children": [
|
||||
{"attributes": {"text": "row", "bounds": "[0,0,5,5]"}, "children": []},
|
||||
{"attributes": {"text": "row"}, "children": []}
|
||||
]}`
|
||||
)
|
||||
|
||||
// TestReachCountsStructuresNotObservations: five observations of three
|
||||
// structures, one of them revisited and one differing only in where it sits on
|
||||
// screen, so the answer is three by construction.
|
||||
func TestReachCountsStructuresNotObservations(t *testing.T) {
|
||||
reach := measureRun(t, 7, home, homeScrolled, ledger, home, ledgerWithRow)
|
||||
|
||||
if reach.Observations != 5 {
|
||||
t.Fatalf("observations = %d, want 5", reach.Observations)
|
||||
}
|
||||
if reach.Distinct != 3 {
|
||||
t.Fatalf("distinct states = %d, want 3", reach.Distinct)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFinalizeRecordIsNoObservation(t *testing.T) {
|
||||
directory := writeRun(t, 7, home, ledger)
|
||||
appendFinalize(t, directory, 3)
|
||||
|
||||
run, err := tracecorpus.Load(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
reach := measure(run)
|
||||
if reach.Observations != 2 || reach.Unobserved != 1 {
|
||||
t.Fatalf("observations = %d, unobserved = %d, want 2 and 1",
|
||||
reach.Observations, reach.Unobserved)
|
||||
}
|
||||
}
|
||||
|
||||
// TestStoredTreeHashesAsTheLiveTreeDid is the equivalence the whole measure
|
||||
// rests on: the hash the settle path computed from the live tree and the hash
|
||||
// this tool computes from the stored one are the same string, so a reach
|
||||
// number counts the same state boundaries the drivers wait on.
|
||||
func TestStoredTreeHashesAsTheLiveTreeDid(t *testing.T) {
|
||||
live, err := hierarchy.Parse(home)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
liveHash := ioscompanion.StructuralHash(live)
|
||||
if liveHash == "" {
|
||||
t.Fatal("live hash is empty, so the test would pass on any stored tree")
|
||||
}
|
||||
|
||||
run, err := tracecorpus.Load(writeRun(t, 7, home))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stored := ioscompanion.StructuralHash(run.Steps[0].Hierarchy)
|
||||
if stored != liveHash {
|
||||
t.Fatalf("stored hash differs from the live one:\n live=%q\n stored=%q",
|
||||
liveHash, stored)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFirstDivergenceNamesTheObservationThatDiffers(t *testing.T) {
|
||||
reference := measureRun(t, 7, home, ledger, home, ledger)
|
||||
replay := measureRun(t, 7, home, ledger, ledgerWithRow, ledger)
|
||||
|
||||
divergence := diverge(reference, replay)
|
||||
if !divergence.Diverged || divergence.Step != 3 {
|
||||
t.Fatalf("divergence = %+v, want the third observation", divergence)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAReplayThatMatchesIsCensoredAtTheSharedLength(t *testing.T) {
|
||||
reference := measureRun(t, 7, home, ledger, home, ledger)
|
||||
replay := measureRun(t, 7, home, ledger, home)
|
||||
|
||||
divergence := diverge(reference, replay)
|
||||
if divergence.Diverged {
|
||||
t.Fatalf("identical observations must not report a divergence: %+v", divergence)
|
||||
}
|
||||
if divergence.Step != 3 || divergence.Compared != 3 {
|
||||
t.Fatalf("divergence = %+v, want censoring at the third observation", divergence)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMedianDivergenceHoldsCensoredReplaysApart(t *testing.T) {
|
||||
median, censored := medianDivergence([]Divergence{
|
||||
{Step: 9, Diverged: true},
|
||||
{Step: 3, Diverged: true},
|
||||
{Step: 40, Diverged: false},
|
||||
})
|
||||
if median != 9 {
|
||||
t.Fatalf("median = %v, want 9", median)
|
||||
}
|
||||
if censored != 1 {
|
||||
t.Fatalf("censored = %d, want 1", censored)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCorpusStatesAreTheUnionNotTheSum(t *testing.T) {
|
||||
first := measureRun(t, 7, home, ledger)
|
||||
second := measureRun(t, 11, ledger, ledgerWithRow)
|
||||
|
||||
if got := corpusDistinct([]Reach{first, second}); got != 3 {
|
||||
t.Fatalf("corpus distinct = %d, want 3", got)
|
||||
}
|
||||
}
|
||||
|
||||
func measureRun(t *testing.T, seed int64, dumps ...string) Reach {
|
||||
t.Helper()
|
||||
run, err := tracecorpus.Load(writeRun(t, seed, dumps...))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return measure(run)
|
||||
}
|
||||
|
||||
func writeRun(t *testing.T, seed int64, dumps ...string) string {
|
||||
t.Helper()
|
||||
directory := t.TempDir()
|
||||
writer, err := trace.NewWriter(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := writer.WriteMeta(trace.Meta{Seed: seed, Platform: "web"}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for index, dump := range dumps {
|
||||
tree, err := hierarchy.Parse(dump)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := writer.WriteStep(trace.Step{Index: index + 1, Hierarchy: tree}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := writer.Close(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return directory
|
||||
}
|
||||
|
||||
func appendFinalize(t *testing.T, directory string, index int) {
|
||||
t.Helper()
|
||||
writer, err := trace.NewWriter(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := writer.WriteStep(trace.Step{
|
||||
Index: index,
|
||||
Violations: []string{"someTransactionExists"},
|
||||
}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := writer.Close(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,482 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"math/rand"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// The test binary doubles as the stub implementation's server and as the
|
||||
// fetcher the stub campaign uses, so the sweep drives a real preview process
|
||||
// over a real port and the served URL is answered by a real HTTP server.
|
||||
func TestMain(m *testing.M) {
|
||||
switch {
|
||||
case os.Getenv("SWEEP_TEST_SERVE_PORT") != "":
|
||||
serveUntilKilled(os.Getenv("SWEEP_TEST_SERVE_PORT"))
|
||||
case os.Getenv("SWEEP_TEST_FETCH_URL") != "":
|
||||
recordFetch(
|
||||
os.Getenv("SWEEP_TEST_FETCH_URL"),
|
||||
os.Getenv("SWEEP_TEST_FETCH_LOG"),
|
||||
)
|
||||
default:
|
||||
os.Exit(m.Run())
|
||||
}
|
||||
}
|
||||
|
||||
func serveUntilKilled(port string) {
|
||||
handler := http.HandlerFunc(
|
||||
func(writer http.ResponseWriter, request *http.Request) {
|
||||
fmt.Fprintf(writer, "%s %s", port, request.URL.RequestURI())
|
||||
},
|
||||
)
|
||||
http.ListenAndServe("localhost:"+port, handler)
|
||||
}
|
||||
|
||||
func recordFetch(url, logPath string) {
|
||||
line := ""
|
||||
response, err := http.Get(url)
|
||||
if err != nil {
|
||||
line = fmt.Sprintf("%s -> error %v\n", url, err)
|
||||
} else {
|
||||
body, _ := io.ReadAll(response.Body)
|
||||
response.Body.Close()
|
||||
line = fmt.Sprintf("%s -> %d %s\n", url, response.StatusCode, body)
|
||||
}
|
||||
logFile, err := os.OpenFile(
|
||||
logPath,
|
||||
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
|
||||
0o644,
|
||||
)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
logFile.WriteString(line)
|
||||
logFile.Close()
|
||||
}
|
||||
|
||||
// stubBun answers install, fails to build impl-02, and serves the preview from
|
||||
// the test binary on the port it was given.
|
||||
const stubBun = `#!/bin/sh
|
||||
echo "$PWD $*" >> "%[1]s"
|
||||
if [ "$1" = "run" ] && [ "$2" = "build" ]; then
|
||||
case "$PWD" in *impl-02) echo "TS2322: type error" >&2; exit 1 ;; esac
|
||||
exit 0
|
||||
fi
|
||||
if [ "$1" = "run" ] && [ "$2" = "preview" ]; then
|
||||
port=""
|
||||
previous=""
|
||||
for argument in "$@"; do
|
||||
if [ "$previous" = "--port" ]; then port="$argument"; fi
|
||||
previous="$argument"
|
||||
done
|
||||
SWEEP_TEST_SERVE_PORT="$port" exec "%[2]s"
|
||||
fi
|
||||
exit 0
|
||||
`
|
||||
|
||||
// stubCampaign records the argv it was handed and whether the sweep manifest
|
||||
// was already on disk when it ran, fetches the URL it was told to drive, writes
|
||||
// the campaign directory the real tool would write, and fails impl-03 seed 4.
|
||||
const stubCampaign = `#!/bin/sh
|
||||
output=""
|
||||
seed=""
|
||||
url=""
|
||||
previous=""
|
||||
for argument in "$@"; do
|
||||
case "$previous" in
|
||||
--output) output="$argument" ;;
|
||||
--seeds) seed="$argument" ;;
|
||||
--bundle-id) url="$argument" ;;
|
||||
esac
|
||||
previous="$argument"
|
||||
done
|
||||
manifest=missing
|
||||
if [ -f "%[1]s" ]; then manifest=present; fi
|
||||
echo "manifest=$manifest argv: $*" >> "%[2]s"
|
||||
SWEEP_TEST_FETCH_URL="$url" SWEEP_TEST_FETCH_LOG="%[3]s" "%[4]s"
|
||||
mkdir -p "$output"
|
||||
printf '{"arm":"stub","seeds":[%%s]}\n' "$seed" > "$output/campaign.json"
|
||||
echo "stub campaign seed=$seed url=$url"
|
||||
case "$output" in
|
||||
*impl-03/seed-4) exit 1 ;;
|
||||
esac
|
||||
exit 0
|
||||
`
|
||||
|
||||
func TestRun_EndToEndAgainstStubBunAndCampaign(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
implementations := filepath.Join(root, "implementations")
|
||||
for _, name := range []string{"impl-01", "impl-02", "impl-03"} {
|
||||
if err := os.MkdirAll(filepath.Join(implementations, name), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
specPath := filepath.Join(root, "relay.ts")
|
||||
if err := os.WriteFile(specPath, []byte("export const properties = [];\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
output := filepath.Join(root, "campaigns")
|
||||
bunLog := filepath.Join(root, "bun.log")
|
||||
campaignLog := filepath.Join(root, "campaign.log")
|
||||
fetchLog := filepath.Join(root, "fetch.log")
|
||||
testBinary, err := filepath.Abs(os.Args[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
bunPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-bun"),
|
||||
fmt.Sprintf(stubBun, bunLog, testBinary),
|
||||
)
|
||||
campaignPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-campaign"),
|
||||
fmt.Sprintf(
|
||||
stubCampaign,
|
||||
filepath.Join(output, manifestFileName),
|
||||
campaignLog,
|
||||
fetchLog,
|
||||
testBinary,
|
||||
),
|
||||
)
|
||||
sanderlingPath := writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-sanderling"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
)
|
||||
basePort := freePortRange(t, 3)
|
||||
|
||||
var stdout bytes.Buffer
|
||||
err = run([]string{
|
||||
"--implementations", implementations,
|
||||
"--spec", specPath,
|
||||
"--seeds", "4-5",
|
||||
"--max-steps", "40",
|
||||
"--duration", "30s",
|
||||
"--concurrency", "2",
|
||||
"--base-port", fmt.Sprint(basePort),
|
||||
"--output", output,
|
||||
"--bun", bunPath,
|
||||
"--campaign", campaignPath,
|
||||
"--sanderling", sanderlingPath,
|
||||
}, &stdout, io.Discard)
|
||||
if err == nil ||
|
||||
!strings.Contains(err.Error(), "1 of 3 implementations never ran") {
|
||||
t.Fatalf(
|
||||
"expected the failed build and the failed campaign to be reported, got %v",
|
||||
err,
|
||||
)
|
||||
}
|
||||
|
||||
var recorded manifest
|
||||
manifestBody, err := os.ReadFile(filepath.Join(output, manifestFileName))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := json.Unmarshal(manifestBody, &recorded); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !slices.Equal(recorded.Seeds, []int64{4, 5}) {
|
||||
t.Errorf("intended seeds: got %v", recorded.Seeds)
|
||||
}
|
||||
if len(recorded.Implementations) != 3 {
|
||||
t.Fatalf("intended implementations: got %v", recorded.Implementations)
|
||||
}
|
||||
for index, planned := range recorded.Implementations {
|
||||
wantPort := basePort + index
|
||||
if planned.Port != wantPort {
|
||||
t.Errorf(
|
||||
"%s port: got %d, want %d",
|
||||
planned.Name,
|
||||
planned.Port,
|
||||
wantPort,
|
||||
)
|
||||
}
|
||||
wantURL := fmt.Sprintf("http://localhost:%d/?seed={seed}", wantPort)
|
||||
if planned.URLTemplate != wantURL {
|
||||
t.Errorf(
|
||||
"%s url template: got %q, want %q",
|
||||
planned.Name,
|
||||
planned.URLTemplate,
|
||||
wantURL,
|
||||
)
|
||||
}
|
||||
}
|
||||
if recorded.Generator != "seeded" || recorded.MaxSteps != 40 {
|
||||
t.Errorf(
|
||||
"manifest generator/budget: got %q/%d",
|
||||
recorded.Generator,
|
||||
recorded.MaxSteps,
|
||||
)
|
||||
}
|
||||
|
||||
// impl-02 fails its build, so the two implementations either side of it
|
||||
// still have to reach the campaign tool with both seeds.
|
||||
campaignLines := readLines(t, campaignLog)
|
||||
if len(campaignLines) != 4 {
|
||||
t.Fatalf(
|
||||
"campaign invocations: got %d, want 4:\n%s",
|
||||
len(campaignLines),
|
||||
strings.Join(campaignLines, "\n"),
|
||||
)
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
for _, line := range campaignLines {
|
||||
if !strings.HasPrefix(line, "manifest=present") {
|
||||
t.Errorf(
|
||||
"a campaign ran before the sweep manifest was written: %q",
|
||||
line,
|
||||
)
|
||||
}
|
||||
arguments := strings.Fields(strings.SplitN(line, "argv: ", 2)[1])
|
||||
arm := argumentValue(arguments, "--arm")
|
||||
seed := argumentValue(arguments, "--seeds")
|
||||
bundle := argumentValue(arguments, "--bundle-id")
|
||||
port := basePort + slices.Index(
|
||||
[]string{"impl-01", "impl-02", "impl-03"},
|
||||
arm,
|
||||
)
|
||||
wantBundle := fmt.Sprintf("http://localhost:%d/?seed=%s", port, seed)
|
||||
if bundle != wantBundle {
|
||||
t.Errorf(
|
||||
"%s seed %s: bundle id %q, want %q",
|
||||
arm,
|
||||
seed,
|
||||
bundle,
|
||||
wantBundle,
|
||||
)
|
||||
}
|
||||
if got := argumentValue(arguments, "--output"); got != filepath.Join(
|
||||
output,
|
||||
arm,
|
||||
"seed-"+seed,
|
||||
) {
|
||||
t.Errorf("%s seed %s: campaign output %q", arm, seed, got)
|
||||
}
|
||||
if got := argumentValue(arguments, "--sanderling"); got != sanderlingPath {
|
||||
t.Errorf("%s seed %s: sanderling path %q", arm, seed, got)
|
||||
}
|
||||
seen[arm+"/"+seed] = true
|
||||
}
|
||||
for _, want := range []string{"impl-01/4", "impl-01/5", "impl-03/4", "impl-03/5"} {
|
||||
if !seen[want] {
|
||||
t.Errorf("%s never reached the campaign tool", want)
|
||||
}
|
||||
}
|
||||
|
||||
// What the served page actually saw: the right port for the
|
||||
// implementation, carrying the same seed the campaign was given.
|
||||
fetched := readLines(t, fetchLog)
|
||||
for _, want := range []string{
|
||||
fmt.Sprintf("http://localhost:%d/?seed=4 -> 200 %d /?seed=4", basePort, basePort),
|
||||
fmt.Sprintf("http://localhost:%d/?seed=5 -> 200 %d /?seed=5", basePort, basePort),
|
||||
fmt.Sprintf("http://localhost:%d/?seed=4 -> 200 %d /?seed=4", basePort+2, basePort+2),
|
||||
fmt.Sprintf("http://localhost:%d/?seed=5 -> 200 %d /?seed=5", basePort+2, basePort+2),
|
||||
} {
|
||||
if !slices.Contains(fetched, want) {
|
||||
t.Errorf(
|
||||
"the served page never saw %q:\n%s",
|
||||
want,
|
||||
strings.Join(fetched, "\n"),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
records := readRecords(t, filepath.Join(output, recordsFileName))
|
||||
if len(records) != 3 {
|
||||
t.Fatalf("implementation records: got %d, want 3", len(records))
|
||||
}
|
||||
byName := map[string]implementationRecord{}
|
||||
for _, record := range records {
|
||||
byName[record.Name] = record
|
||||
}
|
||||
failed := byName["impl-02"]
|
||||
if failed.FailedStage != stageBuild || len(failed.Runs) != 0 {
|
||||
t.Errorf(
|
||||
"impl-02: got stage %q with %d runs, want a build failure and no runs",
|
||||
failed.FailedStage,
|
||||
len(failed.Runs),
|
||||
)
|
||||
}
|
||||
if !strings.Contains(failed.Error, "build.log") {
|
||||
t.Errorf("impl-02 error should point at its log: %q", failed.Error)
|
||||
}
|
||||
buildLog, err := os.ReadFile(
|
||||
filepath.Join(output, "impl-02", stageBuild+".log"),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(buildLog), "TS2322") {
|
||||
t.Errorf("impl-02 build log lost the compiler error: %q", buildLog)
|
||||
}
|
||||
for _, name := range []string{"impl-01", "impl-03"} {
|
||||
record := byName[name]
|
||||
if record.FailedStage != "" || len(record.Runs) != 2 {
|
||||
t.Errorf(
|
||||
"%s: stage %q with %d runs, want no failure and 2 runs",
|
||||
name,
|
||||
record.FailedStage,
|
||||
len(record.Runs),
|
||||
)
|
||||
}
|
||||
if record.MonotonicMillis <= 0 {
|
||||
t.Errorf(
|
||||
"%s took %d ms, so nothing timed how long it worked",
|
||||
name,
|
||||
record.MonotonicMillis,
|
||||
)
|
||||
}
|
||||
for _, run := range record.Runs {
|
||||
if run.MonotonicMillis <= 0 {
|
||||
t.Errorf(
|
||||
"%s seed %d took %d ms, so nothing timed the campaign",
|
||||
name,
|
||||
run.Seed,
|
||||
run.MonotonicMillis,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
if exit := byName["impl-03"].Runs[0].ExitCode; exit != 1 {
|
||||
t.Errorf("impl-03 seed 4 exit code: got %d, want 1", exit)
|
||||
}
|
||||
if exit := byName["impl-03"].Runs[1].ExitCode; exit != 0 {
|
||||
t.Errorf(
|
||||
"impl-03 seed 5 ran after seed 4 failed and should have exited 0, got %d",
|
||||
exit,
|
||||
)
|
||||
}
|
||||
|
||||
for _, name := range []string{"impl-01", "impl-03"} {
|
||||
for _, seed := range []string{"4", "5"} {
|
||||
directory := filepath.Join(output, name, "seed-"+seed)
|
||||
if _, err := os.Stat(filepath.Join(directory, "campaign.json")); err != nil {
|
||||
t.Errorf(
|
||||
"%s seed %s: no campaign directory: %v",
|
||||
name,
|
||||
seed,
|
||||
err,
|
||||
)
|
||||
}
|
||||
log, err := os.ReadFile(filepath.Join(directory, "campaign.log"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(log), "stub campaign seed="+seed) {
|
||||
t.Errorf(
|
||||
"%s seed %s: campaign output was not captured: %q",
|
||||
name,
|
||||
seed,
|
||||
log,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
installed := readLines(t, bunLog)
|
||||
for _, name := range []string{"impl-01", "impl-02", "impl-03"} {
|
||||
if !slices.Contains(
|
||||
installed,
|
||||
filepath.Join(implementations, name)+" install",
|
||||
) {
|
||||
t.Errorf(
|
||||
"%s was never installed:\n%s",
|
||||
name,
|
||||
strings.Join(installed, "\n"),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// Every preview server the sweep started is gone with it: a leaked one
|
||||
// holds its port, and the next sweep would be served by the old build.
|
||||
client := &http.Client{Timeout: 2 * time.Second}
|
||||
for _, port := range []int{basePort, basePort + 2} {
|
||||
if response, err := client.Get(readinessURL(port)); err == nil {
|
||||
response.Body.Close()
|
||||
t.Errorf("port %d is still served after the sweep finished", port)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(stdout.String(), "failed at build") {
|
||||
t.Errorf(
|
||||
"progress output does not name the build failure: %q",
|
||||
stdout.String(),
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func writeScript(t *testing.T, path, body string) string {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(path, []byte(body), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
func readLines(t *testing.T, path string) []string {
|
||||
t.Helper()
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var lines []string
|
||||
scanner := bufio.NewScanner(strings.NewReader(string(body)))
|
||||
for scanner.Scan() {
|
||||
if line := strings.TrimSpace(scanner.Text()); line != "" {
|
||||
lines = append(lines, line)
|
||||
}
|
||||
}
|
||||
return lines
|
||||
}
|
||||
|
||||
func readRecords(t *testing.T, path string) []implementationRecord {
|
||||
t.Helper()
|
||||
var records []implementationRecord
|
||||
for _, line := range readLines(t, path) {
|
||||
var record implementationRecord
|
||||
if err := json.Unmarshal([]byte(line), &record); err != nil {
|
||||
t.Fatalf("%s: %v", line, err)
|
||||
}
|
||||
records = append(records, record)
|
||||
}
|
||||
return records
|
||||
}
|
||||
|
||||
// freePortRange finds count consecutive free ports, which is what the sweep
|
||||
// hands out: one port per implementation from --base-port upwards.
|
||||
func freePortRange(t *testing.T, count int) int {
|
||||
t.Helper()
|
||||
for range 100 {
|
||||
base := 20000 + rand.Intn(20000)
|
||||
if portsAreFree(base, count) {
|
||||
return base
|
||||
}
|
||||
}
|
||||
t.Fatalf("no run of %d free ports", count)
|
||||
return 0
|
||||
}
|
||||
|
||||
func portsAreFree(base, count int) bool {
|
||||
for offset := range count {
|
||||
listener, err := net.Listen(
|
||||
"tcp",
|
||||
fmt.Sprintf("localhost:%d", base+offset),
|
||||
)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
listener.Close()
|
||||
}
|
||||
return true
|
||||
}
|
||||
@@ -0,0 +1,296 @@
|
||||
// Command implementation-sweep runs one identical campaign against every model
|
||||
// implementation of a single requirement. It installs, builds and serves each
|
||||
// implementation on its own port, then hands the campaign tool the same seed
|
||||
// slice, the same step budget and the same generator for all of them, so a
|
||||
// difference between implementations is not a difference in exploration.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/seedspec"
|
||||
)
|
||||
|
||||
// The generator and the platform are fixed rather than exposed: the
|
||||
// pre-registration runs the seeded policy against a served web build, and a
|
||||
// sweep that could quietly run something else records a comparison nobody made.
|
||||
const (
|
||||
generator = "seeded"
|
||||
platform = "web"
|
||||
)
|
||||
|
||||
// defaultConcurrency is how many implementations are built, served and swept at
|
||||
// once. fleet.md measured eight concurrent web campaigns clean at about 1.1 GB
|
||||
// resident each, parallel efficiency 0.83 at eight against 0.87 at six, and a
|
||||
// knee at twelve to sixteen, on a contended laptop it says to re-measure before
|
||||
// trusting anything above eight. Six sits at the better efficiency, costs about
|
||||
// 7 GB of a 64 GB host, and leaves the Android emulator farm that shares that
|
||||
// host its four to six slots. Each worker here also carries a vite server the
|
||||
// fleet measurement did not include.
|
||||
const defaultConcurrency = 6
|
||||
|
||||
// defaultBasePort is the first port handed out. Vite's own defaults are 5173
|
||||
// and 4173, so a sweep starting here does not collide with a dev server someone
|
||||
// left running.
|
||||
const defaultBasePort = 5300
|
||||
|
||||
type config struct {
|
||||
implementationsDirectory string
|
||||
specPath string
|
||||
outputDirectory string
|
||||
seeds []int64
|
||||
maxSteps int
|
||||
duration time.Duration
|
||||
concurrency int
|
||||
basePort int
|
||||
bunPath string
|
||||
campaignPath string
|
||||
sanderlingPath string
|
||||
extraArguments []string
|
||||
}
|
||||
|
||||
const usage = `implementation-sweep runs one seeded campaign per model implementation.
|
||||
|
||||
Usage:
|
||||
implementation-sweep --implementations <dir> --spec <path> --seeds <spec>
|
||||
--max-steps <n> --output <dir> [flags]
|
||||
[-- <sanderling test flags>]
|
||||
|
||||
Each impl-* directory under --implementations is installed, built and served on
|
||||
its own port, and every one is swept with the same seeds and the same step
|
||||
budget. An implementation that fails to install, build or serve is recorded and
|
||||
the sweep moves on to the next one.
|
||||
|
||||
Everything after a bare -- reaches every sanderling test call through the
|
||||
campaign tool.
|
||||
`
|
||||
|
||||
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
flagSet := flag.NewFlagSet("implementation-sweep", flag.ContinueOnError)
|
||||
flagSet.SetOutput(stderr)
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprint(stderr, usage)
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
var configuration config
|
||||
var seedSpecification string
|
||||
flagSet.StringVar(
|
||||
&configuration.implementationsDirectory,
|
||||
"implementations",
|
||||
"",
|
||||
"directory holding impl-01 to impl-NN (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.specPath,
|
||||
"spec",
|
||||
"",
|
||||
"path to the property set every implementation is run against (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&seedSpecification,
|
||||
"seeds",
|
||||
"",
|
||||
"seeds every implementation runs: ranges and lists, e.g. 1-10,20 (required)",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.maxSteps,
|
||||
"max-steps",
|
||||
0,
|
||||
"per-run step budget, identical across implementations (required, must be positive)",
|
||||
)
|
||||
flagSet.DurationVar(
|
||||
&configuration.duration,
|
||||
"duration",
|
||||
5*time.Minute,
|
||||
"per-run wall-clock ceiling passed to each campaign",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.concurrency,
|
||||
"concurrency",
|
||||
defaultConcurrency,
|
||||
"implementations built, served and swept at once",
|
||||
)
|
||||
flagSet.IntVar(
|
||||
&configuration.basePort,
|
||||
"base-port",
|
||||
defaultBasePort,
|
||||
"first port served; each implementation takes the next one in name order",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.outputDirectory,
|
||||
"output",
|
||||
"",
|
||||
"campaign tree to create (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.bunPath,
|
||||
"bun",
|
||||
"bun",
|
||||
"bun binary that installs, builds and serves each implementation",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.campaignPath,
|
||||
"campaign",
|
||||
"campaign",
|
||||
"campaign binary to invoke per implementation and seed",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.sanderlingPath,
|
||||
"sanderling",
|
||||
"sanderling",
|
||||
"sanderling binary each campaign invokes",
|
||||
)
|
||||
if err := flagSet.Parse(arguments); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
configuration.extraArguments = flagSet.Args()
|
||||
|
||||
// Every missing flag is named together, in flag order: stopping at the
|
||||
// first turns one rerun into one rerun per missing flag.
|
||||
var missing []error
|
||||
for _, required := range []struct {
|
||||
name string
|
||||
value string
|
||||
}{
|
||||
{"--implementations", configuration.implementationsDirectory},
|
||||
{"--spec", configuration.specPath},
|
||||
{"--seeds", seedSpecification},
|
||||
{"--output", configuration.outputDirectory},
|
||||
} {
|
||||
if required.value == "" {
|
||||
missing = append(
|
||||
missing,
|
||||
fmt.Errorf("%s is required", required.name),
|
||||
)
|
||||
}
|
||||
}
|
||||
if err := errors.Join(missing...); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
if configuration.maxSteps <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--max-steps must be positive: every implementation needs the same step budget",
|
||||
)
|
||||
}
|
||||
if configuration.duration <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--duration must be positive: %s",
|
||||
configuration.duration,
|
||||
)
|
||||
}
|
||||
if configuration.concurrency <= 0 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--concurrency must be positive: %d",
|
||||
configuration.concurrency,
|
||||
)
|
||||
}
|
||||
if configuration.basePort < 1024 || configuration.basePort > 65535 {
|
||||
return config{}, fmt.Errorf(
|
||||
"--base-port %d is outside 1024-65535",
|
||||
configuration.basePort,
|
||||
)
|
||||
}
|
||||
seeds, err := seedspec.Parse(seedSpecification)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("--seeds: %w", err)
|
||||
}
|
||||
configuration.seeds = seeds
|
||||
for name, value := range map[string]*string{
|
||||
"--implementations": &configuration.implementationsDirectory,
|
||||
"--spec": &configuration.specPath,
|
||||
"--output": &configuration.outputDirectory,
|
||||
} {
|
||||
absolute, err := filepath.Abs(*value)
|
||||
if err != nil {
|
||||
return config{}, fmt.Errorf("%s: %w", name, err)
|
||||
}
|
||||
*value = absolute
|
||||
}
|
||||
return configuration, nil
|
||||
}
|
||||
|
||||
// servedURL is the one place a seed becomes a URL. The same seed is also handed
|
||||
// to the campaign as --seeds, which reaches sanderling as --seed and fixes the
|
||||
// exploration, while the scaffold reads ?seed= and fixes the latency and the
|
||||
// outcome of every send. A violation replays only when both carry the same
|
||||
// number, so both come from the seed argument here and never from two flags.
|
||||
func servedURL(port int, seed string) string {
|
||||
return fmt.Sprintf("http://localhost:%d/?seed=%s", port, seed)
|
||||
}
|
||||
|
||||
func readinessURL(port int) string {
|
||||
return fmt.Sprintf("http://localhost:%d/", port)
|
||||
}
|
||||
|
||||
func campaignDirectory(
|
||||
configuration config,
|
||||
target implementation,
|
||||
seed string,
|
||||
) string {
|
||||
return filepath.Join(
|
||||
configuration.outputDirectory,
|
||||
target.Name,
|
||||
"seed-"+seed,
|
||||
)
|
||||
}
|
||||
|
||||
// campaignArguments builds one campaign invocation. The seed is a string
|
||||
// because it lands in two arguments, --seeds and the ?seed= of --bundle-id,
|
||||
// and passing it once keeps them from drifting apart.
|
||||
func campaignArguments(
|
||||
configuration config,
|
||||
target implementation,
|
||||
seed string,
|
||||
) []string {
|
||||
arguments := []string{
|
||||
"--spec", configuration.specPath,
|
||||
"--bundle-id", servedURL(target.Port, seed),
|
||||
"--platform", platform,
|
||||
"--arm", target.Name,
|
||||
"--generator", generator,
|
||||
"--max-steps", strconv.Itoa(configuration.maxSteps),
|
||||
"--duration", configuration.duration.String(),
|
||||
"--seeds", seed,
|
||||
"--sanderling", configuration.sanderlingPath,
|
||||
"--output", campaignDirectory(configuration, target, seed),
|
||||
}
|
||||
if len(configuration.extraArguments) > 0 {
|
||||
arguments = append(arguments, "--")
|
||||
arguments = append(arguments, configuration.extraArguments...)
|
||||
}
|
||||
return arguments
|
||||
}
|
||||
|
||||
func run(arguments []string, stdout, stderr io.Writer) error {
|
||||
configuration, err := parseArguments(arguments, stderr)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ctx, cancel := signal.NotifyContext(
|
||||
context.Background(),
|
||||
os.Interrupt,
|
||||
syscall.SIGTERM,
|
||||
)
|
||||
defer cancel()
|
||||
return runSweep(ctx, configuration, stdout)
|
||||
}
|
||||
|
||||
func main() {
|
||||
if err := run(os.Args[1:], os.Stdout, os.Stderr); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(os.Stderr, "error: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,241 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"io"
|
||||
"net/url"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func baseArguments() []string {
|
||||
return []string{
|
||||
"--implementations", "/e4/implementations",
|
||||
"--spec", "/e4/relay.ts",
|
||||
"--seeds", "1-3",
|
||||
"--max-steps", "400",
|
||||
"--output", "/campaigns/e4",
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseArguments_DefaultsAndSeeds(t *testing.T) {
|
||||
configuration, err := parseArguments(baseArguments(), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !slices.Equal(configuration.seeds, []int64{1, 2, 3}) {
|
||||
t.Errorf("seeds: got %v", configuration.seeds)
|
||||
}
|
||||
if configuration.concurrency != defaultConcurrency {
|
||||
t.Errorf(
|
||||
"concurrency default: got %d, want %d",
|
||||
configuration.concurrency,
|
||||
defaultConcurrency,
|
||||
)
|
||||
}
|
||||
if configuration.basePort != defaultBasePort {
|
||||
t.Errorf(
|
||||
"base port default: got %d, want %d",
|
||||
configuration.basePort,
|
||||
defaultBasePort,
|
||||
)
|
||||
}
|
||||
if configuration.duration != 5*time.Minute {
|
||||
t.Errorf("duration default: got %s", configuration.duration)
|
||||
}
|
||||
for name, got := range map[string]string{
|
||||
"bun": configuration.bunPath,
|
||||
"campaign": configuration.campaignPath,
|
||||
"sanderling": configuration.sanderlingPath,
|
||||
} {
|
||||
if got != name {
|
||||
t.Errorf("%s path default: got %q", name, got)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseArguments_Rejections(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
arguments []string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
"missing implementations",
|
||||
[]string{
|
||||
"--spec",
|
||||
"s",
|
||||
"--seeds",
|
||||
"1",
|
||||
"--max-steps",
|
||||
"10",
|
||||
"--output",
|
||||
"o",
|
||||
},
|
||||
"--implementations is required",
|
||||
},
|
||||
{
|
||||
"missing spec",
|
||||
[]string{
|
||||
"--implementations",
|
||||
"i",
|
||||
"--seeds",
|
||||
"1",
|
||||
"--max-steps",
|
||||
"10",
|
||||
"--output",
|
||||
"o",
|
||||
},
|
||||
"--spec is required",
|
||||
},
|
||||
{
|
||||
"missing output",
|
||||
[]string{
|
||||
"--implementations",
|
||||
"i",
|
||||
"--spec",
|
||||
"s",
|
||||
"--seeds",
|
||||
"1",
|
||||
"--max-steps",
|
||||
"10",
|
||||
},
|
||||
"--output is required",
|
||||
},
|
||||
{
|
||||
"zero max steps",
|
||||
append(baseArguments(), "--max-steps", "0"),
|
||||
"--max-steps must be positive",
|
||||
},
|
||||
{
|
||||
"zero concurrency",
|
||||
append(baseArguments(), "--concurrency", "0"),
|
||||
"--concurrency must be positive",
|
||||
},
|
||||
{
|
||||
"privileged base port",
|
||||
append(baseArguments(), "--base-port", "80"),
|
||||
"outside 1024-65535",
|
||||
},
|
||||
{
|
||||
"seed zero",
|
||||
append(baseArguments(), "--seeds", "0,1"),
|
||||
"not reproducible",
|
||||
},
|
||||
}
|
||||
for _, testCase := range cases {
|
||||
_, err := parseArguments(testCase.arguments, io.Discard)
|
||||
if err == nil {
|
||||
t.Errorf("%s: expected error", testCase.name)
|
||||
continue
|
||||
}
|
||||
if !strings.Contains(err.Error(), testCase.want) {
|
||||
t.Errorf(
|
||||
"%s: got %q, want it to contain %q",
|
||||
testCase.name,
|
||||
err,
|
||||
testCase.want,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Three flags missing is one rerun, not three: the operator is told about all
|
||||
// of them at once, in flag order, whatever order the check happened to walk.
|
||||
func TestParseArguments_NamesEveryMissingRequiredFlagInFlagOrder(t *testing.T) {
|
||||
_, err := parseArguments(
|
||||
[]string{"--spec", "s", "--max-steps", "10"},
|
||||
io.Discard,
|
||||
)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want every missing flag named")
|
||||
}
|
||||
message := err.Error()
|
||||
previous := -1
|
||||
for _, name := range []string{"--implementations", "--seeds", "--output"} {
|
||||
at := strings.Index(message, name)
|
||||
if at < 0 {
|
||||
t.Fatalf("got %q, want %s named", message, name)
|
||||
}
|
||||
if at < previous {
|
||||
t.Errorf("got %q, want the flags named in flag order", message)
|
||||
}
|
||||
previous = at
|
||||
}
|
||||
if strings.Contains(message, "--spec") {
|
||||
t.Errorf("got %q, want the supplied --spec left out", message)
|
||||
}
|
||||
}
|
||||
|
||||
// The seed reaches two independent things, the campaign's own seed and the
|
||||
// scaffold's failure stream, and a replay reproduces neither unless they carry
|
||||
// the same number.
|
||||
func TestCampaignArguments_OneSeedReachesBothTheCampaignAndTheURL(
|
||||
t *testing.T,
|
||||
) {
|
||||
configuration, err := parseArguments(
|
||||
append(baseArguments(), "--", "--clear-data=false"),
|
||||
io.Discard,
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
target := implementation{
|
||||
Name: "impl-07",
|
||||
Directory: "/e4/implementations/impl-07",
|
||||
Port: 5306,
|
||||
}
|
||||
for _, seed := range []string{"1", "42"} {
|
||||
arguments := campaignArguments(configuration, target, seed)
|
||||
if got := argumentValue(arguments, "--seeds"); got != seed {
|
||||
t.Errorf("--seeds: got %q, want %q", got, seed)
|
||||
}
|
||||
bundle := argumentValue(arguments, "--bundle-id")
|
||||
parsed, err := url.Parse(bundle)
|
||||
if err != nil {
|
||||
t.Fatalf("--bundle-id %q: %v", bundle, err)
|
||||
}
|
||||
if got := parsed.Query().Get("seed"); got != seed {
|
||||
t.Errorf(
|
||||
"served URL seed: got %q, want %q (from %q)",
|
||||
got,
|
||||
seed,
|
||||
bundle,
|
||||
)
|
||||
}
|
||||
if parsed.Host != "localhost:"+strconv.Itoa(target.Port) {
|
||||
t.Errorf(
|
||||
"served host: got %q, want the implementation's own port %d",
|
||||
parsed.Host,
|
||||
target.Port,
|
||||
)
|
||||
}
|
||||
for flagName, want := range map[string]string{
|
||||
"--arm": "impl-07",
|
||||
"--platform": "web",
|
||||
"--generator": "seeded",
|
||||
"--max-steps": "400",
|
||||
"--spec": "/e4/relay.ts",
|
||||
"--output": filepath.Join("/campaigns/e4", "impl-07", "seed-"+seed),
|
||||
} {
|
||||
if got := argumentValue(arguments, flagName); got != want {
|
||||
t.Errorf("%s: got %q, want %q", flagName, got, want)
|
||||
}
|
||||
}
|
||||
if arguments[len(arguments)-2] != "--" ||
|
||||
arguments[len(arguments)-1] != "--clear-data=false" {
|
||||
t.Errorf("passthrough flags lost: %v", arguments)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func argumentValue(arguments []string, name string) string {
|
||||
index := slices.Index(arguments, name)
|
||||
if index < 0 || index+1 >= len(arguments) {
|
||||
return ""
|
||||
}
|
||||
return arguments[index+1]
|
||||
}
|
||||
@@ -0,0 +1,117 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"time"
|
||||
)
|
||||
|
||||
const (
|
||||
manifestFileName = "sweep.json"
|
||||
recordsFileName = "implementations.jsonl"
|
||||
seedPlaceholder = "{seed}"
|
||||
)
|
||||
|
||||
// plannedImplementation is one implementation the sweep intends to run, with
|
||||
// the port it is served on and the URL every seed is served at.
|
||||
type plannedImplementation struct {
|
||||
Name string `json:"name"`
|
||||
Directory string `json:"directory"`
|
||||
Port int `json:"port"`
|
||||
URLTemplate string `json:"url_template"`
|
||||
}
|
||||
|
||||
// manifest is sweep.json: what the sweep INTENDED to run, written before the
|
||||
// first install so a host that dropped an implementation or a seed shows up as
|
||||
// a missing run rather than as a smaller sample.
|
||||
type manifest struct {
|
||||
Generator string `json:"generator"`
|
||||
Platform string `json:"platform"`
|
||||
SpecPath string `json:"spec_path"`
|
||||
MaxSteps int `json:"max_steps"`
|
||||
DurationMillis int64 `json:"duration_millis"`
|
||||
Seeds []int64 `json:"seeds"`
|
||||
Implementations []plannedImplementation `json:"implementations"`
|
||||
Concurrency int `json:"concurrency"`
|
||||
Host string `json:"host"`
|
||||
BunPath string `json:"bun_path"`
|
||||
CampaignPath string `json:"campaign_path"`
|
||||
SanderlingPath string `json:"sanderling_path"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
}
|
||||
|
||||
func buildManifest(
|
||||
configuration config,
|
||||
implementations []implementation,
|
||||
host string,
|
||||
startedAt time.Time,
|
||||
) manifest {
|
||||
planned := make([]plannedImplementation, 0, len(implementations))
|
||||
for _, target := range implementations {
|
||||
planned = append(planned, plannedImplementation{
|
||||
Name: target.Name,
|
||||
Directory: target.Directory,
|
||||
Port: target.Port,
|
||||
URLTemplate: servedURL(target.Port, seedPlaceholder),
|
||||
})
|
||||
}
|
||||
return manifest{
|
||||
Generator: generator,
|
||||
Platform: platform,
|
||||
SpecPath: configuration.specPath,
|
||||
MaxSteps: configuration.maxSteps,
|
||||
DurationMillis: configuration.duration.Milliseconds(),
|
||||
Seeds: configuration.seeds,
|
||||
Implementations: planned,
|
||||
Concurrency: configuration.concurrency,
|
||||
Host: host,
|
||||
BunPath: configuration.bunPath,
|
||||
CampaignPath: configuration.campaignPath,
|
||||
SanderlingPath: configuration.sanderlingPath,
|
||||
StartedAt: startedAt,
|
||||
}
|
||||
}
|
||||
|
||||
func writeManifest(directory string, value manifest) error {
|
||||
body, err := json.MarshalIndent(value, "", " ")
|
||||
if err != nil {
|
||||
return fmt.Errorf("marshal manifest: %w", err)
|
||||
}
|
||||
return os.WriteFile(
|
||||
filepath.Join(directory, manifestFileName),
|
||||
append(body, '\n'),
|
||||
0o644,
|
||||
)
|
||||
}
|
||||
|
||||
// runRecord is one campaign, which is one implementation at one seed.
|
||||
type runRecord struct {
|
||||
Seed int64 `json:"seed"`
|
||||
URL string `json:"url"`
|
||||
ExitCode int `json:"exit_code"`
|
||||
LaunchError string `json:"launch_error,omitempty"`
|
||||
CampaignDirectory string `json:"campaign_directory"`
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
}
|
||||
|
||||
// implementationRecord is one line of implementations.jsonl. FailedStage names
|
||||
// the step that stopped this implementation, and an implementation that never
|
||||
// got past install, build or serve carries no runs at all.
|
||||
type implementationRecord struct {
|
||||
Name string `json:"implementation"`
|
||||
Directory string `json:"directory"`
|
||||
Port int `json:"port"`
|
||||
FailedStage string `json:"failed_stage,omitempty"`
|
||||
Error string `json:"error,omitempty"`
|
||||
StartedAt time.Time `json:"started_at"`
|
||||
MonotonicMillis int64 `json:"monotonic_millis"`
|
||||
Runs []runRecord `json:"runs"`
|
||||
}
|
||||
|
||||
const (
|
||||
stageInstall = "install"
|
||||
stageBuild = "build"
|
||||
stageServe = "serve"
|
||||
)
|
||||
@@ -0,0 +1,109 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"os"
|
||||
"os/exec"
|
||||
"syscall"
|
||||
"time"
|
||||
)
|
||||
|
||||
const (
|
||||
// serverStartTimeout covers a cold vite start on a host already running
|
||||
// five other implementations.
|
||||
serverStartTimeout = 90 * time.Second
|
||||
serverPollInterval = 250 * time.Millisecond
|
||||
serverShutdownGrace = 10 * time.Second
|
||||
)
|
||||
|
||||
// server is one implementation's preview server. It runs in its own process
|
||||
// group so that stopping it takes the whole vite tree with it: a leaked server
|
||||
// holds its port, and the next sweep against that implementation would be
|
||||
// served by the previous build.
|
||||
type server struct {
|
||||
command *exec.Cmd
|
||||
logFile *os.File
|
||||
exited chan struct{}
|
||||
}
|
||||
|
||||
func startServer(
|
||||
ctx context.Context,
|
||||
configuration config,
|
||||
target implementation,
|
||||
logPath string,
|
||||
) (*server, error) {
|
||||
logFile, err := os.Create(logPath)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
command := exec.CommandContext(ctx, configuration.bunPath,
|
||||
"run", "preview", "--port", fmt.Sprint(target.Port), "--strictPort")
|
||||
command.Dir = target.Directory
|
||||
command.Stdout = logFile
|
||||
command.Stderr = logFile
|
||||
command.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
|
||||
if err := command.Start(); err != nil {
|
||||
logFile.Close()
|
||||
return nil, err
|
||||
}
|
||||
running := &server{
|
||||
command: command,
|
||||
logFile: logFile,
|
||||
exited: make(chan struct{}),
|
||||
}
|
||||
go func() {
|
||||
command.Wait()
|
||||
close(running.exited)
|
||||
}()
|
||||
return running, nil
|
||||
}
|
||||
|
||||
// waitReady polls the served page until it answers. A server that exits first
|
||||
// is reported as such rather than waited on for the full timeout, because the
|
||||
// usual cause is a port already taken and that answer is in the log.
|
||||
func (s *server) waitReady(ctx context.Context, url string) error {
|
||||
client := &http.Client{Timeout: 5 * time.Second}
|
||||
deadline := time.Now().Add(serverStartTimeout)
|
||||
for {
|
||||
select {
|
||||
case <-s.exited:
|
||||
return fmt.Errorf("server exited before it answered %s", url)
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
default:
|
||||
}
|
||||
response, err := client.Get(url)
|
||||
if err == nil {
|
||||
io.Copy(io.Discard, response.Body)
|
||||
response.Body.Close()
|
||||
if response.StatusCode == http.StatusOK {
|
||||
return nil
|
||||
}
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
return fmt.Errorf(
|
||||
"server did not answer %s within %s",
|
||||
url,
|
||||
serverStartTimeout,
|
||||
)
|
||||
}
|
||||
time.Sleep(serverPollInterval)
|
||||
}
|
||||
}
|
||||
|
||||
func (s *server) stop() {
|
||||
if s.command.Process != nil {
|
||||
group := -s.command.Process.Pid
|
||||
syscall.Kill(group, syscall.SIGTERM)
|
||||
select {
|
||||
case <-s.exited:
|
||||
case <-time.After(serverShutdownGrace):
|
||||
syscall.Kill(group, syscall.SIGKILL)
|
||||
<-s.exited
|
||||
}
|
||||
}
|
||||
s.logFile.Close()
|
||||
}
|
||||
@@ -0,0 +1,119 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"syscall"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// bunSpawningAServer serves from a child process and then waits, which is the
|
||||
// shape of `bun run preview`: the port belongs to something below the process
|
||||
// the sweep started.
|
||||
const bunSpawningAServer = `#!/bin/sh
|
||||
port=""
|
||||
previous=""
|
||||
for argument in "$@"; do
|
||||
if [ "$previous" = "--port" ]; then port="$argument"; fi
|
||||
previous="$argument"
|
||||
done
|
||||
SWEEP_TEST_SERVE_PORT="$port" "%[1]s" &
|
||||
wait
|
||||
`
|
||||
|
||||
func TestServerStop_TakesTheProcessBelowItWithTheServer(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
testBinary, err := filepath.Abs(os.Args[0])
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
bunPath := writeScript(
|
||||
t,
|
||||
filepath.Join(directory, "stub-bun"),
|
||||
fmt.Sprintf(bunSpawningAServer, testBinary),
|
||||
)
|
||||
port := freePortRange(t, 1)
|
||||
target := implementation{Name: "impl-01", Directory: directory, Port: port}
|
||||
|
||||
// Background rather than the test context: only stop() may end this
|
||||
// server, or a leak would be hidden by the context being cancelled.
|
||||
running, err := startServer(
|
||||
context.Background(),
|
||||
config{bunPath: bunPath},
|
||||
target,
|
||||
filepath.Join(directory, "serve.log"),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(
|
||||
func() { syscall.Kill(-running.command.Process.Pid, syscall.SIGKILL) },
|
||||
)
|
||||
if err := running.waitReady(context.Background(), readinessURL(port)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
running.stop()
|
||||
|
||||
client := &http.Client{Timeout: time.Second}
|
||||
deadline := time.Now().Add(5 * time.Second)
|
||||
for time.Now().Before(deadline) {
|
||||
response, err := client.Get(readinessURL(port))
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
response.Body.Close()
|
||||
time.Sleep(100 * time.Millisecond)
|
||||
}
|
||||
t.Fatalf(
|
||||
"port %d is still served after stop(): the server below bun outlived the sweep and holds the port",
|
||||
port,
|
||||
)
|
||||
}
|
||||
|
||||
func TestServerWaitReady_ReportsAServerThatExited(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
bunPath := writeScript(
|
||||
t,
|
||||
filepath.Join(directory, "stub-bun"),
|
||||
"#!/bin/sh\necho 'port is already in use' >&2\nexit 1\n",
|
||||
)
|
||||
port := freePortRange(t, 1)
|
||||
|
||||
running, err := startServer(
|
||||
context.Background(),
|
||||
config{bunPath: bunPath},
|
||||
implementation{
|
||||
Name: "impl-01",
|
||||
Directory: directory,
|
||||
Port: port,
|
||||
},
|
||||
filepath.Join(directory, "serve.log"),
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer running.stop()
|
||||
|
||||
started := time.Now()
|
||||
err = running.waitReady(context.Background(), readinessURL(port))
|
||||
if err == nil {
|
||||
t.Fatal(
|
||||
"a server that exited should not be waited on until the start timeout",
|
||||
)
|
||||
}
|
||||
if elapsed := time.Since(started); elapsed > 30*time.Second {
|
||||
t.Errorf("waited %s for a server that had already exited", elapsed)
|
||||
}
|
||||
log, err := os.ReadFile(filepath.Join(directory, "serve.log"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if string(log) == "" {
|
||||
t.Error("serve.log did not capture why the server exited")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,399 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
)
|
||||
|
||||
// implementation is one directory under --implementations, with the port it
|
||||
// owns for the whole sweep. The port comes from the implementation's position
|
||||
// in name order rather than from a pool, so the manifest can name the URL every
|
||||
// run was served from before anything has been served.
|
||||
type implementation struct {
|
||||
Name string
|
||||
Directory string
|
||||
Port int
|
||||
}
|
||||
|
||||
const implementationPrefix = "impl-"
|
||||
|
||||
func discoverImplementations(
|
||||
directory string,
|
||||
basePort int,
|
||||
) ([]implementation, error) {
|
||||
entries, err := os.ReadDir(directory)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var found []implementation
|
||||
for _, entry := range entries {
|
||||
if !entry.IsDir() ||
|
||||
!strings.HasPrefix(entry.Name(), implementationPrefix) {
|
||||
continue
|
||||
}
|
||||
found = append(found, implementation{
|
||||
Name: entry.Name(),
|
||||
Directory: filepath.Join(directory, entry.Name()),
|
||||
})
|
||||
}
|
||||
if len(found) == 0 {
|
||||
return nil, fmt.Errorf(
|
||||
"no %s* directories in %s",
|
||||
implementationPrefix,
|
||||
directory,
|
||||
)
|
||||
}
|
||||
sort.Slice(
|
||||
found,
|
||||
func(i, j int) bool { return found[i].Name < found[j].Name },
|
||||
)
|
||||
if basePort+len(found)-1 > 65535 {
|
||||
return nil, fmt.Errorf(
|
||||
"--base-port %d leaves no room for %d implementations",
|
||||
basePort,
|
||||
len(found),
|
||||
)
|
||||
}
|
||||
for index := range found {
|
||||
found[index].Port = basePort + index
|
||||
}
|
||||
return found, nil
|
||||
}
|
||||
|
||||
// resolveBinaries turns bun, campaign and sanderling into absolute paths before
|
||||
// anything is installed. Each campaign runs from the sweep's own directory
|
||||
// rather than the implementation's, so a relative --sanderling would otherwise
|
||||
// resolve against the wrong one, and a binary that is missing altogether has to
|
||||
// stop the sweep here rather than fail once per implementation and seed. Every
|
||||
// one that is missing is named together, in flag order: stopping at the first
|
||||
// turns that single stop into one rerun per missing binary.
|
||||
func resolveBinaries(configuration *config) error {
|
||||
var missing []error
|
||||
for _, binary := range []struct {
|
||||
name string
|
||||
value *string
|
||||
}{
|
||||
{"--bun", &configuration.bunPath},
|
||||
{"--campaign", &configuration.campaignPath},
|
||||
{"--sanderling", &configuration.sanderlingPath},
|
||||
} {
|
||||
resolved, err := exec.LookPath(*binary.value)
|
||||
if err != nil {
|
||||
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
|
||||
continue
|
||||
}
|
||||
absolute, err := filepath.Abs(resolved)
|
||||
if err != nil {
|
||||
missing = append(missing, fmt.Errorf("%s: %w", binary.name, err))
|
||||
continue
|
||||
}
|
||||
*binary.value = absolute
|
||||
}
|
||||
return errors.Join(missing...)
|
||||
}
|
||||
|
||||
type sweep struct {
|
||||
configuration config
|
||||
stdout io.Writer
|
||||
records io.Writer
|
||||
mutex sync.Mutex
|
||||
stalled int
|
||||
failedRuns int
|
||||
totalRuns int
|
||||
}
|
||||
|
||||
func runSweep(
|
||||
ctx context.Context,
|
||||
configuration config,
|
||||
stdout io.Writer,
|
||||
) error {
|
||||
if _, err := os.Stat(filepath.Join(configuration.outputDirectory, manifestFileName)); err == nil {
|
||||
return fmt.Errorf(
|
||||
"%s already exists in %s: pick a fresh --output so two sweeps do not share a directory",
|
||||
manifestFileName,
|
||||
configuration.outputDirectory,
|
||||
)
|
||||
}
|
||||
implementations, err := discoverImplementations(
|
||||
configuration.implementationsDirectory,
|
||||
configuration.basePort,
|
||||
)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := resolveBinaries(&configuration); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := os.Stat(configuration.specPath); err != nil {
|
||||
return fmt.Errorf("--spec: %w", err)
|
||||
}
|
||||
if err := os.MkdirAll(configuration.outputDirectory, 0o755); err != nil {
|
||||
return fmt.Errorf("create sweep dir: %w", err)
|
||||
}
|
||||
host, _ := os.Hostname()
|
||||
if err := writeManifest(configuration.outputDirectory, buildManifest(configuration, implementations, host, time.Now().UTC())); err != nil {
|
||||
return fmt.Errorf("write %s: %w", manifestFileName, err)
|
||||
}
|
||||
|
||||
recordsFile, err := os.OpenFile(
|
||||
filepath.Join(configuration.outputDirectory, recordsFileName),
|
||||
os.O_CREATE|os.O_WRONLY|os.O_APPEND,
|
||||
0o644,
|
||||
)
|
||||
if err != nil {
|
||||
return fmt.Errorf("open %s: %w", recordsFileName, err)
|
||||
}
|
||||
defer recordsFile.Close()
|
||||
|
||||
running := &sweep{
|
||||
configuration: configuration,
|
||||
stdout: stdout,
|
||||
records: recordsFile,
|
||||
}
|
||||
fmt.Fprintf(
|
||||
stdout,
|
||||
"sweep: %d implementations, %d seeds each, %d at a time, %s\n",
|
||||
len(
|
||||
implementations,
|
||||
),
|
||||
len(configuration.seeds),
|
||||
configuration.concurrency,
|
||||
configuration.outputDirectory,
|
||||
)
|
||||
running.work(ctx, implementations)
|
||||
|
||||
fmt.Fprintf(
|
||||
stdout,
|
||||
"sweep complete: %d of %d implementations never ran, %d of %d campaigns failed\n",
|
||||
running.stalled,
|
||||
len(implementations),
|
||||
running.failedRuns,
|
||||
running.totalRuns,
|
||||
)
|
||||
if running.stalled > 0 || running.failedRuns > 0 {
|
||||
return fmt.Errorf(
|
||||
"%d of %d implementations never ran and %d of %d campaigns failed; see %s",
|
||||
running.stalled,
|
||||
len(implementations),
|
||||
running.failedRuns,
|
||||
running.totalRuns,
|
||||
recordsFileName,
|
||||
)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (s *sweep) work(ctx context.Context, implementations []implementation) {
|
||||
queue := make(chan implementation, len(implementations))
|
||||
for _, target := range implementations {
|
||||
queue <- target
|
||||
}
|
||||
close(queue)
|
||||
|
||||
workers := min(s.configuration.concurrency, len(implementations))
|
||||
var waitGroup sync.WaitGroup
|
||||
for range workers {
|
||||
waitGroup.Add(1)
|
||||
go func() {
|
||||
defer waitGroup.Done()
|
||||
for target := range queue {
|
||||
if ctx.Err() != nil {
|
||||
return
|
||||
}
|
||||
s.report(s.runImplementation(ctx, target))
|
||||
}
|
||||
}()
|
||||
}
|
||||
waitGroup.Wait()
|
||||
}
|
||||
|
||||
// runImplementation carries one implementation from install to its last seed.
|
||||
// Every failure it can meet is returned in the record: one implementation that
|
||||
// cannot install, build or serve must not cost the other twenty-three their
|
||||
// runs.
|
||||
func (s *sweep) runImplementation(
|
||||
ctx context.Context,
|
||||
target implementation,
|
||||
) (record implementationRecord) {
|
||||
record = implementationRecord{
|
||||
Name: target.Name,
|
||||
Directory: target.Directory,
|
||||
Port: target.Port,
|
||||
StartedAt: time.Now().UTC(),
|
||||
}
|
||||
started := time.Now()
|
||||
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
|
||||
|
||||
directory := filepath.Join(s.configuration.outputDirectory, target.Name)
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
record.FailedStage = stageInstall
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
for _, step := range []struct {
|
||||
stage string
|
||||
arguments []string
|
||||
}{
|
||||
{stageInstall, []string{"install"}},
|
||||
{stageBuild, []string{"run", "build"}},
|
||||
} {
|
||||
logPath := filepath.Join(directory, step.stage+".log")
|
||||
exitCode, err := runCommand(
|
||||
ctx,
|
||||
target.Directory,
|
||||
s.configuration.bunPath,
|
||||
step.arguments,
|
||||
logPath,
|
||||
)
|
||||
if err != nil {
|
||||
record.FailedStage = step.stage
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
if exitCode != 0 {
|
||||
record.FailedStage = step.stage
|
||||
record.Error = fmt.Sprintf(
|
||||
"bun %s exited %d, see %s",
|
||||
strings.Join(step.arguments, " "),
|
||||
exitCode,
|
||||
logPath,
|
||||
)
|
||||
return record
|
||||
}
|
||||
}
|
||||
|
||||
running, err := startServer(
|
||||
ctx,
|
||||
s.configuration,
|
||||
target,
|
||||
filepath.Join(directory, "serve.log"),
|
||||
)
|
||||
if err != nil {
|
||||
record.FailedStage = stageServe
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
defer running.stop()
|
||||
if err := running.waitReady(ctx, readinessURL(target.Port)); err != nil {
|
||||
record.FailedStage = stageServe
|
||||
record.Error = err.Error()
|
||||
return record
|
||||
}
|
||||
|
||||
for _, seed := range s.configuration.seeds {
|
||||
if ctx.Err() != nil {
|
||||
return record
|
||||
}
|
||||
record.Runs = append(record.Runs, s.runSeed(ctx, target, seed))
|
||||
}
|
||||
return record
|
||||
}
|
||||
|
||||
func (s *sweep) runSeed(
|
||||
ctx context.Context,
|
||||
target implementation,
|
||||
seed int64,
|
||||
) (record runRecord) {
|
||||
seedText := strconv.FormatInt(seed, 10)
|
||||
directory := campaignDirectory(s.configuration, target, seedText)
|
||||
record = runRecord{
|
||||
Seed: seed,
|
||||
URL: servedURL(target.Port, seedText),
|
||||
CampaignDirectory: directory,
|
||||
}
|
||||
started := time.Now()
|
||||
defer func() { record.MonotonicMillis = time.Since(started).Milliseconds() }()
|
||||
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
record.ExitCode = -1
|
||||
record.LaunchError = err.Error()
|
||||
return record
|
||||
}
|
||||
exitCode, err := runCommand(
|
||||
ctx,
|
||||
"",
|
||||
s.configuration.campaignPath,
|
||||
campaignArguments(
|
||||
s.configuration,
|
||||
target,
|
||||
seedText,
|
||||
),
|
||||
filepath.Join(directory, "campaign.log"),
|
||||
)
|
||||
record.ExitCode = exitCode
|
||||
if err != nil {
|
||||
record.LaunchError = err.Error()
|
||||
}
|
||||
return record
|
||||
}
|
||||
|
||||
// runCommand runs one step of the pipeline with its output in logPath. An
|
||||
// empty directory keeps the sweep's own working directory, which is what the
|
||||
// campaign tool gets: it has no reason to run inside an implementation.
|
||||
func runCommand(
|
||||
ctx context.Context,
|
||||
directory, binary string,
|
||||
arguments []string,
|
||||
logPath string,
|
||||
) (int, error) {
|
||||
logFile, err := os.Create(logPath)
|
||||
if err != nil {
|
||||
return -1, err
|
||||
}
|
||||
defer logFile.Close()
|
||||
command := exec.CommandContext(ctx, binary, arguments...)
|
||||
command.Dir = directory
|
||||
command.Stdout = logFile
|
||||
command.Stderr = logFile
|
||||
err = command.Run()
|
||||
if err == nil {
|
||||
return 0, nil
|
||||
}
|
||||
var exitError *exec.ExitError
|
||||
if errors.As(err, &exitError) {
|
||||
return exitError.ExitCode(), nil
|
||||
}
|
||||
return -1, err
|
||||
}
|
||||
|
||||
func (s *sweep) report(record implementationRecord) {
|
||||
s.mutex.Lock()
|
||||
defer s.mutex.Unlock()
|
||||
if record.FailedStage != "" {
|
||||
s.stalled++
|
||||
}
|
||||
s.totalRuns += len(record.Runs)
|
||||
for _, run := range record.Runs {
|
||||
if run.ExitCode != 0 {
|
||||
s.failedRuns++
|
||||
}
|
||||
}
|
||||
if err := json.NewEncoder(s.records).Encode(record); err != nil {
|
||||
fmt.Fprintf(s.stdout, "warning: %s record: %v\n", record.Name, err)
|
||||
}
|
||||
elapsed := time.Duration(record.MonotonicMillis) * time.Millisecond
|
||||
if record.FailedStage != "" {
|
||||
fmt.Fprintf(s.stdout, "%s port=%d failed at %s: %s (%s)\n",
|
||||
record.Name, record.Port, record.FailedStage, record.Error, elapsed)
|
||||
return
|
||||
}
|
||||
failed := 0
|
||||
for _, run := range record.Runs {
|
||||
if run.ExitCode != 0 {
|
||||
failed++
|
||||
}
|
||||
}
|
||||
fmt.Fprintf(s.stdout, "%s port=%d campaigns=%d failed=%d elapsed=%s\n",
|
||||
record.Name, record.Port, len(record.Runs), failed, elapsed)
|
||||
}
|
||||
@@ -0,0 +1,154 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestDiscoverImplementations_NameOrderAndOnePortEach(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
for _, name := range []string{"impl-03", "impl-01", "impl-10", "impl-02", "scaffold", ".DS_Store"} {
|
||||
if err := os.MkdirAll(filepath.Join(directory, name), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(directory, "impl-notes.md"), []byte("not a directory"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
found, err := discoverImplementations(directory, 5300)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
want := []implementation{
|
||||
{Name: "impl-01", Port: 5300},
|
||||
{Name: "impl-02", Port: 5301},
|
||||
{Name: "impl-03", Port: 5302},
|
||||
{Name: "impl-10", Port: 5303},
|
||||
}
|
||||
if len(found) != len(want) {
|
||||
t.Fatalf(
|
||||
"got %d implementations, want %d: %v",
|
||||
len(found),
|
||||
len(want),
|
||||
found,
|
||||
)
|
||||
}
|
||||
for index, target := range found {
|
||||
if target.Name != want[index].Name || target.Port != want[index].Port {
|
||||
t.Errorf(
|
||||
"position %d: got %s on %d, want %s on %d",
|
||||
index,
|
||||
target.Name,
|
||||
target.Port,
|
||||
want[index].Name,
|
||||
want[index].Port,
|
||||
)
|
||||
}
|
||||
if target.Directory != filepath.Join(directory, want[index].Name) {
|
||||
t.Errorf("%s directory: got %q", target.Name, target.Directory)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestDiscoverImplementations_EmptyDirectoryIsRefused(t *testing.T) {
|
||||
_, err := discoverImplementations(t.TempDir(), 5300)
|
||||
if err == nil || !strings.Contains(err.Error(), "no impl-* directories") {
|
||||
t.Fatalf("got %v, want a refusal naming impl-*", err)
|
||||
}
|
||||
}
|
||||
|
||||
// A binary that is not there fails once, before anything is installed, rather
|
||||
// than twenty-four times after the sweep has spent its build time.
|
||||
func TestRunSweep_StopsBeforeItInstallsAnythingWhenABinaryIsMissing(
|
||||
t *testing.T,
|
||||
) {
|
||||
root := t.TempDir()
|
||||
implementations := filepath.Join(root, "implementations")
|
||||
if err := os.MkdirAll(filepath.Join(implementations, "impl-01"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
output := filepath.Join(root, "campaigns")
|
||||
configuration := config{
|
||||
implementationsDirectory: implementations,
|
||||
outputDirectory: output,
|
||||
basePort: 5300,
|
||||
concurrency: 1,
|
||||
bunPath: writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-bun"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
),
|
||||
campaignPath: "campaign-that-is-not-installed",
|
||||
sanderlingPath: writeScript(
|
||||
t,
|
||||
filepath.Join(root, "stub-sanderling"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
),
|
||||
}
|
||||
err := runSweep(t.Context(), configuration, os.Stdout)
|
||||
if err == nil || !strings.Contains(err.Error(), "--campaign") {
|
||||
t.Fatalf("got %v, want the missing campaign binary named", err)
|
||||
}
|
||||
if _, err := os.Stat(output); !os.IsNotExist(err) {
|
||||
t.Errorf(
|
||||
"the sweep created %s before it checked it could run: %v",
|
||||
output,
|
||||
err,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// Two binaries missing is one rerun, not two: the operator is told about both
|
||||
// at once, in flag order, whatever order the check happened to walk.
|
||||
func TestResolveBinaries_NamesEveryMissingBinaryInFlagOrder(t *testing.T) {
|
||||
configuration := config{
|
||||
bunPath: writeScript(
|
||||
t,
|
||||
filepath.Join(t.TempDir(), "stub-bun"),
|
||||
"#!/bin/sh\nexit 0\n",
|
||||
),
|
||||
campaignPath: "campaign-that-is-not-installed",
|
||||
sanderlingPath: "sanderling-that-is-not-installed",
|
||||
}
|
||||
|
||||
err := resolveBinaries(&configuration)
|
||||
if err == nil {
|
||||
t.Fatal("got no error, want both missing binaries named")
|
||||
}
|
||||
message := err.Error()
|
||||
campaign := strings.Index(message, "--campaign")
|
||||
sanderling := strings.Index(message, "--sanderling")
|
||||
if campaign < 0 || sanderling < 0 {
|
||||
t.Fatalf("got %q, want both --campaign and --sanderling named", message)
|
||||
}
|
||||
if campaign > sanderling {
|
||||
t.Errorf("got %q, want --campaign named before --sanderling", message)
|
||||
}
|
||||
if strings.Contains(message, "--bun") {
|
||||
t.Errorf("got %q, want the bun that resolved left out", message)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunSweep_RefusesADirectoryThatAlreadyHoldsASweep(t *testing.T) {
|
||||
implementations := t.TempDir()
|
||||
if err := os.MkdirAll(filepath.Join(implementations, "impl-01"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
output := t.TempDir()
|
||||
if err := os.WriteFile(filepath.Join(output, manifestFileName), []byte("{}"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
configuration := config{
|
||||
implementationsDirectory: implementations,
|
||||
outputDirectory: output,
|
||||
basePort: 5300,
|
||||
concurrency: 1,
|
||||
}
|
||||
if err := runSweep(t.Context(), configuration, os.Stdout); err == nil ||
|
||||
!strings.Contains(err.Error(), "already exists") {
|
||||
t.Fatalf("got %v, want a refusal to reuse the directory", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,549 @@
|
||||
// Command label-coverage reports how much of an app's interactive surface a
|
||||
// spec can address, from the hierarchies a run already recorded.
|
||||
package main
|
||||
|
||||
import (
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"io/fs"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
type element struct {
|
||||
ResourceID string `json:"resourceId"`
|
||||
Text string `json:"text"`
|
||||
Description string `json:"description"`
|
||||
Class string `json:"class"`
|
||||
Package string `json:"package"`
|
||||
Clickable bool `json:"clickable"`
|
||||
Editable bool `json:"editable"`
|
||||
Enabled bool `json:"enabled"`
|
||||
}
|
||||
|
||||
type step struct {
|
||||
Index int `json:"step"`
|
||||
Screen string `json:"screen"`
|
||||
Hierarchy *struct {
|
||||
Elements []element `json:"elements"`
|
||||
} `json:"hierarchy"`
|
||||
}
|
||||
|
||||
// Counts splits a screen's interactive elements by the strongest selector that
|
||||
// can reach them. Text is separated from identifier and description because a
|
||||
// row labelled only by the customer name it displays is addressable in one run
|
||||
// and gone in the next, which is not the same thing as being addressable. A
|
||||
// description that carries the data with it is separated for the same reason.
|
||||
type Counts struct {
|
||||
Screen string `json:"screen"`
|
||||
Observations int `json:"observations"`
|
||||
Elements int `json:"elements"`
|
||||
Interactive int `json:"interactive"`
|
||||
ByIdentifier int `json:"by_identifier"`
|
||||
ByDataID int `json:"by_data_carrying_identifier"`
|
||||
ByDescription int `json:"by_description"`
|
||||
ByVolatile int `json:"by_volatile_description"`
|
||||
ByTextOnly int `json:"by_text_only"`
|
||||
Unaddressable int `json:"unaddressable"`
|
||||
AmbiguousIDs int `json:"ambiguous_identifiers"`
|
||||
// Needing lists the interactive elements no durable selector reaches, which
|
||||
// is the work list for a label pass rather than a statistic about it.
|
||||
Needing []string `json:"needing_labels,omitempty"`
|
||||
}
|
||||
|
||||
func (c Counts) stableShare() float64 {
|
||||
if c.Interactive == 0 {
|
||||
return 0
|
||||
}
|
||||
return float64(
|
||||
c.ByIdentifier+c.ByDescription,
|
||||
) / float64(
|
||||
c.Interactive,
|
||||
) * 100
|
||||
}
|
||||
|
||||
func main() {
|
||||
jsonOut := flag.Bool("json", false, "emit JSON instead of a table")
|
||||
show := flag.Int(
|
||||
"show",
|
||||
0,
|
||||
"list up to this many controls per screen that no durable selector reaches",
|
||||
)
|
||||
pkg := flag.String(
|
||||
"package",
|
||||
"",
|
||||
"count only elements belonging to this package; system UI is dropped either way",
|
||||
)
|
||||
flag.Usage = func() {
|
||||
fmt.Fprintln(
|
||||
os.Stderr,
|
||||
"usage: label-coverage [--json] [--show N] <trace.jsonl | run directory> ...",
|
||||
)
|
||||
}
|
||||
flag.Parse()
|
||||
if flag.NArg() == 0 {
|
||||
flag.Usage()
|
||||
os.Exit(2)
|
||||
}
|
||||
|
||||
traces, err := collect(flag.Args())
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
if len(traces) == 0 {
|
||||
fmt.Fprintln(os.Stderr, "no trace.jsonl found under the given paths")
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
byScreen := map[string]*Counts{}
|
||||
for _, path := range traces {
|
||||
if err := accumulate(path, byScreen, *pkg); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
screens := make([]Counts, 0, len(byScreen))
|
||||
for _, counts := range byScreen {
|
||||
screens = append(screens, *counts)
|
||||
}
|
||||
sort.Slice(
|
||||
screens,
|
||||
func(i, j int) bool { return screens[i].Screen < screens[j].Screen },
|
||||
)
|
||||
|
||||
if *jsonOut {
|
||||
encoder := json.NewEncoder(os.Stdout)
|
||||
encoder.SetIndent("", " ")
|
||||
if err := encoder.Encode(screens); err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
return
|
||||
}
|
||||
render(os.Stdout, screens, len(traces), *show)
|
||||
}
|
||||
|
||||
func collect(paths []string) ([]string, error) {
|
||||
var traces []string
|
||||
for _, path := range paths {
|
||||
info, err := os.Stat(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if !info.IsDir() {
|
||||
traces = append(traces, path)
|
||||
continue
|
||||
}
|
||||
err = filepath.WalkDir(
|
||||
path,
|
||||
func(candidate string, entry fs.DirEntry, err error) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if !entry.IsDir() && entry.Name() == "trace.jsonl" {
|
||||
traces = append(traces, candidate)
|
||||
}
|
||||
return nil
|
||||
},
|
||||
)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
sort.Strings(traces)
|
||||
return traces, nil
|
||||
}
|
||||
|
||||
// accumulate folds one trace into byScreen, counting each distinct hierarchy
|
||||
// once. A run that idles on a screen observes it many times, and summing those
|
||||
// observations would report the screen the explorer sat on rather than the
|
||||
// screen with the most unlabelled controls.
|
||||
// systemPackages own nodes that share the screen with the app under test. A
|
||||
// status bar contributes three well-labelled controls to every capture, and
|
||||
// counting them lifts an app with no identifiers at all off the floor.
|
||||
var systemPackages = map[string]bool{
|
||||
"com.android.systemui": true,
|
||||
"android": true,
|
||||
"com.google.android.inputmethod.latin": true,
|
||||
"com.android.inputmethod.latin": true,
|
||||
"com.google.android.apps.nexuslauncher": true,
|
||||
"com.google.android.googlequicksearchbox": true,
|
||||
}
|
||||
|
||||
// scoped drops the windows the app under test does not own. Most of an app's
|
||||
// own nodes carry no package attribute at all, since only the window roots are
|
||||
// stamped with one, so an empty package is treated as belonging to the app
|
||||
// rather than filtered out. Filtering on an exact match alone discards the
|
||||
// entire application and reports a clean zero.
|
||||
func scoped(elements []element, pkg string) []element {
|
||||
kept := make([]element, 0, len(elements))
|
||||
for _, item := range elements {
|
||||
if item.Package == "" {
|
||||
kept = append(kept, item)
|
||||
continue
|
||||
}
|
||||
if pkg != "" {
|
||||
if item.Package == pkg {
|
||||
kept = append(kept, item)
|
||||
}
|
||||
continue
|
||||
}
|
||||
if !systemPackages[item.Package] {
|
||||
kept = append(kept, item)
|
||||
}
|
||||
}
|
||||
return kept
|
||||
}
|
||||
|
||||
func accumulate(path string, byScreen map[string]*Counts, pkg string) error {
|
||||
file, err := os.Open(path)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
seen := map[string]bool{}
|
||||
decoder := json.NewDecoder(file)
|
||||
for {
|
||||
var recorded step
|
||||
if err := decoder.Decode(&recorded); err != nil {
|
||||
if errors.Is(err, io.EOF) {
|
||||
break
|
||||
}
|
||||
return err
|
||||
}
|
||||
if recorded.Hierarchy == nil || len(recorded.Hierarchy.Elements) == 0 {
|
||||
continue
|
||||
}
|
||||
elements := scoped(recorded.Hierarchy.Elements, pkg)
|
||||
if len(elements) == 0 {
|
||||
continue
|
||||
}
|
||||
screen := recorded.Screen
|
||||
if screen == "" {
|
||||
screen = shape(elements)
|
||||
}
|
||||
key := screen + "\x00" + signature(elements)
|
||||
if seen[key] {
|
||||
continue
|
||||
}
|
||||
seen[key] = true
|
||||
|
||||
observed := measure(elements)
|
||||
observed.Screen = screen
|
||||
counts, ok := byScreen[screen]
|
||||
if !ok {
|
||||
observed.Observations = 1
|
||||
byScreen[screen] = &observed
|
||||
continue
|
||||
}
|
||||
if observed.Interactive > counts.Interactive {
|
||||
observed.Observations = counts.Observations
|
||||
*counts = observed
|
||||
}
|
||||
counts.Observations++
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// shape names a screen the app did not name itself, by the set of distinct
|
||||
// controls it shows. Repetition is dropped deliberately: a ledger holding three
|
||||
// rows and the same ledger holding five is one screen, not two.
|
||||
func shape(elements []element) string {
|
||||
distinct := map[string]bool{}
|
||||
for _, item := range elements {
|
||||
key := item.ResourceID
|
||||
if key == "" {
|
||||
key = item.Class + "/" + item.Description
|
||||
}
|
||||
distinct[key] = true
|
||||
}
|
||||
keys := make([]string, 0, len(distinct))
|
||||
for key := range distinct {
|
||||
keys = append(keys, key)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
sum := sha256.Sum256([]byte(strings.Join(keys, "\x00")))
|
||||
return "shape:" + hex.EncodeToString(sum[:])[:8]
|
||||
}
|
||||
|
||||
// volatile reports whether a description carries the data it labels. Compose
|
||||
// merges a row's children into one description, so a ledger row arrives
|
||||
// labelled with the customer name, the balance and a relative age that reprices
|
||||
// itself every month. Such a description exists, which is why a presence check
|
||||
// scores it as a selector, and it is not one: the run that recorded it is the
|
||||
// only run it matches. The test is a heuristic and deliberately blunt, since
|
||||
// the alternative is to call every one of them stable.
|
||||
func volatile(description string) bool {
|
||||
trimmed := strings.TrimSpace(description)
|
||||
// A one-character label is an avatar initial, and it is the first letter of
|
||||
// a name the run happened to observe. Renaming the record changes it. The
|
||||
// digit test below catches an initial drawn from a numeric name and would
|
||||
// miss every alphabetic one, which is how four of them were counted as
|
||||
// durable before this was noticed on a real ledger.
|
||||
if len([]rune(trimmed)) == 1 &&
|
||||
(unicode.IsLetter([]rune(trimmed)[0]) || unicode.IsDigit([]rune(trimmed)[0])) {
|
||||
return true
|
||||
}
|
||||
for _, character := range trimmed {
|
||||
if unicode.IsDigit(character) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
lowered := strings.ToLower(trimmed)
|
||||
for _, marker := range []string{" ago", "since ", "yesterday", "today", "tomorrow", "last ", "due ", "minute", "hour", "day", "week", "month", "year"} {
|
||||
if strings.Contains(lowered, marker) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// dataCarryingID reports whether an identifier embeds the record it names. A
|
||||
// list row tagged `customer_row_<uuid>` names its role durably and its instance
|
||||
// not at all, so an exact match on it survives exactly one run. The prefix is
|
||||
// still worth something, which is what `idPrefix:` is for, but a measure that
|
||||
// counts the whole string as a durable selector overstates what a spec can say.
|
||||
// Scoped to the local name so an Android package prefix cannot trip it, and
|
||||
// tuned to leave ordinary names like `button2` alone.
|
||||
func dataCarryingID(identifier string) bool {
|
||||
local := identifier
|
||||
if index := strings.LastIndex(local, "/"); index >= 0 {
|
||||
local = local[index+1:]
|
||||
}
|
||||
digits := 0
|
||||
for _, character := range local {
|
||||
if unicode.IsDigit(character) {
|
||||
digits++
|
||||
if digits >= 4 {
|
||||
return true
|
||||
}
|
||||
continue
|
||||
}
|
||||
digits = 0
|
||||
}
|
||||
hex, groups := 0, 0
|
||||
for _, character := range local + "-" {
|
||||
if isHexDigit(character) {
|
||||
hex++
|
||||
continue
|
||||
}
|
||||
if hex >= 4 {
|
||||
groups++
|
||||
}
|
||||
hex = 0
|
||||
}
|
||||
return groups >= 2
|
||||
}
|
||||
|
||||
func isHexDigit(character rune) bool {
|
||||
return unicode.IsDigit(character) ||
|
||||
(character >= 'a' && character <= 'f') ||
|
||||
(character >= 'A' && character <= 'F')
|
||||
}
|
||||
|
||||
// keypadDigits reports whether this screen shows a numeric keypad, in which case
|
||||
// its single-character digit labels name fixed keys rather than the first
|
||||
// character of somebody's name. Without the screen for context the two are
|
||||
// indistinguishable: an avatar initial and a calculator key are both one
|
||||
// character, and only one of them survives a data change.
|
||||
func keypadDigits(elements []element) bool {
|
||||
seen := map[rune]bool{}
|
||||
for _, item := range elements {
|
||||
label := strings.TrimSpace(item.Description)
|
||||
if label == "" {
|
||||
label = strings.TrimSpace(item.Text)
|
||||
}
|
||||
runes := []rune(label)
|
||||
if len(runes) == 1 && unicode.IsDigit(runes[0]) {
|
||||
seen[runes[0]] = true
|
||||
}
|
||||
}
|
||||
return len(seen) >= 6
|
||||
}
|
||||
|
||||
func isSingleDigit(label string) bool {
|
||||
runes := []rune(strings.TrimSpace(label))
|
||||
return len(runes) == 1 && unicode.IsDigit(runes[0])
|
||||
}
|
||||
|
||||
func measure(elements []element) Counts {
|
||||
counts := Counts{Elements: len(elements)}
|
||||
keypad := keypadDigits(elements)
|
||||
identifiers := map[string]int{}
|
||||
for _, item := range elements {
|
||||
if !item.Clickable && !item.Editable {
|
||||
continue
|
||||
}
|
||||
counts.Interactive++
|
||||
switch {
|
||||
case item.ResourceID != "" && !dataCarryingID(item.ResourceID):
|
||||
counts.ByIdentifier++
|
||||
identifiers[item.ResourceID]++
|
||||
case item.ResourceID != "":
|
||||
counts.ByDataID++
|
||||
counts.Needing = append(counts.Needing, describe(item))
|
||||
case item.Description != "" && keypad && isSingleDigit(item.Description):
|
||||
counts.ByDescription++
|
||||
case item.Description != "" && !volatile(item.Description):
|
||||
counts.ByDescription++
|
||||
case item.Description != "":
|
||||
counts.ByVolatile++
|
||||
counts.Needing = append(counts.Needing, describe(item))
|
||||
case item.Text != "":
|
||||
counts.ByTextOnly++
|
||||
counts.Needing = append(counts.Needing, describe(item))
|
||||
default:
|
||||
counts.Unaddressable++
|
||||
counts.Needing = append(counts.Needing, describe(item))
|
||||
}
|
||||
}
|
||||
for _, repeats := range identifiers {
|
||||
if repeats > 1 {
|
||||
counts.AmbiguousIDs += repeats
|
||||
}
|
||||
}
|
||||
return counts
|
||||
}
|
||||
|
||||
func describe(item element) string {
|
||||
class := item.Class
|
||||
if class == "" {
|
||||
class = "(no class)"
|
||||
}
|
||||
switch {
|
||||
case item.ResourceID != "" && dataCarryingID(item.ResourceID):
|
||||
return fmt.Sprintf(
|
||||
"%s id=%q (names its record, not its role)",
|
||||
class,
|
||||
item.ResourceID,
|
||||
)
|
||||
case item.Description != "":
|
||||
return fmt.Sprintf(
|
||||
"%s desc=%q (carries its own data)",
|
||||
class,
|
||||
item.Description,
|
||||
)
|
||||
case item.Text != "":
|
||||
return fmt.Sprintf("%s text=%q", class, item.Text)
|
||||
default:
|
||||
return class + " (no label at all)"
|
||||
}
|
||||
}
|
||||
|
||||
func signature(elements []element) string {
|
||||
var builder strings.Builder
|
||||
for _, item := range elements {
|
||||
builder.WriteString(item.Class)
|
||||
builder.WriteByte('|')
|
||||
builder.WriteString(item.ResourceID)
|
||||
builder.WriteByte('|')
|
||||
builder.WriteString(item.Description)
|
||||
builder.WriteByte(';')
|
||||
}
|
||||
return builder.String()
|
||||
}
|
||||
|
||||
func render(out io.Writer, screens []Counts, traces int, show int) {
|
||||
fmt.Fprintf(out, "%d trace(s), %d screen(s)\n\n", traces, len(screens))
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"%-28s %5s %5s %5s %4s %5s %5s %5s %5s %7s\n",
|
||||
"screen",
|
||||
"obs",
|
||||
"inter",
|
||||
"id",
|
||||
"id~",
|
||||
"desc",
|
||||
"vol",
|
||||
"text",
|
||||
"none",
|
||||
"stable%",
|
||||
)
|
||||
total := Counts{Screen: "TOTAL"}
|
||||
for _, screen := range screens {
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"%-28s %5d %5d %5d %4d %5d %5d %5d %5d %6.1f%%\n",
|
||||
truncate(
|
||||
screen.Screen,
|
||||
28,
|
||||
),
|
||||
screen.Observations,
|
||||
screen.Interactive,
|
||||
screen.ByIdentifier,
|
||||
screen.ByDataID,
|
||||
screen.ByDescription,
|
||||
screen.ByVolatile,
|
||||
screen.ByTextOnly,
|
||||
screen.Unaddressable,
|
||||
screen.stableShare(),
|
||||
)
|
||||
total.Observations += screen.Observations
|
||||
total.Elements += screen.Elements
|
||||
total.Interactive += screen.Interactive
|
||||
total.ByIdentifier += screen.ByIdentifier
|
||||
total.ByDataID += screen.ByDataID
|
||||
total.ByDescription += screen.ByDescription
|
||||
total.ByVolatile += screen.ByVolatile
|
||||
total.ByTextOnly += screen.ByTextOnly
|
||||
total.Unaddressable += screen.Unaddressable
|
||||
total.AmbiguousIDs += screen.AmbiguousIDs
|
||||
}
|
||||
fmt.Fprintf(out, "%-28s %5d %5d %5d %4d %5d %5d %5d %5d %6.1f%%\n",
|
||||
total.Screen, total.Observations, total.Interactive, total.ByIdentifier,
|
||||
total.ByDataID, total.ByDescription, total.ByVolatile, total.ByTextOnly,
|
||||
total.Unaddressable, total.stableShare())
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"\n%d interactive element(s) carry an identifier, which is the number that survives a data change\n",
|
||||
total.ByIdentifier,
|
||||
)
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"%d share a resource id with another on the same screen\n",
|
||||
total.AmbiguousIDs,
|
||||
)
|
||||
if show <= 0 {
|
||||
return
|
||||
}
|
||||
for _, screen := range screens {
|
||||
if len(screen.Needing) == 0 {
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"\n%s, %d control(s) no durable selector reaches:\n",
|
||||
screen.Screen,
|
||||
len(screen.Needing),
|
||||
)
|
||||
for index, item := range screen.Needing {
|
||||
if index == show {
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
" ... and %d more\n",
|
||||
len(screen.Needing)-show,
|
||||
)
|
||||
break
|
||||
}
|
||||
fmt.Fprintf(out, " %s\n", item)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func truncate(value string, width int) string {
|
||||
if len(value) <= width {
|
||||
return value
|
||||
}
|
||||
return value[:width-1] + "~"
|
||||
}
|
||||
@@ -0,0 +1,446 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestMeasureRanksSelectorsByDurability(t *testing.T) {
|
||||
counts := measure([]element{
|
||||
{
|
||||
Class: "Button",
|
||||
ResourceID: "app:id/save",
|
||||
Text: "Save",
|
||||
Clickable: true,
|
||||
},
|
||||
{
|
||||
Class: "Button",
|
||||
Description: "Filter",
|
||||
Text: "Filter",
|
||||
Clickable: true,
|
||||
},
|
||||
{Class: "View", Text: "Ramesh Kumar", Clickable: true},
|
||||
{Class: "View", Clickable: true},
|
||||
{Class: "EditText", ResourceID: "app:id/amount", Editable: true},
|
||||
{Class: "TextView", Text: "Balance", Clickable: false},
|
||||
})
|
||||
if counts.Elements != 6 {
|
||||
t.Fatalf("elements: want 6, got %d", counts.Elements)
|
||||
}
|
||||
if counts.Interactive != 5 {
|
||||
t.Fatalf("interactive: want 5, got %d", counts.Interactive)
|
||||
}
|
||||
if counts.ByIdentifier != 2 {
|
||||
t.Errorf("by identifier: want 2, got %d", counts.ByIdentifier)
|
||||
}
|
||||
if counts.ByDescription != 1 {
|
||||
t.Errorf("by description: want 1, got %d", counts.ByDescription)
|
||||
}
|
||||
if counts.ByTextOnly != 1 {
|
||||
t.Errorf("by text only: want 1, got %d", counts.ByTextOnly)
|
||||
}
|
||||
if counts.Unaddressable != 1 {
|
||||
t.Errorf("unaddressable: want 1, got %d", counts.Unaddressable)
|
||||
}
|
||||
}
|
||||
|
||||
// The row description a merged Compose semantics node produces carries the
|
||||
// balance and a relative age that reprices itself every month. It is present,
|
||||
// so a presence check scores it as a selector; it matches only the run that
|
||||
// recorded it.
|
||||
func TestVolatileDescriptionsAreNotCountedAsStable(t *testing.T) {
|
||||
row := "Ramesh ji, 95, Pending Collection Since 15 months, Due"
|
||||
if !volatile(row) {
|
||||
t.Fatalf("expected %q to be treated as data-carrying", row)
|
||||
}
|
||||
counts := measure([]element{
|
||||
{Class: "Button", Description: row, Clickable: true},
|
||||
{Class: "Button", Description: "Filter", Clickable: true},
|
||||
})
|
||||
if counts.ByVolatile != 1 {
|
||||
t.Errorf("volatile: want 1, got %d", counts.ByVolatile)
|
||||
}
|
||||
if counts.ByDescription != 1 {
|
||||
t.Errorf("stable description: want 1, got %d", counts.ByDescription)
|
||||
}
|
||||
if got := counts.stableShare(); got != 50 {
|
||||
t.Fatalf("stable share: want 50, got %.1f", got)
|
||||
}
|
||||
if len(counts.Needing) != 1 ||
|
||||
!strings.Contains(counts.Needing[0], "carries its own data") {
|
||||
t.Fatalf(
|
||||
"the volatile row belongs on the work list, got %v",
|
||||
counts.Needing,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestVolatileLeavesPlainLabelsAlone(t *testing.T) {
|
||||
for _, label := range []string{"Filter", "Search", "Add Relationship", "Share", "Skip"} {
|
||||
if volatile(label) {
|
||||
t.Errorf("%q should count as a durable label", label)
|
||||
}
|
||||
}
|
||||
for _, label := range []string{"₹95", "2 days ago", "Due today", "Pending Since 15 months", "Edited on 11 Jun 2026"} {
|
||||
if !volatile(label) {
|
||||
t.Errorf("%q carries data and should not count as durable", label)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestStableShareExcludesDataDependentText(t *testing.T) {
|
||||
counts := measure([]element{
|
||||
{ResourceID: "app:id/add", Clickable: true},
|
||||
{Text: "Ramesh Kumar", Clickable: true},
|
||||
{Text: "Suresh Patel", Clickable: true},
|
||||
{Clickable: true},
|
||||
})
|
||||
if got := counts.stableShare(); got != 25 {
|
||||
t.Fatalf("stable share: want 25, got %.1f", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMeasureCountsRepeatedIdentifiersAsAmbiguous(t *testing.T) {
|
||||
counts := measure([]element{
|
||||
{ResourceID: "app:id/row", Text: "Ramesh", Clickable: true},
|
||||
{ResourceID: "app:id/row", Text: "Suresh", Clickable: true},
|
||||
{ResourceID: "app:id/add", Clickable: true},
|
||||
})
|
||||
if counts.AmbiguousIDs != 2 {
|
||||
t.Fatalf("ambiguous: want 2, got %d", counts.AmbiguousIDs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAccumulateKeepsRichestObservationPerScreen(t *testing.T) {
|
||||
trace := writeTrace(
|
||||
t,
|
||||
`{"step":0,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true}]}}
|
||||
{"step":1,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true},{"text":"Ramesh","clickable":true},{"class":"View","clickable":true}]}}
|
||||
{"step":2,"screen":"home","hierarchy":{"elements":[{"description":"Filter","clickable":true}]}}
|
||||
`,
|
||||
)
|
||||
byScreen := map[string]*Counts{}
|
||||
if err := accumulate(trace, byScreen, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ledger := byScreen["ledger"]
|
||||
if ledger.Interactive != 3 {
|
||||
t.Fatalf("ledger interactive: want 3, got %d", ledger.Interactive)
|
||||
}
|
||||
if ledger.Observations != 2 {
|
||||
t.Fatalf("ledger observations: want 2, got %d", ledger.Observations)
|
||||
}
|
||||
if ledger.Unaddressable != 1 {
|
||||
t.Errorf("ledger unaddressable: want 1, got %d", ledger.Unaddressable)
|
||||
}
|
||||
if byScreen["home"].ByDescription != 1 {
|
||||
t.Errorf(
|
||||
"home by description: want 1, got %d",
|
||||
byScreen["home"].ByDescription,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// An idling probe observes one screen many times. Counting each observation
|
||||
// would report how long the explorer sat there rather than what it could reach.
|
||||
func TestAccumulateCountsIdenticalHierarchiesOnce(t *testing.T) {
|
||||
line := `{"step":0,"screen":"ledger","hierarchy":{"elements":[{"resourceId":"app:id/add","clickable":true}]}}` + "\n"
|
||||
byScreen := map[string]*Counts{}
|
||||
if err := accumulate(writeTrace(t, strings.Repeat(line, 5)), byScreen, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got := byScreen["ledger"].Observations; got != 1 {
|
||||
t.Fatalf("observations: want 1, got %d", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAccumulateSkipsStepsWithoutHierarchy(t *testing.T) {
|
||||
byScreen := map[string]*Counts{}
|
||||
err := accumulate(writeTrace(t, `{"step":0,"screen":"ledger"}
|
||||
{"step":1,"screen":"ledger","hierarchy":{"elements":[]}}
|
||||
`), byScreen, "")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(byScreen) != 0 {
|
||||
t.Fatalf("want no screens, got %d", len(byScreen))
|
||||
}
|
||||
}
|
||||
|
||||
func TestShapeIgnoresRepetitionButNotComposition(t *testing.T) {
|
||||
threeRows := shape([]element{
|
||||
{ResourceID: "app:id/list"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
})
|
||||
fiveRows := shape([]element{
|
||||
{ResourceID: "app:id/list"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/row"},
|
||||
})
|
||||
if threeRows != fiveRows {
|
||||
t.Fatalf(
|
||||
"row count changed the screen identity: %s vs %s",
|
||||
threeRows,
|
||||
fiveRows,
|
||||
)
|
||||
}
|
||||
withDialog := shape([]element{
|
||||
{ResourceID: "app:id/list"},
|
||||
{ResourceID: "app:id/row"},
|
||||
{ResourceID: "app:id/confirm_dialog"},
|
||||
})
|
||||
if withDialog == threeRows {
|
||||
t.Fatal(
|
||||
"a screen showing a dialog must not collapse into the screen behind it",
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAccumulateSeparatesUnnamedScreensByShape(t *testing.T) {
|
||||
byScreen := map[string]*Counts{}
|
||||
err := accumulate(
|
||||
writeTrace(
|
||||
t,
|
||||
`{"step":0,"hierarchy":{"elements":[{"resourceId":"app:id/ledger","clickable":true}]}}
|
||||
{"step":1,"hierarchy":{"elements":[{"resourceId":"app:id/add_transaction","clickable":true}]}}
|
||||
`,
|
||||
),
|
||||
byScreen,
|
||||
"",
|
||||
)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(byScreen) != 2 {
|
||||
t.Fatalf("want 2 screens, got %d", len(byScreen))
|
||||
}
|
||||
for name := range byScreen {
|
||||
if !strings.HasPrefix(name, "shape:") {
|
||||
t.Errorf("unnamed screen should be keyed by shape, got %q", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectFindsTracesUnderRunDirectories(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
for _, run := range []string{"run-1", "run-2"} {
|
||||
directory := filepath.Join(root, run)
|
||||
if err := os.MkdirAll(directory, 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte("{}\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
traces, err := collect([]string{root})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(traces) != 2 {
|
||||
t.Fatalf("want 2 traces, got %d: %v", len(traces), traces)
|
||||
}
|
||||
}
|
||||
|
||||
func writeTrace(t *testing.T, content string) string {
|
||||
t.Helper()
|
||||
path := filepath.Join(t.TempDir(), "trace.jsonl")
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return path
|
||||
}
|
||||
|
||||
// A status bar contributes three well-labelled controls to every capture. An
|
||||
// app with no identifiers of its own scores 100 percent if they are counted.
|
||||
func TestAccumulateDropsSystemUIByDefault(t *testing.T) {
|
||||
trace := writeTrace(
|
||||
t,
|
||||
`{"step":0,"screen":"ledger","hierarchy":{"elements":[
|
||||
{"resourceId":"com.android.systemui:id/clock","package":"com.android.systemui","clickable":true},
|
||||
{"resourceId":"com.android.systemui:id/battery","package":"com.android.systemui","clickable":true},
|
||||
{"class":"android.view.View","package":"in.okcredit.merchant.debug","clickable":true}]}}
|
||||
`,
|
||||
)
|
||||
byScreen := map[string]*Counts{}
|
||||
if err := accumulate(trace, byScreen, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
counts := byScreen["ledger"]
|
||||
if counts.Interactive != 1 {
|
||||
t.Fatalf(
|
||||
"interactive: want 1 after dropping system UI, got %d",
|
||||
counts.Interactive,
|
||||
)
|
||||
}
|
||||
if counts.ByIdentifier != 0 {
|
||||
t.Fatalf(
|
||||
"system identifiers leaked into the app's score: %d",
|
||||
counts.ByIdentifier,
|
||||
)
|
||||
}
|
||||
if got := counts.stableShare(); got != 0 {
|
||||
t.Fatalf("stable share: want 0, got %.1f", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAccumulatePackageFlagPinsExactly(t *testing.T) {
|
||||
trace := writeTrace(
|
||||
t,
|
||||
`{"step":0,"screen":"ledger","hierarchy":{"elements":[
|
||||
{"resourceId":"other:id/x","package":"com.other.app","clickable":true},
|
||||
{"class":"android.view.View","package":"in.okcredit.merchant.debug","clickable":true}]}}
|
||||
`,
|
||||
)
|
||||
byScreen := map[string]*Counts{}
|
||||
if err := accumulate(trace, byScreen, "in.okcredit.merchant.debug"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got := byScreen["ledger"].Interactive; got != 1 {
|
||||
t.Fatalf("interactive: want 1, got %d", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Only window roots carry a package attribute in a hierarchy dump. Treating an
|
||||
// unstamped node as foreign discards the application and reports a clean zero.
|
||||
func TestScopedKeepsUnstampedNodes(t *testing.T) {
|
||||
elements := []element{
|
||||
{
|
||||
Package: "com.android.systemui",
|
||||
ResourceID: "sysui:id/clock",
|
||||
Clickable: true,
|
||||
},
|
||||
{Package: "in.okcredit.merchant.debug", ResourceID: "app:id/root"},
|
||||
{Class: "android.view.View", Clickable: true},
|
||||
}
|
||||
if got := len(scoped(elements, "")); got != 2 {
|
||||
t.Fatalf("denylist mode: want 2 kept, got %d", got)
|
||||
}
|
||||
if got := len(scoped(elements, "in.okcredit.merchant.debug")); got != 2 {
|
||||
t.Fatalf("pinned mode: want 2 kept, got %d", got)
|
||||
}
|
||||
if got := len(scoped(elements, "com.other.app")); got != 1 {
|
||||
t.Fatalf(
|
||||
"pinned to a foreign package: want 1 unstamped node kept, got %d",
|
||||
got,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// An avatar renders the first letter of the name beside it, so a one-character
|
||||
// description is the name in disguise. Four of these were scored as durable on
|
||||
// a real ledger before the rule existed.
|
||||
func TestSingleCharacterDescriptionsAreAvatarInitials(t *testing.T) {
|
||||
for _, initial := range []string{"R", "C", "T", "A", "9", " R "} {
|
||||
if !volatile(initial) {
|
||||
t.Errorf(
|
||||
"%q is an avatar initial and changes when the record is renamed",
|
||||
initial,
|
||||
)
|
||||
}
|
||||
}
|
||||
for _, keypad := range []string{"+", "=", "AC"} {
|
||||
if volatile(keypad) {
|
||||
t.Errorf(
|
||||
"%q is a fixed control label and should count as durable",
|
||||
keypad,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A row tagged with its record's uuid names its role durably and its instance
|
||||
// not at all, so an exact match on it survives exactly one run.
|
||||
func TestDataCarryingIdentifiers(t *testing.T) {
|
||||
for _, identifier := range []string{
|
||||
"customer_row_5f338c10-feef-411c-a070-8999b4890a62",
|
||||
"in.okcredit.merchant.debug:id/customer_row_5f338c10-feef-411c-a070-8999b4890a62",
|
||||
"txn_20260812",
|
||||
} {
|
||||
if !dataCarryingID(identifier) {
|
||||
t.Errorf("%q embeds the record it names", identifier)
|
||||
}
|
||||
}
|
||||
for _, identifier := range []string{
|
||||
"customer_supplier_list",
|
||||
"summary_card",
|
||||
"in.okcredit.merchant.debug:id/buttonLogin",
|
||||
"button2",
|
||||
"add_relationship",
|
||||
} {
|
||||
if dataCarryingID(identifier) {
|
||||
t.Errorf("%q names a role and should count as durable", identifier)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestMeasureSeparatesDataCarryingIdentifiers(t *testing.T) {
|
||||
counts := measure([]element{
|
||||
{Class: "View", ResourceID: "summary_card", Clickable: true},
|
||||
{
|
||||
Class: "View",
|
||||
ResourceID: "customer_row_5f338c10-feef-411c-a070-8999b4890a62",
|
||||
Clickable: true,
|
||||
},
|
||||
{
|
||||
Class: "View",
|
||||
ResourceID: "customer_row_8ea26690-f809-4d23-8560-9cca6d1a5bcd",
|
||||
Clickable: true,
|
||||
},
|
||||
})
|
||||
if counts.ByIdentifier != 1 {
|
||||
t.Errorf("durable identifiers: want 1, got %d", counts.ByIdentifier)
|
||||
}
|
||||
if counts.ByDataID != 2 {
|
||||
t.Errorf("data-carrying identifiers: want 2, got %d", counts.ByDataID)
|
||||
}
|
||||
if got := counts.stableShare(); got < 33 || got > 34 {
|
||||
t.Errorf("stable share: want about 33.3, got %.1f", got)
|
||||
}
|
||||
if len(counts.Needing) != 2 {
|
||||
t.Fatalf(
|
||||
"both row identifiers belong on the work list, got %v",
|
||||
counts.Needing,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// A calculator key and an avatar initial are both one character. Only the
|
||||
// screen they sit on tells them apart, so the keypad rule needs that context.
|
||||
func TestKeypadDigitsAreDurableButAvatarInitialsAreNot(t *testing.T) {
|
||||
keypad := make([]element, 0, 10)
|
||||
for _, key := range []string{"1", "2", "3", "4", "5", "6", "7", "8", "9", "0"} {
|
||||
keypad = append(
|
||||
keypad,
|
||||
element{Class: "Button", Description: key, Clickable: true},
|
||||
)
|
||||
}
|
||||
counts := measure(keypad)
|
||||
if counts.ByDescription != 10 {
|
||||
t.Fatalf(
|
||||
"keypad keys are fixed labels: want 10 durable, got %d",
|
||||
counts.ByDescription,
|
||||
)
|
||||
}
|
||||
|
||||
counts = measure([]element{
|
||||
{Class: "Button", Description: "R", Clickable: true},
|
||||
{Class: "Button", Description: "9", Clickable: true},
|
||||
{Class: "Button", Description: "Filter", Clickable: true},
|
||||
})
|
||||
if counts.ByVolatile != 2 {
|
||||
t.Fatalf(
|
||||
"avatar initials carry data: want 2 volatile, got %d",
|
||||
counts.ByVolatile,
|
||||
)
|
||||
}
|
||||
if counts.ByDescription != 1 {
|
||||
t.Fatalf("durable descriptions: want 1, got %d", counts.ByDescription)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,223 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
const maxStepBytes = 64 * 1024 * 1024
|
||||
|
||||
// loadedRun is one run directory: its meta, the steps in file order, and the
|
||||
// synthetic end-of-run record the runner writes when Finalize convicted a
|
||||
// liveness obligation.
|
||||
type loadedRun struct {
|
||||
Directory string
|
||||
Meta trace.Meta
|
||||
Steps []trace.Step
|
||||
Finalize *trace.Step
|
||||
}
|
||||
|
||||
// loadRun reads a run directory and refuses anything E3 cannot replay. A step
|
||||
// written before the format change carries no element depths, so its hierarchy
|
||||
// decodes with a nil root and every selector resolves to nothing: the refusal
|
||||
// has to name the version rather than let the replay report an empty screen.
|
||||
func loadRun(directory string) (loadedRun, error) {
|
||||
metaBody, err := os.ReadFile(filepath.Join(directory, "meta.json"))
|
||||
if err != nil {
|
||||
return loadedRun{}, fmt.Errorf("read meta: %w", err)
|
||||
}
|
||||
var meta trace.Meta
|
||||
if err := json.Unmarshal(metaBody, &meta); err != nil {
|
||||
return loadedRun{}, fmt.Errorf("decode meta: %w", err)
|
||||
}
|
||||
|
||||
file, err := os.Open(filepath.Join(directory, "trace.jsonl"))
|
||||
if err != nil {
|
||||
return loadedRun{}, fmt.Errorf("open trace: %w", err)
|
||||
}
|
||||
defer file.Close()
|
||||
|
||||
run := loadedRun{Directory: directory, Meta: meta}
|
||||
scanner := bufio.NewScanner(file)
|
||||
scanner.Buffer(make([]byte, 0, 1024*1024), maxStepBytes)
|
||||
line := 0
|
||||
for scanner.Scan() {
|
||||
line++
|
||||
if len(scanner.Bytes()) == 0 {
|
||||
continue
|
||||
}
|
||||
var step trace.Step
|
||||
if err := json.Unmarshal(scanner.Bytes(), &step); err != nil {
|
||||
return loadedRun{}, fmt.Errorf(
|
||||
"decode step on line %d: %w",
|
||||
line,
|
||||
err,
|
||||
)
|
||||
}
|
||||
if step.TraceVersion != trace.TraceVersion {
|
||||
return loadedRun{}, fmt.Errorf(
|
||||
"step %d is trace_version %d, and E3 replays version %d only: "+
|
||||
"an older step stores no element depths, so its hierarchy decodes "+
|
||||
"with a nil root and every selector resolves to nothing on replay",
|
||||
step.Index, step.TraceVersion, trace.TraceVersion)
|
||||
}
|
||||
run.Steps = append(run.Steps, step)
|
||||
}
|
||||
if err := scanner.Err(); err != nil {
|
||||
return loadedRun{}, fmt.Errorf("read trace: %w", err)
|
||||
}
|
||||
if len(run.Steps) == 0 {
|
||||
return loadedRun{}, fmt.Errorf("trace has no steps")
|
||||
}
|
||||
if last := run.Steps[len(run.Steps)-1]; len(run.Steps) > 1 &&
|
||||
last.Hierarchy == nil &&
|
||||
len(last.Violations) > 0 {
|
||||
run.Finalize = &last
|
||||
run.Steps = run.Steps[:len(run.Steps)-1]
|
||||
}
|
||||
if meta.Seed == 0 {
|
||||
return loadedRun{}, fmt.Errorf(
|
||||
"meta records no seed, so the run's bundle cannot be reproduced",
|
||||
)
|
||||
}
|
||||
return run, nil
|
||||
}
|
||||
|
||||
// finalizeIndex is the step index the runner gives its end-of-run record: one
|
||||
// past the last step it wrote.
|
||||
func (r loadedRun) finalizeIndex() int {
|
||||
return r.Steps[len(r.Steps)-1].Index + 1
|
||||
}
|
||||
|
||||
// extractorFold reconstructs every extractor's value at each step from the
|
||||
// recorded per-step diffs. A run's trace stores what changed, so the value an
|
||||
// extractor held at a step is the last change at or before it; an extractor
|
||||
// that never changed held JSON null throughout, which is what an unwritten
|
||||
// diff means.
|
||||
func extractorFold(
|
||||
steps []trace.Step,
|
||||
names []string,
|
||||
) ([]map[int]json.RawMessage, error) {
|
||||
index := make(map[string]int, len(names))
|
||||
for position, name := range names {
|
||||
index[name] = position
|
||||
}
|
||||
current := make(map[int]json.RawMessage, len(names))
|
||||
for position := range names {
|
||||
current[position] = json.RawMessage("null")
|
||||
}
|
||||
folded := make([]map[int]json.RawMessage, len(steps))
|
||||
for step := range steps {
|
||||
for name, change := range steps[step].ExtractorChanges {
|
||||
position, ok := index[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf(
|
||||
"step %d records extractor %q, which the spec does not register; "+
|
||||
"the trace and the spec are not the same bundle",
|
||||
steps[step].Index,
|
||||
name,
|
||||
)
|
||||
}
|
||||
current[position] = change.Curr
|
||||
}
|
||||
snapshot := make(map[int]json.RawMessage, len(current))
|
||||
for position, value := range current {
|
||||
snapshot[position] = value
|
||||
}
|
||||
folded[step] = snapshot
|
||||
}
|
||||
return folded, nil
|
||||
}
|
||||
|
||||
// lastActionFor rebuilds the action the runner had applied before the next
|
||||
// step observed. An action the runner chose but never dispatched left
|
||||
// state.lastAction null, and the recorded skip reason is what says so.
|
||||
func lastActionFor(step trace.Step) *verifier.Action {
|
||||
if step.NextAction == nil || step.ActionSkipped != "" {
|
||||
return nil
|
||||
}
|
||||
recorded := *step.NextAction
|
||||
action := verifier.Action{
|
||||
Kind: verifier.ActionKind(recorded.Kind),
|
||||
On: recorded.Selector,
|
||||
Text: recorded.Text,
|
||||
X: recorded.X,
|
||||
Y: recorded.Y,
|
||||
FromX: recorded.FromX,
|
||||
FromY: recorded.FromY,
|
||||
ToX: recorded.ToX,
|
||||
ToY: recorded.ToY,
|
||||
Key: recorded.Key,
|
||||
DurationMillis: recorded.DurationMillis,
|
||||
}
|
||||
return &action
|
||||
}
|
||||
|
||||
func traceLogs(entries []trace.LogEntry) []verifier.LogEntry {
|
||||
if len(entries) == 0 {
|
||||
return nil
|
||||
}
|
||||
logs := make([]verifier.LogEntry, 0, len(entries))
|
||||
for _, entry := range entries {
|
||||
logs = append(logs, verifier.LogEntry{
|
||||
UnixMillis: entry.UnixMillis,
|
||||
Level: entry.Level,
|
||||
Tag: entry.Tag,
|
||||
Message: entry.Message,
|
||||
})
|
||||
}
|
||||
return logs
|
||||
}
|
||||
|
||||
func traceExceptions(entries []trace.Exception) []verifier.Exception {
|
||||
if len(entries) == 0 {
|
||||
return nil
|
||||
}
|
||||
exceptions := make([]verifier.Exception, 0, len(entries))
|
||||
for _, entry := range entries {
|
||||
exceptions = append(exceptions, verifier.Exception{
|
||||
Class: entry.Class,
|
||||
Message: entry.Message,
|
||||
StackTrace: entry.StackTrace,
|
||||
UnixMillis: entry.UnixMillis,
|
||||
})
|
||||
}
|
||||
return exceptions
|
||||
}
|
||||
|
||||
// discoverRuns finds every run directory at or below root, a run directory
|
||||
// being one holding both meta.json and trace.jsonl.
|
||||
func discoverRuns(root string) ([]string, error) {
|
||||
var directories []string
|
||||
err := filepath.Walk(
|
||||
root,
|
||||
func(path string, info os.FileInfo, err error) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if !info.IsDir() {
|
||||
return nil
|
||||
}
|
||||
if _, statErr := os.Stat(filepath.Join(path, "trace.jsonl")); statErr != nil {
|
||||
return nil
|
||||
}
|
||||
if _, statErr := os.Stat(filepath.Join(path, "meta.json")); statErr != nil {
|
||||
return nil
|
||||
}
|
||||
directories = append(directories, path)
|
||||
return nil
|
||||
},
|
||||
)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sort.Strings(directories)
|
||||
return directories, nil
|
||||
}
|
||||
@@ -0,0 +1,162 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
func writeRunFiles(t *testing.T, directory, meta string, steps ...string) {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(filepath.Join(directory, "meta.json"), []byte(meta), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
body := strings.Join(steps, "\n") + "\n"
|
||||
if err := os.WriteFile(filepath.Join(directory, "trace.jsonl"), []byte(body), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadRunRefusesAStepFromBeforeTheFormatChange(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeRunFiles(
|
||||
t,
|
||||
directory,
|
||||
`{"seed": 3}`,
|
||||
`{"step":1,"timestamp":"2026-08-15T12:00:00Z"}`,
|
||||
)
|
||||
|
||||
_, err := loadRun(directory)
|
||||
|
||||
if err == nil {
|
||||
t.Fatal("a version-0 step must be refused")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "trace_version 0") ||
|
||||
!strings.Contains(err.Error(), "depths") {
|
||||
t.Errorf(
|
||||
"the refusal must name the version and why it cannot replay: %v",
|
||||
err,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadRunRefusesARunWithNoSeed(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeRunFiles(
|
||||
t,
|
||||
directory,
|
||||
`{}`,
|
||||
`{"step":1,"trace_version":1,"timestamp":"2026-08-15T12:00:00Z"}`,
|
||||
)
|
||||
|
||||
_, err := loadRun(directory)
|
||||
|
||||
if err == nil || !strings.Contains(err.Error(), "seed") {
|
||||
t.Errorf("a run without a seed cannot be bundled as it was: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadRunSplitsOffTheEndOfRunRecord(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeRunFiles(
|
||||
t,
|
||||
directory,
|
||||
`{"seed": 3}`,
|
||||
`{"step":1,"trace_version":1,"timestamp":"2026-08-15T12:00:00Z","hierarchy":{"elements":[],"depths":[]}}`,
|
||||
`{"step":2,"trace_version":1,"timestamp":"2026-08-15T12:00:01Z","violations":["reachable"]}`,
|
||||
)
|
||||
|
||||
run, err := loadRun(directory)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if len(run.Steps) != 1 {
|
||||
t.Fatalf("observation steps: got %d, want 1", len(run.Steps))
|
||||
}
|
||||
if run.Finalize == nil || run.Finalize.Index != 2 {
|
||||
t.Fatalf("finalize record: %+v", run.Finalize)
|
||||
}
|
||||
if run.finalizeIndex() != 2 {
|
||||
t.Errorf("finalize index: got %d, want 2", run.finalizeIndex())
|
||||
}
|
||||
}
|
||||
|
||||
func TestExtractorFoldCarriesAValueForwardUntilItChanges(t *testing.T) {
|
||||
steps := []trace.Step{
|
||||
{Index: 1, ExtractorChanges: map[string]trace.ExtractorChange{
|
||||
"count": {
|
||||
Prev: json.RawMessage("null"),
|
||||
Curr: json.RawMessage("1"),
|
||||
},
|
||||
}},
|
||||
{Index: 2},
|
||||
{Index: 3, ExtractorChanges: map[string]trace.ExtractorChange{
|
||||
"count": {Prev: json.RawMessage("1"), Curr: json.RawMessage("4")},
|
||||
}},
|
||||
}
|
||||
|
||||
folded, err := extractorFold(steps, []string{"count", "unseen"})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
want := []string{"1", "1", "4"}
|
||||
for position, expected := range want {
|
||||
if got := string(folded[position][0]); got != expected {
|
||||
t.Errorf(
|
||||
"count at step %d: got %s, want %s",
|
||||
position+1,
|
||||
got,
|
||||
expected,
|
||||
)
|
||||
}
|
||||
if got := string(folded[position][1]); got != "null" {
|
||||
t.Errorf(
|
||||
"an extractor that never changed must fold to null, got %s",
|
||||
got,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestExtractorFoldRefusesAValueTheSpecCannotPlace(t *testing.T) {
|
||||
steps := []trace.Step{
|
||||
{Index: 1, ExtractorChanges: map[string]trace.ExtractorChange{
|
||||
"gone": {Curr: json.RawMessage("1")},
|
||||
}},
|
||||
}
|
||||
|
||||
_, err := extractorFold(steps, []string{"count"})
|
||||
|
||||
if err == nil || !strings.Contains(err.Error(), "gone") {
|
||||
t.Errorf("an unplaceable extractor value must be refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLastActionIsEmptyWhenTheRunnerNeverDispatchedIt(t *testing.T) {
|
||||
dispatched := trace.Step{
|
||||
NextAction: &trace.Action{Kind: "Tap", Selector: "id:save", X: 4, Y: 5},
|
||||
}
|
||||
skipped := trace.Step{
|
||||
NextAction: &trace.Action{Kind: "Tap", Selector: "id:save"},
|
||||
ActionSkipped: foregroundLossReason,
|
||||
}
|
||||
|
||||
action := lastActionFor(dispatched)
|
||||
if action == nil || action.Kind != verifier.ActionKindTap ||
|
||||
action.On != "id:save" ||
|
||||
action.X != 4 {
|
||||
t.Fatalf("dispatched action: %+v", action)
|
||||
}
|
||||
if lastActionFor(skipped) != nil {
|
||||
t.Error(
|
||||
"an action the runner threw away never reached state.lastAction",
|
||||
)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,312 @@
|
||||
// Command oracle-reduction re-evaluates stored traces offline under four
|
||||
// oracles and reports what each one refutes: the full engine, a crash-only
|
||||
// detector, a single-state check, and a single-step property triple. The
|
||||
// oracles vary while the traces stay fixed, which is what separates a defect an
|
||||
// oracle cannot express from one an explorer never reached.
|
||||
//
|
||||
// The offline engine has to reproduce the verdicts each run recorded. A
|
||||
// disagreement is a bug here or a gap in the trace, so it is reported as a
|
||||
// mismatch and exits nonzero rather than being counted as a finding.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/testrun"
|
||||
)
|
||||
|
||||
const usage = `oracle-reduction replays stored traces under reduced oracles.
|
||||
|
||||
Usage:
|
||||
oracle-reduction --runs <dir> [--output <path>] [--spec <path>]
|
||||
|
||||
--runs is scanned recursively; every directory holding meta.json and
|
||||
trace.jsonl is one trace. Each trace's spec is bundled from the path its
|
||||
meta.json records unless --spec overrides it.
|
||||
|
||||
Exit status is 2 when the offline engine disagreed with any run's recorded
|
||||
verdicts, which blocks the experiment rather than reporting the difference as
|
||||
noise.
|
||||
|
||||
A reduced oracle whose rewrite no longer states a property reports
|
||||
cannot_express for it rather than a verdict, and the property is temporal-only
|
||||
when the other reductions also fail to state or refute it. A property whose
|
||||
window is longer than the two observations a triple spans is reported with
|
||||
single_step_truncates_window, which is the form that decides it.
|
||||
`
|
||||
|
||||
type config struct {
|
||||
runsRoot string
|
||||
outputPath string
|
||||
specPath string
|
||||
// hostExtractors re-runs the spec's extractor getters over each stored
|
||||
// hierarchy instead of replaying the values a web run's page computed. It
|
||||
// asks a different question of the same trace: whether the stored tree
|
||||
// alone carries what the properties read.
|
||||
hostExtractors bool
|
||||
}
|
||||
|
||||
func parseArguments(arguments []string, stderr io.Writer) (config, error) {
|
||||
flagSet := flag.NewFlagSet("oracle-reduction", flag.ContinueOnError)
|
||||
flagSet.SetOutput(stderr)
|
||||
flagSet.Usage = func() {
|
||||
fmt.Fprint(stderr, usage)
|
||||
flagSet.PrintDefaults()
|
||||
}
|
||||
var configuration config
|
||||
flagSet.StringVar(
|
||||
&configuration.runsRoot,
|
||||
"runs",
|
||||
"",
|
||||
"directory tree holding the run directories to replay (required)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.outputPath,
|
||||
"output",
|
||||
"",
|
||||
"file to write the per-trace JSONL to (default stdout)",
|
||||
)
|
||||
flagSet.StringVar(
|
||||
&configuration.specPath,
|
||||
"spec",
|
||||
"",
|
||||
"spec to bundle instead of the one each meta.json records",
|
||||
)
|
||||
flagSet.BoolVar(
|
||||
&configuration.hostExtractors,
|
||||
"host-extractors",
|
||||
false,
|
||||
"re-run the spec's extractors over each stored hierarchy instead of replaying the values a web run's page computed",
|
||||
)
|
||||
if err := flagSet.Parse(arguments); err != nil {
|
||||
return config{}, err
|
||||
}
|
||||
if configuration.runsRoot == "" {
|
||||
return config{}, fmt.Errorf("--runs is required")
|
||||
}
|
||||
return configuration, nil
|
||||
}
|
||||
|
||||
func main() {
|
||||
configuration, err := parseArguments(os.Args[1:], os.Stderr)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "oracle-reduction: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
code, err := run(configuration, os.Stdout, os.Stderr)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "oracle-reduction: %v\n", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
os.Exit(code)
|
||||
}
|
||||
|
||||
func run(configuration config, stdout, stderr io.Writer) (int, error) {
|
||||
directories, err := discoverRuns(configuration.runsRoot)
|
||||
if err != nil {
|
||||
return 1, err
|
||||
}
|
||||
if len(directories) == 0 {
|
||||
return 1, fmt.Errorf(
|
||||
"no run directories under %s",
|
||||
configuration.runsRoot,
|
||||
)
|
||||
}
|
||||
|
||||
output := stdout
|
||||
if configuration.outputPath != "" {
|
||||
file, createErr := os.Create(configuration.outputPath)
|
||||
if createErr != nil {
|
||||
return 1, createErr
|
||||
}
|
||||
defer file.Close()
|
||||
output = file
|
||||
}
|
||||
encoder := json.NewEncoder(output)
|
||||
|
||||
var reports []runReport
|
||||
rejected := 0
|
||||
for _, directory := range directories {
|
||||
loaded, loadErr := loadRun(directory)
|
||||
if loadErr != nil {
|
||||
rejected++
|
||||
fmt.Fprintf(stderr, "skipped %s: %v\n", directory, loadErr)
|
||||
continue
|
||||
}
|
||||
specPath := loaded.Meta.SpecPath
|
||||
if configuration.specPath != "" {
|
||||
specPath = configuration.specPath
|
||||
}
|
||||
bundle, bundleErr := testrun.BundleSpec(specPath, loaded.Meta.Seed)
|
||||
if bundleErr != nil {
|
||||
return 1, fmt.Errorf(
|
||||
"%s: bundle %s: %w",
|
||||
directory,
|
||||
specPath,
|
||||
bundleErr,
|
||||
)
|
||||
}
|
||||
if loaded.Meta.BundleSHA256 != "" &&
|
||||
bundle.SHA256 != loaded.Meta.BundleSHA256 {
|
||||
// The bundler writes each module's path into the output relative to
|
||||
// the working directory, so this differs whenever the replay is
|
||||
// invoked from somewhere else than the run was. What a changed spec
|
||||
// would actually break is caught by the property set, the extractor
|
||||
// names and the residual comparison below.
|
||||
fmt.Fprintf(stderr,
|
||||
"note: %s bundles to %s here and the run recorded %s\n",
|
||||
directory, bundle.SHA256[:12], loaded.Meta.BundleSHA256[:12])
|
||||
}
|
||||
report, replayErr := replay(
|
||||
loaded,
|
||||
string(bundle.JavaScript),
|
||||
configuration.hostExtractors,
|
||||
)
|
||||
if replayErr != nil {
|
||||
return 1, fmt.Errorf("%s: %w", directory, replayErr)
|
||||
}
|
||||
if err := encoder.Encode(report); err != nil {
|
||||
return 1, err
|
||||
}
|
||||
reports = append(reports, report)
|
||||
}
|
||||
|
||||
summarize(reports, rejected, stderr)
|
||||
for _, report := range reports {
|
||||
if !report.Valid {
|
||||
return 2, nil
|
||||
}
|
||||
}
|
||||
if len(reports) == 0 {
|
||||
return 1, fmt.Errorf(
|
||||
"every run directory was rejected; nothing was replayed",
|
||||
)
|
||||
}
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
func summarize(reports []runReport, rejected int, out io.Writer) {
|
||||
invalid := 0
|
||||
crashed := 0
|
||||
weakest := map[string]int{}
|
||||
byClass := map[string]map[string]int{}
|
||||
inexpressible := map[string]map[string]bool{}
|
||||
unmatched := 0
|
||||
engineRefutations := 0
|
||||
for _, report := range reports {
|
||||
if !report.Valid {
|
||||
invalid++
|
||||
}
|
||||
if report.CrashOnly.Fired {
|
||||
crashed++
|
||||
}
|
||||
for _, property := range report.Properties {
|
||||
recordInexpressible(
|
||||
inexpressible,
|
||||
"single-state",
|
||||
property.SingleState,
|
||||
property.Property,
|
||||
)
|
||||
recordInexpressible(
|
||||
inexpressible,
|
||||
"single-step",
|
||||
property.SingleStep,
|
||||
property.Property,
|
||||
)
|
||||
if !property.Engine.Refuted {
|
||||
if property.SingleState.Refuted || property.SingleStep.Refuted {
|
||||
unmatched++
|
||||
}
|
||||
continue
|
||||
}
|
||||
engineRefutations++
|
||||
weakest[property.Weakest]++
|
||||
if byClass[property.Class] == nil {
|
||||
byClass[property.Class] = map[string]int{}
|
||||
}
|
||||
byClass[property.Class][property.Weakest]++
|
||||
}
|
||||
}
|
||||
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"\ntraces replayed: %d (rejected: %d)\n",
|
||||
len(reports),
|
||||
rejected,
|
||||
)
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"validity: %d of %d reproduced the recorded verdicts exactly\n",
|
||||
len(reports)-invalid,
|
||||
len(reports),
|
||||
)
|
||||
fmt.Fprintf(out, "traces where crash-only fired: %d\n", crashed)
|
||||
fmt.Fprintf(out, "engine refutations: %d\n", engineRefutations)
|
||||
if engineRefutations > 0 {
|
||||
fmt.Fprintf(out, "weakest refuting oracle: %s\n", counts(weakest))
|
||||
for _, class := range sortedKeys(byClass) {
|
||||
fmt.Fprintf(out, " %s: %s\n", class, counts(byClass[class]))
|
||||
}
|
||||
fmt.Fprintf(out, "temporal-only fraction: %.3f\n",
|
||||
float64(weakest["temporal-only"])/float64(engineRefutations))
|
||||
}
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"properties a reduced oracle cannot express: %s\n",
|
||||
counts(distinct(inexpressible)),
|
||||
)
|
||||
fmt.Fprintf(
|
||||
out,
|
||||
"reduced-oracle refutations the engine did not make: %d\n",
|
||||
unmatched,
|
||||
)
|
||||
}
|
||||
|
||||
// recordInexpressible counts a property once per oracle however many traces it
|
||||
// appears on, because whether an oracle can state a property is a fact about
|
||||
// the property and not about the run.
|
||||
func recordInexpressible(
|
||||
seen map[string]map[string]bool,
|
||||
oracle string,
|
||||
finding refutation,
|
||||
property string,
|
||||
) {
|
||||
if !finding.CannotExpress {
|
||||
return
|
||||
}
|
||||
if seen[oracle] == nil {
|
||||
seen[oracle] = map[string]bool{}
|
||||
}
|
||||
seen[oracle][property] = true
|
||||
}
|
||||
|
||||
func distinct(seen map[string]map[string]bool) map[string]int {
|
||||
sizes := map[string]int{}
|
||||
for oracle, properties := range seen {
|
||||
sizes[oracle] = len(properties)
|
||||
}
|
||||
return sizes
|
||||
}
|
||||
|
||||
func counts(values map[string]int) string {
|
||||
parts := make([]string, 0, len(values))
|
||||
for _, key := range sortedKeys(values) {
|
||||
parts = append(parts, fmt.Sprintf("%s=%d", key, values[key]))
|
||||
}
|
||||
return strings.Join(parts, " ")
|
||||
}
|
||||
|
||||
func sortedKeys[V any](values map[string]V) []string {
|
||||
keys := make([]string, 0, len(values))
|
||||
for key := range values {
|
||||
keys = append(keys, key)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
return keys
|
||||
}
|
||||
@@ -0,0 +1,549 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"sort"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/ltl"
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
// foregroundLossReason is the skip reason internal/runner records when the app
|
||||
// was no longer the foreground process at action time. It is the only
|
||||
// foreground signal a stored trace carries, and it is what the crash-only
|
||||
// oracle reads for "the application is no longer the foreground process".
|
||||
const foregroundLossReason = "app_left_foreground"
|
||||
|
||||
// refutation is one oracle's finding for one property on one trace. A reduced
|
||||
// oracle has three of them and not two: CannotExpress says its rewrite no
|
||||
// longer states the property, which is neither a refutation nor a clean bill.
|
||||
type refutation struct {
|
||||
Refuted bool `json:"refuted"`
|
||||
CannotExpress bool `json:"cannot_express,omitempty"`
|
||||
// Step is the observation whose evaluation produced the violation;
|
||||
// OriginStep is the observation whose obligation failed, which for a
|
||||
// deferred one is earlier.
|
||||
Step int `json:"step,omitempty"`
|
||||
OriginStep int `json:"origin_step,omitempty"`
|
||||
Reason string `json:"reason,omitempty"`
|
||||
IsError bool `json:"is_error,omitempty"`
|
||||
}
|
||||
|
||||
type propertyReport struct {
|
||||
Property string `json:"property"`
|
||||
Class string `json:"class"`
|
||||
TopLevel string `json:"top_level"`
|
||||
Engine refutation `json:"engine"`
|
||||
SingleState refutation `json:"single_state"`
|
||||
SingleStep refutation `json:"single_step"`
|
||||
// SingleStepTruncatesWindow marks a property whose window the triple had to
|
||||
// shorten to two observations, so what its single-step column refutes is a
|
||||
// stronger property than the one the author wrote.
|
||||
SingleStepTruncatesWindow bool `json:"single_step_truncates_window,omitempty"`
|
||||
// Weakest names the weakest oracle that refutes this property on this
|
||||
// trace, and is empty unless the engine refuted it. "temporal-only" means
|
||||
// no reduced oracle did.
|
||||
Weakest string `json:"weakest_refuting_oracle,omitempty"`
|
||||
}
|
||||
|
||||
type crashReport struct {
|
||||
Fired bool `json:"fired"`
|
||||
FirstStep int `json:"first_step,omitempty"`
|
||||
ExceptionSteps []int `json:"exception_steps,omitempty"`
|
||||
ForegroundLossSteps []int `json:"foreground_loss_steps,omitempty"`
|
||||
ErrorLogSteps []int `json:"error_log_steps,omitempty"`
|
||||
}
|
||||
|
||||
// mismatch is one disagreement between the offline engine and the verdicts the
|
||||
// run recorded. Any of these blocks the trace's result.
|
||||
type mismatch struct {
|
||||
Property string `json:"property"`
|
||||
Field string `json:"field"`
|
||||
Online string `json:"online"`
|
||||
Offline string `json:"offline"`
|
||||
}
|
||||
|
||||
type runReport struct {
|
||||
Run string `json:"run"`
|
||||
Seed int64 `json:"seed"`
|
||||
Platform string `json:"platform"`
|
||||
Arm string `json:"arm,omitempty"`
|
||||
Generator string `json:"generator,omitempty"`
|
||||
// ReplayMode says where the extractor values under replay came from:
|
||||
// "page-extractor-values" reuses what the page computed in V8 and the run
|
||||
// evaluated against, "host-extractors" re-runs the spec's getters over the
|
||||
// stored hierarchy.
|
||||
ReplayMode string `json:"replay_mode"`
|
||||
StepsObserved int `json:"steps_observed"`
|
||||
StepsSkipped int `json:"steps_skipped"`
|
||||
Valid bool `json:"valid"`
|
||||
Mismatches []mismatch `json:"mismatches,omitempty"`
|
||||
// ResidualMismatches counts every step whose replayed residual formula
|
||||
// differed from the recorded one, of which Mismatches carries the first
|
||||
// few. The recorded residual is the engine's whole pending state, so
|
||||
// agreement on it is a stronger claim than agreement on the verdicts.
|
||||
ResidualMismatches int `json:"residual_mismatches"`
|
||||
CrashOnly crashReport `json:"crash_only"`
|
||||
Properties []propertyReport `json:"properties"`
|
||||
ActionsUnbounded int `json:"actions_without_resolved_bounds"`
|
||||
ScrollActions int `json:"scroll_last_actions"`
|
||||
}
|
||||
|
||||
// witnessRecord is one violation as either side reports it: the step it was
|
||||
// recorded at, plus the witness the engine attached.
|
||||
type witnessRecord struct {
|
||||
RecordedStep int
|
||||
OriginStep int
|
||||
DetectedStep int
|
||||
Reason string
|
||||
IsError bool
|
||||
}
|
||||
|
||||
// replay re-evaluates one run offline under all four oracles. The engine's
|
||||
// offline verdicts are compared against the ones the run recorded, and the
|
||||
// report is marked invalid on any disagreement: a mismatch is a bug here or a
|
||||
// gap in the trace, not a finding.
|
||||
func replay(
|
||||
run loadedRun,
|
||||
bundleJavaScript string,
|
||||
hostExtractors bool,
|
||||
) (runReport, error) {
|
||||
engine, err := verifier.New(
|
||||
verifier.WithSeed(uint64(run.Meta.Seed)),
|
||||
verifier.WithPlatform(run.Meta.Platform),
|
||||
verifier.WithAppPackage(run.Meta.BundleID),
|
||||
)
|
||||
if err != nil {
|
||||
return runReport{}, fmt.Errorf("verifier: %w", err)
|
||||
}
|
||||
if err := engine.Load(bundleJavaScript); err != nil {
|
||||
return runReport{}, fmt.Errorf("load spec: %w", err)
|
||||
}
|
||||
formulas, err := engine.PropertyFormulas()
|
||||
if err != nil {
|
||||
return runReport{}, err
|
||||
}
|
||||
|
||||
singleState := map[string]*ltl.Evaluator{}
|
||||
singleStep := map[string]*ltl.Evaluator{}
|
||||
for name, formula := range formulas {
|
||||
singleState[name] = ltl.NewEvaluator(singleStateFormula(formula))
|
||||
singleStep[name] = ltl.NewEvaluator(singleStepFormula(formula))
|
||||
}
|
||||
|
||||
usePageValues := run.Meta.Platform == "web" && !hostExtractors
|
||||
report := runReport{
|
||||
Run: run.Directory,
|
||||
Seed: run.Meta.Seed,
|
||||
Platform: run.Meta.Platform,
|
||||
Arm: run.Meta.Arm,
|
||||
Generator: run.Meta.Generator,
|
||||
ReplayMode: "host-extractors",
|
||||
}
|
||||
var folded []map[int]json.RawMessage
|
||||
if usePageValues {
|
||||
report.ReplayMode = "page-extractor-values"
|
||||
folded, err = extractorFold(run.Steps, engine.ExtractorNames())
|
||||
if err != nil {
|
||||
return runReport{}, err
|
||||
}
|
||||
}
|
||||
|
||||
offline := map[string]witnessRecord{}
|
||||
stateFired := map[string]int{}
|
||||
stepFired := map[string]int{}
|
||||
var residualMismatches []mismatch
|
||||
var lastAction *verifier.Action
|
||||
for position, step := range run.Steps {
|
||||
if step.NextAction != nil && step.NextAction.Selector != "" &&
|
||||
step.NextAction.ResolvedBounds == nil {
|
||||
report.ActionsUnbounded++
|
||||
}
|
||||
if step.SkippedVerification {
|
||||
report.StepsSkipped++
|
||||
residualMismatches = append(
|
||||
residualMismatches,
|
||||
compareResiduals(step, engine.Residuals())...)
|
||||
lastAction = lastActionFor(step)
|
||||
continue
|
||||
}
|
||||
if lastAction != nil && lastAction.Kind == verifier.ActionKindScroll {
|
||||
report.ScrollActions++
|
||||
}
|
||||
if err := engine.PushSnapshot(verifier.SnapshotInput{
|
||||
Tree: step.Hierarchy,
|
||||
LastAction: lastAction,
|
||||
StepTime: step.Timestamp,
|
||||
StepIndex: step.Index,
|
||||
RunStart: run.Meta.StartedAt,
|
||||
Logs: traceLogs(step.Logs),
|
||||
Exceptions: traceExceptions(step.Exceptions),
|
||||
}); err != nil {
|
||||
return runReport{}, fmt.Errorf("step %d push: %w", step.Index, err)
|
||||
}
|
||||
if usePageValues {
|
||||
skipped, overrideErr := engine.OverrideExtractorValues(
|
||||
folded[position],
|
||||
)
|
||||
if overrideErr != nil {
|
||||
return runReport{}, fmt.Errorf(
|
||||
"step %d override: %w",
|
||||
step.Index,
|
||||
overrideErr,
|
||||
)
|
||||
}
|
||||
if skipped > 0 {
|
||||
return runReport{}, fmt.Errorf(
|
||||
"step %d: %d recorded extractor values fell outside the spec's extractor list",
|
||||
step.Index,
|
||||
skipped,
|
||||
)
|
||||
}
|
||||
}
|
||||
engine.EvaluateProperties()
|
||||
for _, name := range engine.NewlyViolatedProperties() {
|
||||
offline[name] = witnessFrom(engine.Witness(name), step.Index)
|
||||
}
|
||||
residualMismatches = append(
|
||||
residualMismatches,
|
||||
compareResiduals(step, engine.Residuals())...)
|
||||
for name := range formulas {
|
||||
recordFiring(
|
||||
stateFired,
|
||||
name,
|
||||
singleState[name].ObserveAtStep(step.Timestamp, step.Index),
|
||||
step.Index,
|
||||
)
|
||||
recordFiring(
|
||||
stepFired,
|
||||
name,
|
||||
singleStep[name].ObserveAtStep(step.Timestamp, step.Index),
|
||||
step.Index,
|
||||
)
|
||||
}
|
||||
report.StepsObserved++
|
||||
lastAction = lastActionFor(step)
|
||||
}
|
||||
|
||||
finalizeIndex := run.finalizeIndex()
|
||||
for _, name := range engine.Finalize() {
|
||||
offline[name] = witnessFrom(engine.Witness(name), finalizeIndex)
|
||||
}
|
||||
for name := range formulas {
|
||||
recordFiring(
|
||||
stateFired,
|
||||
name,
|
||||
singleState[name].Finalize(),
|
||||
finalizeIndex,
|
||||
)
|
||||
recordFiring(
|
||||
stepFired,
|
||||
name,
|
||||
singleStep[name].Finalize(),
|
||||
finalizeIndex,
|
||||
)
|
||||
}
|
||||
|
||||
report.CrashOnly = crashOnly(run)
|
||||
report.Mismatches = compareVerdicts(onlineVerdicts(run), offline)
|
||||
report.ResidualMismatches = len(residualMismatches)
|
||||
if len(residualMismatches) > reportedResiduals {
|
||||
residualMismatches = residualMismatches[:reportedResiduals]
|
||||
}
|
||||
report.Mismatches = append(report.Mismatches, residualMismatches...)
|
||||
report.Valid = len(report.Mismatches) == 0
|
||||
|
||||
names := make([]string, 0, len(formulas))
|
||||
for name := range formulas {
|
||||
names = append(names, name)
|
||||
}
|
||||
sort.Strings(names)
|
||||
for _, name := range names {
|
||||
property := propertyReport{
|
||||
Property: name,
|
||||
Class: propertyClass(formulas[name]),
|
||||
TopLevel: topLevelForm(formulas[name]),
|
||||
Engine: engineRefutation(offline, name),
|
||||
SingleState: reducedRefutation(
|
||||
singleStateExpresses(formulas[name]),
|
||||
singleState[name],
|
||||
stateFired[name],
|
||||
),
|
||||
SingleStep: reducedRefutation(
|
||||
singleStepExpresses(formulas[name]),
|
||||
singleStep[name],
|
||||
stepFired[name],
|
||||
),
|
||||
|
||||
SingleStepTruncatesWindow: truncatesWindow(formulas[name]),
|
||||
}
|
||||
property.Weakest = weakestOracle(property, report.CrashOnly)
|
||||
report.Properties = append(report.Properties, property)
|
||||
}
|
||||
return report, nil
|
||||
}
|
||||
|
||||
// reportedResiduals caps how many residual differences one trace lists. The
|
||||
// count of all of them is reported alongside; the first few are what says
|
||||
// where the two engines parted.
|
||||
const reportedResiduals = 5
|
||||
|
||||
// compareResiduals checks the replayed pending state against the one the step
|
||||
// recorded. A step the verifier skipped still recorded the residual it was
|
||||
// holding, so those steps assert that a skipped step advanced nothing.
|
||||
func compareResiduals(
|
||||
step trace.Step,
|
||||
replayed map[string]ltl.Formula,
|
||||
) []mismatch {
|
||||
if len(step.Residuals) == 0 {
|
||||
return nil
|
||||
}
|
||||
var mismatches []mismatch
|
||||
names := make([]string, 0, len(step.Residuals))
|
||||
for name := range step.Residuals {
|
||||
names = append(names, name)
|
||||
}
|
||||
sort.Strings(names)
|
||||
for _, name := range names {
|
||||
formula, ok := replayed[name]
|
||||
if !ok {
|
||||
mismatches = append(mismatches, mismatch{
|
||||
Property: name,
|
||||
Field: fmt.Sprintf("residual at step %d", step.Index),
|
||||
Online: string(step.Residuals[name]),
|
||||
Offline: "property not registered",
|
||||
})
|
||||
continue
|
||||
}
|
||||
encoded, err := json.Marshal(formula)
|
||||
if err != nil {
|
||||
encoded = []byte(fmt.Sprintf("%q", err.Error()))
|
||||
}
|
||||
if !bytes.Equal(encoded, step.Residuals[name]) {
|
||||
mismatches = append(mismatches, mismatch{
|
||||
Property: name,
|
||||
Field: fmt.Sprintf("residual at step %d", step.Index),
|
||||
Online: string(step.Residuals[name]),
|
||||
Offline: string(encoded),
|
||||
})
|
||||
}
|
||||
}
|
||||
return mismatches
|
||||
}
|
||||
|
||||
func witnessFrom(witness *verifier.Witness, recordedStep int) witnessRecord {
|
||||
record := witnessRecord{RecordedStep: recordedStep}
|
||||
if witness == nil {
|
||||
return record
|
||||
}
|
||||
record.OriginStep = witness.Step
|
||||
record.DetectedStep = witness.DetectedStep
|
||||
record.Reason = witness.Reason
|
||||
record.IsError = witness.IsError
|
||||
return record
|
||||
}
|
||||
|
||||
// onlineVerdicts reads the violations the run recorded, including the
|
||||
// end-of-run record a finalized liveness obligation is written to.
|
||||
func onlineVerdicts(run loadedRun) map[string]witnessRecord {
|
||||
recorded := map[string]witnessRecord{}
|
||||
steps := run.Steps
|
||||
if run.Finalize != nil {
|
||||
steps = append(append([]trace.Step(nil), steps...), *run.Finalize)
|
||||
}
|
||||
for _, step := range steps {
|
||||
for _, name := range step.Violations {
|
||||
record := witnessRecord{RecordedStep: step.Index}
|
||||
if witness, ok := step.Witnesses[name]; ok {
|
||||
record.OriginStep = witness.Step
|
||||
record.DetectedStep = witness.DetectedStep
|
||||
record.Reason = witness.Reason
|
||||
record.IsError = witness.IsError
|
||||
}
|
||||
recorded[name] = record
|
||||
}
|
||||
}
|
||||
return recorded
|
||||
}
|
||||
|
||||
func compareVerdicts(online, offline map[string]witnessRecord) []mismatch {
|
||||
var mismatches []mismatch
|
||||
names := map[string]bool{}
|
||||
for name := range online {
|
||||
names[name] = true
|
||||
}
|
||||
for name := range offline {
|
||||
names[name] = true
|
||||
}
|
||||
ordered := make([]string, 0, len(names))
|
||||
for name := range names {
|
||||
ordered = append(ordered, name)
|
||||
}
|
||||
sort.Strings(ordered)
|
||||
|
||||
for _, name := range ordered {
|
||||
recorded, wasRecorded := online[name]
|
||||
replayed, wasReplayed := offline[name]
|
||||
switch {
|
||||
case wasRecorded && !wasReplayed:
|
||||
mismatches = append(mismatches, mismatch{
|
||||
Property: name, Field: "violated",
|
||||
Online: fmt.Sprintf(
|
||||
"violated at step %d",
|
||||
recorded.RecordedStep,
|
||||
),
|
||||
Offline: "not violated",
|
||||
})
|
||||
continue
|
||||
case !wasRecorded && wasReplayed:
|
||||
mismatches = append(mismatches, mismatch{
|
||||
Property: name, Field: "violated",
|
||||
Online: "not violated",
|
||||
Offline: fmt.Sprintf(
|
||||
"violated at step %d",
|
||||
replayed.RecordedStep,
|
||||
),
|
||||
})
|
||||
continue
|
||||
case !wasRecorded:
|
||||
continue
|
||||
}
|
||||
for _, field := range []struct {
|
||||
name string
|
||||
online string
|
||||
offline string
|
||||
}{
|
||||
{"step", fmt.Sprint(recorded.RecordedStep), fmt.Sprint(replayed.RecordedStep)},
|
||||
{"origin_step", fmt.Sprint(recorded.OriginStep), fmt.Sprint(replayed.OriginStep)},
|
||||
{"detected_step", fmt.Sprint(recorded.DetectedStep), fmt.Sprint(replayed.DetectedStep)},
|
||||
{"reason", recorded.Reason, replayed.Reason},
|
||||
} {
|
||||
if field.online != field.offline {
|
||||
mismatches = append(mismatches, mismatch{
|
||||
Property: name, Field: field.name,
|
||||
Online: field.online, Offline: field.offline,
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
return mismatches
|
||||
}
|
||||
|
||||
// crashOnly fires where the application left the foreground or an error
|
||||
// surface was recorded. Error-level log lines are reported alongside rather
|
||||
// than folded in: a console error is not a crash, and an analysis that wants
|
||||
// the looser detector can read the steps from here.
|
||||
func crashOnly(run loadedRun) crashReport {
|
||||
report := crashReport{}
|
||||
for _, step := range run.Steps {
|
||||
if len(step.Exceptions) > 0 {
|
||||
report.ExceptionSteps = append(report.ExceptionSteps, step.Index)
|
||||
}
|
||||
if step.ActionSkipped == foregroundLossReason {
|
||||
report.ForegroundLossSteps = append(
|
||||
report.ForegroundLossSteps,
|
||||
step.Index,
|
||||
)
|
||||
}
|
||||
for _, entry := range step.Logs {
|
||||
if entry.Level == "E" || entry.Level == "F" {
|
||||
report.ErrorLogSteps = append(report.ErrorLogSteps, step.Index)
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
first := 0
|
||||
for _, step := range append(append([]int(nil), report.ExceptionSteps...), report.ForegroundLossSteps...) {
|
||||
if first == 0 || step < first {
|
||||
first = step
|
||||
}
|
||||
}
|
||||
report.Fired = first != 0
|
||||
report.FirstStep = first
|
||||
return report
|
||||
}
|
||||
|
||||
func engineRefutation(
|
||||
offline map[string]witnessRecord,
|
||||
name string,
|
||||
) refutation {
|
||||
record, ok := offline[name]
|
||||
if !ok {
|
||||
return refutation{}
|
||||
}
|
||||
return refutation{
|
||||
Refuted: true,
|
||||
Step: record.RecordedStep,
|
||||
OriginStep: record.OriginStep,
|
||||
Reason: record.Reason,
|
||||
IsError: record.IsError,
|
||||
}
|
||||
}
|
||||
|
||||
// reducedRefutation reports a reduced oracle's finding, or its inability to
|
||||
// state the property at all. The evaluator is driven either way so that the
|
||||
// verdict a reduction would have reached is never what decides whether it was
|
||||
// entitled to reach one.
|
||||
func reducedRefutation(
|
||||
expresses bool,
|
||||
evaluator *ltl.Evaluator,
|
||||
firedAt int,
|
||||
) refutation {
|
||||
if !expresses {
|
||||
return refutation{CannotExpress: true}
|
||||
}
|
||||
return evaluatorRefutation(evaluator, firedAt)
|
||||
}
|
||||
|
||||
func evaluatorRefutation(evaluator *ltl.Evaluator, firedAt int) refutation {
|
||||
violation := evaluator.Violation()
|
||||
if violation == nil {
|
||||
return refutation{}
|
||||
}
|
||||
return refutation{
|
||||
Refuted: true,
|
||||
Step: firedAt,
|
||||
OriginStep: violation.Step,
|
||||
Reason: violation.Reason,
|
||||
IsError: violation.IsError,
|
||||
}
|
||||
}
|
||||
|
||||
// recordFiring keeps the first observation at which a reduced oracle latched,
|
||||
// which the evaluator itself does not carry.
|
||||
func recordFiring(
|
||||
fired map[string]int,
|
||||
name string,
|
||||
verdict ltl.Verdict,
|
||||
step int,
|
||||
) {
|
||||
if verdict != ltl.VerdictViolated {
|
||||
return
|
||||
}
|
||||
if _, ok := fired[name]; ok {
|
||||
return
|
||||
}
|
||||
fired[name] = step
|
||||
}
|
||||
|
||||
// weakestOracle names the weakest oracle that refutes a defect the engine
|
||||
// refuted, in the order a crash detector, a single-state check and a
|
||||
// single-step triple are weaker than the engine.
|
||||
func weakestOracle(property propertyReport, crash crashReport) string {
|
||||
if !property.Engine.Refuted {
|
||||
return ""
|
||||
}
|
||||
switch {
|
||||
case crash.Fired:
|
||||
return "crash-only"
|
||||
case property.SingleState.Refuted:
|
||||
return "single-state"
|
||||
case property.SingleStep.Refuted:
|
||||
return "single-step"
|
||||
default:
|
||||
return "temporal-only"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,433 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/hierarchy"
|
||||
"github.com/priyanshujain/sanderling/internal/trace"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
// fixtureSpec counts rows on screen and asserts three things about them: a
|
||||
// bound a single observation can check, a reachability goal, and a growth step
|
||||
// that only two consecutive observations can see.
|
||||
const fixtureSpec = `
|
||||
const rows = __sanderling__.extract((s) => s.ax.findAll({ text: "row" }).length).named("rows");
|
||||
globalThis.properties = {
|
||||
fewRows: __sanderling__.always(() => rows.current < 3),
|
||||
rowsAppear: __sanderling__.eventually(() => rows.current > 0).within(500, "steps"),
|
||||
rowsKeepGrowing: __sanderling__.always(
|
||||
__sanderling__.now(() => rows.current === 1).implies(
|
||||
__sanderling__.next(() => rows.current === 2))),
|
||||
};
|
||||
`
|
||||
|
||||
func rowTree(t *testing.T, rows int) *hierarchy.Tree {
|
||||
t.Helper()
|
||||
children := make([]string, 0, rows)
|
||||
for range rows {
|
||||
children = append(
|
||||
children,
|
||||
`{"attributes": {"text": "row", "bounds": "[0,0,10,10]"}, "children": []}`,
|
||||
)
|
||||
}
|
||||
source := fmt.Sprintf(
|
||||
`{"attributes": {"resource-id": "root", "bounds": "[0,0,100,100]"}, "children": [%s]}`,
|
||||
strings.Join(children, ","),
|
||||
)
|
||||
tree, err := hierarchy.Parse(source)
|
||||
if err != nil {
|
||||
t.Fatalf("parse tree: %v", err)
|
||||
}
|
||||
return tree
|
||||
}
|
||||
|
||||
// writeFixtureRun records a run the way internal/runner does: every step's
|
||||
// hierarchy, the violations that fired at it, their witnesses, and the residual
|
||||
// each property was left holding.
|
||||
func writeFixtureRun(t *testing.T, directory string, rowCounts []int) {
|
||||
t.Helper()
|
||||
engine, err := verifier.New(verifier.WithPlatform("android"))
|
||||
if err != nil {
|
||||
t.Fatalf("verifier: %v", err)
|
||||
}
|
||||
if err := engine.Load(fixtureSpec); err != nil {
|
||||
t.Fatalf("load spec: %v", err)
|
||||
}
|
||||
writer, err := trace.NewWriter(directory)
|
||||
if err != nil {
|
||||
t.Fatalf("trace writer: %v", err)
|
||||
}
|
||||
defer writer.Close()
|
||||
|
||||
runStart := time.Date(2026, 8, 15, 12, 0, 0, 0, time.UTC)
|
||||
if err := writer.WriteMeta(trace.Meta{
|
||||
Seed: 11,
|
||||
SpecPath: "fixture.ts",
|
||||
Platform: "android",
|
||||
BundleID: "com.example.fixture",
|
||||
StartedAt: runStart,
|
||||
}); err != nil {
|
||||
t.Fatalf("meta: %v", err)
|
||||
}
|
||||
|
||||
lastIndex := 0
|
||||
for position, rows := range rowCounts {
|
||||
index := position + 1
|
||||
lastIndex = index
|
||||
stepTime := runStart.Add(time.Duration(index) * time.Second)
|
||||
tree := rowTree(t, rows)
|
||||
if err := engine.PushSnapshot(verifier.SnapshotInput{
|
||||
Tree: tree,
|
||||
StepTime: stepTime,
|
||||
StepIndex: index,
|
||||
RunStart: runStart,
|
||||
}); err != nil {
|
||||
t.Fatalf("push step %d: %v", index, err)
|
||||
}
|
||||
engine.EvaluateProperties()
|
||||
violations := engine.NewlyViolatedProperties()
|
||||
step := trace.Step{
|
||||
Index: index,
|
||||
Timestamp: stepTime,
|
||||
Hierarchy: tree,
|
||||
Violations: violations,
|
||||
Witnesses: fixtureWitnesses(engine, violations, index),
|
||||
Residuals: fixtureResiduals(t, engine),
|
||||
}
|
||||
if err := writer.WriteStep(step); err != nil {
|
||||
t.Fatalf("write step %d: %v", index, err)
|
||||
}
|
||||
}
|
||||
if ended := engine.Finalize(); len(ended) > 0 {
|
||||
if err := writer.WriteStep(trace.Step{
|
||||
Index: lastIndex + 1,
|
||||
Timestamp: runStart.Add(time.Duration(lastIndex+1) * time.Second),
|
||||
Violations: ended,
|
||||
Witnesses: fixtureWitnesses(engine, ended, lastIndex+1),
|
||||
}); err != nil {
|
||||
t.Fatalf("write finalize step: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func fixtureWitnesses(
|
||||
engine *verifier.Verifier,
|
||||
properties []string,
|
||||
index int,
|
||||
) map[string]trace.Witness {
|
||||
if len(properties) == 0 {
|
||||
return nil
|
||||
}
|
||||
witnesses := map[string]trace.Witness{}
|
||||
for _, name := range properties {
|
||||
witness := engine.Witness(name)
|
||||
if witness == nil {
|
||||
continue
|
||||
}
|
||||
detected := witness.DetectedStep
|
||||
if detected == 0 {
|
||||
detected = index
|
||||
}
|
||||
witnesses[name] = trace.Witness{
|
||||
Reason: witness.Reason,
|
||||
IsError: witness.IsError,
|
||||
Step: witness.Step,
|
||||
DetectedStep: detected,
|
||||
Extractors: witness.Extractors,
|
||||
}
|
||||
}
|
||||
return witnesses
|
||||
}
|
||||
|
||||
func fixtureResiduals(
|
||||
t *testing.T,
|
||||
engine *verifier.Verifier,
|
||||
) map[string]json.RawMessage {
|
||||
t.Helper()
|
||||
residuals := map[string]json.RawMessage{}
|
||||
for name, formula := range engine.Residuals() {
|
||||
body, err := json.Marshal(formula)
|
||||
if err != nil {
|
||||
t.Fatalf("marshal residual %q: %v", name, err)
|
||||
}
|
||||
residuals[name] = body
|
||||
}
|
||||
return residuals
|
||||
}
|
||||
|
||||
func replayFixture(t *testing.T, directory string) runReport {
|
||||
t.Helper()
|
||||
loaded, err := loadRun(directory)
|
||||
if err != nil {
|
||||
t.Fatalf("load run: %v", err)
|
||||
}
|
||||
report, err := replay(loaded, fixtureSpec, false)
|
||||
if err != nil {
|
||||
t.Fatalf("replay: %v", err)
|
||||
}
|
||||
return report
|
||||
}
|
||||
|
||||
func propertyByName(
|
||||
t *testing.T,
|
||||
report runReport,
|
||||
name string,
|
||||
) propertyReport {
|
||||
t.Helper()
|
||||
for _, property := range report.Properties {
|
||||
if property.Property == name {
|
||||
return property
|
||||
}
|
||||
}
|
||||
t.Fatalf("property %q missing from the report", name)
|
||||
return propertyReport{}
|
||||
}
|
||||
|
||||
func TestReplayReproducesTheRecordedVerdicts(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
|
||||
|
||||
report := replayFixture(t, directory)
|
||||
|
||||
if !report.Valid {
|
||||
t.Fatalf("replay disagreed with the run: %+v", report.Mismatches)
|
||||
}
|
||||
if report.ResidualMismatches != 0 {
|
||||
t.Errorf("residual mismatches: %d", report.ResidualMismatches)
|
||||
}
|
||||
if report.StepsObserved != 4 {
|
||||
t.Errorf("steps observed: got %d, want 4", report.StepsObserved)
|
||||
}
|
||||
|
||||
growth := propertyByName(t, report, "rowsKeepGrowing")
|
||||
if !growth.Engine.Refuted || growth.Engine.Step != 3 {
|
||||
t.Errorf("engine on rowsKeepGrowing: %+v", growth.Engine)
|
||||
}
|
||||
if growth.SingleState.Refuted {
|
||||
t.Errorf(
|
||||
"single-state refuted a growth step it cannot see: %+v",
|
||||
growth.SingleState,
|
||||
)
|
||||
}
|
||||
if !growth.SingleStep.Refuted {
|
||||
t.Error("single-step did not refute a one-step obligation")
|
||||
}
|
||||
if growth.Weakest != "single-step" {
|
||||
t.Errorf("weakest oracle for rowsKeepGrowing: got %q", growth.Weakest)
|
||||
}
|
||||
|
||||
bound := propertyByName(t, report, "fewRows")
|
||||
if !bound.SingleState.Refuted || bound.Weakest != "single-state" {
|
||||
t.Errorf(
|
||||
"fewRows: single-state %+v weakest %q",
|
||||
bound.SingleState,
|
||||
bound.Weakest,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReplayReportsAViolationTheRunNeverRecorded(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
|
||||
dropRecordedViolation(t, directory, "rowsKeepGrowing")
|
||||
|
||||
report := replayFixture(t, directory)
|
||||
|
||||
if report.Valid {
|
||||
t.Fatal("a trace missing a recorded violation must not replay as valid")
|
||||
}
|
||||
var found bool
|
||||
for _, entry := range report.Mismatches {
|
||||
if entry.Property == "rowsKeepGrowing" && entry.Field == "violated" {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Errorf(
|
||||
"no mismatch named the dropped violation: %+v",
|
||||
report.Mismatches,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReplayReportsAResidualTheRunNeverHeld(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
|
||||
rewriteResidual(t, directory, 1, "fewRows", `{"op":"false"}`)
|
||||
|
||||
report := replayFixture(t, directory)
|
||||
|
||||
if report.Valid {
|
||||
t.Fatal(
|
||||
"a trace whose recorded residual differs must not replay as valid",
|
||||
)
|
||||
}
|
||||
if report.ResidualMismatches != 1 {
|
||||
t.Errorf(
|
||||
"residual mismatches: got %d, want 1",
|
||||
report.ResidualMismatches,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// dropRecordedViolation removes one property from every step's violations,
|
||||
// which is what a trace that failed to record a verdict looks like.
|
||||
func dropRecordedViolation(t *testing.T, directory, property string) {
|
||||
t.Helper()
|
||||
rewriteSteps(t, directory, func(step *trace.Step) {
|
||||
kept := step.Violations[:0]
|
||||
for _, name := range step.Violations {
|
||||
if name != property {
|
||||
kept = append(kept, name)
|
||||
}
|
||||
}
|
||||
step.Violations = kept
|
||||
delete(step.Witnesses, property)
|
||||
})
|
||||
}
|
||||
|
||||
func rewriteResidual(
|
||||
t *testing.T,
|
||||
directory string,
|
||||
index int,
|
||||
property, residual string,
|
||||
) {
|
||||
t.Helper()
|
||||
rewriteSteps(t, directory, func(step *trace.Step) {
|
||||
if step.Index == index {
|
||||
step.Residuals[property] = json.RawMessage(residual)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
func rewriteSteps(t *testing.T, directory string, edit func(*trace.Step)) {
|
||||
t.Helper()
|
||||
path := filepath.Join(directory, "trace.jsonl")
|
||||
body, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("read trace: %v", err)
|
||||
}
|
||||
var rewritten strings.Builder
|
||||
for _, line := range strings.Split(strings.TrimSpace(string(body)), "\n") {
|
||||
var step trace.Step
|
||||
if err := json.Unmarshal([]byte(line), &step); err != nil {
|
||||
t.Fatalf("decode step: %v", err)
|
||||
}
|
||||
edit(&step)
|
||||
encoded, err := json.Marshal(step)
|
||||
if err != nil {
|
||||
t.Fatalf("encode step: %v", err)
|
||||
}
|
||||
rewritten.Write(encoded)
|
||||
rewritten.WriteByte('\n')
|
||||
}
|
||||
if err := os.WriteFile(path, []byte(rewritten.String()), 0o644); err != nil {
|
||||
t.Fatalf("write trace: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCrashOnlyFiresOnAnErrorSurfaceAndOnLeavingTheForeground(t *testing.T) {
|
||||
run := loadedRun{Steps: []trace.Step{
|
||||
{Index: 1, Logs: []trace.LogEntry{{Level: "E", Message: "noisy"}}},
|
||||
{Index: 2, ActionSkipped: foregroundLossReason},
|
||||
{Index: 3, Exceptions: []trace.Exception{{Class: "TypeError"}}},
|
||||
}}
|
||||
|
||||
report := crashOnly(run)
|
||||
|
||||
if !report.Fired || report.FirstStep != 2 {
|
||||
t.Errorf("crash-only: %+v", report)
|
||||
}
|
||||
if len(report.ExceptionSteps) != 1 || report.ExceptionSteps[0] != 3 {
|
||||
t.Errorf("exception steps: %v", report.ExceptionSteps)
|
||||
}
|
||||
if len(report.ErrorLogSteps) != 1 || report.ErrorLogSteps[0] != 1 {
|
||||
t.Errorf("error log steps: %v", report.ErrorLogSteps)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCrashOnlyIgnoresAnErrorLogOnItsOwn(t *testing.T) {
|
||||
run := loadedRun{
|
||||
Steps: []trace.Step{{Index: 1, Logs: []trace.LogEntry{{Level: "E"}}}},
|
||||
}
|
||||
|
||||
if crashOnly(run).Fired {
|
||||
t.Error("an error-level log line is not a crash")
|
||||
}
|
||||
}
|
||||
|
||||
// A reachability goal the run reaches at the fourth observation is clean under
|
||||
// the window its author wrote and refuted by the same window shortened to two
|
||||
// observations. Reporting that shortened refutation would convict every clean
|
||||
// trace, so the single-step column has to admit it cannot state the property.
|
||||
func TestSingleStepDoesNotConvictACleanTraceOfATruncatedWindow(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeFixtureRun(t, directory, []int{0, 0, 0, 1})
|
||||
|
||||
report := replayFixture(t, directory)
|
||||
|
||||
if !report.Valid {
|
||||
t.Fatalf("replay disagreed with the run: %+v", report.Mismatches)
|
||||
}
|
||||
appear := propertyByName(t, report, "rowsAppear")
|
||||
if appear.Engine.Refuted {
|
||||
t.Fatalf(
|
||||
"the trace reaches the goal inside the 500-step window: %+v",
|
||||
appear.Engine,
|
||||
)
|
||||
}
|
||||
if appear.SingleStep.Refuted {
|
||||
t.Errorf(
|
||||
"single-step refuted a clean trace by shortening the window: %+v",
|
||||
appear.SingleStep,
|
||||
)
|
||||
}
|
||||
if !appear.SingleStep.CannotExpress {
|
||||
t.Error(
|
||||
"single-step must record a window it cannot state as inexpressible",
|
||||
)
|
||||
}
|
||||
if !appear.SingleStepTruncatesWindow {
|
||||
t.Error("the truncated-window marker must stay visible on the property")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStateAdmitsWhatOneObservationCannotState(t *testing.T) {
|
||||
directory := t.TempDir()
|
||||
writeFixtureRun(t, directory, []int{0, 1, 1, 3})
|
||||
|
||||
report := replayFixture(t, directory)
|
||||
|
||||
growth := propertyByName(t, report, "rowsKeepGrowing")
|
||||
if !growth.SingleState.CannotExpress {
|
||||
t.Errorf(
|
||||
"single-state kept a next obligation it erases to nothing: %+v",
|
||||
growth.SingleState,
|
||||
)
|
||||
}
|
||||
appear := propertyByName(t, report, "rowsAppear")
|
||||
if !appear.SingleState.CannotExpress {
|
||||
t.Errorf(
|
||||
"single-state kept a reachability goal it erases to nothing: %+v",
|
||||
appear.SingleState,
|
||||
)
|
||||
}
|
||||
bound := propertyByName(t, report, "fewRows")
|
||||
if bound.SingleState.CannotExpress {
|
||||
t.Error("single-state states a bound read at one observation")
|
||||
}
|
||||
if !bound.SingleState.Refuted || bound.Weakest != "single-state" {
|
||||
t.Errorf(
|
||||
"fewRows: single-state %+v weakest %q",
|
||||
bound.SingleState,
|
||||
bound.Weakest,
|
||||
)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,281 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/ltl"
|
||||
)
|
||||
|
||||
// tripleWindow is the horizon a property triple has: an obligation armed at one
|
||||
// observation must discharge at the next, and nothing outlives that.
|
||||
const tripleWindow = 2
|
||||
|
||||
// singleStateFormula is what a checker holding one observation can refute.
|
||||
// Every operator that defers or repeats an obligation is replaced by the
|
||||
// constant that makes the formula around it trivially satisfied, so the only
|
||||
// refutation left is a predicate read at a single observation. The outermost
|
||||
// always survives because it is what says "check this at each observation",
|
||||
// which costs no history.
|
||||
func singleStateFormula(formula ltl.Formula) ltl.Formula {
|
||||
if always, ok := formula.(ltl.AlwaysFormula); ok {
|
||||
always.Inner = stateless(always.Inner, true)
|
||||
return always
|
||||
}
|
||||
return stateless(formula, true)
|
||||
}
|
||||
|
||||
// singleStepFormula is the property triple: one step of history, and no
|
||||
// obligation surviving past the next observation. A next keeps its one-step
|
||||
// deferral, and an eventually of any window shrinks to the two observations a
|
||||
// triple spans.
|
||||
func singleStepFormula(formula ltl.Formula) ltl.Formula {
|
||||
if always, ok := formula.(ltl.AlwaysFormula); ok {
|
||||
always.Inner = oneStep(always.Inner, true)
|
||||
return always
|
||||
}
|
||||
return oneStep(formula, true)
|
||||
}
|
||||
|
||||
// singleStateExpresses and singleStepExpresses decide, from the property's form
|
||||
// alone, whether the reduced oracle can still state the property after its
|
||||
// rewrite. An oracle that cannot state a property does not get to refute it,
|
||||
// and is reported as silent on it rather than as either verdict.
|
||||
//
|
||||
// Two shapes defeat a reduction. The rewrite shortens an obligation window the
|
||||
// oracle's horizon cannot hold, leaving it to refute a property strictly
|
||||
// stronger than the one the author wrote. Or the rewrite leaves a formula whose
|
||||
// verdict no longer depends on the trace, leaving it to report the same answer
|
||||
// everywhere. Either way the column would carry a verdict about a property
|
||||
// nobody wrote, which is worth less than an admission that the oracle is out of
|
||||
// its depth.
|
||||
func singleStateExpresses(formula ltl.Formula) bool {
|
||||
return dependsOnTrace(singleStateFormula(formula))
|
||||
}
|
||||
|
||||
func singleStepExpresses(formula ltl.Formula) bool {
|
||||
return !truncatesWindow(formula) &&
|
||||
dependsOnTrace(singleStepFormula(formula))
|
||||
}
|
||||
|
||||
// truncatesWindow reports whether the single-step rewrite had to shorten a
|
||||
// window to fit a triple. Where it did, the reduced oracle is checking a
|
||||
// stronger property than the author wrote, so a refutation of it is not the
|
||||
// same event as a refutation of the property.
|
||||
func truncatesWindow(formula ltl.Formula) bool {
|
||||
if always, ok := formula.(ltl.AlwaysFormula); ok {
|
||||
return shortensWindow(always.Inner, true)
|
||||
}
|
||||
return shortensWindow(formula, true)
|
||||
}
|
||||
|
||||
// shortensWindow walks the same shape oneStep rewrites, and reports only the
|
||||
// shortening that strengthens the formula. Polarity is what separates the two:
|
||||
// a shorter window under an even number of negations demands the same thing
|
||||
// sooner, so a refutation of it need not be a refutation of the property, while
|
||||
// under an odd number it asks for less and its refutations stay sound. The
|
||||
// operators oneStep erases rather than shortens are erased to the constant its
|
||||
// position is satisfied by, which also only asks for less.
|
||||
func shortensWindow(formula ltl.Formula, positive bool) bool {
|
||||
switch concrete := formula.(type) {
|
||||
case ltl.EventuallyFormula:
|
||||
return positive && windowOutlastsTriple(concrete)
|
||||
case ltl.NowFormula:
|
||||
return shortensWindow(concrete.Inner, positive)
|
||||
case ltl.NotFormula:
|
||||
return shortensWindow(concrete.Inner, !positive)
|
||||
case ltl.AndFormula:
|
||||
return shortensWindow(concrete.Left, positive) || shortensWindow(concrete.Right, positive)
|
||||
case ltl.OrFormula:
|
||||
return shortensWindow(concrete.Left, positive) || shortensWindow(concrete.Right, positive)
|
||||
case ltl.ImpliesFormula:
|
||||
return shortensWindow(concrete.Antecedent, !positive) || shortensWindow(concrete.Consequent, positive)
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
// windowOutlastsTriple asks whether an obligation can still be open after the
|
||||
// two observations a triple spans. A window counted in time is outside what a
|
||||
// triple can state whatever its length, because a triple has no clock: it can
|
||||
// say "at the next observation" and nothing about when that arrives.
|
||||
func windowOutlastsTriple(formula ltl.EventuallyFormula) bool {
|
||||
if formula.HasStepBound {
|
||||
return formula.StepBound > tripleWindow
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// dependsOnTrace reports whether a rewritten formula can still read the trace.
|
||||
// Constants fold through the connectives, and one that folds away to a constant
|
||||
// answers the same on every trace: silence that reads as "did not refute", or a
|
||||
// refutation of everything. Neither is a verdict about the run.
|
||||
func dependsOnTrace(formula ltl.Formula) bool {
|
||||
_, constant := fold(formula).(ltl.PureFormula)
|
||||
return !constant
|
||||
}
|
||||
|
||||
// fold propagates the constants the rewrites substituted for erased operators.
|
||||
// An obligation whose inner formula folded to false keeps its shape, because a
|
||||
// deferred false is a refutation still owed; only the side that can no longer
|
||||
// fail folds away.
|
||||
func fold(formula ltl.Formula) ltl.Formula {
|
||||
switch concrete := formula.(type) {
|
||||
case ltl.AlwaysFormula:
|
||||
concrete.Inner = fold(concrete.Inner)
|
||||
if pure, ok := concrete.Inner.(ltl.PureFormula); ok {
|
||||
return pure
|
||||
}
|
||||
return concrete
|
||||
case ltl.EventuallyFormula:
|
||||
concrete.Inner = fold(concrete.Inner)
|
||||
if isPure(concrete.Inner, true) {
|
||||
return ltl.Pure(true)
|
||||
}
|
||||
return concrete
|
||||
case ltl.NextFormula:
|
||||
inner := fold(concrete.Inner)
|
||||
if isPure(inner, true) {
|
||||
return ltl.Pure(true)
|
||||
}
|
||||
return ltl.Next(inner)
|
||||
case ltl.NowFormula:
|
||||
inner := fold(concrete.Inner)
|
||||
if _, ok := inner.(ltl.PureFormula); ok {
|
||||
return inner
|
||||
}
|
||||
return ltl.Now(inner)
|
||||
case ltl.NotFormula:
|
||||
inner := fold(concrete.Inner)
|
||||
if pure, ok := inner.(ltl.PureFormula); ok {
|
||||
return ltl.Pure(!pure.Value)
|
||||
}
|
||||
return ltl.Not(inner)
|
||||
case ltl.AndFormula:
|
||||
left, right := fold(concrete.Left), fold(concrete.Right)
|
||||
switch {
|
||||
case isPure(left, false) || isPure(right, false):
|
||||
return ltl.Pure(false)
|
||||
case isPure(left, true):
|
||||
return right
|
||||
case isPure(right, true):
|
||||
return left
|
||||
}
|
||||
return ltl.And(left, right)
|
||||
case ltl.OrFormula:
|
||||
left, right := fold(concrete.Left), fold(concrete.Right)
|
||||
switch {
|
||||
case isPure(left, true) || isPure(right, true):
|
||||
return ltl.Pure(true)
|
||||
case isPure(left, false):
|
||||
return right
|
||||
case isPure(right, false):
|
||||
return left
|
||||
}
|
||||
return ltl.Or(left, right)
|
||||
case ltl.ImpliesFormula:
|
||||
antecedent, consequent := fold(concrete.Antecedent), fold(concrete.Consequent)
|
||||
switch {
|
||||
case isPure(antecedent, false) || isPure(consequent, true):
|
||||
return ltl.Pure(true)
|
||||
case isPure(antecedent, true):
|
||||
return consequent
|
||||
}
|
||||
return ltl.Implies(antecedent, consequent)
|
||||
default:
|
||||
return formula
|
||||
}
|
||||
}
|
||||
|
||||
func isPure(formula ltl.Formula, value bool) bool {
|
||||
pure, ok := formula.(ltl.PureFormula)
|
||||
return ok && pure.Value == value
|
||||
}
|
||||
|
||||
// stateless erases every temporal operator. The replacement constant follows
|
||||
// the position's polarity: under an even number of negations a temporal
|
||||
// sub-formula is dropped as satisfied, and under an odd number as failed, so
|
||||
// that in both cases its negation cannot refute anything either.
|
||||
func stateless(formula ltl.Formula, positive bool) ltl.Formula {
|
||||
switch concrete := formula.(type) {
|
||||
case ltl.AlwaysFormula, ltl.NextFormula, ltl.EventuallyFormula:
|
||||
return ltl.Pure(positive)
|
||||
case ltl.NowFormula:
|
||||
return ltl.Now(stateless(concrete.Inner, positive))
|
||||
case ltl.NotFormula:
|
||||
return ltl.Not(stateless(concrete.Inner, !positive))
|
||||
case ltl.AndFormula:
|
||||
return ltl.And(stateless(concrete.Left, positive), stateless(concrete.Right, positive))
|
||||
case ltl.OrFormula:
|
||||
return ltl.Or(stateless(concrete.Left, positive), stateless(concrete.Right, positive))
|
||||
case ltl.ImpliesFormula:
|
||||
return ltl.Implies(
|
||||
stateless(concrete.Antecedent, !positive),
|
||||
stateless(concrete.Consequent, positive),
|
||||
)
|
||||
default:
|
||||
return formula
|
||||
}
|
||||
}
|
||||
|
||||
func oneStep(formula ltl.Formula, positive bool) ltl.Formula {
|
||||
switch concrete := formula.(type) {
|
||||
case ltl.AlwaysFormula:
|
||||
return ltl.Pure(positive)
|
||||
case ltl.NextFormula:
|
||||
return ltl.Next(stateless(concrete.Inner, positive))
|
||||
case ltl.EventuallyFormula:
|
||||
return ltl.EventuallyWithinSteps(stateless(concrete.Inner, positive), tripleWindow)
|
||||
case ltl.NowFormula:
|
||||
return ltl.Now(oneStep(concrete.Inner, positive))
|
||||
case ltl.NotFormula:
|
||||
return ltl.Not(oneStep(concrete.Inner, !positive))
|
||||
case ltl.AndFormula:
|
||||
return ltl.And(oneStep(concrete.Left, positive), oneStep(concrete.Right, positive))
|
||||
case ltl.OrFormula:
|
||||
return ltl.Or(oneStep(concrete.Left, positive), oneStep(concrete.Right, positive))
|
||||
case ltl.ImpliesFormula:
|
||||
return ltl.Implies(
|
||||
oneStep(concrete.Antecedent, !positive),
|
||||
oneStep(concrete.Consequent, positive),
|
||||
)
|
||||
default:
|
||||
return formula
|
||||
}
|
||||
}
|
||||
|
||||
// propertyClass splits safety from liveness by the property's top-level form:
|
||||
// a reachability goal is liveness, everything else is a safety obligation
|
||||
// re-asserted at each observation.
|
||||
func propertyClass(formula ltl.Formula) string {
|
||||
if _, ok := formula.(ltl.EventuallyFormula); ok {
|
||||
return "liveness"
|
||||
}
|
||||
return "safety"
|
||||
}
|
||||
|
||||
func topLevelForm(formula ltl.Formula) string {
|
||||
switch concrete := formula.(type) {
|
||||
case ltl.AlwaysFormula:
|
||||
return "always" + boundSuffix(concrete.HasStepBound, concrete.StepBound, concrete.Duration)
|
||||
case ltl.EventuallyFormula:
|
||||
return "eventually" + boundSuffix(concrete.HasStepBound, concrete.StepBound, concrete.Duration)
|
||||
default:
|
||||
return "predicate"
|
||||
}
|
||||
}
|
||||
|
||||
func boundSuffix(
|
||||
hasStepBound bool,
|
||||
stepBound int,
|
||||
duration time.Duration,
|
||||
) string {
|
||||
switch {
|
||||
case hasStepBound:
|
||||
return fmt.Sprintf(" within %d steps", stepBound)
|
||||
case duration > 0:
|
||||
return fmt.Sprintf(" within %s", duration)
|
||||
default:
|
||||
return ""
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,222 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/ltl"
|
||||
)
|
||||
|
||||
// observe drives an evaluator over a fixed sequence of predicate readings and
|
||||
// returns the observation at which it latched, or 0 if it never did.
|
||||
func observe(formula ltl.Formula, steps int) int {
|
||||
evaluator := ltl.NewEvaluator(formula)
|
||||
base := time.Unix(0, 0)
|
||||
for step := 1; step <= steps; step++ {
|
||||
if evaluator.ObserveAtStep(
|
||||
base.Add(time.Duration(step)*time.Second),
|
||||
step,
|
||||
) == ltl.VerdictViolated {
|
||||
return step
|
||||
}
|
||||
}
|
||||
if evaluator.Finalize() == ltl.VerdictViolated {
|
||||
return steps + 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func constant(value bool) ltl.Formula {
|
||||
return ltl.ThunkNamed("p", func() (bool, error) { return value, nil })
|
||||
}
|
||||
|
||||
func TestSingleStateHoldsWhereRefutationNeedsTheNextObservation(t *testing.T) {
|
||||
property := ltl.Always(
|
||||
ltl.Implies(ltl.Now(constant(true)), ltl.Next(constant(false))),
|
||||
)
|
||||
|
||||
if engine := observe(property, 4); engine == 0 {
|
||||
t.Fatal(
|
||||
"the engine was expected to refute a next obligation that never held",
|
||||
)
|
||||
}
|
||||
if reduced := observe(singleStateFormula(property), 4); reduced != 0 {
|
||||
t.Errorf(
|
||||
"single-state refuted at observation %d, and one observation cannot see a next",
|
||||
reduced,
|
||||
)
|
||||
}
|
||||
if reduced := observe(singleStepFormula(property), 4); reduced == 0 {
|
||||
t.Error("single-step was expected to refute a one-step obligation")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStateRefutesAPredicateReadAtOneObservation(t *testing.T) {
|
||||
property := ltl.Always(constant(false))
|
||||
|
||||
if reduced := observe(singleStateFormula(property), 3); reduced != 1 {
|
||||
t.Errorf(
|
||||
"single-state latched at %d, want the first observation",
|
||||
reduced,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStateKeepsANegatedTemporalHarmless(t *testing.T) {
|
||||
property := ltl.Always(ltl.Not(ltl.Eventually(constant(false))))
|
||||
|
||||
if reduced := observe(singleStateFormula(property), 3); reduced != 0 {
|
||||
t.Errorf(
|
||||
"single-state refuted at observation %d; erasing a negated eventually must not manufacture a violation",
|
||||
reduced,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStepShrinksAReachabilityGoalToTwoObservations(t *testing.T) {
|
||||
property := ltl.EventuallyWithinSteps(constant(false), 500)
|
||||
|
||||
if engine := observe(property, 4); engine != 5 {
|
||||
t.Fatalf(
|
||||
"a 500-step window closes only at run end, latched at %d",
|
||||
engine,
|
||||
)
|
||||
}
|
||||
if reduced := observe(singleStepFormula(property), 4); reduced != 2 {
|
||||
t.Errorf(
|
||||
"single-step latched at %d, want the second observation",
|
||||
reduced,
|
||||
)
|
||||
}
|
||||
if !truncatesWindow(property) {
|
||||
t.Error(
|
||||
"a 500-step window shortened to two observations must be reported as truncated",
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
// sequence returns a predicate that reads the given values, one per call, so
|
||||
// each evaluator gets its own reading counter.
|
||||
func sequence(readings []bool) ltl.Formula {
|
||||
position := 0
|
||||
return ltl.ThunkNamed("p", func() (bool, error) {
|
||||
value := readings[position]
|
||||
position++
|
||||
return value, nil
|
||||
})
|
||||
}
|
||||
|
||||
func TestSingleStepConvictsAGoalTheEngineSeesReached(t *testing.T) {
|
||||
readings := []bool{false, false, true, true}
|
||||
|
||||
if engine := observe(ltl.EventuallyWithinSteps(sequence(readings), 4), 4); engine != 0 {
|
||||
t.Fatalf(
|
||||
"the engine latched at %d, and the goal is reached inside its window",
|
||||
engine,
|
||||
)
|
||||
}
|
||||
if reduced := observe(singleStepFormula(ltl.EventuallyWithinSteps(sequence(readings), 4)), 4); reduced != 2 {
|
||||
t.Errorf(
|
||||
"single-step latched at %d; an obligation armed at the first observation must not reach the third",
|
||||
reduced,
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTruncatesWindowIgnoresAnObligationThatAlreadyFits(t *testing.T) {
|
||||
property := ltl.Always(
|
||||
ltl.Implies(ltl.Now(constant(true)), ltl.Next(constant(true))),
|
||||
)
|
||||
|
||||
if truncatesWindow(property) {
|
||||
t.Error("a next spans two observations already and is not truncated")
|
||||
}
|
||||
}
|
||||
|
||||
func TestPropertyClassAndFormComeFromTheTopLevel(t *testing.T) {
|
||||
safety := ltl.Always(constant(true))
|
||||
liveness := ltl.EventuallyWithinSteps(constant(true), 575)
|
||||
|
||||
if got := propertyClass(safety); got != "safety" {
|
||||
t.Errorf("class of an always: got %q", got)
|
||||
}
|
||||
if got := propertyClass(liveness); got != "liveness" {
|
||||
t.Errorf("class of an eventually: got %q", got)
|
||||
}
|
||||
if got := topLevelForm(liveness); got != "eventually within 575 steps" {
|
||||
t.Errorf("form: got %q", got)
|
||||
}
|
||||
if got := topLevelForm(ltl.EventuallyWithin(constant(true), 3*time.Second)); got != "eventually within 3s" {
|
||||
t.Errorf("form: got %q", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStepCannotExpressAWindowLongerThanATriple(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
name string
|
||||
property ltl.Formula
|
||||
expresses bool
|
||||
}{
|
||||
{"reachability goal in steps", ltl.EventuallyWithinSteps(constant(true), 575), false},
|
||||
{"reachability goal in time", ltl.EventuallyWithin(constant(true), 3*time.Second), false},
|
||||
{"unbounded reachability goal", ltl.Eventually(constant(true)), false},
|
||||
{"window a triple spans exactly", ltl.EventuallyWithinSteps(constant(true), tripleWindow), true},
|
||||
{"deadline nested under an always", ltl.Always(ltl.Implies(
|
||||
ltl.Now(constant(true)), ltl.EventuallyWithin(constant(true), 3*time.Second))), false},
|
||||
{"one-step obligation under an always", ltl.Always(ltl.Implies(
|
||||
ltl.Now(constant(true)), ltl.Next(constant(true)))), true},
|
||||
{"predicate under an always", ltl.Always(constant(true)), true},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
if singleStepExpresses(testCase.property) != testCase.expresses {
|
||||
t.Errorf("single-step expresses %s: got %t, want %t",
|
||||
testCase.name, !testCase.expresses, testCase.expresses)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A shorter window under a negation asks for less rather than more, so the
|
||||
// triple's refutations of it stay sound and the property stays expressible.
|
||||
// Reporting it as inexpressible would push a defect the single-step oracle can
|
||||
// genuinely catch into the temporal-only column.
|
||||
func TestSingleStepStillExpressesANegatedWindow(t *testing.T) {
|
||||
property := ltl.Always(
|
||||
ltl.Not(ltl.EventuallyWithinSteps(constant(false), 575)),
|
||||
)
|
||||
|
||||
if !singleStepExpresses(property) {
|
||||
t.Error(
|
||||
"a window shortened under a negation is weakened, not strengthened",
|
||||
)
|
||||
}
|
||||
if truncatesWindow(property) {
|
||||
t.Error(
|
||||
"the truncation marker must not fire where shortening only weakens",
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSingleStateCannotExpressWhatItsRewriteEmpties(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
name string
|
||||
property ltl.Formula
|
||||
expresses bool
|
||||
}{
|
||||
{"predicate at one observation", ltl.Always(constant(true)), true},
|
||||
{"reachability goal", ltl.EventuallyWithinSteps(constant(true), 575), false},
|
||||
{"next obligation under an always", ltl.Always(ltl.Implies(
|
||||
ltl.Now(constant(true)), ltl.Next(constant(true)))), false},
|
||||
{"deadline under an always", ltl.Always(ltl.Implies(
|
||||
ltl.Now(constant(true)), ltl.EventuallyWithin(constant(true), 3*time.Second))), false},
|
||||
{"negated eventually under an always", ltl.Always(ltl.Not(ltl.Eventually(constant(false)))), false},
|
||||
{"predicate conjoined with a next", ltl.Always(ltl.And(constant(true), ltl.Next(constant(true)))), true},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
if singleStateExpresses(testCase.property) != testCase.expresses {
|
||||
t.Errorf("single-state expresses %s: got %t, want %t",
|
||||
testCase.name, !testCase.expresses, testCase.expresses)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -35,7 +35,10 @@ type testOptions struct {
|
||||
output string
|
||||
clearData bool
|
||||
generator string
|
||||
labelSource string
|
||||
exitOnViolation bool
|
||||
allowNoProperties bool
|
||||
allowNoGeneratorActions bool
|
||||
}
|
||||
|
||||
const topUsage = `sanderling is a property-based UI fuzzer for mobile apps.
|
||||
@@ -71,7 +74,10 @@ func parseTestArgs(args []string, stderr io.Writer) (testOptions, error) {
|
||||
flagSet.BoolVar(&options.clearData, "clear-data", true, "clear app data before launching so each run starts from a fresh install; pass --clear-data=false to resume prior state")
|
||||
flagSet.StringVar(&options.arm, "arm", "", "experiment cell label, recorded in meta.json so a directory of runs can be attributed to a cell")
|
||||
flagSet.StringVar(&options.generator, "generator", "seeded", "action generator: seeded (weighted random) or llm (model picks from the same candidate set; requires generator = llm() in the spec)")
|
||||
flagSet.StringVar(&options.labelSource, "label-source", "visible-text", "how candidates are named to the llm generator: visible-text (what a user reads) or resource-id (the identifier the app assigned). The seeded generator picks by index and ignores this")
|
||||
flagSet.BoolVar(&options.exitOnViolation, "exit-on-violation", false, "stop the run at the first property violation and exit 2, so CI can tell a found bug (2) from a broken harness (1)")
|
||||
flagSet.BoolVar(&options.allowNoProperties, "allow-no-properties", false, "run a spec that registers no properties. Such a run judges nothing and can only report no violations, so it is refused by default; pass this when the run measures what the spec extracts")
|
||||
flagSet.BoolVar(&options.allowNoGeneratorActions, "allow-no-generator-actions", false, "finish a run the action generator never drove. Such a run judged whatever screen the spec's setup left it on and explored nothing, so it is refused by default; pass this when the run measures where the generator reaches and reaching nothing is the measurement")
|
||||
if err := flagSet.Parse(args); err != nil {
|
||||
return testOptions{}, err
|
||||
}
|
||||
@@ -94,6 +100,14 @@ func parseTestArgs(args []string, stderr io.Writer) (testOptions, error) {
|
||||
default:
|
||||
return testOptions{}, fmt.Errorf("unsupported generator: %q (seeded, llm)", options.generator)
|
||||
}
|
||||
// Rejected here rather than defaulted, because a campaign that finishes with
|
||||
// the wrong labelling and a plausible output directory is worse than one
|
||||
// that never starts.
|
||||
switch options.labelSource {
|
||||
case "visible-text", "resource-id":
|
||||
default:
|
||||
return testOptions{}, fmt.Errorf("unsupported label source: %q (visible-text, resource-id)", options.labelSource)
|
||||
}
|
||||
return options, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -10,6 +10,7 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/priyanshujain/sanderling/internal/testrun"
|
||||
"github.com/priyanshujain/sanderling/internal/verifier"
|
||||
)
|
||||
|
||||
func TestParseTestArgs_Defaults(t *testing.T) {
|
||||
@@ -132,6 +133,63 @@ func TestParseTestArgs_RejectsUnknownGenerator(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The label-source cases compare against the verifier's own constants, because
|
||||
// the flag and the code that reads it are the two halves of one contract: a
|
||||
// rename on either side would otherwise leave every run silently labelled by
|
||||
// the default channel while meta.json claimed the other one.
|
||||
func TestParseTestArgs_LabelSourceDefaultsToVisibleText(t *testing.T) {
|
||||
options, err := parseTestArgs([]string{"--spec", "s.ts", "--bundle-id", "com.example"}, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
if options.labelSource != verifier.LabelSourceVisibleText {
|
||||
t.Fatalf("label source default: got %q, want %q", options.labelSource, verifier.LabelSourceVisibleText)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTestArgs_AcceptsResourceIDLabelSource(t *testing.T) {
|
||||
options, err := parseTestArgs([]string{
|
||||
"--spec", "s.ts",
|
||||
"--bundle-id", "com.example",
|
||||
"--label-source", verifier.LabelSourceResourceID,
|
||||
}, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
if options.labelSource != verifier.LabelSourceResourceID {
|
||||
t.Fatalf("label source: got %q, want %q", options.labelSource, verifier.LabelSourceResourceID)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTestArgs_RejectsUnknownLabelSource(t *testing.T) {
|
||||
_, err := parseTestArgs([]string{
|
||||
"--spec", "s.ts",
|
||||
"--bundle-id", "com.example",
|
||||
"--label-source", "resource_id",
|
||||
}, io.Discard)
|
||||
if err == nil || !strings.Contains(err.Error(), "unsupported label source") {
|
||||
t.Fatalf("expected unsupported-label-source error, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPipelineOptionsCarriesTheExperimentCell(t *testing.T) {
|
||||
options, err := parseTestArgs([]string{
|
||||
"--spec", "s.ts",
|
||||
"--bundle-id", "com.example",
|
||||
"--generator", "llm",
|
||||
"--label-source", verifier.LabelSourceResourceID,
|
||||
"--arm", "llm-resource-id",
|
||||
}, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
pipeline := pipelineOptions(options)
|
||||
if pipeline.Generator != "llm" || pipeline.LabelSource != verifier.LabelSourceResourceID || pipeline.Arm != "llm-resource-id" {
|
||||
t.Errorf("cell lost between the flags and the pipeline: generator=%q labelSource=%q arm=%q",
|
||||
pipeline.Generator, pipeline.LabelSource, pipeline.Arm)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseTestArgs_RejectsUnknownPlatform(t *testing.T) {
|
||||
_, err := parseTestArgs([]string{
|
||||
"--spec", "s.ts",
|
||||
@@ -343,6 +401,62 @@ func TestParseTestArgs_ExitOnViolation(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseTestArgs_AllowNoPropertiesReachesThePipeline pins the opt-out the
|
||||
// extraction and portability sweeps pass. Dropped here, the guard is either
|
||||
// unreachable or permanent: a sweep that deliberately judges nothing cannot ask
|
||||
// for it, and every other run keeps the false green the guard exists to stop.
|
||||
func TestParseTestArgs_AllowNoPropertiesReachesThePipeline(t *testing.T) {
|
||||
base := []string{"--spec", "s.ts", "--bundle-id", "com.example"}
|
||||
|
||||
options, err := parseTestArgs(base, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if pipelineOptions(options).AllowNoProperties {
|
||||
t.Error("allowNoProperties default: got true, want false")
|
||||
}
|
||||
|
||||
options, err = parseTestArgs(append(base, "--allow-no-properties"), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !pipelineOptions(options).AllowNoProperties {
|
||||
t.Error("--allow-no-properties never reached the pipeline options")
|
||||
}
|
||||
}
|
||||
|
||||
// The exploration sweeps ask for a run the generator never drove, and they ask
|
||||
// for that alone. Riding on --allow-no-properties made one flag name two
|
||||
// unrelated waivers, so a sweep that wanted the property-free one silently lost
|
||||
// the dead-run detector as well.
|
||||
func TestParseTestArgs_AllowNoGeneratorActionsIsItsOwnFlag(t *testing.T) {
|
||||
base := []string{"--spec", "s.ts", "--bundle-id", "com.example"}
|
||||
|
||||
options, err := parseTestArgs(base, io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if pipelineOptions(options).AllowNoGeneratorActions {
|
||||
t.Error("allowNoGeneratorActions default: got true, want false")
|
||||
}
|
||||
|
||||
options, err = parseTestArgs(append(base, "--allow-no-generator-actions"), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !pipelineOptions(options).AllowNoGeneratorActions {
|
||||
t.Error("--allow-no-generator-actions never reached the pipeline options")
|
||||
}
|
||||
|
||||
options, err = parseTestArgs(append(base, "--allow-no-properties"), io.Discard)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if pipelineOptions(options).AllowNoGeneratorActions {
|
||||
t.Error("--allow-no-properties still waives the dead-run refusal it does not name")
|
||||
}
|
||||
}
|
||||
|
||||
// TestExitCode_SeparatesFoundBugsFromBrokenHarnesses pins the three statuses CI
|
||||
// reads: 0 clean, 2 the run found violations, 1 everything else. A workflow
|
||||
// that asserts "the known bug is still found" is only meaningful while 2 and 1
|
||||
@@ -358,6 +472,7 @@ func TestExitCode_SeparatesFoundBugsFromBrokenHarnesses(t *testing.T) {
|
||||
{"help", flag.ErrHelp, 0, ""},
|
||||
{"violations found", testrun.ViolationsError{Count: 2}, 2, "violations: 2"},
|
||||
{"broken harness", errors.New("launch app: no device"), 1, "error: launch app"},
|
||||
{"spec judges nothing", testrun.NoPropertiesError{Spec: "s.ts"}, 1, "registers no properties"},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
var stderr bytes.Buffer
|
||||
|
||||
@@ -8,8 +8,17 @@ import (
|
||||
)
|
||||
|
||||
func runTestPipeline(ctx context.Context, options testOptions, stdout io.Writer) error {
|
||||
return testrun.Execute(ctx, testrun.Options{
|
||||
return testrun.Execute(ctx, pipelineOptions(options), stdout)
|
||||
}
|
||||
|
||||
// pipelineOptions maps the parsed flags onto the pipeline's options. A field
|
||||
// dropped on the way through here is a run that executes one experiment cell
|
||||
// and records another, which is worth being able to test on its own.
|
||||
func pipelineOptions(options testOptions) testrun.Options {
|
||||
return testrun.Options{
|
||||
Spec: options.spec,
|
||||
AllowNoProperties: options.allowNoProperties,
|
||||
AllowNoGeneratorActions: options.allowNoGeneratorActions,
|
||||
BundleID: options.bundleID,
|
||||
Platform: options.platform,
|
||||
AVD: options.avd,
|
||||
@@ -23,7 +32,8 @@ func runTestPipeline(ctx context.Context, options testOptions, stdout io.Writer)
|
||||
Output: options.output,
|
||||
ClearData: options.clearData,
|
||||
Generator: options.generator,
|
||||
LabelSource: options.labelSource,
|
||||
Arm: options.arm,
|
||||
ExitOnViolation: options.exitOnViolation,
|
||||
}, stdout)
|
||||
}
|
||||
}
|
||||
@@ -33,27 +33,28 @@ enum Snapshot {
|
||||
let application = XCUIApplication(bundleIdentifier: bundleIdentifier)
|
||||
let root = try application.snapshot()
|
||||
var result: [[String: Any]] = []
|
||||
walk(root, into: &result)
|
||||
walk(root, depth: 0, into: &result)
|
||||
return result
|
||||
}
|
||||
|
||||
private static func walk(_ node: XCUIElementSnapshot, into result: inout [[String: Any]]) {
|
||||
private static func walk(_ node: XCUIElementSnapshot, depth: Int, into result: inout [[String: Any]]) {
|
||||
// The keyboard subtree is pruned: the legacy accessibility bridge
|
||||
// never exposed it, its key frames are unreliable as tap targets, and
|
||||
// its shift-state churn destabilizes settle hashing.
|
||||
if node.elementType == .keyboard {
|
||||
return
|
||||
}
|
||||
result.append(serialize(node))
|
||||
result.append(serialize(node, depth: depth))
|
||||
for child in node.children {
|
||||
walk(child, into: &result)
|
||||
walk(child, depth: depth + 1, into: &result)
|
||||
}
|
||||
}
|
||||
|
||||
private static func serialize(_ node: XCUIElementSnapshot) -> [String: Any] {
|
||||
private static func serialize(_ node: XCUIElementSnapshot, depth: Int) -> [String: Any] {
|
||||
let frame = node.frame
|
||||
return [
|
||||
"type": elementTypeName(node.elementType),
|
||||
"depth": depth,
|
||||
"frame": [
|
||||
"x": Double(frame.origin.x),
|
||||
"y": Double(frame.origin.y),
|
||||
|
||||
@@ -46,12 +46,14 @@ enum TextInput {
|
||||
try typeOnFocus(deletes, bundleIdentifier: bundleIdentifier)
|
||||
}
|
||||
|
||||
// pressKey types a single logical key into the focused field. Only return is
|
||||
// supported, matching the simulator companion's key surface.
|
||||
// pressKey types a single logical key into the focused field, matching the
|
||||
// simulator companion's key surface.
|
||||
static func pressKey(key: String, bundleIdentifier: String) throws {
|
||||
switch key {
|
||||
case "return", "enter", "Return", "Enter":
|
||||
try typeOnFocus(XCUIKeyboardKey.return.rawValue, bundleIdentifier: bundleIdentifier)
|
||||
case "escape", "Escape":
|
||||
try typeOnFocus(XCUIKeyboardKey.escape.rawValue, bundleIdentifier: bundleIdentifier)
|
||||
default:
|
||||
throw TextInputError.typingFailed("unsupported key \(key)")
|
||||
}
|
||||
|
||||
+50
-13
@@ -154,21 +154,42 @@ gate_clear_state() {
|
||||
}
|
||||
|
||||
# G4: no doubled text after InputText. For each InputText action targeting a
|
||||
# field, the field's value in the NEXT hierarchy must not contain the input
|
||||
# concatenated with itself (catches append-vs-replace and double-paste bugs).
|
||||
# The action chosen at step N is applied before step N+1 is observed, so the
|
||||
# effect lands in the following snapshot.
|
||||
# field, the field's value in the NEXT hierarchy must not read as the input
|
||||
# applied twice (catches append-vs-replace and double-paste bugs). The action
|
||||
# chosen at step N is applied before step N+1 is observed, so the effect lands
|
||||
# in the following snapshot.
|
||||
#
|
||||
# Two signals, because the typed value is not always in the trace: a target the
|
||||
# platform reports no secure fact for has its typed value written as a fixed
|
||||
# placeholder (internal/verifier/redaction.go), and android reports that fact
|
||||
# for nothing, so on that backend every InputText records the placeholder. The
|
||||
# OBSERVED value is never redacted, and a field holding one string twice over is
|
||||
# the doubling itself, whether that is the whole value or what the value grew by
|
||||
# over the snapshot the action was chosen against. A value that is a single
|
||||
# character repeated is exempt: the input corpus types "a" 4096 times and a pair
|
||||
# of spaces, and neither can be told apart from its own doubling.
|
||||
gate_no_doubled_text() {
|
||||
local run_directory="$1"
|
||||
local trace_file="${run_directory}/trace.jsonl"
|
||||
[[ -f "$trace_file" ]] || return 1
|
||||
local verdict
|
||||
verdict="$(jq -s -r '
|
||||
verdict="$(jq -s -r --arg redacted "[redacted]" '
|
||||
# Map a selector like "testTag:LoginScreen > testTag:LoginEmail" to its
|
||||
# target field id: the token after the final ":" of the last segment.
|
||||
# target field id: the token after the final ":" of the last segment. An
|
||||
# action typed at coordinates carries no selector, and jq splits an empty
|
||||
# string into no segments at all, so the fallbacks keep such a step from
|
||||
# aborting the whole analysis.
|
||||
def target_field(selector):
|
||||
(selector | split(">") | last | gsub("^\\s+|\\s+$";"")) as $last
|
||||
| ($last | split(":") | last);
|
||||
((selector | split(">") | last // "") | gsub("^\\s+|\\s+$";"")) as $last
|
||||
| ($last | split(":") | last // "");
|
||||
|
||||
def self_doubled(value):
|
||||
(value | length) as $n
|
||||
| ($n / 2 | floor) as $half
|
||||
| $n > 1
|
||||
and ($n % 2 == 0)
|
||||
and (value[0:$half] == value[$half:])
|
||||
and ((value | explode | unique | length) > 1);
|
||||
|
||||
[ .[] | select(.hierarchy != null) ] as $steps
|
||||
| reduce range(0; ($steps | length)) as $i ([];
|
||||
@@ -180,16 +201,24 @@ gate_no_doubled_text() {
|
||||
| ($current.next_action.text // "") as $typed
|
||||
| (($next.hierarchy.elements // [])
|
||||
| map(select(.resourceId == $field)) | first) as $element
|
||||
| if $element != null and $typed != ""
|
||||
| (($current.hierarchy.elements // [])
|
||||
| map(select(.resourceId == $field)) | first) as $before
|
||||
| if $element != null
|
||||
then
|
||||
(($element.attrs.text // $element.text // "")) as $value
|
||||
| if ($value | contains($typed + $typed))
|
||||
then . + [{field: $field, typed: $typed, value: $value}]
|
||||
| (($before.attrs.text // $before.text // "")) as $previous
|
||||
| (if $previous != "" and ($value | startswith($previous))
|
||||
then $value[($previous | length):] else "" end) as $appended
|
||||
| if self_doubled($value)
|
||||
or self_doubled($appended)
|
||||
or ($typed != "" and $typed != $redacted
|
||||
and ($value | contains($typed + $typed)))
|
||||
then . + [{field: $field, value: $value}]
|
||||
else . end
|
||||
else . end
|
||||
else . end)
|
||||
| if length == 0 then "pass"
|
||||
else "fail:" + (.[0].field) + ":" + (.[0].value) end
|
||||
else "fail:" + (.[0].field) + ":" + (.[0].value[0:60]) end
|
||||
' "$trace_file")"
|
||||
[[ "$verdict" == "pass" ]]
|
||||
}
|
||||
@@ -493,8 +522,16 @@ self_test() {
|
||||
|
||||
assert "G4 pass no doubling" PASS \
|
||||
"$(gate_no_doubled_text "${testdata}/pass" && echo PASS || echo FAIL)"
|
||||
assert "G4 doubled text caught" FAIL \
|
||||
assert "G4 doubled text caught with the typed value redacted" FAIL \
|
||||
"$(gate_no_doubled_text "${testdata}/g4-doubled-text" && echo PASS || echo FAIL)"
|
||||
assert "G4 doubling appended to existing content caught" FAIL \
|
||||
"$(gate_no_doubled_text "${testdata}/g4-appended-doubling" && echo PASS || echo FAIL)"
|
||||
assert "G4 repeated character is not doubling" PASS \
|
||||
"$(gate_no_doubled_text "${testdata}/g4-repeated-character" && echo PASS || echo FAIL)"
|
||||
assert "G4 doubled text caught from the recorded value" FAIL \
|
||||
"$(gate_no_doubled_text "${testdata}/g4-recorded-text" && echo PASS || echo FAIL)"
|
||||
assert "G4 input without a selector does not abort the gate" PASS \
|
||||
"$(gate_no_doubled_text "${testdata}/g4-selectorless-input" && echo PASS || echo FAIL)"
|
||||
|
||||
assert "seeds default to one distinct value per run" "101 202 303 404 505" \
|
||||
"$(select_seeds 5 "101 202 303 404 505" >/dev/null 2>&1 && echo "${seed_list[*]}")"
|
||||
|
||||
+2
-2
@@ -1,4 +1,4 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[email protected]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "ledger123", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 3, "timestamp": "2026-06-06T12:00:01.800000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
{"step": 4, "timestamp": "2026-06-06T12:00:02.400000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
+2
-2
@@ -1,4 +1,4 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[email protected]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "ledger123", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 3, "timestamp": "2026-06-06T12:00:01.800000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
{"step": 4, "timestamp": "2026-06-06T12:00:02.400000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
+2
-2
@@ -1,4 +1,4 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[email protected]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "ledger123", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginPassword"}}
|
||||
{"step": 3, "timestamp": "2026-06-06T12:00:01.800000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
{"step": 4, "timestamp": "2026-06-06T12:00:02.400000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "ledger123"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
+1
-1
@@ -1,2 +1,2 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "leftover"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[email protected]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "leftover"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": "leftover"}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
@@ -0,0 +1 @@
|
||||
0
|
||||
@@ -0,0 +1 @@
|
||||
run complete
|
||||
@@ -0,0 +1,2 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "demo@"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
+1
-1
@@ -1,2 +1,2 @@
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[email protected]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 1, "timestamp": "2026-06-06T12:00:00.600000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}, "next_action": {"kind": "InputText", "text": "[redacted]", "selector": "testTag:LoginScreen > testTag:LoginEmail"}}
|
||||
{"step": 2, "timestamp": "2026-06-06T12:00:01.200000+05:30", "hierarchy": {"elements": [{"resourceId": "LoginScreen", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginScreen", "class": "android.view.View", "text": ""}}, {"resourceId": "LoginEmail", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginEmail", "class": "android.widget.EditText", "text": "[email protected]@folio.app"}}, {"resourceId": "LoginPassword", "class": "android.widget.EditText", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginPassword", "class": "android.widget.EditText", "text": ""}}, {"resourceId": "LoginSubmit", "class": "android.view.View", "editable": true, "bounds": {"left": 34, "top": 87, "right": 286, "bottom": 135}, "attrs": {"resource-id": "LoginSubmit", "class": "android.view.View", "text": ""}}]}}
|
||||
Loaded 100 of 243 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user