mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
9fb121e9d09c1245a4ff1f3d00ccb25846f8ca85
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
76dce1a75e |
experiment instrumentation: step budgets, arm labels, campaign runner (#72)
* feat(cli): add --max-steps for step-bounded runs runner.Options.MaxSteps already worked but was unreachable from the command line. A step budget is what makes two generators comparable: one making a model call per step and one drawing from a PRNG are not comparable per second. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(trace): record arm membership and host in meta.json meta.json recorded the seed but not which picker ran, how it was configured, what budget it was given, or which machine produced it. A directory of runs cannot be attributed to an experiment cell without those, which makes any factorial computed from such a directory unanalysable after the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(cli): add --arm and populate run meta from it Model and instructions are recorded only when the LLM picker is the one that will actually run, so a spec declaring generator = llm() that is run under the seeded picker does not label its trace with a model it never called. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): sweep seeds for one experiment cell campaign.json lists the seeds a sweep intended to run and is written before the first run, so a host that dropped runs shows up as missing seeds rather than as a smaller sample. Seed 0 is rejected: sanderling test reads it as "derive a seed from the clock", which is why conformance/gates.sh controls nothing today. Each run contributes one runs.jsonl line carrying steps to first violation by origin step, the step that armed the failed obligation, so the survival analysis never reopens a trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(runner): no silent generator fallback, and llm on web --generator llm against a spec declaring no generator = llm(...) logged a warning and ran the seeded picker. For a comparison campaign that is silent arm corruption: the run completes, the directory looks correct, and the wrong policy drove it. It is now fatal. pickSources also returned the V8 source for both action and extractor on web before it looked at the generator, so the llm policy was unreachable there. The two axes are now independent: the driver picks the extractor source, the flag picks the action source, and llmSource composes with either because the runner populates the candidate list and screenshot on every platform. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): make the hierarchy dump agree with the web runtime Three facts differed between the dump the goja host reads and the DOM the V8 host reads, so the two enumerated different candidates on one page. scrollable was never emitted, and worker.go reads exactly that attribute while targets.ts requires it for scrolls, so the goja host could not offer a single web scroll. clickable tested el.onclick, which React assigns to its root container for event delegation, making the whole viewport a tap target here and in no other enumeration. Both now resolve through the selector sets in pkg/spec/src/web-runtime.ts. The dump also rooted at body while collectTargets walks querySelectorAll("*"), so the goja host never saw html, where page-level scrolling lives. It now roots at documentElement and skips the head subtree, which is all zero-bounds and would otherwise carry script and title text into the trace. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(conformance): give the gate reproducible seeds SEED defaulted to 0 and sanderling test reads --seed 0 as "derive a seed from the clock", so the tunable controlled nothing and a gate failure could not be re-run. SEEDS now takes one explicit non-zero seed per run, recorded in the results table so a failing row names its stream. The five runs stay on five different streams: a gate that scored one path five times would catch less than one that scores five. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): emit editable as a plain boolean editable was emitted as `isEditable || null`, and an absent field sends internal/hierarchy into the native fallback, which reads any class name containing "EditText" as an Android text widget. On web that is just a CSS class, so a page styling a div with it was editable to the goja host and not to the web runtime, and the model policy could be offered typing into a div. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(spec): leave the head subtree out of the web target walk collectTargets walked querySelectorAll("*") while the hierarchy dump skips head, so the two hosts enumerated different element sets on every page with a <head>. No candidate changes: builtinCandidates pushes only for targets acceptsTarget admits, and head elements have no positive bounds, so the list the draw ranges over is untouched. What changes is that targetIndex now means the same thing on both hosts. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * test(chrome): compare the facts both hosts derive from one DOM The existing parity harness hand-authors the facts on both sides, so it proves that given identical facts both hosts select identical candidates, and says nothing about the two code paths that derive those facts from a real page. Four divergences lived in that blind spot and it passed throughout. This drives one real page and compares clickable, enabled, editable, scrollable and positiveBounds element by element, plus the element sets themselves, which is what catches a host that omits html or includes head. Reverting any of the four fixes makes it fail naming the element and the fact. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * chore(make): run the browser packages one at a time Both launch Chrome and launching two at once has failed with "Launch: context canceled". Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style: remove every em-dash and en-dash Eighteen occurrences across fourteen files. Each sentence was repunctuated to suit what the dash was doing rather than swapped for a hyphen, which produces comma splices. The minus sign in folio-web's ledger is a minus sign and stays. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(chrome): honor the caller context in Launch Launch and clearState ran against d.tabCtx, so a target that accepts the connection and never answers wedged the process past its own --duration and through SIGTERM, needing SIGKILL. Unattended that is a campaign worker lost for the rest of the sweep with no diagnostic. The browser is still allocated against d.tabCtx first, because chromedp starts Chrome under whichever context calls Run first and allocating under a caller deadline would kill the browser when Launch returns. Everything after allocation goes through runCtx. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * fix(sidecarassets): publish the extracted jar through a rename Extract wrote a 96 MB jar with a plain WriteFile into a temp path every sanderling process on the host shares. On a cold host several concurrent workers all miss the checksum and all write the same path, and O_TRUNC lets one spawn a JVM against another's half-written archive. A fresh experiment host is exactly a cold host. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * feat(campaign): kill a run that outlives --run-timeout A wedged run holds its worker for the rest of the sweep, and on an unattended host nothing else will send it a signal. Defaults to three times --duration and must exceed it. A killed run is recorded as timed_out rather than as a generic failure, so the analysis can tell a lost cell from a real crash. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX * style(test): gofmt browser_test.go Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX |
||
|
|
7343085614 |
llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker * feat(spec): make llm marker inert on the JS picker * feat(spec): expose __sanderlingSampleInput__ corpus draw * feat(openrouter): minimal chat-completions client * test(openrouter): cover request shape, parse, and errors * feat(verifier): thread screenshot + capture corpus sampler * feat(verifier): LLM accessors — candidates, config, sampler * test(verifier): cover AllCandidates, LLMConfig, SampleInput * feat(trace): record action Source and LLMReasoning * feat(runner): thread step screenshot into PushSnapshot * feat(runner): llmSource selects actions via OpenRouter * feat(runner): wire llmSource selection and trace stamping * test(runner): cover llmSource selection, mapping, downscale * docs(folio): add llm action-backend example spec * docs(folio): document the LLM action backend run * feat(llmclient): support OPENAI_API_KEY, openrouter wins * refactor(runner): rename openrouter package to llmclient * docs: both api keys, example model gpt-5.4-nano * docs: add pr style rules to claude.md * fix(runner): explain action kinds in llm prompt to stop swipe loops * feat(trace): record llm ranked list and chosen rank * feat(runner): stamp llm ranked list and chosen rank on trace * fix(runner): tap by selector to survive layout shift after observe * revert(runner): drop selector-first tap; broke path/testTag selectors * feat(spec): llm() accepts optional instructions * feat(verifier): read llm instructions off config * feat(runner): append spec instructions to llm system prompt * docs(folio): describe app in llm spec instructions * feat(bundler): map generator export to globalThis.generator * feat(verifier): read llm config off globalThis.generator * feat(runner): gate llm source on --generator flag * feat(cmd): add --generator llm|seeded flag * test: cover --generator flag parsing and pickSources gating * feat(verifier): enumerate llm candidates by walking actionsRoot collect-walk the weighted action tree: recurse weighted branches accumulating selection probability, call authored leaves once for concrete actions, enumerate builtins per element. label controls by visible text (borrowing descendant text), fold gestures into directional scrolls over scrollable containers, drop disabled, dedup descriptions. * test(verifier): cover candidate enumeration walk * feat(verifier): add SetupAction to walk setup without the seeded root * test(verifier): cover SetupAction setup-only precedence * refactor(llmclient): make JSONSchema.Schema raw json for pinned field order * feat(trace): record llm choice number and chosen_action echo * feat(runner): llm picks one number from weighted candidates drop the seeded-root call for a setup-only precedence path, render a numbered weighted candidate list, pin a reasoning-first choice schema, strict-skip when chosen_action does not echo the numbered entry, and let the model supply typed values (corpus fallback when empty). * test(runner): cover choice schema, strict-skip, and setup precedence * refactor(verifier): drop the superseded AllCandidates enumeration * feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts * fix(verifier): label editable fields by hint, not the typed value an editable field's own text is its transient content; prefer the hint so the field is named by purpose and the label stays stable. * test(runner): cover weight-suffixed echo and stripWeightSuffix * fix(runner): accept chosen_action echo that carries the weight suffix real runs showed the model copies the whole numbered line including the trailing (w34) weight annotation, so strict-skip rejected ~91% of picks and the llm was paralyzed. strip the weight suffix before comparing. also nudge the prompt to stress-test repeated submissions (idempotency). * fix(verifier): skip llm enumeration on cross-fade frames a navhost mid-transition carries >1 route *Screen in a collapsed coordinate space; acting on it taps garbage (soft keyboard). real runs showed the llm acting on 44% of steps being such frames. skip them so the llm re-observes a settled frame next step. * feat(folio): show current balance on the add-transaction screen renders the account's balance (testTag TxnCurrentBalance) below the account name, above the credit/debit toggle, so before/after screenshots carry comparison data. * fix(replay): derive device space from screen extent, not first node the first positive-bounds element is often a short status-bar node (320x24 on android); using it gave a 320/24 aspect ratio that squashed the screenshot overlay into a grey horizontal band. use the max extent across elements (like the runner's screenBounds) instead. * fix(folio): show balance as a compact one-line label per review: one line, account-name-sized, e.g. "Balance: $0.00" instead of a large balance card. * fix(folio): move balance into the header, one compact line under the account name * fix(replay): attribute deferred violations to the causing step, not detection * fix(replay): show a step's own violations in both panels, no next-step bleed * refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check * chore: ignore .playwright-mcp scratch output * docs: document the llm generator and --generator flag * docs(spec): correct the llm() comment; config reads off globalThis.generator * docs: add pr description rules |
||
|
|
94d9511312 |
test: full test-suite refactor sweep (#61)
* chore(test): start test-suite refactor sweep * test(ltl): pin exact multi-obligation residual AST * test(ltl): table-test finalize Kleene connective combinations * test(ltl): pin reduce over pending inner for bound, Or, Not * test(ltl): marshal bounded Always steps/duration/deadline * test(verifier): cover LTL combinator verdict transitions and within unit panic * test(verifier): table-test DecodeAction kinds and lastAction field exposure * test(verifier): assert WithPlatform(ios) reaches the picker host and key pool * test(verifier): widen weighted-selection assertion to a 5x skew margin * test(verifier): un-skip ax-find round trip with a committed tree fixture * test(runner): pin isWDADrop to sidecar reconnect-failed message origin * test(runner): assert PressKey/Wait trace encoding records kind-specific fields * test(runner): cover RenderSummary unsupported-verbs surfacing branch * test(trace): set Hierarchy in round-trip and lock lossy Tree contract Also add a -race concurrent WriteStep test that asserts N well-formed JSONL lines, catching torn lines if the writer mutex is dropped. * test(trace): round-trip witnesses/changes/metrics/exceptions, pin step-0 witness * test(trace): document ViolationsAreGreppable grep contract and lock-free WriteScreenshot * test(hierarchy): cover invalid-JSON and malformed-bounds parser paths * test(trace): guard writer mutex via WriteStep/Close race on w.file * test(replay): drop unfailable assets and devproxy assertions * test(replay): cache reuses on equal mtime, reparses after append * test(replay): violation marker falls back to detection step when attributed missing * test(replay): corrupt meta/trace dirs return 500 with error body * test(replay): SSE client receives runs.changed after a broadcast * test(replay): Run coalesces creates, ignores write/chmod, closes subs on cancel * fix(sidecar): synchronize health fixture writes and exercise healthError * test(sidecar): cover swipe/longpress/doubletap/erase/presskey/metrics/logs translations * test(sidecar): cover DoubleTapSelector composition and mid-gesture cancel * test(sidecar): assert gRPC error status surfaces from action RPC * fix(chrome): route action methods through runCtx so caller cancellation aborts CDP * fix(chrome): route hierarchy/screenshot/waitidle/metrics through runCtx * refactor(ios): extract pure simctl JSON parsers * refactor(ios): add command-runner seams for EnsureSimulator * test(ios): table-test simctl parsers and EnsureSimulator seams * test(sidecarassets): cover placeholder build path * test(sidecarassets): assert reuse via sentinel bytes not mtime * test(bundler): cover properties-only spec registration * refactor(testrun): extract prepareBundleInputs from Execute * test(testrun): cover prepareBundleInputs aliases and missing-runtime error * test(testrun): table-test resolveRuntimeSibling search edges * test(testrun): exact-output tests for progressHandler line format * fix(cmd): point bundle-check aliases at pkg/spec/src * test(cmd): smoke-test bundle-check resolves spec aliases * test(cmd): table-test hier-check parse and FindAll on fixture * test(cmd): unit-test buildBrowseURL deep-link vs root * test(cmd): drop flaky TestRun_Doctor that launched real Chromium * test(cmd): pin pipeline error to bundle resolution on web platform * test(replay-ui): add bun test script * ci(replay-ui): run bun test via make web-test target * ci(replay-ui): point bun cache key at replay-ui/bun.lock * test(replay-ui): exercise real URL encoding and non-ok throw in getJson * refactor(replay-ui): extract snapshot flatten/getAtPath into lib module * test(replay-ui): pin snapshot flatten/getAtPath path round-trip * refactor(replay-ui): extract action selector/format into lib module * test(replay-ui): pin action selector parse and row formatting * refactor(replay-ui): share one statusFor between panels * refactor(replay-ui): extract run-history derivation into lib module * test(replay-ui): pin shared statusFor precedence and ordering * test(replay-ui): pin run-history derivation alignment * refactor(replay-ui): export clampIndex for testing * refactor(replay-ui): extract keyboard-nav dispatch into pure module * refactor(replay-ui): extract metrics formatters into lib module * test(replay-ui): pin clampIndex step boundaries * test(replay-ui): pin keyboard-nav ownership and key routing * test(replay-ui): pin metrics formatters and path gap handling * refactor(sidecar): expose device-output parsers as internal for testing * test(sidecar): table-test device-output parsers against malformed input * test(sidecar): cover logcat parsing year inference and line skipping * test(sidecar): pin pressKey keycode mapping and unknown-key rejection * test(sidecar): metrics bundleId falls back to launched app and honors override * test(sidecar): loosen deadline upper bound to tolerate slow CI scheduling * test(web-runtime): export selector builders for unit tests * test(web-runtime): guard sanitize cycle, function, and depth limits * test(web-runtime): table-test selector builder quoting and escaping * test(sidecar): collapse scalar-forwarding RPC tests into a table * test(replay-ui): dedup step/summary fixtures into shared module * test(ios): collapse pickSimulator point-tests into a table |
||
|
|
410602d2e1 |
Fix action-system audit findings (#60)
* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
|
||
|
|
b44077afde |
replay ui fix (#56)
* refactor: rename inspect to replay across the codebase Renames inspect-ui/ to replay-ui/, internal/inspect/ to internal/replay/, the CLI subcommand from `sanderling inspect` to `sanderling replay`, and updates all references in docs, Makefile, README, and Go comments. * feat(replay-ui): show spec filename with full path on hover RunList and RunDetail now render the basename of spec_path (e.g. login.spec.ts) with the full path available as a title tooltip. |