mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-04 12:07:09 +00:00
26a1eda68d1dd1869317f1129ed4560110b7fc44
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7343085614 |
llm action-selection backend (#68)
* feat(spec): add llm() action-backend marker * feat(spec): make llm marker inert on the JS picker * feat(spec): expose __sanderlingSampleInput__ corpus draw * feat(openrouter): minimal chat-completions client * test(openrouter): cover request shape, parse, and errors * feat(verifier): thread screenshot + capture corpus sampler * feat(verifier): LLM accessors — candidates, config, sampler * test(verifier): cover AllCandidates, LLMConfig, SampleInput * feat(trace): record action Source and LLMReasoning * feat(runner): thread step screenshot into PushSnapshot * feat(runner): llmSource selects actions via OpenRouter * feat(runner): wire llmSource selection and trace stamping * test(runner): cover llmSource selection, mapping, downscale * docs(folio): add llm action-backend example spec * docs(folio): document the LLM action backend run * feat(llmclient): support OPENAI_API_KEY, openrouter wins * refactor(runner): rename openrouter package to llmclient * docs: both api keys, example model gpt-5.4-nano * docs: add pr style rules to claude.md * fix(runner): explain action kinds in llm prompt to stop swipe loops * feat(trace): record llm ranked list and chosen rank * feat(runner): stamp llm ranked list and chosen rank on trace * fix(runner): tap by selector to survive layout shift after observe * revert(runner): drop selector-first tap; broke path/testTag selectors * feat(spec): llm() accepts optional instructions * feat(verifier): read llm instructions off config * feat(runner): append spec instructions to llm system prompt * docs(folio): describe app in llm spec instructions * feat(bundler): map generator export to globalThis.generator * feat(verifier): read llm config off globalThis.generator * feat(runner): gate llm source on --generator flag * feat(cmd): add --generator llm|seeded flag * test: cover --generator flag parsing and pickSources gating * feat(verifier): enumerate llm candidates by walking actionsRoot collect-walk the weighted action tree: recurse weighted branches accumulating selection probability, call authored leaves once for concrete actions, enumerate builtins per element. label controls by visible text (borrowing descendant text), fold gestures into directional scrolls over scrollable containers, drop disabled, dedup descriptions. * test(verifier): cover candidate enumeration walk * feat(verifier): add SetupAction to walk setup without the seeded root * test(verifier): cover SetupAction setup-only precedence * refactor(llmclient): make JSONSchema.Schema raw json for pinned field order * feat(trace): record llm choice number and chosen_action echo * feat(runner): llm picks one number from weighted candidates drop the seeded-root call for a setup-only precedence path, render a numbered weighted candidate list, pin a reasoning-first choice schema, strict-skip when chosen_action does not echo the numbered entry, and let the model supply typed values (corpus fallback when empty). * test(runner): cover choice schema, strict-skip, and setup precedence * refactor(verifier): drop the superseded AllCandidates enumeration * feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts * fix(verifier): label editable fields by hint, not the typed value an editable field's own text is its transient content; prefer the hint so the field is named by purpose and the label stays stable. * test(runner): cover weight-suffixed echo and stripWeightSuffix * fix(runner): accept chosen_action echo that carries the weight suffix real runs showed the model copies the whole numbered line including the trailing (w34) weight annotation, so strict-skip rejected ~91% of picks and the llm was paralyzed. strip the weight suffix before comparing. also nudge the prompt to stress-test repeated submissions (idempotency). * fix(verifier): skip llm enumeration on cross-fade frames a navhost mid-transition carries >1 route *Screen in a collapsed coordinate space; acting on it taps garbage (soft keyboard). real runs showed the llm acting on 44% of steps being such frames. skip them so the llm re-observes a settled frame next step. * feat(folio): show current balance on the add-transaction screen renders the account's balance (testTag TxnCurrentBalance) below the account name, above the credit/debit toggle, so before/after screenshots carry comparison data. * fix(replay): derive device space from screen extent, not first node the first positive-bounds element is often a short status-bar node (320x24 on android); using it gave a 320/24 aspect ratio that squashed the screenshot overlay into a grey horizontal band. use the max extent across elements (like the runner's screenBounds) instead. * fix(folio): show balance as a compact one-line label per review: one line, account-name-sized, e.g. "Balance: $0.00" instead of a large balance card. * fix(folio): move balance into the header, one compact line under the account name * fix(replay): attribute deferred violations to the causing step, not detection * fix(replay): show a step's own violations in both panels, no next-step bleed * refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check * chore: ignore .playwright-mcp scratch output * docs: document the llm generator and --generator flag * docs(spec): correct the llm() comment; config reads off globalThis.generator * docs: add pr description rules |
||
|
|
410602d2e1 |
Fix action-system audit findings (#60)
* fix(replay-ui): size overlay viewBox from hierarchy root bounds
Tap points are recorded in the hierarchy's coordinate space (iOS points,
Android pixels, web CSS px) while screenshots are device pixels, so the
overlay rendered at 1/3 position on iOS 3x screens. Derive the viewBox
from the root element bounds; natural image size stays the fallback.
* fix(runner): derive trace tap point from resolveCoordinates
stampSelectorTarget preferred possibly-stale action X/Y while dispatch
preferred the fresh tree-resolved center, so the trace could record a
different point than the one tapped. Both now share resolveCoordinates.
* fix(runner): settle after InputText focus tap before key events
The focus tap raises the keyboard; with no settle the keyboard
animation races the erase/type key events on iOS, landing them in the
wrong field or dropping them. Wait for idle after a successful focus
tap, bounded by the run's idle timeout.
* fix(driver): skip pre-erase for replace-on-input drivers
The web driver's InputText already replaces content via select-all, so
the runner's unconditional EraseText was a redundant round-trip on
every InputText. A new optional TextReplacer capability lets a driver
assert replace semantics; the runner skips the erase when asserted.
* fix(hierarchy): rank spatial-fallback matches by specificity
The bounds-containment fallback returned the first pre-order match, so
a screen-sized container could win over the intended small element.
Matches are now ordered smallest-area first; equal-area matches keep
pre-order, preserving the iOS-flat equal-bounds sibling pattern.
* fix(runner): treat an unchanging transitional tree as settled
A UI persistently showing two route-level Screen ids (overlay, both
route ids alive at rest) burned the full retry budget every step and
skipped the verifier forever. A tree byte-identical to the previous
attempt now breaks the retry loop as settled; genuine cross-fades
differ between attempts and keep the retry/skip behavior.
* fix(replay-ui): skip synthetic zero-bounds root in deviceSpaceOf
The iOS hierarchy prepends a zero-bounds node before the real root
window, so elements[0] returned undefined and the overlay fell back to
the screenshot's pixel size. Take the first element with positive
extent instead; pre-order puts the root window before any content.
Verified against a real iOS trace in the replay UI.
* fix(sidecar): never replay non-idempotent actions after reconnect
A dropped connection mid-action (e.g. a read timeout while the device
is still typing) re-ran the whole block after reconnecting, typing the
text twice and double-firing taps. Non-idempotent actions now reconnect
for the next RPC's benefit but surface UNAVAILABLE, which the runner
already treats as transient; idempotent reads keep the replay.
* fix(sidecar): land the second double-tap sequentially on gesture collision
The overlapped second tap can hit the XCTest runner while the first
gesture is still executing ('only one gesture can be performed at a
time'), failing the step. The second tap now waits the first out and
retries once, keeping the tight gap on the happy path.
* fix(sidecar): map non-Exception throwables to INTERNAL status
The vendored iOS client throws failures that do not extend Exception;
runRpc missed them, killing the RPC as a channel-level Unknown the
runner cannot classify. Catch Throwable instead.
* feat(sidecar): close the driver and app under test on shutdown
* test(sidecar): cover service shutdown paths
* fix(testrun): stop the sidecar with SIGTERM before killing
* fix(sidecar): reap orphaned XCTest runner sessions at iOS init
* fix(sidecar): probe channel liveness before restarting the XCTest runner
* test(sidecar): cover WdaRecovery restart and retry policy
* fix(sidecar): absorb first-leg double-tap collision sequentially
* fix(runner): scope WDA-drop detection and cap consecutive transient failures
* chore(sidecar): silence vendored loggers on expected failure paths
* fix(runner): absorb one-off apply errors; only an unbroken streak aborts
* fix(folio): install the current build before the Android fuzz run
* chore(sidecar): silence absorbed view-hierarchy poll noise in Android runs
The driver logs an ERROR for every on-device view-hierarchy fetch that the
device-side server cancels or times out while the UI animates. The stability
poll fetches the hierarchy on a sub-second cadence and swallows those throws
to keep polling, so each line is advisory with no effect on the run. Real
failures still reach the runner as gRPC status errors, so nothing is lost.
|