mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-03 11:37:09 +00:00
* feat(spec): add llm() action-backend marker * feat(spec): make llm marker inert on the JS picker * feat(spec): expose __sanderlingSampleInput__ corpus draw * feat(openrouter): minimal chat-completions client * test(openrouter): cover request shape, parse, and errors * feat(verifier): thread screenshot + capture corpus sampler * feat(verifier): LLM accessors — candidates, config, sampler * test(verifier): cover AllCandidates, LLMConfig, SampleInput * feat(trace): record action Source and LLMReasoning * feat(runner): thread step screenshot into PushSnapshot * feat(runner): llmSource selects actions via OpenRouter * feat(runner): wire llmSource selection and trace stamping * test(runner): cover llmSource selection, mapping, downscale * docs(folio): add llm action-backend example spec * docs(folio): document the LLM action backend run * feat(llmclient): support OPENAI_API_KEY, openrouter wins * refactor(runner): rename openrouter package to llmclient * docs: both api keys, example model gpt-5.4-nano * docs: add pr style rules to claude.md * fix(runner): explain action kinds in llm prompt to stop swipe loops * feat(trace): record llm ranked list and chosen rank * feat(runner): stamp llm ranked list and chosen rank on trace * fix(runner): tap by selector to survive layout shift after observe * revert(runner): drop selector-first tap; broke path/testTag selectors * feat(spec): llm() accepts optional instructions * feat(verifier): read llm instructions off config * feat(runner): append spec instructions to llm system prompt * docs(folio): describe app in llm spec instructions * feat(bundler): map generator export to globalThis.generator * feat(verifier): read llm config off globalThis.generator * feat(runner): gate llm source on --generator flag * feat(cmd): add --generator llm|seeded flag * test: cover --generator flag parsing and pickSources gating * feat(verifier): enumerate llm candidates by walking actionsRoot collect-walk the weighted action tree: recurse weighted branches accumulating selection probability, call authored leaves once for concrete actions, enumerate builtins per element. label controls by visible text (borrowing descendant text), fold gestures into directional scrolls over scrollable containers, drop disabled, dedup descriptions. * test(verifier): cover candidate enumeration walk * feat(verifier): add SetupAction to walk setup without the seeded root * test(verifier): cover SetupAction setup-only precedence * refactor(llmclient): make JSONSchema.Schema raw json for pinned field order * feat(trace): record llm choice number and chosen_action echo * feat(runner): llm picks one number from weighted candidates drop the seeded-root call for a setup-only precedence path, render a numbered weighted candidate list, pin a reasoning-first choice schema, strict-skip when chosen_action does not echo the numbered entry, and let the model supply typed values (corpus fallback when empty). * test(runner): cover choice schema, strict-skip, and setup precedence * refactor(verifier): drop the superseded AllCandidates enumeration * feat(folio): drive spec.ts under --generator llm; drop spec-llm.ts * fix(verifier): label editable fields by hint, not the typed value an editable field's own text is its transient content; prefer the hint so the field is named by purpose and the label stays stable. * test(runner): cover weight-suffixed echo and stripWeightSuffix * fix(runner): accept chosen_action echo that carries the weight suffix real runs showed the model copies the whole numbered line including the trailing (w34) weight annotation, so strict-skip rejected ~91% of picks and the llm was paralyzed. strip the weight suffix before comparing. also nudge the prompt to stress-test repeated submissions (idempotency). * fix(verifier): skip llm enumeration on cross-fade frames a navhost mid-transition carries >1 route *Screen in a collapsed coordinate space; acting on it taps garbage (soft keyboard). real runs showed the llm acting on 44% of steps being such frames. skip them so the llm re-observes a settled frame next step. * feat(folio): show current balance on the add-transaction screen renders the account's balance (testTag TxnCurrentBalance) below the account name, above the credit/debit toggle, so before/after screenshots carry comparison data. * fix(replay): derive device space from screen extent, not first node the first positive-bounds element is often a short status-bar node (320x24 on android); using it gave a 320/24 aspect ratio that squashed the screenshot overlay into a grey horizontal band. use the max extent across elements (like the runner's screenBounds) instead. * fix(folio): show balance as a compact one-line label per review: one line, account-name-sized, e.g. "Balance: $0.00" instead of a large balance card. * fix(folio): move balance into the header, one compact line under the account name * fix(replay): attribute deferred violations to the causing step, not detection * fix(replay): show a step's own violations in both panels, no next-step bleed * refactor(hierarchy): one Tree.Transitional, drop the duplicated cross-fade check * chore: ignore .playwright-mcp scratch output * docs: document the llm generator and --generator flag * docs(spec): correct the llm() comment; config reads off globalThis.generator * docs: add pr description rules
203 lines
8.4 KiB
TypeScript
203 lines
8.4 KiB
TypeScript
import {
|
|
InputText,
|
|
Tap,
|
|
actions,
|
|
always,
|
|
extract,
|
|
from,
|
|
integers,
|
|
llm,
|
|
next,
|
|
weighted,
|
|
whenRoute,
|
|
} from "@sanderling/spec";
|
|
import { defaultActions, doubleTaps } from "@sanderling/spec/defaults";
|
|
import {
|
|
computeHomeTotalBalance,
|
|
parseTypedAmount,
|
|
submitChangesBalanceByTypedAmount,
|
|
} from "./predicates";
|
|
|
|
interface Account {
|
|
name: string;
|
|
balance: number;
|
|
}
|
|
|
|
// Parses formatCents output like "$5.00", "-$1,234.56", "+$0.50" back to integer cents.
|
|
function parseDollarCents(text: string | undefined): number {
|
|
if (!text) return 0;
|
|
const sign = text.startsWith("-") ? -1 : 1;
|
|
const digits = text.replace(/[^0-9]/g, "");
|
|
return digits ? sign * parseInt(digits, 10) : 0;
|
|
}
|
|
|
|
// Route detection via testTag (resource-id on Android, accessibilityIdentifier on iOS)
|
|
const loggedIn = extract("loggedIn", s => s.ax.find({ testTag: "LoginScreen" }) == null);
|
|
const route = extract<string | null>("route", s => {
|
|
if (s.ax.find({ testTag: "LoginScreen" })) return "login";
|
|
if (s.ax.find({ testTag: "AddAccountScreen" })) return "add-account";
|
|
if (s.ax.find({ testTag: "AddTransactionScreen" })) return "add-transaction";
|
|
if (s.ax.find({ testTag: "LedgerScreen" })) return "ledger";
|
|
if (s.ax.find({ testTag: "HomeScreen" })) return "home";
|
|
return null;
|
|
});
|
|
|
|
// Account cards on Home: identity is the AccountName text; balance comes from AccountBalance.
|
|
const accounts = extract<Account[]>("accounts", s =>
|
|
s.ax.findAll([{ testTag: "HomeScreen" }, { testTag: "AccountCard" }]).map(card => ({
|
|
name: card.find({ testTag: "AccountName" })?.text ?? "",
|
|
balance: parseDollarCents(card.find({ testTag: "AccountBalance" })?.text),
|
|
})));
|
|
|
|
// Total balance: sum of AccountCard balances visible on Home. The carrier
|
|
// deliberately tracks only the Home multi-account total. Ledger's
|
|
// LedgerBalance is a single-account number on a different scale and would
|
|
// corrupt cross-screen comparisons if mixed in. Off-Home steps carry forward
|
|
// the last-seen Home sum so `previous` and `current` stay on the same scale.
|
|
let lastHomeTotal = 0;
|
|
const totalBalance = extract("totalBalance", s => {
|
|
const cards = s.ax.findAll([{ testTag: "HomeScreen" }, { testTag: "AccountCard" }]);
|
|
const cardBalanceTexts = cards.map(c => c.find({ testTag: "AccountBalance" })?.text);
|
|
lastHomeTotal = computeHomeTotalBalance({ cardBalanceTexts, previousCarrier: lastHomeTotal });
|
|
return lastHomeTotal;
|
|
});
|
|
|
|
const lastAction = extract("lastAction", s => s.lastAction);
|
|
|
|
const loginEmailField = extract("loginEmailField", s =>
|
|
s.ax.find([{ testTag: "LoginScreen" }, { testTag: "LoginEmail" }]));
|
|
const loginPasswordField = extract("loginPasswordField", s =>
|
|
s.ax.find([{ testTag: "LoginScreen" }, { testTag: "LoginPassword" }]));
|
|
const loginSubmit = extract("loginSubmit", s =>
|
|
s.ax.find([{ testTag: "LoginScreen" }, { testTag: "LoginSubmit" }]));
|
|
const addAccountButton = extract("addAccountButton", s =>
|
|
s.ax.find([{ testTag: "HomeScreen" }, { testTag: "AddAccountButton" }]));
|
|
const accountNameField = extract("accountNameField", s =>
|
|
s.ax.find([{ testTag: "AddAccountScreen" }, { testTag: "AccountNameField" }]));
|
|
const addAccountSubmit = extract("addAccountSubmit", s =>
|
|
s.ax.find([{ testTag: "AddAccountScreen" }, { testTag: "AddAccountSubmit" }]));
|
|
const addTxnButton = extract("addTxnButton", s =>
|
|
s.ax.find([{ testTag: "LedgerScreen" }, { testTag: "AddTransactionButton" }]));
|
|
const txnAmountField = extract("txnAmountField", s =>
|
|
s.ax.find([{ testTag: "AddTransactionScreen" }, { testTag: "TxnAmountField" }]));
|
|
const txnSubmit = extract("txnSubmit", s =>
|
|
s.ax.find([{ testTag: "AddTransactionScreen" }, { testTag: "TxnSubmit" }]));
|
|
const accountCards = extract("accountCards", s =>
|
|
s.ax.findAll([{ testTag: "HomeScreen" }, { testTag: "AccountCard" }]));
|
|
|
|
// Property 1: every newly-appearing account starts with balance === 0.
|
|
// Identity is by visible name. Guard against navigation transitions where
|
|
// accounts vanish from the visible tree.
|
|
const newAccountBalanceIsZero = always(
|
|
next(() => {
|
|
const prev = accounts.previous ?? [];
|
|
const curr = accounts.current;
|
|
if (prev.length === 0 || curr.length === 0) return true;
|
|
const prevNames = new Set(prev.map(a => a.name));
|
|
return curr.filter(a => !prevNames.has(a.name)).every(a => a.balance === 0);
|
|
})
|
|
);
|
|
|
|
// Property 2: a tap on TxnSubmit must move the total balance by exactly the
|
|
// typed amount. A double-submit lands two transactions, so the balance shifts
|
|
// by twice the typed amount and the check fires. The route gate inside the
|
|
// predicate skips off-Home landings where totalBalance.current is the carrier.
|
|
const submitMovesBalanceByTypedAmount = always(
|
|
next(() =>
|
|
submitChangesBalanceByTypedAmount({
|
|
route: route.current,
|
|
lastAction: lastAction.current,
|
|
typedAmount: parseTypedAmount(txnAmountField.previous?.text),
|
|
prevTotalBalance: totalBalance.previous ?? 0,
|
|
currTotalBalance: totalBalance.current,
|
|
}),
|
|
),
|
|
);
|
|
|
|
const DEMO_EMAIL = "[email protected]";
|
|
const DEMO_PASSWORD = "ledger123";
|
|
|
|
// Login: drive the form by reading what's currently in each field, not by
|
|
// inferring intent from which one happens to be focused. A focus-driven
|
|
// approach loops forever if we re-enter the login screen with the password
|
|
// field already focused (e.g. after an exception bounces us back from
|
|
// another screen): it would type the password then tap submit with an
|
|
// empty email field, see no progress, and repeat indefinitely.
|
|
const login = actions(() => {
|
|
if (loggedIn.current) return [];
|
|
const email = loginEmailField.current;
|
|
const pwd = loginPasswordField.current;
|
|
if (email && !email.text) return [InputText({ into: email, text: DEMO_EMAIL })];
|
|
if (pwd && !pwd.text) return [InputText({ into: pwd, text: DEMO_PASSWORD })];
|
|
const submit = loginSubmit.current;
|
|
return submit ? [Tap({ on: submit })] : [];
|
|
});
|
|
|
|
const accountNames = from(["Checking", "Savings", "Travel", "Emergency Fund", "Investments"]);
|
|
|
|
const addAccount = whenRoute(route, ["home", "add-account"], () => {
|
|
if (route.current === "home") {
|
|
const btn = addAccountButton.current;
|
|
return btn ? [Tap({ on: btn })] : [];
|
|
}
|
|
const field = accountNameField.current;
|
|
const submit = addAccountSubmit.current;
|
|
const opts = [];
|
|
if (field) opts.push(InputText({ into: field, text: accountNames.generate() }));
|
|
if (submit) opts.push(Tap({ on: submit }));
|
|
return opts;
|
|
});
|
|
|
|
const amounts = integers().between(1, 500);
|
|
|
|
const addTxn = whenRoute(route, ["home", "ledger", "add-transaction"], () => {
|
|
if (route.current === "home") {
|
|
const cards = accountCards.current;
|
|
if (cards.length === 0) return [];
|
|
return [Tap({ on: from(cards).generate() })];
|
|
}
|
|
if (route.current === "ledger") {
|
|
const btn = addTxnButton.current;
|
|
return btn ? [Tap({ on: btn })] : [];
|
|
}
|
|
const field = txnAmountField.current;
|
|
const submit = txnSubmit.current;
|
|
const opts = [];
|
|
if (field) opts.push(InputText({ into: field, text: String(amounts.generate()) }));
|
|
if (submit) opts.push(Tap({ on: submit }));
|
|
return opts;
|
|
});
|
|
|
|
export const properties = {
|
|
newAccountBalanceIsZero,
|
|
submitMovesBalanceByTypedAmount,
|
|
};
|
|
|
|
export const setup = login;
|
|
|
|
// Weights declare testing intent. The transaction chain is the focus: it is
|
|
// the deepest flow and both balance properties observe it. Account creation
|
|
// stays in the mix because newAccountBalanceIsZero needs fresh accounts to
|
|
// fire. doubleTaps gets explicit weight on every screen because rapid
|
|
// double-submission is a failure mode these forms must be idempotent under.
|
|
// defaultActions adds breadth so the fuzzer wanders the whole app and types
|
|
// edge-case values into every field.
|
|
export const actionsRoot = weighted(
|
|
[25, addAccount],
|
|
[45, addTxn],
|
|
[5, doubleTaps],
|
|
[25, defaultActions],
|
|
);
|
|
|
|
// The LLM generator is orthogonal to actionsRoot: with `--generator llm` a model
|
|
// picks from the SAME weighted candidate set above, reading the screenshot and a
|
|
// numbered, weight-annotated list; the default `--generator seeded` ignores it.
|
|
// instructions describe only WHAT the app is, never HOW to test it — the model
|
|
// figures out how to surface bugs on its own. With a plain OpenAI key, drop the
|
|
// vendor prefix from the model id.
|
|
export const generator = llm({
|
|
model: "gpt-5.4-nano",
|
|
instructions:
|
|
"Folio is a personal-finance ledger app. After signing in, the home screen lists accounts, each with a balance. You can create accounts, open an account to see its ledger, and add transactions; each transaction has an amount and changes that account's balance and the overall total.",
|
|
});
|