From e9935a15b6a9b30df2dc0d41bd18793a88571741 Mon Sep 17 00:00:00 2001 From: PJ Date: Thu, 30 Jul 2026 22:34:50 +0530 Subject: [PATCH] docs: document the llm generator and --generator flag --- docs/manual/cli.md | 1 + docs/manual/spec-language.md | 22 ++++++++++++++++++++-- 2 files changed, 21 insertions(+), 2 deletions(-) diff --git a/docs/manual/cli.md b/docs/manual/cli.md index ce48e73..060675b 100644 --- a/docs/manual/cli.md +++ b/docs/manual/cli.md @@ -23,6 +23,7 @@ Run a spec against an app for a fixed duration. | `--ios-app-path` | optional (ios) | Path to the `.app` bundle for clear-state reinstall (simulator via `simctl`, device via `devicectl`). | | `--duration` | `5m` | Total test duration (`30s`, `5m`, `2h`, `1d`). | | `--seed` | `0` | PRNG seed. `0` uses a random seed and records it in `meta.json`. | +| `--generator` | `seeded` | Who picks each action: `seeded` (the run's PRNG) or `llm` (a vision model). See [the LLM generator](../spec-language/#llm-generator). | | `--output` | `./runs` | Output directory for traces. | | `--clear-data` | `true` | Clear app data before launching so the run starts from a fresh install. Pass `--clear-data=false` to resume prior state. | diff --git a/docs/manual/spec-language.md b/docs/manual/spec-language.md index aeff4cc..b8a5908 100644 --- a/docs/manual/spec-language.md +++ b/docs/manual/spec-language.md @@ -8,14 +8,15 @@ Lookup reference for everything importable from `@sanderling/spec`. For a worked ## Module structure -A spec is a TypeScript module evaluated by the Go runner each step. It exports `properties` and `actionsRoot`, plus an optional `setup`: +A spec is a TypeScript module evaluated by the Go runner each step. It exports `properties` and `actionsRoot`, plus an optional `setup` and `generator`: ```ts import { ... } from "@sanderling/spec"; export const properties = { ... }; export const actionsRoot = weighted(...); -export const setup = login; // optional +export const setup = login; // optional +export const generator = llm(...); // optional, see below ``` `setup` is an `ActionGenerator` the runner consults before `actionsRoot` each step. While it returns actions, they run; when it returns an empty list, the runner falls through to `actionsRoot`. Use it for preconditions like login and onboarding. If the app later regresses across the precondition (a logout mid-run), `setup` re-engages on its own. @@ -266,6 +267,23 @@ InputText({ into: nameField, text: names.generate() }) InputText({ into: amountField, text: String(amounts.generate()) }) ``` +## LLM generator + +By default the run's PRNG picks each action. `--generator llm` swaps out the picker for a vision model and nothing else: same spec, same `actionsRoot`, same weights, same actions. Add the export and pick a model. + +```ts +export const generator = llm({ + model: "gpt-5.4-nano", + instructions: "Folio is a personal-finance ledger app. The home screen lists accounts with balances; you can open an account and add transactions.", +}); +``` + +Set `OPENROUTER_API_KEY` or `OPENAI_API_KEY` (OpenRouter wins if both are set). With a plain OpenAI key, drop the vendor prefix from the model id. The model needs image input and strict `json_schema` structured output. + +Each step it gets a screenshot plus a numbered list of the concrete actions your tree yields right now, each tagged with its weight, and picks one number. `instructions` are appended to the prompt: say what the app is, not how to test it — the model works that part out. Everything else is unchanged. Setup actions still run first, typing still falls back to the edge-case corpus when the model supplies no text, and the trace records the reasoning, the chosen number, and `source: "llm"` so the replay UI can show why each pick happened. + +It is one model call per step, so keep `--duration` modest. + ## Defaults ```ts