mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
docs(manual): value generators are refused under the model policy too
Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX
This commit is contained in:
1 parent
3cab483b96
commit
f307354bc7
1 file changed
+3
-1
@@ -295,7 +295,9 @@ Set `OPENROUTER_API_KEY` or `OPENAI_API_KEY` (OpenRouter wins if both are set).
|
|||||||
|
|
||||||
Each step it gets a screenshot plus a numbered list of the concrete actions your tree yields right now, each tagged with its weight, and picks one number. That list is the seeded picker's own candidate enumeration, so both modes explore the same action space and only the choice differs. `instructions` are appended to the prompt: say what the app is, not how to test it; the model works that part out. Everything else is unchanged. Setup actions still run first, typing still falls back to the edge-case corpus when the model supplies no text, and the trace records the reasoning, the chosen number, and `source: "llm"` so the replay UI can show why each pick happened.
|
Each step it gets a screenshot plus a numbered list of the concrete actions your tree yields right now, each tagged with its weight, and picks one number. That list is the seeded picker's own candidate enumeration, so both modes explore the same action space and only the choice differs. `instructions` are appended to the prompt: say what the app is, not how to test it; the model works that part out. Everything else is unchanged. Setup actions still run first, typing still falls back to the edge-case corpus when the model supplies no text, and the trace records the reasoning, the chosen number, and `source: "llm"` so the replay UI can show why each pick happened.
|
||||||
|
|
||||||
One thing this mode refuses outright: a multi-item sampler inside an `actions()` leaf. `from(...).generate()` draws from the seeded picker's stream, which the model never enters, so it would be offered the first item on every step while a seeded run reaches all of them. The run stops and names the leaf rather than degrade quietly. Return one action per item instead (`cards.map(card => Tap({ on: card }))`); a one-item list never draws and needs no change.
|
One thing this mode refuses outright: a sampler drawn inside an `actions()` leaf. `from(...).generate()` draws from the seeded picker's stream, which the model never enters, so it would be offered the first item on every step while a seeded run reaches all of them. The run stops and names the leaf rather than degrade quietly. Return one action per item instead (`cards.map(card => Tap({ on: card }))`); a one-item list never draws and needs no change.
|
||||||
|
|
||||||
|
The value generators draw from that same stream, so they refuse there too: an authored `InputText({ into: field, text: String(amounts.generate()) })` would type one fixed value on every model step while the seeded arm varies it, which is a different experiment rather than a different action space. Pass a fixed value instead. A generator whose span is a single value (`integers().between(7, 7)`, `strings().length(0, 0)`) never draws and is left alone, and `setup` is untouched: it runs through the picker with its seed under both generators, so a sampler there keeps drawing.
|
||||||
|
|
||||||
Every step of a model-driven run also appends one record to `llm-calls.jsonl` in the run directory, keyed by the step index it shares with `trace.jsonl`: the system and user prompts as sent, the numbered candidate list as the model saw it, the path of the screenshot that went with the call, the raw response, token counts, latency, and an `outcome`. `outcome` is `selected` when the pick became an action; every other value (`echo_mismatch`, `choice_out_of_range`, `unparsable_response`, `request_failed`, `no_candidates`, `setup_action`, ...) names a step that ran no model-chosen action, so a step a guard threw away can never be mistaken for one where the model had nothing to pick. A seeded run writes no such file.
|
Every step of a model-driven run also appends one record to `llm-calls.jsonl` in the run directory, keyed by the step index it shares with `trace.jsonl`: the system and user prompts as sent, the numbered candidate list as the model saw it, the path of the screenshot that went with the call, the raw response, token counts, latency, and an `outcome`. `outcome` is `selected` when the pick became an action; every other value (`echo_mismatch`, `choice_out_of_range`, `unparsable_response`, `request_failed`, `no_candidates`, `setup_action`, ...) names a step that ran no model-chosen action, so a step a guard threw away can never be mistaken for one where the model had nothing to pick. A seeded run writes no such file.
|
||||||
|
|
||||||
|
|||||||
Reference in new issue
Block a user