mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
fix(trace): an action names the generator that produced it
The setup exclusion landed for the model arm only, because only a model pick stamped a source. A seeded run returned setup's action through the same entry with no marker, so its denominator still counted the login while the model arm's did not, and the two are compared. serializeAction names setup and seeded on the wire, so both arms are counted by one rule. An already-recorded trace names nothing and keeps exactly the count it was reported with; unattributed_actions counts those steps so the old denominator cannot pass as the new one. TraceVersion is deliberately unbumped: oracle-reduction refuses a differing version, and a bump would make all 169 recorded runs unreplayable.
This commit is contained in:
1 parent
66fd5bce5d
commit
454988fbc8
19 files changed
+377
-82
No files matched your search
@@ -59,6 +59,8 @@ Preconditions like login run through the spec's `setup` export (see the [case st
|
||||
| 30 min | ~15s | 0.8% |
|
||||
| 1 hour | ~15s | 0.4% |
|
||||
|
||||
Those steps are still actions the app received, so the trace names them: every action it records carries a `source` saying whether `setup`, the seeded picker or the model produced it. Anything measured per action counts the last two.
|
||||
|
||||
## Session state
|
||||
|
||||
Session tokens, keychain entries, shared preferences, and cookies survive the whole run. If the app logs the user out mid-run, the gating extractor flips, `setup` re-engages, and the run logs back in. No retry logic needed in the spec.
|
||||
|
||||
@@ -343,6 +343,8 @@ Set `OPENROUTER_API_KEY` or `OPENAI_API_KEY` (OpenRouter wins if both are set).
|
||||
|
||||
Each step it gets a screenshot plus a numbered list of the concrete actions your tree yields right now, each tagged with its weight, and picks one number. That list is the seeded picker's own candidate enumeration, so both modes explore the same action space and only the choice differs. `instructions` are appended to the prompt: say what the app is, not how to test it; the model works that part out. Everything else is unchanged. Setup actions still run first, typing still falls back to the edge-case corpus when the model supplies no text, and the trace records the reasoning, the chosen number, and `source: "llm"` so the replay UI can show why each pick happened.
|
||||
|
||||
Every dispatched action names its producer in the trace that way, whichever mode drove it: `"setup"` for a step the spec's `setup` drove, `"seeded"` for the seeded picker's own pick, `"llm"` for the model's. Only the last two explored anything, so a rate measured per action divides by those and leaves the login out on both modes.
|
||||
|
||||
One thing this mode refuses outright: a sampler drawn inside an `actions()` leaf. `from(...).generate()` draws from the seeded picker's stream, which the model never enters, so it would be offered the first item on every step while a seeded run reaches all of them. The run stops and names the leaf rather than degrade quietly. Return one action per item instead (`cards.map(card => Tap({ on: card }))`); a one-item list never draws and needs no change.
|
||||
|
||||
The value generators draw from that same stream, so they refuse there too: an authored `InputText({ into: field, text: String(amounts.generate()) })` would type one fixed value on every model step while the seeded arm varies it, which is a different experiment rather than a different action space. Pass a fixed value instead. A generator whose span is a single value (`integers().between(7, 7)`, `strings().length(0, 0)`) never draws and is left alone, and `setup` is untouched: it runs through the picker with its seed under both generators, so a sampler there keeps drawing.
|
||||
|
||||
Reference in new issue
Block a user