mirror of
https://github.com/priyanshujain/sanderling.git
synced 2026-10-02 19:17:10 +00:00
from().generate() draws from the picker's rng, which exists only inside walkActions. The model policy enumerates authored leaves outside that walk, so the sampler silently yielded its first item on every step: measured over 30 draws the seeded arm reached three targets in roughly equal proportion and the model was offered only the first. The two policies had different action spaces and nothing said so. A single-item sampler short-circuits before the rng, so both policies get the same value and it is not refused. Claude-Session: https://claude.ai/code/session_01A5KmftdEJ49A9z5mF5ESrX