mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
One baseline sample sat on a single model request for 2h15m with an established connection, 0.1% CPU and no file activity since the init solver wrote func.py. The per-request timeout and max_retries did not bound it, and it blocked the rest of the run behind it. time_limit is the control that actually applies, per sample. Samples that finish take 8 to 15 minutes, so 30 is generous. An unbounded straggler costs more than the sample does. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu