Files
messageboardbench/docs/swe-scenario-series.md

2.7 KiB

SWE-bench scenario series

This series reuses the baseline's exact 20 fjzzq2002/impossible_swebench records at revision 9c2d34f364b7229e8c0ff807c646100bdc18bbb5, including record hashes, order, and four sequential cohorts of five.

Scenario Status Record
Isolated baseline Complete Analysis
Isolated offline control Prepared Design
Shared scratchpad Complete Analysis
Shared board Analyzed; DeepSeek incomplete Analysis
Board and organizer feedback Complete Analysis
Board and internet Prepared Design
Board and token awareness Running under PJ's control Design

Each scenario runs openrouter/z-ai/glm-5.3-flash, openrouter/meta/muse-spark-1.3-contributor, and openrouter/deepseek/deepseek-v4-pro-0813. Each model has independent communication storage and history. The baseline DeepSeek run used ClinePass, so it is not a provider-matched control for the new DeepSeek runs.

The new isolated control, scratchpad, board, feedback, and token-awareness plans set agent networking to none. The internet plans use a project-specific bridge. All new scenarios use an offline fresh grader, immutable grader image identity, strict null treatment for missing or error target statuses, a 100-message limit, 1,000,000 metered-token limit, 1,800-second episode limit, 600-second grader limit, ten submissions, and temperature 1.

The scratchpad intervention advertises only /workspace/scratch; it does not disclose peer access. Board scenarios use the same optional team prompt and post_message and read_messages tools. Feedback adds only the private write-only organizer tool. Internet adds only outbound agent network access. Token awareness adds the exact budget sentence and check_remaining_tokens tool.

The primary comparisons are scratchpad versus isolated offline, board versus isolated offline, feedback versus board, internet versus board, and token awareness versus board. Scratchpad versus board is descriptive because both the medium and model-visible framing differ.

Three malformed tasks and two satisfiable tasks are outside the valid contradictory-task estimand. Null grader outcomes are infrastructure outcomes rather than behavioral failures. One shared history per model is the propagation inference unit. Publication, exact delivery, attributed adoption, candidate adoption, rejection, and independent discovery remain separate labels.