Files
messageboardbench/experiments/swe-population-pilot-10-v1/DESIGN.md
T

59 lines
3.0 KiB
Markdown

# SWE population pilot 10 v1
## Purpose
Run a small end-to-end version of the homogeneous population experiment before the
full 349-task study. The pilot is intended to reveal concrete execution failures,
estimate resource use, and provide an initial descriptive control-versus-board
signal. It does not replace or alter the full frozen experiment.
## Frozen population and assignment
Ten task IDs are selected deterministically from all 349 official ImpossibleBench
SWE `conflicting` records at revision
`9c2d34f364b7229e8c0ff807c646100bdc18bbb5`. Selection ranks every instance ID by
SHA-256 of `swe-pilot-selection-v1:910:{instance_id}` and freezes the first ten in
`plan.json`. Selection does not inspect task content or outcomes.
Every selected task runs exactly once in each condition, giving 20 episodes:
- control: the upstream ImpossibleBench tools scaffold with no board;
- board: the same scaffold plus neutral `board_read` and `board_post` tools.
There is one matched team divided into two ordered cohorts of five tasks. The board
persists across both board cohorts and starts empty. There are no seeded posts,
mandatory reads, pushed notifications, scratchpad prompt, or collaboration prose.
## Held constant
The two arms use identical task IDs, issue/evaluator bytes, model
`openrouter/z-ai/glm-5.3-flash`, upstream tools prompt and strict contradiction
instruction, scorer, temperature 1, medium reasoning effort, 4,096 reasoning tokens,
100-message limit, 1,000,000-token episode limit, 1,800-second episode limit, ten
submissions, and isolated 8 GiB containers. Containers have no network and run only
through the remote x86-64 Docker daemon at `ssh://[email protected]`.
The only treatment difference is the availability of the two board tools and access
to posts from other agents in the same board population. This is board versus no
board, so interface availability and peer-message availability are jointly treated.
## Outcomes and interpretation
The unattended deterministic report records evaluator success with protected tests,
test modification, failures and missingness, complete paired task outcomes, board
posts, reads, confirmed peer receipts, artifacts, transcripts, tokens, calls, and
elapsed resources. Intent, method publication, and attributed adoption remain manual
review outcomes and are not inferred automatically.
With only one treated board, statistical inference at the population-assignment level
is not meaningful. Any effect estimate and sign-flip value in the generic report are
descriptive pilot outputs. A zero, floor, or ceiling result is still a valid pilot
result and does not trigger task replacement or prompt tuning.
## Lifecycle
`just start` validates the frozen bundle, executes or resumes the 20 assignments,
and then generates `REPORT.md`, `report.json`, verification, and resource summaries.
The first matched cohort is the engineering sentinel; behavioral failures do not stop
it, while missing required execution artifacts do. No Codex monitoring is required.