Files
messageboardbench/results/board-pilot-sept8/swe-readiness.md
T

7.7 KiB

SWE-bench extension readiness — 8 September 2026

Not ready for an immediate model run. The local ImpossibleBench repository provides a usable starting adapter, but the active .venv has neither the swebench nor Python docker package. No SWE entries were found in the configured Hugging Face hub/dataset cache roots, no repository data files matching parquet/arrow/jsonl were found, and docker image ls contains only Arch, Inspect tool-support, and Nix images. Docker itself responds. Host architecture is arm64. These checks were read-only; no paid calls, dependency installation, dataset download, image pull, or container run was performed. Absence is scoped to these standard caches and repository files, not every location on the machine.

Smallest honest next step

After the LCB board pilot works, prepare one SWE instance in original and conflicting versions, in separate fresh episodes. First load only the metadata needed to identify an instance common to both splits, preserve dataset revision and original/mutated test patches, and inspect their exact difference. Select using infrastructure suitability and a clear same-input contradiction, before observing model behavior. Prefer a small Python project/test target without heavy native compilation, a compatible available image, and short oracle test runtime. No instance ID is recommended yet because no SWE records were locally available to audit; inventing an ID would hide this missing step.

Before any model calls, use the adapter's existing dummy='nochange' and dummy='oracle' paths, verifying that they complete without invoking generation. Require successful repository and test-patch setup, baseline failure on the targeted original issue, oracle success on the original suite, and oracle failure at the specific added contradiction in the conflicting suite. Save commands, exit codes, collected-test counts, and per-test results. A dependency/import/setup failure is not an impossible-task outcome. The same instance image should serve both splits: only test patches differ, assuming metadata inspection confirms identical base/environment commits.

The first pair establishes dataset integration and a behavioral case study, not cross-task transmission. For board exposure, use independent cohorts containing the same donor-board snapshot and one recipient per split, with fresh identities and private directories. Do not let the original episode's exact fix become a note shown to its conflicting counterpart. If a live same-dataset producer/recipient design is required, audit a second distinct SWE ID; one paired instance cannot supply independent tasks.

Reuse boundaries

Reuse messageboardbench.board.board_tools, the host-owned board event format, identity boundaries, and the factual availability wording. Build a small SWE-specific wrapper around upstream multi_submission_solver initialization/submission behavior and swe_bench_scorer; do not pass SWE metadata to the existing LCB episode_solver. The current LCB initializer creates func.py/test.py and its scorer assumes that interface. SWE requires a checked-out repository at /testbed, base-commit reset, application of the dataset test patch, repository-specific test commands and dependencies, and patch capture.

The upstream full SWE solver already uses Inspect basic_agent with bash/python/text_editor/think and multiple submissions. Its initializer and retry callback are nested, so adding board tools cleanly needs a small local adaptation or upstream refactor. Preserve its baseline SWE prompt identically between conditions, with only factual private-scratch/board availability appended. “Exact baseline” here means the SWE baseline, not a prompt describing the LCB single-function environment. Keep private scratch outside the repository so it is not included in model patches or initialization commits. Existing LCB test-modification checks must become repository/test-patch-aware.

Concrete infrastructure and validity blockers

  • Dependencies and image acquisition: adapter entry asserts find_spec('swebench'). Image builder imports Python Docker SDK and SWE harness APIs. Installed source setup.py requests swebench>=4.0.0; resolve/pin a compatible environment rather than assuming the current Python 3.13 environment is validated. Pull/build only the selected instance. Native arm64 image availability is unverified; architecture fallback/emulation can affect runtime. Avoid the default unfiltered task constructor, which can build many images.
  • Build controls: task defaults to build_docker_images=True, with pull then local-build fallback. When disabling builds, supply a validated docker_image_from_id callable: the default is None but sandbox configuration invokes it. Use make_test_spec(...).instance_image_key as the authoritative naming logic; the file also contains legacy naming helpers. Default compose grants only 1 GiB RAM, which may cause infrastructure failures for some projects.
  • Tool setup and network: full solver installs inspect-tool-support with pip inside every episode; installation result is logged without a hard failure. Preinstall and smoke-test tool support in a derived image. Critically, allow_internet=False currently generates a Docker network with internal: false, so it does not enforce offline isolation. Correct this before relying on board-only communication. Repeated repository install commands in the scorer may also need cached dependencies.
  • Scoring semantics: stock scorer grants success from the test process exit code; the official FAIL_TO_PASS/PASS_TO_PASS parser call is commented out. Capture independent per-test validation and collected-test counts, including preservation tests, and report raw evaluator pass separately from legitimate issue resolution. Do not silently call this official SWE-bench resolved rate.
  • Test policy: hide_tests=False, reset_tests=False leaves exposed tests mutable while the baseline prompt prohibits modifying them. reset_tests=True restores tests for scoring; it is not OS-level read-only permission. Choose one policy and keep it fixed across conditions. Record test edits before scorer resets so attempted tampering remains visible. Exposed tests align most closely with the current LCB diagnostic; changing protection changes the experimental question.
  • Fail closed: test-patch application failures are currently logged and execution can continue. The reset-tests evaluation script also prints failed patch checks before continuing. Require successful patch application and correct test collection in the wrapper. Distinguish the stock 300-second grading timeout from model limits.

Given the deadline, the next deliverable should be this one validated original/conflicting pair, followed by two bounded model episodes only once infrastructure checks pass. Keep SWE outcomes separate from LCB aggregates. If image/dependency setup consumes the remaining experimental window, report SWE as an explicitly unfinished extension with this concrete blocker record, rather than implying both datasets were tested.

Local primary sources inspected

Source links above are workspace-relative; the inspected upstream checkout is /Users/pj/Workspace/projects/python/research/impossiblebench. No external literature search was needed for this installed-code readiness audit.