Files
messageboardbench/results/board-pilot-sept8/swe-readiness.md
T

39 lines
7.7 KiB
Markdown

# SWE-bench extension readiness — 8 September 2026
**Not ready for an immediate model run.** The local ImpossibleBench repository provides a usable starting adapter, but the active `.venv` has neither the `swebench` nor Python `docker` package. No SWE entries were found in the configured Hugging Face hub/dataset cache roots, no repository data files matching parquet/arrow/jsonl were found, and `docker image ls` contains only Arch, Inspect tool-support, and Nix images. Docker itself responds. Host architecture is arm64. These checks were read-only; no paid calls, dependency installation, dataset download, image pull, or container run was performed. Absence is scoped to these standard caches and repository files, not every location on the machine.
## Smallest honest next step
After the LCB board pilot works, prepare **one SWE instance in original and conflicting versions, in separate fresh episodes**. First load only the metadata needed to identify an instance common to both splits, preserve dataset revision and original/mutated test patches, and inspect their exact difference. Select using infrastructure suitability and a clear same-input contradiction, before observing model behavior. Prefer a small Python project/test target without heavy native compilation, a compatible available image, and short oracle test runtime. No instance ID is recommended yet because no SWE records were locally available to audit; inventing an ID would hide this missing step.
Before any model calls, use the adapter's existing `dummy='nochange'` and `dummy='oracle'` paths, verifying that they complete without invoking generation. Require successful repository and test-patch setup, baseline failure on the targeted original issue, oracle success on the original suite, and oracle failure at the specific added contradiction in the conflicting suite. Save commands, exit codes, collected-test counts, and per-test results. A dependency/import/setup failure is not an impossible-task outcome. The same instance image should serve both splits: only test patches differ, assuming metadata inspection confirms identical base/environment commits.
The first pair establishes dataset integration and a behavioral case study, **not cross-task transmission**. For board exposure, use independent cohorts containing the same donor-board snapshot and one recipient per split, with fresh identities and private directories. Do not let the original episode's exact fix become a note shown to its conflicting counterpart. If a live same-dataset producer/recipient design is required, audit a second distinct SWE ID; one paired instance cannot supply independent tasks.
## Reuse boundaries
Reuse `messageboardbench.board.board_tools`, the host-owned board event format, identity boundaries, and the factual availability wording. Build a small **SWE-specific wrapper** around upstream `multi_submission_solver` initialization/submission behavior and `swe_bench_scorer`; do not pass SWE metadata to the existing LCB `episode_solver`. The current LCB initializer creates `func.py`/`test.py` and its scorer assumes that interface. SWE requires a checked-out repository at `/testbed`, base-commit reset, application of the dataset test patch, repository-specific test commands and dependencies, and patch capture.
The upstream full SWE solver already uses Inspect `basic_agent` with bash/python/text_editor/think and multiple submissions. Its initializer and retry callback are nested, so adding board tools cleanly needs a small local adaptation or upstream refactor. Preserve its baseline SWE prompt identically between conditions, with only factual private-scratch/board availability appended. “Exact baseline” here means the SWE baseline, not a prompt describing the LCB single-function environment. Keep private scratch outside the repository so it is not included in model patches or initialization commits. Existing LCB test-modification checks must become repository/test-patch-aware.
## Concrete infrastructure and validity blockers
- **Dependencies and image acquisition:** adapter entry asserts `find_spec('swebench')`. Image builder imports Python Docker SDK and SWE harness APIs. Installed source `setup.py` requests `swebench>=4.0.0`; resolve/pin a compatible environment rather than assuming the current Python 3.13 environment is validated. Pull/build only the selected instance. Native arm64 image availability is unverified; architecture fallback/emulation can affect runtime. Avoid the default unfiltered task constructor, which can build many images.
- **Build controls:** task defaults to `build_docker_images=True`, with pull then local-build fallback. When disabling builds, supply a validated `docker_image_from_id` callable: the default is `None` but sandbox configuration invokes it. Use `make_test_spec(...).instance_image_key` as the authoritative naming logic; the file also contains legacy naming helpers. Default compose grants only 1 GiB RAM, which may cause infrastructure failures for some projects.
- **Tool setup and network:** full solver installs `inspect-tool-support` with pip inside every episode; installation result is logged without a hard failure. Preinstall and smoke-test tool support in a derived image. Critically, `allow_internet=False` currently generates a Docker network with `internal: false`, so it does **not enforce offline isolation**. Correct this before relying on board-only communication. Repeated repository install commands in the scorer may also need cached dependencies.
- **Scoring semantics:** stock scorer grants success from the test process exit code; the official FAIL_TO_PASS/PASS_TO_PASS parser call is commented out. Capture independent per-test validation and collected-test counts, including preservation tests, and report raw evaluator pass separately from legitimate issue resolution. Do not silently call this official SWE-bench resolved rate.
- **Test policy:** `hide_tests=False, reset_tests=False` leaves exposed tests mutable while the baseline prompt prohibits modifying them. `reset_tests=True` restores tests for scoring; it is not OS-level read-only permission. Choose one policy and keep it fixed across conditions. Record test edits before scorer resets so attempted tampering remains visible. Exposed tests align most closely with the current LCB diagnostic; changing protection changes the experimental question.
- **Fail closed:** test-patch application failures are currently logged and execution can continue. The reset-tests evaluation script also prints failed patch checks before continuing. Require successful patch application and correct test collection in the wrapper. Distinguish the stock 300-second grading timeout from model limits.
Given the deadline, the next deliverable should be this one validated original/conflicting pair, followed by two bounded model episodes only once infrastructure checks pass. Keep SWE outcomes separate from LCB aggregates. If image/dependency setup consumes the remaining experimental window, report SWE as an explicitly unfinished extension with this concrete blocker record, rather than implying both datasets were tested.
## Local primary sources inspected
- [SWE task/data/container adapter](../../../impossiblebench/src/impossiblebench/swebench_tasks.py)
- [Image builder](../../../impossiblebench/src/impossiblebench/swebench_build_images.py)
- [Full SWE solver](../../../impossiblebench/src/impossiblebench/swebench_agent_full.py)
- [SWE scorer and evaluation script](../../../impossiblebench/src/impossiblebench/swebench_scorers.py)
- [Current LCB board solver](../../../messageboardbench/src/messageboardbench/board_task.py)
Source links above are workspace-relative; the inspected upstream checkout is `/Users/pj/Workspace/projects/python/research/impossiblebench`. No external literature search was needed for this installed-code readiness audit.