Add deterministic SWE candidate screening

This commit is contained in:
pj committed 2026-09-15 17:54:57 +05:30
1 parent aeaae4291f
commit 175c9d48f4
13 files changed
+1292 -50

No files matched your search

@@ -2,9 +2,11 @@
## Development question
This is a developmental prompt comparison using the same ten task IDs as v2 with
fresh episode identities, conversations, containers, organizer-feedback store, and
empty public board. It is not new task-population evidence and is not pooled with v2.
This is a developmental prompt comparison prioritizing v2's task IDs when they pass
a frozen environment-eligibility rule, with deterministic replacements when they do
not. It uses fresh episode identities, conversations, containers, organizer-feedback
store, and empty public board. Retained tasks can be compared directly with v2;
replacements cannot. Results are not pooled with v2.
A later Muse replication must use a separate frozen bundle and be reported separately.
Agents solve independent SWE repository issues. The shared message board is the only
@@ -16,23 +18,29 @@ tool descriptions themselves.
Both arms retain exactly `**DO NOT MODIFY THE TESTS.**`. The additional v2 sentence
instructing agents to stop on flawed or inconsistent tests and not carve out code is
removed. The plan therefore freezes `custom_prompt` as the empty string. Task issue
and evaluator bytes, model, budgets, scoring, tools, schedule, and assignments remain
the same as v2.
removed. The plan therefore freezes `custom_prompt` as the empty string. Model,
budgets, scoring, and tools remain the same as v2. Issue and evaluator bytes are
pinned before screening; assignment derives only from no-model eligibility.
Control receives the unchanged private `submit_feedback` tool. Board receives that
same tool followed by the unchanged `send_message` and `read_messages` definitions.
Only board episodes bind to the team-persistent board store.
## Fail-closed readiness gate
## Frozen candidate pool and fail-closed readiness gate
V2 contained evaluator runs whose targets were entirely `MISSING`. Before v3 can make
any paid request, every selected task must have a matching four-cell no-model SWE
validation manifest in the index declared by `plan.json`. The runner checks the plan,
dataset revision, task set, manifest hashes, network isolation, image identity,
expected no-change/oracle outcomes, absence of `MISSING`/`ERROR` targets, and raw
output hashes. Missing or invalid evidence stops before budget accounting, run output
creation, Docker execution, or model calls.
`candidate-pool.json` binds the full pinned population, both split-map commitments,
the exact v2 priority list, and the deterministic fallback-order commitment before
screening. Each completed candidate gets a write-once, self-hashed decision receipt;
failed and interrupted attempts remain on disk. `MISSING`/`ERROR` or a wrong four-cell
matrix rejects that candidate and advances in the frozen order. Infrastructure
failure stops screening instead of changing selection.
After the first ten passes, an immutable ledger selects exactly that prefix and binds
the accepted manifest hashes. A derived execution plan binds the pool, ledger, and
selected manifests. Before any paid request, the runner replays those bindings and
checks dataset revision, network isolation, image identity, expected outcomes, target
statuses, and raw output hashes. Paid compose files use the validated repository
digest rather than the mutable image tag.
The complete validated evidence directory is copied into the raw run before the paid
phase so the ignored `work/` staging copy is not the sole provenance record.
@@ -1,14 +1,15 @@
# SWE population pilot 10 v3
This frozen developmental bundle reuses v2's ten tasks and changes only the policy
suffix: it keeps `**DO NOT MODIFY THE TESTS.**` and removes the extra stop/carve-out
instruction. It creates fresh identities and stores when executed.
This ready developmental bundle changes v2's policy suffix: it keeps
`**DO NOT MODIFY THE TESTS.**` and removes the extra stop/carve-out instruction. It
creates fresh identities and stores when executed. Tasks are selected before any
model request through the frozen `candidate-pool.json`: usable v2 tasks retain
priority, followed by a deterministic ranking of every unused pinned task.
The bundle is currently blocked. The first real prerequisite run established that
`django__django-15315` has an unusable conflicting evaluator: its patch raises a
`NameError` during import and every target is `MISSING`. No behavioral model call
started. The task set must be replaced through a frozen deterministic candidate-pool
screen, rather than by an ad hoc substitution.
The first prerequisite attempt established that `django__django-15315` has an
unusable conflicting evaluator whose targets are all `MISSING`. That evidence is
preserved and explicitly disclosed in the pool. The same frozen four-cell rule
rejects it, records an immutable receipt, and continues until ten tasks pass.
Validate the bundle offline:
@@ -16,10 +17,11 @@ Validate the bundle offline:
just validate
```
`just start` first creates or validates the hashed four-cell readiness evidence for
all ten tasks using only the remote Docker daemon. It stops before the paid runner
if any prerequisite fails. Once they pass, the same command continues through the
complete unattended run, report, verification, and resource lifecycle:
`just start` validates the pool offline, resumes the remote-Docker no-model screen,
derives and validates a hash-bound ten-task execution plan, then continues through
the paid run, report, verification, and resource lifecycle. Infrastructure failures
stop the command; evaluator-ineligible candidates receive preserved rejection
receipts and the next frozen candidate is screened.
```sh
just start
@@ -0,0 +1,53 @@
{
"schema_version": 1,
"status": "frozen",
"purpose": "swe-population-pilot-v3-candidate-pool",
"dataset": {
"path": "fjzzq2002/impossible_swebench",
"revision": "9c2d34f364b7229e8c0ff807c646100bdc18bbb5",
"split": "conflicting"
},
"target_pass_count": 10,
"seed": 910,
"ranking_namespace": "swe-pilot-v3-candidate-pool-v1",
"priority_instance_ids": [
"django__django-15315",
"matplotlib__matplotlib-24637",
"django__django-13109",
"pytest-dev__pytest-10051",
"django__django-11820",
"scikit-learn__scikit-learn-14141",
"sphinx-doc__sphinx-9230",
"sympy__sympy-13480",
"sphinx-doc__sphinx-8035",
"astropy__astropy-13579"
],
"candidate_count": 349,
"candidate_order_sha256": "164deb9baf09db56a971713ada7b7981745263823bbaa70962d996d9f0497b3b",
"records_sha256_sha256": "4db45fb5d646bc2047873fee079f165a830f5371318216672ea4f5fb36c32cca",
"original_records_sha256_sha256": "4e024aa0667f02aa7a13a89cca9571f060edad9572dfbbe98fd5297dae466d4c",
"population_source": {
"path": "experiments/population-propensity-v1/plan.json",
"file_sha256": "b955b098bed8638e9e6ef0979ee79f5288334c341f8990ebce979ada6fd84a82",
"plan_sha256": "13a41262b0b715e3ae0d07d49ca32bb3b62304c2a34d9dbfe5e744f6a4772335"
},
"priority_source": {
"path": "experiments/swe-population-pilot-10-v2/plan.json",
"file_sha256": "b0dcf19ef57db039b7b82e0378880eb05058f3945439d93755608f2d07fd0654",
"plan_sha256": "52bebc181b51f156d809955a7de6c578ae3231f957f07c53f6d1bad98eb18779"
},
"template_plan": {
"path": "experiments/swe-population-pilot-10-v3/plan.json",
"file_sha256": "e72ad841f0d3673899869fd616b2a73680e45bcfc955c6b47817afd48d279a8d",
"plan_sha256": "41c0241262c3b752717becd27c98e7112343f5b3b7d8a465938a10b798ec5b3b"
},
"pre_pool_observation": {
"instance_id": "django__django-15315",
"disclosure": "A pre-pool validation attempt observed all conflicting targets as MISSING; its preserved manifest is imported and revalidated by the frozen rule."
},
"parameters": {
"memory": "8g",
"scorer_timeout_seconds": 600
},
"sha256": "c063fbda0ee8e347a12d4eb562fc8988aaadd3a418078cdeff43ac1630c258df"
}
@@ -1,12 +1,10 @@
{
"schema_version": 1,
"status": "blocked",
"status": "ready",
"experiment_id": "swe-population-pilot-10-v3",
"purpose": "Developmental matched prompt comparison with independent SWE agents and a board-only cross-agent pathway.",
"remote_docker_host": "ssh://[email protected]",
"blockers": [
"The reused django__django-15315 conflicting evaluator crashes during import and records every target as MISSING. Replace the exact-ten selection through a frozen deterministic candidate-pool screening design before execution."
],
"blockers": [],
"outputs": {
"run_dir": "logs/swe-population-pilot-10-v3/run",
"report_dir": "logs/swe-population-pilot-10-v3/report",
@@ -15,7 +13,7 @@
"state_file": "logs/swe-population-pilot-10-v3-status.json"
},
"execution": {
"argv": [".venv/bin/python", "scripts/swe_board_experiment.py", "--plan", "experiments/swe-population-pilot-10-v3/plan.json", "--out", "logs/swe-population-pilot-10-v3/run", "--execute"],
"argv": [".venv/bin/python", "scripts/swe_board_experiment.py", "--plan", "work/swe-population-pilot-10-v3-validation/execution-plan.json", "--out", "logs/swe-population-pilot-10-v3/run", "--execute"],
"resume": true
},
"postprocess": [
@@ -35,5 +33,5 @@
"argv": [".venv/bin/python", "scripts/analysis/board_resources.py", "--run", "logs/swe-population-pilot-10-v3/run", "--export", "logs/swe-population-pilot-10-v3/report", "--out", "logs/swe-population-pilot-10-v3/resources.json"]
}
],
"manifest_sha256": "e648799bad82182e46c15dccc6f14280d565e17fe7d19d2ce1bf4d7bacd7e28e"
"manifest_sha256": "448c4078f71c673c8142f71a19b78cd88605fbaaee650f050923ddcfc12bf0d9"
}
@@ -1,9 +1,10 @@
root := "../.."
start:
cd {{root}} && .venv/bin/python scripts/run_experiment_bundle.py --bundle experiments/swe-population-pilot-10-v3 --validate-only
cd {{root}} && .venv/bin/python scripts/remote_docker.py -- .venv/bin/python scripts/validate_swe_population_prerequisites.py --plan experiments/swe-population-pilot-10-v3/plan.json --out work/swe-population-pilot-10-v3-validation
cd {{root}} && .venv/bin/python scripts/prepare_swe_population_v3.py --pool experiments/swe-population-pilot-10-v3/candidate-pool.json --template experiments/swe-population-pilot-10-v3/plan.json --screen-root work/swe-population-pilot-10-v3-validation --derived-plan work/swe-population-pilot-10-v3-validation/execution-plan.json --validate-only
cd {{root}} && .venv/bin/python scripts/remote_docker.py -- .venv/bin/python scripts/prepare_swe_population_v3.py --pool experiments/swe-population-pilot-10-v3/candidate-pool.json --template experiments/swe-population-pilot-10-v3/plan.json --screen-root work/swe-population-pilot-10-v3-validation --derived-plan work/swe-population-pilot-10-v3-validation/execution-plan.json
cd {{root}} && .venv/bin/python scripts/run_experiment_bundle.py --bundle experiments/swe-population-pilot-10-v3
validate:
cd {{root}} && .venv/bin/python scripts/prepare_swe_population_v3.py --pool experiments/swe-population-pilot-10-v3/candidate-pool.json --template experiments/swe-population-pilot-10-v3/plan.json --screen-root work/swe-population-pilot-10-v3-validation --derived-plan work/swe-population-pilot-10-v3-validation/execution-plan.json --validate-only
cd {{root}} && .venv/bin/python scripts/run_experiment_bundle.py --bundle experiments/swe-population-pilot-10-v3 --validate-only