archive historical results and add token budget logs

This commit is contained in:
pj committed 2026-09-25 22:07:26 +05:30
1 parent 638e978227
commit e94b75c6f5
813 files changed
+12631 -410710

No files matched your search

+1 -1
View File
@@ -20,4 +20,4 @@ The common series setting is an offline agent container, with a separate offline
Preserve the exact initial and final directory state, observed changes during each episode, successful read/write tool calls, file bytes or hashes when captured, and timestamps. A final snapshot alone cannot establish who wrote a transient file or whether a peer read it. Attribute a publication only to an actor-linked successful write with captured content. Claim exact peer receipt only when a later tool response shows the relevant file bytes to that peer; a command that may have read a file without exposing its returned bytes leaves receipt unknown. Behavioral adoption also requires review of the recipient's later actions and patch. Directory listings and self-reads are not peer receipt. Distinguish task diagnosis, method publication, receipt, adoption, rejection, and independent discovery. Scorer passes require artifact and evaluator review; null and invalid-task outcomes remain separate.
The unit of shared history is the model's one directory across cohorts. Twenty episodes in one directory are dependent observations, not independent replications. The earlier [shared-directory pilot](../../results/team-pilot-sept7/RESULTS.md) showed peer-note reads but also used explicit sharing and an added integrity prompt, so it cannot answer this narrower question. Its once-per-second snapshots could miss transient files and its agent notes misstated some actual outcomes; use tool-level and grader evidence rather than note claims. The [baseline review](../baseline-swebench/analysis.md) flags malformed or satisfiable tasks and infrastructure nulls that must not be counted as honest failures.
The unit of shared history is the model's one directory across cohorts. Twenty episodes in one directory are dependent observations, not independent replications. An earlier shared-directory pilot, preserved locally in `archive/results/team-pilot-sept7/`, showed peer-note reads but also used explicit sharing and an added integrity prompt, so it cannot answer this narrower question. Its once-per-second snapshots could miss transient files and its agent notes misstated some actual outcomes; use tool-level and grader evidence rather than note claims. The [baseline review](../baseline-swebench/analysis.md) flags malformed or satisfiable tasks and infrastructure nulls that must not be counted as honest failures.