mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
archive historical results and add token budget logs
This commit is contained in:
1 parent
638e978227
commit
e94b75c6f5
813 files changed
+12631
-410710
No files matched your search
@@ -20,4 +20,4 @@ The common series setting is an offline agent container, with a separate offline
|
||||
|
||||
Preserve the exact initial and final directory state, observed changes during each episode, successful read/write tool calls, file bytes or hashes when captured, and timestamps. A final snapshot alone cannot establish who wrote a transient file or whether a peer read it. Attribute a publication only to an actor-linked successful write with captured content. Claim exact peer receipt only when a later tool response shows the relevant file bytes to that peer; a command that may have read a file without exposing its returned bytes leaves receipt unknown. Behavioral adoption also requires review of the recipient's later actions and patch. Directory listings and self-reads are not peer receipt. Distinguish task diagnosis, method publication, receipt, adoption, rejection, and independent discovery. Scorer passes require artifact and evaluator review; null and invalid-task outcomes remain separate.
|
||||
|
||||
The unit of shared history is the model's one directory across cohorts. Twenty episodes in one directory are dependent observations, not independent replications. The earlier [shared-directory pilot](../../results/team-pilot-sept7/RESULTS.md) showed peer-note reads but also used explicit sharing and an added integrity prompt, so it cannot answer this narrower question. Its once-per-second snapshots could miss transient files and its agent notes misstated some actual outcomes; use tool-level and grader evidence rather than note claims. The [baseline review](../baseline-swebench/analysis.md) flags malformed or satisfiable tasks and infrastructure nulls that must not be counted as honest failures.
|
||||
The unit of shared history is the model's one directory across cohorts. Twenty episodes in one directory are dependent observations, not independent replications. An earlier shared-directory pilot, preserved locally in `archive/results/team-pilot-sept7/`, showed peer-note reads but also used explicit sharing and an added integrity prompt, so it cannot answer this narrower question. Its once-per-second snapshots could miss transient files and its agent notes misstated some actual outcomes; use tool-level and grader evidence rather than note claims. The [baseline review](../baseline-swebench/analysis.md) flags malformed or satisfiable tasks and infrastructure nulls that must not be counted as honest failures.
|
||||
Reference in new issue
Block a user