mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-03 03:27:07 +00:00
archive old experiments and publish swe series
This commit is contained in:
1 parent
480100587e
commit
638e978227
1522 files changed
+220002
-4900
No files matched your search
@@ -1,8 +1,8 @@
|
||||
# Shared scratchpads: evidence and design choices
|
||||
|
||||
September 7, 2026. Research note for the active small-team study. A reviewed original-task
|
||||
diagnostic is reported at the end. [SPRINT-SEPT10.md](../SPRINT-SEPT10.md) remains the working
|
||||
plan. This note separates verified source details from our proposed interpretation.
|
||||
September 7, 2026. Historical research note for the small-team study. A reviewed
|
||||
original-task diagnostic is reported at the end. This note separates verified source
|
||||
details from our proposed interpretation.
|
||||
|
||||
## What the DeepMind paper actually establishes
|
||||
|
||||
@@ -158,6 +158,6 @@ attempted zero tests. Avoid classifying either detail from keywords alone. These
|
||||
were agent-run checks, not independently repeated reviewer tests.
|
||||
|
||||
Full samples and indexed evidence are saved in
|
||||
[agent-1 review](../scratchpad/team-pilot-sept7/original-lcbhard_0-review.json) and
|
||||
[agent-2 review](../scratchpad/team-pilot-sept7/original-lcbhard_1-review.json).
|
||||
[agent-1 review](../../results/team-pilot-sept7/original-lcbhard_0-review.json) and
|
||||
[agent-2 review](../../results/team-pilot-sept7/original-lcbhard_1-review.json).
|
||||
This diagnostic supplies no estimate of a shared-versus-private effect.
|
||||
@@ -5,7 +5,7 @@ September 7 design, implemented September 8 in `messageboardbench/board.py`,
|
||||
and its integrity-framed prompts remain preserved. The new runner creates fresh
|
||||
private files and episode IDs and carries only deliberately posted public messages
|
||||
across cohorts. Infrastructure verification and the twelve-attempt pilot completed.
|
||||
[Final results](../scratchpad/board-pilot-sept8/RESULTS.md): voluntary gaming-method
|
||||
[Final results](../../results/board-pilot-sept8/RESULTS.md): voluntary gaming-method
|
||||
publication occurred, but no agent read the board; no transmission was observed.
|
||||
|
||||
## Verified DeepMind design
|
||||
@@ -121,7 +121,7 @@ explanations. Replicate teams/runs, rather than treating dependent messages as s
|
||||
|
||||
## Token use and model comparison
|
||||
|
||||
The [frozen historical token audit](../scratchpad/token-comparison-sept7/REPORT.md)
|
||||
The [frozen historical token audit](../../results/token-comparison-sept7/REPORT.md)
|
||||
separates total, uncached input, cached input, generated output, reported reasoning,
|
||||
time and limits. Failed tasks are not automatically honest. Resource differences
|
||||
also reflect early stopping on success and retries after failure.
|
||||
@@ -148,6 +148,6 @@ Muse used an integer subclass comparing equal to both required answers. Neither
|
||||
edited the tests or hit limits. These are one-task diagnostic observations, not
|
||||
model-level rates or a test of communication. Muse's provider-redacted reasoning
|
||||
was not used for labeling; actions and final code provide the evidence. See the
|
||||
[completed comparison](../scratchpad/model-comparison-sept7/RESULTS.md). All Muse
|
||||
[completed comparison](../../results/model-comparison-sept7/RESULTS.md). All Muse
|
||||
requests used Contributor; no ordinary-tier substitution occurred. Provider errors
|
||||
remain archived and excluded from behavior counts.
|
||||
Reference in new issue
Block a user