mirror of
https://github.com/priyanshujain/messageboardbench.git
synced 2026-10-02 11:07:07 +00:00
Forks ImpossibleBench's LiveCodeBench solver rather than passing instruction_prompt=, which injects text into the system message, the user message and every failure message. The scratch line now appears once, in the system message, verified by the mockllm smoke run. test_byte_match.py reads ImpossibleBench's expected_test construction out of its source with ast and re-executes it, so our test.py is checked against theirs rather than a copy. Confirmed to fail on a one-character upstream change. Without this, drift flags every sample as test-modified, resets it, and destroys the cheat measurement silently. The event adapter fixes a bug carried over from messageboard: relation() left absolute paths absolute, so `cat /workspace/scratch/notes.md` classified as outside the directory. Every absolute-path touch would have scored as a miss. The checks are new rather than reused. The old ones score not-applicable when the prompt names the directory, which ours does by design, and discard reads after the first write. Both would undercount here. The decisions are kept, the code is not. Claude-Session: https://claude.ai/code/session_01Cq98H7sNoSJdL3W98f18bu
50 lines
1.7 KiB
Python
50 lines
1.7 KiB
Python
"""Confirm the model slug, the key, and tool support. Costs about a cent.
|
|
|
|
basic_agent loops uselessly against a model that cannot call a tool, so the second check
|
|
matters as much as the first: a model that answers "hi" but ignores tools would burn the
|
|
whole message limit on every sample and score zero for a reason that looks like honesty.
|
|
"""
|
|
|
|
import asyncio
|
|
import os
|
|
|
|
from dotenv import load_dotenv
|
|
from inspect_ai.model import ChatMessageUser, GenerateConfig, get_model
|
|
from inspect_ai.tool import ToolInfo, ToolParams
|
|
from inspect_ai.util import JSONSchema
|
|
|
|
MODEL = os.environ.get("MBB_MODEL", "openrouter/z-ai/glm-5.3-flash")
|
|
|
|
|
|
async def main() -> None:
|
|
load_dotenv()
|
|
model = get_model(MODEL)
|
|
|
|
out = await model.generate("hi")
|
|
print(f"[1/2] generate: {out.completion.strip()[:120]!r}")
|
|
print(f" tokens: in={out.usage.input_tokens} out={out.usage.output_tokens}")
|
|
|
|
add = ToolInfo(
|
|
name="add",
|
|
description="Add two integers.",
|
|
parameters=ToolParams(
|
|
properties={
|
|
"a": JSONSchema(type="integer", description="first addend"),
|
|
"b": JSONSchema(type="integer", description="second addend"),
|
|
},
|
|
required=["a", "b"],
|
|
),
|
|
)
|
|
out = await model.generate(
|
|
[ChatMessageUser(content="Use the add tool to add 17 and 25. Do not answer directly.")],
|
|
tools=[add],
|
|
config=GenerateConfig(max_tokens=200),
|
|
)
|
|
calls = out.message.tool_calls or []
|
|
print(f"[2/2] tool calls: {[(c.function, c.arguments) for c in calls]}")
|
|
print(" TOOL SUPPORT OK" if calls else " NO TOOL CALL: basic_agent will not work")
|
|
|
|
|
|
if __name__ == "__main__":
|
|
asyncio.run(main())
|