Skip to content

feat(adapters): AssistantBench suite import and its answer metric as a third stage - #379

Open
nahatav wants to merge 3 commits into
TIGER-AI-Lab:mainfrom
nahatav:feat/assistantbench-adapter
Open

nahatav wants to merge 3 commits into
TIGER-AI-Lab:mainfrom
nahatav:feat/assistantbench-adapter

Conversation

@nahatav

@nahatav nahatav commented Sep 23, 2026 •

Copy link
Copy Markdown

What does this PR do?

Advances #188 (and #72): ClawBench can now run AssistantBench tasks and report AssistantBench's own answer metric next to the existing two stages.

clawbench-assistantbench-adapt --input ./assistantbench/validation.jsonl \
    --output-dir test-cases/assistantbench

clawbench-batch --models <m> --cases-dir test-cases/assistantbench --all-cases

clawbench-assistantbench-score test-output/<m> \
    --gold test-cases/assistantbench/assistantbench-gold.json

Two commands, both new, nothing existing changed beyond a CHANGELOG line, a CLI table row, and two entries in HELP_MODULES.

The design decision worth reviewing

AssistantBench grades a free-text answer. There is no final write request for Stage-1 to intercept, which is the one thing ClawBench's scoring keys on.

The easy answer was to introduce an eval_schema.mode: "judge" discriminator and the runner branch that docs/answer-mode-tasks.md sketches. I did not, because main already solves this shape and ships tasks that use it: the claw-eval port tells the agent to submit its answer at the runtime server's http://127.0.0.1:7878/submit form, that form POSTs to /api/task-submit, the task's eval_schema targets that endpoint, and the judge scores the submitted answer against judge_context. This adapter reuses that path verbatim, same endpoint and same footer wording, so an AssistantBench task runs through the ordinary pipeline with zero runner changes and produces the standard five-layer bundle. A test pins eval_schema and the footer against a bundled claw-eval task so the two ports cannot drift apart.

This is the same decision #356 made for WebVoyager, reached independently. I built on main rather than stacking on #345 so this reviews and merges on its own; the loader is a pure function and drops into src/clawbench/adapters/ unchanged if that lands.

Three scores, not two

Stage Question Where
1 Did the agent commit to an answer at all? run-meta.intercepted
2 Does the answer fulfil the instruction, per the rubric? judge.json
3 Does the answer match gold, by AssistantBench's metric? clawbench-assistantbench-score

Stage 3 has no model in the loop, which is the point of adding it: it is an independent, deterministic check on the LLM judge over the same runs. --write-run-meta records it as assistantbench_answer_score in each run's run-meta.json, which is what #188 asks for. It is off by default because it edits run output you already have.

Mapping

ClawBench AssistantBench
metadata.source_task_id id, verbatim, the real key
metadata.task_id 600000 + n, n over ids in sorted order
metadata.class the expertise field inside metadata
metadata.sites_involved hostnames parsed out of gold_url
instruction task + answer-format block + the claw-eval submit footer
eval_schema fixed POST /api/task-submit
time_limit --time-limit, default 30, since upstream is long-horizon
judge_context.reference_solution answer + explanation + gold_url
judge_context.rubric derived from whether a gold answer ships

Case names are keyed on the source id, not the row's position, so re-exporting the dataset in a different order does not rename every case. A name collision is an error, never a silent overwrite.

Gold answers never enter the container. run.py mounts only the eval_schema block as /eval-schema.json, so metadata and judge_context stay host-side where the judge and the scorer read them. A test asserts the gold answer, its values and its source hostnames appear in no agent-facing field.

The metric, and how I know it is right

assistantbench_score.py re-implements the evaluation/ package of the AssistantBench leaderboard Space (Apache-2.0, string metric from DROP's drop_eval). It is a re-implementation rather than a vendored copy because upstream needs numpy and scipy and ClawBench takes neither. scipy.optimize.linear_sum_assignment is replaced by an exact rectangular Jonker-Volgenant solver; the tests check it against brute force over random matrices in both orientations, and one test pins a case where greedy would lose.

Re-implementing a metric is exactly where a port quietly drifts, so I checked it rather than asserting it. I stood the upstream evaluator up under numpy 2.5.3 and scipy 1.18.1 and ran both scorers over 7249 (prediction, gold) pairs: the 3249-pair exhaustive product of a corpus covering all four answer types, plus 4000 fuzzed pairs.

exhaustive  pairs=  3249  agree=  2835  disagree=   0  upstream_raised= 414  port_raised=   0
fuzz        pairs=  4000  agree=  3746  disagree=   0  upstream_raised= 254  port_raised=   0

TOTAL comparable=6581  agree=6581  disagree=0  upstream_raised=668  port_raised=0

Exact agreement to 1e-9 on accuracy and answer rate for every pair upstream can score. The 668 it cannot are all malformed predictions where upstream raises TypeError/AttributeError: a record object answering a numeric question, null, true, a bare list. The port scores those 0.0 rather than taking down a whole batch's scoring.

That run found two real bugs in my first draft, both now fixed and both pinned by a test:

  • opposite signs. numpy gets to 0 by way of nan, since max(0, nan) is 0; math.log raises instead, so the sign is checked first.
  • a bare number list against a prose answer. Upstream's tokenizer raises on a non-string span and scores that 0. My draft stringified the spans and handed out token overlap upstream never awards.

Every remaining deviation is enumerated in the module docstring, as are the upstream quirks kept on purpose: , read as a decimal point so "1,000" parses as 1.0, the log metric exceeding 1 when both values are negative, precision being recall with the arguments swapped. Those have their own tests saying, in words, not to fix them.

Test plan

  • tests/test_assistantbench_score.py, 55 cases: the assignment solver against brute force, answer-type dispatch for all four types, the preserved quirks, the two bugs above, the leaderboard aggregation formulas, answer extraction from interception.json including a string-encoded body, run discovery for both batch and single-run layouts, all three gold-file shapes, the CLI including the default that leaves run-meta.json alone, and the refusal to pool several models.
  • tests/test_assistantbench_adapter.py, 28 cases: field mapping, the interception contract pinned against a bundled claw-eval task, gold-answer containment, generated tasks validated against test-cases/task.schema.json and through validate_task_data, case-name stability and collision, malformed-export errors that name the line, filters, suite layout, and an adapt-then-score round trip.
  • Full suite: 396 passed, 10 skipped. The one failure on my machine is test_host_tasks.py::...[v1-lite], which is the v1-lite symlinks arriving as plain text files on a Windows checkout. It fails the same way on main and is unrelated to this change; the same checkout artifact is why local ruff reports test-cases/v1-lite/.../solution_code.py.
  • ruff check, ruff format --check, and pyright clean on the new files. uv build and twine check pass.
  • End-to-end smoke with the real CLIs over a synthetic three-row export: adapt produces the suite and the gold sidecar, a reordered record-list answer scores 1.0, a run that never submitted is counted as unanswered rather than wrong so precision stays at 100.0 while accuracy is 66.7, the held-out row with no gold is skipped with a warning, and --write-run-meta updates three files.

Not done here, and the doc says so

Corpus

  • v2
  • v1
  • both
  • not applicable (adds a new generated suite; v1 and v2 are untouched)

Related issues

Advances #188. Related: #72, and #356 which reaches the same answer-submit conclusion for WebVoyager.

Adds two commands:

- clawbench-assistantbench-adapt converts an AssistantBench export into a
  ClawBench task suite over the existing answer-submit interception path, the
  one the bundled claw-eval port already uses: the instruction points the agent
  at the runtime server's /submit form, eval_schema targets /api/task-submit,
  and judge_context carries the gold answer and a rubric. No runner changes.
- clawbench-assistantbench-score reports AssistantBench's own deterministic
  answer metric over a finished batch, as a third stage next to interception
  and the LLM judge: accuracy, answer rate, precision, EM, and accuracy by
  difficulty, in the shape of the upstream leaderboard row.

The metric is a stdlib re-implementation of the upstream evaluator, since that
one needs numpy and scipy. scipy's linear_sum_assignment is replaced by an
exact rectangular Jonker-Volgenant solver checked against brute force. Every
deviation from upstream is enumerated in the module docstring.

Gold answers never enter the container: run.py mounts only the eval_schema
block, so metadata and judge_context stay host-side. A test asserts the gold
answer appears in no agent-facing field.

Advances TIGER-AI-Lab#188.
…dels

Self-review of the scorer against how clawbench-analyze already aggregates.

- discover_runs now uses rglob for run-meta.json and data/interception.json,
  the same discovery analyze.py does, instead of three fixed-depth globs.
- The summary reports runs and distinct tasks separately. Every metric
  averages over runs, as analyze.py's do, so calling the run count "tasks"
  was wrong whenever a case was run more than once. The report says so when
  the two counts differ.
- Pointing at test-output/ rather than test-output/<model> silently averaged
  models into one leaderboard row. That is now an error naming the models it
  found, with --allow-mixed-models to override.
@nahatav

nahatav commented Sep 23, 2026

Copy link
Copy Markdown
Author

Pushed 6db63ee after re-reading the scorer against how clawbench-analyze already aggregates. Three things were wrong and are now fixed, all inside the new scorer:

  • discover_runs used three fixed-depth globs. It now uses rglob over run-meta.json and data/interception.json, which is the discovery analyze.py does, so nesting depth stops mattering.
  • The summary called the run count "tasks". Every metric averages over runs, the way analyze.py's n_runs does, so that label was wrong whenever a case was run more than once. It now reports runs and distinct tasks separately and says so in the report when the two differ.
  • Worse: pointing at test-output/ instead of test-output/<model> silently averaged several models into one leaderboard row. That is now an error naming the models it found, with --allow-mixed-models for anyone who genuinely wants them pooled.

Four more tests cover these. Suite is 396 passed, 10 skipped; the description is updated.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant