Repository navigation
Conversation
Adds two commands: - clawbench-assistantbench-adapt converts an AssistantBench export into a ClawBench task suite over the existing answer-submit interception path, the one the bundled claw-eval port already uses: the instruction points the agent at the runtime server's /submit form, eval_schema targets /api/task-submit, and judge_context carries the gold answer and a rubric. No runner changes. - clawbench-assistantbench-score reports AssistantBench's own deterministic answer metric over a finished batch, as a third stage next to interception and the LLM judge: accuracy, answer rate, precision, EM, and accuracy by difficulty, in the shape of the upstream leaderboard row. The metric is a stdlib re-implementation of the upstream evaluator, since that one needs numpy and scipy. scipy's linear_sum_assignment is replaced by an exact rectangular Jonker-Volgenant solver checked against brute force. Every deviation from upstream is enumerated in the module docstring. Gold answers never enter the container: run.py mounts only the eval_schema block, so metadata and judge_context stay host-side. A test asserts the gold answer appears in no agent-facing field. Advances TIGER-AI-Lab#188.
…dels Self-review of the scorer against how clawbench-analyze already aggregates. - discover_runs now uses rglob for run-meta.json and data/interception.json, the same discovery analyze.py does, instead of three fixed-depth globs. - The summary reports runs and distinct tasks separately. Every metric averages over runs, as analyze.py's do, so calling the run count "tasks" was wrong whenever a case was run more than once. The report says so when the two counts differ. - Pointing at test-output/ rather than test-output/<model> silently averaged models into one leaderboard row. That is now an error naming the models it found, with --allow-mixed-models to override.
Author
|
Pushed 6db63ee after re-reading the scorer against how
Four more tests cover these. Suite is |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Advances #188 (and #72): ClawBench can now run AssistantBench tasks and report AssistantBench's own answer metric next to the existing two stages.
clawbench-assistantbench-adapt --input ./assistantbench/validation.jsonl \ --output-dir test-cases/assistantbench clawbench-batch --models <m> --cases-dir test-cases/assistantbench --all-cases clawbench-assistantbench-score test-output/<m> \ --gold test-cases/assistantbench/assistantbench-gold.jsonTwo commands, both new, nothing existing changed beyond a CHANGELOG line, a CLI table row, and two entries in
HELP_MODULES.The design decision worth reviewing
AssistantBench grades a free-text answer. There is no final write request for Stage-1 to intercept, which is the one thing ClawBench's scoring keys on.
The easy answer was to introduce an
eval_schema.mode: "judge"discriminator and the runner branch thatdocs/answer-mode-tasks.mdsketches. I did not, becausemainalready solves this shape and ships tasks that use it: theclaw-evalport tells the agent to submit its answer at the runtime server'shttp://127.0.0.1:7878/submitform, that formPOSTs to/api/task-submit, the task'seval_schematargets that endpoint, and the judge scores the submitted answer againstjudge_context. This adapter reuses that path verbatim, same endpoint and same footer wording, so an AssistantBench task runs through the ordinary pipeline with zero runner changes and produces the standard five-layer bundle. A test pinseval_schemaand the footer against a bundledclaw-evaltask so the two ports cannot drift apart.This is the same decision #356 made for WebVoyager, reached independently. I built on
mainrather than stacking on #345 so this reviews and merges on its own; the loader is a pure function and drops intosrc/clawbench/adapters/unchanged if that lands.Three scores, not two
run-meta.interceptedjudge.jsonclawbench-assistantbench-scoreStage 3 has no model in the loop, which is the point of adding it: it is an independent, deterministic check on the LLM judge over the same runs.
--write-run-metarecords it asassistantbench_answer_scorein each run'srun-meta.json, which is what #188 asks for. It is off by default because it edits run output you already have.Mapping
metadata.source_task_idid, verbatim, the real keymetadata.task_id600000 + n, n over ids in sorted ordermetadata.classmetadatametadata.sites_involvedgold_urlinstructiontask+ answer-format block + the claw-eval submit footereval_schemaPOST /api/task-submittime_limit--time-limit, default 30, since upstream is long-horizonjudge_context.reference_solutionanswer+explanation+gold_urljudge_context.rubricCase names are keyed on the source id, not the row's position, so re-exporting the dataset in a different order does not rename every case. A name collision is an error, never a silent overwrite.
Gold answers never enter the container.
run.pymounts only theeval_schemablock as/eval-schema.json, sometadataandjudge_contextstay host-side where the judge and the scorer read them. A test asserts the gold answer, its values and its source hostnames appear in no agent-facing field.The metric, and how I know it is right
assistantbench_score.pyre-implements theevaluation/package of the AssistantBench leaderboard Space (Apache-2.0, string metric from DROP'sdrop_eval). It is a re-implementation rather than a vendored copy because upstream needsnumpyandscipyand ClawBench takes neither.scipy.optimize.linear_sum_assignmentis replaced by an exact rectangular Jonker-Volgenant solver; the tests check it against brute force over random matrices in both orientations, and one test pins a case where greedy would lose.Re-implementing a metric is exactly where a port quietly drifts, so I checked it rather than asserting it. I stood the upstream evaluator up under
numpy2.5.3 andscipy1.18.1 and ran both scorers over 7249(prediction, gold)pairs: the 3249-pair exhaustive product of a corpus covering all four answer types, plus 4000 fuzzed pairs.Exact agreement to 1e-9 on accuracy and answer rate for every pair upstream can score. The 668 it cannot are all malformed predictions where upstream raises
TypeError/AttributeError: a record object answering a numeric question,null,true, a bare list. The port scores those 0.0 rather than taking down a whole batch's scoring.That run found two real bugs in my first draft, both now fixed and both pinned by a test:
numpygets to 0 by way ofnan, sincemax(0, nan)is 0;math.lograises instead, so the sign is checked first.Every remaining deviation is enumerated in the module docstring, as are the upstream quirks kept on purpose:
,read as a decimal point so "1,000" parses as 1.0, the log metric exceeding 1 when both values are negative, precision being recall with the arguments swapped. Those have their own tests saying, in words, not to fix them.Test plan
tests/test_assistantbench_score.py, 55 cases: the assignment solver against brute force, answer-type dispatch for all four types, the preserved quirks, the two bugs above, the leaderboard aggregation formulas, answer extraction frominterception.jsonincluding a string-encoded body, run discovery for both batch and single-run layouts, all three gold-file shapes, the CLI including the default that leavesrun-meta.jsonalone, and the refusal to pool several models.tests/test_assistantbench_adapter.py, 28 cases: field mapping, the interception contract pinned against a bundledclaw-evaltask, gold-answer containment, generated tasks validated againsttest-cases/task.schema.jsonand throughvalidate_task_data, case-name stability and collision, malformed-export errors that name the line, filters, suite layout, and an adapt-then-score round trip.396 passed, 10 skipped. The one failure on my machine istest_host_tasks.py::...[v1-lite], which is thev1-litesymlinks arriving as plain text files on a Windows checkout. It fails the same way onmainand is unrelated to this change; the same checkout artifact is why localruffreportstest-cases/v1-lite/.../solution_code.py.ruff check,ruff format --check, andpyrightclean on the new files.uv buildandtwine checkpass.--write-run-metaupdates three files.Not done here, and the doc says so
±2ppreproduction against upstream that feat: adapter — run AssistantBench (realistic time-consuming live-web tasks) under the ClawBench harness #188 asks for is a follow-up once someone runs one. The metric fidelity that check would test is already established above by the differential run.--corpus assistantbenchas a first-class selector. feat: adapter — run AssistantBench (realistic time-consuming live-web tasks) under the ClawBench harness #188 sketchesclawbench run --corpus assistantbench. This ships as a suite generator plus--cases-dir, matching howharborandedgebenchalready work, and leaves the registry question to Support running other tasks suits under ClawBench's infrastructure using adaptor layers. #72.Corpus
Related issues
Advances #188. Related: #72, and #356 which reaches the same answer-submit conclusion for WebVoyager.