Repository navigation
Conversation
…n read it _build_answer_msg reads judge_context.gold_answer and docs/answer-mode-tasks.md documents it, but judge_context in test-cases/task.schema.json declares only rubric, reference_solution and source_task_yaml under additionalProperties: false. No schema-valid task can set gold_answer, so the branch is unreachable and validate-task would reject any task that tried. The two drifted in 68da27a, which added the answer judge and the fourth key without extending the schema that 2e5ccbf had introduced. Declares the property and adds tests that hold the schema and the judges together in both directions: every key a judge reads must be declared, and every declared key must be read by a judge. Both new assertions fail against the schema as it stands on main. The field earns its place next to reference_solution because the judge labels each context value by its key. Across the 19 bundled claw-eval tasks reference_solution holds a worked procedure in some and the answer itself in others, so a judge reading it cannot tell which it has; gold_answer says so.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes #380:
judge_context.gold_answeris read by the answer judge and documented indocs/answer-mode-tasks.md, buttask.schema.jsondeclaresjudge_contextwithadditionalProperties: falseand only three properties, so no schema-valid task can set it. The branch is unreachable, andvalidate-taskwould reject any task that tried to use it.Four lines of schema, one extracted constant, and a test that holds the two lists together so they cannot drift again.
Why it drifted
2e5ccbfaddedjudge_contextto the schema with three properties.68da27aaddedjudge_answerreading a fourth key, plus the doc describing it, without extending the schema. Nothing compared the two lists, andtests/test_answer_judge.py::test_build_answer_msg_includes_answer_and_rubricalready passes{"gold_answer": "X"}and passes, so the suite reads as though the field works.Why declare it rather than delete it
The judge labels each context value by its key (
f"{key}:\n{value}"), so the key name is the only signal the judge gets about what it is looking at. Across the 19 bundled claw-eval tasksreference_solutionholds a worked procedure in some (ce-T046: "1. web_search(...) → CVE details") and the answer itself in others (ce-T062: "2-year CAGR = ... ≈ 22.6%"). A judge readingreference_solutioncannot tell which it has.gold_answersays so, and it is the natural target for the answer-submitting imports in #188, #189 and #190.The change is purely additive. No existing task changes, no bundled task uses the key today (
grep -rl gold_answer test-cases/is empty), andadditionalProperties: falsestill rejects everything else.If the intent was the opposite and
gold_answershould come out ofjudge.pyand the doc instead, say so and I will flip the PR; the anti-drift test is worth keeping either way.The anti-drift test
ANSWER_CONTEXT_KEYSis now a module constant injudge.pyrather than a literal inside_build_answer_msg, so a test can compare it to the schema. Two assertions, running in both directions:judge_context;Both of the new assertions that cover the bug fail against the schema as it stands on
main, which is the check that matters:With the schema fix:
11 passed.Test plan
tests/test_answer_judge.py: the forward lock, the reverse lock, a task carryinggold_answervalidating against the bundled schema, and the value reaching the judge labelled by its key.315 passed, 10 skipped. The one failure on my machine istest_host_tasks.py::...[v1-lite], thev1-litesymlinks arriving as plain text files on a Windows checkout; it fails the same way onmainand is unrelated.ruff check,ruff format --checkandpyrightclean on both changed files.task.schema.jsonis in the diff,validate-taskre-validates every bundled task against the new schema rather than just the changed ones.tests/test_host_tasks.pydoes the same locally and passes for v1, v2 and claw-eval.Corpus
Related issues
Fixes #380. Unblocks the
gold_answermapping for the answer-submitting imports in #188, #189 and #190; #379 carries the AssistantBench one and can adopt the field once this lands.