Skip to content

fix(schema): declare judge_context.gold_answer so the answer judge can read it - #381

Open
nahatav wants to merge 1 commit into
TIGER-AI-Lab:mainfrom
nahatav:fix/judge-context-gold-answer
Open

nahatav wants to merge 1 commit into
TIGER-AI-Lab:mainfrom
nahatav:fix/judge-context-gold-answer

Conversation

@nahatav

@nahatav nahatav commented Sep 23, 2026

Copy link
Copy Markdown

What does this PR do?

Fixes #380: judge_context.gold_answer is read by the answer judge and documented in docs/answer-mode-tasks.md, but task.schema.json declares judge_context with additionalProperties: false and only three properties, so no schema-valid task can set it. The branch is unreachable, and validate-task would reject any task that tried to use it.

Four lines of schema, one extracted constant, and a test that holds the two lists together so they cannot drift again.

Why it drifted

2e5ccbf added judge_context to the schema with three properties. 68da27a added judge_answer reading a fourth key, plus the doc describing it, without extending the schema. Nothing compared the two lists, and tests/test_answer_judge.py::test_build_answer_msg_includes_answer_and_rubric already passes {"gold_answer": "X"} and passes, so the suite reads as though the field works.

Why declare it rather than delete it

The judge labels each context value by its key (f"{key}:\n{value}"), so the key name is the only signal the judge gets about what it is looking at. Across the 19 bundled claw-eval tasks reference_solution holds a worked procedure in some (ce-T046: "1. web_search(...) → CVE details") and the answer itself in others (ce-T062: "2-year CAGR = ... ≈ 22.6%"). A judge reading reference_solution cannot tell which it has. gold_answer says so, and it is the natural target for the answer-submitting imports in #188, #189 and #190.

The change is purely additive. No existing task changes, no bundled task uses the key today (grep -rl gold_answer test-cases/ is empty), and additionalProperties: false still rejects everything else.

If the intent was the opposite and gold_answer should come out of judge.py and the doc instead, say so and I will flip the PR; the anti-drift test is worth keeping either way.

The anti-drift test

ANSWER_CONTEXT_KEYS is now a module constant in judge.py rather than a literal inside _build_answer_msg, so a test can compare it to the schema. Two assertions, running in both directions:

  • every key a judge reads must be a declared property of judge_context;
  • every declared property must be read by a judge, so the schema never promises a field nothing consumes.

Both of the new assertions that cover the bug fail against the schema as it stands on main, which is the check that matters:

$ git stash push test-cases/task.schema.json   # revert the schema to main
$ uv run --frozen pytest tests/test_answer_judge.py -q
FAILED tests/test_answer_judge.py::test_schema_declares_every_key_the_answer_judge_reads
FAILED tests/test_answer_judge.py::test_a_task_carrying_a_gold_answer_validates
2 failed, 9 passed

With the schema fix: 11 passed.

Test plan

  • Four new cases in tests/test_answer_judge.py: the forward lock, the reverse lock, a task carrying gold_answer validating against the bundled schema, and the value reaching the judge labelled by its key.
  • Verified the two new lock assertions fail on the unpatched schema and pass with it, as shown above.
  • Full suite: 315 passed, 10 skipped. The one failure on my machine is test_host_tasks.py::...[v1-lite], the v1-lite symlinks arriving as plain text files on a Windows checkout; it fails the same way on main and is unrelated.
  • ruff check, ruff format --check and pyright clean on both changed files.
  • Because task.schema.json is in the diff, validate-task re-validates every bundled task against the new schema rather than just the changed ones. tests/test_host_tasks.py does the same locally and passes for v1, v2 and claw-eval.

Corpus

  • v2
  • v1
  • both (schema is shared; no task file changes)
  • not applicable

Related issues

Fixes #380. Unblocks the gold_answer mapping for the answer-submitting imports in #188, #189 and #190; #379 carries the AssistantBench one and can adopt the field once this lands.

…n read it

_build_answer_msg reads judge_context.gold_answer and
docs/answer-mode-tasks.md documents it, but judge_context in
test-cases/task.schema.json declares only rubric, reference_solution and
source_task_yaml under additionalProperties: false. No schema-valid task can
set gold_answer, so the branch is unreachable and validate-task would reject
any task that tried.

The two drifted in 68da27a, which added the answer judge and the fourth key
without extending the schema that 2e5ccbf had introduced.

Declares the property and adds tests that hold the schema and the judges
together in both directions: every key a judge reads must be declared, and
every declared key must be read by a judge. Both new assertions fail against
the schema as it stands on main.

The field earns its place next to reference_solution because the judge labels
each context value by its key. Across the 19 bundled claw-eval tasks
reference_solution holds a worked procedure in some and the answer itself in
others, so a judge reading it cannot tell which it has; gold_answer says so.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

judge_context.gold_answer is read by the answer judge and documented, but task.schema.json rejects it

1 participant