Skip to content

Judge each follow-up in the report - #263

Open
matthiola0 wants to merge 1 commit into
sysprog21:mainfrom
matthiola0:report-follow-ups
Open

matthiola0 wants to merge 1 commit into
sysprog21:mainfrom
matthiola0:report-follow-ups

Conversation

@matthiola0

@matthiola0 matthiola0 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

The report's debrief listed every follow-up a problem offers, but never said which ones Jim raised or how the candidate answered them, and the report model was never shown the list. This makes the reviewer judge each follow-up the interviewer was actually handed.

When the coding round completed, the report brief now lists the released follow-ups, numbered, and the report model returns a followUps entry for each: whether an Interviewer line posed it, and for a raised one, one to three sentences on what the answer covered and what a stronger answer would have added. The server merges that into debrief.followUps as {text, raised, assessment}, the way hints carry given. The card and the Markdown export label each follow-up Raised or Not reached and show the assessment under a raised one.

The server, not the model, settles what the model cannot know. A follow-up never handed to the interviewer is raised: false, and nothing the model wrote about it is checked. One that was handed over but not judged is null and shows no label: a lost report, a missing entry, or a false read from a transcript whose opening transcript_for_report cut. An entry whose index is missing, out of range or repeated, or whose raised is not a boolean is dropped instead of refusing the report, and an assessment that is missing, blank, not a string or on an entry not raised becomes null; a kept assessment meets the same length limit and scans as the other narrative fields, with errors at its position in the model's response. Once the repairs run out, an assessment still judging delivery is dropped the way a refused self-review check is, and docs/observable-delivery-policy.md says so.

This is bundle 31: report prompt 18 and report schema 3, with a row in docs/interview-contract-versions.md. The live prompt and rubric are unchanged. Schema 2 reports stay scored, and their bare-text follow-ups read as unjudged.

Testing: each behavior above has a test that fails when the line it covers is broken, and the prompt and report-schema goldens are regenerated. On Windows, cargo test passes and the formatters and linters are clean on the changed files; the browser tests that fail on this checkout fail the same way on an untouched main.

Closes #248

cubic-dev-ai[bot]

This comment was marked as resolved.

Comment thread src/agent/prompts.rs
Comment thread src/agent/report.rs Outdated
Comment thread src/agent/report.rs Outdated
Comment thread src/agent/report.rs Outdated
Comment thread web/render.js Outdated
cubic-dev-ai[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

The debrief listed a problem's follow-ups but never said which ones the
interviewer raised or how they were answered. Once the coding round
completes, the report brief lists the released follow-ups and the
reviewer judges each one. A follow-up never handed over is stamped as
not reached, and one that no report judged, or that was read as unasked
from a cut transcript, as unknown. A malformed judgment costs only its
own entry. Bundle 31 moves the report prompt to 18 and the report schema
to 3; schema 2 reports stay scored.

Closes sysprog21#248
@matthiola0

Copy link
Copy Markdown
Contributor Author

@jserv, CI failed in the editor cache test with 2 !== 1. It seems the repaint from selecting all text can finish after the counter resets, so the test counts an extra paint.

I reproduced the same issue on main with a controlled delay. Waiting for the bracket marks to clear before resetting the counter fixes it locally.

Would you prefer a separate PR for this fix, or should I include it here?

@jserv

jserv commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

I reproduced the same issue on main with a controlled delay. Waiting for the bracket marks to clear before resetting the counter fixes it locally.
Would you prefer a separate PR for this fix, or should I include it here?

Create another pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Show in the report which follow-ups were asked and how they were answered

2 participants