Repository navigation
Judge each follow-up in the report - #263
Open
matthiola0 wants to merge 1 commit into
Open
matthiola0 wants to merge 1 commit into
matthiola0 wants to merge 1 commit into
Conversation
jserv
reviewed
Oct 9, 2026
jserv
reviewed
Oct 9, 2026
jserv
reviewed
Oct 9, 2026
jserv
reviewed
Oct 9, 2026
jserv
reviewed
Oct 9, 2026
matthiola0
force-pushed
the
report-follow-ups
branch
from
October 9, 2026 08:17
d00e28d to
623e948
Compare
matthiola0
force-pushed
the
report-follow-ups
branch
from
October 9, 2026 08:39
623e948 to
2b638a2
Compare
The debrief listed a problem's follow-ups but never said which ones the interviewer raised or how they were answered. Once the coding round completes, the report brief lists the released follow-ups and the reviewer judges each one. A follow-up never handed over is stamped as not reached, and one that no report judged, or that was read as unasked from a cut transcript, as unknown. A malformed judgment costs only its own entry. Bundle 31 moves the report prompt to 18 and the report schema to 3; schema 2 reports stay scored. Closes sysprog21#248
matthiola0
force-pushed
the
report-follow-ups
branch
from
October 9, 2026 09:00
2b638a2 to
d438354
Compare
Contributor
Author
|
@jserv, CI failed in the editor cache test with I reproduced the same issue on Would you prefer a separate PR for this fix, or should I include it here? |
Contributor
Create another pull request. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The report's debrief listed every follow-up a problem offers, but never said which ones Jim raised or how the candidate answered them, and the report model was never shown the list. This makes the reviewer judge each follow-up the interviewer was actually handed.
When the coding round completed, the report brief now lists the released follow-ups, numbered, and the report model returns a
followUpsentry for each: whether an Interviewer line posed it, and for a raised one, one to three sentences on what the answer covered and what a stronger answer would have added. The server merges that intodebrief.followUpsas{text, raised, assessment}, the way hints carrygiven. The card and the Markdown export label each follow-up Raised or Not reached and show the assessment under a raised one.The server, not the model, settles what the model cannot know. A follow-up never handed to the interviewer is
raised: false, and nothing the model wrote about it is checked. One that was handed over but not judged isnulland shows no label: a lost report, a missing entry, or afalseread from a transcript whose openingtranscript_for_reportcut. An entry whoseindexis missing, out of range or repeated, or whoseraisedis not a boolean is dropped instead of refusing the report, and an assessment that is missing, blank, not a string or on an entry not raised becomesnull; a kept assessment meets the same length limit and scans as the other narrative fields, with errors at its position in the model's response. Once the repairs run out, an assessment still judging delivery is dropped the way a refused self-review check is, anddocs/observable-delivery-policy.mdsays so.This is bundle 31: report prompt 18 and report schema 3, with a row in
docs/interview-contract-versions.md. The live prompt and rubric are unchanged. Schema 2 reports stay scored, and their bare-text follow-ups read as unjudged.Testing: each behavior above has a test that fails when the line it covers is broken, and the prompt and report-schema goldens are regenerated. On Windows,
cargo testpasses and the formatters and linters are clean on the changed files; the browser tests that fail on this checkout fail the same way on an untouchedmain.Closes #248