Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,11 @@ can select it.

## The active bundle

Bundle 31: live prompt 23, report prompt 18, rubric 1, report schema 2.
Bundle 32: live prompt 23, report prompt 19, rubric 1, report schema 3.

| Bundle | Introduced |
|---|---|
| 32 | The report judges the problem's follow-ups. When the coding round completed and the interviewer was handed them, the report brief lists them numbered, and the report model returns a `followUps` entry for each: whether an interviewer line posed it, and for one that was raised, a short assessment of the answer addressed to the candidate. An entry whose `index` is missing, out of range or repeated, or whose `raised` is not a boolean, and every entry when the follow-ups were never handed over, is dropped before anything in it is checked rather than refusing the report; an assessment that is missing, blank, not a string or on an entry not raised becomes null. A kept assessment is held to a length limit and the same delivery-judgment and published-name scans as the other narrative fields, with errors named by the entry's position in the response; one still judging delivery once the repairs run out is dropped as a refused self-review check is. The server stamps each `debrief.followUps` entry as `{text, raised, assessment}`: `raised` is false when the follow-ups were never handed over, and null when they were but no report judged them, or when the report read `false` from a transcript whose opening was cut. Report prompt 19 and report schema 3; schema 2 reports, whose follow-ups are bare text, stay scored and read as unjudged. The live prompt and rubric are unchanged. |
| 31 | A candidate may disable code execution for the entire editor interview. Whiteboard sessions ignore this preference because they have no code runner. The signed participant setting persists across languages and reconnects, execution events cannot earn credit, and Test may be recorded from a hand trace of written code with source `candidate_speech`. Live prompts ask for traces rather than Run, preserve completed evidence, and distinguish the choice from a runner outage. Assessment metadata and the server-stamped optional `codeExecution: false` report field record the restriction without inventing executed or passing cases; browser reports, saved history and Markdown retain it. The field is additive under schema 2 and older renderers may ignore it. Older reports without the field retain their existing meaning. |
| 30 | Qualifying browser-reported executions record Test at receipt time and publish the checklist without waiting for a model tool call, including failing runs and runs during a thinking hold. Test is also recorded when an earlier credited run covers the editor again, as when the candidate switches back to the language that ran, and the interviewer is then told so without being asked to reply, along with the follow-ups when that completed the coding round. A delayed model call returns the existing execution row without duplicating it or changing its timestamp, and its reply says nothing was added. Setup errors, empty runs, starter code and runs that no longer cover the editor grant no new Test row, and the counts of a run that earned no credit never choose the next step. The checklist is republished until a send succeeds and again when the candidate rejoins, and a reminder of unrecorded earlier steps counts as given only once the reply carrying it is delivered. A phase judge, a side call on the report model a few seconds after each candidate turn, a turn the recognizer grew, or an editor change while a step is open, records Repeat, Example and Algorithm until Coding is recorded, and Coding and Optimizations while there is code, when the candidate's own words, or for Coding what they typed, complete them. The server accepts a step only while it is open and only with a quote found in one recognized candidate turn (adjacent candidate lines the recognizer split count as one), six words for the spoken steps and four for Optimizations, or for Coding in the editor and not in the starter; an approach or analysis the interviewer said before the candidate did (the interviewer saying it back afterwards does not count against it), a paraphrase, an unrecognized turn, or a judgment about code since rewritten records nothing, and claims or requests to mark a step do not count. The row carries the verified quote, the interviewer is told what was recorded without being asked to reply, along with the follow-ups when the step completed the coding round, and the end waits up to 5 seconds for the judge so the report sees the candidate's last answer, while a round transition is held for up to 5 seconds without stopping the room, which keeps forwarding the candidate's audio, so the round decision sees it too. The judge reads the exercise's contract and constraints, a call that fails leaves its window to be read again, a verified quote stands even when a later candidate turn was not recognized, and a held transition also waits for a judgment of anything said after the last one's snapshot, and stays held while the interview is paused. A judgment asked for by the editor alone waits 30 seconds after the last one and takes at most 12 of the 48 calls. At a whiteboard the judge is told there is no editor and the board is not shown, judges Optimizations from speech once the board has work on it, and leaves Coding to the interviewer, since its quote must come from typed code; no run records Test there, and the whiteboard instructions carry the same evidence check. A rate limit on a judgment does not cool the report key. Conversational phase instructions require grounded evidence calls before acknowledging completion, allow independent calls in a batch, and forbid filling earlier phases from later progress. Editor, board and evidence replies carry the recorded phase ids used by the checklist, including on refusals; the interviewer checks that state before claiming a step is marked, and records a missing step only from earlier turns, never because the candidate asked. Live prompt 22, with the phase judge's instruction and prompt in its golden; the report prompt, rubric and schema are unchanged. |
| 29 | Whiteboard interviews: the live prompt is written for the surface the candidate works on, so a whiteboard session is told it has no editor and no test runner, is given the six steps as drawn work ending in the complexity of the approach on the board, is offered `read_board` in place of `read_editor` and a `log_hint` whose clue arrives with the board again, asks a candidate whose speech stays unclear to write it on the board rather than as a code comment, and reads an unclear mark on the board by what the candidate said while drawing it, asking what it stands for rather than guessing; `board_snapshot` joins the evidence sources and is the only one besides candidate speech a whiteboard session may record as observed (`session_timing` stays available to both surfaces for skipped STAR steps), while an editor session may not record it at all; the phases about written work are gated on strokes on the board rather than on characters in the editor. The report prompt follows the same surface: a whiteboard review is sent labeled images of every completed REACTO phase plus a changed final board, so clearing the live surface does not erase earlier evidence; it is told that nothing ran and that the Coding, Test and Optimizations phases were a hand trace, the cases named against the drawing, and the complexity they confirmed. Its interim notes are told the interview has no editor rather than shown an empty one. Its system instruction is its own, so the scoring rules that outrank the brief define both scores at the board, cap an empty board rather than an empty editor, and cite the board where the other cites the code and the test account. The rubric and the report schema are unchanged, so a report from either surface is scored the same way and against the same ten phases. |
Expand Down
9 changes: 5 additions & 4 deletions docs/observable-delivery-policy.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,10 +47,11 @@ the shared five-call report budget. The repair prompt bounds the number and
length of errors it carries: an error names as many phrases as fit and counts
the rest, and every failing path gets one error before any path gets a second.
When the repairs run out, the latest response that becomes a valid report
once its prohibited self-review checks are dropped and its prohibited success
criteria replaced is kept, even if a later response broke something else: a
list left empty gets one fixed, neutral check, and a criterion is replaced with
fixed text that judges no one. Without such a response, a claim that remains
once its prohibited self-review checks and follow-up assessments are dropped
and its prohibited success criteria replaced is kept, even if a later response
broke something else: a list left empty gets one fixed, neutral check, a
criterion is replaced with fixed text that judges no one, and a follow-up keeps
whether it was raised without an assessment. Without such a response, a claim that remains
after the last repair leaves the report incomplete, with no score or verdict,
and the candidate may request the one regeneration that
[provider cost and degradation](provider-cost-and-degradation.md) describes.
Expand Down
15 changes: 11 additions & 4 deletions src/agent.rs
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ pub use prompts::{
pub(crate) use prompts::{editor_tool_continuity, end_interview_refusal, report_transcript_lines};
pub(crate) use report::{MAX_ERROR_CHARS, Sanitized, sanitize_report_candidate};
pub use report::{
MAX_SUMMARY_TEXT, fallback_report, final_report, names_published_problem,
MAX_SUMMARY_TEXT, ReportRounds, fallback_report, final_report, names_published_problem,
report_response_schema, spelled_words, validate_report, validate_report_candidate,
validate_report_for_round,
};
Expand Down Expand Up @@ -171,11 +171,11 @@ pub const THINKING_CHECK_IN_S: u64 = 120;
pub(crate) const THINKING_RELEASE_COOLDOWN: std::time::Duration =
std::time::Duration::from_secs(10);

pub const INTERVIEW_CONTRACT_BUNDLE_VERSION: u32 = 31;
pub const INTERVIEW_CONTRACT_BUNDLE_VERSION: u32 = 32;
pub const LIVE_PROMPT_VERSION: u32 = 23;
pub const REPORT_PROMPT_VERSION: u32 = 18;
pub const REPORT_PROMPT_VERSION: u32 = 19;
pub const RUBRIC_VERSION: u32 = 1;
pub const REPORT_SCHEMA_VERSION: u32 = 2;
pub const REPORT_SCHEMA_VERSION: u32 = 3;

pub fn interview_contract_json() -> serde_json::Value {
serde_json::json!({
Expand Down Expand Up @@ -2695,6 +2695,13 @@ pub fn transcript_for_report(lines: &[String]) -> String {
transcript_tail(&mark_unrecognized_turns(lines), MAX_TRANSCRIPT_BYTES)
}

/// Whether `transcript_for_report` leaves out the opening of these lines.
/// Asked of the lines rather than of the text it returns, which carries the
/// candidate's own words and so can spell the omission notice itself.
pub(crate) fn report_transcript_cut(lines: &[String]) -> bool {
tail_start(&mark_unrecognized_turns(lines), MAX_TRANSCRIPT_BYTES) > 0
}

/// What an assessment reads in place of a candidate turn the recognizer did
/// not return as English. The report prompt names it, so a change here that
/// left the prompt describing the old marker fails
Expand Down
26 changes: 25 additions & 1 deletion src/agent/prompts.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2317,6 +2317,11 @@ pub struct ReportPromptInput<'a> {
/// question without one, and the transcript then holds an answer the round
/// status says never happened.
pub behavioral_round: BehavioralRound,
/// Whether the interviewer was handed the problem's follow-ups, which
/// happens when the coding round completes. A follow-up it never held
/// cannot have been raised, so the brief lists none and the debrief says
/// none was reached without asking the model.
pub follow_ups_released: bool,
}

/// Where the behavioral round stood when the interview ended, in the three
Expand Down Expand Up @@ -2497,6 +2502,7 @@ fn report_brief(input: &ReportPromptInput<'_>) -> String {
// the contract is measured by, what the candidate produced, and what
// account there is of it running. The editor's wording is what it always
// was; a whiteboard's is its own paragraphs, above.
let follow_ups = report_follow_ups(input);
let whiteboard = input.interview_mode.is_whiteboard();
let graded_by = if whiteboard {
"Contract a correct answer meets"
Expand Down Expand Up @@ -2548,6 +2554,8 @@ HINTS THE INTERVIEWER GAVE: {} total; the candidate reached hint rung {} of 3.
the interviewer helped, but weaker evidence than a requested hint that the
candidate depended on; treat both as context, never as a numeric deduction.

{follow_ups}

{execution}

{behavioral_round}
Expand All @@ -2569,6 +2577,21 @@ candidate depended on; treat both as context, never as a numeric deduction.
)
}

/// The follow-ups the interviewer was handed, numbered as `followUps` cites
/// them, or the instruction to judge none. They are listed on the condition
/// the live prompt hands them over, so the reviewer is never asked whether a
/// follow-up the interviewer did not hold was raised.
fn report_follow_ups(input: &ReportPromptInput<'_>) -> String {
let follow_ups = input.problem.variant().follow_ups;
if !input.follow_ups_released || follow_ups.is_empty() {
return "FOLLOW-UPS: None were released to the interviewer in this interview. Return `followUps` as an empty array.".to_string();
}
format!(
"FOLLOW-UPS: Once the coding round completed, the interviewer was given these, to raise at most two of them in order:\n{}\nReturn one `followUps` entry for each, with its number as `index`. Set `raised` to true only when an Interviewer line in the transcript poses that follow-up, in any wording; the candidate bringing up the same idea unprompted does not raise it. For a raised follow-up, `assessment` is one to three sentences written to the candidate as \"you\": what the answer covered and what a stronger answer would have added, citing what was said. Interviewer agreement does not show the answer was correct. Use null for `assessment` when `raised` is false.",
Comment thread
matthiola0 marked this conversation as resolved.
numbered_list(follow_ups)
)
}

/// Only the platform's transition opens the round, and the round status the
/// report card shows reads the same flag, so an answer to a question asked
/// without it would sit beside a round marked skipped or not configured. The
Expand Down Expand Up @@ -2698,7 +2721,8 @@ impact (high, medium or low), frequency (a positive count of observations in
this session), drill, durationMin (1-30), successCriterion and selfReview; and
frameworkAssessment with rubricVersion {rubric_version} and one phase entry each
for Repeat, Example, Algorithm, Coding, Test, Optimizations, Situation, Task,
Action and Result in that order, each with a score (integer 0-100 or null).
Action and Result in that order, each with a score (integer 0-100 or null); and
followUps, one entry per follow-up the brief lists, as it describes.
Each strengths/improvements list must contain 2 to 4 concrete, specific items
grounded in the rolling assessment, the transcript, and {material}, never generic
filler, and no item may repeat another in the same list. A session with little to praise still holds two
Expand Down
Loading