Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 36 additions & 2 deletions .cargo/mutants.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Functions the mutation gate cannot judge, because `cargo test` cannot reach
# them. Matched against the mutant names that `cargo mutants --list` prints.
#
# EXCLUSIONS: 69
# EXCLUSIONS: 86
#
# That number is checked by `scripts/test.sh`, so adding an entry means editing
# this line too. The point is not the count, it is that the list only ever grows
Expand Down Expand Up @@ -109,7 +109,9 @@
# `publish_framework_progress`, `send_wrap_up_and_wait`, `drop_stale_playout`,
# `close_turns`, `set_agent_state` and `publish_transcript` publish;
# `generate_report` and the two transport functions under it post to Google
# and would need an API key to reach a single line. The last two joined this
# and would need an API key to reach a single line. `generate_task_feedback` is
# task mode's and only names Google's URL; `generate_task_feedback_at`, which
# takes the URL, is tested against a local server. The last two joined this
# list when they moved into livekit/session.rs: never in a diff before, they
# had never been scored, and the move made every line of them new. What they
# decide before they write is tested on its own: `close_turn`,
Expand Down Expand Up @@ -299,6 +301,21 @@
# builds the Google URL for `generate_phase_judgment_at`, which is tested
# against a local server; reaching the real one needs a live key.
#
# The task room in src/livekit/tasks.rs is `run_room`'s counterpart for task
# mode, and its I/O is excluded for the same reason: `run` joins a LiveKit room,
# and the `TaskRoom` handlers listed below (`on_data`, `on_room_event`,
# `on_tick`, `recover`, `ask_wrap_up`, `on_audio_frame`, `publish_state`,
# `settle_page_loss`, `serve`, `send_or_queue`, `deliver_review`, `close`) each
# take the room or the Gemini socket and answer by publishing or sending, so
# none of them runs without a LiveKit server. `publish`, `publish_error`,
# `silence_output` and `LiveUsage::drain` are the same, one layer down. Its
# `on_gemini_event` and `on_turn_complete` were already matched by the entries
# above. What these handlers decide is not excluded: it was moved out of the
# room so `cargo test` reaches it, into `TaskSession::receive` and
# `TaskSession::answer_tool` (every page message and tool call),
# `TaskSession::page_left` and `page_returned` (the page leaving and joining),
# `page_loss`, `wrap_up_due` and `speaking`, each tested in tests/unit.
#
# Keep this list short and each entry justified. An entry that is really "we
# never got around to testing this" belongs in a test, not here.
exclude_re = [
Expand Down Expand Up @@ -338,6 +355,7 @@ exclude_re = [
"set_agent_state",
"publish_transcript",
"generate_report",
"replace generate_task_feedback -> ",
"report_packet",
"publish_with_recovery",
"LiveRecoveryRoom",
Expand Down Expand Up @@ -371,4 +389,20 @@ exclude_re = [
"settle_phase_judge",
"flush_framework_progress",
"generate_phase_judgment_with_keys",
"TaskRoom<'_>::on_data",
"TaskRoom<'_>::on_room_event",
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
"TaskRoom<'_>::on_tick",
"TaskRoom<'_>::recover",
"TaskRoom<'_>::ask_wrap_up",
"TaskRoom<'_>::on_audio_frame",
"TaskRoom<'_>::publish_state",
"TaskRoom<'_>::settle_page_loss",
"TaskRoom<'_>::serve",
"TaskRoom<'_>::send_or_queue",
"TaskRoom<'_>::deliver_review",
"TaskRoom<'_>::close",
"^src/livekit/tasks\\.rs:\\d+:\\d+: (replace run |.* in run$)",
"^src/livekit/tasks\\.rs:\\d+:\\d+: (replace publish(_error)? |.* in publish(_error)?$)",
"^src/livekit/tasks\\.rs:\\d+:\\d+: replace silence_output ",
"LiveUsage::drain",
]
30 changes: 25 additions & 5 deletions .github/workflows/check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -281,6 +281,14 @@ jobs:
cargo-audit --version 2>/dev/null | grep -qF " $CARGO_AUDIT_VERSION" \
|| cargo install cargo-audit --locked --version "$CARGO_AUDIT_VERSION" --force

- name: Install instructor packaging dependencies
run: |
# Outside target/, which is cached: a restored venv would keep a
# package the requirements file has since dropped.
python3 -m venv "$RUNNER_TEMP/task-tools"
"$RUNNER_TEMP/task-tools/bin/pip" install -r scripts/requirements-task.txt
echo "PYTHON=$RUNNER_TEMP/task-tools/bin/python" >> "$GITHUB_ENV"

# The Python here generates the problem bank and drives two integration
# harnesses, and `ruff check` is its only linter. The gate skips when the
# binary is missing, which is right on a laptop and wrong here. pipx is on
Expand Down Expand Up @@ -369,9 +377,9 @@ jobs:
# one runner testing eight. `--list` parses the tree without building it, one
# second here, so the count is cheap to ask for and the width can follow it.
#
# One shard per 25 mutants, capped at four. The cap is not about arithmetic:
# past that the setup cost of another runner stops paying for itself against
# the mutants it would take.
# One shard per 25 mutants, under the ceiling the step explains. The ceiling
# is not about arithmetic: past it the setup cost of another runner stops
# paying for itself against the mutants it would take.
#
# A pull request from a fork's main is refused in `check`, and this job does
# not wait on that one, so it repeats the condition rather than plan and fan
Expand Down Expand Up @@ -433,8 +441,12 @@ jobs:
# shard is what the line above asks for and what fits the job's
# timeout; a ceiling of four turned a 764-mutant diff into 191 a
# shard, which is seven times the design load and forty-five minutes
# of work reported as nothing at all.
[ "$width" -gt 16 ] && width=16
# of work reported as nothing at all. Sixteen did the same to a diff
# that adds a whole subsystem: every shard ran out of its budget
# before reporting. Forty-eight keeps the design load up to twelve
# hundred mutants; shards beyond the account's concurrent jobs queue
# rather than fail.
[ "$width" -gt 48 ] && width=48
{
echo "count=$count"
echo "width=$width"
Expand Down Expand Up @@ -687,8 +699,13 @@ jobs:
- name: Verify the assets about to be embedded
run: ./scripts/verify-vendor.sh
shell: bash
# The device sign-in task mode uses needs a GitHub OAuth App's client
# ID compiled in. It is public, not a secret; a repository without the
# variable builds a release with no device sign-in, which says so.
- name: Build the self-contained executable
if: runner.os != 'Linux'
env:
CODETRIAL_GITHUB_CLIENT_ID: ${{ vars.CODETRIAL_GITHUB_CLIENT_ID }}
run: cargo build --release --locked --target ${{ matrix.target }}
# GitHub retired the Ubuntu 20.04 runner, but its glibc 2.31 remains the
# useful Linux release baseline. Build inside Bullseye, which carries the
Expand Down Expand Up @@ -724,6 +741,8 @@ jobs:
- name: Build Linux executable against the compatibility baseline
if: runner.os == 'Linux'
shell: bash
env:
CODETRIAL_GITHUB_CLIENT_ID: ${{ vars.CODETRIAL_GITHUB_CLIENT_ID }}
run: |
docker run --rm \
--user "$(id -u):$(id -g)" \
Expand All @@ -734,6 +753,7 @@ jobs:
--env CARGO_HOME=/cargo \
--env CXX=/clang/bin/clang++ \
--env GITHUB_SHA \
--env CODETRIAL_GITHUB_CLIENT_ID \
${{ matrix.image }} \
cargo build --release --locked --target ${{ matrix.target }}
# A successful link does not prove the Objective-C categories survived.
Expand Down
4 changes: 2 additions & 2 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ unsafe_code = "forbid"
all = "deny"

[dependencies]
chrono = { version = "0.4.45", default-features = false, features = ["std"] }
axum = { version = "0.8.9", default-features = false, features = ["http1", "json", "original-uri", "query", "tokio"] }
base64 = "0.23.1"
futures-util = { version = "0.3.34", default-features = false, features = ["sink", "std"] }
Expand All @@ -38,7 +39,7 @@ rusqlite = { version = "0.40.2", default-features = false, features = ["bundled"
rust-embed = { version = "8.12.0", features = ["include-exclude"] }
serde_json = "1.0.151"
serde = { version = "1.0.229", features = ["derive"] }
tokio = { version = "1.53.1", features = ["fs", "macros", "net", "rt-multi-thread", "sync", "time"] }
tokio = { version = "1.53.1", features = ["fs", "macros", "net", "rt-multi-thread", "sync", "time", "signal"] }
tokio-tungstenite = { version = "0.29.0", default-features = false, features = ["connect", "rustls-tls-webpki-roots"] }
tokio-util = "0.7.19"
tower-http = { version = "0.6.11", default-features = false, features = ["compression-gzip"] }
Expand All @@ -61,7 +62,6 @@ tree-sitter-python = "=0.25.0"
# Independent implementations verify production signatures and digest fixtures.
hmac = "0.13.0"
sha2 = "0.11.0"
chrono = { version = "0.4.45", default-features = false, features = ["std"] }
livekit-protocol = "0.8.0"
# The paused clock, so the Live socket's keepalive and idle limits are tested
# at their real values instead of waiting them out.
Expand Down
21 changes: 21 additions & 0 deletions docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -290,3 +290,24 @@ Checks that need credentials or a running service stay out of the gate and are
run on their own: `scripts/server-check.sh`, `gemini-check.sh`,
`parity-check.sh`, `report-parity-check.sh`, `visual-parity-check.sh`,
`recording-integration.sh`, and `recording-provision-check.sh`.

## Instructor task packaging

The instructor tool and its acceptance tests use the pinned Python dependency
in `scripts/requirements-task.txt`. `scripts/task-package.py` finds its own
interpreter: it needs Python 3.9 or newer, and reruns itself under a newer
`python3.x` on PATH when started on an older one; `build` also needs the
dependency, and reruns under `target/task-tools`, which it sets up from that
file the first time (`CODETRIAL_TASK_TOOLS` names another place). The gate uses
`target/task-tools` too when it exists, so once one `build` has run,
`./scripts/test.sh` needs no `PYTHON`. To set it up by hand:

```sh
python3 -m venv target/task-tools
target/task-tools/bin/pip install -r scripts/requirements-task.txt
```

`scripts/task-package.py --help` lists references, new, validate,
render-prompt, preview, build and publish-scan;
[task-mode.md](task-mode.md#instructor-workflow) says what each does. The
tooling never uploads or pushes files.
33 changes: 32 additions & 1 deletion docs/integrity-evidence.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ the camera required, because the recording notice describes a video recording.

## Response windows

The replay page lists a *response window* for each turn the interviewer took:
The replay page lists a _response window_ for each turn the interviewer took:
the time between the interviewer finishing that turn and the interviewer
speaking again, taken from the `avatar` state rows the browser recorded. A
window no candidate transcript turn was attributed to says so; the turn itself
Expand Down Expand Up @@ -118,6 +118,37 @@ candidate who is thinking is slow too, and the number cannot tell the two
apart. It also cannot separate the silence before an answer from the answer
itself, because a transcript turn is recorded when it ends.

## Task mode rules

Everything above holds for interviews. [Task mode](task-mode.md) is a learning
exercise with rules the instructor publishes and the learner accepts before it
starts, and there two things change.

First, breaking an announced rule ends the attempt and marks it invalid. The
rules are observable facts: the workspace left fullscreen, the page became
hidden, or the learner's face was out of frame or tilted down past their own
calibrated limit for longer than the set allows. The first two have no grace;
the third warns with a countdown and a sound before it ends anything. All of
them are listed on the instructor's page and again before Start, with what is
not detected.

Second, the camera rule reads more than a count. Task mode estimates head pitch
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
from where the nose sits between the eyes and the mouth, as the detector's
keypoints place them (or from the face box when it gives none), against a calibration the learner
runs before the attempt, because the rule's purpose is sustained attention
somewhere below the screen, such as a phone in the lap. It still does not read
eye gaze, expression, or anything that would support a claim about what a
person is thinking, and the reasons in the next section are why: an overlay
below the webcam or a phone held at screen height defeats it, so it deters
rather than detects. A detector that stops answering is a technical
interruption, never a violation.

What does not change is the claim. An invalid attempt records which rule ended
it, when and the measurement, and says the attempt ended under the announced
rules; it never says or implies that the learner cheated, and it scores
nothing. Task mode runs on the learner's own machine, so all of this is a
deterrent the learner could disable, and the task contract says so.

## What CodeTrial does not look for

CodeTrial does not attempt to detect interview-assistant overlays, and this is
Expand Down
9 changes: 9 additions & 0 deletions docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,3 +85,12 @@ migration; refresh the prompt and report goldens; add this bundle's table row;
cover successful, incomplete, legacy, malformed, and future reports; verify
HTML, Markdown, history and progress, and replay provenance; then run the
complete local test suite.

## Task mode

[Task mode](task-mode.md) owns a separate versioned task package, sidecar,
wire protocol, rubric and Learning review. Its design version is 1, revised
before any distribution for learner-local execution; the pilot implementation
is under code and release-run verification. The active interview tuple
and its prompt/report goldens are unchanged. Task reviews use an explicit
assessment mode and never pass through the interview scorer.
7 changes: 7 additions & 0 deletions docs/observable-delivery-policy.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,13 @@ transcript that notes, reports and recovery read; the interviewer hears the
audio itself. None of this is a guarantee that provider-generated transcripts
are accurate.

In [task mode](task-mode.md), an attempt that breaks one of the set's announced
rules (fullscreen, page visibility, or a sustained look-away measured against
the learner's own calibration) ends and is marked invalid, with no ratings.
That is the rule's consequence, stated as a fact about the session; it is not
an assessment of the learner and never a claim about intent. Everything else
above applies to task reviews unchanged.

## What enforces it

The report prompt states the boundary, and the server independently scans every
Expand Down
37 changes: 28 additions & 9 deletions docs/provider-cost-and-degradation.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,6 +210,25 @@ saves it on arrival like any report, and the final outcome replaces it under
the same report id, so a tab lost during the window keeps the failure rather
than nothing.

Two exceptions cover [task mode](task-mode.md), which runs on the learner's
own machine for that learner alone.

The latest attempt's result (a Learning review, an unavailable envelope, or an
invalid or interrupted record) stays in process memory with the attempt's final
code and revision ID, so the learner's page can collect it after a room
disconnect or a reload. It is returned only to the signed-in learner with
`Cache-Control: no-store`, bounded by `reviewBytes`, and dropped once
`reviewRetentionSeconds` has passed (at the next task request, so an idle
process can hold it until then), when the next attempt ends, or on process
exit. It is
never persisted and never an input to another attempt or a later review. The
attempt's other revisions, turns, runs and review prompt are not kept with it.

The instructor's package is static problem data. Its manifest and encrypted
`tasks.enc` are downloaded at unlock and decoded in process memory; nothing
from it, encrypted or not, is written to disk, and the PIN and the key derived
from it are discarded once the set is decoded.

## What the candidate sees

The browser exposes distinct accessible states for connecting, live,
Expand Down Expand Up @@ -376,19 +395,19 @@ editor arm typed code and the whiteboard arm drew eight figures, 30 seconds
apart. The counts are the server's own `live_turn_usage` lines, read with
`scripts/analyze-gemini-usage.py`.

| Opening | Prompt tokens, first generation | Runs |
|---|---|---|
| Editor | 5,431 | 3 of 3 identical |
| Whiteboard | 5,408 | 3 of 3 identical |
| Opening | Prompt tokens, first generation | Runs |
| ---------- | ------------------------------- | ---------------- |
| Editor | 5,431 | 3 of 3 identical |
| Whiteboard | 5,408 | 3 of 3 identical |

Per text, counted with `countTokens` on `gemini-3.1-flash-lite`. The parts sum
to the live difference exactly (47 - 56 - 14 = -23):

| Text | Editor | Whiteboard | Difference | Billed |
|---|---|---|---|---|
| System instructions | 4,408 | 4,455 | +47 | every generation |
| Tool declarations | 429 | 373 | -56 | every generation |
| Greeting | 170 | 156 | -14 | every generation |
| Text | Editor | Whiteboard | Difference | Billed |
| ------------------- | ------ | ---------- | ---------- | ---------------- |
| System instructions | 4,408 | 4,455 | +47 | every generation |
| Tool declarations | 429 | 373 | -56 | every generation |
| Greeting | 170 | 156 | -14 | every generation |

The fixed overhead is therefore 23 tokens below the editor's: the whiteboard
instructions are longer, and `read_board` is declared in fewer tokens than
Expand Down
Loading