Skip to content

Sync upstream Soju06/codex-lb main - #32

Merged
maisi merged 144 commits into
mainfrom
sync/upstream-main
Sep 29, 2026
Merged

maisi merged 144 commits into
mainfrom
sync/upstream-main

Conversation

@maisi

@maisi maisi commented Sep 10, 2026

Copy link
Copy Markdown
Owner

Automated daily upstream sync could not land on main cleanly (git conflict, or it would leave multiple Alembic migration heads). main was left untouched. Resolve the conflicts here — add a merge migration if the Alembic graph has more than one head — then merge.

Soju06 and others added 30 commits September 10, 2026 04:04
…06#2264)

* feat(settings): manage Codex session prewarm from the dashboard

Move http_responses_session_bridge_codex_prewarm_enabled from env-only to
dashboard-managed (slop-removal campaign 0908, M3): a nullable BOOLEAN
dashboard_settings column of the same name (NULL = inherit the deprecated
CODEX_LB_* alias, then the code default, off), exposed on GET/PUT
/api/settings with provenance and the tri-state null-clear contract, and a
"Session bridge" card with an InheritBadge switch under Settings -> Advanced
(en/ko/zh-CN). The label states the behaviour (one generate=false warm-up on
the first turn of a new Codex session) and the cost (one extra upstream
request per new session).

The single consumer, _http_bridge_prewarm_enabled, resolves the switch from a
dashboard-settings snapshot through resolve_inheritable.
_maybe_prewarm_http_bridge_session awaits the settings-cache snapshot once per
new Codex session BEFORE prewarm_lock and reads no settings under the lock
(issues Soju06#1971/Soju06#1972 wedged keyed submits on that pattern); a spy test pins
exactly one snapshot read outside the lock and zero session-factory calls.

tiers.MIGRATING drops the entry; the env field stays one release as a
deprecated alias, so the Helm configmap/values keep rendering it. openspec
change dashboard-managed-codex-prewarm modifies the deployment-installation
"Removed tunables" requirement (switch alone, dashboard-managed, resolved
before the prewarm lock) and the responses-api-compat context/ops notes name
the dashboard setting. docs/reference/settings.md regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): read the prewarm switch through the dashboard overlay and correct the lock invariant wording

Review follow-up on the M3 codex prewarm change.

The switch now rides the existing per-request overlay instead of a new
settings-cache read: http_responses_session_bridge_codex_prewarm_enabled joins
DASHBOARD_OVERRIDE_SETTINGS, so DashboardOverridesMiddleware -- which already
takes one snapshot per request -- folds the non-NULL column over the deprecated
env alias before the proxy facade returns Settings.
_http_bridge_prewarm_enabled goes back to its one-argument shape and stays a
plain memory read, and _maybe_prewarm_http_bridge_session drops the await it
had added before prewarm_lock. Provenance and the resolver in
settings/service.py are unchanged.

The spec delta and the code comment claimed an invariant the code does not
hold: the prewarm_lock body already called
_acquire_request_state_response_create_admission and
_reconnect_http_bridge_session, both of which read the settings cache under the
lock. Both now state what is true -- the switch is resolved before the lock and
this change adds no settings read under it. The pre-existing in-lock reads are
left alone and noted in the PR body as a follow-up.

Tests: the eligibility and prewarm-path tests drive the overlay
(dashboard_overrides_bound / with_dashboard_overrides) instead of passing a
snapshot; the bridge integration harness points the middleware at the same fake
row it gives the proxy, so the dashboard switch reaches the bridge the way it
does in production. The "no DB under the lock" spy now drives a full prewarm
body (warm-up built, admitted, sent to a fake upstream, completed) and scopes
its assertion by call site: no snapshot read is attributed to the prewarm path
itself, the only read under the lock is the pre-existing admission gate's, and
no database session is opened. The disabled path asserts it never reaches the
settings cache at all. docs/configuration.md notes that an "Inherited from
environment" badge shows the serving replica's environment.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…hboard (Soju06#2279)

* feat(settings): manage stream and bridge request budgets from the dashboard

Moves `http_responses_stream_request_budget_seconds` (the 2-hour ceiling of a
streaming Responses turn) and `http_responses_session_bridge_request_budget_seconds`
(the 2-hour ceiling of an HTTP session bridge request) from env-only to
dashboard-managed, following the C2-1 pattern (Soju06#2221). Two nullable
`dashboard_settings` columns: NULL keeps today's behaviour (environment, else
code default); a value saved in the dashboard wins on every replica without a
restart, and `provenance` says where each value comes from.

Both names join `DASHBOARD_TIMEOUT_SETTINGS`, so the request-path readers pick
the dashboard value up through the existing `with_dashboard_overrides` overlay
with no call-site change, and `PUT /api/settings` evaluates the timeout
invariants on the effective values. One new rule,
`upstream-connect-within-stream-budget`, joins the existing
`admission-wait-within-stream-budget` and
`bridge-stuck-gate-retire-within-bridge-budget` rules.

Two background readers move off the bare environment value: the quota warm-up
claim lease floor now resolves from the scheduler's per-tick dashboard
snapshot, and the HTTP bridge stale-operation abandonment sweep takes one
`SettingsCache` snapshot per heartbeat pass (environment fallback, warn-only,
when the snapshot cannot be read).

Part of the slop-removal campaign 0908 (MIGRATING triage, M1).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(quota-planner): stream the warm-up probe under the dashboard budget

codex review R1 (P2): the scheduler passed its dashboard snapshot only to the
claim-TTL calculation, so the probe's own stream still used the environment
budget. A dashboard budget below the environment value could therefore expire
the claim lease while the probe was in flight and let another replica reclaim
the decision (duplicate warm-up). Bind the snapshot around the probe, the same
way the proxy warm-up path binds it around each submission.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): bound the connect timeout by the bridge budget too

Adversarial review of Soju06#2279:
- New invariant `upstream-connect-within-bridge-budget`: the HTTP session
  bridge path spends the connect timeout inside the bridge request budget, so
  a connect timeout above it can never be honoured (connect 700 / bridge 650
  was accepted before). Mirrored client-side by adding the bridge budget to
  the card's connect-budget list.
- The card also mirrors `admission-wait-within-stream-budget`: a stream budget
  below the fixed 10 s admission wait is blocked before the PUT, with strings
  in en/ko/zh-CN.
- Both new fields join the Upstream timeouts card's remount key so a saved
  value re-seeds the draft.
- openspec: the pre-response silence budget block now names the fixed 300 s
  stuck gate (constantized by Soju06#2256) and scopes the dashboard-managed clause to
  the two settings-derived terms, so it is a superset of the pending
  constantize-session-bridge-tunables block. Archive simulations in both orders
  on a copy of main's openspec/: constantize-then-M1 applies cleanly; the
  reverse aborts with openspec's own "refresh the change spec" guard rather
  than silently dropping either edit.
- Stale anchors: `ADMISSION_WAIT` → work_admission.py:25, `STREAM_BUDGET` →
  streaming/helpers.py:927.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…oju06#2265)

* feat(settings): per-model context window overrides in the dashboard

`CODEX_LB_MODEL_CONTEXT_WINDOW_OVERRIDES` (`slug -> reported context
window`) was environment-only: raising or reverting a model's advertised
window meant editing the environment on every replica and restarting,
while a too-large window turns every request for that slug into a
pre-stream 400 until the next restart. That is the T3 defect
`configuration-tiers` describes, and the field was M4 of the MIGRATING
backlog.

The value is a mapping, so its dashboard home is a small override table
(`model_context_window_overrides`: slug PK, window, timestamps) rather
than a `dashboard_settings` column, with a `/api/settings/
model-context-window-overrides` sub-API (GET merged list, PUT/DELETE per
slug). Precedence is per slug: a dashboard row wins, a slug without a row
inherits the environment entry, a slug with neither has no override. The
migration never copies the environment dict into rows.

`app/core/config/context_window_overrides.py` is the single resolver and
holds the row-snapshot cache (TTL plus the cross-replica `settings`
invalidation namespace, like `SettingsCache`); the catalog endpoints
resolve the merged map once per build, outside the per-model loops, and
`_resolved_context_window` keeps the upstream `max_context_window` clamp.
The new "Model catalogue" card in Settings -> Advanced adds, edits and
removes rows with a per-row provenance badge; environment-inherited rows
are read-only until overridden.

The environment variable stays as a deprecated per-slug fallback for one
release; its `MIGRATING` row is replaced by a `DASHBOARD_HOMES` mapping to
the new table.

* fix(settings): reject non-integer context windows and untrimmed slugs

codex review round 1 (both P2): the upsert body used a lax pydantic `int`,
so a numeric string or a bool was coerced into a stored window, and the
slug validator stripped before checking, so " gpt-5.4 " was accepted and
silently stored under a different slug than the operator wrote. Both
contradicted the change's own requirement. `StrictInt` and raw-segment
validation, with the rejected shapes pinned in the spec and tests.

* fix(settings): harden the context window override writes and cache

codex review round 2 and the adversarial pass, all on the new sub-API:

- The upsert request took an unbounded `StrictInt` while the column is a
  database `Integer`, so `2**31` passed validation and then aborted the
  PostgreSQL commit as a 500. Bounded at 2147483647, rejected as a 422.
- `invalidate()` cleared the snapshot outside the lock and the poller
  registered the bare `clear`, so a catalog build already awaiting the
  database could reinstall its pre-write rows and keep serving the old
  window for the rest of the TTL. Clear under the lock and register
  `invalidate(propagate=False)`, mirroring `SettingsCache`.
- The upsert was a read-then-insert, so two concurrent creates of the same
  absent slug raced into a primary-key 500 instead of the documented
  create-or-replace. One dialect-native `on_conflict_do_update`.
- Neither write emitted an audit event, unlike `PUT /api/settings` and the
  model-source CRUD, though a row changes what the catalog advertises to
  every client. Both writes log `settings_changed` with the slug.
- A database blip past the TTL turned `GET /v1/models` into a 500. The
  cache now serves its last known rows on a load failure, like
  `SettingsCache.cached_row`, without refreshing the timestamp.
- The MSW handlers trimmed the slug and routed a single segment, both
  contradicting the backend. `:slug*`, no trim.

* fix(settings): keep the override snapshot as a fallback and wait for it before the firewall scroll

codex review round 3, both P2:

- `invalidate()` dropped the rows, not just their freshness, so the very
  next read after any settings bump had no fallback left: a transient
  database failure then turned `GET /v1/models` into a 500. It now expires
  freshness only (with `-inf`, since a `0.0` marker reads as fresh during
  the first TTL seconds of process uptime) and keeps the last known rows;
  `clear()` remains the full reset for test teardown.
- The Model catalogue card sits above Firewall and grows by a table row per
  override, but its query was missing from `FIREWALL_LAYOUT_QUERY_KEYS`, so
  a late overrides response could push the `#firewall` deeplink target back
  out of view after the one-shot scroll.

Also re-chains this branch's migration after rebasing onto M1 (Soju06#2279):
`down_revision = "20260909_080000_dashboard_stream_bridge_budgets"`, single
head.
…6#2281)

* feat(settings): pause background schedulers from the dashboard

Moves the three background job switches — Auth Guardian, the automations
scheduler and reset-credit polling — from env-only to dashboard-managed.
Each becomes a nullable `dashboard_settings` column of the same name
(NULL = inherit the `CODEX_LB_*` variable, then the code default), is
exposed with provenance on `GET`/`PUT /api/settings`, and gets a
"Background jobs" switch group under Settings → Advanced plus a
"Pause all automations" toggle in the Automations page header.

Every loop now always starts and reads the settings-cache snapshot at the
start of each cycle (guardian refresh pass, automations tick /
`run_due_jobs`, reset-credit refresh cycle), so a toggle applies on the
next tick on every replica without a restart; `run_now()` is refused with
`409 automations_paused` while paused. The Auth Guardian keeps its
topology gate (multi-replica ring without leader election), which the
dashboard cannot override and which the API reports as
`authGuardianBlockedByTopology`.

Part of the slop-removal campaign 0908 (migrating-triage M2).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): harden the background job toggles after review

Review follow-ups on the dashboard-managed background job switches:

- The per-tick settings read now sits inside each loop's exception guard
  (automations, reset credits), so one transient database error can no
  longer kill the loop task — which would also abort the shutdown stop
  chain. Tests cover "loop survives a transient toggle-read failure".
- The automations tick takes one snapshot before its lock and threads it
  into the leader-gated body, so nothing awaits the settings cache under
  the scheduler lock.
- The reset-credit auto-redeem gate is symmetric: disabling polling while
  the opt-in is on is refused with the same error code, instead of
  silently persisting a pair that can never run. Spec scenario and API
  test added.
- Migration upgrade/downgrade test for the new columns (parent derived
  from the script directory, so a re-chain needs no test edit).
- The guardian's per-pass topology-blocked log drops to DEBUG (the
  builder already warns once); the settings service reads the topology
  fields defensively instead of type-checking the startup settings; the
  Automations pause switch and its reset action are disabled without
  write permission; a paused run-now maps to its own localized toast; and
  the inheritance badge is captioned so its on/off is not read as the
  inverse "paused".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…2275)

* feat(settings): conversation archive toggle in the dashboard

Move conversation_archive_enabled from env-only to dashboard-managed: one
nullable BOOLEAN dashboard_settings column of the same name (NULL = inherit
the deprecated CODEX_LB_CONVERSATION_ARCHIVE_ENABLED alias, then the code
default off), exposed on GET/PUT /api/settings with provenance and the
tri-state null-clear contract, plus a read-only admin-only
conversationArchiveDir (T1, this replica's local shard).

Enabling turns the proxy into a full prompt/response recorder readable by the
same dashboard admin, so the migration carries three safeguards:

- the dashboard only sends true through a confirmation dialog ("All
  prompt/response bodies will be written to each replica's local archive
  directory ..."); off saves immediately and "Reset to inherited" is blocked
  when the env alias would silently start recording;
- the settings API writes a dedicated conversation_archive_toggled audit
  event (enabled, source, actor, actor_role, actor IP) on every effective
  on/off flip, in addition to settings_changed.changed_fields;
- the card shows the archive directory read-only with the per-replica
  local shard limitation stated.

archive_enabled() resolves the toggle from SettingsCache.cached_row() (env
layer before the first load) through a shared resolve_archive_enabled(); the
15 + 5 archive_* call sites in the upstream HTTP/WebSocket clients are
unchanged and a grep gate keeps them that way. The telemetry feature flag
reports the effective value. tiers.MIGRATING drops the entry; the env field
stays one release as a deprecated alias. openspec change
dashboard-managed-conversation-archive adds deltas on
proxy-runtime-observability and audit-logging. docs/reference/settings.md
regenerated.

Part of the slop-removal campaign 0908 (M5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(settings): assert guests never see the archive directory

The conversation archive directory is a filesystem path of this replica and
the settings response only fills it in for admin principals. Pin that in the
guest-access integration test next to the archive-records 403, and record the
admin-only exposure in the spec delta.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): keep the last settings row after a cache invalidation

``SettingsCache.invalidate`` dropped the cached row outright, so
``cached_row()`` returned None between an invalidation and the next load. The
conversation archive gate reads that row synchronously and falls back to the
deprecated CODEX_LB_CONVERSATION_ARCHIVE_ENABLED alias when it is None: on a
host that still sets the alias to true, every settings mutation reopened a
window in which an archive an operator had just turned off in the dashboard
resumed recording — indefinitely on a replica serving only long-lived
WebSocket traffic, which loads a snapshot once at connection start.

Keep the last loaded row in a separate slot that ``invalidate`` does not
clear; ``get()`` still reloads because only the freshness slot is expired.
This also restores the documented contract of the dashboard-overrides
middleware ("only before the first successful load does the environment
fallback apply"), which had the same gap whenever the database was unreadable
right after an invalidation.

Found by codex review.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): refresh the settings snapshot off the invalidation bus

The conversation archive gate reads the last loaded dashboard row and cannot
await, so a replica that is carrying nothing but already-open streams never
pulled a new snapshot in: an operator disabling the recorder on one replica
left the others recording until some request happened to reload the cache.
Register ``SettingsCache.refresh`` as a second ``settings`` invalidation
callback (the ``account_routing`` namespace already refreshes rather than only
evicting), so the new row lands within the poll interval on every replica;
the invalidate registered before it still runs, so a failed refresh degrades
to the ordinary TTL reload instead of serving a stale value.

Also from review:

- ``tests/conftest.py`` resets the new fallback slot, so a dashboard row can
  no longer leak between tests (the local fixture that hid this is gone);
- the confirmation dialog and the recording notice no longer promise that the
  audit line carries "your identity" unconditionally — password auth has no
  per-user identity, so the three locales now say "when the authentication
  mode provides one", and an integration test pins the actor under
  trusted-header auth;
- ``archive_enabled()`` hoists the field default out of the per-frame path;
- the migration test reads its parent revision from the alembic graph instead
  of a literal (this revision is the tail of a stack that gets re-chained),
  an unrelated PUT is asserted to emit no toggle event, and the M5 ownership
  audit folds into the existing resilience-toggle loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(main): teach the lifespan settings-cache doubles the refresh callback

The seven test_otel lifespan tests stub get_settings_cache with a
SimpleNamespace; the new settings invalidation callback needs refresh
on it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* docs(openspec): archive the round-3 standalone landed changes

Archive the five changes from the 2026-09-09 slop-removal round whose PRs
merged before the settings campaign: backoff-upstream-burst-429-on-http-stream
(Soju06#2240), document-helm-pre-1-13-shim-deprecation (Soju06#2230),
background-paths-read-dashboard-snapshot (Soju06#2233), retention-inherit-badge
(Soju06#2231) and automation-run-pinned-budget (Soju06#2241).

Every MODIFIED delta matched the current spec text; none needed a rebase.
Spec files shared with a later batch (configuration-tiers, automations,
deployment-installation) are merged in the batch that touches them last.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(openspec): archive the retier and constantize changes

Archive retier-topology-settings (Soju06#2252), constantize-core-tunables (Soju06#2261)
and constantize-session-bridge-tunables (Soju06#2256) in merge order, so the
constantized text is in the specs before the dashboard-managed changes that
edit the same requirements.

The two constantize changes carry explicit REMOVED blocks for the
requirements whose settings disappeared (live-usage-ingestion "Live ingestion
is decoupled and switchable", rate-limit-reset-credits "Reset credit polling
interval is configurable", responses-api-compat "Proxy-generated prompt cache
key derivation is operator-toggleable" and the two renamed session-bridge
requirements); no heading is lost without one. No delta needed a rebase.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(openspec): archive the dashboard-managed settings changes

Archive the five dashboard-managed changes in merge order:
dashboard-managed-codex-prewarm (Soju06#2264), dashboard-managed-stream-bridge-budgets
(Soju06#2279), dashboard-managed-context-window-overrides (Soju06#2265),
dashboard-managed-background-jobs (Soju06#2281) and
dashboard-managed-conversation-archive (Soju06#2275).

dashboard-managed-codex-prewarm's deployment-installation delta was rebased
before archiving: it was authored against the pre-Soju06#2261 text, so it dropped the
"never-tuned core tunables constantized by constantize-core-tunables" bullet
list and reverted the removed-settings example that Soju06#2261 had updated. Both
current sentences were restored and only the prewarm edits (the dashboard
switch paragraph, the two scenario rewrites and the new env-alias scenario)
were kept. The other four deltas already matched the current text.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…cally unavailable (Soju06#2234)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…n to constant windows (Soju06#2235)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…oju06#2244)

* feat(proxy): trace terminal-frame delivery on Responses SSE streams

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(proxy): stamp active pipelined requests and commit terminal delivery only after the full frame

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…06#2263)

* refactor(egress): use native transport for direct usage fetches

* fix(egress): preserve direct usage transport semantics
…ource (ship-dark) (Soju06#2257)

Decision at route admission via a read-only exhaustion probe; dispatch through the hardened direct source route with a SourceDispatch owner.
Thread pins, SDK anchors and tombstones with neutral release; per-source breaker, bulkhead and fast-decline; not_portable_history hint on unservable pinned history.
WebSocket evidence routing (WP-D folded in): handshake 426 on pin/tombstone/bounce evidence, in-band 503 plus bounce row, anchor routing on open sockets.
Source-direction telemetry header filter; client store intent restored on the source body (anchor iff store is not false).
SHIP-DARK: with dashboard_settings.subscription_overflow_* NULL the request path is byte-identical and performs no probe or pin lookup.
Production flip is gated on a separate-deployment canary; codex P2 (anchor when content precedes response.created) tracked as a pre-flip follow-up on Soju06#2123.

Part of Soju06#2123 (WP-C2 + WP-D).

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…TTL (Soju06#2236)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* perf(request-logs): keep SQLite facet scans on value indexes

* docs(openspec): archive verified SQLite facet query fix
…ries (Soju06#2246)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…atency + per-model output throughput) in weighted selection (Soju06#2258)

* feat(load-balancer): discount relatively slow first-token accounts in weighted selection

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(load-balancer): exclude websocket replays and admission waits from TTFT cohort samples

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(load-balancer): weight fresh draws by per-model output throughput and gate cohort transitions per signal

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(load-balancer): prepare live selection states in sticky_selection

`LoadBalancer._prepare_sticky_selection_states` now delegates to
`sticky_selection.prepare_selection_states` (reclaim stale leases, prune
the runtime, build the states, narrow to the pinned account). The method,
its signature and the `_build_states` seam tests patch on the balancer
module are unchanged; no behaviour changes. `load_balancer.py` goes from
3029 to 3009 lines, back under the 3021 architecture ceiling that
origin/main sits exactly on, with room for the latency cohort hook.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(load-balancer): measure throughput to the upstream terminal, not settlement

`_finalize_websocket_request_state` computed `latency_ms` only after
awaiting the API-key settlement (and the deferred backoff / health writes),
so for keyed WebSocket and bridge turns the throughput cohort sample counted
local DB contention as generation time: 400 tokens over 10 s followed by a
10 s settlement read 20 tok/s instead of 40 and discounted a healthy
account relative to HTTP or unkeyed traffic.

The upstream reader now stamps `upstream_terminal_at` on the turn when it
parses the terminal frame; the finalizer captures it (falling back to its
own entry time) before the gate release and the settlement and hands the
funnel a `latency_upstream_terminal_ms`, which `record_tps_sample` uses as
the end of the span. The persisted `latency_ms` is unchanged. HTTP stream
rows are written before settlement, so their `latency_ms` already ends at
the terminal and they keep using it.

Regression: a delayed-settlement WebSocket turn through the real finalizer
and request-log funnel records 40 tok/s (with and without the reader's
stamp) while its row keeps the 22 s wall latency. OpenSpec delta names the
upstream-terminal span.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(load-balancer): stamp upstream terminals on HTTP and bridge paths and sample TTFT from upstream send

Review follow-ups on the latency cohort sampling seams (PR Soju06#2258).

Throughput span end. `_stream_once` yields the terminal frame and only writes
its row in the generator's `finally`, so a slow downstream drain or an
upstream connection lingering after `response.completed` landed in the span;
the HTTP-to-WebSocket bridge consumes upstream events through its own reader,
so `upstream_terminal_at` was never set and the finalizer's entry-time
fallback included the durable alias / operation / circuit-settlement writes
that run before it. Every path now stamps the terminal at its parse site: the
HTTP stream on `_StreamSettlement.upstream_terminal_at` via
`streaming.helpers._stamp_terminal` at both terminal-detection sites (before
the yield) and passes `_upstream_terminal_latency_ms(settlement,
attempt_started_at)`; the bridge reader on the matched request state right
after the terminal pop. `streaming/mixin.py` stays at 1100/1100 (two
two-line terminal checks folded into the helper call, two init lines merged).

First-token sample start. The bridge row's `latency_first_token_ms` starts at
request-state creation, before session-lease reacquisition, prewarm, image
inlining and slimming, of which `bridge_queue_wait_ms` captures only the
final gate. The funnel now takes `latency_upstream_send_ms` (start ->
`response.create` send; the finalizer derives it from
`response_create_sent_at`, the HTTP stream passes 0 since its attempt clock
is re-anchored right before the send) and `record_ttft_sample` records
`latency_first_token_ms - latency_upstream_send_ms`; rows without a send
anchor or with a send after the first token are never sampled. The persisted
`latency_first_token_ms` is unchanged.

Transition log gate uses `>=` so a move of exactly `_LOG_DELTA` is logged, as
the comment says.

Regressions through the real paths: `_stream_once` with a 5 s downstream
drain plus a 5 s lingering upstream records 40 tok/s and a 2000 ms TTFT
sample while the row keeps 22 s, and an upstream without a terminal frame
records nothing; `_process_http_bridge_upstream_text` stamps the terminal
10 s before the finalizer is entered; the finalizer + funnel records a
1700 ms sample for a bridge turn with 3 s of pre-send work (row keeps
4700 ms), nothing for an unanchored turn, and 40 tok/s for a bridge turn
with 10 s of bookkeeping before the finalizer plus a 10 s settlement (row
keeps 32 s). OpenSpec delta makes the send anchor and per-path terminal
stamp normative with scenarios; proposal, tasks and docs/routing.md updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* docs: add codex-lb-status community companion

* docs: add @VictorStatko as a contributor

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
* fix(warmup): warm unused Free quota once

* fix(warmup): stop after lost initial claim

* fix(warmup): preserve window locks during rolling upgrades

* chore(ci): rerun transient browser smoke

---------

Co-authored-by: Hulian Felipe Muller Buligon <hulian@MacBook-Pro-de-Hulian.local>
* docs(companions): add Codex LB for Omarchy

* docs: add @janaki-sasidhar as a contributor

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
* docs: add @NikitaMGrimm as a contributor

* fix(db): preserve rollback after SQLite invalidation
…#2159)

* chore(deps): bump vitest from 4.1.11 to 5.0.0 in /frontend

Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.0.
- [Release notes](https://github.com/vitest-dev/vitest/releases)
- [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md)
- [Commits](https://github.com/vitest-dev/vitest/commits/v5.0.0/packages/vitest)

---
updated-dependencies:
- dependency-name: vitest
  dependency-version: 5.0.0
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* chore(deps): upgrade Vitest runner and coverage together

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Soju06 <qlskssk@gmail.com>
Bumps [httpx2](https://github.com/pydantic/httpx2) from 2.10.0 to 2.12.0.
- [Release notes](https://github.com/pydantic/httpx2/releases)
- [Changelog](https://github.com/pydantic/httpx2/blob/main/src/httpx2/CHANGELOG.md)
- [Commits](pydantic/httpx2@v2.10.0...v2.12.0)

---
updated-dependencies:
- dependency-name: httpx2
  dependency-version: 2.12.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…oju06#1621)

* spec: TLA+ model of ownership/continuity/timeout core with TLC negative controls

Model covers account lease, API-key reservation, bridge session, websocket
turn, continuity anchor provenance, owner epoch, gate waiters, freshness
versions, and shutdown drain (3 replicas, 2 accounts, 2 turns).

spec/check.sh: full model = 5,792,958 distinct states, zero invariant
violations, deadlock checking on; 7 weakening configs each reproduce a
distinct historical bug class as a TLC counterexample (mapped to fix
commits in spec/README.md).

* spec: make TLA checks non-vacuous

* spec: extend the TLA+ model with the 2026-08-06 live failure classes

Every new bug class extends the leg that should have caught it, so the three
classes proven on the live stack today each get model state, an invariant or
liveness property, and a negative control that reproduces the failure.

Anchor account ownership (PR Soju06#1638). A continuity anchor now carries the
account that owns it. Upstream accepts a request carrying a
previous_response_id owned by a different account and then never sends
response.created, so UpstreamRespondsTo gates StartStream, CompleteTurn and
ClaimCompletedDelivery: a foreign-anchored turn can only leave the
pre-response phase through a timer or a cancel. Inv10AnchorAccountOwnership
forbids dispatching with a foreign-owned anchor;
weak-cross-account-anchor.cfg reproduces the wedge.

Pre-response eventless phase (PR Soju06#1633). The "active" phase - dispatched
upstream, response.created not seen yet - now has its own bound instead of
sharing the request/stream-idle budget, and ExpireDeadline picks both its
bound and its budget label per phase. Inv11PreResponseBudget requires the
pre-response bound to be the minimum of the named gate-retire and stream-idle
budgets, at or above the keepalive cadence floor, strictly below the
post-start budget, and forbids reporting a kill under the post-start
stream-idle budget while the response had not started.
weak-conflated-timers.cfg collapses the two timer names and TLC produces a
healthy pre-start wait killed under the wrong budget.

Bounded client retry backoff (PR Soju06#1634). A turn killed in the pre-response
phase tears the client session; ClientRetryAttempt repairs it and is fair,
but only while retryBackoff stays inside MaxRetryBackoff. The new liveness
property TearEventuallyRecovers states that a recoverable tear is eventually
recovered; weak-unbounded-backoff.cfg lets the backoff grow past every
deadline in the model, and TLC produces a behaviour where the client stays
torn forever - the 29-hour client sleep observed today.

check.sh gains the three mappings and a PROPERTY:<Name> expectation form: a
liveness control must violate exactly the temporal property its config
declares and must not violate any invariant.

bash spec/check.sh exits 0. Full model: 16696096 states generated, 3606740
distinct, depth 23, zero violations, deadlock checking enabled. All 13
weakenings produce their mapped counterexample.

* spec: model anchored dispatch and phase-specific deadlines

* spec: reset phase deadlines and retry state per turn

* fix: preserve relative TLA timeout budgets

* fix(spec): refine core ownership formal phases

* fix(spec): close the model's vacuous checks

Rebased onto current main and addressed the four open review threads on the
formal model:

- UseAnchor now fires at most once per anchor value, and AcquireTurn/StartStream
  reset that budget when they write a new anchor. The action used to stay
  enabled forever and reproduce the current state, giving every anchored turn a
  Next self-loop that hid deadlocks from TLC.
- AcquireTurn consumes the route that RouteFromSnapshot produced, including its
  replica and account, so no dispatch can bypass the freshness check.
  snapshotRouteAttempted becomes snapshotRoute and carries the routed pair.
- A mismatched-lineage anchor is now a reachable input: same account, current
  owner epoch, safe provenance, incompatible lineage. It is the only input that
  isolates the lineage conjunct of AnchorSafe, and weak-ignore-anchor-lineage.cfg
  is the negative control that accepts it and violates Inv1AnchorCurrent.
- Inv5SingleOwnerCAS pins each live turn's replica and epoch to owner/ownerEpoch
  instead of only counting live turns, so reassigning the durable owner under a
  live turn is a violation rather than an unchecked state.

* fix(spec): close formal-model review gaps

* chore(spec): allow formal model at repository root

* fix(formal): model idle progress and late terminal producers

* fix(spec): detect mismatched lineage at anchor use

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
* refactor(egress): interpret WebSocket routing metadata in Rust

* fix(egress): isolate oversized WebSocket integers from IPC
Bumps [astral-sh/uv](https://github.com/astral-sh/uv) from 0.12.10 to 0.12.12.
- [Release notes](https://github.com/astral-sh/uv/releases)
- [Changelog](https://github.com/astral-sh/uv/blob/main/CHANGELOG.md)
- [Commits](astral-sh/uv@0.12.10...0.12.12)

---
updated-dependencies:
- dependency-name: astral-sh/uv
  dependency-version: 0.12.12
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…backfill (Soju06#2298)

* feat(metadata): automate pricing and Codex version updates with cost backfill

* fix(metadata): address review findings and keep generated references in sync
* chore(deps): bump @playwright/test

Bumps the frontend-minor-patch group in /frontend with 1 update: [@playwright/test](https://github.com/microsoft/playwright).


Updates `@playwright/test` from 1.62.1 to 1.63.0
- [Release notes](https://github.com/microsoft/playwright/releases)
- [Commits](microsoft/playwright@v1.62.1...v1.63.0)

---
updated-dependencies:
- dependency-name: "@playwright/test"
  dependency-version: 1.63.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: frontend-minor-patch
...

Signed-off-by: dependabot[bot] <support@github.com>

* build(nix): refresh Playwright dependency derivations

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Soju06 <qlskssk@gmail.com>
* fix(proxy): sanitize rebuilt HTTP request hop-by-hop headers

Reapply the fork's local hop-by-hop header sanitation on current upstream main while preserving upstream routing-hint behavior and isolating regression coverage from active upstream test files. Validation is intentionally left pending on this rebased candidate.

* test(proxy): cover remaining hop-by-hop headers

* chore(contributors): credit header sanitation and verify current main

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
Soju06 and others added 29 commits September 18, 2026 13:40
…un.nix (Soju06#2453)

* fix(ci): unbreak main on the rustls advisory and a missing contributor

Two independent failures are red on main at 947c8f6, and both are required
checks, so every open PR is blocked behind them.

RUSTSEC-2026-0285 (published 2026-09-14) covers rustls <0.23.45: TLS 1.3
handshake messages were accepted across encryption level boundaries. The
workspace pinned `=0.23.36`. 0.23.45 requires aws-lc-rs ^1.18, so that pin
moves from `=1.16.2` to `=1.18.1` in the same change; cargo relocked
aws-lc-sys 0.39.1 -> 0.45.0 and rustls-webpki 0.103.13 -> 0.103.15 with it.

`Contributors attribution` fails because GitHub attributes two commits
(Soju06#731, Soju06#1127) to felixcake618, who is a different account from the already
listed Felix201209. Added the entry.

`Nix flake check` is also red on a stale frontend/bun.nix; that needs a nix
toolchain to regenerate and is handled separately.

* fix(nix): regenerate frontend/bun.nix from the current bun.lock

Soju06#2435 bumped the frontend-minor-patch group in bun.lock without regenerating
bun.nix, so `Nix flake check` is red on main.

Delta is 39 added / 38 removed package entries and no hash changes on retained
ones: the @rolldown/binding-* set 1.2.5 -> 1.2.8, @typescript-eslint/* 8.69.0 ->
8.70.0, @types/{node,react,react-dom}, lucide-react 1.41.0 -> 1.45.0, picomatch
dropped, and @oxc-project/types added.

bun2nix is distributed only through the flake and this host has no nix, so the
file was regenerated by deriving each entry from bun.lock. The derivation was
validated by reproducing all 878 pre-existing entries byte-for-byte before
applying the delta, and the diff touches package blocks only. CI regenerates
with the real bun2nix and diffs, so it remains the authority.

---------

Co-authored-by: Soju06 <soju06@users.noreply.github.com>
List the Home Assistant integration in community companions so operators can discover pool and per-account quota sensors.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Soju06 <qlskssk@gmail.com>
Co-authored-by: Soju06 <soju06@users.noreply.github.com>
…Soju06#2454)

Soju06#2453 moved rustls =0.23.36 -> =0.23.45 for RUSTSEC-2026-0285 and aws-lc-rs
=1.16.2 -> =1.18.1 to satisfy it, but test_native_egress_lockfile_pins_codex_release_family
asserts the literal versions against Cargo.lock, so main's unit lane is red.

Soju06#2453's own pytest lanes reported pass in 3s because the `backend` path filter
does not select Cargo.toml / Cargo.lock, so the lane was skipped rather than
run - the failure only surfaced on the next PR that did touch backend paths.
Worth a follow-up: the backend filter should include Cargo.lock, since a
lockfile change can break a backend test.

Co-authored-by: Soju06 <soju06@users.noreply.github.com>
…06#2316)

* ci: classify contributor PRs and newly opened issues additively

* docs: sync and archive contributor labeling contract

* test(ci): skip the issue-labeler contract test without node

Matches the tests/unit/test_helm_external_secrets.py precedent so `make test-unit`
does not fail in the repo's own flake.nix devShell, where node is absent.

---------

Co-authored-by: Soju06 <soju06@users.noreply.github.com>
Co-authored-by: ubuntu <u@c>
Co-authored-by: Soju06 <qlskssk@gmail.com>
…2431)

Add a SCIM 2.0 `/scim/v2/Users` endpoint so an identity provider can
create, update and deactivate dashboard accounts without an operator
inviting each person by hand, plus the `/api/scim-tokens` management
surface and the Organisation settings card that issues the credential.

The credential is a hashed row: only the SHA-256 digest is stored, in a
unique index, so verification is an index equality and no readable copy
survives the response that issued it. A SCIM token authenticates SCIM
and nothing else — it never becomes a `DashboardPrincipal`, never reaches
the dashboard or the proxy, and carries no grants. Because it carries
none, issuing one is itself an `ADMIN_GRANTS` delegation, and a pushed
role slug is bounded to the non-admin presets: an identity provider push
may not create or promote to an admin, which would otherwise inflate
`last_admin_protected` with an account that has no password and so can
never be a qualifying break-glass holder.

Provisioning reuses the existing `IdentityResolver` rather than forking
it: `_provision` and `_link_invited` become public entry points that the
sign-in path also calls, so the slug rules, the collision walk, the
reserved `admin` name and the retry all exist once.

Deprovisioning goes through the shared `deactivate_user()`, which this
change completes: an `invited` account is now deactivated rather than
silently ignored — its invite is deleted in the same transaction and the
row moves to `disabled`, closing an `sso_only` invitation that otherwise
never expires — and a conditional write that matches no row now re-reads
the counts and names the invariant that actually refused instead of
always reporting the last-admin one.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(db): drop the withdrawn subscription-overflow schema

The feature was withdrawn in Soju06#2416: a 60-body capture sweep of real Codex
traffic found nothing the portability predicate would ever pass, and making
it pass would mean translating request bodies per provider. That change
removed the code and deliberately left the storage alone so the tree matched
the deployed schema at every commit; this one takes the storage.

Production never wrote a row -- `model_source_pins` empty,
`subscription_overflow_source_id` and `subscription_overflow_drain_until`
NULL, no request log carrying a `model_source_id` -- so the upgrade cannot
lose data. Every step is guarded, and the downgrade rebuilds exactly what
`20260908_000000_add_subscription_overflow` and
`20260911_000000_model_source_pins_kind_expires_index` built, so the pair
round-trips.

The three historical revisions stay: an install stranded below them still
upgrades through them, and deleting a revision would fork the graph. Their
tests keep asserting what they build at their own revision; only the
assertions about the state *at head* are inverted, because head is now past
the withdrawal. The merge-revision test's data-preservation claim moves to
the merge itself -- the last revision where both branches' schema coexists --
and gains the withdrawal as a separate fact. The pin index's query-plan test
is deleted rather than re-pointed: its subject was a query the dashboard no
longer makes against a table that no longer reaches head.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(db): anchor the overflow merge spec and repair pin indexes on downgrade

Addresses both review findings on the retirement revision.

P1, the OpenSpec gate: `database-migrations` still carries "Overflow and
transport migration heads converge without rewriting history", whose scenarios
promise that a populated parent walked to `head` keeps its rows and lands on
"the single merge head". Both clauses were written when
20260908_020000_merge_overflow_transport_heads was head; it is an interior node
now, and this revision deliberately drops the pin rows that sentence protects.
The new change folder re-anchors the claims at the merge revision — the same
correction the merge revision's own tests already needed — adds a scenario for
continuing past it, and states the retirement contract: what the revision
drops, that discarding pins is authorized because they were routing state for a
router removed one revision earlier, that every step is guarded, and that the
downgrade restores both originals' work.

P2, a real downgrade defect: the two CREATE INDEX calls rode along inside the
CREATE TABLE guard, so a downgrade interrupted between the table and its
indexes — or a partial restore carrying the table alone — finished stamped at
the parent with neither index and nothing later to build them. Each index is
now reflected on its own, exactly as 20260908_000000_add_subscription_overflow
does it. The regression test plants the table without indexes at the retirement
revision, downgrades, and asserts both indexes and both settings columns are
back; it fails on the pre-fix revision with zero indexes.

* docs(deploy): require stopping pre-withdrawal replicas before the overflow drop

The change folder only documented the rollback direction, and that left the
larger window undescribed: the upgrade itself is not rolling-safe. Every
release below this one maps both `dashboard_settings.subscription_overflow_*`
columns and loads the settings row as one entity -- `get_or_create()` issues
`session.get(DashboardSettings, 1)`, and the upstream-proxy resolver issues
`select(DashboardSettings)` -- so a pod of an earlier release that is still
serving when the drop commits fails every settings read with `UndefinedColumn`.
The chart's migration Job is a `pre-upgrade` hook on every values branch of
`codex-lb.migrationHookPhases`, so it runs before the new pods roll and
therefore before the old ones drain: an ordinary `helm upgrade` leaves that
window open, exactly as it did for 20260912_010000, which already documents the
stop on the same page.

So the requirement is now stated where an operator will meet it: a sibling
section in `docs/deployment/kubernetes.md` naming both columns and pointing at
the remedies the credential section already lists, a normative paragraph and
scenario in the spec delta, the ordering bullet in the proposal, and the
revision's own docstring.

It also says what it does not do. This drop logs no drain warning of its own:
`check_legacy_credential_drop()` is written around the credential columns --
its message, its sentinel and its fresh-install evidence all name them -- and
turning it into a table of not-rolling-safe drops is a separate concern from
retiring this storage. The docs section says so rather than leaving an operator
to expect a warning that will not come.

No behavior change: `openspec validate --strict` passes on the change and
`--specs` stays 65/0, `mkdocs build --strict` builds, and the migration suite is
untouched at 52 passed / 9 skipped.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Soju06 <r@r.local>
…ually sends (Soju06#2449)

Upstream reports a spent account in two delivery forms: an HTTP body, which
carries a status and usually a code, and a serialized `response.failed` frame,
which carries no status and no error code at all. A code-less envelope
normalizes to `upstream_error`, which is in no transport retry list and in
neither account-health code set, so both streaming gates that read such a frame
answered from a code table that the sentence upstream did send never reached:

- the retry decision surfaced the rejection on the first account, while the
  identical rejection delivered as a body walked the pool;
- the account-health write left the account ACTIVE, while the identical *coded*
  frame benched it.

Measured on three accounts inside one request, the coded frame produced
{a: rate_limited, b: rate_limited, c: rate_limited} and the code-less frame
produced {a: rate_limited, b: rate_limited, c: ACTIVE}. The account that had
just said its window was spent came out the pool's healthiest and was selected
first next time. It bites the last account of a walk, reached with retry no
longer permitted so its frame is terminal rather than retried, and separately
any frame arriving after a lifecycle event is already downstream.

One predicate now answers for every delivery form, read from the message after
folding non-alphanumeric runs to single spaces so punctuation and wrapping
cannot defeat it. It is deliberately status-free -- the same rejection arrives
with a status and without one -- and may decide only for `upstream_error`, the
code a missing code normalizes to. `invalid_request_error` is excluded: it is
also upstream's catch-all for request-shaped failures, and the HTTP paths pass
their status to the health write as evidence only, so nothing downstream could
separate a real rejection from an echoed sentence. A message-derived usage
limit classifies `rate_limit` rather than `retryable_transient`, so a spent
account is benched instead of being waited out on itself; as a consequence a
code-less 429 proving the usage limit is no longer treated as a burst.

The WebSocket transport had the same defect from the same cause, and it is the
transport most of the traffic uses. There the health write is also what retires
the socket and releases the turn for replay, so a code-less frame left the
spent account unbenched and still connected. Its replay-code and account-health
gates now come from the same predicate, answering under the usage-limit code so
an owner-pinned turn takes the coded form's path. The immediate coalesced usage
refresh is likewise triggered by the rejection rather than the literal code.
…29 (Soju06#2373)

* docs(openspec): propose walking the account pool before surfacing a 429

A single pooled account's 429 is returned to the client as though it were the
pool's 429. The walk stops at a fixed per-transport attempt constant (3/3/2)
unrelated to pool size, and the failure surfaced is the first rejecting
account's verbatim upstream body, so a 28-account fleet gives up after three.

This change proposes: bounding the walk by the pool instead of a constant
(request deadline + runaway ceiling + monotone exclusion progress), deciding
the client-visible failure at the end of the walk through the existing
exhaustion probe, and teaching the single classifier to answer "is this
account exhaustion" beside its unchanged failure class so a model-capacity
rejection cannot rotate a healthy pool.

Spec only; no code in this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): correct three defects the first implementation exposed

The exclusion answer was specified as an exhaustion answer. "Is this account
exhausted" and "may the walk move off this account" are different questions and
the walkable set includes retryable_transient, so a code-less burst 429 — which
this change's own scenario requires to be excluded — and a model-capacity 429,
which must not be, both come back false from an exhaustion field. The
requirement now states the selection predicate directly.

The terminal probe could not answer inside the request. pool_usage_exhaustion
needs a rate-limited status AND an at-or-above-limit usage sample;
handle_quota_exceeded writes the sample, handle_rate_limit does not, and the
only other supplier is a debounced background refresh that cannot land before
the terminal probe of the request that provoked it. "Every account exhausted
yields the canonical pool rejection" was therefore unreachable through the
mandated mechanism, for coded rejections too. Added as its own requirement.

The runaway ceiling was smaller than the fleet. Sixteen attempts on a
28-account pool is not a fence, it is the fixed cap this change removes wearing
another name; the ceiling must now exceed the largest supported pool.

Also: the message override is narrowed to the two codes that carry no
classification decision of their own, so it cannot silently reverse
"Model-capacity messages are retryable transient failures" or the
overloaded_error rule; the monotone-progress proof gets its capacity-recovery
carve-out, which has a live counterexample in the target function; and a walk
must record which bound ended it, since every bound returns "surface" today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): name the relocation change by its actual slug

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): align the proposal with the corrected exclusion answer

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): prove pool exhaustion from the walk, not from account state

The previous revision asked handle_rate_limit to record the usage sample
pool_usage_exhaustion needs, mirroring handle_quota_exceeded. Implementing it
showed the mirror is inert: both write to the transient AccountState that
_state_for builds, _sync_runtime_state copies across only the fields
RuntimeState declares, and a usage sample is not one of them. A re-read returns
None and the probe still reports a healthy pool, so the scenario this change
exists for stayed unreachable — through the original mechanism and through the
repair alike.

The walk already knows which accounts it attempted and why it excluded each.
That evidence is now what proves exhaustion, with the persisted-state probe
authoritative only for requests that never attempted the whole pool.

Also resolves a contradiction inside this change folder: the terminal
requirement demanded the probe be consulted on every walk end, while the budget
scenario demands upstream_request_timeout, which is only reachable by skipping
it. Both that bound and a non_retryable failure now answer without the probe,
and the canonical pool rejection must carry a retry hint when it has no
resets_at to offer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): declare the usage-refresh trigger change the classifier forces

Moving the immediate-refresh gate from the literal usage_limit_reached code to
the classification widens an existing requirement that is worded on the code.
Declared as a MODIFIED entry rather than left implicit, with the throttling and
quota carve-out restated so the widening cannot be read as covering them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): let the health write decide whether an account is benched

The capacity carve-out was written to apply to every failure class, including
the ones whose health write benches the account. Implementing it showed the two
answers then contradict each other for the one envelope that carries both a
benching code and the capacity message: mark_rate_limit has already persisted a
status and a reset deadline, so "do not exclude" hands the walk a candidate
selection will not offer. Running the delta as written against the branch
produced four failures.

The carve-out now applies only to classes whose health write leaves the account
selectable. A rate_limit or quota classification excludes even under a capacity
message, because the bench is already a fact and the walk must not disagree
with it.

Also drops the stale claim that the settings ratchet sits at 96; it is 95 with
zero headroom, so the requirement is that it does not move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): do not read a request rejection from its own message

The message override admitted invalid_request_error alongside the code-less
envelope. That envelope is how upstream rejects a request, and a rejection
frequently quotes the request back, so a message reading the usage-limit phrase
out of content the client supplied would bench an account that is perfectly
healthy — and benching is the expensive direction, since the account leaves the
pool until its reset deadline.

An envelope carrying no code at all cannot be quoting anything, so the override
is now limited to that one case. Two implementation branches had already
diverged on this point, each with tests asserting the opposite answer; this
settles it toward the safer one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): stop two changes from modifying one requirement

Both this change and classify-usage-limit-terminal-frames carried a MODIFIED
block for the burst-cooldown requirement and for the immediate-usage-refresh
requirement. A MODIFIED block replaces the whole requirement, so archiving the
second would silently clobber the first — and the two were written against the
same pre-change text, so neither contained the other's edit.

The sibling lands first and owns the delivery-form half of both: it is the
change that teaches the terminal-frame path to read a usage-limit rejection at
all. This change's burst-cooldown block is now written against the text that
sibling leaves behind, keeping its usage-limit carve-out and adding only the
owner-bound/unbound scoping the pool walk needs. The duplicate usage-refresh
delta is deleted outright.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: rvw <r@r.local>
…ju06#2374)

* docs(openspec): propose relocating anchored turns across accounts

A Codex continuation sends a delta plus previous_response_id, and that anchor
is owned by the account that produced it. When the owner is exhausted or gone
the proxy fails the turn closed and the conversation dies, even though the
pool has usable accounts and the durable spool holds every parent turn's
request body and terminal response.

get_replayable_transcript already walks that chain fail-closed and has zero
production callers: Soju06#2336 deleted the caller and kept the storage. This change
proposes one shared relocation decision for every transport, a strict rebuild
from the spool for definitive pre-dispatch rejections, and one fenced
at-least-once relocation for ambiguous eventless dispatches.

Partially reverses aae61f6 (Soju06#2336) for the ambiguous case, without restoring
the deleted four-valued recovery-mode setting. Blocked on Soju06#2366, which deletes
the fence primitives this change is the caller for.

Spec only; no code in this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): ground anchored relocation in production measurement

7 days on the production fleet, requests that died with no response event while
carrying a conversation: 2,352 turns across 993 distinct conversations in the
ambiguous class against 1,883 turns across 367 conversations in the definitive
class. The ambiguous class is not a tail case, so the fenced lane earns its
fences.

Two findings corrected the spec rather than confirming it.

`upstream_operation_status_unknown` fired zero times in that window: the
bounded 503 lives on the bridge submit path, while the ambiguous eventless
failures that actually occur surface on the streaming path as
`stream_incomplete`. A refused claim now terminates through whatever
fail-closed outcome its own transport already produces, and must not be
reported as pool exhaustion or turned into a second dispatch.

The definitive class concentrates ~5 dead turns per affected conversation
against ~2.4 for the ambiguous class, which is the client retry loop failing
identically against a dead anchor. Recorded as a verification signal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): scope anchored relocation to the lane that has the material

Three corrections the implementation found, all of which made the proposal
claim more than the mechanism can deliver.

The chain order was stated backwards. get_replayable_transcript calls
turns.reverse() before returning, so it hands back oldest first; the spec said
"oldest turn last". A wiring author following the normative text would dispatch
a chronologically reversed conversation that passes every structural predicate
and fails silently.

The rebuild has material only where the spool is written. record_operation has
exactly one caller, the HTTP bridge submit path; the WebSocket path imports the
coordinator for owner lookup and records nothing. Split by transport the
measured dead turns are 2,151 definitive and 1,768 ambiguous on downstream
WebSocket against 30 and 622 on HTTP, so this change reaches the smaller lane.
A transport without durable material must now report an absent transcript and
keep today's behaviour rather than relocate on partial context.

The fenced lane leaned on a dedupe contract that only one transport has.
tool_call_dedupe is wired from the bridge submit path and nowhere else, so on
the other transports the duplicate case 2 knowingly risks would be unbounded in
kind rather than in tokens. The fenced lane is now restricted to transports
carrying that contract.

Also records that Soju06#2366 merged and was reverted by Soju06#2383, and aligns the
overlap-dedupe task with where the projection has to run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): decide the rebuild boundary by shape, not by content equality

Three implementations of the chain/client join, three different ways to delete
a message the user had just sent: a leading content match that was a
coincidence, then a tail-anchored match that was a coincidence, then the same
class again by another path. The failure is silent by construction — the
shortened conversation passes every structural predicate and the strict
account-neutral projection accepts it.

Content equality cannot tell a restatement from a coincidence, so the
requirement stops asking it to. The shape of the client's input decides:
a continuation delta is appended to the chain, a full resend is the authority
and the chain is unused, and an input that is provably neither refuses. Content
comparison may verify a boundary the shape established; it may not discover one.

Refusing costs a conversation that could have been recovered, which is the
failure this change already has today. Guessing costs a message the user wrote.

Also pins two properties that survived mutation in round 3: the transcript byte
bound is a whole-transcript bound rather than a per-turn one, and a terminal
event reporting failure must not make a failed turn read as answered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): name the anchor as the signal that decides the join

The previous revision said to decide the join by shape rather than by content
equality, but did not say what establishes the shape. A fourth implementation
answered that with another content predicate — one that reads "the input holds
no model-authored item" as "this is a continuation delta" — and it classified a
full client resend as a delta, joined the durable chain to it, and doubled the
conversation at the public entry point. A client may legitimately restate a
history that contains no assistant turn.

previous_response_id is the signal, and it is definitional rather than
inferred: present means the input is that turn's delta, absent means the body
is the whole conversation. Two cases, no search, no third case.

Also adds two things the fourth round showed were missing: the rebuild must be
told which anchor it is rebuilding and must verify the chain terminates there,
because internally consistent parent links do not prove the chain belongs to
this request; and the shapes of the turns inside the chain are not a gate, since
requiring each to look like a delta disables relocation for any thread whose
first recorded turn was a resend.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): classify the join by self-containment, per stored request

Naming previous_response_id as the signal was wrong twice over, and the fifth
review measured both. The proxy injects anchors onto requests that did not
arrive with one, so the anchor on the wire proves nothing about the input. And
an anchored full resend is not a contradiction — it is a shape this repository
already verifies, in responses-api-compat :1214 and :2741. The rule defined that
shape out of existence, so the implementation joined the chain to it and doubled
the conversation in six measured forms, one of them spending the one-shot
recovery budget on a doubled body.

The discriminator was already in the module. Self-containment, decided per
stored request rather than once for the join: a self-contained history
supersedes what came before it, a delta that does not restate prior output
appends, a partial restatement refuses because its boundary is only recoverable
by matching content. Applying it to every chain turn as well as the client's
turn is what makes a parent turn that restated the conversation harmless — it
supersedes at its own position instead of being concatenated onto the history it
restates.

Also records the cycle guard the chain walk still needs, and drops the stale
overlap-dedupe instruction the task list kept after the spec forbade it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): join on the asymmetry, and stop classifying the client's intent

Six revisions of this join failed the same way, and five failed because this
change told the implementation to classify something that cannot be classified.
The sixth review's truth table made it plain: the discrimination came out
inverted, refusing the full resends Codex actually sends — developer-led,
tool-pair tail — while superseding the partial restatements the text said must
refuse, and in one measured shape silently dropping six items of a user's
conversation while satisfying "no client item lost, no turn twice".

The question has no answer. A client re-sending its last exchange plus a new
turn is byte-identical to one whose conversation began at that exchange.
previous_response_id cannot separate them because the proxy injects anchors
itself; "no model-authored items" cannot; and
responses_input_items_are_self_contained_fresh_replay cannot, because it answers
whether an input references state we do not hold, not whether it is complete.

The asymmetry can. The chain is our reconstruction; the client's input is what
the client is sending now. Where they overlap the client's copy is
authoritative, so the overlap comes off the chain, never off the client. A
chain turn the client just re-supplied is still dispatched — in the client's
words — so discarding it costs nothing, while discarding a client item costs a
message the user wrote, past every structural check. Both readings of an
ambiguous input then produce the same correct conversation, so no view about
intent is needed.

The overlap is anchored at the accumulated tail; a match in the middle of the
chain is a coincidence and must shorten nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): normalize for the comparison, and require it to be linear

The tail-overlap rule was right about what to discard and silent about two
things the seventh review measured.

The chain's items have been through the account-neutral projection and the
client's have not, so comparing them as they stand makes the overlap depend on
fields the projection normalizes away. One legal difference — an assistant
message omitting status, which the wire allows — collapsed the overlap to zero
and doubled an eight-turn conversation through the public entry point. The
comparison is now specified on the projected form of both sides, with the
client's verbatim items still the ones dispatched.

The search also has to be linear. The transcript caps bound turns and bytes but
not items, so a chain at 72% of the byte cap carries enough of them to cost
6.6 s of blocking CPU on a single-worker event loop, extrapolating to roughly
16 minutes at the cap — inside a failover path whose purpose is to be faster
than losing the conversation.

Also states what the anchor is still for: a request naming a prior response is
owed durable material and must fail closed without it, rather than dispatching
its own input as though it were the whole conversation. That is the one read of
previous_response_id the join permits, and it decides whether material is owed,
never what the input contains.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): identify items positively, and bound the item count

Two rounds built the join's comparison key by projecting an item and then
subtracting the fields a recording does not keep. Both lists were reasonable;
both were incomplete. The first missed status, the second missed phase and
internal_chat_message_metadata_passthrough, and each omission doubled a
four-turn conversation through the public entry point — the second on a field
carried by the branch's own production fixture.

Subtraction cannot be finished, because it requires knowing every field the wire
may carry that the recording may drop. The key must instead be built from a
positive enumeration of what identifies an item: a message's role and content, a
tool call's identity and arguments, a tool output's call and result. A field
nobody anticipated is then ignored rather than read as a difference.

Also bounds the item count. Turns and bytes do not bound items — the smallest
legal item repeated until the byte budget is spent yields six figures of them —
and items are what every per-item cost in the rebuild scales with, including the
canonicalization the linear search depends on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(openspec): make the identity enumeration recurse

The positive key fixed the defect at an item's top level and reproduced it one
level down. Round 9 identified additional_tools by serializing the whole tools
value, which makes every field inside a declaration identity-bearing: a client
that rewords a tool's description, changes a format, or adjusts a search
configuration between turns reads as a different turn and has its conversation
dispatched twice. Seven such fields were measured doubling on the unmutated
head.

It is the subtractive failure again. The reasoning that forbids a list of
fields to ignore equally forbids comparing raw anything the enumeration has not
reached, so the enumeration now applies to every nested structure the key
touches, and a structure it does not reach makes the item unidentifiable rather
than compared on its raw form — failing closed instead of guessing.

Also requires the enumeration to be checked against the item kinds the strict
predicate admits, so a kind added later is identified rather than compared by
accident.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…u06#2461)

Two revisions took the 20260914_000000 slot on different lineages, so the
graph forks and every job that migrates a database fails with MultipleHeads.
The topology check's authoring-time remedy does not apply once both ids are
published, so it now says so instead of advising a re-stamp that would orphan
alembic_version rows.

Co-authored-by: rvw <r@r.local>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ll timeouts (Soju06#2495)

Two timing hazards that fail on the runner rather than on the code under test.

`dashboard.spec.ts` measured document geometry immediately after
`page.setViewportSize`, which resolves when the resize is dispatched, not when
the browser has reflowed for it. Run 35314320645 on Soju06#2422 read a 1488px
scrollWidth against a 1440px viewport; the same commit passed on re-run.
Measurements now wait for the document width to stop moving, with a frame cap
so a page that truly overflows still fails its assertion.

`_wait_until_draining` polled the shutting-down fixture server with a 0.2s
client timeout and no exception guard, so any transport error or slow response
during SIGTERM handling raised out of the retry loop. Its sibling
`_wait_until_ready` already treats those as "not yet"; this makes the pair
symmetric.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
`tests/unit/test_native_egress_packaging.py` asserts on `Cargo.toml`,
`Cargo.lock`, `crates/**` and both container build files, but the `backend`
area filter matched none of them, so a pull request touching only those paths
satisfied the required pytest contexts with the placeholder step.

Soju06#2453 bumped the Cargo manifest and lockfile; its head `Tests (pytest, unit)`
context reported success in three seconds, the merge landed, and the same job
then failed on `main` in push run 35307892544. Soju06#2454 followed 52 minutes later
to re-pin the test.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

* chore(metadata): refresh model pricing and Codex version

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
…2479)

Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 4.3.0 to 4.4.0.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](docker/setup-qemu-action@1f40c72...9901266)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: 4.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…2481)

Bumps [docker/build-push-action](https://github.com/docker/build-push-action) from 7.3.0 to 7.4.0.
- [Release notes](https://github.com/docker/build-push-action/releases)
- [Commits](docker/build-push-action@53b7df9...c3c9e26)

---
updated-dependencies:
- dependency-name: docker/build-push-action
  dependency-version: 7.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
….2 (Soju06#2477)

Bumps [github/codeql-action/upload-sarif](https://github.com/github/codeql-action) from 4.38.0 to 4.38.2.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](github/codeql-action@b96794f...2892aa5)

---
updated-dependencies:
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [astral-sh/uv](https://github.com/astral-sh/uv) from 0.12.13 to 0.12.19.
- [Release notes](https://github.com/astral-sh/uv/releases)
- [Changelog](https://github.com/astral-sh/uv/blob/main/CHANGELOG.md)
- [Commits](astral-sh/uv@0.12.13...0.12.19)

---
updated-dependencies:
- dependency-name: astral-sh/uv
  dependency-version: 0.12.17
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…Soju06#2489)

* fix(accounts): preserve quota chart samples and account display state

* docs(accounts): illustrate missing quota sample regression

* fix(accounts): interpolate missing quota chart samples

* fix(accounts): merge quota observations by instant

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
…06#2468)

* fix(quota): validate planner timezones and tolerate legacy keys

* docs(openspec): record quota planner full verification

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
* fix(proxy): preserve routed file failover provenance

* docs(openspec): record routed file failover full verification

* fix(proxy): prevent file finalize replay after prior poll

* docs(openspec): archive verified file finalize replay guard

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
…2484)

* perf(usage): cap SQLite bulk usage history reads per account

* fix(usage): normalize SQLite history row caps

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
…of a pre-handshake close (Soju06#2255)

* fix(shutdown): deny WebSocket upgrades during drain with 503 instead of a pre-handshake close

`InFlightMiddleware` rejected new WebSocket connections during drain by
sending `websocket.close` (1013) before the handshake. ASGI servers
surface a pre-handshake close as a bare HTTP 403, so clients saw an
access error (`"WebSocket /backend-api/codex/responses" 403` in the
access log, dozens per restart on a single replica) and SDK-style
clients that only retry 5xx treated the drain as terminal.

Deny the upgrade with the same retryable 503 + Retry-After response the
overload bulkhead already uses (`deny_websocket_with_http_response`),
keeping the graceful-shutdown contract: the connection is rejected
without invoking the route handler and never counts as in flight.

* test(shutdown): cover generic drain denial and document contracts

* docs(openspec): record the drain-time WebSocket 503 denial contract

Soju06's review on Soju06#2255 asked for a focused OpenSpec change: the graceful-shutdown
spec only promised rejection without route invocation or in-flight growth, while the
fix changes the externally visible WebSocket drain rejection from a pre-accept close
(surfaced as HTTP 403) to an HTTP denial response with 503, Retry-After: 5, and the
proxy_unavailable envelope (or a generic detail body on non-proxy paths).

Validated with openspec validate deny-drain-websocket-upgrades-with-503 --strict and
openspec validate --specs.

---------

Co-authored-by: Abaddollyon <>
Co-authored-by: Soju06 <qlskssk@gmail.com>
)

A bare credits.has_credits=true flag with no positive balance kept an
account whose weekly window was exhausted active and selectable; every
request on it failed. Treat credits as usable only when unlimited or
balance > 0. apply_usage_quota already owns that rule for both the
account-summary mapper and proxy state derivation, so the callers are
unchanged and main's primary/secondary precedence order is preserved:
both windows exhausted without spendable credits stays quota_exceeded
with the secondary reset.

The OpenSpec MODIFIED block carries the complete requirement, including
the restored primary-precedence and operator-disabled scenarios, so the
strict change validation added by Soju06#2306 passes.

Co-authored-by: Soju06 <qlskssk@gmail.com>
* chore(db): add advisory migration benchmark for Soju06#1471

* test(db): keep benchmark checks valid across future migrations

* docs(db): document advisory benchmark contract and usage

* docs(openspec): archive verified advisory migration benchmark

* fix(db): protect benchmark input and retain step stamps

* docs(db): archive verified benchmark review repair

* test(db): verify benchmark failure-test role

* style(db): wrap benchmark role diagnostic

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
…u06#2513)

* fix(proxy): send a single Content-Type on codex control requests

Extracted from Soju06#2065 (commit a788898) so the fix can land on its own.
codex_control_request rewrote the inbound media type into a second
"Content-Type" field, so standalone search and realtime control calls
reached upstream with two content-type headers. Replace the header
case-insensitively and in place, and drop every spelling when the
request has no body.

Refs Soju06#2128

Co-authored-by: nhdong1993 <nhdong1993@users.noreply.github.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(contributors): add nhdong1993 to the roster

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: nhdong1993 <nhdong1993@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: rvw <r@r.local>
… reads it (Soju06#2460)

* fix(dashboard-users): take the owner row before the owned-key cascade reads it

On PostgreSQL a key could stay active on a disabled owner. The disable
cascade reads which keys are active and then writes them; the API-key
page's `isActive: true` reads the owner's status `FOR UPDATE` and then
writes the key. The two took their rows in opposite orders, so under READ
COMMITTED the cascade's scan could miss a key a PATCH turned on and
committed a moment later, and nothing re-checked either read. SQLite hid
it: its writers are serialised by the database write lock.

Both directions of the cascade now take the owner's row before their first
read, so the two writers agree on one order and whoever loses the row reads
the winner's committed state. The API-key gate moves ahead of the first
field assignment for the same reason: assigning a field first would
autoflush the key row's UPDATE before the owner lock and deadlock against
the cascade instead of refusing.

`reactivate_owner_disabled_keys` gets the same row, which closes the SCIM
`active: true` rejoin: it commits the status flip and restores the keys in
a second transaction, where the conditional UPDATE alone could still read
an owner a concurrent disable was about to turn off.

CI never saw it: `tests/integration/test_dashboard_users_api.py` was in no
PostgreSQL target, and the race passes on SQLite. Its two race tests and
the concurrent-admin one join `POSTGRES_PYTEST_TARGETS`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(dashboard-users): assert the end state after the delete/reactivate race

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: codex-lb-agent <r@r.local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…6#2480)

Bumps [docker/setup-buildx-action](https://github.com/docker/setup-buildx-action) from 4.3.0 to 4.4.1.
- [Release notes](https://github.com/docker/setup-buildx-action/releases)
- [Commits](docker/setup-buildx-action@37fe631...f87e599)

---
updated-dependencies:
- dependency-name: docker/setup-buildx-action
  dependency-version: 4.4.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* chore(ci): bump astral-sh/setup-uv from 10.0.1 to 10.2.0

Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 10.0.1 to 10.2.0.
- [Release notes](https://github.com/astral-sh/setup-uv/releases)
- [Commits](astral-sh/setup-uv@20cfd1b...c18668a)

---
updated-dependencies:
- dependency-name: astral-sh/setup-uv
  dependency-version: 10.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* test(ci): pin the setup-uv action SHA asserted by the workflow test to 10.2.0

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Soju06 <qlskssk@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(accounts): omit unsupported field from force probe

* chore(contributors): add tubededentifrice-dev to the roster

---------

Co-authored-by: Soju06 <qlskssk@gmail.com>
@maisi
maisi merged commit f5c3b99 into main Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.