Skip to content

chore(upstream): integrate beta6 while preserving fork features - #31

Merged
maisi merged 23 commits into
mainfrom
chore/sync-upstream-beta6
Sep 9, 2026
Merged

maisi merged 23 commits into
mainfrom
chore/sync-upstream-beta6

Conversation

@maisi

@maisi maisi commented Sep 9, 2026 •

Copy link
Copy Markdown
Owner

Integrates the 20 upstream commits after PR #29 through c0beaaadd96a89f0240582b5449bf4dd50647c7d, including beta6, permanent report aggregates, dashboard routing controls, HTTP/WebSocket promotion, and native HTTP stream completion.

Preserves token vending, per-key account ranking, prompt-cache continuation, forced usage, cache reports, warmup behavior, and fork release/contribution policies. A merge migration joins upstream and deployed fork history. The lightweight report options endpoint retains historical API-key choices after raw-log pruning.

This is the explicitly requested upstream integration concern; upstream ancestry is preserved with a merge commit. Issue #27 remains in focused PR #30.

Validation: 1,831 routing/proxy unit tests, 47 report API/aggregation tests, 38 migration tests, 330 native SSE wire tests and 24 report-page tests passed locally. Rust workspace tests, frontend build, Python lint/type checks and 64 strict OpenSpec specs passed. The complete GitHub implementation-head CI passed, including PostgreSQL, all backend shards, frontend coverage, browser smoke, Rust, Docker, Helm/kind, Nix and packaging. Final documentation-head checks are required before merge.

No new fork environment settings. Upstream moves behavior tunables to dashboard settings and removes obsolete internal settings; the four existing token-vending settings retain their documented topology/secret tiers. The settings budget follows upstream reductions plus those four fork settings.

OpenSpec: openspec/changes/archive/2026-09-09-integrate-upstream-v1-25-beta6/.

Before (beta5):
Beta5 routing settings

After (beta6, new dashboard routing controls):
Beta6 routing settings

Soju06 and others added 22 commits September 9, 2026 17:07
…e dashboard (Soju06#2224)

Squash of f7d3cd9e5, 2faf2df85, cf38d55ff, 3556a9638 (PR Soju06#2224) ahead of the rebase onto Soju06#2220/Soju06#2221.
…ju06#2227)

* perf(reports): serve historical reports from permanent aggregates

* test(reports): cover dedicated options in date-range flows

* fix(reports): isolate cache hits from slow computations
* fix(proxy): restore trusted Codex routing hints

Build subscription hints from normalized request state across HTTP, WebSocket, bridge reconnects, and client-request compaction. Preserve inbound-header isolation and actual response tiers.

Adapted from minpeter's routing-hint changes at addaac0, with compaction propagation and main-based regression coverage.

* test(proxy): accept routing context in websocket observability fixture
… them through (Soju06#2240)

A code-less upstream HTTP 429 (per-account burst/concurrency rejection whose
body carries only a message) normalized to upstream_error and was treated as
a plain transient failure: the account stayed fully selectable, sticky routing
sent client retries straight back to it, and for owner-bound payloads (codex
multi-turn requests carrying reasoning items) the failover_next decision
could never be executed, so the original 429 was surfaced immediately.

- record a replica-local burst cooldown (RuntimeState.burst_backoff_until,
  5s default, upstream Retry-After honored up to 30s) at the
  _handle_stream_error funnel; fresh/unbound selection steers away through
  the existing overload soft-backoff predicate, established owners are kept,
  persisted account status is never touched
- failover_decision learns owner_bound: owner-bound requests never report
  failover_next; a burst 429 retries the same owner after 1s/2s/4s waits
  (Retry-After as floor, 10s cap) while the HTTP startup wait keeps holding
  headers, then surfaces the original 429 with Retry-After
- _wait_for_first_stream_probe re-arms its post-ready window on a newer wait
  marker so consecutive bounded waits keep the header hold
- OpenSpec change backoff-upstream-burst-429-on-http-stream

Fixes Soju06#2239

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
…pose statements (Soju06#2228)

Twenty-two capability specs under openspec/specs still opened with the
archive tool's placeholder ("TBD - created by archiving change ..." or
the manual-sync variant). Each Purpose is now a short statement of what
the capability governs and why it exists, written from the originating
archived change proposal and the spec's current requirements.

No requirement, scenario, or behaviour text changes.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…deprecation (Soju06#2230)

* docs(helm): document the pre-1.13 StatefulSet migration shim and its deprecation

Follow-up to Soju06#2211, which kept the Helm legacy-deployment hooks. The chart
has carried a three-piece shim since 1.13.0 (Soju06#363) for releases first
installed on the old Deployment controller: the legacy-deployment
prepare/cleanup hook Jobs, the lookup-based Service selector auto mode and
migration.serviceSelectorMode. The chart README only mentioned the
workload-name split in one bullet.

Add an "Upgrading" section to deploy/helm/codex-lb/README.md:

- From chart versions older than 1.13.0: what the two hooks do, the
  selector-mode knob, kubectl verification, and the caveats (hooks run as
  no-ops on every upgrade, do not force `workload` before cutover,
  client-side dry runs cannot lookup the Service).
- Deprecation statement: planned removal of the shim in the first minor
  release after 1.26, announced here, with the intermediate-upgrade path
  for releases still on a pre-1.13 chart.
- Upgrading across 1.24 -> 1.25: the env vars the chart no longer
  templates (CODEX_LB_UPSTREAM_STREAM_TRANSPORT via Soju06#2192,
  CODEX_LB_OPENAI_CACHE_AFFINITY_MAX_AGE_SECONDS and the other
  Soju06#2190 removals) with their dashboard replacements, and the two chart
  values (config.upstreamConnectTimeout, config.circuitBreakerEnabled)
  whose env vars became deprecated dashboard aliases in Soju06#2220/Soju06#2221/Soju06#2224.

values.yaml gains the deprecation note on migration.serviceSelectorMode and
alias notes on the two config values. No template or default changes; no
OpenSpec requirement covers the chart upgrade path, so this is docs-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(helm): correct shim caveats, fullname-keyed verification, and record the deprecation in OpenSpec

Review follow-ups on the Upgrading section:

- The shim does not deliver a zero-downtime cutover: the chart no longer
  renders the legacy Deployment and nothing marks it resource-policy keep, so
  Helm removes it during the resource sync before the post-upgrade hook.
  Say so and tell operators to plan a maintenance window.
- migration.serviceSelectorMode=legacy only controls the rendered selector;
  the cleanup hook patches the live Service to workload regardless.
- Verification commands use the chart fullname (release `codex-lb` renders
  `codex-lb`, not `codex-lb-codex-lb`), with the concrete names spelled out.
- Only the _ENVIRONMENT_INHERITABLE_SETTINGS aliases log the "environment
  value(s) ignored because the dashboard owns the setting" WARN; the three
  resilience toggles are overridden silently. Scope the sentence accordingly.
- "for at least one release" wording for the removed-setting WARN, and the
  full Settings -> Advanced -> Routing path for the stream transport.
- New OpenSpec change document-helm-pre-1-13-shim-deprecation recording the
  compatibility commitment (shim retained through 1.26, removal announced,
  removal is its own change) as an ADDED deployment-installation requirement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…Soju06#2238)

* test: make http-bridge file-affinity tests provision their own schema (Soju06#2209)

Two tests in tests/unit/test_proxy_http_bridge.py passed only when an
earlier module had already reset the shared test database:

- ..._fails_closed_before_file_affinity_when_previous_response_owner_misses
  writes a file pin through the real SessionLocal and failed alone with
  "no such table: file_account_pins". It now requests the existing
  db_setup fixture from tests/conftest.py.
- ..._waiter_propagates_terminal_inflight_proxy_error let the
  registration gate query ring membership against the unprovisioned
  database and timed out its 0.1s wait_for. It now pins
  _http_bridge_should_wait_for_registration closed like the sibling
  inflight tests; the waiter path under test never needed the DB.

Also fence the leaked audit-log tasks reported in Soju06#2213's verification:
tests/unit/test_dashboard_auth_password_service.py drives the real
verify_password/verify_guest_password/verify_totp paths, whose
fire-and-forget AuditService.log_async tasks could never complete
against the in-memory fakes and outlived the test on the session loop,
so test_otel's lifespan drain test found a non-empty _AUDIT_LOG_TASKS
before it started. The module now stubs log_async, and tests/conftest.py
gains a sync autouse fence (same shape as the live-usage ingestor
reaper) that cancels and drops any audit task a test leaves pending.

Fixes Soju06#2209

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: describe the audit-log leak mechanism precisely in the login stub fixture

The leaked audit-log-login_* task is scheduled and then never stepped
because the test returns synchronously; it is not blocked on a missing
schema. Reword the fixture docstring to state the actual mechanism.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…round paths (Soju06#2233)

* fix(routing): resolve soft-drain from the dashboard snapshot on background paths

Follow-up to Soju06#2220 / Soju06#2224: the quota planner tick and forecast endpoint and
the usage-refresh recovery reconciliation built account states without a
dashboard-settings snapshot, so their health tier inherited
CODEX_LB_SOFT_DRAIN_ENABLED and the routing tunables came from the
environment while the dashboard said otherwise.

- QuotaPlannerScheduler._run_once_as_leader takes one SettingsCache snapshot
  per tick (before its session, outside any runtime lock) and
  GET /api/quota-planner/forecast one per request; both pass
  effective_routing_tunables(snapshot) and the resolved soft_drain_enabled
  into _build_states.
- reconcile_recoverable_account_statuses accepts the dashboard_settings row
  the refresh cycle already read and resolves both once;
  background_recovery_state_from_account forwards them to _state_from_account.
- openspec change background-paths-read-dashboard-snapshot modifies the two
  account-routing requirements that documented the background env fallback.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(quota-planner): take the forecast dashboard snapshot before the first query

A settings-cache refresh opens its own session; reading the snapshot after
the forecast queries made the request hold a pooled connection while a
second checkout waited (bounded-pool timeout). Read it first, like the
scheduler tick, and pin the order in the forecast integration test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ion settings (Soju06#2231)

The Data retention card rendered its own "Not configured: effective N days"
hint while every other inheritable setting on the page already uses the shared
InheritBadge / useInheritableSetting affordance. The backend reports
provenance for request_log_retention_days and usage_history_retention_days
(source default | dashboard), so the card now renders `Default (0)` while a
window is unconfigured and `Reset to inherited` (tri-state PUT with an
explicit null for that window only) while the dashboard owns it. The tri-state
input, the 0 = disabled semantics and the 30/45-day floor messages are
unchanged; the bespoke hint and its `settings.retention.inheritedHint` string
are removed from en, ko and zh-CN.

configuration-tiers: the temporary "retention MAY keep its effective-value
hint" allowance is removed (openspec/changes/retention-inherit-badge).

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…tion suites (Soju06#2243)

The recurring CI failure (issue Soju06#1949) is not lock contention in the
startup sentinel stamps: every failing test took exactly the 30 s
busy_timeout and logged event_loop_lag=30 s. An async test that already
holds the async_client lifespan entered starlette's blocking TestClient
inline; TestClient.__enter__ froze the session loop, which owns the first
lifespan's background SQLite writers. A write those writers had in flight
(reproduced: the cache-invalidation poller flushing the bumps queued by the
test's own setup mutations, INSERT executed, COMMIT never dispatched) could
not release the writer slot, so the portal lifespan's startup INSERTs waited
out busy_timeout and failed. Retries or BEGIN IMMEDIATE cannot help because
the holder is frozen by the caller waiting on it.

- tests/integration/off_loop_test_client.py: `async with
  off_loop_test_client(app)` flushes the loop-owned deferred writers and
  enters/exits the blocking client from a worker thread so the loop keeps
  draining the holder.
- test_proxy_realtime_live.py: the four async tests at the CI failure sites
  use the helper; bodies unchanged.
- tests/conftest.py: autouse guard makes TestClient.__enter__ on a running
  loop fail immediately with the explanation instead of flaking after 30 s.
- tests/integration/test_off_loop_test_client.py: pins startup completing
  while a loop-owned write transaction is open, the pending-bump flush, the
  guard, and sync-test compatibility.

Refs Soju06#1949

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Soju06#2229)

Archive the four OpenSpec changes whose PRs have already merged, re-run on
top of current main (a97ffc3, after Soju06#2227/Soju06#2240/Soju06#2242/Soju06#2228) so the
merged spec text reflects the specs as they stand today:

- dashboard-managed-resilience-toggles (Soju06#2220)
- dashboard-managed-upstream-timeouts (Soju06#2221)
- dashboard-managed-routing-overload (Soju06#2224)
- split-release-guards-from-ci-matrix (Soju06#2212)

`openspec list` now shows only the changes still in flight
(add-subscription-overflow-model-source, add-request-log-usage-rollups,
backoff-upstream-burst-429-on-http-stream from Soju06#2240).

split-release-guards-from-ci-matrix task 3.3 is ticked with the Actions
evidence gathered after Soju06#2212 merged.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…Soju06#2247)

* fix(proxy): promote agentic HTTP continuations to reusable WebSockets

* fix(proxy): avoid counting redundant bridge outage bypasses

* test(proxy): cover Chat trailing slash capability policy
…ptured at run start (Soju06#2241)

* fix(automations): pin the stale-claim reclaim window to the budget captured at run start

The stale-claim reclaim window (compact budget + 30 s) was computed from the
compact request budget in effect at reclaim time. Since the budget became a
live dashboard setting (Soju06#2221), lowering it while a run is in flight shrank
the window under that run and let the scheduler (or another replica) reclaim
it and start a second compact ping for the same slot.

Every claim and reclaim now stores the effective compact budget on the run
row (automation_runs.claim_budget_seconds, nullable). Staleness is judged per
run from the stored value on every path (due-cycle discovery, scheduled
reclaim, due manual-run discovery, manual reclaim) and the run's compact
request executes with the same stored budget, so the two can never disagree.
Rows claimed before the column existed (NULL) keep using the current budget.

Migration 20260909_060000_automation_run_claim_budget adds the column on
sqlite and PostgreSQL; downgrade drops it.

Deferred from Soju06#2221 (codex R2 P2).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(automations): widen the pinned reclaim window to the current budget and page due manual runs

- Re-chain the migration after 20260909_060000_add_report_rollup (Soju06#2227) as
  20260909_070000_automation_run_claim_budget so the graph keeps a single head.
- The reclaim window now covers max(pinned, current) + 30 s: lowering the
  dashboard budget still cannot shrink the window under an in-flight run, and a
  raised budget covers an attempt started by a pre-pin writer that advanced
  started_at without refreshing the pin (rolling deploy). Review finding.
- list_due_manual_runs pages candidates until `limit` eligible rows are
  collected so an in-flight claim inside its own window never consumes a batch
  slot owed to a due row behind it. Review finding.
- Spec delta: compact request timeout bounded by (not equal to) the captured
  budget; new scenario for a raised budget over a lower pin.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…u06#2245)

Move recognized native HTTP Responses stream termination into Rust. Flush completed, failed, and incomplete terminal events, release the upstream body without waiting for EOF, and let Python retire the exchange without a cancel round trip.

Negotiate the completion capability, validate final-fragment metadata, preserve cancellation and shared-helper isolation, and document the lifecycle boundary in OpenSpec. Regression coverage includes direct/routed and SDK/native streams with fragmented terminals and invalid trailing bytes.
… lifespan or unregistering the outer poller (Soju06#2254)

Follow-up to Soju06#2243. Three defects in tests/integration/off_loop_test_client.py:

- Nested-lifespan startup could hang forever. The app's process-global
  anyio.Lock singletons (settings cache, sqlite writer section, ...) are
  shared by the test loop and the portal loop once both run. anyio hands
  lock ownership to the next waiter by set_result() on that waiter's future
  from the releasing task's thread; a release on the test loop for a waiter
  on the portal loop is queued through a non-thread-safe call_soon that the
  portal loop, asleep in select with no timers yet, never runs. Captured
  hung: the ring heartbeat blocked in settings_cache.get() -> Lock.acquire,
  the portal lifespan suspended in startup, both loops idle. The nested
  lifespan now runs behind KeepPortalLoopAwake, an ASGI pass-through with a
  50 ms ticker task on the portal loop, so queued wake-ups run within a tick.
- The nested lifespan clears the cache-invalidation poller registration on
  shutdown, leaving the outer lifespan's poller running but unregistered;
  the helper restores it (codex review P2).
- Cancelling the awaiting task during startup left the worker thread to
  finish TestClient.__enter__ and leak the portal lifespan; the startup call
  is shielded, awaited to completion and exited before the cancellation
  propagates (codex review P2).

Contract tests cover all three (the ticker test uses a bounded dummy
lifespan so a regression fails a timing assertion instead of hanging).

Refs Soju06#2243
Refs Soju06#1949

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ju06#2250)

* test: stop gating background loops through CODEX_LB_*_ENABLED env in the test and smoke harness

tests/conftest.py and scripts/run_dashboard_browser_smoke.py disabled the app
lifespan's background loops by exporting CODEX_LB_*_ENABLED=false. Upcoming
settings constantization would silently turn those exports into no-ops and
start the real loops on the first test lifespan tick (SQLite lock flakes,
no_op planner decision rows, public GitHub/npm catalog lookups).

Replace the exports with an in-process seam: one autouse fixture
(_disable_background_loop_schedulers) swaps every app.main build_*_scheduler
for a no-op factory, folding the four per-loop fixtures and the duplicate
reset-credits patch in app_instance, and additionally covering the auth
guardian. The smoke backend applies the same builder list plus a no-op
live-usage ingestor before uvicorn imports the app. A new unit test pins the
builder tuple against the names app.main imports (patched set plus the three
always-on maintenance loops), so a new loop must be classified or the suite
fails, and asserts neither harness exports the retired env names.

CODEX_LB_USAGE_REFRESH_ENABLED=false stays in tests/conftest.py because the
field also gates request-path refreshes (account import, usage_limit_reached),
which otherwise cost 20-30 s per imported account against the unreachable
test upstream; CODEX_LB_HTTP_RESPONSES_SESSION_BRIDGE_ENABLED=false stays
because the bridge is a request-path feature whose field is retained as a T4
kill switch. The ops.md runbook command no longer relies on the env prefix.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: guard the retained usage-refresh env export against field removal

Review follow-ups for the harness seam PR:

- tests/unit/test_background_loop_harness.py: add a hand-off guard that fails
  the moment usage_refresh_enabled leaves Settings while tests/conftest.py
  still exports CODEX_LB_USAGE_REFRESH_ENABLED (Settings ignores unknown env
  names, so the export would otherwise become a silent no-op and every
  account-importing test would pay the 20-30 s example.invalid refresh).
  Document why the env-kill-switch assertion checks the live process
  environment and point shell users at the fix.
- scripts/run_dashboard_browser_smoke.py: the disabled live-usage path also
  resets the hub publisher to None; say the no-op is equivalent in a fresh
  backend process rather than an exact match.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…metry fallback (Soju06#2252)

* chore(settings): retier topology-bound env settings and document telemetry fallback

Move proxy_unauthenticated_client_cidrs, dashboard_trust_loopback_host_header_for_long_sessions
and proxy_response_create_limit from T3 to T1 in SETTING_TIERS and drop their MIGRATING rows:
they are per-replica socket/network-namespace facts and a per-process asyncio.Semaphore
capacity (same class as bulkhead_proxy_limit), so they were never dashboard migration
candidates.

telemetry_enabled already has its dashboard home (dashboard_settings.telemetry_consent,
persisted > env > default since Soju06#2186). Add the DASHBOARD_HOMES registry so a T3 field
can declare a table.column home the checker verifies against the SQLAlchemy metadata,
give telemetry_enabled that mapping, drop its MIGRATING row and fix the docs/spec wording
that still described the env variable as a seed/override.

No Settings field added, removed or renamed; no runtime behaviour change.
docs/reference/settings.md regenerated (Tier column only).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): warn instead of fail on a stale DASHBOARD_HOMES entry whose column is gone

check_t3_dashboard_home validated the table.column target before checking
whether the field still exists, so removing a Settings field together with
its column produced an error instead of the stale-entry warning the
configuration-tiers spec requires (removals may land in either order).
Check staleness first; the target only matters for a live field.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(settings): scope the live DASHBOARD_HOMES assertions to fields that still exist

A field removed before its DASHBOARD_HOMES entry is a checker warning, not
an error; the live-tree test must not turn that staggered-removal state
into a CI failure.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(settings): classify redundant DASHBOARD_HOMES entries before validating their target

An entry made redundant by a same-name column or a re-tier is due for
deletion, so a malformed or dropped target must warn like the stale-field
case instead of failing make lint mid-cleanup.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(settings): constantize never-tuned core tunables

Turn 27 never-tuned MIGRATING settings into module-level constants equal
to their previous defaults (issue Soju06#1340 precedent, slop-removal 0908 K1):
upstream SSE/websocket frame and response.create budgets, OAuth and
token-refresh timeouts, the refresh claim TTL (now a code helper), the
refresh-failure cooldown, the admission wait and the token-refresh /
websocket-connect / compact gates, usage fetch timeout and retries, the
usage refresh interval with its derived freshness horizon, the usage
auth-failure cooldown, the reset-credits polling interval, the always-on
usage refresh / live ingestion / sticky cleanup / model registry / quota
planner switches (the planner keeps the dashboard mode "off" as its only
switch), the HTTP ingress body budgets (shared with --ws-max-size), inline
image fetching and its never-populated host allowlist, the public default
image model and prompt-cache-key derivation.

The env names join _REMOVED_SETTINGS for their one-release warning; every
test injection seam is kept (module constants are monkeypatched, scheduler
dataclasses and WorkAdmissionController keep their kwargs) and the request
path usage refresh is neutralised in tests by a fixture instead of the
removed env toggle. Helm stops rendering the two removed keys, the settings
reference is regenerated and the settings-field budget drops 130 -> 103.

token_refresh_interval_days stays: the fast canary suite sets
CODEX_LB_TOKEN_REFRESH_INTERVAL_DAYS in its failure-matrix environment.

OpenSpec: constantize-core-tunables.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(settings): cover the fixed claim TTL and wording in the constantize-core-tunables deltas

Add MODIFIED deltas for the usage-refresh-policy claim-serialization
requirement (the claim TTL is the fixed helper, so the derive-or-reject
settings scenario becomes a fixed-TTL scenario), outbound-http-clients
(fixed 16 MiB event cap) and proxy-runtime-observability (fixed 10 s
admission wait); note them in the proposal and tasks.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(settings): rebase constantize-core-tunables onto main after K0

Port the tests Soju06#2245 and Soju06#2247 added on main that vary the upstream SSE
event budget through Settings (a silent no-op once the field is a
constant) to the MAX_SSE_EVENT_BYTES seam the rest of the suite already
patches; the promotion payload-size bypass patches the constant to 3 MiB
so the derived 1 MiB websocket payload budget trips again.

Refresh the deployment-installation delta onto main's current requirement
text (Soju06#2229 archived C2: resilience toggle alias scenario and the
dashboard-owned resilience switches paragraph) and list the two removed
chart values in the Helm README "Upgrading across 1.24 -> 1.25" table
(Soju06#2230).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…les (Soju06#2256)

* refactor(settings): constantize never-tuned HTTP session bridge tunables

Seven HTTP session bridge tunables were never tuned in any deployment, yet
each cost a Settings field, tier and MIGRATING rows, a docs row, a Helm key
and a getattr(..., default) seam at every read site. Replace them with fixed
module constants (PRINCIPLES.md P2) and list the env names in
_REMOVED_SETTINGS for the one-release WARN:

- idle_ttl_seconds -> helpers.HTTP_BRIDGE_IDLE_TTL_SECONDS (120s)
- codex_idle_ttl_seconds -> helpers.HTTP_BRIDGE_CODEX_IDLE_TTL_SECONDS (900s)
- stuck_gate_retire_after_seconds -> helpers.HTTP_BRIDGE_STUCK_GATE_RETIRE_AFTER_SECONDS (300s)
- anchor_poison_failure_threshold -> retry_circuit._HTTP_BRIDGE_RETRY_CIRCUIT_FAILURE_THRESHOLD (2);
  the value was already clamped to the circuit threshold at all 13 read
  sites, so the capping helper and the configured_threshold parameter go
- server_recovery_max_attempts -> api.HTTP_BRIDGE_SERVER_RECOVERY_MAX_ATTEMPTS (6)
- clean_close_retry_jitter_max_seconds -> request_submit._HTTP_BRIDGE_CLEAN_CLOSE_RETRY_JITTER_MAX_SECONDS (2.0s)
- operation_ledger_enabled -> always on; the three gate branches are deleted

http_responses_session_bridge_enabled is retiered T3 -> T4 (kill switch,
not a tunable) and leaves MIGRATING. The timeout invariant
"2x stuck gate < bridge budget" reads the constant through _expr. Helm drops
the two idle-TTL keys and injects pod identity env unconditionally.
docs/reference/settings.md is regenerated; [settings_fields].max 130 -> 123.
Tests keep every injection seam by monkeypatching the module constants.

OpenSpec: openspec/changes/constantize-session-bridge-tunables

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* refactor(proxy): describe the fixed anchor-poison threshold in the bridge spec delta

The MODIFIED delta still promised to honour a configured anchor-poison
threshold below the circuit threshold, which no longer exists. Reword the
five clauses to the fixed circuit threshold and drop the now-unreachable
below-threshold quarantine re-check in the strike path (the opening branch
already arms it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(openspec): state the K2 operator-visible behaviour changes and re-state the anchor-poison requirement at the fixed threshold

Review follow-ups after rebasing onto main (K0 Soju06#2250, T Soju06#2252, K1 Soju06#2261):

- proposal.md Impact now says explicitly what changes for operators who had
  overridden one of the three values whose override used to change behaviour:
  ANCHOR_POISON_FAILURE_THRESHOLD=1 now poisons on the second strike,
  CLEAN_CLOSE_RETRY_JITTER_MAX_SECONDS=0 can no longer disable jitter, and
  OPERATION_LEDGER_ENABLED=false no longer exists as a kill switch.
- The "Repeated zero-event idle failures poison dead anchors" requirement
  still spoke of a configured threshold and had a scenario whose GIVEN
  ("configured poison threshold is greater than two") is now impossible.
  openspec archive refuses a MODIFIED block that drops a scenario, so the
  requirement is REMOVED and re-ADDED as "... at the circuit threshold",
  based on the current main text, with that scenario replaced by the
  circuit-open probe scenario. Verified with a simulated
  `openspec archive --yes` on a copy of main's openspec/: +2 ~3 -2, no
  stale wording left, archived specs validate --strict.
- Helm README 1.24 -> 1.25 Upgrading table gains rows for the two removed
  idle-TTL chart values / configmap keys and the five extraEnv-only names.
- Settings surface after the rebase is 96 (103 - 7); the proposal and tasks
  say so instead of 130 -> 123.
- api.py: `settings` in `_http_bridge_recovery_request_eligible` became
  unused once K1 constantized the payload budget; drop the read.
- P3: one non-default monkeypatch test for HTTP_BRIDGE_IDLE_TTL_SECONDS
  proving the runtime config reads the constant at call time.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
@maisi
maisi marked this pull request as ready for review September 9, 2026 21:18
@maisi
maisi merged commit a5e4de0 into main Sep 9, 2026
34 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants