Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
85 commits
Select commit Hold shift + click to select a range
b224704
feat(coref): port the co-reference index and component onto main
amiddavid Sep 1, 2026
24a8a9e
feat(extract_llm_sweep): fill the evidence seam and add the economic …
amiddavid Sep 1, 2026
a40f827
docs(coref): port the coref pages and the experiment record onto main
amiddavid Sep 1, 2026
292375e
test(coref): point the two skipped kept-verbatim tests at issue #175
amiddavid Sep 1, 2026
143cf73
docs(experiments): pre-register iteration 022 — three arms against a …
amiddavid Sep 1, 2026
913528a
docs(experiments): iteration 022 amendment 1 — the pre-flight caught …
amiddavid Sep 1, 2026
f724872
chore(harbor): commit iteration 022's rig — the band is a parameter, …
amiddavid Sep 1, 2026
bb80cee
docs(experiments): iteration 022 results — the mechanism works, the l…
amiddavid Sep 2, 2026
ab9d397
docs(experiments): iteration 022 — record that summarize never fired,…
amiddavid Sep 2, 2026
9714b95
docs(experiments): iteration 022 — arm B is confounded; withdraw the …
amiddavid Sep 2, 2026
a39413f
fix(extract_llm_sweep): give cg.sweep.ask a join key, and correct two…
amiddavid Sep 2, 2026
d2d005f
docs(experiments): record that the sweep's floor buys peers, not mass
amiddavid Sep 2, 2026
bd438f9
fix(extract_llm_sweep): price a drop on DEPTH, not size, and stop min…
amiddavid Sep 2, 2026
79bbab8
feat(harbor): iteration 023 — extract_llm unpinned, sweep floor at 100
amiddavid Sep 2, 2026
81a6e56
docs(experiments): iteration 023 results — coverage holds, the confou…
amiddavid Sep 2, 2026
d7ae8a1
docs(experiments): correct iteration 023 — the confound framing and t…
amiddavid Sep 2, 2026
abe8786
docs(experiments): iteration 023 — the pooled counters are root-cause…
amiddavid Sep 2, 2026
be79553
Merge origin/main into feat/coref-recut — take #178's per-component a…
amiddavid Sep 2, 2026
cba9e7c
docs(experiments): pre-register iteration 024 — cost and latency, mea…
amiddavid Sep 2, 2026
026e2a7
docs(experiments): iteration 024 — the first significant reward resul…
amiddavid Sep 3, 2026
952700c
docs(experiments): iteration 024 — withdraw the recovery-channel mech…
amiddavid Sep 3, 2026
779303b
docs(experiments): iteration 024 — the latency finding is confounded …
amiddavid Sep 3, 2026
9b0c6e0
docs(experiments): iteration 024 — the recovery channel is explained,…
amiddavid Sep 3, 2026
bda7a93
docs(experiments): iteration 024 — the recovery channel may not be fr…
amiddavid Sep 3, 2026
ae2e5ed
docs(experiments): withdraw an invalid inference — unmarked does not …
amiddavid Sep 3, 2026
3241705
docs(experiments): cross-reference #199, and note the sizing needs a …
amiddavid Sep 3, 2026
cad6b44
docs(experiments): cite the counter-splitting rule rather than re-arg…
amiddavid Sep 3, 2026
93f2763
docs(experiments): the 9,478/3,522 expansion figures are floors, not …
amiddavid Sep 3, 2026
794baef
docs(experiments): measure the expansion gap properly instead of hedg…
amiddavid Sep 3, 2026
7f68857
docs(experiments): withdraw the combined 1.68 — summing a clean gate …
amiddavid Sep 3, 2026
d62962f
Merge origin/main into feat/coref-recut — and fix a dangling-marker b…
amiddavid Sep 8, 2026
2285a5f
fix(sweep): charge the adjudication to the econ trigger that authoris…
amiddavid Sep 8, 2026
e7ffa4e
docs(sweep): explain which cost declines the econ trigger, and why no…
amiddavid Sep 8, 2026
5d58575
feat(sweep): record WHERE in the window each trigger decision was taken
amiddavid Sep 8, 2026
d44e07f
docs(iter025): pre-register the run before any number beyond the pre-…
amiddavid Sep 8, 2026
1b68ff7
docs(iter025): amendment 1 — reuse iteration 024's arm A for four of …
amiddavid Sep 8, 2026
d184f83
fix(iter025): verify the frozen binary before every pass, not once at…
amiddavid Sep 8, 2026
c63499c
test(iter025): the drift check, and its power limit recorded before i…
amiddavid Sep 8, 2026
1a87395
fix(sweep_when): print the over-window bucket instead of dropping it
amiddavid Sep 8, 2026
aaeec2f
test(iter025): an errored baseline task is unpairable, not a zero
amiddavid Sep 8, 2026
1393794
docs(iter025): results — the reward result does not replicate through…
amiddavid Sep 8, 2026
1ca9784
docs(iter025): record that the 64k band was not held, and correct 568…
amiddavid Sep 8, 2026
26af194
docs(iter025): correct the headline — 55% of decisions were taken whe…
amiddavid Sep 8, 2026
dfb0f65
docs(iter025): withdraw the reward conclusion — the treatment did not…
amiddavid Sep 8, 2026
d91ec6b
docs(iter025): the clearing was configured correctly and still did no…
amiddavid Sep 8, 2026
8279582
fix(sweep): put the econ decision on the session-aware logger
amiddavid Sep 8, 2026
5007edc
fix(sweep): pass the session id explicitly on the econ row
amiddavid Sep 8, 2026
4889840
docs(limits): the benchmark's compaction never fired — the gateway do…
amiddavid Sep 9, 2026
ff829a2
docs(sweep): define the econ trigger's vocabulary, and what the compo…
amiddavid Sep 9, 2026
85f2e05
feat(sweep): two opportunity floors, so the ask ledger stops calibrat…
amiddavid Sep 9, 2026
ff772a1
feat(iter026): the deferral corridor — sweep from 0.70, summarize at …
amiddavid Sep 9, 2026
fd77dc9
docs(summarize): the trigger reads the POST-pipeline request, not the…
amiddavid Sep 9, 2026
69e3e03
fix(summarize): count the checkpoint-reuse path — closes the uncounte…
amiddavid Sep 9, 2026
a310cc9
docs(summarize): summary_covered_span_changed is a CHURN signal, not …
amiddavid Sep 9, 2026
8568bce
docs(limits): clusters give the power, seeds make each cluster trustw…
amiddavid Sep 9, 2026
7a08c1f
docs(iter026): pre-register the deferral run — interleaved, with chec…
amiddavid Sep 9, 2026
52bc346
feat(sweep): price the break-even for what a removal delivers, and st…
amiddavid Sep 10, 2026
386ebcc
docs(iter027): amend the probe's fixtures — request size, not step co…
amiddavid Sep 11, 2026
ec6177b
fix(iter027): make the benchmark pipeline the one being claimed, and …
amiddavid Sep 11, 2026
fe11344
feat(sweep): stop re-buying a verdict already given, and stop asking …
amiddavid Sep 11, 2026
7854f2d
fix(adjudicate): ask whether the CONTENTS are needed again, and tell …
amiddavid Sep 14, 2026
08a48c9
feat(cg-selarm): count every way a batch fails to produce a judgement
amiddavid Sep 14, 2026
1ffdf2f
fix(extract): accept a newline-delimited verdict reply instead of dis…
amiddavid Sep 14, 2026
8ee9e47
test(offload): simulate the econ gate by band, and ask the harness th…
amiddavid Sep 14, 2026
f980bc5
fix(offload): price approval on mass and per inventory regime, not on…
amiddavid Sep 14, 2026
dad2130
test(cg-selarm): select decision points by inventory size
amiddavid Sep 14, 2026
5026234
docs(loca): iteration 027 results, and the ground-truth class that ca…
amiddavid Sep 14, 2026
8d250b3
docs(loca): correct iteration 027 §1 — the result does belong to the …
amiddavid Sep 14, 2026
f091eb9
test(offload): simulate the proposed configuration end to end
amiddavid Sep 14, 2026
affea91
docs(loca): preregister iteration 028 — does reward track removal vol…
amiddavid Sep 14, 2026
a9b8586
docs(loca): amend iteration 028 to two arms with a per-seed futility …
amiddavid Sep 15, 2026
3f1fbcb
feat(harbor): iteration 028 runner — one seed per invocation, then stop
amiddavid Sep 15, 2026
33ff068
chore(harbor): pin the iteration 028 binary hash in the runner
amiddavid Sep 15, 2026
eedf539
fix(harbor): make run028's skip path say what to do with an invalid pass
amiddavid Sep 15, 2026
81ec2dd
fix(offload): credit the removal in the pruner's horizon too (#232, s…
amiddavid Sep 16, 2026
35cae5c
feat(harbor): declare iteration 028 at 128k, not 64k
amiddavid Sep 16, 2026
6054895
docs(loca): amend iteration 028 for the 128k band and the credited pr…
amiddavid Sep 16, 2026
2bb5926
feat(harbor): make the agent's model an env var in stage022
amiddavid Sep 16, 2026
4ef4ea8
feat(harbor): add a baseline-only step and a per-environment readout
amiddavid Sep 16, 2026
f29467b
fix(harbor): resolve a pass by its own log, and separate a killed epi…
amiddavid Sep 16, 2026
67cecf9
feat(harbor): bound the agent's thinking, and stop reaching for a big…
amiddavid Sep 16, 2026
49dc2cb
docs(loca): stop iteration 028, and say what enabling the sweep on re…
amiddavid Sep 16, 2026
4a874e4
Merge origin/main into feat/coref-recut
amiddavid Sep 16, 2026
2fa664e
style: gofmt two files the lint gate rejected
amiddavid Sep 16, 2026
d972078
Merge remote-tracking branch 'origin/main' into feat/coref-recut
amiddavid Sep 17, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
511 changes: 511 additions & 0 deletions cmd/cg-selarm/main.go

Large diffs are not rendered by default.

72 changes: 72 additions & 0 deletions cmd/cg-selarm/main_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
package main

import "testing"

// Every case below is a reply shape MEASURED on this corpus or on the gateway, not an invented one.
// The taxonomy exists because the first live run of this tool folded all four into "no verdicts" and
// reported failures=0 while 17 of 30 batches never produced a judgement.
func TestClassifyReplySeparatesTheFourOutcomes(t *testing.T) {
for _, tc := range []struct {
name string
reply string
candidates int
want string
wantN int
}{{
name: "a complete verdict array is the only measurement",
reply: `[{"i":0,"needed_by":"none","verdict":"drop"},{"i":1,"needed_by":"none","verdict":"keep"}]`,
candidates: 2, want: replyOK, wantN: 2,
}, {
// The shape that produced the 94-row file: prose, no array at all. Remedy is the prompt.
name: "prose with no array is unparseable, not truncated",
reply: "I have considered the outputs but will not produce a verdict array here.",
candidates: 3, want: replyUnparseable,
}, {
// Remedy is the opposite -- raise max_tokens -- so it must not be reported as unparseable.
name: "an array opened and never closed is truncated",
reply: `[{"i":0,"needed_by":"none","verdict":"drop"},{"i":1,"needed_by":"a","quote":"I will run the su`,
candidates: 3, want: replyTruncated,
}, {
// The QUIET failure: it parses, so nothing errors, but the omitted candidates would read as
// deliberate keeps in a scorer that defaults absent to keep.
name: "a short but valid array is partial",
reply: `[{"i":0,"needed_by":"none","verdict":"drop"}]`,
candidates: 6, want: replyPartial, wantN: 1,
}, {
// A deliberate keep-all. It MUST NOT read as junk: ParseVerdicts documents this distinction
// as the one a live arm turned on.
name: "a closed empty array is a decision, not a failure",
reply: `[]`,
candidates: 0, want: replyOK,
}} {
t.Run(tc.name, func(t *testing.T) {
vs, kind := classifyReply(tc.reply, tc.candidates)
if kind != tc.want {
t.Errorf("kind = %q, want %q", kind, tc.want)
}
if len(vs) != tc.wantN {
t.Errorf("verdicts = %d, want %d", len(vs), tc.wantN)
}
})
}
}

// An unparseable reply must yield NO verdicts, so the caller cannot write decisions derived from a
// reply it already classified as junk.
//
// HONEST LIMIT ON THIS TEST: it cannot fail against today's code, and a mutation run confirmed that
// -- changing `return nil, replyUnparseable` to `return vs, replyUnparseable` leaves every test
// green, because `ParseVerdicts` itself returns `nil, false` on failure, so the two lines are
// identical. It is kept as a contract assertion on that callee: if ParseVerdicts is ever changed to
// salvage a partial slice from a reply it rejects, this is what trips. Read it as documentation of a
// dependency, NOT as evidence that classifyReply drops verdicts on its own.
func TestAnUnparseableReplyYieldsNoVerdicts(t *testing.T) {
for _, r := range []string{
"no array here",
`[{"i":0,"needed_by":"none","verdict":"dr`,
} {
if vs, kind := classifyReply(r, 4); len(vs) != 0 {
t.Errorf("%q classified %s but returned %d verdicts", r, kind, len(vs))
}
}
}
30 changes: 26 additions & 4 deletions cmd/context-guru-proxy/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -1188,10 +1188,32 @@ func modelWindows() modelinfo.Resolver {
}
return chain
}
return append(chain,
modelinfo.NewLiteLLM(os.Getenv("MODEL_INFO_URL"), nil, 0),
modelinfo.DefaultStatic(),
)
url := strings.TrimSpace(os.Getenv("MODEL_INFO_URL"))
live := modelinfo.NewLiteLLM(url, nil, 0)
if url == "" {
// The public map. A fetch failure here is a network inconvenience, not a misconfiguration, so
// the embedded table remains the right answer and the chain keeps its fallback.
return append(chain, live, modelinfo.DefaultStatic())
}
// AN EXPLICIT MODEL_INFO_URL IS AN AUTHORITY THE OPERATOR NAMED, and guessing past it is wrong for
// the same reason MODEL_PRICES is fatal rather than skipped a few lines up: a silently-absent
// document is indistinguishable from one that says something different, and every fraction-based
// threshold in the pipeline would then be evaluated against a window nobody chose. Probed
// SYNCHRONOUSLY here because the resolver's own fetch is a background refresh — by the time it
// fails, requests are already being served against DefaultStatic's 1,000,000 for a claude model.
//
// This is the check that would have stopped iteration 024. Its rig passed a correct 64k document on
// a URL the proxy could not reach; the proxy resolved 1,000,000 on all 2,207 requests, summarize's
// 0.78 trigger became 780,000 and never fired once in either arm, the econ trigger's horizon came
// out 16x too long and authorised 626 asks, and the run completed with every counter healthy. Six
// hours of benchmark time and $207 per arm bought numbers that described no configuration.
if err := live.Load(context.Background()); err != nil {
log.Fatalf("MODEL_INFO_URL %s: no context window could be resolved from it (%v). Refusing to "+
"start: every fraction-based trigger would be evaluated against a built-in default instead "+
"of the document you configured, and nothing downstream can tell the difference.", url, err)
}
// The fallback stays BEHIND the probed document, for models the document does not list.
return append(chain, live, modelinfo.DefaultStatic())
}

// cheapModelFromEnv builds the static "config"-source LLM client for NeedsModel
Expand Down
15 changes: 15 additions & 0 deletions components/offload/commitgate_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,21 @@ func gateCases() []gateCase {
body := strings.Repeat("a line of log output that goes on for a while\n", 80)
return c.(components.Offload), toolMsgs(body, body+"tail\n"), ctxFor(st)
}},
{"coref", func(t *testing.T, st store.Store) (components.Offload, *bschemas.BifrostChatRequest, *components.Ctx) {
// Driven rather than exempted, because coref DID have the defect this table exists to
// catch: it ignored commitMark's return and spliced plus froze regardless.
//
// min_later_turns AND min_batch_frac are both defaulted out of the way, the same as
// corefFor does. Leaving the opportunity floor at its default of 8 protects this
// fixture's tail candidate, nothing is cut, and the saturated case below would then
// pass vacuously -- which this table's own precondition catches.
c, err := newCoref([]byte("min_tokens: 20\nmin_batch_frac: 0\nmin_later_turns: 0\n" +
"break_even: false\n"))
if err != nil {
t.Fatal(err)
}
return c.(components.Offload), corefReq(), corefCtx(st)
}},
{"dedup", func(t *testing.T, st store.Store) (components.Offload, *bschemas.BifrostChatRequest, *components.Ctx) {
c, err := newDedup([]byte("min_tokens: 20\n"))
if err != nil {
Expand Down
Loading
Loading