Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 74 additions & 0 deletions docs/handoffs/2026-09-09-the-full-cohort-through-the-loop.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# The full senktide and TTX cohorts through the loop — the pointer

> **Written straight into this directory on 2026-09-09 — nothing in it is half-done.** The
> run is complete and its artifacts are in the darkroom. It was never a root signal.
>
> **This file carries no number derived from the recordings** (FOUNDATIONS §5: anything
> derived from real data stays machine-local). Every measured result, every figure and the
> full run record are in `<darkroom>/bugarach/2026-09-09-full-cohort-senktide-ttx/`;
> resolve `<darkroom>` with `bugarach.paths.darkroom()`, or `python -m bugarach.paths`.
>
> **Not murderboarded**, and neither is the run record it points at — both are working
> material for sessions in this tree, the same standing as `docs/pipeline.md`. The one
> artifact addressed to another team, `for_fireflies/README.md`, is **marked as a draft and
> must not be sent** until it has been through `/murderboard` and Tony has released it.

Successor to [the APV+CNQX+GZ pilot](2026-09-07-the-pilot-cohort-through-the-loop.md),
whose MAHICE and slow-bench caveats are **not** superseded by this run.

## What is different from the pilot, and it is most of what mattered

Both of `pipeline.md`'s blockers closed on 2026-09-08 (#507, #508, #509), and this is the
first run to spend them:

| the pilot could not | this run did |
|---|---|
| apply tuned settings from the command line — the six ran at shipped operating points | detected at a calibration derived from this cohort's own simulated data |
| run a trained model on real data — nothing persisted a model | ran all six learned models from checkpoints, in a separate process from training |
| fit background heterogeneity — six baselines were too few | fitted it from the folder; **burstiness is now the last inherited generator quantity** |
| scan for field-step artifacts — the folder was UNCHECKED | read the producer's artifact-excluded export |

## What the run found before it started

**The producer's newest export had been on disk for six days and no file in this repo named
it.** Its own README says *"for any new analysis, use this folder."* Landed as **#511**,
which declares it as three roles beside `default` rather than instead of it — moving
`default` re-points every existing analysis and every parity fixture, and that is a
decision, not a housekeeping edit.

That PR also carries the three things this run needed and did not have: a **floor under a
percentage K**, a `derive_spec` path that **aggregates across recordings at each one's own
resolved K** instead of selecting one column of the scan, and **group facets** on the
before/after figure. Two silent defects turned up in existing code on the way — a
hardcoded clip width that dropped the rightmost facet of any wide page, and a
positional-tuple read across a module boundary.

## Still open, in the order I would take it

1. **MAHICE has still never been run on an approved folder.** Skipped here on instruction,
as in the pilot. Everything downstream inherits `RESET.md` §1 — a coordination number
nobody looked at is not a result. Expert attention, not compute.
2. **An edge-of-grid threshold refuses on the coded branch and only warns on the learned
one**, and it decided what the best learned model shipped at here.
[The todo](../todo/2026-09-09-an-edge-of-grid-threshold-refuses-on-one-branch-and-warns-on-the-other.md)
names the one measurement that settles it.
3. **Six folds cannot support a corrected pairwise winner** — that is arithmetic, not a
property of the detectors, and the threshold is eight. Relevant to
[`performance_table.md`](../performance_table.md) §1, whose replacement argument is still
unwritten.
4. **Burstiness is still inherited.** `pipeline.md` says fit it from the user's own
baseline; rate heterogeneity now is, and this one is not.
5. **No slow bench.** Every operating point is fast-derived and applied to a stream whose
measured coactivity is several times higher. Unchanged from the pilot.

## Traps this run hit

- **zsh does not word-split an unquoted parameter.** Building `--model a --model b` into a
shell variable and passing it unquoted sends **one** argument; argparse reports every
flag as unrecognised and the leading empty string is the tell. Write the flags out.
- **A recording missing from a detections file is drawn at zero**, where it cannot be told
from a detector that ran and found nothing. The before/after figure now counts and names
them on the page; it caught a `--limit` smoke run immediately.
- **A wide page loses its right-hand column silently.** Fixed in the shared renderer, but
the lesson generalises: the clip was a pixel constant copied from the viewport it was
meant to track.
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
---
status: open
filed: 2026-09-09
---

# An optimum at the edge of its grid refuses on the coded branch and only warns on the learned one

> **Not murderboarded** — working material for sessions in this tree. **If any of it
> reaches an outside reader, murderboard that artifact first.**

Found while running the full senktide + TTX cohorts through the loop, where it decided
what the best-performing learned model shipped at.

## The same condition, two verdicts

This project has a written position on an optimum that lands on the boundary of the
grid it was searched over. `bench.EdgeOfRange`'s own docstring states it:

> *"Not a warning. An optimum at the edge is not an optimum — it is the search telling
> you it stopped too early, and reporting it as a calibrated point is how a boundary
> value once got published upstream as one."*

`pick_operating_point` raises it, `tools/refit.py` reports it as an outcome, and
`tools/settings_from_bakeoff.py` will not emit such a value without `--allow-edge`,
which stamps the file to say so. That is the coded branch.

On the learned branch the same condition is a `RuntimeWarning` from
`learn/train.py`, the fit continues, the checkpoint is written, and the model runs on
real recordings. `learned_settings.csv` does record `threshold_at_grid_edge`, which is
the honest half — the value is knowable. Nothing acts on it.

## What it cost here, concretely

On this cohort's own generator spec, `tube` is the best learned model and second overall
on the bake-off (F1 0.678, fold range 0.665–0.716). Its fitted threshold is **0.9999,
the top of the searched grid `[0.0001, 0.9999]`**, flagged `threshold_at_grid_edge: yes`.
It was saved to a checkpoint and run on all 67 recordings.

So the best learned detector in this run is deployed at a threshold the coded branch
would have refused to publish, and the refusal that exists to catch exactly this never
fired because it is not wired to this branch.

`trace` and `tiny` are at the other edge (0.0001) and also flagged, but that is their
known degenerate-control behaviour rather than a search that stopped early — 78 calls
across 29 recordings, one span per analysis window. Two different problems wearing one
flag, which is its own reason to want a verdict rather than a boolean.

## Why this is a decision and not a patch

Making the learned branch refuse is a one-line change and it is **not obviously right**:

- The threshold grid is a probability in `(0, 1)`, so "widen it" means going to
`1 - 1e-5` and further, and there is always another decimal place. The coded knobs have
natural ranges; this one does not, so an edge here may mean *"the model separates
cleanly and any high threshold works"* rather than *"the search stopped too early"*.
Those two want opposite responses.
- `tube` at 0.9999 has a **probe rate of 1.37/min**, inside its own three-run history
(0.75 pilot, 2.05 published). It is not behaving like a detector at a runaway operating
point, which is what an edge-of-grid value would normally predict.
- A refusal would have blocked this run's learned half entirely, and the learned half is
what `#508`/`#509` existed to make reachable.

So the question is not "should it refuse" but **which of the two meanings an edge has
for a probability threshold**, and that is answerable by measurement: sweep `tube`'s
threshold past 0.9999 on the bench and see whether F1 is flat (separates cleanly) or
still climbing (stopped too early). Nobody has.

## What to do

1. **Measure it.** Extend the threshold grid for the tube variants past 0.9999 and
report whether F1 is flat or climbing there. One probe, and it settles the question.
2. **Then decide the verdict**, and make it the same kind of object on both branches —
a refusal, or a stamped-and-allowed value like `--allow-edge`, not a warning nothing
reads.
3. **Separate the two flags.** `threshold_at_grid_edge` currently means both "search
stopped early" and "this control is degenerate". Those need different names.

Related: [`2026-09-08-the-ratio-tube-cannot-count-cells.md`](2026-09-08-the-ratio-tube-cannot-count-cells.md)
— the ratio arms read 0.00/min on the probe by arithmetic rather than by merit, which is
the other place a learned model's number means less than it looks like it means.
Loading