diff --git a/docs/handoffs/2026-09-09-the-full-cohort-through-the-loop.md b/docs/handoffs/2026-09-09-the-full-cohort-through-the-loop.md new file mode 100644 index 00000000..c5330c91 --- /dev/null +++ b/docs/handoffs/2026-09-09-the-full-cohort-through-the-loop.md @@ -0,0 +1,74 @@ +# The full senktide and TTX cohorts through the loop — the pointer + +> **Written straight into this directory on 2026-09-09 — nothing in it is half-done.** The +> run is complete and its artifacts are in the darkroom. It was never a root signal. +> +> **This file carries no number derived from the recordings** (FOUNDATIONS §5: anything +> derived from real data stays machine-local). Every measured result, every figure and the +> full run record are in `/bugarach/2026-09-09-full-cohort-senktide-ttx/`; +> resolve `` with `bugarach.paths.darkroom()`, or `python -m bugarach.paths`. +> +> **Not murderboarded**, and neither is the run record it points at — both are working +> material for sessions in this tree, the same standing as `docs/pipeline.md`. The one +> artifact addressed to another team, `for_fireflies/README.md`, is **marked as a draft and +> must not be sent** until it has been through `/murderboard` and Tony has released it. + +Successor to [the APV+CNQX+GZ pilot](2026-09-07-the-pilot-cohort-through-the-loop.md), +whose MAHICE and slow-bench caveats are **not** superseded by this run. + +## What is different from the pilot, and it is most of what mattered + +Both of `pipeline.md`'s blockers closed on 2026-09-08 (#507, #508, #509), and this is the +first run to spend them: + +| the pilot could not | this run did | +|---|---| +| apply tuned settings from the command line — the six ran at shipped operating points | detected at a calibration derived from this cohort's own simulated data | +| run a trained model on real data — nothing persisted a model | ran all six learned models from checkpoints, in a separate process from training | +| fit background heterogeneity — six baselines were too few | fitted it from the folder; **burstiness is now the last inherited generator quantity** | +| scan for field-step artifacts — the folder was UNCHECKED | read the producer's artifact-excluded export | + +## What the run found before it started + +**The producer's newest export had been on disk for six days and no file in this repo named +it.** Its own README says *"for any new analysis, use this folder."* Landed as **#511**, +which declares it as three roles beside `default` rather than instead of it — moving +`default` re-points every existing analysis and every parity fixture, and that is a +decision, not a housekeeping edit. + +That PR also carries the three things this run needed and did not have: a **floor under a +percentage K**, a `derive_spec` path that **aggregates across recordings at each one's own +resolved K** instead of selecting one column of the scan, and **group facets** on the +before/after figure. Two silent defects turned up in existing code on the way — a +hardcoded clip width that dropped the rightmost facet of any wide page, and a +positional-tuple read across a module boundary. + +## Still open, in the order I would take it + +1. **MAHICE has still never been run on an approved folder.** Skipped here on instruction, + as in the pilot. Everything downstream inherits `RESET.md` §1 — a coordination number + nobody looked at is not a result. Expert attention, not compute. +2. **An edge-of-grid threshold refuses on the coded branch and only warns on the learned + one**, and it decided what the best learned model shipped at here. + [The todo](../todo/2026-09-09-an-edge-of-grid-threshold-refuses-on-one-branch-and-warns-on-the-other.md) + names the one measurement that settles it. +3. **Six folds cannot support a corrected pairwise winner** — that is arithmetic, not a + property of the detectors, and the threshold is eight. Relevant to + [`performance_table.md`](../performance_table.md) §1, whose replacement argument is still + unwritten. +4. **Burstiness is still inherited.** `pipeline.md` says fit it from the user's own + baseline; rate heterogeneity now is, and this one is not. +5. **No slow bench.** Every operating point is fast-derived and applied to a stream whose + measured coactivity is several times higher. Unchanged from the pilot. + +## Traps this run hit + +- **zsh does not word-split an unquoted parameter.** Building `--model a --model b` into a + shell variable and passing it unquoted sends **one** argument; argparse reports every + flag as unrecognised and the leading empty string is the tell. Write the flags out. +- **A recording missing from a detections file is drawn at zero**, where it cannot be told + from a detector that ran and found nothing. The before/after figure now counts and names + them on the page; it caught a `--limit` smoke run immediately. +- **A wide page loses its right-hand column silently.** Fixed in the shared renderer, but + the lesson generalises: the clip was a pixel constant copied from the viewport it was + meant to track. diff --git a/docs/todo/2026-09-09-an-edge-of-grid-threshold-refuses-on-one-branch-and-warns-on-the-other.md b/docs/todo/2026-09-09-an-edge-of-grid-threshold-refuses-on-one-branch-and-warns-on-the-other.md new file mode 100644 index 00000000..4dfc9b15 --- /dev/null +++ b/docs/todo/2026-09-09-an-edge-of-grid-threshold-refuses-on-one-branch-and-warns-on-the-other.md @@ -0,0 +1,80 @@ +--- +status: open +filed: 2026-09-09 +--- + +# An optimum at the edge of its grid refuses on the coded branch and only warns on the learned one + +> **Not murderboarded** — working material for sessions in this tree. **If any of it +> reaches an outside reader, murderboard that artifact first.** + +Found while running the full senktide + TTX cohorts through the loop, where it decided +what the best-performing learned model shipped at. + +## The same condition, two verdicts + +This project has a written position on an optimum that lands on the boundary of the +grid it was searched over. `bench.EdgeOfRange`'s own docstring states it: + +> *"Not a warning. An optimum at the edge is not an optimum — it is the search telling +> you it stopped too early, and reporting it as a calibrated point is how a boundary +> value once got published upstream as one."* + +`pick_operating_point` raises it, `tools/refit.py` reports it as an outcome, and +`tools/settings_from_bakeoff.py` will not emit such a value without `--allow-edge`, +which stamps the file to say so. That is the coded branch. + +On the learned branch the same condition is a `RuntimeWarning` from +`learn/train.py`, the fit continues, the checkpoint is written, and the model runs on +real recordings. `learned_settings.csv` does record `threshold_at_grid_edge`, which is +the honest half — the value is knowable. Nothing acts on it. + +## What it cost here, concretely + +On this cohort's own generator spec, `tube` is the best learned model and second overall +on the bake-off (F1 0.678, fold range 0.665–0.716). Its fitted threshold is **0.9999, +the top of the searched grid `[0.0001, 0.9999]`**, flagged `threshold_at_grid_edge: yes`. +It was saved to a checkpoint and run on all 67 recordings. + +So the best learned detector in this run is deployed at a threshold the coded branch +would have refused to publish, and the refusal that exists to catch exactly this never +fired because it is not wired to this branch. + +`trace` and `tiny` are at the other edge (0.0001) and also flagged, but that is their +known degenerate-control behaviour rather than a search that stopped early — 78 calls +across 29 recordings, one span per analysis window. Two different problems wearing one +flag, which is its own reason to want a verdict rather than a boolean. + +## Why this is a decision and not a patch + +Making the learned branch refuse is a one-line change and it is **not obviously right**: + +- The threshold grid is a probability in `(0, 1)`, so "widen it" means going to + `1 - 1e-5` and further, and there is always another decimal place. The coded knobs have + natural ranges; this one does not, so an edge here may mean *"the model separates + cleanly and any high threshold works"* rather than *"the search stopped too early"*. + Those two want opposite responses. +- `tube` at 0.9999 has a **probe rate of 1.37/min**, inside its own three-run history + (0.75 pilot, 2.05 published). It is not behaving like a detector at a runaway operating + point, which is what an edge-of-grid value would normally predict. +- A refusal would have blocked this run's learned half entirely, and the learned half is + what `#508`/`#509` existed to make reachable. + +So the question is not "should it refuse" but **which of the two meanings an edge has +for a probability threshold**, and that is answerable by measurement: sweep `tube`'s +threshold past 0.9999 on the bench and see whether F1 is flat (separates cleanly) or +still climbing (stopped too early). Nobody has. + +## What to do + +1. **Measure it.** Extend the threshold grid for the tube variants past 0.9999 and + report whether F1 is flat or climbing there. One probe, and it settles the question. +2. **Then decide the verdict**, and make it the same kind of object on both branches — + a refusal, or a stamped-and-allowed value like `--allow-edge`, not a warning nothing + reads. +3. **Separate the two flags.** `threshold_at_grid_edge` currently means both "search + stopped early" and "this control is degenerate". Those need different names. + +Related: [`2026-09-08-the-ratio-tube-cannot-count-cells.md`](2026-09-08-the-ratio-tube-cannot-count-cells.md) +— the ratio arms read 0.00/min on the probe by arithmetic rather than by merit, which is +the other place a learned model's number means less than it looks like it means.