Three documents argue from the flat field, and the grep that found the third is the rule we wrote after Kreuz - #505
Merged
Conversation
…e third is the rule we wrote after Kreuz
Another session reported two: `performance_table.md` §1 argues from the superseded
flat-field result, and `learned/bakeoff.md` says the pooled trace has "three of its four
folds at the floor" where its own JSON says one. Both confirmed against the artifacts.
`78ebe26` re-measured the background axis on the FITTED field at twelve seeds: one detector
leads at all seven grid points, largest rank change two. The same seeds on the flat field
still give three winners and a change of three, which is how the artefact was identified
rather than assumed -- and `tests/test_background_curve.py` asserts both curves side by side
so it cannot drift back. So the repo's own tests contradict what three of its documents say.
THE THIRD WAS NOT REPORTED. `forks.md`'s background-curve fork says CoactDetect "goes from
first to fifth across the grid" -- the four-place move `78ebe26` names as the flat field's.
It turned up by grepping the tree for the same shape before closing the report out, which is
CLAUDE.md's own rule from the Kreuz incident, where "that was one grep and nobody ran it".
First time the rule has been used since it was written, and it paid.
WHAT IS FIXED AND WHAT DELIBERATELY IS NOT. The bakeoff prose is corrected outright: one
fold, and the thresholds are printed so the next reader can check without opening the JSON.
The other two get headers naming the superseding measurement, and NOTHING is re-derived --
restating stale numbers as current is the failure the headers exist to stop.
The separation that matters is in forks.md. That fork is about SPREAD, not order: mean
own-range fell 0.185 -> 0.136 and the gate still refuses for all six, so the retraction takes
the ranking claim and leaves the finding standing. `performance_table.md` §1 is the opposite
case -- its entire "why there is no ranking" argument was the winner-swapping, and the
replacement argument ("no ranking because the spread is large") is a different one that
nobody has written.
Also flagged: `learned/background_curve.png` is the flat field's, its Panel B draws a
crossing that no longer reproduces, and it has not been regenerated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Another session reported two defects. Both confirmed against the artifacts, and a third
found by grepping for the same shape.
78ebe26re-measured the background axis on the fitted field at twelve seeds: onedetector leads at all seven grid points, largest rank change two. The same seeds on the
flat field still give three winners and a change of three — which is how the artefact was
identified rather than assumed, and
tests/test_background_curve.pyasserts both curvesside by side. So the repo's own tests contradict what three of its documents say.
performance_table.md§1forks.md, background-curve forklearned/bakeoff.mdlearned/background_curve.pngThe third was not reported.
forks.mdturned up by grepping the tree for the same shapebefore closing the report out — CLAUDE.md's own rule from the Kreuz incident, where "that
was one
grepand nobody ran it". First use since it was written.What is deliberately not done. Nothing is re-derived.
performance_table.md§1'sconclusion may well survive, but for a different reason: the axis narrowed rather than went
dead (mean own-range 0.185 → 0.136, still several times the 0.017 gap the bake-off asks
readers to believe). "No ranking because the spread is large" is a different argument from
"no ranking because the winner swaps", and nobody has written it.
The separation matters most in
forks.md: that fork is about spread, not order, so theretraction takes the ranking claim and leaves the finding.
🤖 Generated with Claude Code