New post (scheduled 2026-08-26): what an AI code reviewer was told to ignore - #613
Merged
Conversation
… ignore Founder/ICP-E stream, arrival-purposed, 1,410 words. The calendar picked the stream and the date before the topic existed: the last 7 days held 11 posts, all Rails Technical, while ICP-E had been quiet since 08-08. The thesis survived being wrong twice, which is the only reason it is worth publishing. **First version.** "An AI reviewer is weakest where AI code is weakest." Then Uber's published numbers arrived: only 51% of their HUMAN review comments are judged valid and addressed, against over 65% for uReview, with 75% rated useful across >90% of ~65,000 weekly diffs. On that evidence AI review is not the weak link - human review is. The post now leads with that, because a reader who thinks this is an anti-AI piece stops reading. **What the post actually turns on** is Xiong & Zhang (ISSTA 2026): filtering takes SAST false positives from >92% down to 6.3%, and the same configuration "incorrectly suppresses 22.25% of real vulnerabilities", with miss rates that "exceed 50% for cryptography- and policy-related categories". The tuning that makes an AI reviewer bearable to developers is what drops the findings a founder would care most about. **Second correction, caught by the claim gate.** The draft said a Tencent result was "close to the opposite" of that finding. Reading to the end of that paper: "recall remains below the enterprise-expected threshold of 90%, indicating that some degree of manual review is still necessary". Two teams, opposite-looking headlines, same conclusion - do not run it unattended. I had manufactured a tension that the sources do not support; the convergence is the better point and is now what the section says. Every figure was re-fetched at its primary before it was written down. Two reached the outline through search excerpts and both changed on contact: the per-CWE miss rates I had (77.17%, 84.50%) come from a secondary catalog, not the paper, so the post quotes the paper's own ">50%" instead. Also in this commit, from the same session's corrections: - `blog-next` gains Stage A0 (calendar/spacing/stream-and-stack balance before topic selection), a fourth exit SCHEDULED (a same-week collision is a spacing problem, not a verdict - it previously produced DO-NOT-WRITE and lost the topic), three sources (Reddit via WebSearch, X, changelogs), and a MANDATORY invocation of `blog-write` rather than a recommendation to run it. - Reddit's .json endpoint is documented as BLOCKED (403 + HTML block page; old.reddit 302), verified, so nobody scripts it. - Stage B now fires NotebookLM deep research FIRST and in parallel; mirrored into blog-pipeline.md. `research_import` documented as mandatory - without it `notebook_query` answers from an empty notebook. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures (fabrication + uncitedness ratchets; the post carries 5 citations). Zero banned words, zero em dashes. Both internal links confirmed to resolve in the built output rather than trusted from a sibling post. NOT run, and the post should not be merged as if they were: no independent verifier saw this draft. Agent spawning was unavailable, so the four-lens cold-eyes panel did not run and the gates above are self-administered. No cover image yet. `rake test:links` cannot validate this post at all while it is future-dated - production skips future content - so quoting a green link run here would be evidence of nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The draft shipped at 1,410 words with zero images, which trips the cognitive-load rule for anything over 800 words. Measured before adding anything: a 216-word unbroken prose run in the section carrying the Uber comparison, and no visual break across seven H2s. Both diagrams carry an argument rather than decorating one, and both sit where the brick was: - `reviewers.svg` breaks the 216-word run at the counter-intuitive fact - 51% of human review comments get acted on against over 65% for uReview. That is the number a scanning reader should hit first, because it is the one that stops the post reading as anti-AI. - `suppressed.svg` carries the finding the post turns on: the filter takes false alarms from over 92% to 6.3%, and takes 22.25% of real vulnerabilities with it, over 50% in crypto and policy categories. Labels sit inside the diagram. Longest run is now 154 words, down from 216. Both SVGs use the house palette (#0e0e14 ground, ruby #cc342d, purple #a855f7) and carry title/desc for screen readers. Cover rendered to the `.stitch/design.md` layout at the specified 2400x1260 and exported with rsvg-convert - the PNG is the artifact, per the standing rule. Looked at the render rather than trusting it: the first pass had the year pill text overflowing its rounded rect on both sides, so the pill went 300px -> 420px and it was re-rendered and checked again. Hugo picked it up and generated the responsive webp/jpg variants. Also fixed in the prose while placing the second image: "Weak encryption. Weak password hashing. Trust boundaries." was impersonal fragment stacking AND a rule of three, both banned in 90.11. It is one sentence now. Verified served, not assumed: reviewers.svg 200, suppressed.svg 200, cover.png 200. The `<img` grep said the SVGs were missing from the built page and was WRONG - Hugo wraps the tag, so the filenames were checked in the HTML instead. That is the second time that grep has lied this session. Gates: bin/hugo-build green. Still no independent reviewer and no browser screenshot pass at 1280x800 / 390x844 - the four-criteria new-media gate has not been scored by anyone but me. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
Paul's row, verbatim: "AI should generate code, but verify - NOPE!" Rails/eng audience, 1,165 words, scheduled a week clear of the 08-26 post because theme recency - not spacing - was the binding constraint: four of the previous six days' posts already touched AI verification. Thesis: the advice sounds like a division of labour and is a transfer. The word doing the damage is "then", which puts checking second, and things that come second sound smaller. THE RESEARCH CORRECTED THE POST TWICE. 1. The GitClear figures I was handed were WRONG. Secondary write-ups say churn "doubled" 3.3% -> 7.1%, duplication "4x", 211M lines. GitClear's own 2026 research says two-week churn is +15% against 2022, block duplication +81% over 2023 (40.3 -> 73.0 per million changed lines), across 623 million analyzed changes. The second URL those write-ups cite 404s. Every one of those numbers would have shipped wrong. 2. The real finding is better than the garbled one and is now the post's centre: moved code - GitClear's proxy for reorganising work - was 21% of changed lines in 2022, 13% in 2023, and 3.8% year-to-date in 2026. Duplication you can find later; the habit is harder to restart. Load-bearing claim verified at primary before drafting: Harness State of Engineering Excellence 2026, "81% say developers spend more time in code review since adopting AI coding tools, with 28% reporting a significant increase of more than 30%" - 700 practitioners and managers, five countries, Sapio Research, April 2026. The post states the sample and invites the reader to discount it. METR was BANNED from this post and is absent: it appears 5x in what-senior-developers-catch-that-ai-misses, and reusing it is the cross-post repetition an ICP reviewer called "reading one post three times". The automation-trap section deliberately does NOT re-argue the 08-26 post - it links once and moves on. A "63% of developers spent more time debugging AI code" stat was dropped entirely: its attribution chain runs through a vendor blog to three reports without landing in any of them. Visuals, because the last post shipped with none: refactoring.svg carries the 21% -> 3.8% collapse, placed to break what was a text brick. Cover built from the house template and rendered with rsvg-convert at the 2400x1260 spec, then looked at. Two self-review defects fixed before commit: a 208-word unbroken run, and a "What we do instead" section written as four parallel bold-led items, which is the banned bold-inline-header shape. Both were the same section; it is now a decision table plus prose. Longest run 208 -> 165. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. Zero banned words, zero em dashes. Section endings varied - no repeated aphorism landing. Both internal links and both assets confirmed present in the built output. NOT run: no independent verifier. Agent spawning unavailable, so the four-lens cold-eyes panel did not run and every gate above is self-administered. No browser pass at 1280x800 / 390x844. rake test:links cannot validate a future-dated post, so it was not run and would prove nothing about this post if it had been. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
…3-59% Paul's row: "setup highly-autonomus team of agents is real already today, but how many got success?" It turns out to have a measured answer, which is why it earned a slot. Cemri et al. (arXiv 2503.13657, NeurIPS 2025 D&B; Zaharia, Gonzalez and Stoica among the authors) published success rates across six multi-agent frameworks: AG2 59.0%, MetaGPT 40.0%, Magentic-One 38.0%, ChatDev 33.3%, HyperAgent 25.3%, AppWorld 13.3%. The post prints their figure caption alongside the table - "Performances are measured on different benchmarks, therefore they are not directly comparable" - because presenting that spread as a league table is the obvious way to misuse it. The counterweight is quoted at equal length rather than buried: Anthropic's multi-agent research system "outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval", at about 15x the tokens of a chat, with token usage alone explaining 80% of the variance. And their own sentence on why it suits research and not our day job: "most coding tasks involve fewer truly parallelizable tasks than research". The story is the convergence. Cognition published "Don't Build Multi-Agents" in June 2025 and, ten months later, "Multi-Agents: What's Actually Working", conceding a narrower class: "setups where multiple agents contribute intelligence to a task while writes stay single-threaded." That is Anthropic's read-only-subagent design in different vocabulary. Many agents may think, one agent writes - and writers.svg carries exactly that contrast. Trace count deliberately absent: sources disagree (1,600+ annotated / 150 rigorous / 200 in the body) and none was confirmed in the PDF, so the post writes around it rather than picking one. Dedup held: multi-agent-llm-rails-rubyllm (08-20) shares the vocabulary and not the subject - it is the Rails product feature, this is the engineering workflow. Linked once for implementation, never re-argued. A MEASUREMENT ERROR OF MINE, FOUND AND RECORDED. I had been flagging "208-word text bricks" using a metric I invented - words between structural elements. The canonical rule in voice-rules.md is a PARAGRAPH over ~700 source chars, and that file already warns that a tighter threshold "flagged normal prose". Measured properly, the longest paragraphs in these three posts are 361, 427 and 485 chars. There were never any bricks. I rewrote prose in the previous post to satisfy an instrument nobody had calibrated - the second time this session a self-invented measurement drove edits to correct work. Warning added to voice-rules.md where the rule lives. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. Zero banned words, zero em dashes. Both internal links and both assets confirmed in the built output. Cover re-rendered after the first attempt silently produced the WRONG post's cover - a sed against the wrong source file matched nothing and copied it verbatim; caught only by looking at the PNG. NOT run: no independent verifier, no four-lens panel, no browser pass at 1280x800 / 390x844. rake test:links cannot validate a future-dated post. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
…dings
Paul's verdict: "all your today's posts are too AI slops". He is right, and the
diagnosis was already written down in our own reference file - I reproduced the
exact failure it documents.
Measured against blog-writer-reference-samples.md's own audit table, the four
posts scored: code blocks 0 (competitor floor is 8), named peer engaged 0,
disclosed interest 0 - while every one closed by selling a rescue call. That
file's verdict on the previous batch applies verbatim to these: "an essay with
no code, no diagram, no named source and no disclosed interest is asking to be
believed on tone alone, which is the one thing a sceptical founder will not
extend."
Three structural changes, copied from Arkency and Evil Martians rather than
invented:
1. CARDS ON THE TABLE, UP FRONT. Every post now declares the commercial interest
in its opening lines, then makes its argument survive without the sale.
Fiedler's move: declare the conflict, then argue in bare ActiveRecord so the
reasoning does not depend on the product.
2. AN ARTIFACT THE READER CAN USE. Each post now carries something copyable
rather than a conclusion to admire: the before/after reviewer brief that
actually changed our review output, four questions to paste into an email to
your shop, an estimate written the way we write it, and the four-line rule
for who may write.
3. ENDINGS THAT DIFFER. All four previously closed diagnosis-then-CTA, the
cross-post repetition an ICP reviewer once described as "reading one post
three times". They now end on an experiment, a standard, a warning and a
question. Two carry no CTA at all.
FOUR ERRORS I INTRODUCED WHILE REWRITING AND CAUGHT BEFORE COMMIT:
- Invented a personal failure for Paul ("that second failure is mine, from
earlier this year"). A fabricated specific, which claims-canon bans. Removed.
- A new opening promised "the last section links to a call" in the post where I
had just deleted the CTA. Fixed.
- Wrote "our engineering practices are public" linking to the rescue page. That
page is a service page. Checked it, the sentence was false, rewritten.
- Wrote "neither did we, until we started counting", implying we measure agent
run success rates. No evidence we do. Softened.
All four were written in one sitting by one hand, which is the condition under
which a shared skeleton becomes invisible to its author.
Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0
failures. Zero em dashes across all four.
STILL NOT DONE: four cold-eyes reviewer agents were spawned for this and had not
reported when this landed, so their findings are not reflected here. The
competitor-standard lens was tasked with fetching current Arkency and Evil
Martians posts and judging ours against them - the check that would say whether
this pass closed the gap or only performed it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
…e it Paul: "for us we found the formula and were able to deliver highly autonomus agents for ours local projects." That is the first-party material the competitor-standard reviewer said all four posts were missing - it counted zero first-party numbers across them and called this post's "review findings kept getting lost between agents" paragraph unfalsifiable as written. The claim now sits where that paragraph was, scoped exactly as he scoped it: our own codebases, not client delivery. The loop runs end to end - agents pick work, write, review each other, open PRs - against a short written list of decisions that come back to a human. I verified that list is real rather than describing it from memory: it is in .claude/skills/blog-next/SKILL.md line 395, "what genuinely needs Paul". The post says "that is our own codebases, not client delivery" in its own voice, because a capability claim that quietly widens its own scope is the thing claims-canon exists to stop. TWO NUMBERS I WROTE AND THEN REMOVED, both in this commit's own drafting: - "higher than that for us once we stopped letting agents write in parallel" - implies a measured success rate. Paul gave me a capability, not a percentage. - "Ours was" (the impression being better than the number) - implies we measured and compared. We have not. The closing now says plainly that we have not published a percentage because we have not measured one, and that quoting a feeling beside Cemri's benchmark figures would be the exact move the post is complaining about. That is a better ending than the one I cut. Also fixed from the competitor review, in this post only: - "**Many agents may think. One agent writes.**" - the best line in the four posts, delivered as a bolded slogan. Unbolded, and the mechanism attached: parallel writers surface their disagreements at integration, a single writer surfaces them as review comments you can act on. - "Fewer than the demos suggest, more than the sceptics claim" - a both-sides closer restating the headline. Replaced. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. STILL OPEN: the competitor reviewer flagged 9 passages across the other three posts that survived the humanizer pass, and diagnosed the section rhythm as arithmetic rather than style - 6-7 H2s over ~1250 words leaves ~180 words a section, too little to demonstrate anything, so sections close by asserting. That fix is fewer sections, not better sentences, and it is not done. The ICP and voice lenses went idle without reporting and have been pinged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
…orning Three of four reviewers reported. The blocking finding is mine and it is bad: all four posts opened "Cards on the table", which voice-guide 90.11 §6c bans BY NAME at line 560, quoting Paul's own correction from earlier the same day - "readers are not interested." I wrote that section into the guide this morning after he corrected me on that exact phrase, then reintroduced the phrase four times this evening while trying to fix a different problem. §6c's own text names the mechanism: "the failure mode when a correction lands is over-applying it." I over-applied a CTA correction into four identical disclosure paragraphs. Both reviewers reached it independently. The ICP lens: "the first one disarmed me. The second made me notice. By the fourth it is a house format, and a format cannot be candour - the thing candour is made of is that it was not scheduled." The voice lens counted the reassurance beat landing in the same position in all four. In every case the paragraph delayed a stronger sentence; "We had a reviewer that approved everything" had been demoted to paragraph two. CUT, this commit: - all four opener paragraphs - the anti-sell refrain they spawned, three instances across three posts - the Anthropic "not yet great at coordinating" quote from the second of the two posts carrying it - the ICP lens said the pair "looks like one article split in two" - "Nobody designed that. It is just where two reasonable decisions happen to meet." and three more epigram closers - "the author list includes Matei Zaharia, Joseph Gonzalez and Ion Stoica" - credentialing by names the ICP reader cannot evaluate - "So, how many got success?" as a heading. Not grammatical. Caught by the voice lens, missed by me across three passes. - "actually" in the delegation post: 6 -> 2 WHAT THE PANEL FOUND THAT NO GATE COULD, and what is still open: The section-rhythm diagnosis is arithmetic, not style. 16 of 26 sections across the four close on a standalone verdict paragraph, and the voice lens located the cause underneath it: nine instances of the writer appending a judgment sentence after an exhibit, telling the reader how much the exhibit is worth. "That is the part I would worry about." "The third row is the one I would not compromise on." That is the committee voice. Fixing the closers without fixing that habit would just move it. Notably the delegation post scored 2 of 7 and the voice lens said not to touch it - the two code fences fixed its rhythm as a side effect. Exhibits break sections better than sentences do. STILL NOT DONE, and both are structural rather than line-level: - Three of four posts are written for someone running agents, not for the ICP. The ICP reader's verdict on agent-team-success-rate: "I do not have agent runs. This post reads as written for my developers, not for me." - Zero first-party numbers across all four. The competitor lens measured ours against Evil Martians publishing 268,000 tracked requests from their own logs, including the finding that embarrassed them. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. Zero em dashes. All four carry a code block. Also banked: 20.09 §13n, Paul's client story - release day as the most stressful day, daily releases as "romantic fantasy". Deliberately undated and BLOCKED on how far the client may be described, because a date would create pressure to write it with a composite client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
claims-lens - the lens that checks every figure against its primary - went idle without reporting, was pinged with the four claims I most doubted, and went idle again without reporting. Two silent idles is not a third ping, so I ran its job. VERIFIED at the primary PDF (arxiv.org/pdf/2503.13657v2, fetched and extracted rather than read from a search excerpt, which is how these numbers reached the post in the first place): - All six success rates present and correct: AG2 59.0, MetaGPT 40.0, Magentic-One 38.0, ChatDev 33.3, HyperAgent 25.3, AppWorld 13.3. - The caption I quote is verbatim: "Performances are measured on different benchmarks, therefore they are not directly comparable." AND THE CAPTION CARRIED SOMETHING I HAD DROPPED. Its full text is "Failure rates of six popular Multi-Agent LLM Systems **with GPT-4o and Claude-3**." The post never said which models. A reader in 2026 would reasonably take a table in a post about 2026 as current-model numbers; these are not. The post now says so, and says the open question the paper does not answer: frontier models have moved, whether the coordination failures moved with them is unmeasured. That is the "sentence around the quote" failure again - the quote was right, the frame around it was not - and it is the third time today. ORPHANED CITATION, self-inflicted an hour ago: I cut the duplicated Anthropic quote from the delegation post and left Anthropic in its Sources list. Zero mentions in the body, first entry in the list - the exact "cited but never spent" defect. Fixed by making the citation earn its place: the post's claim that bad briefs hurt agents more than people needed support, and that is what Anthropic's write-up actually supports. Attributed now rather than quoted, so it no longer duplicates the sentence used in the other post. Also trimmed the ADP 6-0 sources note, which was explaining a non-citation at a length the ICP reviewer called "a preoccupation". CHECKED MECHANICALLY across all four: every remaining listed source is engaged in its body, every internal link resolves in the built output. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
…an AI Paul on agent-team-success-rate: "so bad, rephrase each paragraph." Rewritten end to end rather than patched. What was actually wrong, now that I read it as prose instead of as a checklist: A BROKEN SENTENCE had been sitting in paragraph two since I cut the author names: "A group at UC Berkeley - collected execution traces". An orphaned dash, shipped through three review passes and two of my own reads. EVERY PARAGRAPH ENDED ON A VERDICT. Five of six sections closed on a standalone judgment - "Not a rounding error." / "which is a form you can act on." / "There is nothing else to it." Now one of six does. The rest end on a quote, an explanation, an instruction or a question, because those are what a paragraph ends on when it has finished saying something rather than finished performing. THE COMMITTEE VOICE was in the connective tissue, not the claims. "Two things travel with those numbers." "The three categories are worth memorising." "That constraint is doing specific work." "Then the sentence that should decide this for most teams." Nine instances of telling the reader how to weight the thing they just read. All cut. The evidence carries itself or it does not belong. PLAINER WHERE IT WAS ABSTRACT. "The failures are organisational, not mathematical" became "The failures look like a bad org chart", and the three categories are now given in plain words - the specification was wrong, the agents talked past each other, nobody checked the result - before the taxonomy names them. Cohen's kappa is gone; it added nothing a reader could use. Also: "Cognition changed its mind in public" replaces "Both camps arrived at the same rule". The competitor reviewer pointed out we stage those two labs agreeing, which is the less interesting version. One of them publicly narrowed a position ten months after taking it, and that is the story. SEPARATELY, delegate-the-goal-not-the-task, Paul: "'We had a reviewer that approved everything' - do you mean reviewer agent?" Yes, and the post did not say so until two sections later. Read as written, the opening line sounds like I am publicly criticising a colleague, which is both wrong and a bad look. It now reads "We had an AI code reviewer that approved everything", and the later callback no longer re-introduces the same fact as if it were new. Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0 failures. 1250 words, zero em dashes, code block and diagram intact, internal link resolves in the built output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
… folds
Four cold-eyes agents (ICP, voice, competitor standard, claim verification)
reviewed the four scheduled posts. This applies their findings to three of
them; delegate-the-goal-not-the-task is held pending a decision from Paul.
FACTUAL (these changed what a post asserts):
- dev-shop: the 51%-vs-65% comparison was invalid. Uber's 65% is the system
re-running itself five times on the final commit; the 51% is a human
author's verdict. Both quotes verbatim, the comparison between them is not
supported. Rewritten to state each instrument, and reviewers.svg deleted -
the chart rendered the false equivalence as a bar pair.
- dev-shop: the ISSTA paper measures LLM agents triaging SAST alerts, not an
AI reviewer's comments. Named the tool class instead of bridging silently.
- dev-shop: 22.25% is "on the OWASP Benchmark positives" - synthetic only,
not the real-world Java dataset the disclaimer implied.
- dev-shop: Uber suppresses categories of "historically low developer value",
not ones that "annoyed developers". Tencent's 94-98% is 433 alarms, three
bug types, one product - the symmetry with Monash was overstated.
- generate-then-verify: two-week code churn +15% is indexed to 2023, not
2022; the "vs 2022 levels" in the source governs the reuse signals. The
sentence carrying the wrong baseline is cut.
- agent-team: the 80%-variance gloss dropped two named factors and read
variance-explained as causal attribution. Restated.
- agent-team: Anthropic's subagents do not write code, so their system has no
writes to serialise; the equivalence with Yan's rule was ours, asserted as
a restatement of theirs.
- agent-team: two internal contradictions - "nobody else writes" vs "agents
write", and "unattended" vs a loop that escalates anything published
outward. Both were true of different scopes; both now say which.
- agent-team: the cited figure is captioned "Failure rates" while its legend
labels the plotted values Success. Said so, and used the v3 model name.
VOICE (20 of 26 section closers ended on a declarative verdict; 13 on an
is/are copula):
- Cut ten valuation sentences that told the reader what the evidence was
worth before showing it ("The number worth sitting with", "Their third
stated lesson is blunt enough to quote whole", "The third row is the one I
would not compromise on", "The reason it works is unglamorous").
- Rewrote the verdict-closers that generalised rather than landing on
evidence or an action.
- Cut the signposting paragraphs and the two unattributed strawmen.
- generate-then-verify no longer opens on the connotation of "then".
MECHANICAL:
- dev-shop and agent-team shipped 3-column tables against the 2-column cap
in .okf/content/voice-rules.md:79 - at 390px the third column breaks the
mobile scroll gate, and Magentic-One / SWE-Bench Lite are exactly the
unbreakable tokens the rule names. Both folded to 2.
GATES: bin/hugo-build clean. test:critical 38 runs / 126 assertions / 0
failures, snap_diff 55 screenshots no failures. test:links 0 errors across
31,948 unique links - but all four posts are future-dated and were NOT in
that build, so the link evidence does not cover them; their five internal
links were verified by hand instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
Three claim-verification findings on this post, independent of the fabricated caching example that still blocks it. - Anthropic ranks nothing. The source says LLM agents are "not yet great at coordinating and delegating to other agents in real time"; the post had promoted that to "the thing these systems are still worst at". Quoted instead of paraphrased upward. - "and has for longer than software has existed" was the exact lineage claim the post's own Sources entry declined to defend. Cut the clause, kept the term. - Dropped the ADP 6-0 Sources entry. An entry whose body is a disclaimer that the author has not read the literature is not a source, and it was the one bullet across all four posts that no sentence engaged. Post remains HELD: its central before/after prompt pair describes a tenant-scoped caching layer that does not exist in this repo, and the closing section asserts it as a real instruction. Needs the real one from Paul. Gates: bin/hugo-build clean. Prose-only diff, no template/CSS/body HTML. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
Durable learning from the four-lens panel, riding the same PR as the fixes. fabrication-ratchet gains a section: a claim can be false while every quote in it is verbatim, when the two figures being compared came from different instruments. The ratchet counts invented shapes and cannot reach this, because the defect lives in the join rather than in either half. Second half of the lesson: correcting the prose did not correct the claim - the comparison was also a chart, which survived the sentence edit and stayed published in the page bundle. Claim fixes now include grepping for a diagram carrying the claim. Also logged: the 20-of-26 verdict-closer count as the measurable form of "reads like AI" (set-level, so unreachable by per-post review), and the reminder that a link-check green over a build excluding future-dated posts is vacuous evidence for those posts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
Same-commit rule debt from this PR - four posts entered flight and one hit a blocker with no STATUS row. Closing it. - Blog/SEO row: four posts scheduled 08-26 through 09-16, the four-lens panel and what it found, and the new defect class now in fabrication-ratchet. - Blocked on Paul: the real reviewer anecdote for delegate-the-goal-not-the-task, which is HELD until it arrives. Plus the optional dev-shop cover chip, whose 22.25% no longer matches the scope the body now states. - Named the two gaps the panel found that I did NOT close: 0 of 4 posts carry a runnable code block, 0 carry a first-party number. Recording them so they are declined rather than merely absent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
The post is dated 2026-09-16. Merging it while HELD would arm a time bomb: if the real anecdote never arrives, a known fabrication publishes itself in three weeks. draft: true removes the deadline without losing the work. Verified in both directions rather than assumed - the production build alone does not discriminate, since the post is future-dated and would be absent either way: hugo --buildFuture -> ABSENT (the draft flag excludes it) hugo --buildFuture --buildDrafts -> present (so the check can see it) Resume: replace the invented before/after prompt pair with the real one, then flip draft to false. The frontmatter carries the same note, and STATUS.md > Blocked on Paul carries the ask. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
New post (scheduled 2026-08-26): what an AI code reviewer was told to ignore
Founder/ICP-E stream, arrival-purposed, 1,410 words. The calendar picked the
stream and the date before the topic existed: the last 7 days held 11 posts, all
Rails Technical, while ICP-E had been quiet since 08-08.
The thesis survived being wrong twice, which is the only reason it is worth
publishing.
First version. "An AI reviewer is weakest where AI code is weakest." Then
Uber's published numbers arrived: only 51% of their HUMAN review comments are
judged valid and addressed, against over 65% for uReview, with 75% rated useful
across >90% of ~65,000 weekly diffs. On that evidence AI review is not the weak
link - human review is. The post now leads with that, because a reader who
thinks this is an anti-AI piece stops reading.
What the post actually turns on is Xiong & Zhang (ISSTA 2026): filtering
takes SAST false positives from >92% down to 6.3%, and the same configuration
"incorrectly suppresses 22.25% of real vulnerabilities", with miss rates that
"exceed 50% for cryptography- and policy-related categories". The tuning that
makes an AI reviewer bearable to developers is what drops the findings a founder
would care most about.
Second correction, caught by the claim gate. The draft said a Tencent result
was "close to the opposite" of that finding. Reading to the end of that paper:
"recall remains below the enterprise-expected threshold of 90%, indicating that
some degree of manual review is still necessary". Two teams, opposite-looking
headlines, same conclusion - do not run it unattended. I had manufactured a
tension that the sources do not support; the convergence is the better point and
is now what the section says.
Every figure was re-fetched at its primary before it was written down. Two
reached the outline through search excerpts and both changed on contact: the
per-CWE miss rates I had (77.17%, 84.50%) come from a secondary catalog, not the
paper, so the post quotes the paper's own ">50%" instead.
Also in this commit, from the same session's corrections:
blog-nextgains Stage A0 (calendar/spacing/stream-and-stack balance beforetopic selection), a fourth exit SCHEDULED (a same-week collision is a spacing
problem, not a verdict - it previously produced DO-NOT-WRITE and lost the
topic), three sources (Reddit via WebSearch, X, changelogs), and a MANDATORY
invocation of
blog-writerather than a recommendation to run it.old.reddit 302), verified, so nobody scripts it.
blog-pipeline.md.
research_importdocumented as mandatory - without itnotebook_queryanswers from an empty notebook.Gates: bin/hugo-build green. marketing_copy_test 5 runs / 13 assertions / 0
failures (fabrication + uncitedness ratchets; the post carries 5 citations).
Zero banned words, zero em dashes. Both internal links confirmed to resolve in
the built output rather than trusted from a sibling post.
NOT run, and the post should not be merged as if they were: no independent
verifier saw this draft. Agent spawning was unavailable, so the four-lens
cold-eyes panel did not run and the gates above are self-administered. No cover
image yet.
rake test:linkscannot validate this post at all while it isfuture-dated - production skips future content - so quoting a green link run
here would be evidence of nothing.
🤖 Generated with Claude Code
https://claude.ai/code/session_016PUkwFTsiv7EB2DYKogbpg