Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .claude/skills/blog-operator/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,6 +81,17 @@ Concrete, in the order these tend to pay:
Volume is not on that list. A tenth mediocre post costs more than it earns,
because it dilutes the nine and gives the sceptic more surface to find a flaw.

## Budget for the panel, not just the draft

Writing a post is the cheap half. On 2026-08-22 three posts passed every check
their author could run and a four-lens cold-eyes panel then found wrong numbers,
footer-only citations, a broken shell command, and a section-level rhythm
identical across all three.

So when you sequence work, a post is not one unit. It is draft, then panel, then
a fix pass that waits for all four reviewers before touching anything. Plan for
the panel or you will ship the draft.

## The gates are not yours to waive

Both hands carry their own blocking gates and they stay blocking. You may
Expand Down
62 changes: 62 additions & 0 deletions .claude/skills/blog-write/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,6 +156,68 @@ One dev server per session, never 1313:
PORT=$((20000 + RANDOM % 20000)) bin/dev
```

## What four cold-eyes reviewers found, and what it costs to skip them

Three posts shipped on 2026-08-22 having passed my own checks. A four-lens panel
then found defects in all three, and every category below is something no script
and no self-review caught.

**Numbers I was confident about were wrong.** A tool output was reproduced inside
a fence with one field silently edited. A `[snap_diff] 287 screenshots compared`
line took 287 from the *unit-test run count*. A link figure from today's run was
used to illustrate a fix that landed weeks earlier at a different number. Every
one felt remembered rather than invented, which is exactly why they survived.

> Before you put a number in a fence, open the artifact that produced it and copy
> the line. Numbers you "remember measuring" are the dangerous ones - you did
> measure something, just not this.

**Cited but never spent.** Two of three posts listed a source in the footer that
the body never engages. A citation nobody uses is dressing, and one of ours had
its claim on the cover image while the post never told the story.

**A source's own words checked, its attribution not.** "Responsible Statecraft
traced it" - they reported it; a NewsGuard analyst traced it. "Both publish their
error bars" - one does. Verifying the quote is not verifying the sentence around
the quote.

**Commands shipped without running them.** A `for f in content/blog/**/index.md`
loop matches nothing in bash without `shopt -s globstar` - it reports success
having checked zero files, in a post about checks that check nothing. Run every
command you publish, in a clean shell, and show what it returned.

**The macro-rhythm, which is the actual slop.** Every H2 ended on a one-sentence
aphorism paragraph. Sentence lengths varied; the choreography did not -
"evidence, evidence, punchline, white space, next heading", five times a post,
three posts. And all three opened on the identical two-beat reveal: flat claim,
one-line rug-pull. Individually each works. As a set, a reader hears the machine.

> Read two consecutive sections aloud. If they land the same way, one of them
> has to change shape - not wording.

**Cross-post repetition.** The same Laravel anecdote carried two same-day posts;
all three closed on the same pitch. The ICP reviewer's verdict: "as a subscriber
I'd feel I read one post three times."

## Run the panel, do not self-review

Four lenses, spawned as separate agents of a DIFFERENT type than the writer:

1. **ICP cold read** - give it `90.10`, ask where it stops reading and what it
would cut. Require three cuts per post.
2. **Voice and AI tells** - give it `90.11` and STEP 5a. Ask for what a regex
cannot see: prose that is grammatical, on-pattern and lifeless.
3. **Competitor standard** - give it `blog-writer-reference-samples.md` and have
it FETCH two current competitor posts, so it compares against writing rather
than a summary.
4. **Claim verification** - every external claim fetched at the primary, every
in-repo number checked against the artifact that produced it.

Brief each with goal and artifact, never with your conclusions. Require each to
name something it would cut; a panel returning three approvals was not asked a
real question. Run them in parallel and wait for all four before editing, or you
will fix the same post four times.

## Three exits, and only three

- **SHIPPED** - committed, PR open, gate verdicts quoted with their numbers.
Expand Down
2 changes: 1 addition & 1 deletion .okf/build/test-gates.md
Original file line number Diff line number Diff line change
Expand Up @@ -847,7 +847,7 @@ Skipping step 1 has cost this repo repeatedly:
spare hits swallowed a planted banned adjective whole. A ratchet with slack
is a gate that has already been disarmed.
- `rake test:links` excluded 133,874 of 149,516 links (production renders
absolute URLs; `--offline` drops every http(s) URI) and was green for a year
absolute URLs; `--offline` drops every http(s) URI) and was green from the day it shipped (2026-07-21) until 2026-08-22
on a site with five real broken links, one of them a conversion path and one
a post's own canonical pointing at a 404.

Expand Down
2 changes: 1 addition & 1 deletion Rakefile
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ namespace :test do
# emits is globbed from disk and passed as an explicit input, so no PAGE is
# skipped.
#
# That is not the same as no LINK being skipped, and for a year it wasn't:
# That is not the same as no LINK being skipped, and from the day the gate shipped (2026-07-21) it wasn't:
# the production build renders internal links absolute
# (https://jetthoughts.com/...) and `--offline` excludes every http(s) URI by
# design, so 133,874 of 149,516 links were excluded and the job was green
Expand Down
27 changes: 13 additions & 14 deletions content/blog/how-to-audit-content-you-didnt-write/index.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "How to Audit Content You Didn't Write"
description: "Someone spent $900,000 publishing fake research so chatbots would repeat it. The same economics apply to the blog your agency built. Four checks you can run."
description: "A fake think tank published 100+ reports built to be repeated by chatbots, behind a $900,000 government contract. The same economics reached your blog. Four checks you can run."
date: 2026-08-22
draft: false
author: "Paul Keen"
Expand All @@ -16,19 +16,17 @@ canonical_url: 'https://jetthoughts.com/blog/how-to-audit-content-you-didnt-writ
related_posts: false
---

The Hanover Institute for Public Policy published more than a hundred research reports, complete with footnotes, tables of contents, and the flat neutral register that policy writing has.
There is no Hanover Institute for Public Policy. There is a website carrying more than a hundred reports under that name - footnotes, tables of contents, the flat neutral register that policy writing has - and behind it a marketing firm working on a government contract.

It does not exist.

[Responsible Statecraft traced it](https://responsiblestatecraft.org/israel-influence-chatgpt/) to Piro, Inc., and Politico found the Department of Justice filing showing $900,000 of Israeli government funding behind it. GPTZero flagged eleven of twelve sampled articles as machine-written.
NewsGuard analyst Alice Lee connected it to Piro, Inc., [reported by Responsible Statecraft](https://responsiblestatecraft.org/israel-influence-chatgpt/); Politico first reported the Department of Justice filing, under which Piro has received $900,000 from the Israeli government for its work. GPTZero flagged all twelve sampled articles as AI-written - eleven with high confidence, one moderate.

The interesting part is who the reports were written for. Piro's founder said it on LinkedIn: "When someone asks ChatGPT, Gemini, or Perplexity about your category, an answer comes back in one confident paragraph... we spent months reverse-engineering it."

Those reports were never aimed at readers. Their audience was the machine that answers readers, and the reports were shaped to be the thing it repeats.

## Your blog runs on the same economics

Nobody spent $900,000 on your content.
Nobody put a government contract behind your blog.

That is the point. They did not have to, and neither did whoever produced yours, because manufacturing text that reads like expertise stopped being expensive somewhere around the middle of 2023.

Expand All @@ -40,7 +38,7 @@ Then a tool started drafting, and the person approving its output was not equipp

You would think there is a number. There are several and they disagree - ten percent, a third, or half, depending on whose sample and whose detector.

[Graphite](https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans) put the crossover, more machine-written articles than human ones, in November 2024. [Pew](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) ran ~490,000 Common Crawl pages through Open Pangram this month and found 10% carrying AI-authorship signals, rising to over a third among pages published after ChatGPT shipped. Both publish their error bars, which is the habit worth stealing whatever you make of the figures.
[Graphite](https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans) put the crossover, more machine-written articles than human ones, in November 2024. [Pew](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) ran ~490,000 Common Crawl pages through Open Pangram this month and found 10% carrying AI-authorship signals, rising to over a third among pages published after ChatGPT shipped. Graphite states its own 4.2% false-positive rate and notes on its own page that this finding has since been superseded by a study averaging three detectors. Pew publishes no error bars at all - only the caveat that detectors "sometimes misclassify individual documents" and hold up in aggregate. Graphite tells you how wrong it might be; Pew tells you it might be wrong.

Pew also split it by domain, and that is where you come in:

Expand Down Expand Up @@ -82,12 +80,11 @@ That shape is greppable:
```bash
# every case-study heading in the archive
grep -rniE '^#{2,4} .*case stud' content/

# the anonymous-subject tell, right after one
grep -rniE 'a (mid-siz|medium-siz|large)|\(anonymous' content/
```

Run the first one and read every hit. Real client work names the client or does not get published, and you can apply that test without understanding a word of the subject matter.
Run it and read every hit. Real client work names the client or does not get published, and you can apply that test without understanding a word of the subject matter.

Do not try to automate the second half. We tried: an anonymous-subject pattern (`a mid-sized`, `a large`) returned ten matches that were nearly all legitimate, because "a large number of" and "a large team" are ordinary English. The heading is the cheap signal; the judgement stays human.

**3. Ask whether a claim can be checked at all.**

Expand All @@ -101,22 +98,24 @@ Count yours:

```bash
# posts over 400 words carrying zero outbound links to anywhere but your own domain
for f in content/blog/**/index.md; do
find content/blog -name '*.md' | while read -r f; do
words=$(wc -w < "$f")
links=$(grep -oE '\]\(https?://[^)]+\)' "$f" | grep -vc 'yourdomain.com')
[ "$words" -gt 400 ] && [ "$links" -eq 0 ] && echo "$words words, 0 sources: $f"
done
```

Use `find`, not `content/blog/**/*.md` - bash does not expand `**` recursively unless `shopt -s globstar` is set, so the glob version silently matches nothing and reports success. That is the exact defect this post is about, and it was in my first draft of this command.

A post making technical claims with zero citations is not a red flag about that post's accuracy so much as a flag that accuracy was never tested.

**4. Check whether the advice has expired.**

Any post with a version number in the title has a shelf life its author never wrote down.

A migration guide that recommends Laravel 11 today is sending readers onto a release whose security support ended in March 2026. Nothing in that guide has to be invented for it to do damage.
Every framework you write about has a support table, and every one of your version-numbered posts is silently betting that the version it recommends is still on it. When that stops being true, the post does not change and nothing in it becomes false - it just starts pointing readers at an unpatched release.

It was true when written, and became harmful without a word of it changing.
That is the whole failure. No invention required.

```bash
# every post whose title names a version - each one has an expiry date
Expand Down
22 changes: 12 additions & 10 deletions content/blog/what-senior-developers-catch-that-ai-misses/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,10 +19,12 @@ related_posts: false
Here is a change we caught before it merged. It is small, it is plausible, and it is wrong in a way you cannot see without knowing Rails.

```diff
- Propshaft is dramatically faster than Sprockets: precompilation drops from
- 45-60 seconds to under 5 seconds on a medium app.
+ Propshaft drops the transpilation and concatenation stages entirely, so asset
+ precompilation stops being a build step that scales with your asset count.
- Propshaft replaces Sprockets as the default asset pipeline in Rails 8, and the
- difference is dramatic: in our experience, build times drop from 45-60 seconds
- to under 5 seconds for medium-sized apps.
+ Propshaft replaces Sprockets as the default asset pipeline in Rails 8. It drops
+ the transpilation and concatenation stages entirely, so asset precompilation
+ stops being a build step that scales with your asset count.
```

The deletion is correct. That timing figure had no measurement behind it and deserved to go.
Expand All @@ -49,6 +51,8 @@ You catch it by already knowing what `assets:precompile` does. Sean Goedecke put

He calls the thing experts do "steering" - you recognise a suboptimal suggestion and redirect it. This diff is that mechanism running backwards. Without someone who knows the asset pipeline, there is nothing to steer against and the confident answer wins by default.

Senko Rašić pushed back on that framing a fortnight later, insisting that "creating good code is a craft that requires skill, patience, attention to detail, experience and wisdom". The diff above argues for his side better than it does for Goedecke's. Writing that sentence took no craft at all. Knowing it was wrong took every item on his list.

Note what the change was *for*. The task was removing an unsourced number, and the same edit introduced a new defect while completing it. Cleanup is where this happens most, because a correction feels like tidying rather than authorship.

## What actually caught it
Expand All @@ -57,13 +61,11 @@ A second model, told to attack the diff.

The fix was not a better model or a longer prompt. It was a different one, with a brief that made disagreement its job rather than a risk.

It came back with four findings. This was one, stated flatly:

> On applications with many assets, Propshaft still enumerates, fingerprints, and copies every asset during `assets:precompile`, so its work still scales with asset count. Removing transpilation and concatenation reduces the per-asset cost but does not make the build independent of asset count; the new wording gives readers an incorrect performance expectation.
It came back with four findings. One was this: on applications with many assets, Propshaft still enumerates, fingerprints and copies every asset during `assets:precompile`, so the work still scales with asset count. Dropping transpilation and concatenation lowers the cost per asset without making the build independent of how many there are, and the new wording promised the wrong thing.

Then a person had to decide whether the objection was correct, and that step took either already knowing the answer or being willing to go and read the Propshaft source until you did.

Three links in that chain, and only one of them is automatable. A model wrote, another model challenged, and someone with domain knowledge adjudicated.
Three links, and the last one is the only one you cannot automate. A model wrote, another model challenged, and someone with domain knowledge decided which was right.

![Three links in the chain: a model writes, a second model challenges, a person referees. Only the first two are automatable.](chain.svg)

Expand Down Expand Up @@ -115,7 +117,7 @@ We wrote about the [team structure that makes this hold up](/blog/claude-code-xp

## Sources

- Sean Goedecke, ["LLMs reward expertise"](https://www.seangoedecke.com/llms-reward-expertise/) - [HN discussion](https://news.ycombinator.com/item?id=49161518), 573 comments
- Senko Rašić, ["'Code was never the hard part' is an insult to all programmers"](https://blog.senko.net/code-was-never-the-hard-part-is-an-insult-to-all-programmers) - [HN discussion](https://news.ycombinator.com/item?id=49222189), 590 comments
- Sean Goedecke, ["LLMs reward expertise"](https://www.seangoedecke.com/llms-reward-expertise/)
- Senko Rašić, ["'Code was never the hard part' is an insult to all programmers"](https://blog.senko.net/code-was-never-the-hard-part-is-an-insult-to-all-programmers)
- METR, ["Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity"](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) - the 19% slowdown, the 24% predicted speedup, the 20% believed speedup, and METR's own limits on what it shows
- [Propshaft](https://github.com/rails/propshaft) - the asset pipeline whose behaviour the claim got wrong. Its README settles both halves. It is faster: "a dramatically simpler and faster asset pipeline compared to previous options, like Sprockets." And it still does per-asset work: "All assets in the load path will be copied (or compiled) in a precompilation step for production that also stamps all of them with a digest hash." The original sentence took the first half as the reason for the second.
6 changes: 3 additions & 3 deletions content/blog/when-did-a-test-last-fail-on-purpose/checked.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading