Waiter, there's a watermark in my slop

speckx1 pts0 comments

Waiter, there’s a watermark in my slop

20 Aug 2026

Waiter, there’s a watermark in my slop

David Gerard covers the new text<br>watermarking in Anthropic&rsquo;s LLM output: Anthropic<br>watermarks AI output — and the AI bros yell.

Nobody who claims they can tell good writing from bad should care<br>this much. If they&rsquo;re not using the slopbot instead of doing the<br>writing.

The watermarked text doesn&rsquo;t even make it into the final writing<br>project, if the LLM was just used for assistance, with methods such as<br>Seth Godin&rsquo;s AI tear<br>down. So it shouldn&rsquo;t really matter. Maybe Matt Birchler is right,<br>and the<br>reader&rsquo;s desire to know when text is AI-generated is more important than<br>the AI user&rsquo;s desire for the opposite. Maybe Dave Winer is right in<br>A rebuttal to Doc Searls&rsquo; piece about<br>watermarks and the risk that a watermark might be detected could<br>help discourage people from getting a long bit of writing from the<br>bot, and then pasting it into a comment thread in a GitHub repo I<br>run.

The bigger question, though, isn&rsquo;t about the negligible<br>difference between watermarked slop and somehow purer non-watermarked<br>slop.

If LLM output can have watermarks, what else can it have?

Other identifiers. If Anthropic can get a 1-bit<br>yes/no watermark into 200 words of text, how many words does an LLM<br>provider need to encode a 32-bit identifier? That&rsquo;s probably infeasible<br>in the lengths of text that most people use LLMs for in practice. Maybe<br>they would be able to find the authors of all those slop<br>clone books on Amazon dot com, but for most business writing and<br>article-length text it wouldn&rsquo;t work. SynthID, described in Scalable<br>watermarking for identifying large language model outputs, encodes a<br>single bit of information by over-weighting token choice from one<br>red list and one green list. Splitting the list into 2n<br>categories would give you more bits but require a lot more text to<br>encode it.

But what about just a few more bits? Now that governments and Big<br>Tech know that encoding information into LLM-generated text is possible,<br>what&rsquo;s next?

USA Citizen/non-citizen? This would be a obvious one for a<br>company trying to get in good with the Federal government here.

Age verification? I don&rsquo;t know, they&rsquo;re sticking age verification<br>on everything now, someone will try it.

Marketing-related data? The<br>highest-earning 10% of people in the USA buy 50% of the stuff, so<br>all kind of uses for a flag to show which side of the K-shaped economy<br>the user is on.

Infringing or plagiarizing text. The big AI<br>companies have a pretty well-defined political program, and can identify<br>likely opponents to some or all of that program. Some of those opponents<br>are consistent &ldquo;AI vegans&rdquo; in their own personal IT choices, but if<br>someone is politically inconvenient and an AI user, well, it&rsquo;s<br>possible to tune the likelihood that an LLM&rsquo;s output contains material<br>straight out of the training set.

Could the LLM&rsquo;s output contain deliberate infringement bombs or<br>plagiarism bombs? Deliberately giving users some text that would get<br>them in trouble later is not the kind of thing that a human developer<br>would risk—too much risk that the commit messages would come out in<br>discovery—but it is the kind of trick that a heavily &ldquo;agentic&rdquo; automated<br>software process would come up with. An AI agent can<br>already hack a gym to get its owner a spot in pilates class, so a<br>program of compromising political opponents who are also users seems<br>feasible.

Marketing side effects. Some LLMs are serving ads<br>now, which means a lot of hard-to-predict ML-driven ad placements. And<br>this stuff will be a lot weirder and more indirect than current projects<br>like the attribution<br>cartel, which is pretty clearly going to cause ML to come up with<br>privacy-violating ways to juice the apparent results from Big Tech<br>advertising. ML systems going for other goals are going to get a lot<br>weirder. Meta&rsquo;s ad ML is already doing<br>pretty weird (and creepy) stuff to maximize engagement, and it&rsquo;s<br>only going to get weirder.

What happens when more ad-revenue-maxing ML is in the loop in more<br>places? What if you can sell a renter&rsquo;s insurance policy to user A by<br>giving their friend, user B, some bad household tips that result in an<br>expensive bill from their landlord? Consumer-facing LLMs are run by the<br>same companies that find themselves in a desperate<br>squeeze to keep raising ad revenue at startup-like growth rates.<br>Corners will be cut. Other ways will be looked. Slop advice will reflect<br>the need to achieve ambitious business goals, even at the user&rsquo;s<br>expense.

Anyway, the difference between watermarked slop and non-watermarked<br>slop is tiny compared to some of the text that&rsquo;s going to be in the LLM<br>output. Enjoy.

Bonus links

Hundreds<br>of Fake VPNs Are Flooding the Chrome Web Store by Ritoban Mukherjee.<br>(Ad blockers are another category with similar problems. If...

rsquo text slop output from user

Related Articles