Waiter, there’s a watermark in my slop
20 Aug 2026
Waiter, there’s a watermark in my slop
David Gerard covers the new text<br>watermarking in Anthropic’s LLM output: Anthropic<br>watermarks AI output — and the AI bros yell.
Nobody who claims they can tell good writing from bad should care<br>this much. If they’re not using the slopbot instead of doing the<br>writing.
The watermarked text doesn’t even make it into the final writing<br>project, if the LLM was just used for assistance, with methods such as<br>Seth Godin’s AI tear<br>down. So it shouldn’t really matter. Maybe Matt Birchler is right,<br>and the<br>reader’s desire to know when text is AI-generated is more important than<br>the AI user’s desire for the opposite. Maybe Dave Winer is right in<br>A rebuttal to Doc Searls’ piece about<br>watermarks and the risk that a watermark might be detected could<br>help discourage people from getting a long bit of writing from the<br>bot, and then pasting it into a comment thread in a GitHub repo I<br>run.
The bigger question, though, isn’t about the negligible<br>difference between watermarked slop and somehow purer non-watermarked<br>slop.
If LLM output can have watermarks, what else can it have?
Other identifiers. If Anthropic can get a 1-bit<br>yes/no watermark into 200 words of text, how many words does an LLM<br>provider need to encode a 32-bit identifier? That’s probably infeasible<br>in the lengths of text that most people use LLMs for in practice. Maybe<br>they would be able to find the authors of all those slop<br>clone books on Amazon dot com, but for most business writing and<br>article-length text it wouldn’t work. SynthID, described in Scalable<br>watermarking for identifying large language model outputs, encodes a<br>single bit of information by over-weighting token choice from one<br>red list and one green list. Splitting the list into 2n<br>categories would give you more bits but require a lot more text to<br>encode it.
But what about just a few more bits? Now that governments and Big<br>Tech know that encoding information into LLM-generated text is possible,<br>what’s next?
USA Citizen/non-citizen? This would be a obvious one for a<br>company trying to get in good with the Federal government here.
Age verification? I don’t know, they’re sticking age verification<br>on everything now, someone will try it.
Marketing-related data? The<br>highest-earning 10% of people in the USA buy 50% of the stuff, so<br>all kind of uses for a flag to show which side of the K-shaped economy<br>the user is on.
Infringing or plagiarizing text. The big AI<br>companies have a pretty well-defined political program, and can identify<br>likely opponents to some or all of that program. Some of those opponents<br>are consistent “AI vegans” in their own personal IT choices, but if<br>someone is politically inconvenient and an AI user, well, it’s<br>possible to tune the likelihood that an LLM’s output contains material<br>straight out of the training set.
Could the LLM’s output contain deliberate infringement bombs or<br>plagiarism bombs? Deliberately giving users some text that would get<br>them in trouble later is not the kind of thing that a human developer<br>would risk—too much risk that the commit messages would come out in<br>discovery—but it is the kind of trick that a heavily “agentic” automated<br>software process would come up with. An AI agent can<br>already hack a gym to get its owner a spot in pilates class, so a<br>program of compromising political opponents who are also users seems<br>feasible.
Marketing side effects. Some LLMs are serving ads<br>now, which means a lot of hard-to-predict ML-driven ad placements. And<br>this stuff will be a lot weirder and more indirect than current projects<br>like the attribution<br>cartel, which is pretty clearly going to cause ML to come up with<br>privacy-violating ways to juice the apparent results from Big Tech<br>advertising. ML systems going for other goals are going to get a lot<br>weirder. Meta’s ad ML is already doing<br>pretty weird (and creepy) stuff to maximize engagement, and it’s<br>only going to get weirder.
What happens when more ad-revenue-maxing ML is in the loop in more<br>places? What if you can sell a renter’s insurance policy to user A by<br>giving their friend, user B, some bad household tips that result in an<br>expensive bill from their landlord? Consumer-facing LLMs are run by the<br>same companies that find themselves in a desperate<br>squeeze to keep raising ad revenue at startup-like growth rates.<br>Corners will be cut. Other ways will be looked. Slop advice will reflect<br>the need to achieve ambitious business goals, even at the user’s<br>expense.
Anyway, the difference between watermarked slop and non-watermarked<br>slop is tiny compared to some of the text that’s going to be in the LLM<br>output. Enjoy.
Bonus links
Hundreds<br>of Fake VPNs Are Flooding the Chrome Web Store by Ritoban Mukherjee.<br>(Ad blockers are another category with similar problems. If...