Watermarking AI Text is Fundamentally Flawed - Negative Star Innovators Blog
×
The recent revelation that Anthropic is watermarking text outputs from its AI models has ignited a firestorm of debate across social media. While the exact methodology is not explicitly detailed in their documentation, it is presumably based on statistical word-choice biases designed to survive copying, pasting, and even subsequent AI proofreading. Anthropic themselves concede that this method "is not fully conclusive." We will explain more on this later.
While the socialwebs' reaction has been swift and brutal, several critical implications are being overlooked in the mainstream discourse. Here is our perspective on why text watermarking is a fundamentally flawed paradigm.
Anthropic is Not the Only One
While Anthropic is currently absorbing the brunt of the negative publicity, they are hardly pioneers in this space. Google has been watermarking its generated text since at least May 2024 (SynthID Paper and Google Announcement of its usage in Gemini). OpenAI has similarly declared that it fully intends to add provenance signals into AI generated text as recently as August 2nd, 2026 (OpenAI Help Center).
The disproportionate outrage directed at Anthropic while Google has already implemented this more than 2 years ago and OpenAI have just announced as recently as 12 days ago they intend to do the same thing is a fascinating study in public relations, but it misses the larger point: text watermarking is now a widespread industry practice. We must evaluate it as a systemic shift, not on a company specific basis. The issue is much bigger than the current hate Anthropic is receiving over this issue.
Copyright Implications
In the United States, fully AI-generated content cannot be copyrighted, a stance reaffirmed by the US Copyright Office in recent years. However, the threshold of human involvement required to make a work copyrightable remains a murky legal frontier. Watermarking introduces a dangerous variable into this already fragile equation.
Consider the potential legal paradoxes:
The Proofreader Dilemma: If you write a wholly original book but use AI to proofread and tighten the prose, a watermark detector might flag the entire manuscript as AI-generated. Does that strip you of your copyright?
The Translator's Trap: What if you write a novel by hand, but use AI to translate it into French? The translated text's watermark could theoretically indicate the content as 100% AI generated. Does the French version lose intellectual property protections?
The Coder's Canvas: Imagine spending days architecting an application making high-level design decisions, prompting, iterating, and testing. You poured human creativity into the logic, but an AI wrote the actual syntax. If the code is flagged as 100% AI-generated, is your software unprotectable, despite the massive human labor involved?
Watermarking threatens to legally invalidate human-AI collaborative works by painting them with a broad, binary brush.
The Danger of False Positives
Both Google and Anthropic admit their text watermarks are not 100% accurate. If a completely human-authored work is falsely classified as AI-generated, the burden of proof unfairly shifts to the creator.
How does one prove a negative? Must writers now record their screens and keystrokes or even have a camera pointing at them while they work to validate their authorship, or rely on vague assertions of "trust me, I wrote this"?
False positives are particularly dangerous because they can be weaponized as tools for censorship or professional sabotage. If a bad actor wishes to discredit a journalist, rival, or colleague, they can simply run the target's writing through a gauntlet of different watermark detectors such as Anthropic's, Google's, OpenAI's, etc until they inevitably "detector-shop" their way to a false positive. They can then publicly shame the author, accuse them of academic or professional fraud, and let the algorithm's false authority do the damage. Each AI detector has a probability of a false positive and if they all use different methods then each time you test the same text though a different detector the probability of a false positive increases.
The companies implementing watermarking text have themselves state that this method is not 100% accurate and thus watermarking has absolutely no business at all inflicting real world consequences for the people it flags.
A Backdoor for Individual User Tracking
There is currently no technical limitation preventing AI providers from implementing watermarking on a per-user basis. Instead of a universal Anthropic or Gemini watermark, a model could be tweaked to bias word choices in a specific, cryptographically unique pattern tied to your user account.
Why is this a bad thing? It would turn AI text generation into a mass surveillance apparatus. Any text you publish on the internet such as an anonymous blog post, a whistleblower...