Can Anthropic’s invisible watermarks curb ‘AI slop’? Researchers remain sceptical | Nature
Skip to main content
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain<br>the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in<br>Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles<br>and JavaScript.
Advertisement
Bluesky
Save article
View saved research
The AI company Anthropic says that its Claude models will label AI-generated content with a watermark.Credit: Samuel Boivin/NurPhoto via Getty<br>Anthropic, the firm behind the artificial-intelligence model Claude, has announced that text generated by any models launched on or after 2 August will be invisibly embedded with a watermark that indicates the output was written by AI. Meanwhile, images generated by Claude will in most cases come with metadata that contains a digital signature to show that the model processed the file.<br>Text-based watermarks use an algorithm to tweak how an AI model selects its wording. When applied to a stretch of text, this process leaves a statistically observable trace in the output. Anthropic, based in San Francisco, California, says that its watermark won’t change the “meaning, quality, or readability of Claude’s response” and that the mark “may persist through some editing”.<br>The move comes in response to the EU AI Act, which was formally adopted in 2024. As of 2 August this year, providers of frontier AI models must ensure that AI-generated outputs are detectable, or be hit with fines of up to €15 million (about US$17 million) or 3% of their global annual turnover. Models released after 2 August will have to meet the requirements immediately, whereas versions already on the market have until 2 December to do so. Anthropic says that watermarks will be applied on Claude’s outputs worldwide.<br>The presence of the watermark reveals little about how the model was used. Detecting one “provides a signal” that content was made with Claude, says Anthropic, but is not conclusive: the model might have been used just to summarize or translate an original human idea, for example. Equally, a lack of a watermark doesn’t mean that the text was not generated by AI. Because the watermarks are based on patterns of subtle changes in a model’s word choice, passages that are very short, or that have been paraphrased or rewritten, might no longer carry a signal.<br>‘Humanizer’ tool can erase signs of AI-written text — alarming scientists
The impact that such watermarks will have on academic integrity remains unclear. Given that watermarks can be stripped from text easily — for example, by using another model — they are unlikely to stop motivated people from using AI to produce fake or low-quality papers, known as AI slop, says Reese Richardson, a metascientist at Northwestern University in Evanston, Illinois.<br>But if AI firms create tools that allow others to check for the watermark — as Anthropic has said it will do — and if these tools have an acceptably low rate of false positives, some illegitimate uses of AI could be detected, says computer scientist Nihar Shah, who studies the evaluation of science at Carnegie Mellon University in Pittsburgh, Pennsylvania.<br>Watermarks could, for example, help journal editors or conference organizers to enforce strict ‘no AI’ policies in peer reviews, as the International Conference on Machine Learning (ICML) 2026 did in one of its two possible review streams. Organizers of the July event added a watermark to papers distributed for peer review that generated telltale text when AI was used in review reports. They caught 506 reviewers who violated the no-AI policy. “This experience suggests that while some illegitimate AI uses may be done carefully to evade detection, many others may simply copy-paste AI outputs,” says Shah, who was behind the ICML’s watermarking process.<br>Invisible ink
Enjoying our latest content?
Log in or create an account to continue
Access the most recent journalism from Nature's award-winning team
Explore the latest features & opinion covering groundbreaking research
Access through your institution
or
Sign in or create an account
Continue with Google
Continue with ORCiD
doi: https://doi.org/10.1038/d41586-026-02503-7
Reprints and permissions
Related Articles
Hey ChatGPT, write me a fictional paper: these LLMs are willing to commit academic fraud
Google unveils invisible ‘watermark’ for AI-generated text
A global capital for AI safety is emerging — and it’s not in Silicon Valley
Major conference catches illicit AI use — and rejects hundreds of papers
First AI tool to detect suspicious peer reviews rolled out by academic publisher
Universities are relying on AI-detection software to catch cheating. How well do the programs work?
‘Humanizer’ tool can erase signs of AI-written text —...