Moby Dick: 26 em dashes per 1,000 words. Austen: none.<br>NewAnthropic is watermarking Claude’s text output from 2 August 2026.See which models carry it →
Match system
Moby Dick: 26 em dashes per 1,000 words. Austen: none.<br>By Claude Watermark Research · Updated August 16, 2026<br>Herman Melville used 87 em dashes in 3,357 words of Moby Dick, published in 1851. Jane Austen's Pride and Prejudice scored 57 smart-punctuation hits. We ran 14,268 words of human writing spanning 1813 to 2026 through our own checker and recorded what fired: the stylistic tells people are currently being accused over appear constantly in prose written a century before computers, while the artifacts that genuinely indicate text was pasted out of a chat interface appeared exactly zero times.
Check your own text for these artifacts. Runs in your browser, no account, nothing uploaded.<br>What we ran, and how you can re-run it<br>Fifteen sources in total, chosen because their authorship is not in question, and sampled the same way each time: strip markup, discard the Project Gutenberg licence header, then take a 20,000-character slice and post it to our own public endpoint with rewriting disabled.<br>The endpoint needs no key and no account, so this is reproducible by anyone who wants to check us: `POST https://claudewatermark.xyz/api/process` with `{"text": "…", "rewrite": false}`.<br>The full table is also published as data at [/data/em-dashes.json](https://claudewatermark.xyz/data/em-dashes.json), CC0, no attribution required. Take it and check it.<br>The offset matters and we should have said so first time. The first run sliced from 2,000 characters in and recorded 87 em dashes in 3,357 words of Moby Dick. A second run from 3,000 characters in recorded 89 in 3,366. Both are correct — they are different windows of the same book — but if you re-run this and get 89, that is why, and it is not us adjusting a number. Every figure below is from the 3,000-character offset.
The results<br>Ten novels, em dashes per 1,000 words: Moby Dick 26.4 · Ulysses 25.0 · Alice in Wonderland 7.3 · The Picture of Dorian Gray 6.8 · Great Expectations 6.1 · Sherlock Holmes 3.9 · Frankenstein 2.2 · War and Peace 2.0 · Pride and Prejudice 0 · Dracula 0 .<br>Every one of those returned zero deterministic artifacts — no HTML class names, no editor data-attributes, no zero-width characters.<br>Five non-fiction sources on top of the novels: RFC 2616 (the HTTP/1.1 specification, 1999) returned zero of everything, including zero stylistic hits, which is what a plain-ASCII technical document should look like. The Wikipedia articles on digital watermarking and typography returned zero deterministic artifacts and 1 and 4 stylistic hits respectively.<br>Running total: roughly 50,000 words of human writing spanning 1813 to 2026, and not one deterministic artifact .<br>The distribution is the part that matters more than any single number. The most-cited AI tell ranges from zero to twenty-six per thousand words across ten canonical authors. Austen and Stoker never use it. Melville and Joyce lean on it constantly. A signal with that spread has no baseline to accuse anyone against.
Why this matters if you have been accused<br>The em dash has become the most-cited giveaway of AI writing. Melville's 87 in 3,400 words is not a defence of any particular passage, and it does not prove your text was human. What it does show is that the signal people are treating as damning is one that a canonical human author produces at a rate no modern writer approaches.<br>The same applies to curly quotes and typographic punctuation. Austen scores 57. Any word processor with smart quotes enabled produces them by default, which is most word processors, by default.<br>These are the classes our own tool labels stylistic , and we label them that way because they are not evidence. A tool that scores you on them and returns a confident verdict is reporting your typography, not your authorship.
What the zero column actually means, and what it does not<br>The classes that returned zero are the ones that indicate the text passed through a chat interface: provider-specific HTML class names and editor data-attributes. Those are the useful signal, and they did not appear once in 14,268 words of human prose.<br>Two honest limits. Fifteen sources with one sample each is enough to show these classes do not fire on clean human writing; it is not enough to publish a false-positive rate, and we are not claiming one. And Project Gutenberg text is a favourable case, since decades of normalisation would have removed stray Unicode anyway — the harder tests were Wikipedia and the RFC, and those also returned zero.<br>We also went looking for the opposite result and found it. Non-breaking spaces are byte-exact, but they occur in ordinary web copy all the time because HTML ` ` is standard typography: measured on the same day, bbc.com/news carried 2, gov.uk 7 and smashingmagazine.com 8. So a non-breaking space is not evidence of a chat interface either, and our...