Anthropic says text watermarking scheme relies on inconsequential words
And other AI model makers are expected to deploy something similar
In an effort to "watermark" text that Claude has generated and comply with the EU AI Act, Anthropic unveiled a plan on Friday to modify its bots' choice of words in a way that would be detectable as AI.
Traditional watermarks are patterns or images overlaid on currency, postage, or official documents as an assertion of authenticity. In the digital realm, the term is more flexible and can refer to a variety of techniques for applying an identifier to electronic data.
Anthropic's approach involves influencing inconsequential word choices made by its models, a technique introduced in Google DeepMind's SynthID-Text paper.
To oversimply things, large language models work by predicting the next word in a sequence of words. The AI biz explains that while composing sentence output like "The weather today was cold and…", a model like Claude might respond with words like "cold" or "gray" and would be unlikely to respond with a word like "sugary."
That's the theory, but when actually asked to complete that sentence, Claude Opus 4.8 went a bit overboard: "…crisp, the kind of cold that nips at your fingertips and turns your breath to little clouds. The sky was a pale, washed-out blue, and everything felt sharp and clear.
"Want me to take it somewhere specific — cozy, gloomy, cheerful? Or keep going with the same tone?"
But remove whatever training has been applied to promote engagement and simulate literary style, and that's basically what Claude is doing here – predicting the next word in a sequence.
Anthropic asserts that in most cases, the example sentence could be completed by either "cold" or "gray" and "the meaning of the sentence is largely the same either way."
The watermark gets generated by deviating from the predicted word to something else. A different source of randomness is used and that can be detected with a digital key.
As Google DeepMind researchers explain in their paper, "Generative watermarking works by carefully modifying the next-token sampling procedure to inject subtle, context-specific modifications into the generated text distribution. Such modifications introduce a statistical signature into the generated text; during the watermark detection phase, the signature can be measured to determine whether the text was indeed generated by the watermarked LLM."
Anthropic insists this will be done with low-stakes passages in a way that won't alter the meaning. "In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text," the company said, adding that in a controlled study, human raters saw no difference in quality between watermarked and unwatermarked answers.
That assumption hinges on not applying the watermark to any consequential text. As Anthropic puts it, "Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text." The biz goes on to say that the situation is similar with code – the watermarking algorithm can't simply start swapping method names.
In the context of literature, the notion that some words are interchangeable is likely to raise a few hackles. While it may be a satisfying thought experiment to imagine Claude emitting, "It was the best of times, it was the least of times…" or "Telephone me Ishmael", anyone trying to pass off generated text as serious writing probably should face whatever social backlash watermarking may entail.
On the plus side, Anthropic's flavor of watermarking isn't excessively intrusive. It doesn't involve any personally identifying information and only serves to indicate that Claude was probably involved at some stage of the creation of the marked text.
What's more, the technique is expected to be only semi-effective. In its FAQs, Anthropic points out that some amount of editing should erase the watermark.
"Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will," the company said. "In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated."
In all likelihood, Anthropic doesn't care if its watermarking scheme can be defeated. The company's post makes clear that it is doing so as a matter of legal compliance and has chosen a solution that doesn't raise costs.
"Watermarking has a negligible impact on the speed of models, and because it produces no extra tokens, the model is the same price to serve and use," the biz said.
Hey Claude, what's another word for performative compliance? ®
Originally published on The Register


