Anthropic says text watermarking scheme relies on inconsequential words
And other AI model makers are expected to deploy something similar
Anthropic has introduced a text watermarking scheme designed to watermark AI-generated text in compliance with the EU AI Act. This watermark will subtly alter the choice of words in Claude's responses, making them detectable as AI-generated. Traditional watermarks, like those on banknotes or official documents, are more rigid patterns or images. In contrast, digital watermarks can take many forms, and Anthropic's approach involves influencing the word choices made by its large language models.
Large language models, such as Claude, predict the next word in a sequence of words. For instance, when asked to complete the sentence "The weather today was cold and...", Claude might typically respond with "cold" or "gray," but less likely with "sugary." However, when Claude is forced to complete the sentence in a different way, it demonstrates the technique behind Anthropic's watermarking.
By removing the training that encourages engaging and literary-style responses, Claude essentially predicts the next word in a sequence, introducing subtle, context-specific modifications into the generated text.
Anthropic asserts that this watermarking process will not significantly impact the content, creativity, or readability of Claude's text. In internal testing, they found no difference in quality between watermarked and unwatermarked answers when measured by human raters. However, Anthropic emphasizes that this watermarking technique is intended to be applied only to inconsequential text, such as less critical or factual passages, and not to code or high-stakes literature.
They state that in most cases, the example sentences could be completed using interchangeable words without altering the overall meaning.
While Anthropic's watermarking scheme is not overly intrusive and does not involve any personally identifying information, it is expected to be only semi-effective. In their FAQs, they note that some editing should be sufficient to remove the watermark, while a complete rewrite of the text would eliminate the watermarked distinction. Anthropic views this watermarking as a matter of legal compliance, with minimal impact on the speed of their models and no additional cost for serving and using the model.
Written by urgent.news from The Register's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.