Even Claude says AI watermarking is no ‘silver bullet’
Claude has questioned whether watermarking AI-generated text can ever provide a foolproof way to identify machine-written material, just days after its maker, Anthropic, began embedding invisible marks into the chatbot’s own output. Asked by City AM whether AI-generated content should be watermarked, Claude warned that simple marks can be “cropped, screenshotted, or edited out”, while [...]
Claude, the AI chatbot developed by Anthropic, has expressed skepticism about the effectiveness of watermarking AI-generated text as a foolproof method for identifying machine-written material. Just days after Anthropic began embedding invisible marks into Claude's output, the chatbot warned that simple watermarks can be removed through cropping, screenshots, or editing.
More sophisticated text watermarks may be defeated through paraphrasing or reformatting. While Claude supports watermarking, it emphasizes that it is only one component of a broader system for identifying AI content, rather than a "silver bullet." Claude noted that the technology is weakest against individuals who would intentionally misuse AI-generated content for fraud, disinformation, or non-consensual imagery, as they are likely to employ methods to strip away the watermark.
The chatbot explained that these individuals could strip the watermark by rephrasing the text or translating it into a different language and back again. Anthropic, the company behind Claude, has begun watermarking Claude-generated text to comply with new EU transparency rules. The company has also made the technology available for third-party verification.
However, Anthropic stresses that finding the watermark does not definitively prove that Claude authored a piece of content. Genuine human-written work can be watermarked after being translated or substantially edited by Claude, while text initially generated by the chatbot may lose its detectable signal if heavily rewritten. Claude acknowledged that reliable text watermarking that withstands editing is not yet "close to solved."
The chatbot expressed cautious optimism about applying the technology to images or video, where invisible marks are more difficult to remove without degrading the content and where deepfakes or fake evidence pose greater risks. For text, Claude emphasized the importance of disclosure, provenance standards, and media literacy when publishing AI-assisted material.
Anthropic's move aligns with the EU AI Act, which pushes developers to make machine-generated material identifiable. The company has signed the EU's voluntary Code of Practice and plans to extend watermarking to older Claude models. Other AI firms, such as OpenAI and Google, are also working on marking text outputs as part of their commitments to the European rules.
However, Elon Musk's xAI, the only major LLM maker that did not sign the EU code, will still need to comply with applicable AI Act requirements. Some Claude users have criticized Anthropic's decision, expressing concern that genuine human work could be mistakenly identified as AI-assisted if the chatbot subsequently edits it. Anthropic stressed that its watermark does not identify a particular person, company, or conversation.
Written by urgent.news from City AM's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.