AI firms are watermarking generated text – here’s why it won’t work
AI companies have begun embedding watermarks in the output of their models to improve transparency and help crack down on misinformation, disinformation and cheating on homework, but there are limitations to this approach
EU legislation has prompted numerous AI companies to add watermarks to text, images and videos generated by their models, with the aim of helping identify AI-created content and promote transparency. However, experts argue that text watermarking is unlikely to be effective. The EU's AI Act, implemented on August 2nd, mandates watermarking to be in place by December 2nd.
Companies like OpenAI, behind ChatGPT, and Anthropic, which created Claude, are already working on text watermarking for future models. The AI Act allows a grace period for already deployed models, but all must include watermarking by the end of the year.
James Padolsey, from AI safety firm NOPE, believes that watermarking text will be challenging due to the nature of language models. These models generate sentences by selecting the most likely next word, so introducing detectable patterns could be easily altered with minor edits. Padolsey's online tool, declaude, demonstrates this by changing AI-generated text in ways that existing AI detectors fail to recognize, including watermarked text.
However, even without such tools, open-source AI models that do not comply with EU rules will persist, accessible to malicious actors looking to generate AI content in bulk. Furthermore, the EU's rules might not effectively address false positives, where text is mistakenly identified as AI-generated, or when models only play a minor role in creating content. Detecting watermark removal or tampering could also lead to unintended consequences, such as flagging student work that simply utilized an AI tool for assistance.
Written by urgent.news from New Scientist's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.