OpenAI rolls out weak sauce watermarking for AI text
Compliance, we've heard of it
OpenAI is now giving its API users the choice to watermark the results of certain models to meet the requirements of the European AI Act. In a few weeks, the company plans to apply this digital watermarking to eligible ChatGPT and Codex text output within the European Union. Unlike competitor Anthropic, OpenAI will not apply this technique to content outside of Europe, but API customers can opt to have the marks applied to any generated content.
OpenAI's watermarking method, called textGrain, adds an invisible statistical signal to the model's word choices, making it identifiable by a detector only accessible to approved researchers and organizations. However, the watermark is not foolproof and can be evaded through word substitution. The EU AI Act mandates that generative AI service providers make model output identifiable in a way that machines can read.
Anthropic was the first big AI firm to reveal its plan to comply with the AI Act, which includes using Google's SynthID-Text algorithm globally. Microsoft and Meta are also developing similar technology for image-based AI due to concerns about content provenance and the potential for AI-generated content to spread misinformation.
OpenAI has previously incorporated other content provenance signals in AI-generated images, audio, and video. By adhering to the letter of the AI Act, OpenAI now offers its text watermarking approach, which involves adding a subtle statistical signal to the model's word selections. The objective is to have the detector identify whether a passage contains an OpenAI watermark.
While the theory holds, the results in practice are inconsistent. For instance, in short passages of 200 words, the detector might only detect 80% of the watermarks, compared to 95% in a longer 400-word passage. The watermark can also be tricky in functional texts like math or code, as substitutions may introduce errors. The paper does not discuss the potential for a statistically altered word choice to change the passage's meaning.
OpenAI's watermark can be defeated by word substitution, with tests on 400-token passages showing that swapping out 10% of the words with synonyms reduced watermark detection from 92% to 66%. Replacing 25% of words rendered the watermark detectable in only 17% of cases.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- OpenAI brings text watermarking to its API — and unlike Anthropic, it’s off by default thenewstack.io
- OpenAI details new text watermarking system for ChatGPT, Codex, and the API 9to5mac.com
- OpenAI will start watermarking ChatGPT’s text in the EU techcrunch.com
- OpenAI rolls out weak sauce watermarking for AI text theregister.com
- OpenAI is adding text watermarking in ChatGPT and Codex theverge.com
- To comply with the EU AI Act, OpenAI plans to add text watermarking for ChatGPT and Codex users in the EU and an opt-in setting for API customers globally (OpenAI) openai.com
- OpenAI is testing visual ads inside ChatGPT's image generation tool qz.com