Urgent.News

What's breaking now, across thousands of outlets.

AI

Why AI Watermarks and Detectors Could Backfire

AI watermarks and detectors may leave us worse off by creating a false sense of confidence in content marked as genuine, writes Nadav Ziv.

Why AI Watermarks and Detectors Could Backfire

AI watermarks and detectors are gaining popularity as trust in online content is becoming increasingly difficult to maintain. Instances such as the Will Smith eating spaghetti video demonstrate the advancements in AI-generated content. By 2025, AI was capable of producing highly realistic content. Deepfakes have become so convincing that experts suggest families establish secret code words to identify one another.

Despite this, research consistently shows that humans are not good at identifying AI-generated content. This can be particularly concerning given the overwhelming amount of AI-generated content on the internet. You might think you can detect offending content, but it often only lasts briefly. Any discernible signal will likely be circumvented by sophisticated actors.

We have witnessed this trend before. During the early days of the internet, visual polish indicated that a website was legitimate. However, this advice has remained largely unchanged even as content creation tools have become more accessible. A 2022 study found that 96% of America's leading colleges and universities still follow outdated advice on evaluating online information.

The most dangerous outcome of this focus on visual cues is the "inverse illusion" - the tendency to believe that the absence of a signal proves something is genuine. Just because a site has no apparent flaws does not mean it is trustworthy. The same principle applies to AI-generated content. Even if visible flaws are present, their absence does not guarantee authenticity.

However, experts often provide simplistic clues for identifying AI-generated content, which can be misleading. For example, during the 2024 election, Stanford Professor Sam Wineburg and the author warned against public officials advising citizens to look for specific visual cues, even though AI-generated content had long since stopped making these mistakes.

Many guides on spotting AI content still follow this poor advice. AI watermarks and detectors, which rely on hidden signals within content, present themselves as a solution to the problem. While these tools cannot always detect signs of AI, they promise that algorithms can. It is important to note that stripping metadata from AI-generated images is relatively easy, and watermarks like SynthID can be removed through various means.

Google admits that detecting watermarked AI text becomes significantly less accurate when users extensively rewrite their content, and it is not designed to prevent motivated adversaries from causing harm. Moreover, open-weight AI models that can run locally, outside platform terms and conditions, ensure the proliferation of unmarked content.

Third-party detectors also have a questionable track record. The author regularly ran AI-generated text through various detectors, with mixed results. The biggest challenge for detectors and watermarks is the "inverse illusion." Just because content lacks a watermark does not mean it was not produced or edited with AI. In one study, 117 students were shown a confident chatbot answer with fabricated facts.

Half of the students did not trust the information, while one student acknowledged that AI-generated content can be both right and wrong at different times. However, relying on AI detectors or hunting for visual clues is not a reliable solution. Instead, it is more effective to consider the reputation and context of the content.

It is relatively easy to create AI-generated content, but faking a strong, credible reputation requires significant effort. Before accepting unfamiliar online content, question the source and check if reputable individuals and organizations confirm its accuracy.

Written by urgent.news from Time's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at time.com →

More in AI

More from Tuesday 25 August →