Urgent.News

What's breaking now, across thousands of outlets.

AI

Who Is Claude Trying to Kid?

Confessing to being mechanical is the most mechanical trick of all The post Who Is Claude Trying to Kid? appeared first on Nautilus .

Who Is Claude Trying to Kid?

In August 2026, Anthropic unveiled Claude, an AI model that embeds a secret watermark into its generated text and attaches signed provenance metadata to supported files across the globe. This move is in response to the European Union AI Act's transparency code, but it has a more curious purpose. Every sentence Claude writes now carries a cryptic message: "It was me."

The author argues that we should add a fourth law to Asimov's robotics principles: "A robot must not deceive humans by impersonating a human." This watermarking goes beyond mere compliance with laws and regulations. It suggests that we can no longer differentiate between human and machine-generated text just by reading it. The only reliable proof of a human hand is now the presence of errors, anomalies, or hesitations.

This idea is rooted in the uncanny valley, a concept introduced by Japanese roboticist Masahiro Mori in 1970. As robots become more human-like, our affinity for them rises, then drops sharply just before they become indistinguishable from humans. The valley represents the uncanny or unsettling feeling we experience when confronted with something that is almost, but not quite, human.

The valley has opened up in the realm of video, images, and text. Consider the phrase "AI slop." It's a term of disgust, much like poor text or images. However, it's not the low-quality content that elicits this reaction, but the nearly perfect content that's almost, but not quite, human. The repetitive bullet points, sycophantic throat-clearing, and the overuse of style guides like the em dash have all contributed to our uncanny experience of AI-generated text.

The author notes that humans have been laboring to prove that they aren't machines for decades, from typing fire hydrant descriptions to identifying Sarah Connor in photos. Now, the question has shifted from "Can machines think?" to "Can humans still demonstrate that they typed this themselves?" The answer is becoming increasingly clear: humans cannot reliably prove their authorship in the age of AI.

To detect synthetic text, information scientists Maurice Jakesch, Jeffrey Hancock, and Mor Naaman found that people rely on shallow heuristics, such as the use of first-person pronouns, contractions, and informal language. However, these heuristics are easily exploited by language models, leading to machine-generated text being rated as more human than human-written text.

The watermarking solution is ironic and clever. It embeds a deliberate, statistically detectable irregularity in the text, much like the subtle musical notation added to piano synthesis to give it character and prevent it from sounding perfect and fake. This hidden mark, which the machine willingly adopts, is an admission that we can no longer rely on human intuition to detect AI-generated content.

Written by urgent.news from Nautilus's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at nautil.us →

More in AI

How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise

Is the model actually getting worse, or did I just get unlucky on a handful of prompts? That question is why threads like "is it just me or is it dumber today" keep recurring, and it is the question a…

  • Freeze prompts and pin harness to ensure consistent treatment of the model over time
  • Calibrate panel with 78 sometimes-right questions for statistical power
  • Compare item scores with clustered standard errors to avoid confounding factors

More from Wednesday 30 September →