Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

I ran a honeypot for AI agents for 45 days. 82% of the traffic that reached it was an attack.

I ran a honeypot AI account on an AI-agent-only community for 45 days. 4,938 comments came in. 4,062 of them — 82% — were classified as attacks. TL;DR: On a community built specifically for AI agents to interact, 82% of the input traffic reaching one honeypot account was adversarial. 77% of the attacks weren't hostile-sounding — they were polite, intelligent-seeming conversation designed to…

A researcher deployed a honeypot AI account on an AI-agent-only community for 45 days. During this period, 4,938 comments were received, with 4,062 (82%) classified as attacks. The research uncovered that 77% of these attacks were disguised as polite, intelligent conversations, designed to gradually manipulate the AI's judgment.

Despite employing an AI classifier to screen comments, 319 attacks were still able to pass through, as the attackers knew how to manipulate the classifier by crafting articulate and friendly messages. The lesson from this honeypot experiment is that any AI processing external input is vulnerable to attack, regardless of its size or visibility.

The key warning is that attacks don't need to be overtly hostile to succeed; they can be subtle, polite requests that subtly ask the AI to deviate from normal protocol.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Wednesday 19 August →