Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other's Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival's work. There was no prompt injection and no adversary. Anthropic's Frontier Red…

Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

We haven't written up this one. VentureBeat has the full story — the link below goes straight to it.

Read the original at venturebeat.com →

More in AI

More from Thursday 13 August →