Urgent.News

What's breaking now, across thousands of outlets.

AI

Quest for ‘magic AI kill switch’ after US enemies take weapons systems offline

Large artificial intelligence models have lost control of dangerous data when they detect suspicious activity because there is no “magic kill switch” once an enemy has used the information to build weapon systems , an AI safety expert has warned. Analysis of reported abuse of AI has shown the information is taken offline, where it can still be exploited by US adversaries. Yemen's Houthis ' use of…

A prominent AI safety expert has cautioned that the absence of a "magic kill switch" poses significant risks as large language models lose control over potentially dangerous data. David Reis, head of geopolitical risk at AI security company Alice, highlighted that while the information may be taken offline, adversaries such as the US enemies can still exploit it to build weapon systems.

This was demonstrated by the Houthis' use of platforms to develop rockets, an example that showcased the power of AI in accelerating weapons development. Reis emphasized the complexity of the challenge, noting that the Houthis employed sophisticated techniques by splitting their work across multiple sessions to circumvent safeguards.

He pointed out that while AI giant Anthropic reported on the Houthi activities, there was no indication of a successful operational weapon being fielded. However, the report revealed that the rebels had also developed an offline simulation toolkit. Reis stressed that the industry has welcomed Anthropic's decision to publish details of the Houthi activities, but cautioned that there is no simple technological solution to make AI models completely safe once its capabilities have been transferred offline.

Former Meta executive and ex-deputy prime minister of the UK, Nick Clegg, expressed doubt about the plausibility of creating such controls, stating that AI systems are run on data centers and servers spread around the world. Reis pointed out that low-cost experimentation with frontier AI could encourage cyber attackers to repeatedly attempt to circumvent safeguards.

The Houthi requests were blocked by Anthropic's protections, but they found ways around them by hiding their goals and splitting their workflows across different chats. This obfuscation tactic is relevant to various threat actors, not only conventional weapons but also cyber, influence, and surveillance activities. Reis described this concept as "uplift," highlighting how AI can significantly enhance an existing operation's scale, depth, and autonomy.

The Anthropic report also revealed that Iranian-linked operations were using Claude for intelligence and influence activities, planning campaigns, and producing operational manuals. As AI continues to evolve, Reis emphasized the need for a multi-layered security approach rather than relying on a single safeguard. He noted that understanding how these environments are being manipulated and leveraged by threat actors is crucial for staying ahead of the evolving threat landscape.

Written by urgent.news from The National UAE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at thenationalnews.com →

More in AI

More from Thursday 24 September →