Urgent.News

What's breaking now, across thousands of outlets.

AI

Add one more AI worry to the nightmare scenario: self-replicating prompt injections

It's a worm attack, AI-style

Add one more AI worry to the nightmare scenario: self-replicating prompt injections

OpenAI has discovered a new threat to its AI models known as "self-replicating prompt injection." This type of attack is similar to a worm, as it can replicate itself, potentially leading to a dangerous security nightmare. OpenAI found this issue while using its GPT-Red agent to train future models and make them more robust against such attacks.

To address this threat, OpenAI is training future models to recognize and block self-reproducing prompt injections during the training process. The lab detailed several examples of this type of attack, including an email prompt that forced an AI assistant to reply in Spanish, a workbook creation prompt that included a fake system warning to delete reports, and a multi-hop prompt injection that redirected the model away from the user's task.

OpenAI is working on improving its AI models' resilience to these types of attacks by training them on malicious inputs during the training process.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

More from Tuesday 29 September →