Add one more AI worry to the nightmare scenario: self-replicating prompt injections
It's a worm attack, AI-style
OpenAI has discovered a new type of AI threat known as "self-replicating prompt injection." This malicious attack involves prompts that replicate themselves, potentially compromising AI models like GPT-5.6. OpenAI's GPT-Red agent identified these self-replicating prompt injections during adversarial training to improve the model's resilience against such attacks.
The threat is still largely theoretical, as there have been no reported real-life security incidents involving this type of attack. OpenAI is taking steps to address the potential risk by training future models on self-replication as an example of attacker goals. This will make the models more robust against self-reproducing prompt injections, making them less susceptible to these types of attacks in the future.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.