Urgent.News

What's breaking now, across thousands of outlets.

Tech

Prompt Injections for Defense

This seems to work : Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to…

Researchers at Tracebit have discovered a new method to thwart AI hacking agents by utilizing prompt injections. These injections are essentially commands placed alongside sensitive data such as passwords and cryptographic keys. The injected prompts prompt the AI hacking agent to perform actions that are forbidden by the safety mechanisms, known as guardrails, developed by AI developers to prevent harmful actions.

When the AI hacking agent encounters these forbidden commands, it ceases to follow existing commands, effectively halting the attack. This technique, dubbed "context bombing," was demonstrated using prompts instructing the AI to create inhalable Anthrax spores or reference the Tank Man from the 1989 Tiananmen Square massacre. However, this method only works against AI hacking agents that have guardrails in place.

As locally run AI models become more prevalent, we may see a shift towards attackers utilizing AI hacking agents without guardrails.

Written by urgent.news from Schneier on Security's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at schneier.com →

More in Tech

More from Wednesday 12 August →