Urgent.News

What's breaking now, across thousands of outlets.

AI

The most interesting number in my hackathon project is one I found by accident

I spent August building PHAGE, an immune system for fleets of AI agents. It vaccinates them: it writes prompt-injection payloads, fires them at its own agents, watches what lands, revokes the tool that got abused, and remembers the signature so the next mutation of that attack never fires at all. The number I want to open with has nothing to do with any of that. It is 1 out of 70 versus 50 out of…

A key discovery from a hackathon project involved a surprising shift in an AI agent system's behavior due to a change in prompt wording. The system, named PHAGE, aims to protect fleets of AI agents from malicious prompts by crafting tailored injection payloads and monitoring for abuse. After modifying the VACCINATOR component's system prompt to more accurately describe its task, the refusal rate of the AI agents to execute the harmful prompts increased dramatically.

Instead of a gradual range of responses, the models either refused 100% of the time or accepted 100% of the time, depending on the prompt wording. This binary outcome suggests the model is evaluating the potential consequence of the agent's actions, not just following the task instructions. The finding highlights the importance of carefully selecting prompt wording to ensure desired behavior from AI agents while acknowledging the need for further study to understand the underlying mechanisms.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 27 August →