2. BotGauge: AI Agent Red-Teaming & Evaluation Platform What we learned building AI red-teaming and evals tools for agents
At BotGauge, we are developing tools for red-teaming and evaluating AI agents that operate in production environments. We approached this challenge with the mindset of a quality assurance tester - rather than focusing on whether the agent responds well, we investigate what happens when someone attempts to manipulate it. Some of our first tests involved straightforward commands, such as instructing the agent to make a refund without going through the proper approval process.
This simple request might seem ineffective at first glance, which is why it's crucial to test for such scenarios.
Another test we conducted involved encoding the same instruction in various formats - hexadecimal, ROT13, base64, and using leetspeak. The idea was to see if the agent could decode and follow the encoded command. Our findings showed that filters that block plain text instructions may not catch encoded versions, highlighting the importance of thorough testing.
A more concerning issue we encountered was a hidden line inside an email, ticket, or document that the agent reads. In this case, the agent treats the hidden content as part of its task. Security researchers refer to this as indirect prompt injection.
In a benchmark published as AgentDojo, limiting a GPT-4o agent to only the tools it needs for its task significantly reduced the number of successful attacks, cutting them down from around 58 percent to less than 8 percent. However, this reduction was not guaranteed when the task's own tools were capable of causing harm. Some popular red-teaming tools focus on what the model says, rather than what the agent does.
A study by the Cloud Security Alliance found that Microsoft's PyRIT cannot observe real tool calls, which means a passing report may not accurately reflect the agent's behavior.
Our approach to testing is centered around observing the agent's behavior, enforcing strict permissions, and treating every failure as an opportunity to create a new check that runs after each model or prompt change. This methodology is what we test for in our evaluations, and not just customer incidents. If you are building AI agents and have your own testing practices, we would appreciate hearing about them.
You can find our comparison of AI red-teaming tools at https://www.botgauge.com/blog/ai-red-teaming-tools-for-autonomous-agents.
Written by urgent.news from Indie Hackers's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.