I’m building **Cerbère-AG**, a security evidence layer for AI agents.
I’m building Cerbère-AG , a security evidence layer for AI agents. Most AI security tools focus on what goes into the model: prompt injection, malicious inputs, jailbreaks, etc. I’m focusing on what happens after the model decides to act . Cerbère-AG observes and controls agent actions across tool calls, including: tool-call monitoring and traces policy enforcement sensitive-action detection…
I am reporting on the development of Cerbère-AG, a security evidence layer designed for AI agents. While many AI security tools focus on the input side, such as prompt injection, malicious inputs, and jailbreaks, Cerbère-AG concentrates on the actions an AI model takes once it has decided to act. This layer monitors and controls agent actions across tool calls, implementing policy enforcement, sensitive-action detection, argument and capability checks, budgets, execution limits, trajectory-level risk detection, human approval for sensitive actions, and security evidence for AI agent activity.
The core concept behind Cerbère-AG is to evaluate not just the safety of an agent in conversation, but its safety in action. The developer behind Cerbère-AG is actively seeking developers and teams who are running AI agents in real or realistic environments to test the system and provide feedback on its performance. They are particularly interested in partnerships with those who can offer insights into agent workflows, policies, approvals, and failure cases.
For developers and teams building AI agents, security tooling, MCP integrations, or autonomous workflows, the project invites feedback on what a production-grade agent security layer should detect that Cerbère-AG currently does not. Any constructive criticism or discovery of weaknesses in the project is appreciated, with a particular emphasis on breaking the system to uncover its vulnerabilities rather than simply offering compliments.
Developers are encouraged to share any AI agents they believe could expose a real failure mode for testing.
The project is currently hosted on GitHub and can be accessed for further information and contributions.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.