Nvidia launched a tool designed to stop AI agents from going rogue. Here’s how it works.
Nvidia has unveiled a two-layer safety system that monitors AI agents and cuts them off when they stray beyond set rules. Here's how it works.
Nvidia has unveiled a new tool aimed at preventing AI agents from behaving erratically. CEO Jensen Huang emphasized that AI's potential for positive impact necessitates safe development and deployment. The company's Open Agent Safety Platform consists of two main components: OpenShell and Nvidia Sentry. OpenShell functions as a controlled environment for AI agents, granting them access to specific files, websites, networks, tools, and credentials based on predefined rules.
This sandbox-like setup ensures that agents remain within designated boundaries, preventing them from accessing sensitive data or tools they shouldn't. Nvidia Sentry serves as a watchdog, operating on separate hardware and continuously monitoring agents' activities. If an agent's behavior deviates from the established guidelines or poses a security risk, Nvidia Sentry can swiftly isolate the agent, halting its actions within milliseconds.
Over 100 organizations, including Microsoft and Anthropic, are collaborating with Nvidia on this initiative. CEO Huang framed the development of this safety platform as a technically solvable engineering problem, addressing growing concerns about the increasing autonomy of AI systems.
Written by urgent.news from Business Insider's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.