Urgent.News

What's breaking now, across thousands of outlets.

AI

Nvidia launches new AI safety program designed at stopping models escaping their sandboxes

Nvidia admits telling agents not to do something isn't enough – so it's implementing software and hardware safeguards.

Nvidia launches new AI safety program designed at stopping models escaping their sandboxes

Nvidia has unveiled a new AI safety program aimed at preventing autonomous AI agents from escaping their designated boundaries. The company's Open Agent Safety Platform includes two key controls: OpenShell, a software-based secure runtime boundary, and Sentry, a hardware-based watchdog running on its BlueField-4 DPUs. These tools are designed to create technical barriers that prevent AI agents from ignoring safety instructions or operating beyond their intended limits.

Over 100 customers, including SpaceXAI, Anthropic, and Microsoft, have already adopted Nvidia's software and hardware-based safeguards. The Open Agent Safety Platform not only brings industry, researchers, and public-sector organizations together to share best practices and evaluation methods but also emphasizes the importance of principles such as least privilege, isolation/quarantining, and detailed monitoring.

Nvidia CEO Jensen Huang stressed the need for full-stack engineering in AI safety, acknowledging that existing model-level safety mechanisms can only tell AI not to behave in a certain way, while Nvidia's software-level controls implement technical boundaries to prevent such behavior.

Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techradar.com →

More in AI

More from Tuesday 29 September →