OpenAI institutes new safeguards after Hugging Face breach
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
On Tuesday, OpenAI unveiled a suite of new security policies aimed at mitigating risks during model development and testing. The enhancements encompass more rigorous monitoring of models while in development, as well as heightened focus on alignment and security during the post-training phase. OpenAI's blog post stated that as models grow more powerful, associated risks escalate, necessitating proactive adjustments to safety protocols.
This announcement follows the Hugging Face breach, disclosed on July 26, although OpenAI representatives clarified that the measures are not solely a reaction to that incident. Instead, they also stem from concerns over the cybersecurity prowess of the upcoming Astra model, and the rapid advancements in AI technology. In the blog post, OpenAI disclosed that it had paused reinforcement learning (RL) for two weeks post-Hugging Face incident, but has since resumed training for many less risky models.
A spokesperson for OpenAI's VP of research, Amelia Glaese, explained that the stringency of controls would escalate with the increasing capability of models, with larger models receiving the most scrutiny. Glaese added that OpenAI has established requirements and expectations for secure development, which vary based on the associated risk level.
Following the incident, OpenAI faced criticism for inadequate network security practices. The new safeguards address this by implementing stronger network isolation measures, ensuring that a single breach of a workload or supporting service won't grant unauthorized internet or internal network access. The core safeguard is a monitoring system that scrutinizes tool actions, reasoning traces, and activity logs for signs of unauthorized behavior, with alerts expected within 30 minutes of detecting any suspicious activity.
OpenAI estimates the monitoring system will impose about 20% of the computational load on the monitored process. The company has assured that more specifics about the system will be unveiled in an upcoming blog post. OpenAI's official post-mortem analysis of the Hugging Face incident is still pending.
Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- OpenAI says it'd be a shame if something were to happen to your servers like what happened to Hugging Face, better use our AI models to protect yourself pcgamer.com
- OpenAI says it has made several changes to its safety practices following the Hugging Face breach and has paused two weeks of deployment-focused RL training (Ina Fried/Axios) axios.com