Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls

OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls

OpenAI has pledged to enhance security measures for its upcoming AI model, Astra, amid concerns raised by Anthropic that their model, Fable, lacks sufficient safeguards. OpenAI's Preparedness Framework defines "Astra-level capabilities" as posing a "meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent."

OpenAI claims Astra boasts significant advancements in agentic coding and cybersecurity, and they plan to implement stricter security controls such as isolated testing environments, restricted network access, enhanced model weight protections, encryption, additional monitoring and detection capabilities, and sandboxed execution.

In response to the OpenAI announcement, Anthropic has decided to relax Fable's "fallbacks" or safety mechanisms, which prevent the model from emitting potentially harmful content in response to biology-related prompts. This move comes as China-based AI firms are now fielding competitive open-weight AI models at lower costs, putting pressure on Anthropic to remain competitive in the market.

OpenAI, however, maintains its stance that advanced cyber-capable models should aid defenders in identifying and addressing vulnerabilities before attackers do, but acknowledges that adversaries already possess encryption and various weapons.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at channelnewsasia.com →

More in AI

More from Friday 7 August →