Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

Or how I learned to stop worrying and love dangerous AI

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

OpenAI has pledged to enhance security measures for its upcoming AI model, Astra, amid concerns raised by Anthropic that their model, Fable, lacks sufficient safeguards. OpenAI's Preparedness Framework defines "Astra-level capabilities" as posing a "meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent."

OpenAI claims Astra boasts significant advancements in agentic coding and cybersecurity, and they plan to implement stricter security controls such as isolated testing environments, restricted network access, enhanced model weight protections, encryption, additional monitoring and detection capabilities, and sandboxed execution.

In response to the OpenAI announcement, Anthropic has decided to relax Fable's "fallbacks" or safety mechanisms, which prevent the model from emitting potentially harmful content in response to biology-related prompts. This move comes as China-based AI firms are now fielding competitive open-weight AI models at lower costs, putting pressure on Anthropic to remain competitive in the market.

OpenAI, however, maintains its stance that advanced cyber-capable models should aid defenders in identifying and addressing vulnerabilities before attackers do, but acknowledges that adversaries already possess encryption and various weapons.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

More from Friday 7 August →