Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

Or how I learned to stop worrying and love dangerous AI

OpenAI pledges to add Astra security as Anthropic loosens Fable's leash

OpenAI has pledged to enhance security measures for its upcoming AI model, Astra, following concerns over potential cyber capabilities in unreleased models. The company defines "critical cyber capabilities" as advanced features that could introduce new threats, requiring robust safeguards during development. OpenAI claims its internal evaluations of Astra show significant progress in areas like agentic coding and cybersecurity.

To address these issues, OpenAI plans to introduce stricter security controls, including isolated testing environments, restricted access to networks and tools, enhanced protections for model weights, increased monitoring, and sandboxed execution. The company will also pause Astra testing in environments lacking these security measures and provide recommendations to third-party testing partners on safely conducting high-risk evaluations.

Additionally, OpenAI intends to implement thought policing during Astra's pre-release stage, focusing on monitoring risky actions and misalignment. However, this internal commitment may not necessarily translate to commercial operation.

In contrast, Anthropic has loosened restrictions on its model, Fable, allowing for increased likelihood of interactions in sensitive areas like biology. Anthropic's decision comes amidst increased competition from Chinese AI firms producing open-weight models at lower costs. OpenAI argues that advanced cyber-capable models can help defenders identify and address vulnerabilities before attackers take action.

Nonetheless, the effectiveness of such an approach remains to be seen, as adversaries already possess encryption and various other weapons. OpenAI may believe it can offer exclusive access to its most capable models, but history suggests any such advantage is temporary. Instead, the company should prioritize building robust defenses over perpetually playing catch-up.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

More from Friday 7 August →