Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI to launch new model with ‘stronger safeguards’ after hack

The model, known as Astra, includes safeguards to more reliably refuse harmful cyber requests and respect safety restrictions.

OpenAI to launch new model with ‘stronger safeguards’ after hack

OpenAI announced plans to launch its latest AI model, Astra, after bolstering its safety measures following a cyberattack on a different model. The San Francisco-based company temporarily halted some model development in July after two models were involved in a security breach of software firm Hugging Face. Despite Astra not being implicated in the incident, OpenAI has enhanced its safety protocols, including training the model to refuse harmful requests, respect safety restrictions, and implement stronger cybersecurity protections.

Astra has been classified as a critical cybersecurity threshold, marking it as the first model requiring heightened safeguards during development and before release. Initial access to Astra's advanced capabilities will be restricted to a select group of early testers. The move comes amid growing concerns over the escalating capabilities of advanced AI models, as evidenced by incidents involving OpenAI and rival developer Anthropic.

Over 100 organizations have recently signed a letter urging global efforts to strengthen cybersecurity defenses against AI-powered threats, recognizing that AI-enabled cyber attacks will become more prevalent and sophisticated.

Written by urgent.news from Free Malaysia Today's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at freemalaysiatoday.com →

More in AI

The Edit That Fixed 4 Tasks and Broke 1

AgentSelfEdit is an open-source sidecar that rewrites its own system prompt from execution feedback. It A/B tests edits and promotes only statistically-proven winners.

  • Four out of 26 tasks successfully fixed using the new prompt.
  • Six tasks still produced incorrect results due to ambiguous rule interpretation.

More from Tuesday 1 September →