Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI to launch new model with ‘stronger safeguards’ after hack

The model, known as Astra, includes safeguards to more reliably refuse harmful cyber requests and respect safety restrictions.

OpenAI to launch new model with ‘stronger safeguards’ after hack

Following a cyberattack on an AI model separate from OpenAI's, the San Francisco-based company announced plans to launch its new, powerful model, Astra, equipped with enhanced safety measures. After pausing model development for two weeks in the summer due to two testing models involved in a security breach at software company Hugging Face, OpenAI has implemented stronger safeguards for Astra.

These measures include training the model to refuse harmful cyber requests, implementing additional protections against misuse, and monitoring to stop unauthorized activity. Astra has been classified as reaching a critical cybersecurity threshold, signifying OpenAI's belief in its ability to discover and exploit cybersecurity weaknesses.

As Astra is slated for release, access to certain capabilities will be restricted, with advanced features made available to a select group of early testers.

Written by urgent.news from Free Malaysia Today's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at freemalaysiatoday.com →

More in AI

The Edit That Fixed 4 Tasks and Broke 1

AgentSelfEdit is an open-source sidecar that rewrites its own system prompt from execution feedback. It A/B tests edits and promotes only statistically-proven winners.

  • Four out of 26 tasks successfully fixed using the new prompt.
  • Six tasks still produced incorrect results due to ambiguous rule interpretation.

More from Tuesday 1 September →