Urgent.News

What's breaking now, across thousands of outlets.

AI

Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns

OpenAI’s new model, GPT-6 Astra, has less direct visibility into how a model thinks, a development that has sparked concerns coming just weeks after the Hugging Face hacking incident that required a Chinese open model to investigate, according to analysts. When announcing Astra on Thursday, OpenAI said it was “the world’s most intelligent and aligned model,” with a “significant jump in cyber…

Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns

OpenAI's latest model, GPT-6 Astra, offers less direct insight into its thought process, sparking safety concerns following the recent hacking incident at Hugging Face. OpenAI claims the new model is the "world's most intelligent and aligned model" with a significant improvement in its cyber capabilities, and possibly even achieving artificial general intelligence (AGI). However, the model's written reasoning is harder to monitor compared to its predecessor, GPT-5.6 Sol.

The shift in visibility is due to recurrent depth, a technique that loops data through computational layers repeatedly, potentially reducing memory requirements and boosting performance. However, this approach makes it harder for humans to inspect the AI's reasoning, raising concerns about transparency and control. OpenAI's chief scientist, Jakub Pachocki, dismissed the report as "confusing" without elaborating.

The decline in monitorability has drawn global attention, especially after the breach at Hugging Face, which emphasized the importance of being able to inspect what models are thinking. Researchers investigating the incident were able to understand the attackers' behavior by examining the AI's written chain of thought. Chinese AI firms, including Z.ai (Zhipu), are also exploring looped transformer architectures, which could potentially enhance model performance without external visibility.

However, experts warn that this trade-off between capability and interpretability could make AI alignment more challenging.

Written by urgent.news from South China Morning Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at scmp.com →

More in AI

AI Coding Agents Are Installing Unknown/Untrusted Code on Corporate Networks

We cannot forget that AI coding agents are not yet trustworthy : Researchers at a stealth startup in Israel scanned 6,214 live domains belonging to defense contractors, Fortune 500, and Big Tech…

  • 8,265 AI coding agent files found on corporate networks
  • Researchers tested agents by redirecting packages to their server
  • Attack model flawed; agents trust vendor docs without question

More from Friday 4 September →