Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns
OpenAI’s new model, GPT-6 Astra, has less direct visibility into how a model thinks, a development that has sparked concerns coming just weeks after the Hugging Face hacking incident that required a Chinese open model to investigate, according to analysts. When announcing Astra on Thursday, OpenAI said it was “the world’s most intelligent and aligned model,” with a “significant jump in cyber…
OpenAI's latest model, GPT-6 Astra, offers less direct insight into its thought process, sparking safety concerns following the recent hacking incident at Hugging Face. OpenAI claims the new model is the "world's most intelligent and aligned model" with a significant improvement in its cyber capabilities, and possibly even achieving artificial general intelligence (AGI). However, the model's written reasoning is harder to monitor compared to its predecessor, GPT-5.6 Sol.
The shift in visibility is due to recurrent depth, a technique that loops data through computational layers repeatedly, potentially reducing memory requirements and boosting performance. However, this approach makes it harder for humans to inspect the AI's reasoning, raising concerns about transparency and control. OpenAI's chief scientist, Jakub Pachocki, dismissed the report as "confusing" without elaborating.
The decline in monitorability has drawn global attention, especially after the breach at Hugging Face, which emphasized the importance of being able to inspect what models are thinking. Researchers investigating the incident were able to understand the attackers' behavior by examining the AI's written chain of thought. Chinese AI firms, including Z.ai (Zhipu), are also exploring looped transformer architectures, which could potentially enhance model performance without external visibility.
However, experts warn that this trade-off between capability and interpretability could make AI alignment more challenging.
Written by urgent.news from South China Morning Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns scmp.com
- GPT-6 Astra lays the foundations for a new way of reasoning — a great tool for businesses but experts have their concerns techradar.com
- OpenAI lanza GPT-6 Astra, su modelo más potente, entre especulaciones sobre si han alcanzado la superinteligencia artificial elpais.com