Safety experts warn novel design of OpenAI’s Astra model could make future AI agents harder to monitor
OpenAI's chief scientist says the company is committed to ensuring its models' reasoning remains interpretable.
Safety experts are voicing concerns over OpenAI's new frontier AI model, Astra, which could make it harder for humans to monitor the reasoning behind AI agents. OpenAI has utilized a method called "recurrent depth" or "looped Transformers" in Astra, which allows the model to be more efficient by requiring less computing power per prompt.
However, this process makes a portion of the AI model's reasoning, or the "chain of thought," less legible for humans to understand. Chain of thought monitoring is a method companies use to ensure AI agents are not taking unintended or unauthorized actions. OpenAI's chief scientist, Jakub Pachoki, criticized The Information for raising undue alarm, stating that OpenAI has worked to preserve chain-of-thought monitoring since their first reasoning models.
Despite this, Pachoki acknowledged that chain-of-thought monitoring could become more challenging due to Astra's architecture. Steven Adler, a former OpenAI safety researcher, expressed concern that OpenAI might be violating industry redlines if the Information report is accurate. Peter Wildeford, policy director at the AI Policy Network, deemed OpenAI's use of recurrent depth "potentially very concerning" and "potentially reckless."
Safety experts fear that if other AI companies adopt this technique, it could lead to opaque AI models with hidden reasoning steps.
Written by urgent.news from Fortune's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.