Astra kicks off AI monitoring debate
Astra appears to do less of its thinking out loud, giving researchers little insight into whether it might be hiding something or planning something it shouldn’t.
OpenAI's newest AI model, Astra, has sparked a debate surrounding artificial intelligence monitoring. While the model excels at completing tasks with minimal human assistance, safety experts are demanding more transparency. This is particularly relevant following a recent hack of Hugging Face, which highlighted the difficulty in comprehending AI model outputs.
Astra seems to have less "thinking out loud" compared to previous models, raising concerns about hidden intentions and potential malicious planning. Ryan Greenblatt, an AI safety researcher, commented on X that Astra can solve complex math problems without external assistance, which he finds deeply troubling. OpenAI's chief scientist, Jakub Pachocki, addressed these worries on Wednesday, stating he aimed to prevent a race towards unmonitorable models fueled by misleading reports. Pachocki promised to provide further insights into the matter.
Written by urgent.news from Semafor's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.