Your next OpenAI API timeout might not be a timeout at all
OpenAI said Tuesday that its upcoming Astra model is the company’s first to reach the Critical cybersecurity threshold in its The post Your next OpenAI API timeout might not be a timeout at all appeared first on The New Stack .
OpenAI announced on Tuesday that its upcoming Astra model is the company's first to achieve the highest cybersecurity threshold in its Preparedness Framework, a designation reserved for models that can identify vulnerabilities and create exploits with minimal human assistance. Astra will receive more stringent monitoring as a result, with the ability to interrupt an agent mid-task, depending on the platform it's running on.
For ChatGPT and Codex, users may be asked to review the paused action before the agent resumes, while API jobs will simply cease. OpenAI has yet to publish Astra's system card, leaving unclear what happens when an API job is stopped and whether it can be resumed. This is particularly significant for Astra, as it's designed for extended periods of open-ended research and security tasks, potentially accumulating hours of work before OpenAI intervenes.
Developers will need to understand why the job was halted, as a timeout could typically be retried, but if OpenAI halted the job for safety reasons, restarting it might immediately encounter the same issue. The stricter limitations stem from Astra's impressive capabilities: it scored 100% on ExploitBench, an exploit benchmark, and found two previously unknown vulnerabilities while being tested against 20 high-severity vulnerabilities disclosed between June and August.
Astra also managed to exploit a browser, escape its sandbox, and execute commands on the host during security expert testing, and combined vulnerabilities in a hardened operating system to escalate privileges from unprivileged to root. OpenAI previously indicated in August that it could no longer rule out Astra reaching its Critical cybersecurity threshold.
Currently, access to Astra's advanced cybersecurity features will be limited to a select group of testers before broader release through Daybreak Blue. However, monitoring safety issues comes at a cost. OpenAI estimated that this monitoring adds about 20% to the inference compute of affected workloads, meaning some of the computational power behind Astra will be spent observing the model's actions rather than completing the tasks itself.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.