Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI’s new Astra model can code better. But here’s why its cybersecurity skills matter as much

OpenAI’s new Astra model can code better. But here’s why its cybersecurity skills matter as much

OpenAI's latest artificial intelligence (AI) model, GPT-6 Astra, has been unveiled, boasting improvements in computer usage, coding, long-running tasks, and cybersecurity capabilities. This marks Astra as the first OpenAI system to surpass what the company deems its "critical" cybersecurity capability threshold, enabling the model to identify previously unknown security flaws and devise ways to exploit them in well-protected systems without human intervention at each stage.

This newfound ability is particularly crucial in light of recent occurrences where AI agents have carried out unauthorized actions and exploited weaknesses in external software.

Astra was introduced on September 3, 2024, initially to a select group of organizations, with broader availability to follow for ChatGPT Plus, Pro, Business, and Enterprise users, as well as via its API and Amazon Web Services. Unlike traditional AI models, Astra is designed for end-to-end work, allowing it to interact with computers and browsers, fill online forms, update customer records, manage calendars, conduct web research, draft documents and emails, analyze scientific data, generate plots, build websites, install and test software, and troubleshoot problems visible on a screen.

Additionally, it can create documents, spreadsheets, and presentations while adhering to a specific template or visual style.

Astra outperforms its predecessor, GPT-5.6 Sol, in benchmark tests, particularly in computer use and software engineering. It scored 72.6% on OSWorld 2.0, an AI agent's ability to use a computer environment, compared to 65.7% for GPT-5.6 Sol, and completed tasks 47% faster. In terms of screen interaction, Astra achieved a score of 92.7% on ScreenSpot-Pro, against 76.9% for GPT-5.6 Sol.

On AutomationBench, Astra scored 41.4%, compared to 18.1% for GPT-5.6 Sol, while on Terminal-Bench 4.0, which tests terminal-based tasks, it scored 57.9% versus 37.3%.

The cybersecurity capabilities of Astra are noteworthy. It has been classified at the critical level in OpenAI's Preparedness Framework, a position not previously achieved by GPT-5.6 Sol. When tested without production safeguards, Astra achieved 100% on ExploitBench, a test that evaluates models' ability to transform known software vulnerabilities into functional exploits, compared to 78.5% for GPT-5.6 Sol.

Furthermore, in an internal benchmark using vulnerabilities disclosed between June and August 2026, Astra outperformed GPT-5.6 Sol. During testing, Astra discovered and utilized two previously unknown zero-day vulnerabilities, which OpenAI promptly disclosed to the affected maintainers. Expert-led tests also revealed that Astra could identify previously unknown vulnerabilities in a hardened browser and develop an exploit chain that achieved unsandboxed code execution, meaning it could execute the exploit without testing it in a separate environment.

This powerful function is available for defensive tasks such as secure code review, but more advanced requests, like generating proof-of-concept exploits, are restricted. These defensive capabilities are only accessible separately to authorized defenders through OpenAI's Daybreak program.

Written by urgent.news from The Indian Express's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at indianexpress.com →

More in AI

More from Friday 4 September →