Google Unveils Gemini 4 Argon, Retaking Benchmark Lead Over OpenAI and Anthropic
Google has unveiled Gemini 4 Argon, a new frontier AI model that it says leads or ties rivals on 13 of 18 disclosed benchmarks. The model is initially being released only to trusted cyber defenders and select pre-release testers, with broader availability planned later. VentureBeat reports: Google is not claiming that Gemini 4 Argon wins every benchmark. But across the benchmark table disclosed…
Google has introduced Gemini 4 Argon, an advanced AI model that leads or matches rivals on 13 out of 18 benchmark tests, according to the company's embargoed materials. Initially, the model will only be accessible to trusted cybersecurity experts and select pre-release testers, with wider release planned later. VentureBeat highlights that Google does not assert Gemini 4 Argon as the best model across all benchmarks.
However, the model outperforms GPT-6 Astra and Claude Opus 5.5 in the majority of disclosed benchmarks. Specifically, Argon leads outright on 12 of the 18 benchmarks, and ties for the top score in one. In comparison, GPT-6 Astra leads on three benchmarks, and ties with Argon on one. Claude Opus 5.5 leads on two benchmarks. Therefore, Gemini 4 Argon emerges as the leading frontier model by total number of benchmark leads or top scores in Google's comparison set, despite a closely contested race where OpenAI and Anthropic still hold advantages in certain technical areas.
This development gives Google its strongest claim to overall frontier leadership by benchmark count. Additionally, Argon extends Google's output capacity, supporting an industry-leading 1 million output tokens, a significant upgrade from the previous 64,000-token limit. This enhancement is particularly beneficial for agentic software engineering, audit, migration, and legal-review tasks that rely on the model's ability to maintain a chain of work before requiring human intervention.
Written by urgent.news from Slashdot's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.