Urgent.News

What's breaking now, across thousands of outlets.

AI

GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence

GPT-6.1 Sol has taken over the role of GPT-6 Sol just seven days after its release, boasting near-Astra intelligence levels. Despite scoring 1 point lower than GPT-6 Astra in the Intelligence Index, it achieves this at less than a quarter of the cost per task. The pricing mirrors that of GPT-6 Sol at $2 per 10 million input/output tokens, with a cache read discount increasing from 90% to 95%.

This new model outperforms GPT-6 Sol in several key areas, including agentic knowledge work. It scores 4 points higher than GPT-6 Sol on AA-Briefcase v1.1 and GDPval-AA v2.1, while also achieving a notable 12-point improvement in Terminal-Bench 4.0, a 5-point jump in Humanity's Last Exam, a 6-point increase in GDP.pdf, and an 8-point gain in AA-Omniscience Accuracy, with hallucination rates dropping from 60% to 54%.

In terms of cost efficiency, GPT-6.1 Sol is significantly cheaper than its counterparts. At maximum effort, it costs less than a quarter of GPT-6 Astra ($0.72 vs $3.26) per task, 31% less than GPT-6 Sol ($1.05), and 64% less than GPT-5.6 Sol ($1.99). It pushes the cost efficiency Pareto frontier, meaning no cheaper model offers the same level of intelligence. However, it uses slightly more output tokens than GPT-6 Sol (~10-30%), although its low and medium effort levels are still Pareto optimal for token efficiency.

The model also shows improvements in the Coding Agent Index, gaining 3 points over GPT-6 Sol at maximum effort and sitting 2 points below GPT-6 Astra. It dominates the lower-price range of the Pareto frontier for this index, with the xhigh effort setting outperforming the max effort setting by 3 points. Despite using 10-30% more output tokens than GPT-6 Sol, its performance at lower effort levels makes it an efficient choice for coding tasks.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at artificialanalysis.ai →

More in AI

More from Wednesday 30 September →