Grok 4.6 narrows frontier AI performance gap
SpaceXAI has released Grok 4.6, pushing its flagship artificial intelligence model into the top tier of independent performance rankings and intensifying competition with OpenAI and Anthropic on advanced reasoning, coding and autonomous agent tasks. The model, released on 12 August, scored 61 on the Artificial Analysis Intelligence Index, matching OpenAI’s GPT-5.6 Sol at its highest tested…
SpaceXAI has launched Grok 4.6, positioning its primary artificial intelligence model among industry leaders in terms of performance. This latest release, announced on 12 August, has matched OpenAI's GPT-5.6 Sol in the Artificial Analysis Intelligence Index, with both scoring 61. Grok 4.6 demonstrated a five-point improvement over its predecessor, Grok 4.5, which had scored 56. However, Anthropic's Claude Opus 5 and Claude Fable 5 have maintained a slight edge, with scores of 63 and 62, respectively.
The development of Grok 4.6 marks a significant step towards creating advanced AI models capable of handling long, multi-step tasks instead of merely responding to individual prompts. The model has been specifically optimized for long-running agents, software development, knowledge work, and interactive or visual projects. Its strongest results were observed in tasks that necessitate a series of model operations.
Grok 4.6 achieved an Elo score of 1,753 in the GDPVal-AA v2 knowledge-work evaluation, slightly outperforming GPT-5.6 Sol's Elo score of 1,728. On CursorBench 3.2, which gauges coding performance, Grok 4.6 scored 69.9%, compared to GPT-5.6 Sol's 67.2% and Claude Fable 5's 70%. Performance varied across different tests; however, Grok 4.6 still showcased competitive performance on agentic workloads.
It achieved a 50.7% score on a multi-turn banking benchmark involving customer-service tasks and tool use, and an 88.4% Terminal-Bench v2.1 score, placing it alongside leading models for software tasks executed via a terminal environment.
Price could prove to be a significant competitive advantage for Grok 4.6. The model starts at $2 per million input tokens and $6 per million output tokens, identical to Grok 4.5's pricing. A faster version is available at twice these rates. Despite its comparable benchmark scores, Grok 4.6 remains considerably cheaper on headline API pricing compared to several rival models offering similar performance.
Measured costs reaffirm this advantage, with Grok 4.6 requiring approximately $0.84 per task in Artificial Analysis testing, placing it among the more efficient high-performing models.
SpaceXAI further enhanced Grok 4.6 through extended supplemental training compared to its predecessor. This process incorporated curated model-generated reasoning data, engineering material, and alterations to the training optimization process. Following supervised fine-tuning, reinforcement learning was applied to knowledge work, general coding, kernel optimization, web development, and computer-aided design.
The company has also placed increased emphasis on models that can create complete applications and work products. During internal testing, Grok 4.6 engaged in researching unfamiliar areas, structuring applications, implementing core functions, and refining results through multiple rounds of feedback. This direction reflects a broader trend in frontier AI development, as seen in OpenAI's recent strengthening of GPT-5.6 Sol for complex research, coding, and professional workflows, alongside improvements in factual reliability and consistency in ChatGPT.
Written by urgent.news from Arabian Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.