Grok 4.6
Article URL: https://x.ai/news/grok-4-6 Comments URL: https://news.ycombinator.com/item?id=49274027 Points: 304 # Comments: 315
Grok 4.6 is the latest version of the Grok model, building on the improvements of its predecessor, Grok 4.5. This update places a strong emphasis on long-running agents and more complex interactive and visual tasks. Grok 4.6 maintains its capability to handle intricate tasks spanning multiple steps, whether that involves researching, analyzing information, working with codebases, or transforming ideas into polished applications or work artifacts.
The model excels at complex tasks across several agentic coding and knowledge work benchmarks, matching the performance of GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite score of nine benchmarks. Grok 4.6 is accessible in Cursor and Grok Build, with an offer of double the included usage for the first week to allow users to immediately test out the new version.
The model underwent an extended supplemental training phase compared to Grok 4.5, incorporating curated model-generated data to enhance reasoning abilities and understanding of advanced technical concepts. High-quality engineering data was also integrated. Following the training, Grok 4.5 was utilized to regenerate SFT (Supervised Fine-Tuning) trajectories across various domains, including STEM, software engineering, and knowledge work.
This process involved filtering out any problematic traces using model-based checks, resulting in a stronger foundation for subsequent stages.
Grok 4.6 has been trained on a diverse range of agentic RL (Reinforcement Learning) tasks, covering areas such as general coding, knowledge work, and domain-specific environments like kernel optimization, web development, and computer-aided design. During testing, the model demonstrated exceptional performance, particularly in transforming broad product ideas into working first versions.
It can research unfamiliar domains, structure applications, implement core interactions, and refine results through multiple iterations of feedback.
In addition to its strengths in long-running tasks, Grok 4.6 shows notable improvements in visual and interactive projects. Given a concrete product idea, it can establish structure and visual language for an application in a single pass, making it especially useful for projects where the fastest route to a satisfactory result involves starting with a substantial foundation and then iterating through the loop.
The model also features enhanced and calibrated safeguards aligned with its expanded capabilities. The safety stack prioritizes utility and security across legitimate use cases, allowing Grok 4.6 to be both helpful and safe in various domains, including vulnerability patching, accelerating the engineering design cycle, and augmenting AI research.
Comprehensive pre-deployment testing, including the widest-ever suite of capabilities and safeguard evaluations, as well as extensive post-deployment and third-party testing, has been conducted to ensure the model's robustness and reliability. Grok 4.6 is now available in Cursor, Grok Build, the API, and other partners like OpenRouter, Vercel, and Cloudflare.
Pricing starts at $2 per million input tokens and $6 per million output tokens, with a faster variant available at twice the cost.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.