SpaceXAI trained Grok 4.6 on something most AI labs throw away
SpaceXAI released Grok 4.6 on Wednesday, less than a month after Grok 4.5. The company says Grok 4.6 can research The post SpaceXAI trained Grok 4.6 on something most AI labs throw away appeared first on The New Stack .
SpaceXAI unveiled Grok 4.6, an AI model designed for research and app development, on Wednesday. The updated model can research new topics, navigate large codebases, and transform product ideas into functional applications. SpaceXAI observed Grok 4.6 taking extra precautions during extended tasks, checking its work more frequently before proceeding. The model's training places a stronger emphasis on identifying and rectifying mistakes while maintaining focus on the original task.
SpaceXAI expanded Grok 4.6's training beyond its predecessor by combining model-generated reasoning, technical material, and engineering data with additional reinforcement learning focused on general coding, kernel optimization, web development, and computer-aided design. The model was rewarded for successfully completing larger tasks rather than merely generating plausible code. As a result, Grok 4.6 is more likely to pause during lengthy tasks to verify the accuracy of its work before continuing.
Benchmarks reveal mixed results for Grok 4.6, with improvements over Grok 4.5 evident in several areas but still trailing behind other models like OpenAI's GPT-5.6 Sol Max and Anthropic's Fable 5 Max in specific categories. Grok 4.6's performance on longer tasks outperformed competitors, and it achieved notable gains in APEX-Agents, Terminal-Bench, and Artificial Analysis Intelligence Index scores.
Despite retaining the same API pricing as Grok 4.5, costs may vary depending on the specific task and the frequency of tool usage, file re-reading, or task restarts. Grok 4.6 is available through SpaceXAI's API, OpenRouter, Vercel, Cloudflare, Grok Build, and Cursor. The release of Grok 4.6 underscores SpaceXAI's shift from a chatbot model towards building infrastructure for long-running agents, reflecting a broader trend in the coding-model race.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.