Grok 4.7 was built to work for hours. It still fails most of the time.
A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through The post Grok 4.7 was built to work for hours. It still fails most of the time. appeared first on The New Stack .
SpaceXAI has released Grok 4.7, a coding agent designed to operate for extended periods. The model underwent a specially tailored reinforcement learning run, focusing on complex tasks that could span several hours. This training method aimed to improve Grok's self-verification abilities and its capacity to manage extensive context.
Following the release, Grok 4.7 demonstrated significant improvements in performance, excelling in tasks such as Terminal-Bench 4.0, CursorBench 4.0, and the multi-hour evaluation, AA Briefcase v1.1.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.