{
  "id": 8990156,
  "title": "Grok 4.7 was built to work for hours. It still fails most of the time.",
  "url": "https://urgent.news/2026/09/21/grok-4-7-was-built-to-work-for-hours-it-still-fails-most-of-the-time",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-21T19:13:23.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/grok-4-7-agent-stamina/"
  },
  "original_language": "en",
  "account": "SpaceXAI has released Grok 4.7, a coding agent designed to operate for extended periods. The model underwent a specially tailored reinforcement learning run, focusing on complex tasks that could span several hours. This training method aimed to improve Grok's self-verification abilities and its capacity to manage extensive context. Following the release, Grok 4.7 demonstrated significant improvements in performance, excelling in tasks such as Terminal-Bench 4.0, CursorBench 4.0, and the multi-hour evaluation, AA Briefcase v1.1.",
  "summary": "A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through The post Grok 4.7 was built to work for hours. It still fails most of the time. appeared first on The New Stack .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}