{
  "id": 3051815,
  "title": "Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents",
  "url": "https://urgent.news/2026/08/24/up-to-30x-more-work-per-watt-nvidia-vera-rubin-nvl72-sets-a-new",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-24T15:00:19.000Z",
  "source": {
    "name": "NVIDIA Blog",
    "slug": "nvidia-blog",
    "url": "https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/"
  },
  "original_language": "en",
  "account": "NVIDIA Vera Rubin NVL72 systems are delivering up to 30 times more work per watt compared to NVIDIA GB300 NVL72 on agentic workloads, according to new measured performance data. This significant increase in throughput per megawatt is a result of NVIDIA's extreme codesign across every layer of the platform, including disjointed serving, rate matching, large-scale expert parallelism, distributed KV-caching, and fused CUDA kernels. Agentic AI workloads, which involve complex research and reasoning across multiple steps, consume much more computational power than simple chat requests, with context growing exponentially throughout the session. NVIDIA's Vera Rubin NVL72 platform is designed to handle this increased demand efficiently, offering up to 35 times lower token cost per million tokens than its predecessor, GB300 NVL72. This allows AI factories to run agentic workloads continuously at scale, maximizing revenue potential while minimizing energy consumption.",
  "summary": "According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "NVIDIA Blog",
        "title": "With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents",
        "url": "https://urgent.news/2026/08/24/with-groq-3-lpx-in-full-production-nvidia-extends-vera-rubin",
        "published": "2026-08-24T15:00:41.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}