Urgent.News

What's breaking now, across thousands of outlets.

AI

Nvidia ties AI factory economics to tokens and power efficiency

AI factory economics increasingly depend on more than access to high-performance graphics processing units. As agentic systems draw on multiple models, databases and tools, the entire data center must work as one computing system. That transition is shifting attention from individual chips to the infrastructure that turns computing capacity into useful intelligence. Networking, storage,…

Nvidia ties AI factory economics to tokens and power efficiency

Nvidia has shifted focus from individual graphics processing units (GPUs) to data center infrastructure that maximizes computing capacity while minimizing power consumption, according to Ian Buck, vice president and general manager of hyperscale and high-performance computing (HPC) at Nvidia. Buck explained that as agentic systems leverage multiple models, databases, and tools, the entire data center must function as a single computing unit.

Instead of traditional metrics like cars or devices, the key performance indicators are tokens, which are revenue-generating, durable, and productive assets. An AI factory's commercial output stems from inference, where deployed models process requests and generate tokens. However, inference does not replace training, as organizations continuously update deployed models based on evolving data and market conditions.

This continuous refinement constitutes a form of training, with reinforcement learning and online alignment being areas of active work. Low-latency workloads, which require faster reasoning, represent another economic tier, with Nvidia's Groq 3 LPX inference accelerator enhancing token generation rates for time-sensitive tasks when paired with its Vera Rubin platform.

Power constraints at the data center level dictate the scale of computing infrastructure, making tokens per watt a crucial metric for AI factory economics. Nvidia has consistently improved performance with each GPU generation, achieving significant efficiency gains, such as a 30x improvement in tokens per watt with the Blackwell architecture.

This shift in focus also alters how systems are designed and operated, with solutions like CoreWeave offering customers pre-configured or higher-level inference services to navigate the myriad options available.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

I ran six coding agents on seven local models, 30 times each

Last week I posted a small benchmark on whether coding agents still work when the model you run yourself is shaky at tool calls.

  • Six coding agents tested on seven local models
  • Tests conducted 30 times each on RTX 5080 GPU
  • Polyglot agent showed consistent performance

More from Thursday 1 October →