Urgent.News

What's breaking now, across thousands of outlets.

AI

Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say

The industry is spending billions on more efficient chips, cooling, and grid connections. But what about making computers do less work – or doing it at a better time?

Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say

Hyperscalers could see a significant reduction in energy consumption by focusing on software improvements rather than solely investing in hardware upgrades, according to researchers. The International Energy Agency projects that by 2030, AI will demand 945TWh of electricity, equivalent to Japan's current usage. Currently, the industry has primarily concentrated on enhancing hardware capabilities, such as creating more efficient GPUs and securing a stable power supply from a constrained grid.

A potential alternative to this hardware-centric approach is optimizing software and algorithms to reduce the workload on servers. Currently, the average power usage effectiveness (PUE) has remained unchanged for six consecutive years, indicating that cooling and power systems are not as efficient as desired. Since servers account for about 60% of a data center's electricity consumption, while cooling ranges from 7% in efficient hyperscale sites to over 30% in less efficient facilities, optimizing software could be a logical next step.

Jae-Won Chung, a University of Michigan PhD candidate in computer science and engineering, suggests that computers should be viewed as a stack, with hardware at the bottom, followed by systems software, algorithms, and applications. He claims that gains made in efficiency at these higher levels can compound and significantly impact overall energy use.

ML.Energy’s tests on the Alibaba Qwen 3 235B A22B Thinking model revealed that running inference in FP8 format consumed 33% less energy compared to bfloat16 versions on problem-solving tasks. Chung's Perseus training optimizer identifies less critical parts of a large-model training job and slows them down, allowing busier parts to finish simultaneously. This approach reduced training energy consumption by up to 30% without affecting throughput or hardware.

Major GPU manufacturers, such as Nvidia, are already employing similar strategies. Nvidia's Blackwell power profiles fine-tune various aspects of GPU operation, like compute and memory frequencies, power limits, NVLink states, and cache settings, to optimize energy usage. This approach can potentially save up to 15% of energy while maintaining 97% or more of performance, enabling power-constrained facilities to run more GPUs and boost throughput by up to 13%.

Written by urgent.news from Tom's Hardware's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at tomshardware.com →

More in AI

More from Thursday 8 October →