Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

The AI inference race moves beyond GPUs to reshape data center infrastructure

AI inference infrastructure is becoming a system-level challenge as organizations move generative and agentic applications into production. Graphics processing unit performance remains essential, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens. The requirements also vary by workload. Interactive chat…

The AI inference race moves beyond GPUs to reshape data center infrastructure

As generative and agentic applications move into production, the AI inference infrastructure is becoming a system-level challenge, according to IBM's Ka Wai Leung. While GPU performance remains crucial, storage latency, network bandwidth, data movement, and power consumption are increasingly determining the cost and speed of producing tokens. The requirements vary by workload, with interactive chat prioritizing latency, batch inference focusing on throughput, and agentic systems creating expanding contexts.

Leung emphasized the importance of understanding the type of workload and building a system that caters to its characteristics. Storage plays a vital role in this infrastructure, with low-latency random reads and writes being crucial due to the continuous retrieval of proprietary or recently updated information. This requirement becomes even more complex when dealing with structured, unstructured, and multimodal information spread across various enterprise environments.

Power efficiency is another significant consideration in AI inference infrastructure as data centers face energy and physical capacity constraints. Kioxia's BiCS8-based CM9 drives showed significant improvements in random-read and write input/output operations per second per unit of power compared to their previous CM7 generation.

Supermicro's reference design integrates an Nvidia HGX B300 compute environment with IBM Storage Scale Erasure Code Edition and Kioxia drives. This architecture employs a high-performance storage tier to support latency-sensitive workloads and can accommodate a capacity tier for less frequently accessed data. By providing a unified solution from multiple vendors, Supermicro aims to simplify the process of building a total data center solution that includes servers, liquid cooling, and storage.

To demonstrate the benefits of IBM Storage Scale, IBM, Nvidia, and Supermicro tested it as a shared KV cache. The cache allowed previously computed context to be reused without consuming limited GPU memory or system RAM. The test achieved subsecond time-to-first-token responses across various prompt lengths. When heavy traffic was added to Supermicro's Spectrum-X network to test its throughput advantage over an uncached baseline, it was found that the number of requests per second decreased slightly, but the efficiency improved by 18 times compared to the uncached baseline.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

The compliance frameworks were written before AI coding tools existed

ISO 27001, SOC 2, and NIST assume humans made decisions and left a paper trail. That assumption is breaking. Imagine an auditor sitting across from your engineering team.

  • Compliance frameworks predate AI coding tools
  • Three main gaps identified in current frameworks
  • Organizations face short-term compliance challenges

OpenAI Adds Zero Data Retention and Private Safety Processing for Enterprise AI

OpenAI has announced Zero Data Retention (ZDR) for frontier-model deployments and a new Private Safety Processing layer for enterprise customers.

  • OpenAI introduces Zero Data Retention (ZDR) for enterprise AI deployments.
  • ZDR ensures customer prompts and model responses are not retained after a single request.
  • Private Safety Processing adds a safety layer for monitoring related interactions.

More from Wednesday 19 August →