{
  "id": 2015812,
  "title": "The AI inference race moves beyond GPUs to reshape data center infrastructure",
  "url": "https://urgent.news/2026/08/19/the-ai-inference-race-moves-beyond-gpus-to-reshape-data-center",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-19T20:59:28.000Z",
  "source": {
    "name": "SiliconANGLE",
    "slug": "siliconangle",
    "url": "https://siliconangle.com/2026/08/19/ai-inference-infrastructure-requires-full-stack-coordination-supermicroopenstoragesummit/"
  },
  "original_language": "en",
  "account": "As generative and agentic applications move into production, the AI inference infrastructure is becoming a system-level challenge, according to IBM's Ka Wai Leung. While GPU performance remains crucial, storage latency, network bandwidth, data movement, and power consumption are increasingly determining the cost and speed of producing tokens. The requirements vary by workload, with interactive chat prioritizing latency, batch inference focusing on throughput, and agentic systems creating expanding contexts.\n\nLeung emphasized the importance of understanding the type of workload and building a system that caters to its characteristics. Storage plays a vital role in this infrastructure, with low-latency random reads and writes being crucial due to the continuous retrieval of proprietary or recently updated information. This requirement becomes even more complex when dealing with structured, unstructured, and multimodal information spread across various enterprise environments.\n\nPower efficiency is another significant consideration in AI inference infrastructure as data centers face energy and physical capacity constraints. Kioxia's BiCS8-based CM9 drives showed significant improvements in random-read and write input/output operations per second per unit of power compared to their previous CM7 generation.\n\nSupermicro's reference design integrates an Nvidia HGX B300 compute environment with IBM Storage Scale Erasure Code Edition and Kioxia drives. This architecture employs a high-performance storage tier to support latency-sensitive workloads and can accommodate a capacity tier for less frequently accessed data. By providing a unified solution from multiple vendors, Supermicro aims to simplify the process of building a total data center solution that includes servers, liquid cooling, and storage.\n\nTo demonstrate the benefits of IBM Storage Scale, IBM, Nvidia, and Supermicro tested it as a shared KV cache. The cache allowed previously computed context to be reused without consuming limited GPU memory or system RAM. The test achieved subsecond time-to-first-token responses across various prompt lengths. When heavy traffic was added to Supermicro's Spectrum-X network to test its throughput advantage over an uncached baseline, it was found that the number of requests per second decreased slightly, but the efficiency improved by 18 times compared to the uncached baseline.",
  "summary": "AI inference infrastructure is becoming a system-level challenge as organizations move generative and agentic applications into production. Graphics processing unit performance remains essential, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens. The requirements also vary by workload. Interactive chat…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}