Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

Multi-tier storage rewrites the economics of AI inference

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine flash, object storage and disk-based capacity tiers, enabling enterprises to serve training and inference workflows while maximizing GPU productivity and economic savings. Super Micro Computer Inc. has…

Multi-tier storage rewrites the economics of AI inference

Multi-tier storage architectures are gaining popularity in AI infrastructure to control costs and improve performance, as inference has become the primary workload. Supermicro has partnered with various companies to tackle the challenge of efficiently serving AI agents' key-value, or KV, cache demands, where workflow data is stored.

Paul McLeod, Supermicro's product director of storage, highlighted that monolithic solutions are not effective as requirements change, and software-defined partners have been successful in addressing specific needs. Intel has addressed this issue with its QuickAssist Technology, a hardware accelerator built into select Intel processors, which moves compression and encryption from software to hardware, boosting CPU cycles, reducing latency, and improving power efficiency at scale.

Western Digital has expanded its hard disk drive, Ultrastar, portfolio to provide necessary capacity and performance for high-intensity AI workloads. Scality has integrated GPU-direct storage access into its platform, allowing AI training and inference pipelines to stream data directly from object storage into GPU memory. Samsung Semiconductor has developed scalable memory expansion for the KV cache storage tier to maintain GPU performance.

Samsung's PM1723 drive, a Gen 6 drive, offers 28.4 gigabytes per second of sequential read throughput and up to 6.6 million random read IOPS. Enterprises must carefully select optimal storage technology to balance performance and cost, addressing the tradeoff between performance and cost for AI inference workloads.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

More from Tuesday 11 August →