Urgent.News

What's breaking now, across thousands of outlets.

AI

AI inference gets a new tier as context windows grow

AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference. Agentic AI is also changing the shape of the data problem. Interactions are growing longer […] The post AI inference gets a new…

AI inference gets a new tier as context windows grow

AI inference is experiencing a new tier as context windows grow larger, according to experts at Solidigm, Vast Data, and Supermicro. As agentic AI systems analyze, act and reassess, they generate longer contexts and more data that need to be accessed quickly during inference. This creates new storage and memory demands for AI workloads, beyond what graphics processing unit memory can handle.

The companies discussed how storage infrastructure can provide additional tiers to balance proximity, capacity, and speed for the data load. Solidigm supplies fast SSDs for cached data, Supermicro integrates them into rack-scale systems, and Vast Data's AI Operating System provides network storage and data services. The middle tier, enabled by Non-Volatile Memory Express SSDs, allows context data to live in places previously considered unsuitable.

KV cache, a significant optimization for AI inference, can replace compute with storage, reducing GPU expense and latency. Supermicro's Context Memory eXtension (CMX) targets large AI clusters with substantial data demands, while other deployments may require different combinations of memory, local SSDs, and network storage. Vast Data has tested KV cache offload, achieving 20 times faster time-to-first-token and 90% savings in GPU time.

Solidigm's SSDs address different points in the hierarchy, and Supermicro integrates them into customizable systems. The storage infrastructure's ability to match storage performance and capacity to the workload is crucial for AI applications.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

Your Agent Pipeline's Review Gate Should Be Code, Not a Prompt Convention

Your Agent Pipeline's Review Gate Should Be Code, Not a Prompt Convention ZOdyssey is a code-enforced orchestration pipeline for ZCode (for now — the pattern is harness-portable), and its center of…

  • ZOdyssey is a code-enforced orchestration pipeline for ZCode
  • Review node is the only barrier to execution in the pipeline
  • Review phase enforces a nonce-bound gate, preventing unauthorized code execution

More from Tuesday 25 August →