Vast uses tiered storage to ease AI agent memory demands
AI agent memory is creating new demands on infrastructure as agents run longer sessions and spread across the enterprise. Retaining that context and making it available when needed puts pressure on memory capacity and data movement. Those demands extend beyond the context held during an individual interaction. Enterprise agents also need shared knowledge that persists […] The post Vast uses…
AI agents are requiring more memory storage as they function over extended periods and expand throughout the enterprise. Keeping this context accessible when needed places strain on memory capacity and data transfer. Enterprise agents also need shared knowledge that stays consistent across sessions, according to Alon Horev, co-founder and chief technology officer of Vast.
Horev explained that there are various types of memory for agents, including long-term memory for agents to understand past conversations and interactions. At Fully Connected 2026, Horev spoke with John Furrier and Dave Vellante about AI agent memory, key value cache offloading, and the shift towards data movement as the next bottleneck.
Inference pressure manifests with long-running sessions holding their key value cache in graphics processing unit memory. If this memory can be stretched and sessions offloaded to storage, it can prevent repetitive calculations. Vast's approach utilizes memory in tiers, starting with GPU memory, then central processing unit memory, and finally persistent media capable of holding petabytes of KV cache, with Nvidia's Dynamo software managing the process.
This allows for moving sessions between busy and less busy GPUs, moving KV cache over the network or reading it from Vast. As companies deploy thousands of agents handling sensitive data, AI agent memory becomes a governance and performance asset. Vast has also introduced a confidential computing service for sensitive workloads.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.