Urgent.News

What's breaking now, across thousands of outlets.

AI

How Everpure plans to stop AI from starving without data

SPONSORED FEATURE: The vendor's AI solutions are dedicated to increasing GPU utilization and avoiding costly GPUs doing nothing while waiting for data

How Everpure plans to stop AI from starving without data

Everpure, a company focused on AI infrastructure, aims to address the issue of AI models starving for data. In a typical scenario, an AI agent, operating within a high-performance Nvidia SuperPOD system, receives a query from a customer regarding insurance coverage. To answer this query, the agent needs to access and retrieve relevant policy data from a vast storage system.

This data is stored in petabytes or exabytes, making it crucial for the storage system to quickly deliver the information to the GPU servers, which are expensive and prone to idle time, costing the company around $25 per minute of inactivity.

Par Botes, Everpure's VP of AI Infrastructure, emphasizes the need for a modern data processing approach tailored to AI. Traditional data processing methods, while still relevant, have evolved to accommodate the unique access patterns of AI systems. Metadata, containing information about the structure, state, and semantics of data, has become increasingly important, as it can outperform the data itself in terms of enabling efficient queries.

To overcome the challenges posed by scale, Everpure has developed an AI Data Platform based on the NVIDIA AI Data Platform reference design. This platform focuses on three key bottlenecks:

1. Throughput starvation: To keep thousands of GPUs fully utilized, the storage hardware must be capable of handling high bandwidths, ranging from 10 TB/sec to even more. Everpure offers solutions such as FlashBlade//S for high-end enterprises and FlashBlade//EXA for extremely large-scale operations. These platforms provide exceptional performance, with capabilities up to 450 million IOPS, 220 GB/sec bandwidth, and 4.6 billion metadata operations per second.

2. KV cache prefill tax: When an AI model accesses a document, the system must preprocess and store it in a key-value cache. If a second user requests access to the same document, the system must repeatedly perform the prefill process, wasting computational resources. Everpure's Key Value Accelerator (KVA) addresses this issue by directly transferring cached token states to shared flash memory via NVIDIA GPUDirect Storage (GDS) using Remote Direct Memory Access (RDMA). This eliminates the need for redundant computations and reduces GPU overhead.

3. Eliminating irrelevant data silos: As enterprises accumulate data across various systems, storing and managing it in a single, centralized location becomes impractical. Everpure's Data Stream platform enables data re-use across silos, ensuring that data remains valid and up-to-date without the need for expensive data copying. This approach streamlines the process of providing AI agents with the necessary information while minimizing the time and resources spent on data consolidation.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

More from Tuesday 15 September →