Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine
Running large language model inference at scale forces a KV cache trade-off: oversized GPU instances or slow time-to-first-token. This post builds a tiered KV cache on Amazon SageMaker HyperPod that extends the cache into a shared, distributed NVMe pool with Curvine, so replicas reuse cache at near-local-disk speeds on cost-efficient instances.
Large language model (LLM) inference at scale often necessitates trade-offs between KV cache size and performance. Teams deploying foundation models across various endpoints face higher infrastructure costs and degraded user experiences due to this trade-off. During generation, vLLM stores processed token attention keys and values in a KV cache, preventing recomputation on subsequent steps.
Prefix caching reuses this cache across prompts sharing the same initial tokens. However, prefix cache capacity is limited on cost-efficient instances. The solution proposed is a tiered KV cache architecture on Amazon SageMaker HyperPod, which extends the cache hierarchy beyond GPU and CPU memory into a shared, distributed NVMe pool.
This architecture utilizes HyperPod Managed Tiered KV Cache, Intelligent Routing, and Curvine as the shared L2 tier. This approach enables reuse of KV cache across replicas at near-local-disk speeds, achieving up to 100% cross-Pod cache hit rate and reducing time-to-first-token by up to 2.7 times. This architecture also allows workloads that previously required more expensive instances to run on lower-cost G6e instances, reducing per-endpoint costs depending on model size and traffic profile.
Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.