DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1 Flash has been released, showcasing remarkable speed and compression capabilities for Long-horizon Agent Workflows. This new iteration aims to push the limits of KVCache compression, addressing the growing need for persistent storage, reuse, and transfer of large KVCache in handling ultra-long-context processing. The model's architecture incorporates a combination of Sparse Attention and Sliding Window Attention (SWA) to process long sequences efficiently.
DeepSeek-V4.1 Flash supports multimodal input and contexts of up to 1 million tokens, with a parameter scale of 552B, and maintains high-quality task completion while achieving a 4x compression of KVCache. This compression is achieved through joint optimizations in model architecture, cache precision, and deployment strategy, including the use of FP4 quantization and enhanced global compressed attention (CSA/HCA).
The model also features a Causal Encoder-Decoder (CED) architecture, which reduces long-context Prefill computation and leverages Engram conditional memory with DSpark speculative decoding.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.