{
  "id": 7937671,
  "title": "DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression",
  "url": "https://urgent.news/2026/09/17/deepseek-v4-1-flash-pushing-the-limits-of-kv-cache-compression",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-17T01:39:47.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://zartbot.github.io/blog/model_arch/dsv41flash_arch/en.html"
  },
  "original_language": "en",
  "account": "DeepSeek-V4.1 Flash has been released, showcasing remarkable speed and compression capabilities for Long-horizon Agent Workflows. This new iteration aims to push the limits of KVCache compression, addressing the growing need for persistent storage, reuse, and transfer of large KVCache in handling ultra-long-context processing. The model's architecture incorporates a combination of Sparse Attention and Sliding Window Attention (SWA) to process long sequences efficiently. DeepSeek-V4.1 Flash supports multimodal input and contexts of up to 1 million tokens, with a parameter scale of 552B, and maintains high-quality task completion while achieving a 4x compression of KVCache. This compression is achieved through joint optimizations in model architecture, cache precision, and deployment strategy, including the use of FP4 quantization and enhanced global compressed attention (CSA/HCA). The model also features a Causal Encoder-Decoder (CED) architecture, which reduces long-context Prefill computation and leverages Engram conditional memory with DSpark speculative decoding.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}