{
  "id": 10982147,
  "title": "The AI Race Just Got Awkward",
  "url": "https://urgent.news/2026/09/30/the-ai-race-just-got-awkward",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T15:50:11.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://insufferable.dev/posts/the-ai-race-just-got-awkward/"
  },
  "original_language": "en",
  "account": "The AI race has taken an unexpected turn, with Chinese labs seemingly offering a helping hand to their Western counterparts. While Western labs often claim that Chinese labs are distilling their models and posing a danger to humanity, the tables have turned. The Western AI companies, such as Anthropic and OpenAI, are now quietly adopting the Chinese labs' advances instead of accusing them of theft.\n\nOne of the most significant breakthroughs comes from DeepSeek, a Chinese lab that has generously shared their KV cache optimization techniques with the world. This optimization has resulted in a staggering reduction of the KV cache footprint for certain use cases, like coding, by approximately 437 times compared to DeepSeek-V1. DeepSeek was the first to release the MLA architecture, which compressed the cache by roughly 15 times, followed by 'Compressed Sparse Attention' and 'Heavily Compressed Attention.' The latest DeepSeek-V4.1-Flash takes it even further, utilizing cross-layer cache reuse, a causal encoder-decoder architecture, and FP4 caching, bringing the global KV cache down to an astonishing 890 bytes per token.\n\nThe significance of these optimizations lies in their impact on serving long-context models. One of the largest costs in serving these models is the VRAM required to hold the cache in GPU memory. The graph provided in the source material clearly demonstrates the extent of the improvements. For example, DeepSeek-V4.1-Flash costs only 890 bytes per token, a dramatic decrease from what the same tier would have cost roughly two months ago.\n\nThese developments have caught the Western AI companies off guard. Both Anthropic and OpenAI have released models that utilize these optimizations, and their user reviews have been overwhelmingly positive. The quality of the models does not seem to be far off from their flagship models, with cache read costs dropping significantly. Anthropic's Claude Opus 5.5 reduced cache-read pricing by 60% compared to Opus 5, while OpenAI's GPT-6.1 Sol reduced it by 80% compared to GPT-5.6 Sol's late-July pricing.\n\nThe reasons behind the Chinese labs' decision to freely share these breakthroughs remain unclear. However, it is evident that the constraints on access to advanced GPUs have pushed Chinese labs to prioritize performance optimization. As a result, they have thrown a lifeline to their Western loss-making counterparts, and the AI race has taken an awkward turn.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Hacker News Best",
        "title": "The AI Race Just Got Awkward",
        "url": "https://urgent.news/2026/09/30/the-ai-race-just-got-awkward-10994423",
        "published": "2026-09-30T15:50:11.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}