Urgent.News

What's breaking now, across thousands of outlets.

AI

The AI Race Just Got Awkward

The AI race has taken an unexpected turn, with Chinese labs seemingly offering a helping hand to their Western counterparts. While Western labs often claim that Chinese labs are distilling their models and posing a danger to humanity, the tables have turned. The Western AI companies, such as Anthropic and OpenAI, are now quietly adopting the Chinese labs' advances instead of accusing them of theft.

One of the most significant breakthroughs comes from DeepSeek, a Chinese lab that has generously shared their KV cache optimization techniques with the world. This optimization has resulted in a staggering reduction of the KV cache footprint for certain use cases, like coding, by approximately 437 times compared to DeepSeek-V1. DeepSeek was the first to release the MLA architecture, which compressed the cache by roughly 15 times, followed by 'Compressed Sparse Attention' and 'Heavily Compressed Attention.'

The latest DeepSeek-V4.1-Flash takes it even further, utilizing cross-layer cache reuse, a causal encoder-decoder architecture, and FP4 caching, bringing the global KV cache down to an astonishing 890 bytes per token.

The significance of these optimizations lies in their impact on serving long-context models. One of the largest costs in serving these models is the VRAM required to hold the cache in GPU memory. The graph provided in the source material clearly demonstrates the extent of the improvements. For example, DeepSeek-V4.1-Flash costs only 890 bytes per token, a dramatic decrease from what the same tier would have cost roughly two months ago.

These developments have caught the Western AI companies off guard. Both Anthropic and OpenAI have released models that utilize these optimizations, and their user reviews have been overwhelmingly positive. The quality of the models does not seem to be far off from their flagship models, with cache read costs dropping significantly. Anthropic's Claude Opus 5.5 reduced cache-read pricing by 60% compared to Opus 5, while OpenAI's GPT-6.1 Sol reduced it by 80% compared to GPT-5.6 Sol's late-July pricing.

The reasons behind the Chinese labs' decision to freely share these breakthroughs remain unclear. However, it is evident that the constraints on access to advanced GPUs have pushed Chinese labs to prioritize performance optimization. As a result, they have thrown a lifeline to their Western loss-making counterparts, and the AI race has taken an awkward turn.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at insufferable.dev →

More in AI

More from Wednesday 30 September →