Urgent.News

What's breaking now, across thousands of outlets.

More in AI

The sm_120 shared-memory cliff: why FP8 KV cache crashes vLLM on workstation Blackwell

If you run vLLM with --kv-cache-dtype fp8 on a DeepSeek-family (MLA) model and your GPU is a GB10, an RTX PRO 6000, or any workstation or consumer Blackwell card, there is a decent chance the engine…

  • vLLM crashes on Blackwell cards due to fp8 KV cache exceeding shared memory limit
  • Recommended solution: compute tile's shared-memory requirement vs device's sharedmemoryperblockoptin

More from Tuesday 25 August →