DeepSeek V4 Flash on a Single AMD MI300X
Article URL: https://github.com/ryanzhou/deepseek-v4-flash-mi300x Comments URL: https://news.ycombinator.com/item?id=49166386 Points: 272 # Comments: 60
The DeepSeek AI repository provides a configuration and patches to run the DeepSeek-V4-Flash-0731 model on a single AMD MI300X in production. The stack includes Docker Compose, SHA-256-pinned file overlays, reference diffs against upstream, JIT-compiled gfx942 kernel sources, and tuning tables. The checkpoint runs without weight quantization or offload.
The repository documents the tuning journey, with dated reports available. The MI300X has 192 GB of HBM3, 5.3 TB/s memory bandwidth, and 2.4× the HBM capacity of an H100 SXM5. The 304B-parameter checkpoint allows for a simple single-GPU deployment on the MI300X. The stack uses a digest-pinned official vLLM ROCm nightly with specific hardware and software configurations.
The repository includes fixes for FP8 format, MoE routing, checkpoint expert-activation clamps, causal speculative verification, CPU-KV synchronization, and prefill and decode kernel tuning. The correct FP8 format was identified as the first priority, and performance tuning followed. The model can reach up to 11.53K tokens per second using the current configuration.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.