Run DeepSeek V4 on Your Own Hardware With DwarfStar, the Redis Creator's New Inference Engine
The creator of Redis thinks your old GPU is not obsolete. Salvatore Sanfilippo, better known as antirez, published a project called DwarfStar (the repo is antirez/ds4 ) that runs DeepSeek V4 Flash, GLM 5.x, and DeepSeek V4 PRO entirely on consumer hardware: Macs, DGX Spark boxes, Strix Halo desktops, and older NVIDIA cards like Ada Lovelace and L40S. It hit the front page of Hacker News this week…
DeepSeek V4, a powerful language model, can now be run on consumer hardware using DwarfStar, an inference engine created by the inventor of Redis, antirez. The project is based on quantized versions of DeepSeek V4 Flash, GLM 5.x, and DeepSeek V4 PRO models, and it is designed to run efficiently on a variety of hardware, including Macs, DGX Spark boxes, Strix Halo desktops, and older NVIDIA cards.
Unlike generic inference engines, DwarfStar is a narrow, self-contained native C program with no dependency on GGML. It supports a limited set of models but ships its own quantized weights and tests the entire stack, from tensor loading to HTTP server, to ensure optimal performance.
DwarfStar offers support for various backends, including Metal on Apple Silicon, CUDA (including multi-GPU setups), and ROCm on Strix Halo systems. The project is currently in beta quality, with models being removed and replaced as better alternatives become available.
The model selection is intentional and aims to deliver excellent performance on specific models, rather than supporting a broad range of models. DeepSeek V4 Flash and GLM 5.x models utilize mixture-of-experts (MoE) architecture, which allows for aggressive quantization without sacrificing performance. This enables the models to fit within the memory constraints of consumer-grade hardware.
The hardware requirements vary depending on the specific setup, but a 96 to 128 GB Apple Silicon Mac is recommended as the baseline. For older NVIDIA cards like Ada Lovelace and L40S, the project supports running the models using CUDA. In a cluster of two 128 GB Macs, it is possible to run 4-bit DeepSeek Flash or GLM 5.3 Flash using tensor parallelism over RDMA.
Full-resident DeepSeek V4 PRO Q4 can be run on a 512 GB workstation, such as the M3 Ultra. The project also mentions streaming capabilities, where models can be preloaded from a fast local SSD, enabling efficient processing of large context windows.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.