Urgent.News

What's breaking now, across thousands of outlets.

AI

From the creator of Redis; run LLM locally with ds4

DwarfStar 4 is an innovative narrow C inference engine designed to run large language models locally on high-memory Mac, CUDA, and ROCm machines. This powerful tool supports DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next models, offering a complete stack of text and vision models, local APIs, a command-line interface, and a native agent.

The key differentiator of DwarfStar 4 lies in its ability to run local inference engines, a departure from the traditional remote serving approach. To achieve this, the engine employs asymmetric quantization techniques to optimize the routing of experts while preserving crucial computational paths. This enables the models to become practical on high-memory machines, unlocking their full potential.

DwarfStar 4 exposes a user-friendly interface, including a command-line interface, HTTP APIs, and a native agent. These components share the same model state and cache, making it convenient for developers to integrate the engine into their workflows. The engine's performance is demonstrated by its ability to accommodate various models at different memory capacities. For instance, at 128 GB of memory, GLM 5.3 Q2 and Qwen Q4 models can be run locally, while V4.1 Q2 models can stream from the SSD.

The engine's capabilities are backed by extensive benchmark data, which can be found in the provided reference. To get started, the guide recommends checking the hardware matrix and connecting your preferred editor, agent, or API client to the local server. The quickstart guide provides a helpful introduction to the installation process.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dwarfstar.sh →

More in AI

More from Friday 2 October →