Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Article URL: https://github.com/Niko1221/Strata Comments URL: https://news.ycombinator.com/item?id=49953495 Points: 210 # Comments: 99
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100 tokens per second
The 125-billion-parameter AI model, Qwen 3.8 Flash Next, can now run on consumer gaming PCs with NVIDIA or AMD graphics cards (12 GB or more) using Strata, a free and open-source inference engine. The model, usually requiring server resources, now operates on a normal PC, capable of tasks such as chatting, writing code, reading images, and working with applications.
Strata runs on Windows or Linux and requires a token to be approximately three-quarters of a word. Speed varies based on VRAM capacity, with an RTX 3090 (24 GB) achieving around 100-140 tokens per second. Multi-GPU setups are also possible, allowing for sharing the model between two or three cards. The entire installation process is automated, downloading the 70 GB model and setting up the appropriate engine for the user's graphics card.
Strata requires around 35-55 GB of RAM and locks part of it for the graphics card, which is normal during the initial load. After the model starts, a Strata app can be accessed at http://127.0.0.1:8080, and the PC can be used while the model runs. Updating Strata is also straightforward, with an UPDATE.bat or ./update.sh script available.
Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.