Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
The Strata open source AI inference engine enables running the 125-billion-parameter Qwen3.8-Flash-Next model on consumer-grade gaming PCs with RTX 4090 or similar NVIDIA/AMD graphics cards. Users can install the free software on Windows or Linux, which handles setting up the required environment automatically. The model can perform tasks like chatting, code generation, image analysis, and integration with user apps, all while keeping data within the local machine.
A 125-billion parameter model typically requires substantial VRAM, with an RTX 3090 handling about 100-140 tokens per second on a single card. Strata optimizes this by distributing the workload across the entire PC. The installation process detects the graphics card and installs the optimal engine configuration. Users simply run the provided batch file to start the model, which downloads a 70GB model and begins execution.
The first run may require 1-3 minutes of system slowdown as large parts of the model are loaded into RAM.
Strata provides a user-friendly interface to monitor progress and performance. Smaller model sizes are faster, while larger sizes offer more capabilities. Additional models can be added later. Full documentation on installation, troubleshooting, and model details is available. Strata is open source under the MIT License, with various licenses for individual components. The project is supported by donations to help with development.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.