{
  "id": 12130920,
  "title": "Running Qwen 3.8 Flash Next (125B) on a RTX 4090 – 100 T/s on a Desktop",
  "url": "https://urgent.news/2026/10/05/running-qwen-3-8-flash-next-125b-on-a-rtx-4090-100-t-s-on-a-desktop",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-05T10:47:57.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/naveenmalothu/running-qwen-38-flash-next-125b-on-a-rtx-4090-100-ts-on-a-desktop-2ao7"
  },
  "original_language": "en",
  "account": "Niko1221’s Strata repository reveals that the new Qwen 3.8 Flash Next (125B) model can run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second (T/s). This breakthrough effectively democratizes access to massive language models (LLMs), as previously only attainable via expensive API contracts or multi‑node GPU clusters. By quantizing the model to 4‑bits and utilizing torch.compile along with a custom CUDA kernel, Strata squeezes optimal performance from the 24 GB VRAM card of the RTX 4090.\n\nThe implications for developers are profound. Running inference on-premises eliminates costly cloud token fees and reduces latency to sub‑second levels, enabling real‑time assistants, low-latency code completion, and other edge‑deployments. Understanding the quantization and compilation techniques opens opportunities to apply similar methods to other large models such as LLaMA‑3 or Gemma‑2.\n\nTo utilize the model, one must first install the necessary dependencies including Python 3.11, PyTorch, and Bitsandbytes. Next, clone the Strata repository and install the package within the virtual environment. Download the 4‑bit quantized model from the Hugging Face repository, which requires approximately 30 GB of disk space. Finally, run the inference script, which demonstrates the model’s ability to generate text at near‑real‑time speeds on an RTX 4090.\n\nStrata’s approach demonstrates that quantization is the key to unlocking high performance on consumer GPUs, while torch.compile optimizes the execution pipeline. The operational simplicity of deploying a single node also lowers the barriers for small teams and startups. While there are caveats such as GPU memory fragmentation, the overall benefits for rapid prototyping, data-sensitive workloads, and cost-conscious teams make this development a significant milestone in accessible AI infrastructure.",
  "summary": "1. What was released / announced Niko1221’s Strata repo shows that the new Qwen 3.8 Flash Next (125B) model can be run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second (T/s) . In practice, that means you can generate text at near‑real‑time speed without a multi‑GPU server or a cloud‑based inference endpoint. The repo ships a set of scripts, quantisation tricks, and a…",
  "key_points": [
    "Qwen 3.8 Flash Next (125B) model runs on RTX 4090 at 100 T/s",
    "Model quantized to 4-bits with torch.compile and custom CUDA kernel",
    "Enables on-premises inference, eliminating cloud token fees and reducing latency"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}