{
  "id": 11954274,
  "title": "Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s",
  "url": "https://urgent.news/2026/10/04/run-qwen-3-8-flash-next-125b-on-consumer-hardware-rtx-4090-at-100t-s-11954274",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-04T12:51:53.000Z",
  "source": {
    "name": "Hacker News Best",
    "slug": "hacker-news-best",
    "url": "https://github.com/Niko1221/Strata"
  },
  "original_language": "en",
  "account": "Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100 tokens per second\n\nThe 125-billion-parameter AI model, Qwen 3.8 Flash Next, can now run on consumer gaming PCs with NVIDIA or AMD graphics cards (12 GB or more) using Strata, a free and open-source inference engine. The model, usually requiring server resources, now operates on a normal PC, capable of tasks such as chatting, writing code, reading images, and working with applications. Strata runs on Windows or Linux and requires a token to be approximately three-quarters of a word. Speed varies based on VRAM capacity, with an RTX 3090 (24 GB) achieving around 100-140 tokens per second. Multi-GPU setups are also possible, allowing for sharing the model between two or three cards. The entire installation process is automated, downloading the 70 GB model and setting up the appropriate engine for the user's graphics card. Strata requires around 35-55 GB of RAM and locks part of it for the graphics card, which is normal during the initial load. After the model starts, a Strata app can be accessed at http://127.0.0.1:8080, and the PC can be used while the model runs. Updating Strata is also straightforward, with an UPDATE.bat or ./update.sh script available.",
  "summary": "Article URL: https://github.com/Niko1221/Strata Comments URL: https://news.ycombinator.com/item?id=49953495 Points: 210 # Comments: 99",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Hacker News",
        "title": "Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s",
        "url": "https://urgent.news/2026/10/04/run-qwen-3-8-flash-next-125b-on-consumer-hardware-rtx-4090-at-100t-s",
        "published": "2026-10-04T12:51:53.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}