Urgent.News

What's breaking now, across thousands of outlets.

AI

Running Qwen 3.8 Flash Next (125B) on a RTX 4090 – 100 T/s on a Desktop

1. What was released / announced Niko1221’s Strata repo shows that the new Qwen 3.8 Flash Next (125B) model can be run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second (T/s) . In practice, that means you can generate text at near‑real‑time speed without a multi‑GPU server or a cloud‑based inference endpoint. The repo ships a set of scripts, quantisation tricks, and a…

Niko1221’s Strata repository reveals that the new Qwen 3.8 Flash Next (125B) model can run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second (T/s). This breakthrough effectively democratizes access to massive language models (LLMs), as previously only attainable via expensive API contracts or multi‑node GPU clusters. By quantizing the model to 4‑bits and utilizing torch.compile along with a custom CUDA kernel, Strata squeezes optimal performance from the 24 GB VRAM card of the RTX 4090.

The implications for developers are profound. Running inference on-premises eliminates costly cloud token fees and reduces latency to sub‑second levels, enabling real‑time assistants, low-latency code completion, and other edge‑deployments. Understanding the quantization and compilation techniques opens opportunities to apply similar methods to other large models such as LLaMA‑3 or Gemma‑2.

To utilize the model, one must first install the necessary dependencies including Python 3.11, PyTorch, and Bitsandbytes. Next, clone the Strata repository and install the package within the virtual environment. Download the 4‑bit quantized model from the Hugging Face repository, which requires approximately 30 GB of disk space. Finally, run the inference script, which demonstrates the model’s ability to generate text at near‑real‑time speeds on an RTX 4090.

Strata’s approach demonstrates that quantization is the key to unlocking high performance on consumer GPUs, while torch.compile optimizes the execution pipeline. The operational simplicity of deploying a single node also lowers the barriers for small teams and startups. While there are caveats such as GPU memory fragmentation, the overall benefits for rapid prototyping, data-sensitive workloads, and cost-conscious teams make this development a significant milestone in accessible AI infrastructure.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Move a Claude Code task to a teammate without sharing a login

A coding task often stalls because the person who wrote the brief cannot run it right now. Passing a provider login to someone else blurs who executed the work and what access they had.

  • Handoff file includes task description, outcome, scope, and starting details
  • Executor uses own account with Claude Code access, no shared login
  • Reviewer checks changed files, commit/PR URL, test results, and visual check

I Run Claude Code on $0/Month: A Practical Guide to Free LLM APIs

I Run Claude Code on $0/Month: A Practical Guide to Free LLM APIs AI coding assistants are incredible — until the API bill arrives.

  • Free LLM APIs available for coding without credit card
  • Google AI Studio and Groq recommended for Claude Code
  • Free tiers have daily request limits and varying quality

The 'Terminator scenario': How realistic is the AI apocalypse?

OpenAI boss Sam Altman and Anthropic chief Dario Amodei warn of the risks of overly rapid AI development. More experts now think AI could wipe out humanity. How realistic is that?

  • Skynet AI system breaks free in 2029, leading to nuclear war and autonomous killer robots.
  • Artificial intelligence expert Reinhard Karger warns of unintended catastrophic consequences.

More from Monday 5 October →