Urgent.News

What's breaking now, across thousands of outlets.

Tech

Qwen4 Isn’t Here Yet, but Qwen3.8-Flash-Next Tells Us a Lot

Qwen4 still isn’t officially here, but Qwen3.8-Flash-Next gives us something more useful than another round of release-date rumors. It gives us a look at the direction Qwen seems to be taking with the next generation. And the part that caught my attention isn’t the total parameter count. It’s how little of the model needs to be active at once. The 6B active number is more interesting than 125B…

Qwen4 has yet to be officially launched, yet the preview model Qwen3.8-Flash-Next offers valuable insights into the upcoming generation. Instead of focusing on the total parameter count of 125B, the noteworthy aspect is the limited number of active parameters during inference. With Qwen3.8-Flash-Next employing a sparse Mixture-of-Experts architecture, only around 6B parameters are active for each token.

This efficiency shifts the perspective from "bigger model equals higher costs" to a model capable of delivering robust reasoning, coding, and tool use without incurring the full inference expense of a dense model. When Qwen4 eventually arrives, the key to watch will be the proportion of active capacity during inference, routing behavior under real-world workloads, and whether this efficiency holds beyond benchmark conditions.

Another indication of Qwen's focus on long-context efficiency is the model's capacity to handle large context windows, extending up to the 1M-token range. However, the sheer number of tokens doesn't matter as much as the model's ability to stay useful when filled with extensive data. Instead of treating these results as Qwen4 benchmarks, they should be viewed as a preview of the direction Qwen is heading.

When Qwen4 releases, testing should focus on real-world applications such as coding, agent behavior, multimodal workflows, and cost efficiency. Comparing metrics like task success rate, latency, token usage, retries, tool-call failures, and cost per accepted task will provide a clearer picture than just looking at the parameter count.

Qwen4 might indeed be a large model, but the more intriguing aspect will likely be how much of it actually runs for each token. This is the aspect I will be keeping an eye on.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Does Brave Browser Load Faster Than Chrome, Firefox, and Microsoft Edge?

The Brave web browser published a new round of benchmarks claiming its desktop browser uses less system resources than Chrome, Microsoft Edge, and Firefox — and also loads pages faster.

  • Brave Browser uses 44% less CPU than Chrome, Edge, and Firefox
  • It consumes 28% less memory and 10% less energy
  • Page loads are 20% faster than Chrome, Edge, and Firefox

x402 Payment Required

AI stopped being a tool you use, it's becoming something that acts. Agents that don't just answer questions, but make decisions, execute tasks, and increasingly, act autonomously without a human…

More from Tuesday 8 September →