Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4”
Alibaba this week unveiled Qwen3.8-Flash, an open-weight, multimodal Mixture-of-Experts (MoE) model. Hot on the heels of Qwen 3.8 Max, which The post Alibaba just released Qwen3.8-Flash: “An early preview of the architecture in Qwen4” appeared first on The New Stack .
Alibaba has recently launched Qwen3.8-Flash, an open-weight multimodal Mixture-of-Experts model, as an early preview of the Qwen4 architecture. This 125-billion-parameter AI model serves as a performance and value-for-money alternative to Qwen 4. By releasing Qwen3.8-Flash early, Alibaba aims to enable the community to examine and test the mechanics, constructs, and components of the architecture before the full Qwen4 model is developed.
The model is positioned to offer superior capabilities in coding and office tasks, while balancing capability, latency, and cost. Alibaba claims that Qwen3.8-Flash has a systematic improvement across four aspects: attention, residual, embedding, and optimization. This optimization leads to enhanced model capability, computational efficiency, model capacity, and training stability.
Benchmarking Qwen3.8-Flash against rival models, such as DeepSeek-V4-Flash and Claude-Opus-4.6, shows that it performs well in agentic coding, long-horizon agent tasks, and multimodal intelligence. Qwen3.8-Flash-Next, the open-weight research frontier model, features a 125B parameter main model, supplemented by an additional 51B N-gram embeddings, and supports up to 1,000,000 tokens of context.
Alibaba asserts that Qwen3.8-Flash requires only around one-ninth of the training resources of Qwen3.7-Plus, while delivering superior performance and significantly reducing both training and inference costs.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.