GLM-5.3-Flash: Z.ai Reveals Ox Alpha Was Its Open Multimodal Model
For the past week, developers have been puzzling over a model called Ox Alpha. It appeared on OpenCode and OpenRouter on August 20 with no owner attached, free to use, with a 1M-token context window and support for image and video input. Independent researchers fingerprinted its tokenizer, ran compression analyses, and traced it to Z.ai's GLM family with high confidence. On August 26, Z.ai…
Z.ai recently announced that its Ox Alpha model, which had been circulating online with no owner credit, was actually the GLM-5.3-Flash model. This model is the first to be natively multimodal in the GLM-5 series, offering support for text, images, and video. The model features a 320 billion parameter Mixture-of-Experts architecture and a 1 million token context window. Z.ai claims the model offers frontier-adjacent performance for roughly one-tenth the price of its predecessor.
The model's architecture is particularly noteworthy, as it combines sparse and linear attention methods to handle the computational cost of processing 1 million tokens. This allows the model to maintain high performance while being relatively inexpensive to serve. Z.ai also developed an indexing technique called IndexPool to further reduce the memory and latency overhead of the attention mechanism.
In terms of performance, GLM-5.3-Flash scores highly on various benchmark tests. On the Artificial Analysis Intelligence Index v4.1.1, it scored 57 at a discounted price, making it comparable in intelligence to models that previously cost significantly more. The model also outperformed several other models in automation, tool usage, and coding tasks. However, it falls short in some vision-based benchmarks, particularly those involving raw video understanding.
Overall, GLM-5.3-Flash represents a significant step forward in multimodal AI, offering powerful capabilities at a fraction of the cost of previous cutting-edge models. As Z.ai continues to refine and expand the model's capabilities, it will be interesting to see how it performs in more real-world applications.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.