Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash
Z.ai Co. today released the code for GLM-5.3-Flash, a large language model that is ten times more cost-efficient than its predecessor. The algorithm made its original debut last week under the codename Ox Alpha. LLM marketplace operator OpenRouter Inc. launched a free hosted version of Ox Alpha and didn’t disclose its developer, which drew a […] The post Z.ai open-sources ‘Ox Alpha’ model as…
Z.ai has released the source code for its new large language model, GLM-5.3-Flash. This model is ten times more cost-effective than its predecessor, Ox Alpha. OpenRouter Inc. provided a free hosted version of Ox Alpha, generating considerable industry interest. Z.ai's GLM-5.3-Flash features a mixture of experts architecture with 320 billion parameters and can process up to 1 million tokens of text, images, or video in user requests, while generating up to 131,072 tokens in its responses.
The model's attention mechanism has been redesigned to reduce hardware overhead by analyzing only the most relevant tokens, using sparse and linear attention techniques. These innovations allow GLM-5.3-Flash to operate at a fraction of the cost and with superior efficiency compared to Z.ai's earlier LLMs. The model has demonstrated strong performance in various AI benchmarks, outperforming competitors such as Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash.
Z.ai trained GLM-5.3-Flash on a dataset of 30 trillion tokens using a technique called mHC to optimize the workflow, which minimizes the risk of gradient distortion during training. The model's weights are now accessible on Hugging Face.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- Alibaba releases Qwen3.8-Flash, an open-weight, 125B-parameter model built on its next-gen Qwen 4 architecture, saying it rivals Opus 4.6 and V4-Flash (Luz Ding/Bloomberg) bloomberg.com
- Z.ai releases GLM-5.3-Flash, the first natively multimodal GLM-5 series model, with 320B parameters, saying it outperforms GLM-5.2 at "one-tenth the price" (Z.ai) z.ai
- Alibaba's Qwen launches Qwen3.8-Flash AI model with lower training costs economictimes.indiatimes.com