Z.ai’s GLM-5.3-Flash is cheap, good, and served on Chinese chips
Ox-alpha, the stealth model that quickly became the most popular model on OpenRouter in the last few days, is actually The post Z.ai’s GLM-5.3-Flash is cheap, good, and served on Chinese chips appeared first on The New Stack .
Z.ai's GLM-5.3-Flash model, a hybrid with 320 billion parameters (including 18 billion active parameters), has been revealed by the team. This model is specifically designed for ultra-low-cost inference and is now available on Hugging Face under the MIT license. The model is also accessible on several inference platforms, including OpenRouter, where users can access it at $0.075 per million input tokens and $0.25 per million output tokens, with a 50% discount.
In terms of performance, GLM-5.3-Flash competes with models like Anthropic's Claude Fable 5, achieving similar benchmark scores to GPT-5.6 Terra, Google's Gemini 3.7 Flash, Meta's Muse Spark 1.2, and Qwen 3.8 2.4T A95B. Additionally, GLM-5.3-Flash excels in driving AI agents and visual tasks, making it suitable for real-world use cases.
Z.ai's model also benefits from the low inference costs, which is advantageous for smaller models that require more reasoning steps to reach this level of performance. One significant advantage of Z.ai's GLM-5.3-Flash is its pricing, which is considerably lower than other competing models. This low cost, combined with the availability of the model on Chinese AI chips, demonstrates that Chinese chips can efficiently and economically support frontier-model inference at scale.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.