Urgent.News

What's breaking now, across thousands of outlets.

AI

Zhipu AI shares jump as viral Ox Alpha model revealed as GLM-5.3-Flash on Chinese chips

China’s Zhipu AI has launched its latest open-weight model, GLM-5.3-Flash – previously code-named Ox Alpha – saying that the system ran entirely on a cluster of 100,000 domestically produced chips during a high-profile stealth trial. The announcement followed a week of heavy traffic on artificial intelligence model marketplace OpenRouter and agent platform OpenCode, where the model processed 62…

Zhipu AI shares jump as viral Ox Alpha model revealed as GLM-5.3-Flash on Chinese chips

Zhipu AI, based in China, has unveiled its newest open-weight model, GLM-5.3-Flash, which was previously known as the code name Ox Alpha. The company claims the model was run entirely on a cluster of 100,000 domestically produced chips during a high-profile test. The announcement came after a week of heavy usage on AI model marketplaces OpenRouter and OpenCode, with the model handling 62 trillion tokens before its official release.

This achievement highlights China's capacity to manage large-scale global inference workloads using locally manufactured hardware, as Beijing aims to decrease dependency on advanced processors from US leader Nvidia due to stringent export controls. Upon its release, GLM-5.3-Flash soared to the top of global usage rankings on OpenRouter, processing 10.3 trillion tokens, or nearly 31 percent of the platform's weekly volume.

To contend with the limited memory capacity and bandwidth of individual Chinese chips, Zhipu developed a specialized inference engine that segregated processing stages into independent computing pools. This architecture increased end-to-end serving performance by three times compared to its initial baseline, achieving hardware efficiency and per-token costs equivalent to mainstream Nvidia accelerators. However, these claims have yet to be independently verified.

GLM-5.3-Flash boasts 320 billion total parameters, activating 18 billion per request to minimize computing overhead. It is the first GLM-5 series model capable of processing visual information alongside text. Benchmarking firm Artificial Analysis rated the model 57 on its Intelligence Index, positioning it tenth globally and third among open-weight models, behind Moonshot AI's Kimi K3 and Alibaba Group Holding's Qwen3.8 2.4T A95B.

Despite heavy traffic during the beta period, early developer feedback was mixed. While users commended the model's ability to debug complex code – handling 28 percent of 175 LiveCodeBench problems in community tests – others reported occasional hallucinations, dropped tasks, and sluggish generation. Artificial Analysis also noted that GLM-5.3-Flash's output speed was slower than the industry average.

Zhipu has made the model weights globally available and integrated GLM-5.3-Flash into its application programming interface, ZCode platform, and GLM Coding Plan. The launch occurs amidst heightened competition in China's open-source ecosystem, with Alibaba releasing Qwen3.8-Flash-Next, a multimodal preview of its upcoming Qwen4, which reportedly activates 6 billion of its 125 billion parameters to reduce inference costs.

Written by urgent.news from South China Morning Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at scmp.com →

More in AI

100 signs of Silicon Shock

We are biased. Given our faith in the rapidly rising utility of AI that can only be done on a completely new type of hardware, we wrote in January about the theme of this era: Silicon Shock.

More from Thursday 27 August →