{
  "id": 3705388,
  "title": "Zhipu AI shares jump as viral Ox Alpha model revealed as GLM-5.3-Flash on Chinese chips",
  "url": "https://urgent.news/2026/08/27/zhipu-ai-shares-jump-as-viral-ox-alpha-model-revealed-as-glm-5-3-3705388",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-27T08:30:12.000Z",
  "source": {
    "name": "South China Morning Post",
    "slug": "south-china-morning-post",
    "url": "https://www.scmp.com/tech/big-tech/article/3365433/zhipu-ai-shares-jump-viral-ox-alpha-model-revealed-glm-53-flash-chinese-chips"
  },
  "original_language": "en",
  "account": "Zhipu AI, based in China, has unveiled its newest open-weight model, GLM-5.3-Flash, which was previously known as the code name Ox Alpha. The company claims the model was run entirely on a cluster of 100,000 domestically produced chips during a high-profile test. The announcement came after a week of heavy usage on AI model marketplaces OpenRouter and OpenCode, with the model handling 62 trillion tokens before its official release.\n\nThis achievement highlights China's capacity to manage large-scale global inference workloads using locally manufactured hardware, as Beijing aims to decrease dependency on advanced processors from US leader Nvidia due to stringent export controls. Upon its release, GLM-5.3-Flash soared to the top of global usage rankings on OpenRouter, processing 10.3 trillion tokens, or nearly 31 percent of the platform's weekly volume.\n\nTo contend with the limited memory capacity and bandwidth of individual Chinese chips, Zhipu developed a specialized inference engine that segregated processing stages into independent computing pools. This architecture increased end-to-end serving performance by three times compared to its initial baseline, achieving hardware efficiency and per-token costs equivalent to mainstream Nvidia accelerators. However, these claims have yet to be independently verified.\n\nGLM-5.3-Flash boasts 320 billion total parameters, activating 18 billion per request to minimize computing overhead. It is the first GLM-5 series model capable of processing visual information alongside text. Benchmarking firm Artificial Analysis rated the model 57 on its Intelligence Index, positioning it tenth globally and third among open-weight models, behind Moonshot AI's Kimi K3 and Alibaba Group Holding's Qwen3.8 2.4T A95B.\n\nDespite heavy traffic during the beta period, early developer feedback was mixed. While users commended the model's ability to debug complex code – handling 28 percent of 175 LiveCodeBench problems in community tests – others reported occasional hallucinations, dropped tasks, and sluggish generation. Artificial Analysis also noted that GLM-5.3-Flash's output speed was slower than the industry average.\n\nZhipu has made the model weights globally available and integrated GLM-5.3-Flash into its application programming interface, ZCode platform, and GLM Coding Plan. The launch occurs amidst heightened competition in China's open-source ecosystem, with Alibaba releasing Qwen3.8-Flash-Next, a multimodal preview of its upcoming Qwen4, which reportedly activates 6 billion of its 125 billion parameters to reduce inference costs.",
  "summary": "China’s Zhipu AI has launched its latest open-weight model, GLM-5.3-Flash – previously code-named Ox Alpha – saying that the system ran entirely on a cluster of 100,000 domestically produced chips during a high-profile stealth trial. The announcement followed a week of heavy traffic on artificial intelligence model marketplace OpenRouter and agent platform OpenCode, where the model processed 62…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 6,
    "also_reported_by": [
      {
        "outlet": "Techmeme",
        "title": "Alibaba releases Qwen3.8-Flash, an open-weight, 125B-parameter model built on its next-gen Qwen 4 architecture, saying it rivals Opus 4.6 and V4-Flash (Luz Ding/Bloomberg)",
        "url": "https://urgent.news/2026/08/26/alibaba-releases-qwen3-8-flash-an-open-weight-125b-parameter-model",
        "published": "2026-08-26T12:51:34.000Z"
      },
      {
        "outlet": "Techmeme",
        "title": "Z.ai releases GLM-5.3-Flash, the first natively multimodal GLM-5 series model, with 320B parameters, saying it outperforms GLM-5.2 at \"one-tenth the price\" (Z.ai)",
        "url": "https://urgent.news/2026/08/26/z-ai-releases-glm-5-3-flash-the-first-natively-multimodal-glm-5",
        "published": "2026-08-26T14:26:48.000Z"
      },
      {
        "outlet": "Economic Times Tech",
        "title": "Alibaba's Qwen launches Qwen3.8-Flash AI model with lower training costs",
        "url": "https://urgent.news/2026/08/26/alibabas-qwen-launches-qwen3-8-flash-ai-model-with-lower-training",
        "published": "2026-08-26T15:02:52.000Z"
      },
      {
        "outlet": "SiliconANGLE",
        "title": "Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash",
        "url": "https://urgent.news/2026/08/27/z-ai-open-sources-ox-alpha-model-as-glm-5-3-flash",
        "published": "2026-08-27T00:26:24.000Z"
      },
      {
        "outlet": "SCMP Tech",
        "title": "Zhipu AI shares jump as viral Ox Alpha model revealed as GLM-5.3-Flash on Chinese chips",
        "url": "https://urgent.news/2026/08/27/zhipu-ai-shares-jump-as-viral-ox-alpha-model-revealed-as-glm-5-3",
        "published": "2026-08-27T08:30:12.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}