{
  "id": 3622648,
  "title": "Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash",
  "url": "https://urgent.news/2026/08/27/z-ai-open-sources-ox-alpha-model-as-glm-5-3-flash",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-27T00:26:24.000Z",
  "source": {
    "name": "SiliconANGLE",
    "slug": "siliconangle",
    "url": "https://siliconangle.com/2026/08/26/z-ai-open-sources-ox-alpha-model-as-glm-5-3-flash/"
  },
  "original_language": "en",
  "account": "Z.ai has released the source code for its new large language model, GLM-5.3-Flash. This model is ten times more cost-effective than its predecessor, Ox Alpha. OpenRouter Inc. provided a free hosted version of Ox Alpha, generating considerable industry interest. Z.ai's GLM-5.3-Flash features a mixture of experts architecture with 320 billion parameters and can process up to 1 million tokens of text, images, or video in user requests, while generating up to 131,072 tokens in its responses. The model's attention mechanism has been redesigned to reduce hardware overhead by analyzing only the most relevant tokens, using sparse and linear attention techniques. These innovations allow GLM-5.3-Flash to operate at a fraction of the cost and with superior efficiency compared to Z.ai's earlier LLMs. The model has demonstrated strong performance in various AI benchmarks, outperforming competitors such as Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash. Z.ai trained GLM-5.3-Flash on a dataset of 30 trillion tokens using a technique called mHC to optimize the workflow, which minimizes the risk of gradient distortion during training. The model's weights are now accessible on Hugging Face.",
  "summary": "Z.ai Co. today released the code for GLM-5.3-Flash, a large language model that is ten times more cost-efficient than its predecessor. The algorithm made its original debut last week under the codename Ox Alpha. LLM marketplace operator OpenRouter Inc. launched a free hosted version of Ox Alpha and didn’t disclose its developer, which drew a […] The post Z.ai open-sources ‘Ox Alpha’ model as…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 4,
    "also_reported_by": [
      {
        "outlet": "Techmeme",
        "title": "Alibaba releases Qwen3.8-Flash, an open-weight, 125B-parameter model built on its next-gen Qwen 4 architecture, saying it rivals Opus 4.6 and V4-Flash (Luz Ding/Bloomberg)",
        "url": "https://urgent.news/2026/08/26/alibaba-releases-qwen3-8-flash-an-open-weight-125b-parameter-model",
        "published": "2026-08-26T12:51:34.000Z"
      },
      {
        "outlet": "Techmeme",
        "title": "Z.ai releases GLM-5.3-Flash, the first natively multimodal GLM-5 series model, with 320B parameters, saying it outperforms GLM-5.2 at \"one-tenth the price\" (Z.ai)",
        "url": "https://urgent.news/2026/08/26/z-ai-releases-glm-5-3-flash-the-first-natively-multimodal-glm-5",
        "published": "2026-08-26T14:26:48.000Z"
      },
      {
        "outlet": "Economic Times Tech",
        "title": "Alibaba's Qwen launches Qwen3.8-Flash AI model with lower training costs",
        "url": "https://urgent.news/2026/08/26/alibabas-qwen-launches-qwen3-8-flash-ai-model-with-lower-training",
        "published": "2026-08-26T15:02:52.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}