OpenClaw + GLM 5.3 Flash — Opus-Class Scores at Flash Cost
For a week an anonymous model called ox-alpha sat at the top of OpenRouter and OpenCode with nobody knowing who made it. On 26 August 2026 Z.ai (Zhipu AI) revealed it as GLM 5.3 Flash and open-sourced the weights under MIT the same day. Here is what it actually is, what the benchmarks support, and how to run it on an OpenClaw agent. What GLM 5.3 Flash Is GLM 5.3 Flash is a mixture-of-experts…
On 26 August 2026, Z.ai (Zhipu AI) unveiled GLM 5.3 Flash, an open-source multimodal model with 320 billion total parameters and approximately 18 billion active per token. This is the first natively multimodal model in the GLM-5 series, capable of processing text, images, and video, and generating text as output. GLM 5.3 Flash is designed to maintain frontier-level capability while reducing the cost of serving it.
It achieves this through hybrid sparse + linear attention, which combines local dependencies and global context retrieval. The model also incorporates mHC for scaling efficiency and IndexPool to compress key vectors for affordable 1M-token contexts. GLM 5.3 Flash utilizes a 30T-token multimodal pre-training corpus and has 45 layers compared to the similar-sized GLM-4.5's 92 layers.
The model was tested anonymously as ox-alpha on OpenCode and OpenRouter before release to gather unbiased user feedback. When running GLM 5.3 Flash on OpenClaw, users can simply add their OpenRouter key under BYOK settings in their dashboard, select the model, and it will apply to the running agent immediately.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.