Urgent.News

What's breaking now, across thousands of outlets.

AI

Z.ai's secret AI model Ox Alpha was running on Chinese chips all along

The model topped OpenRouter usage charts anonymously before Z.ai claimed it and released its weights on Wednesday

Z.ai's secret AI model Ox Alpha was running on Chinese chips all along

In January 2026, Z.ai chairman and cofounder Liu Debing marked the company's debut on the Hong Kong Stock Exchange. The mystery AI model Ox Alpha was actually Z.ai's new GLM-5.3-Flash model, which it revealed after days of speculation. Ox Alpha was free to try, impressing users with its coding abilities and remaining mysterious enough to leave developers guessing.

The Beijing AI company confirmed that Ox Alpha was its GLM-5.3-Flash model, which runs on Chinese AI chips despite US technology restrictions. Z.ai was founded in 2019 based on technology developed at Tsinghua University in Beijing and went public in January, with its shares trading at around HK$1,100, nine times its IPO price. The company reports its results on June 15th, 2026.

Written by urgent.news from Business Insider's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at qz.com →

More in AI

5 Undocumented Rules for Gemini Structured Output, Measured in Production

We run a document extraction pipeline on Gemini with a native responseSchema attached, not a "please reply with JSON" instruction in the prompt text.

  • Order required properties as per required array, then optional properties
  • Measure order effect on gemini-3-flash-preview model, adjust if needed
  • Enum value count limit exists, exceeded schemas rejected in production

Even an AI cost-management vendor can lose control of its agent spending

In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.

  • AI agents can cause unexpected expenses if not monitored properly
  • Revenium experienced $3,762 cost from AI coding assistant running 4 days
  • Top 1% of runs drive nearly half of total AI spending

More from Friday 28 August →