For nearly a week, a mystery AI model called “Ox Alpha” quietly topped coding-model leaderboards on OpenRouter and OpenCode without anyone knowing who had built it. On August 26, Chinese AI company Z.ai (formerly Zhipu AI) ended the guessing game, confirming that Ox Alpha was an early look at its new large language model, GLM-5.3-Flash, and releasing the weights on Hugging Face under a permissive MIT license.

A Stealth Run, Then a Reveal

Since around August 20, the unbranded “ox-alpha” model ran anonymously on the two platforms developers use to test unreleased models, processing tens of trillions of tokens and, at its peak, accounting for roughly 31% of OpenRouter’s weekly coding-model traffic, according to data cited by Bloomberg. Z.ai did not confirm authorship until the model’s public launch.

What GLM-5.3-Flash Brings

GLM-5.3-Flash is the first natively multimodal entry in Z.ai’s GLM-5 family, built to handle text, images and video. It uses a sparse mixture-of-experts design with 320 billion total parameters but only 18 billion active per query, paired with a context window beyond 1.3 million tokens. According to Z.ai, the model beats its predecessor, GLM-5.2, on coding and agentic benchmarks and approaches Anthropic’s Claude Opus 4.8, at a fraction of the cost. The company priced API access at $0.15 per million input tokens and $0.50 per million output tokens, cut 50% through September 9 as a launch promotion.

Built on Chinese Chips

Bloomberg reported that Z.ai trained and served the stealth model entirely on a cluster of domestically produced Chinese AI chips rather than Nvidia hardware — a detail that helped push Zhipu’s Hong Kong-listed shares up roughly 12%, reflecting investor interest in China’s push toward chip self-sufficiency amid US export controls.

The release also draws a line under a bumpier chapter for Z.ai: its flagship GLM-5.3 had its own open-weight release delayed earlier this month after engineers found a 45-year-old math bug it had inadvertently surfaced.

Read also