~/wiki / novosti / ox-alpha-glm-5-3-flash

Ox Alpha was GLM-5.3-Flash: how the stealth model Z.ai topped the charts OpenRouter

Main chat

A chat for vibe coders: news, guides, live cases, marketplace, and finding executors.

$ cd section/ $ join vibe dev
Ox Alpha was GLM-5.3-Flash: how the stealth model Z.ai topped the charts OpenRouter - обложка

Mysterious stealth model Ox Alpha, which since August 20 free distributed on OpenRouter and OpenCode and along the way climbed to the top of the usage charts, was GLM-5.3-Flash from Z.ai. On August 26, the company revealed the authorship and posted the weights under the MIT license on Hugging Face.

  • Multimodal MoE model 320B parameters, 18B active (320B-A18B).
  • Context: 1,048,576 tokens (1M), input – text, images and video.
  • Why do you need: a powerful and very cheap tool for coding and agency tasks, with open weights.
  • Caution: Free stealth mode is over, public prices and limits are different.

What happened

On August 20, 2026, an anonymous model named stealth/ox-alpha appeared on OpenRouter. No label, no vendor, no description – only a card with a giant context in 1M tokens, support for pictures and videos at the entrance and a price of $ 0 for tokens and entry and exit. In parallel, it began to distribute in OpenCode.

Then happened what should happen to a powerful free model: a few days Ox Alpha climbed to the top of the usage charts, in places ahead of DeepSeek. The developers chased it on real tasks, compared it with Claude and GPT, wondering who is behind the release. By August 25, Bloomberg wrote about the “mysterious model”.

On August 26, Z.ai closed the intrigue: Ox Alpha is a preview run of the GLM-5.3-Flash, which was tested anonymously to collect honest feedback without brand effect. At the same time, a named release was released, and free stealth routes were replaced by ordinary ones - already with public prices and limits.

Why did you need an anonymous launch

Adopting with an “unnamed” model is not a marketing trick for the sake of hype, but a way to remove the distortion of expectations from the evaluation. When the card says “GLM” or “OpenAI,” part of the quality judgment is formed in advance. Anonymous ox-alpha made people judge by the result: how the model actually fixes bugs, holds a long context and behaves in an agency cycle.

A side effect is an honest infrastructure stress test. Free access instantly overtakes the load, and the vendor sees the behavior of the model under real traffic, not in the laboratory. For Z.ai, it’s also a way to show that the previous GLM-5 was not a one-off success.

What is known about GLM-5.3-Flash

According to Z.ai and the first analysis, the key characteristics look like this:

  • Architecture: Mixture-of-Experts, 320B parameters, 18B active per token. This “Flash” line is a bet on speed and cost, not on maximum “raw” intelligence.
  • Multimodality: native, not twisted sideways. Text, images and videos are accepted at the entrance.
  • Context: 1,048,576 tokens. This is enough to keep a large codebase or a long session of the agent in the window without aggressive slicing.
  • License: MIT, weights published on Hugging Face – can be run locally and used commercially.
  • Iron: Reportedly (particularly Quartz), the model is trained on Chinese chips, which continues the GLM-5 line trained on Huawei Ascend.
$ cd ../ ← back to News

$ nav --prev

OpenAI Opens Free ChatGPT Chats with GPT-5.6 Luna

$ nav --next

OpenAI launches GPT-6 Astra in what the company calls the “AGI era”