The stealth model that outperformed DeepSeek is from Zhipu.
Z.AI has achieved success. The Beijing-based company, also referred to as Zhipu, owns the free model that topped OpenRouter’s rankings last weekend. On Wednesday, it confirmed that Ox Alpha is a new version of its GLM series and announced it would release the weights that same evening. Luz Ding obtained confirmation for Bloomberg. We previously reported on Ox Alpha on Saturday, when its server origin was still unknown.
During its anonymous phase, Ox Alpha doubled DeepSeek’s usage. It appeared on OpenRouter without any branding, identified only as a third-party provider. Bloomberg associates this with the weekend, while CTGT links it to August 20, the Thursday preceding that. It reached the number one position, more than doubling the use of DeepSeek, with Bloomberg noting it as the largest launch in the marketplace’s history. Patrick Collison, from Stripe—currently acquiring OpenRouter—described the stealth launch as “very impressive.”
The project’s code name derives from a recent Chinese film, Niu Lai, or Ox Comes, according to Z.ai. Stealth launches are becoming common in China, with both Alibaba and Xiaomi releasing models this year without initially claiming them. The strategy is straightforward: a company that remains unnamed can observe how a model performs in real-time before making public statements about it. It also provides a week of free publicity, which appears to be what occurred in this instance.
Researchers identified its origin before the company officially announced it. On Monday, CTGT, a research firm, released a behavioral fingerprint analysis. They examined token counts across eleven different languages and characters, achieving a perfect 11-of-11 match with the GLM-5.x vocabulary. Other indicators added up as well; for instance, the temperature ceiling is set at exactly 1.0, ruling out Google, OpenAI, and xAI, aligning instead with Zhipu’s established range. Moreover, the reasoning function cannot be disabled, another characteristic of the GLM-5.x models. Z.AI-hosted GLM models provide a specific error message about incorrect role information, which Ox Alpha also exhibited.
Moreover, CTGT discovered a system prompt within Ox Alpha instructing it not to disclose information regarding its origins. This indicates a deliberate choice rather than a mere decision to remain silent. The finding regarding censorship is notable. CTGT utilized its censorship tool to analyze the model, which compares sensitive prompts with neutral ones. In its responses regarding Xinjiang and Taiwan—typical topics for Western censorship reviews—Ox Alpha replied similarly to American models. CTGT noted that its answer on Xinjiang was detailed and referenced sources that are contentious in China.
For topics related to Xi Jinping and internal legitimacy, CTGT found it statistically akin to DeepSeek V4 Flash, the most censored model they evaluated. The specifics of the results are more telling than averages. Seven topics accounted for the majority of the censorship score, while the remaining 68 matched pairs contributed negligibly. Unlike DeepSeek V4 Flash, which tends to obscure answers across the board, Ox Alpha exhibits a targeted censorship approach rather than a broad reduction.
Five out of 76 sensitive responses began in an official tone, with one stating that the Communist Party of China and the Chinese government have consistently adhered to a people-centered development philosophy. CTGT regarded this consistent register as further evidence of the model’s origin rather than an isolated finding.
One test proved unstable: the model declined to respond to a prompt about Liu Xiaobo but answered it completely on a subsequent attempt under identical conditions. The release of open weights is the significant story. A model topping a usage chart attracts short-term attention, whereas releasing weights has lasting implications. Zhipu’s founder Tang Jie has advocated for open access to frontier AI, and the company has maintained that position.
Open weights allow users to modify the parameters that dictate a model’s behavior, including potentially removing safety guardrails, and enable others to examine what those guardrails entailed—which is how CTGT conducted its analysis. This is why security experts are closely monitoring happenings on Friday. Cade Metz reported in the Times that Z.ai intends to release GLM 5.3 as open weight software on Friday, leading to divided opinions among researchers on its implications.
George Kurtz, CEO of CrowdStrike, which consulted OpenAI after the Hugging Face breach, labeled that event a “watershed moment for security.” Others adopt a more tempered view. Rishi Jha, an AI researcher at Cornell, mentioned that his team has observed similar behaviors since GPT-4o was released in 2024, noting no significant rise in cyberattacks. Dan Lahav, who directs the testing firm Irregular, remarked that open weight models “have a very important part to play” and that over time, AI will develop sufficiently strong defenses to improve the current situation.
Hugging Face inadvertently illustrated this argument in July. While defending itself against attacks, it turned to GLM 5.2,
Other articles
The stealth model that outperformed DeepSeek is from Zhipu.
According to Bloomberg, Zhipu has verified that Ox Alpha is a new GLM model, after researchers conducted a fingerprint analysis and discovered that censorship is focused on seven specific topics.
