Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Z.ai (Zhipu AI)

z.ai

Z.ai, formerly Zhipu AI, is the Beijing lab behind the GLM family and the most credible open-weight challenger to the closed frontier. GLM-5.2 shipped June 13, 2026: a 744 billion parameter mixture-of-experts model with a 1 million token context window, 131K max output, a Max-effort reasoning mode, and MIT-licensed weights, priced around $1.40 input and $4.40 output per million tokens. It sits fourth overall and first among open-weight models on the Artificial Analysis Intelligence Index at 51, with a vendor-reported 62.1 on SWE-Bench Pro that would top GPT-5.5. The strategically important fact is the hardware: the training pipeline ran on roughly 100,000 Huawei Ascend 910B chips with zero Nvidia in the loop, at an estimated $25 million all-in, which makes GLM the proof case that frontier-adjacent training no longer requires US silicon. Reporting puts GLM-5.2 at something like 40 percent of developer tokens flowing through OpenRouter. The practical catch is that full-precision self-hosting needs about 1.5TB of GPU memory, roughly nineteen H100s. Two releases in August 2026 extended the family in different directions. GLM-5.3 launched API-first on August 14 at the same $1.40/$4.40 rate, built by extended post-training on the GLM-5.2 base rather than a new pretrain, and its weights reached Hugging Face on August 28 after a two-week safety review, under a custom GLM-5.3 License that is permissive except for one clause: Model-as-a-Service operators with more than $10 billion in group revenue must pass a Z.ai security review before commercial use. GLM-5.3 Flash followed on August 26 as a separate model rather than a cheaper 5.3: the first natively multimodal GLM-5, a 320 billion parameter MoE activating roughly 18 billion per token, 1M context, MIT weights, and list pricing of $0.15 input and $0.50 output, roughly a tenth of the text flagship on input. A 50 percent launch promotion on Flash ended September 9, 2026. On September 18 Z.ai added GLM-5.3-FlashX, which is the same 320B weights on a faster serving stack rather than a new model: $0.37 input, $0.075 cached, and $1.25 output, roughly two and a half times the base Flash rate, in exchange for an advertised 200 output tokens per second. It is a hosted tier with no separate weight release, so the premium buys the serving stack and nothing else. On September 13, five days after a joint CISA, NSA, and FBI advisory named Z.ai among six China-based labs accused of industrial-scale distillation of US models, the Hong Kong-listed company disclosed roughly $5 billion in new funding: a placement of up to 21.97 million H shares at HK$714 and RMB 20.14 billion of zero-coupon convertible bonds due 2027.

Founded

2019

Headquarters

Beijing, China

CEO

Tang Jie

Models

4 active

Key Products

GLM-5.3GLM-5.3 FlashGLM-5.3-FlashXGLM-5.2GLM-5.1GLM Coding PlanZ.ai chatbot

Strengths

  • ✓First among open-weight models on the Artificial Analysis index
  • ✓Open weights with 1M context (MIT on GLM-5.2 and GLM-5.3 Flash, a custom license on GLM-5.3)
  • ✓Trained end to end on Huawei Ascend silicon
  • ✓Roughly 80 percent cheaper than Opus 4.8
  • ✓GLM-5.3 Flash brings native multimodality at $0.15/$0.50
  • ✓Heavy OpenRouter developer adoption

Z.ai (Zhipu AI) Models

ModelInput / 1MOutput / 1MContextCapabilities
GLM-5.31.404.401.0Mtext, tool-use, code, reasoning
GLM-5.21.404.401Mtext, code, tool-use, reasoning
GLM-5.3-FlashX0.371.251.0Mtext, vision, video, tool-use, code, reasoning
GLM-5.3 Flash0.150.501.0Mtext, vision, video, tool-use, code, reasoning

Prices per 1M tokens in USD. See the full pricing guide for detailed analysis.

Comparisons

Other Providers

Visit Z.ai (Zhipu AI)Compare ModelsAll Models