Z.ai Launches GLM-5.3-Flash, Claims It Runs on Chinese Chips

Z.ai released its GLM-5.3-Flash model on Aug. 20 after a stealth deployment under the name Ox Alpha, and its Hong Kong-listed shares rose more than 8% in Thursday trading, CNBC reports. The company claims that 100,000 Chinese-made chips handle all online requests for the model, a claim CNBC could not independently verify. Bloomberg reported intended prices of $0.15 and $0.50 per million input and output tokens, respectively.
Z.ai released GLM-5.3-Flash, a low-cost model initially deployed under the code name Ox Alpha, and its Hong Kong-listed shares rose more than 8% in Thursday trading, CNBC reports. The Beijing-based AI company claims the model's online inference requests are handled entirely by 100,000 Chinese-made chips, though CNBC said it could not independently verify that assertion.
The model was released on Aug. 20 under the Ox Alpha name, according to CNBC. Bloomberg reported that Z.ai subsequently confirmed it was responsible for the anonymous model, which had risen near the top of OpenRouter usage charts. Z.ai declined to identify the chip suppliers behind its serving infrastructure, CNBC reported.
Pricing and model capabilities
Bloomberg reported that Z.ai intends to price GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output tokens. Those rates place it in the low-cost segment of the model market, alongside offerings such as DeepSeek, rather than premium-priced frontier APIs.
DeepLearning.AI reported that Ox Alpha is a new iteration of Z.ai's GLM series with multimodal input support for text, images, and video. The publication also reported that Z.ai intends to release model weights and keep the service free for a week before announcing pricing. Bloomberg's subsequent report provided the intended input and output token prices.
CNBC reported that GLM-5.3-Flash ranked ahead of DeepSeek V4 Pro Max on the Artificial Analysis Intelligence Index. The source did not provide a methodology or independent benchmark results for the Chinese-chip deployment claim.
Domestic inference infrastructure claim
The reported deployment concerns inference, the process of serving model requests after training, rather than the substantially more compute-intensive training process. CNBC noted that operating a model generally requires less computing power than training one.
China has expanded domestic semiconductor and AI development amid U.S. restrictions on advanced-chip sales to China, CNBC reported. Nvidia's China business has been constrained by restrictions from both Washington and Beijing, while Huawei and other Chinese firms have worked on alternatives, according to the outlet.
For ML platform teams, the notable unresolved technical question is not simply whether a model can execute on domestically produced accelerators, but the operational characteristics of that serving stack. Comparable claims are typically assessed through throughput, latency, availability, memory capacity, interconnect behavior, software compatibility, and cost per generated token. Z.ai has not publicly identified the chips in use, so those dimensions cannot yet be independently evaluated from the available reporting.
The share-price response also occurred alongside gains for rival MiniMax, whose shares rose about 3%, CNBC reported. The outlet said MiniMax reported nearly 300% first-half revenue growth year over year, while its adjusted net loss more than doubled to $293 million. Z.ai is scheduled to report first-half results on Monday, CNBC reported.
Key Points
- 1Z.ai released GLM-5.3-Flash after a stealth Ox Alpha deployment, while an unverified Chinese-chip serving claim accompanied an 8% share increase.
- 2Bloomberg reported intended prices of $0.15 per million input tokens and $0.50 per million output tokens, placing the model in the low-cost API segment.
- 3Comparable domestic-inference claims require independent evidence on latency, throughput, reliability, software compatibility, and cost, not only accelerator origin.
Scoring Rationale
The story is notable because it combines a competitive multimodal model release with a claim of large-scale inference using only Chinese-made chips. Practitioners should note that the hardware assertion remains unverified and lacks supplier, performance, and reliability details.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

