Three prominent Chinese artificial intelligence laboratories have established a new benchmark for the open-weight landscape with the release of trillion-parameter-scale Mixture-of-Experts (MoE) models. Moonshot AI’s Kimi K3, DeepSeek V4 Pro, and Zhipu AI’s GLM-5.2 have emerged as the leading contenders for long-horizon coding and agent-based workloads. Each model features a 1-million-token context window, but they differ significantly in their current accessibility, cost structures, and technical performance.
Capability and Benchmarking
Measured capability varies across the three models, with Kimi K3 currently holding the highest position on the Artificial Analysis Intelligence Index. Scoring approximately 57, Kimi K3 ranks third overall, trailing only Claude Fable 5 and GPT-5.6 Sol. GLM-5.2 follows with a score of 51, while DeepSeek V4 Pro reaches 44. While vendor-reported benchmarks can be difficult to compare directly due to differing testing harnesses, Moonshot’s internal data indicates that Kimi K3 consistently outperforms GLM-5.2 across a variety of coding and reasoning tasks, including SWE-bench and GPQA-Diamond.
DeepSeek V4 Pro remains highly competitive in specific coding applications. It achieved an 80.6% score on SWE-bench Verified, tying it with Gemini 3.1 Pro, and demonstrated strong long-context performance with an 83.5 score on MRCR 1M. GLM-5.2, while trailing Kimi K3 in peak benchmark scores, maintains a strong position as a capable open-weight option that previously led the field before the launch of K3.
Licensing and Accessibility
The practical deployment of these models is currently dictated by their respective licensing terms. DeepSeek V4 Pro and GLM-5.2 are both available under the MIT license, allowing for unrestricted commercial use, fine-tuning, and self-hosting. Their weights are currently accessible via Hugging Face.
Kimi K3 operates under different constraints. As of July 18, 2026, it is available only through API access and Kimi applications. Moonshot AI has committed to releasing the model weights by July 27, 2026, under a Modified MIT license. This license includes an attribution clause that triggers only for entities with more than 100 million monthly active users.
Serving Costs and Technical Requirements
API pricing reveals a significant divide between the models. DeepSeek V4 Pro is the most cost-effective option, with Artificial Analysis reporting a cost per task of $0.04, compared to $0.32 for GLM-5.2 and $0.94 for Kimi K3. In terms of raw output, one dollar buys approximately 1.15 million tokens from DeepSeek V4 Pro, 227,000 from GLM-5.2, and 67,000 from Kimi K3.
Speed and infrastructure requirements also vary. GLM-5.2 is the fastest of the group, measured at approximately 168 tokens per second, while Kimi K3 and DeepSeek V4 Pro operate at roughly 62 tokens per second. Self-hosting these models presents a substantial hardware challenge; GLM-5.2 requires over 1TB of VRAM in BF16, and Kimi K3 is the most resource-intensive, with Moonshot recommending 64 or more accelerators for local deployment. While Kimi K3 utilizes MXFP4 weights and MXFP8 activations to improve hardware compatibility, it remains the most difficult of the three to host locally.

Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!