Skip to main content

China's First 2-Trillion-Parameter AI Model Runs on New Alibaba Cloud Super Node

Alibaba Cloud has officially rolled out its latest super node instance, the Lingjun Zhenwu M890, marking a significant milestone for large-scale AI training and inference in China. The first batch is already on sale in the Ulanqab region, and it's the first in the country to successfully run a large model with over 2 trillion parameters.

What's powering this? The next-generation Zhenwu M890 AI chip from T-head, which supports a range of data precisions from FP32 down to FP4. Combined with the ICN Switch 1.0 interconnection chip, it achieves high-speed 800GB/s connectivity between 64 cards, with a memory pool of 9TB. That's enough to handle the expert parallel demands of massive MoE models.

Already, well-known models like Kimi K3 and Alibaba's own Qwen3.8 Max—with a parameter scale of up to 2.4 trillion—are running on this instance. The performance gains are impressive: in training scenarios like intelligent driving and embodied intelligence, it's up to three times faster than the previous generation Zhenwu 810E. For Agentic reasoning, it delivers up to 1.5 times improvement.

One of the biggest advantages? Enterprise customers don't need to build their own data centers. They can simply activate high-speed interconnected computing units on the cloud. A single instance can support inference for MoE models with up to 100 trillion parameters.

Under the hood, the Lingjun Zhenwu M890 relies on the unified Lingjun Intelligent Computing Platform, equipped with HPN 8.0 training and inference integrated network. Its single cluster can support up to 130,000 heterogeneous computing resources, and it can be flexibly expanded to millions of cards. That's a solid foundation for the explosive growth of AI workloads we're seeing today.

This launch isn't just about hardware—it's about making cutting-edge AI more accessible. By offering this super node on the cloud, Alibaba Cloud is lowering the barrier for enterprises that want to train and deploy massive models without the hefty upfront investment.

Key Points

  • First in China: Successfully runs a large model with over 2 trillion parameters.
  • Hardware: Zhenwu M890 chip with FP32 to FP4 precision, 800GB/s interconnect, 9TB memory pool.
  • Performance: Up to 3x faster in training scenarios, 1.5x in Agentic reasoning.
  • Availability: Now available in Ulanqab region; models like Kimi K3 and Qwen3.8 Max already running.
  • Scalability: Supports up to 130,000 heterogeneous computing resources per cluster, expandable to millions.