Skip to main content

China Telecom's New AI Model Runs on Consumer GPUs, Fully Domestic

China Telecom's New AI Model Runs on Consumer GPUs, Fully Domestic

China Telecom has quietly dropped a bombshell in the AI world. On September 17, the telecom giant unveiled its next-generation Xing4.0-29B-A4B large model, and it's turning heads for two big reasons: it's fully domestic, and it can run on consumer-grade GPUs.

Let's break down what that means.

A Fully Domestic Stack

You've probably heard the phrase "full stack" thrown around, but here it's not just marketing. From the Ascend chips used for training to the domestic frameworks and self-developed architecture, every piece of this model is made in China. No reliance on overseas supply chains. In a world where computing power is often constrained, holding the entire chain in your own hands is a big deal.

Performance That Speaks for Itself

But does it actually perform? In the SuperCLUE evaluation, its agent capabilities scored 93.52 points, landing it in third place overall. The gap between it and the top two models, both from Qwen, is less than one point. That's like a lightweight boxer going toe-to-toe with the heavyweights.

Runs on Your Gaming Rig

Here's where it gets really interesting for developers. After 4-bit quantization, the model's VRAM usage drops to just 15GB—a 75% reduction compared to FP16. That means a consumer-grade card like the RTX 3090 or RTX 4090 with 24GB of memory can run long-context tasks locally. No need for server farms or renting cloud GPUs. Individuals and small teams can now experiment with a powerful model right on their desktop.

Image

What This Means for the AI Community

China Telecom's move isn't just about technical bragging rights. It's a statement: domestic AI can be both powerful and accessible. By open-sourcing the model, they're inviting developers to build on top of it, potentially accelerating innovation across the board.

Of course, challenges remain. The ecosystem around domestic chips and frameworks is still maturing, and competition is fierce. But this release shows that the gap is closing—and fast.

Image

Key Points

  • Fully domestic: From Ascend chips to frameworks, no overseas dependencies.
  • Efficient: 29B parameters, only 4B activated, handles up to 512K context.
  • High performance: Scored 93.52 in SuperCLUE, just shy of top models.
  • Accessible: Runs on consumer GPUs after 4-bit quantization, using only 15GB VRAM.
  • Open source: Available for developers to build upon.

As AI continues to evolve, moves like this could democratize access and fuel innovation from the ground up. Keep an eye on this space—it's moving fast.