Zhipu's GLM-5.3-Flash: A Game-Changer in AI Pricing
In a bold move that could reshape the AI landscape, Zhipu AI has officially released and open-sourced its first native multimodal model of the GLM-5 series: GLM-5.3-Flash (320B-A18B). This isn't just another model drop—it's a statement about accessibility and performance.
The model's capabilities are nothing short of impressive. In the Artificial Analysis Intelligence Index, it scored 57 points, tying with Claude Opus4.8. Its programming skills are on par with larger models, and its overall performance even surpasses its bigger sibling, GLM-5.2. Before the official launch, it was tested anonymously under the alias "Ox-Alpha" and quickly climbed to the top of the usage charts on OpenCode and OpenRouter. That's a strong signal of real-world utility.
But here's the kicker: the price. GLM-5.3-Flash is priced at just one-tenth of the same series' GLM-5.3, and with a limited-time discount, it drops to one-twentieth. That's equivalent to one-fortieth the price of Opus4.8. For developers and businesses that have been priced out of top-tier AI, this is a game-changer. It's not just about affordability—it's about making cutting-edge technology accessible to everyone.

Under the hood, the model employs a hybrid architecture combining sparse attention and linear attention—a first for open-source models. This, along with manifold constraint super connection technology, reduces attention computation by 3.01 times and KV cache by 4.44 times. What does that mean for you? Better long-context handling without skyrocketing inference costs. Plus, it's natively multimodal, meaning it can handle visual tasks like front-end development, 3D modeling, and document creation using visual feedback. In professional settings like financial reports and legal documents, it's already showing strong performance.

One of the most intriguing aspects is that the model's full traffic is supported by domestic chip clusters. The team developed a self-researched inference engine with an EPD separation architecture and multiple memory optimization solutions. The result? A threefold improvement in end-to-end inference performance, with hardware utilization efficiency comparable to mainstream NVIDIA GPUs. This proves that domestic computing power can handle large-scale cutting-edge model inference—a significant milestone for tech self-reliance.
GLM-5.3-Flash is now fully open-sourced, with API interfaces and online experience channels available. It's also been integrated into intelligent Agent platforms like ZCode and AutoClaw. To sweeten the deal, Zhipu is offering daily experience cards for ten thousand uses, inviting global developers to dive in and explore.
Online Experience
- Z.ai: https://chat.z.ai
- Zhipu Qingyan App: https://chatglm.cn
Key Points:
- GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index, matching Claude Opus4.8.
- Priced at 1/40th of Opus4.8, making top-tier AI affordable.
- Hybrid attention architecture reduces computation and cache, balancing performance and cost.
- Native multimodal capabilities for visual tasks like front-end development and 3D modeling.
- Runs on domestic chips with a self-developed inference engine, achieving 3x performance boost.
- Fully open-sourced with API access and integration into Agent platforms.