Skip to main content

Zhipu's GLM-5.3: Same Size, 50% Smarter, and Closing in on Claude

On August 14, Zhipu AI officially released its latest large language model, GLM-5.3. While the base parameter count remains at over 74 billion—the same as its predecessor GLM-5.2—the company has managed to squeeze out a remarkable 50% performance improvement through advanced post-training techniques. This means the rumored jump to 1 trillion parameters didn't happen; that card might be saved for GLM-5.5. But the gains are real, and they're turning heads.

GLM-5.3 has posted significant score increases across several key benchmarks, and the numbers are hard to ignore. On Terminal-Bench 3.0, which tests a model's ability to handle complex tasks in real terminal environments, the score soared from 4.6 to 28.3. On DeepSWE v1.1, a benchmark focused on long-range software engineering and continuous code modification, it jumped from 46.2 to 66.9. And on Agents' Last Exam, which covers real-world professional scenarios requiring cross-tool collaboration and long-range tasks, it climbed from 23.8 to 28.5. In GDPval-AA v2, which spans 44 professions and evaluates high-value knowledge work, GLM-5.3 scored 1769 points, showcasing its professional task execution capabilities built on solid programming skills.

Image

What's particularly telling is Zhipu's own Z.ai Code Bench—a set of evaluations that place the model in a real local development environment, executing end-to-end tasks under different thinking modes. This setup simulates the experience developers have when using a Coding Agent as closely as possible. The results show that GLM-5.3 has found a sweet spot between effectiveness and token usage. In High mode, it achieved an accuracy rate of 31.4%, surpassing Claude Opus 4.8's maximum mode at 29.5%. Even more impressive, each task required only about 50,000 tokens on average, while Opus 4.8 needed roughly 120,000 tokens. That means GLM-5.3 can accomplish the same tasks with a much shorter execution path—a win for both speed and cost.

Image

On the product front, GLM-5.3 is already available on Zhipu's official programming tool ZCode and efficiency tool AutoClaw. It's also open to all users of the GLM Coding Plan and subscription users. Third-party coding platforms like TraeWork, TraeCode, Kousi, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode have also opened early access. The API is slated to launch soon, and the complete model weights will be open-sourced within two weeks, following necessary security reinforcement. Zhipu's strategy here is to limit the model's potential attack capability while preserving its defensive value. At 13:00 today, the quota for all GLM Coding Plan users was reset, and everyone can see their restored quotas in the backend usage statistics.

Key Points

  • Parameter Count Unchanged: GLM-5.3 keeps the same 740B parameters as GLM-5.2, but performance improved by 50% through post-training.
  • Benchmark Surge: Significant gains on Terminal-Bench 3.0, DeepSWE v1.1, Agents' Last Exam, and GDPval-AA v2.
  • Efficiency Win: On Z.ai Code Bench, GLM-5.3 outperforms Claude Opus 4.8 in accuracy while using far fewer tokens.
  • Availability: Now on ZCode, AutoClaw, and third-party platforms; API and open-source weights coming soon.