Skip to main content

Zhipu's GLM-5.3: Same Size, 50% Smarter, and Closing In on Claude

On August 14, Zhipu AI officially unveiled GLM-5.3, the latest addition to its large model lineup. The base parameter count remains at over 74 billion—same as GLM-5.2—so those rumors about a jump to 1 trillion parameters? Not this time. That card might be saved for GLM-5.5. But don't let the unchanged size fool you. Through clever post-training techniques, Zhipu managed to squeeze out a 50% performance boost over the previous generation, making GLM-5.3 the top performer among current open-source models on several major benchmarks.

One of the most striking improvements is in real-world task execution. On Terminal-Bench 3.0, which tests a model's ability to handle complex tasks in actual terminal environments, the score skyrocketed from 4.6 to 28.3. That's a massive jump. On DeepSWE v1.1, which focuses on long-range software engineering and continuous code modification, it climbed from 46.2 to 66.9. And on Agents' Last Exam, a benchmark covering various professional scenarios that require cross-tool collaboration and long-range tasks, it went from 23.8 to 28.5. In GDPval-AA v2, which spans 44 professions and evaluates high-value knowledge work, GLM-5.3 scored 1769 points, showcasing its ability to handle professional tasks that rely on programming skills. Overall, the model has made significant strides in complex software engineering, terminal operations, and broader agent tasks.

Image

But here's where it gets really interesting. Zhipu's own Z.ai Code Bench puts the model in a simulated local development environment, running end-to-end tasks under different thinking modes to mimic what developers experience when using a Coding Agent. The results? GLM-5.3 strikes a better balance between effectiveness and token usage. In High mode, it hit an accuracy rate of 31.4%, surpassing Claude Opus 4.8's maximum mode at 29.5%. And it does this while averaging only about 50,000 tokens per task, compared to Opus 4.8's 120,000. That means GLM-5.3 can get the job done with a much shorter execution path—efficiency matters.

Image

So, when can you try it? GLM-5.3 is already available on Zhipu's official programming tool ZCode and efficiency tool AutoClaw. It's also open to all users of the GLM Coding Plan and subscription users. Third-party coding platforms like TraeWork, TraeCode, Kousi, WorkBuddy, CodeBuddy, Qoder, QwenWork, CatPaw, JoyCode, and OpenCode have also opened early access. The API is coming soon, and the complete model weights will be open-sourced within two weeks, after necessary security reinforcement. Zhipu's approach here is to limit the model's potential attack capability while keeping its defensive value intact. And for those already on the GLM Coding Plan, quotas were reset at 13:00 today, so you should see your usage stats restored.

This release is a clear signal that Zhipu is serious about competing at the highest level. By focusing on post-training improvements rather than just scaling up parameters, they've shown that bigger isn't always better—smarter training can make all the difference. For developers and AI enthusiasts, GLM-5.3 is definitely worth a look.

Key Points

  • Parameter Count: Stays at 74 billion, no trillion-parameter upgrade this time.
  • Performance Leap: 50% improvement over GLM-5.2 through post-training techniques.
  • Benchmark Wins: Top scores on Terminal-Bench 3.0, DeepSWE v1.1, and Agents' Last Exam.
  • Efficiency: Achieves higher accuracy with fewer tokens compared to Claude Opus 4.8.
  • Availability: Now on ZCode, AutoClaw, and third-party platforms; API and open-source weights coming soon.