Skip to main content

Zhipu's GLM-5.3: Same Base, Better Coding, Open-Source Crown Shifts

Zhipu AI officially dropped GLM-5.3 today, and the headline isn't a bigger base model—it's a smarter one. Unlike previous upgrades that swapped out the foundation, this version sticks with the same base as GLM-5.2. All the gains come from cranking up the post-training phase: a tenfold boost in long-range task environments, more diverse scenarios, and extended training time. The result? A model that's significantly sharper on the same hardware.

According to internal evaluations, the programming experience with GLM-5.3 feels about 50% better than with GLM-5.2. That's a massive leap, and it's enough to crown it the strongest open-source model for coding right now. But the proof isn't just in-house—public benchmarks tell the same story. Terminal-Bench 3.0 scores jumped from 4.6 to 28.3, DeepSWE v1.1 climbed from 46.2 to 66.9, and Agents' Last Exam went from 23.8 to 28.5. Even GDPval-AA v2 hit 1769 points. These numbers put GLM-5.3 in the same league as Claude Fable 5, and it's already taken first place in the open-source category on both Terminal Bench 3.0 and Agents' Last Exam (CLI).

Image

Post-Training Scaling: The Base Model's Untapped Potential

Zhipu's team is adamant that these improvements come from post-training, not from tweaking the architecture. They leveraged their IndexShare, SAO, and the evolving next-generation Slime framework to run reinforcement learning efficiently on the same base as GLM-5.2. And they're candid about it: "We may be far from fully exploring the intelligence upper limit of this base." That's a refreshing admission, and it sends a clear signal to the industry—competition isn't always about adding more parameters. Sometimes, it's about squeezing every drop of capability out of what you already have.

Image

Security is another area where GLM-5.3 holds its own. In tasks like white-box code review and vulnerability detection, it performs just as well as Mythos 5, which bodes well for network defense applications. Zhipu plans to release the full model weights two weeks after launch, but only after completing security assessments and hardening the model to limit potential attack capabilities while preserving its defensive value.

The new model is already available on Zhipu's official programming tools—ZCode, AutoClaw, and the GLM Coding Plan—and early access is rolling out on platforms like Trae, Kouzi, WorkBuddy/CodeBuddy, and Qoder. As open-source models close the gap with closed-source flagships in the programming arena, the second half of the domestic large-model race is shifting. It's no longer just about who talks better; it's about who works better.

Key Points:

  • GLM-5.3 keeps the same base model as GLM-5.2, with all improvements from post-training scaling.
  • Programming experience improved ~50% over GLM-5.2, making it the top open-source coding model.
  • Benchmarks: Terminal-Bench 3.0 up to 28.3, DeepSWE v1.1 to 66.9, Agents' Last Exam to 28.5, GDPval-AA v2 at 1769.
  • Security performance matches Mythos 5 in code review and vulnerability detection.
  • Full weights release in two weeks after security checks; available now on ZCode, AutoClaw, and other tools.