Z.ai releases GLM-5.3

17-08-2026

GLM-5.3 is Z.ai's post-training only upgrade to GLM-5.2, released 14 August 2026, with a 1,000,000 token context window, 28.3 percent on Terminal-Bench 3.0 and 84.5 percent on CyberGym on Z.ai's own evaluations, and open weights promised roughly two weeks after launch.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

z.ai releases glm 5.3
Sign up for our Newsletter

Published: 17 August 2026

Z.ai released GLM-5.3 on 14 August 2026, a coding and cybersecurity model that reuses the GLM-5.2 base model unchanged and takes every reported gain from post-training alone. GLM-5.3 has a 1,000,000 token context window, a 128,000 token output limit, and on Z.ai’s own evaluations lifts Terminal-Bench 3.0 from 4.6 percent to 28.3 percent. Access at launch runs through the GLM Coding Plan and the ZCode agent rather than a public API, and Z.ai says the open weights will follow roughly two weeks after launch.

What is new in GLM-5.3?

The headline of the GLM-5.3 release is what Z.ai did not do. In its own words, the model uses the same base model as GLM-5.2 and every gain comes from post-training. Z.ai frames the whole release around that point, writing that scaling post-training is all it did for GLM-5.3. For anyone tracking whether the frontier still depends on ever larger pre-training runs, a jump of this size on a frozen base is the interesting part of the announcement.

The target is long horizon software engineering and cyber defence work. Z.ai titles the announcement around frontier coding with emergent cyber capabilities, and the benchmark table is weighted heavily towards agentic coding, exploit discovery and tool use rather than general knowledge. The model is text in and text out, so unlike Qwen3.8-27B or dots3-note Preview, both released the same day, it has no vision or audio path.

One breaking change matters for anyone migrating. GLM-5.3 exposes three effort levels, low, high and max, with max as the default, and it no longer supports disabling thinking at all. Applications that pass a thinking type of disabled will see the request fail rather than fall back to a non-thinking mode.

GLM-5.3 benchmarks and what Z.ai actually measured

On Terminal-Bench 3.0 GLM-5.3 scores 28.3 percent against 4.6 percent for GLM-5.2. Z.ai publishes the harness in full: Claude Code 2.1.207 at maximum reasoning effort with a 400K context and a 128K output cap, averaged over three rollouts per task, each rollout in an isolated container capped at 600 agent turns and a 10 hour timeout, with tool search disabled and scoring done by the task’s own official verifier. That level of disclosure is unusual and it is worth stating plainly, because it makes the number checkable in a way most vendor benchmarks are not.

The same table shows where GLM-5.3 still sits behind.

  • On Terminal-Bench 3.0 it trails GPT-5.6 Sol at 34.6 percent and Claude Fable 5 at 33.7 percent, and leads Claude Opus 4.8 at 21.1 percent and Kimi K3 at 17.4 percent.
  • On DeepSWE v1.1 it reaches 66.9 percent against 46.2 percent for GLM-5.2, but again behind GPT-5.6 Sol at 72.7 percent and Kimi K3 at 67.5 percent, and ahead of DeepSeek V4 Pro-0813 at 62.7 percent.
  • On the older Terminal Bench 2.1 the field is compressed: GLM-5.3 at 88.2 percent, Kimi K3 at 88.3 percent and GPT-5.6 Sol at 88.8 percent are within a point of each other.

The strongest single result is cybersecurity. On CyberGym GLM-5.3 scores 84.5 percent against 77.2 percent for GLM-5.2, ahead of Claude Fable 5 at 83.8 percent and GPT-5.6 Sol at 83.6 percent, which makes it the leading published score on that benchmark. Z.ai gives the method: Claude Code 2.1.207 at maximum reasoning effort with no web tools, single run pass@1 across 1,507 tasks, agent placed inside the task container with Git information removed and a domain whitelist applied. Three results in the table are explicitly outsourced rather than run by Z.ai: GDPval-AA v2, where it scores 1769, was evaluated by Artificial Analysis; FrontierSWE, at 78.1 percent, by Proximal; and Toolathlon Verified, at 73.0 percent, through the benchmark’s official evaluation service.

The claim to treat carefully is the headline coding improvement. Z.ai’s roughly 50 percent jump comes from Z.ai Code Bench, an in house benchmark the company describes as private specifically to reduce contamination risk. It publishes no task count, no sample size and no table, only figures read off a chart: 31.4 percent at high effort at around 50,000 output tokens per task against 29.5 percent for Claude Opus 4.8 at 120,000, and 34.5 percent at max effort against 23.4 percent for GLM-5.2. A private benchmark is a defensible anti-contamination measure, but a headline number with no methodology behind it is a vendor claim rather than an established result. Two smaller inconsistencies are worth knowing: Z.ai’s own text calls the same competitor Fable 5 in the benchmark table and Mythos 5 in the prose, and its vulnerability disclosure figures describe 1,097 findings as medium to high in the prose while its own widget labels the same 1,097 as critical and high.

GLM-5.3 pricing, access and when the open weights arrive

There is no public API for GLM-5.3 yet. Z.ai’s developer documentation states that the GLM-5.3 API is coming soon, and the model does not appear on the company’s pricing page, which still lists GLM-5.2 at 1.40 dollars per million input tokens and 4.40 dollars per million output tokens. Anyone reporting a per token rate for GLM-5.3 is reporting a rate Z.ai has not published.

What does exist is the subscription route. The GLM Coding Plan starts at 18 dollars per month and now runs on a points based quota, with input, cached input and output tokens counted separately. Calls made outside peak hours consume half the standard points, where peak is defined as 14:00 to 18:00 in UTC+8 from Monday to Friday, so all weekend usage bills at the off peak rate. Z.ai says all plans support GLM-5.3 and that requests for GLM-5.2 and GLM-5.1 are automatically routed to GLM-5.3. The model also runs in Z.ai’s own ZCode agent, in Claude Code and in OpenCode.

On open weights, Z.ai states twice in the announcement that it will release them two weeks after launch, once safety evaluation and hardening are complete, which points to around 28 August 2026. As of 17 August nothing had been published and the announcement page carries only a placeholder where the repository link would sit. Z.ai has also not stated a licence for the coming GLM-5.3 weights. GLM-5.2 shipped under MIT, but that is precedent rather than a commitment, and it should not be reported as the licence for GLM-5.3.

For more details please visit the official Z.AI announcement of GLM-5.3.

Add DataNorth AI to your Google favorites