Back to News
Daily analysis

Zhipu's GLM-5.3 Weights Land Exactly on Schedule

August 28, 2026 · 6 min read

This report comes from Astro's continuous monitoring of the AI market. The period's sources are listed below and every claim links to its origin.


AI Radar · Daily · August 28, 2026

Zhipu's Z.ai published the full weights for GLM-5.3 on Hugging Face today, two weeks to the day after it launched the model API-only on August 14 and said it would withhold the weights pending a security review. The model page is not a placeholder: the repository carries 141 safetensors shards totaling roughly 755GB, matching the model's declared 753 billion parameters, each shard backed by a real LFS hash. The commit log confirms the timing, too: the repo was scaffolded on August 25, the weight-bearing "Initial commit 0828" landed August 27 at 17:16 UTC (past midnight in Beijing, hence the "0828" in the commit title), and two follow-up commits pushed benchmark documentation through 15:22 UTC today.

That is a promise kept on the day it was made, which is the story: labs that delay weights for safety reasons routinely let the date slip. This one didn't. The release is also the top post on Hacker News this window, at 441 points and 158 comments roughly five hours after it went up, per the HN thread.

What actually changed

GLM-5.3 uses the same base model as GLM-5.2; Zhipu says every gain comes from extended post-training, not a new pretrain. The model card's own benchmark table shows the shift is real and concentrated in coding and agentic tasks: Terminal-Bench 3.0 jumps from 4.6 to 28.3, DeepSWE (v1.1) from 46.2 to 66.9, and FrontierSWE from 67.5 to 78.1. Zhipu calls it "the most capable open-weights model for coding," citing a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, a number nobody outside the company can check.

The other jump, the one that triggered the delay, is on offensive-security benchmarks: CyberGym rises from 77.2 to 84.5, edging out Claude Opus 4.8 (78.1) and GPT-5.6 Sol (83.6) by Zhipu's own numbers, and ExploitBench more than doubles, from 24.4 to 54.4. Zhipu's announcement, covered by The Decoder, says the model "began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains," and that Zhipu used it with Chinese security teams to surface 2,436 vulnerabilities across 269 open-source projects, 1,097 rated critical or high, published to a public registry. Worth noting: on the same table, GLM-5.3 still trails Claude Opus 4.8 and the two unreleased comparison models Zhipu lists (Fable 5 and GPT-5.6 Sol) on the harder exploit benchmarks, ExploitBench and ExploitGym, by a wide margin. The CyberGym lead that reportedly justified a two-week hold is narrow; the tasks it doesn't lead on are the ones further up the exploitation chain.

The stated reason, and what the license actually does

Zhipu framed the delay as caution about a capability that "developed faster than we expected" during training. What's notable, hours after the weights landed, is that this rationale does not show up as any cybersecurity-specific restriction in the license. The GLM-5.3 license, published alongside the weights, is otherwise a permissive MIT-style grant: use, copy, modify, fine-tune, redistribute, sell. Its only conditional clause requires a Z.ai security review before commercial use, but only for a licensee whose "Model as a Service" business (and affiliates) clears $10 billion in revenue over any 12 months, a threshold aimed at large cloud resellers, not at vulnerability-discovery or exploit use specifically. Nothing in the license or model card restricts, gates, or requires disclosure for cybersecurity applications of the weights themselves. The two-week wait bought Zhipu a risk review it hasn't described in any detail; it did not buy a usage restriction.

What's still unverified

The release is hours old, and it shows. Every benchmark on the model card, coding and cyber alike, is Zhipu's own number on its own harness (Claude Code 2.1.207, specific reasoning-effort and timeout settings); no independent lab has reproduced Terminal-Bench 3.0 or CyberGym results yet. A VentureBeat report that GLM-5.3 found a "potentially serious vulnerability" in Cursor around the August 14 launch remains unconfirmed; VentureBeat said it had asked Cursor for comment and was awaiting a response. And on the Hugging Face repo itself, a discussion opened this afternoon flags that GLM-5.3 refuses more requests than GLM-5.2, with the original poster noting the refusal language "reads exactly like Claude's," a single anecdote, not a pattern, but a reminder that whatever guardrails did or didn't change between versions haven't been documented by Zhipu either. On Hacker News, the skepticism runs similar lines: commenters question whether the CyberGym lead survives outside a curated benchmark, and whether the weights being public makes fine-tuning away any residual restraint trivial.

Sources

  1. GLM-5.3 - Hugging Face — 2026-08-28
  2. GLM-5.3 is now open-weight - Hacker News — 2026-08-28
  3. Zhipu AI releases GLM-5.3, claims it's the strongest open-weights coding model - The Decoder — 2026-08-28
  4. GLM-5.3 is here with advanced cyber capabilities, and reportedly already found a 'serious vulnerability' in Cursor - VentureBeat — 2026-08-14

The week in AI, in your inbox

AI market intelligence, weekly. The reports published here, delivered to your email.