$ cat wiki/models/glm-5-3.md
GLM-5.3
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Z.ai (Z.ai / Zhipu AI) |
| Released | 2026-08-14 (GLM Coding Plan); open weights 2026-08-28 |
| Announced | 2026-08-14 |
| Context window | 1M (max_position_embeddings: 1048576, vendor config.json, 2026-08-28) |
| Pricing | GLM Coding Plan from $18/month; no per-token row published — the model card carries no price, and Z.ai's API pricing table still lists GLM-5.2 and has no GLM-5.3 entry |
| License | glm-5.3 — MIT-shaped grant with a revenue-gated security-review clause; not MIT |
| Availability | GLM Coding Plan subscribers; open weights on Hugging Face (zai-org/GLM-5.3 FP8, zai-org/GLM-5.3-BF16) |
Context window moved twice, and the second move is the one that settles it. | |
It was unknown until 2026-08-23, then filled from a third-party listing — | |
| Artificial Analysis at 1M (source) — | |
| because Z.ai had published no figure and inferring one from GLM-5.2's | |
shared base is the error this schema's unknown exists to prevent. **On 2026-08-28 | |
the vendor published the number itself**: config.json in the weights repository | |
records max_position_embeddings: 1048576 | |
| (source). The cell now | |
| cites a first-party artefact rather than a listing, and the two agree. |
Not taken from that file: a parameter count. config.json gives 78 layers,
hidden size 6144, 256 routed experts and 8 active per token, and no total. Adding
those up would be this wiki's arithmetic, not Z.ai's figure — so the 743B/744B
discrepancy in ## Conflicting Reports stands unresolved.
Not taken from the same row: a price. Artificial Analysis's Cost per Task USD
is their measured cost of one task, not a per-token price, as that snapshot states
in its own header. Pricing stays as published.
Release Date
2026-08-14, through the GLM Coding Plan (source).
Open weights: 2026-08-28. zai-org/GLM-5.3 was last modified
2026-08-28T15:22:14Z carrying 141 FP8 .safetensors shards, and
zai-org/GLM-5.3-BF16 at 13:46:13Z carrying 282. The repository itself had
been created 2026-08-25T06:42:50Z and stood empty of weights for three days. A
third-party quantisation, unsloth/GLM-5.3-GGUF, was created 2026-08-28T04:05:14Z
(source).
The window closed on the date it was announced for. Z.ai stated on 2026-08-14 that open weights would arrive "in stages after a safety evaluation, roughly two weeks" away (source); the Hugging Face placeholder named 2026-08-28; the weights landed 2026-08-28 UTC. This wiki recorded the window as open and unclosed at 09:00 KST on 2026-08-28, which was correct at the time and is the reason the closure is recorded here with its timestamp rather than as a date.
That is still a change of pattern, not a delay: GLM-5.2 went from subscriber access (2026-06-13) to open weights and standalone API (2026-06-16) in three days. This release separated them by fourteen and attached a condition — a safety evaluation — to the second half. See Open-Weights Policy Fight.
What the safety evaluation was about, stated at last. The model card volunteers a reason the 2026-08-14 material did not: "As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks" (source). That answers the question this page recorded as open on 2026-08-27 — why the gate bound GLM-5.3 and not GLM-5.3-Flash — though the card does not say what the evaluation consisted of, who ran it, or what it concluded.
The licence is glm-5.3, and it is not MIT. The model card's front matter records
license: other, license_name: glm-5.3, and the repository carries a licence file
in English and Chinese (source).
The grant is otherwise MIT-shaped — use, copy, modify, merge, publish, distribute, sublicense, sell, run, deploy, fine-tune, create derivative works, free of charge. One clause is added. A licensee operating a "Model as a Service" business whose aggregate revenue with its affiliates exceeds 10 billion US dollars over any consecutive 12 months must pass Z.AI's security review before any commercial use, with "the scope and method of the security review … reasonably determined by Z.AI". Model as a Service is defined as giving a third party inference or fine-tuning access "in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data", and explicitly excludes end-user products embedding the model in specific features or harnesses, and mere relaying of requests to models hosted by others.
Three things follow, and the page states them separately from the clause itself.
- The threshold selects hyperscalers and the largest model vendors and nobody else. A lab, a startup or a self-hoster is unaffected by clause 2 on its face.
- GLM-5.3-Flash shipped MIT on day one, eight days after GLM-5.3's announcement and two days before its weights. The siblings did not receive the same terms, and nothing read explains the split.
- This is the second condition attached to the same release. The first was timing (a safety evaluation); the second is who may serve it commercially. See Open-Weights Policy Fight, where "open weights" continues to name a range of positions rather than one.
Benchmarks
Vendor-stated. No harness is reported as published for any of these figures,
and none of the five benchmarks appears in any snapshot under sources/evals/
(source):
| Benchmark | GLM-5.2 | GLM-5.3 |
|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 |
| DeepSWE v1.1 | 46.2 | 66.9 |
| AutomationBench | 26.2 | 48.2 |
| Agents' Last Exam (CLI) | unknown | 28.5 |
| CyberGym | unknown | 84.5% |
| The DeepSWE baseline and the AutomationBench row were added 2026-08-21 from a | ||
| second, independent report of Z.ai's figures | ||
| (source). That | ||
| report gives Terminal-Bench 3.0 4.6 → 28.3 and DeepSWE 66.9 identically to | ||
| the 2026-08-14 capture, which is why the two figures it adds are accepted rather | ||
| than filed as a competing account. They remain **vendor-stated with no harness | ||
| published**, like the rest of this table. |
Z.ai's internal evaluations report a ~50% improvement in coding capability over GLM-5.2, and coverage ranks GLM-5.3 first among open-source models on Terminal-Bench 3.0 and Agents' Last Exam (source).
Two figures are worth reading twice:
- Terminal-Bench 3.0: 4.6 → 28.3. A six-fold move, and also under 29% of the tasks. The release material carries the first reading only. Note the benchmark version: Qwen 3.8 27B and DeepSeek V4-Pro-0813 report Terminal-Bench 2.1, a different suite — the numbers are not comparable across the version boundary. See Eval Harness Configuration.
- Agents' Last Exam CLI: 28.5 against GPT-5.6 Sol's 28.6. A gap of one tenth of a point, run by the party that benefits from it being small, with no published harness on either side. Recorded because Z.ai published it, not because it is a measured tie.
On CyberGym, GLM-5.3's 84.5% is reported to slightly surpass Mythos 5 and GPT-5.6 Sol (and Terra, Luna), while gaps remain against closed frontier models on deep exploitation tasks such as ExploitBench (source). The nearest figure this wiki already holds is DeepSeek V4-Pro-0813's vendor-stated 83.3 on the same benchmark — also unharnessed, also first-party.
2026-08-23 — the first independent measurement, and it moves in the direction the vendor claimed. Artificial Analysis's weekly leaderboard lists GLM-5.3 (max) at Artificial Analysis Intelligence Index 60, Cost per Task USD $0.68, Median Tokens/s 102, Latency First Chunk 4.35 s, Total Response 28.84 s, context 1M. It was absent from the 2026-08-16 capture and is new this week (source; the prior week is here).
Read against the same table's GLM-5.2 (max) at 53 — a figure that is $0.44 this week against $0.32 last — the index moves 53 → 60 between two models Z.ai states share an unchanged base. That is the first number in this wiki supporting Post-Training Scaling's central claim from a party with nothing to sell, and it is a composite from a third-party harness rather than any of the five vendor benchmarks above.
Three limits on how far that goes. The index is Artificial Analysis's own
composite of their suite, not an accuracy and not comparable to the vendor table;
(max) is one reasoning setting, so this compares GLM-5.3's top configuration with
GLM-5.2's; and a 7-point move on a composite is not the "~50% improvement in coding
capability" Z.ai's internal evaluations report — it is a smaller, broader, and
independently measured claim. In the same table GLM-5.3 sits level with
Kimi K3 (max) 60 and Grok 4.6 (xhigh) 60, one point
below GPT-5.6 Sol (and Terra, Luna) (max) 61 and Claude Opus 5 (high) 61,
and three below Claude Opus 5 (max/xhigh) 63.
2026-08-28 — the full vendor table, and the harness this page said was missing. The model card publishes 16 benchmark rows against 7 comparison models. Reproduced in full in the capture; the rows that change what this page already held (source):
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4 Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | – | – | 21.1 | 33.7 | 34.6 |
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym (2h / 6h) | 105 / 130 | 29 / 39 | 36 / 70 | – | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | – | 28.8 | 40.0 | 78.0 | 76.5 |
| AutomationBench (v1.0.6) | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents' Last Exam (ALE-CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
| The four figures this page has carried since 2026-08-14 are confirmed unchanged | ||||||||
| — Terminal-Bench 3.0 4.6 → 28.3, DeepSWE 66.9, AutomationBench 48.2, ALE-CLI 28.5 — | ||||||||
| now from the vendor's own model card rather than through coverage. **CyberGym is | ||||||||
published as 84.5**; the table above carries the card's value, and the earlier | ||||||||
| section on this page writes it as a percentage. |
Every row now names its harness, and that is the change. This page said on 2026-08-14 that "no harness is reported as published for any of these figures". The card's footnotes publish one for almost all of them, in the detail Eval Harness Configuration has spent a month arguing for: Claude Code 2.1.207 for Terminal-Bench 2.1 and 3.0, CyberGym, ExploitGym, ExploitBench, PostTrainBench, SWE-Marathon and ALE-CLI; mini-swe-agent for DeepSWE; GPT-5.6-luna (medium) as judge on HLE; Proximal for FrontierSWE; Artificial Analysis for GDPval-AA v2 — with sampling parameters, context lengths, timeouts, turn caps, rollout counts and container policy stated per row.
Three of those footnotes deserve to be read rather than summarised:
- Terminal-Bench 3.0 is avg@3, reasoning effort max, 400K context, 600 agent turns, 10-hour timeout, Tool Search disabled, official separate verifier, each rollout in a container from the task's official image.
- ExploitGym's 2h / 6h budgets are not wall-clock. They are "the API inference time rescaled by per-model tokens per second rate", with TPS taken from Artificial Analysis — 115 for GLM-5.3, 40 for Kimi K3, 47 for Qwen3.8 Max. A model measured as faster is therefore allowed more work inside the same nominal budget, and the three models compared this way are the only three Z.ai ran itself.
- Two benchmarks had their anti-cheat checks removed and replaced. On
PostTrainBench the pattern-matching checks "produced false positives when a
local vLLM endpoint was accessed through the OpenAI SDK"; on SWE-Marathon's
strip-clonethe import detection "could reject valid implementations". Both were replaced with LLM-based inspection. Z.ai discloses this, which is the right behaviour, and it also means two rows were scored under a modified protocol.
Where GLM-5.3 leads and where it does not. It is first in the table on CyberGym (84.5), AutomationBench (48.2) and GDPval-AA v2 (1769), and first among the open-weights entries on Terminal Bench 3.0. It trails GPT-5.6 Sol on Terminal Bench 2.1, Terminal Bench 3.0, DeepSWE, HLE and ALE-CLI, and trails Fable 5 by a wide margin on the exploitation rows — ExploitBench 54.4 against 78.0, and ExploitGym 105 / 130 against 181 / 247. The "state of the art on CyberGym" claim is about vulnerability discovery and the card is careful to say so; the exploitation rows in the same table go the other way. That distinction survives into AI-Enabled Cyberattacks and should not be flattened.
2026-09-29 — the first independent cyber measurement, and it reverses the reading above. Anthropic's Frontier Red Team published GLM-5.3 and the Spread of Advanced Cyber Capabilities — read first-party on this run (source). Authors: Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, Tripp Gallagher.
On ExploitBench, 410 attempts per model:
| Model | End-to-end exploits | Rate |
|---|---|---|
| Claude Mythos Preview | 56 / 410 | 14% |
| GLM-5.3 | 50 / 410 | 12% |
| Claude Opus 4.6 | at or near 0 | ~0% |
| GLM-5.2 | at or near 0 | ~0% |
| Kimi K3 | at or near 0 | ~0% |
| DeepSeek V4.1-Flash | at or near 0 | ~0% |
| And on binary exploitation — full control-flow hijacks — GLM-5.3 4% | ||
| against Claude Mythos Preview's 6%, with every other tested model at 0%. |
This page has said since 2026-08-28 that GLM-5.3 "trails Fable 5 by a wide
margin on the exploitation rows — ExploitBench 54.4 against 78.0". The independent
measurement says it is within two points of Anthropic's most capable cyber model.
Both cannot be a description of the same quantity, and the disagreement is
disclosed under ## Conflicting Reports rather than resolved by preferring the
newer number.
The unaided discovery result is the part with no vendor equivalent. Verbatim:
Over the course of a day (and with limited human attention), GLM-5.3 found several previously unknown vulnerabilities in the browser's JavaScript engine, and chained them together into a working exploit: a webpage that, when visited, reads arbitrary files from the visitor's computer.
That is not a benchmark score. No row in Z.ai's 16-benchmark table measures novel vulnerability discovery against a real target, and AI-Enabled Cyberattacks carries the consequence.
Safeguards, which this page had no figure for at all. Engagement with malicious cyber-attack orders in a simulated environment:
| Condition | GLM-5.3 | Claude |
|---|---|---|
| Bare order | 0% | 0% |
| False cover story | 64% | 0% |
| Prefilled reasoning | 92% | 0% |
| Abliterated | 100% | n/a |
| Abliteration took a team that had never attempted it before about **2,200 GPU | ||
| hours, roughly $4,400**, dropping refusal from above 90% to **JailbreakBench 3% · | ||
| HarmBench 2% · StrongREJECT 12%**. Anthropic states the asymmetry plainly: | ||
| *"Claude models are released with cyber safeguards, and versions with reduced | ||
| safeguards are limited to vetted users. Anyone can download and use GLM-5.3."* | ||
| See Open-Weights Policy Fight. |
A second independent assessment is quoted in the post and not measured by it. NIST's Center for AI Standards and Innovation (CAISI) found GLM-5.3 "the most cyber-capable open-weight model released to date", lagging the US frontier by approximately four months on an aggregate of cyber benchmarks. That is a third-party figure relayed by a competitor and was not independently read on this run.
Use Cases
Coding and long-horizon agentic work, delivered through a subscription coding product rather than a token API — the same channel GLM-5.2 launched into (source). Since 2026-08-28 it is also self-hostable: the card lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers and Unsloth, and — for Ascend NPU — vLLM-Ascend, xLLM and SGLang (source). The Ascend path is worth noting beside Z.ai's disclosure of over 100,000 domestic chip units serving inference.
One inference default that changes what a reproduction means. reasoning_effort
accepts low, high and max, and defaults to max if not passed or set to
anything else; the card instructs "For benchmark and leaderboard reproduction, keep
the default max." Every figure in the table above is therefore a max-effort
figure, which is the same caveat this page already applies to Artificial Analysis's
(max) listing.
The cyber-security framing is Z.ai's own and is the sharpest positioning in the release: it leads with CyberGym rather than with a general coding benchmark. See AI-Enabled Cyberattacks.
2026-08-18 — the access model attached to that framing. GLM-5.3 ships alongside Z.ai's "Shield of Open Source" programme: free security audits, automated code-auditing tools via the ZCode platform, and free model usage quotas for the open-source community, with a restricted tier — "Cybersecurity Trusted Access" — under which the model's most sensitive offensive capabilities are reserved exclusively for verified users (source).
Nothing read reconciles that tier with the open weights, and the weights are now out. This page recorded the tension while the weights were pending; publication resolves it in the direction that makes the gate least effective. A verified-access control on "the model's most sensitive offensive capabilities" does not obviously survive publication of the weights it gates — anyone may now run the model without passing through Z.ai at all. The model card, the licence and the repository make no mention of the Shield of Open Source programme or of Cybersecurity Trusted Access (source), so nothing read states whether the two plans are the same plan, whether the gate applies to the hosted API only, or what verification consists of. The question is now sharper than when it was first recorded, and no less open.
Compared To
- GLM-5.2 — the same base model, unchanged; every difference between these two pages is post-training
- GPT-5.6 Sol (and Terra, Luna) — the named comparison on Agents' Last Exam CLI and CyberGym
- Claude Fable 5 — coverage describes GLM-5.3's coding and agent capability as "approaching" it; no shared benchmark figure was published
- DeepSeek V4-Pro-0813 — the other Chinese-lab agentic release of the same week, and the only other CyberGym figure this wiki holds
- Qwen 3.8 27B — released the same day, and the opposite trade: weights first, subscription never
Sources
- GLM-5.3 and the Spread of Advanced Cyber Capabilities — Anthropic Frontier Red Team, 2026-09-29 → (snapshot) — read first-party; the only independent cyber measurement this page holds, and the source of the safeguard, abliteration and CAISI figures
- GLM-5.3 model card, licence and
config.json— Hugging Face,zai-org/GLM-5.3, 2026-08-28 → (snapshot) — the only first-party artefact this page cites; the source of the licence, the context window, the full benchmark table and every harness footnote - Zhipu AI's answer to Project Glasswing marks shift for Chinese cyber safety — SCMP → (snapshot)
- Zhipu AI releases GLM-5.3, claims it's the strongest open-weights coding model — the-decoder → (snapshot)
- Zhipu releases GLM-5.3 through its coding service, with weights still two weeks away — MLQ News
- GLM-5.3: Post-Training Alone Rebuilt the Coding Ladder — Digital Applied
- GLM-5.3 Launch: Benchmarks, Pricing & Access — explainx.ai
- GLM-5.3 Pricing — emergent.sh
- GLM-5.3: How Chinese labs keep stride with the frontier — Interconnects
- "Death of Params": Jie Tang on GLM-5.3 and a post-training scaling law — Latent Space AINews → (snapshot) — the source of the DeepSWE baseline and the AutomationBench row; see Post-Training Scaling
Conflicting Reports
-
"ExploitBench" names two measurements that disagree by roughly 5x, and this page now carries both. Z.ai's model card reports GLM-5.3 54.4, GLM-5.2 24.4, Kimi K3 32.2, Opus 4.8 40.0, Fable 5 78.0, GPT-5.6 Sol 76.5 (source). Anthropic's Frontier Red Team reports GLM-5.3 12% (50/410), Claude Mythos Preview 14% (56/410), and GLM-5.2, Kimi K3, Opus 4.6 and DeepSeek V4.1-Flash at or near 0% (source).
The rank order is preserved and the magnitudes are not. Both put an Anthropic frontier model narrowly above GLM-5.3. But Z.ai's card has GLM-5.2 at 24.4 where Anthropic has it at ~0%, and Opus 4.8 at 40.0 where Opus 4.6 is ~0% — differences no rounding explains.
What is established: the two runs used different harnesses. Z.ai's footnote names Claude Code 2.1.207 for its ExploitBench row; Anthropic reports a denominator of 410 attempts and names no harness. What is not established: whether the task sets are the same benchmark at all, whether
54.4is a percentage (this page has written it both ways), and which is the better estimate of the capability.Neither figure is withdrawn and neither is preferred.
CLAUDE.mdgives the vendor announcement precedence for the vendor's own model, and Anthropic is a competitor reporting on GLM-5.3 — but Anthropic is also the only party here who measured a model it does not sell. See Eval Harness Configuration: a benchmark name without a harness is not a measurement, and this is the sharpest case this wiki holds of the same name carrying two. -
The published context window disagrees with the catalogue, and the page's figure stands.
spec-checkrun 90 (2026-09-12) reports context 1,000,000 vs catalogue 1,310,720 (source).The Spec row is not changed: the cell above is faithful to the vendor material this page cites, and
CLAUDE.mdgives the vendor announcement precedence over the catalogue. 1,310,720 is 1.25 × 1,048,576, which is a round figure in the catalogue's units rather than an arbitrary one — recorded as an observation and not adopted as an explanation.What is not established: which figure is correct, and whether the catalogue is describing the same deployment.
openrouter.aiis blocked from the daily run's sandbox, so this cannot be settled here.Reported by a red Action on every run since 2026-08-29 — 25 consecutive — and unrecorded on this page until 2026-09-13.
-
Base parameter count. Coverage of this release describes the shared base as 743B (source); this wiki records 744B total for GLM-5.2 from its own June capture (source). Third-party on both sides, one billion apart, and Z.ai states the base is unchanged — so at most one of the two roundings is right. Neither page was edited on the strength of the other.
Referenced by
Sources
- sources/blogs/anthropic-2026-09-29-glm-5-3-cyber-capabilities.md
- sources/evals/spec-check-2026-09-12.md
- sources/blogs/zai-2026-08-28-glm-5-3-open-weights.md
- sources/evals/artificial-analysis-2026-08-23.md
- sources/blogs/zai-2026-08-18-shield-of-open-source.md
- sources/blogs/zai-2026-08-20-post-training-scaling-law.md
- sources/blogs/zai-2026-08-14-glm-5-3.md