$ cat wiki/entities/zai.md
Z.ai
Latest
- 2026-09-29
A competitor measured GLM-5.3's cyber capability independently, and reported it two points behind Anthropic's most capable cyber model — with the saf…
- 2026-09-28
Z.ai says it used its own model to ship its own model, and Jack Clark calls that an outer RSI loop.
- 2026-09-24
**Someone published an eight-stage post-training recipe on a Z.ai base
Overview
Z.ai is the international brand of Zhipu AI (智谱AI), a Beijing-based AI company founded in 2019, spun out of Tsinghua University. Z.ai operates the GLM (General Language Model) series, with GLM-5.2 (June 2026) representing the company's first frontier-class open-weight model. Positioned as "China's open-source frontier competitor" — openly publishing weights under MIT license while running a commercial API.
Parent: Zhipu AI — partially state-backed (backed by Chinese national funds and Alibaba).
Key People
- Zhipeng Jie — CEO, Zhipu AI / Z.ai
- Research team originates from the KEG (Knowledge Engineering Group) at Tsinghua University
Models & Products
- GLM-5.3 — released 2026-08-14 to GLM Coding Plan subscribers; the same 744B base post-trained further, weights withheld for ~2 weeks pending a safety evaluation
- GLM-5.2 — released 2026-06-16, 744B MoE (40B active), MIT license, 1M context, Code Arena #2; ****
- GLM Coding Plan — subscription product giving access to latest GLM models
- Z.ai API — standalone API for GLM models (routes through China-based infrastructure)
- NVIDIA NIM hosted version — available for GLM-5.2 on NVIDIA's inference microservices platform
Recent Activity
-
2026-09-29: A competitor measured GLM-5.3's cyber capability independently, and reported it two points behind Anthropic's most capable cyber model — with the safeguards removable for $4,400. Anthropic's Frontier Red Team published GLM-5.3 and the Spread of Advanced Cyber Capabilities, read first-party on this run (source). ExploitBench: GLM-5.3 50/410 (12%) against Claude Mythos Preview's 56/410 (14%); binary exploitation 4% against 6%; GLM-5.2, Kimi K3, Claude Opus 4.6 and DeepSeek V4.1-Flash at or near 0%. Over a day with limited human attention it found previously unknown browser JavaScript-engine vulnerabilities and chained them into a drive-by file read — a capability no row in Z.ai's own 16-benchmark table measures. Safeguard engagement 64% (false cover story), 92% (prefilled reasoning), 100% (abliterated, ~2,200 GPU hours / ~$4,400). NIST CAISI, quoted in the post, calls GLM-5.3 "the most cyber-capable open-weight model released to date", about four months behind the US frontier.
Two things this page does not do with it. It does not adopt Anthropic's figure over Z.ai's own card, which reports GLM-5.3's ExploitBench at 54.4 and GLM-5.2's at 24.4 where Anthropic has GLM-5.2 near zero — a ~5x magnitude disagreement under one benchmark name, disclosed on GLM-5.3's
## Conflicting Reportsand unresolved. And it does not read the safeguard figures as a retraction of Z.ai's Shield of Open Source position from 2026-08-18, which this page already holds; the two are a disagreement about open release, now with a price attached. See Open-Weights Policy Fight -
2026-09-28: Z.ai says it used its own model to ship its own model, and Jack Clark calls that an outer RSI loop. Reported in Import AI 474: Z.ai published a post describing using its own models to build its own infrastructure, with the launch of GLM-5.3-Flash as the worked case. The loop has three named parts — engineers define the objectives and the system boundaries; an Infra Agent handles analysis, hypothesis generation and code changes; and the experimental environment supplies "layered, timely, and verifiable" feedback. Clark's framing is that this is recursion running through engineers and infrastructure rather than inside a training run, which is why he calls it outer. Why it matters: Frontier Pacing has argued all quarter about speed limits on recursive self-improvement as a thing to negotiate, and R&D Automation Index measures the inner version on a benchmark. This is a lab describing the outer version as a shipped practice, on a release this wiki already holds and dates to 2026-08-26 — so the claim is checkable against an artefact rather than a score. It also lands on a lab whose releases this pipeline has repeatedly captured late through western coverage: yesterday's formal Z.ai check found nothing new, and this arrived a day later through a newsletter feed instead. What is not established, and it is most of it: no figure of any kind — no speed-up, no headcount, no iteration count, no share of changes the agent authored — and nothing read verifies the claim independently; a lab's own account of its own acceleration is exactly the claim that benefits from being unfalsifiable.
jack-clark.netansweredEGRESS_BLOCKED, so the issue was not read first-party; one search pass against its own URL → R&D Automation Index, GLM-5.3-Flash, Frontier Pacing (source) (Import AI 474) -
2026-09-24: Someone published an eight-stage post-training recipe on a Z.ai base model and says it beats Z.ai's own post-trained release. Rufus-Air: An Open LLM Post-Training Recipe — Rufus-Air, an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), run as eight serial stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, RLHF. The stated constraint is the contribution — open-source components and public data, much of it used as released, with no new human annotation and no in-house distillation teacher — and the stated organising principle is a double progression, basic to advanced capabilities and hard verifiable rewards to softer judge-based signals, from which the paper draws the rule that reward reliability is what should order the stages. Why it matters: it is a third party documenting what this lab's base weights are worth in someone else's hands, and it is the first agentic post-training this wiki has seen split by agent type — General, Coding and Search as three separate stages rather than one phase. Whether the authors are affiliated with Z.ai is not established, and it changes what "improves over the official release" means. Not established, and it is the whole comparison: no benchmark figure of any kind appears — "improves over the official GLM-4.5-Air post-trained release" and "competitive with similarly sized open models" name no benchmark, no margin and no comparator, and the abstract states the stagewise results are in the paper, which is not readable from here. Context: this repo's captured Artificial Analysis table lists no
GLM-4.5-Airrow at all — its Z AI entries are GLM-5.3 (max) 45, GLM-5.3-Flash 42 and GLM-5.3 (low) 34 — so the base being improved on is two minor generations behind what that publisher currently lists → Post-Training Scaling, Agentic Reinforcement Learning (source) (source) -
2026-09-21: Z.ai's own coding tool was uploading whole repositories — credentials included — and the fix was to open-source it. An analysis published 2026-09-18 found ZCode, Zhipu's desktop AI coding tool, packaging developers' entire workspaces and uploading them to Alibaba Cloud OSS, including
.gitfolders, LFS caches and reflogs. The worked example: one commercial project snapshot reached 313MB and 42,411 files, 86.6% of it the.gitdirectory, with 564 failed uploads queued for retry on one pass. Data types named include full project source, architecture design, complete version history, database access passwords, cloud service permission credentials and employees' personal information — stated to significantly exceed the collection scope declared in ZCode's own Privacy Policy. The user cannot decrypt what was sent: the payload is encrypted with a symmetric key wrapped under an RSA-OAEP public key issued by the server, and the private key lives only in Z.ai's cloud. Zhipu apologised the same day, attributing it to a feature that was on by default, and on 2026-09-21 open-sourced ZCode alongside a joint audit by CAICT and NSFOCUS reporting thezcode-prodbucket and all data objects in it deleted; v3.14.0 removed the RepoWiki feature and severed the snapshot-and-upload pipeline. Why it matters: on 2026-08-18 this page recorded Z.ai announcing "Shield of Open Source" — free security audits and automated code-auditing tools via ZCode — as its answer to Project Glasswing, treating openness as a security asset. The audit tool was the exfiltration path. It is also the first entry on this wiki where the harm from an agentic coding tool came from its ordinary client-side data handling rather than from model capability: nothing here involves a GLM model deciding anything. What is not established: who published the original analysis and where; the licence and repository of the open-sourced code, and whether it is the uploading version or only v3.14.0; how many users or repositories were affected — 313MB / 42,411 files is one project, and no total exists in anything read; whether any uploaded credential was used; whether affected parties were notified; and whether the audits covered the upload code path at all, since every quoted finding is about the bucket being empty. → Agents (LLM Agents), Open-Weights Policy Fight (source) -
2026-09-14: Z.ai holds the top two open-weight scores put to Congress, in prepared remarks by Nathan Lambert briefing Congressional members and staff on open models in the US–China frame: GLM-5.3 at 45 and GLM-5.3-Flash at 42 on the Artificial Analysis Intelligence Index, against 26 for the leading American open model and 23 for the next. Both figures match this wiki's own capture of that leaderboard dated 2026-09-13 cell for cell. Why it matters: this page's standing has rested on vendor claims, Code Arena placings and one LMArena row; a figure entering a congressional briefing is a different kind of use of the same number. Not established: the hearing, committee or date, and no adoption or download figure appears in anything read → Open-Weights Policy Fight (source) (source)
-
2026-09-10: Named by Anthropic as one of seven China-based labs it disrupted for distillation — Anthropic's September 2026 threat-intelligence report states it has disrupted distillation attacks from seven China-based labs since February 2026, all targeting its generally available models, and reporting names Z.ai among them alongside Alibaba, Moonshot AI, DeepSeek, Xiami and MiniMax. No case identifier, volume, account count or date range is attributed to Z.ai specifically in anything read — the only campaign given figures is Alibaba's GTG-16005. Why it matters: it is an allegation carried at reporting confidence with no per-lab evidence attached, recorded as such rather than as a finding; Z.ai has not responded in anything read. → AI-Enabled Cyberattacks, Anthropic (source) (TechCrunch)
-
2026-09-06: A GLM model appears in this wiki's LMArena top 10 for the first time, and it is not the newest one — the weekly capture places GLM 5.2 (Max) at #10 with 6.23% ±0.77%, displacing Claude Opus 4.7. No GLM row appears in any prior top-10 snapshot this repo holds — 2026-08-09, 08-16, 08-23 and 08-30 all contain zero. It joins Kimi K3 (Max) at #6, so two of the visible ten are now Chinese labs where one was on every earlier capture. Why it matters: this page has carried Z.ai's standing as vendor claims and one third-party composite — Code Arena #2, an Artificial Analysis Intelligence Index of 53 for GLM-5.2 (max) — and a blind human-preference board is a different kind of evidence from either. The interval is the second-narrowest in the top ten (±0.77%, against ±1.53% to ±2.11% for most rows), which is a statement about vote volume rather than quality. What is not established: GLM-5.3 and GLM-5.3-Flash do not appear, though the capture sees only the top 10 the page renders server-side, so nothing here says where they rank. Nothing read explains why the older model is the one that surfaced, and this page infers nothing from it → GLM-5.2, Anthropic (source) (previous week)
-
2026-08-28: The GLM-5.3 weights landed on the day they were promised for, and the licence is not the one everyone assumed —
zai-org/GLM-5.3was last modified 2026-08-28T15:22:14Z with 141 FP8.safetensorsshards, alongsidezai-org/GLM-5.3-BF16at 282 shards; a third-party quantisation (unsloth/GLM-5.3-GGUF) appeared the same day. The licence isglm-5.3, an MIT-shaped grant with one added clause: a "Model as a Service" operator whose group revenue exceeds $10 billion over any consecutive 12 months must pass Z.AI's security review before commercial use, with the scope "reasonably determined by Z.AI". The card also states, for the first time, why the gate existed: "As we scaled post-training, cyber capability developed faster than we expected", with gains "largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks". Why it matters: this page has carried the same question since 2026-08-14 — why the safety gate bound GLM-5.3 and not GLM-5.3-Flash — and both halves are now answered from a first-party artefact. The gate was about cyber capability, and the terms differ too: Flash shipped MIT, GLM-5.3 did not. What the closure does not settle: the Cybersecurity Trusted Access tier recorded here on 2026-08-18 is a verified-access control on offensive capability, and nothing in the card, the licence or the repository mentions it — a gate on a model whose weights anyone may now download is a control with no obvious remaining mechanism. This is the first first-party Z.ai artefact this wiki holds for GLM-5.3; every prior figure came through coverage → GLM-5.3, Open-Weights Policy Fight, Eval Harness Configuration (source) -
2026-08-27: "Entirely on Chinese chips" gets a number, and it turns out to be a claim about serving — Zhipu disclosed that all traffic for GLM-5.3-Flash was handled by domestic chips, with over 100,000 domestic chip units deployed for inference services, and stated that "hardware efficiency and cost per token have reached levels comparable to mainstream NVIDIA GPUs". Why it matters: yesterday's entry recorded Z.ai volunteering "running entirely on Chinese AI chips" as a headline feature with no part, vendor or training-vs-serving scope named. One of those three is now answered — "entirely" means serving — and it is the weaker of the two possible readings: inference on domestic parts does not establish that a frontier model can be trained without NVIDIA. The part and vendor are still unnamed for this model. The vendor lists and the 100,000 Huawei Ascend 910B training cluster reported alongside are stated of the GLM-5 line, not of GLM-5.3-Flash, and are held on the capture rather than the model page so the wrong silicon is not attributed to the wrong model. Nothing published lets the unit count be checked, and "comparable to mainstream NVIDIA GPUs" names no baseline, workload or measurement. Nothing read is first-party:
triviumchina.comrefused CONNECT this run, so the figures come from secondary coverage of Zhipu's disclosure → GLM-5.3-Flash (source) -
2026-08-26: GLM-5.3-Flash — MIT weights on day one, from the lab that is still holding GLM-5.3's — Z.ai released GLM-5.3-Flash, a 320B-A18B natively multimodal MoE (image and video in) with a 1M-token context window, on a hybrid sparse + linear attention architecture, under the MIT License, at $0.15/M input · $0.03/M cached · $0.50/M output with a 50% launch discount. Vendor-stated: DeepSWE v1.1 63.4 (against GLM-5.2's 46.2), Terminal-Bench 2.1 84.3, AutomationBench 48.8. Independent: Artificial Analysis Intelligence Index v4.1.1 = 57 at $0.045 per task. It had already been serving in public as the stealth model "Ox Alpha" on OpenCode, which secondary write-ups characterise as a staged load test; r/LocalLLaMA identified it at 06:28 the same morning. Why it matters: twelve days ago this page recorded GLM-5.3 shipping without its weights — withheld ~2 weeks pending a safety evaluation while being marketed as "the strongest open-weights coding model", with the window closing around 2026-08-28. Its smaller sibling shipped MIT weights immediately, and nothing read explains why the gate applies to one and not the other, or whether the GLM-5.3 window is still open. The other line in the announcement is about silicon, not the model: Z.ai states unprompted that it is "running entirely on Chinese AI chips" — a headline feature rather than a footnote, with no part, vendor, or training-vs-serving scope named. Compare LongCat-2.0, trained on Huawei Ascend 910, where this wiki recorded the same fact as a detail. Nothing read is first-party: every external host refused CONNECT this run, so the figures come from search summaries of Z.ai's X post and secondary coverage → GLM-5.3-Flash, Open-Weights Policy Fight, Post-Training Scaling (source)
-
2026-08-23: GLM-5.3 gets its first independent number, and it points the way Z.ai said it would. Artificial Analysis's weekly capture lists GLM-5.3 (max) at Intelligence Index 60 with Cost per Task USD $0.68 and Median Tokens/s 102 — the model was absent from the 2026-08-16 capture. Against the same table's GLM-5.2 (max) at 53, that is a 7-point move on a third-party composite between two models Z.ai states share an unchanged base, which is the first evidence for the Post-Training Scaling claim above from a party with nothing to sell. It is not the same claim, and smaller: an index composite is not the "~50% coding improvement" of Z.ai's internal evaluations, and
(max)compares each model's top reasoning setting. The weights are still not out — the stated safety-evaluation window closes around 2026-08-28, five days from this entry. → GLM-5.3 (source) -
2026-08-20: Jie Tang names a post-training scaling law — "death of params" — and offers GLM-5.3 as the experiment — Z.ai's CEO is reported to argue that parameter count is meaningful only alongside data, where compute is spent, and who runs the model under what conditions, identifying five scaling knobs including MoE sparsity under a new "XA-YB" notation. The load-bearing distinction: memorization prefers more parameters; reasoning prefers more post-training data and effective depth; advanced skills — the named example is finding software vulnerabilities — do not live in total parameter count once a knowledge-holding threshold is reached. GLM-5.3 is the evidence: GLM-5.2's base untouched, every gain from RL on long-horizon environments. Reported figures, GLM-5.2 → GLM-5.3: Terminal-Bench 3.0 4.6 → 28.3, DeepSWE 46.2 → 66.9, AutomationBench 26.2 → 48.2. Why it matters: this wiki's Weekly Synthesis — W33 (2026-08-10 → 2026-08-16) synthesis read "capability stopped arriving in the weights" off four release notes; this is the first time a lab shipping models has asserted it as a law with a mechanism, and it cuts against the premise of Open-Weights Policy Fight — if the base is the cheap half, publishing weights transfers less than it appears to. Two Terminal-Bench and DeepSWE figures match the independent 2026-08-14 capture, which corroborates the source; the GLM-5.2 baselines are new. All of it is vendor-stated with no harness published, arriving through a newsletter digest rather than a paper or transcript. → Post-Training Scaling, GLM-5.3 (source) (Latent Space)
-
2026-08-18: Z.ai has a Project Glasswing of its own, and it is built the opposite way round — alongside GLM-5.3, Z.ai announced "Shield of Open Source": free security audits to help users patch vulnerabilities, automated code-auditing tools via its ZCode platform, and free model usage quotas for the open-source community — plus a restricted tier, "Cybersecurity Trusted Access", under which the model's most sensitive offensive capabilities are reserved exclusively for verified users. SCMP, quoting a researcher, calls it "Project Glasswing with Chinese characteristics" and frames it as treating openness as an asset rather than a drawback. Why it matters: Anthropic's Glasswing gates a closed model to ~100–150 vetted organisations; Z.ai proposes to publish weights and gate the offensive capability, which are not obviously compatible once the weights are out. Nothing read explains how a verified-access tier is enforced on a downloadable model, or whether GLM-5.3's withheld weights — staged pending a safety evaluation, due back around 2026-08-28 — ship under this programme or are gated by it. → AI-Enabled Cyberattacks, Open-Weights Policy Fight (source) (SCMP) (The Register)
-
2026-08-14: GLM-5.3 released to GLM Coding Plan subscribers, weights held back. Z.ai states the model reuses GLM-5.2's base exactly as it was and that every gain comes from extended post-training alone; open weights and API access are stated for roughly two weeks out, in stages, after a safety evaluation. Vendor-stated: Terminal-Bench 3.0 4.6 → 28.3, DeepSWE v1.1 66.9, Agents' Last Exam CLI 28.5 (against GPT-5.6 Sol's 28.6), CyberGym 84.5%, and a ~50% internal coding improvement. No per-token price is published — Z.ai's API table still has no GLM-5.3 row (source)
-
2026-08-02 → present: GLM-5.5 remains an unconfirmed rumour — a JPMorgan note, a founder's "epic plus" remark, rumoured >1T parameters, no model card, benchmark or endpoint. GLM-5.3 is not that model and nothing read connects them (source)
-
2026-06-13 / 2026-06-16: GLM-5.2 released — GLM Coding Plan subscribers first (June 13), then open weights + standalone API (June 16). MIT-licensed weights with no regional restrictions. See GLM-5.2 for full spec. (source) (VentureBeat)
Strategic Position
- Niche: Open-weight frontier models for coding and agentic tasks — explicitly competing with closed-weights Anthropic and OpenAI models on price/performance
- Differentiation: MIT license with "no regional limits" clause makes GLM-5.2 the highest-performance model available under a truly unrestricted open-source license (as of June 2026). Compare: Llama 4 (Meta custom license restricting > 700M users), Mistral Large 3 (Apache 2.0, but 675B vs GLM-5.2's 744B)
- Data risk: Z.ai API routes through Chinese infrastructure — enterprise and government users in sensitive sectors are advised to use self-hosted open weights instead of the managed API
- China-US dynamic: Z.ai is a Chinese company operating in a geopolitical context where Chinese-origin AI models face scrutiny; GLM-5.2's MIT license is a deliberate positioning choice to reduce adoption friction globally
- Competitive benchmark: GLM-5.2 is the first open-weight model to reach Code Arena #2 — directly behind Opus 4.8, and ahead of all other open-weight models
- Release pattern changed on 2026-08-14. GLM-5.2 put subscriber access and open weights three days apart. GLM-5.3 puts them a stated two weeks apart, and conditions the second on a safety evaluation — while the release is still marketed as "the strongest open-weights coding model" (source). Whether the weights ship on that schedule is the thing to watch: the claim in the headline is currently unbacked by an artefact anyone can download. See Open-Weights Policy Fight
- Two releases, one base. GLM-5.3 is GLM-5.2's base post-trained further, which puts Z.ai on the same footing this wiki recorded for Gemini 3.7 Flash eight days earlier: shipping post-training as a numbered model. See Frontier Pacing
Related
- GLM-5.3 — current flagship, weights pending
- GLM-5.2 — the base both releases share
- Alibaba / Qwen AI Lab — Qwen series (separate Chinese open-source rival); also backed by Alibaba indirectly
- Mistral AI — European open-weight competitor; Mistral Large 3 is closest direct open-weight rival
Conflicting Reports
None currently recorded.
Referenced by
Sources
- sources/blogs/anthropic-2026-09-29-glm-5-3-cyber-capabilities.md
- sources/blogs/zai-2026-09-21-zcode-data-upload-open-source.md
- sources/newsletters/interconnects-2026-09-21-balance-of-power-open-models.md
- sources/evals/artificial-analysis-2026-09-13.md
- sources/blogs/anthropic-2026-09-10-threat-intelligence-report.md
- sources/evals/lmarena-2026-09-06.md
- sources/blogs/zai-2026-08-28-glm-5-3-open-weights.md
- sources/blogs/zai-2026-08-27-glm-5-3-flash-domestic-chips.md
- sources/blogs/zai-2026-08-26-glm-5-3-flash.md
- sources/evals/artificial-analysis-2026-08-23.md
- sources/blogs/zai-2026-08-20-post-training-scaling-law.md
- sources/blogs/zai-2026-08-18-shield-of-open-source.md
- sources/blogs/zai-2026-06-16-glm-5-2.md
- sources/blogs/zai-2026-08-14-glm-5-3.md
- https://z.ai
- https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost
- sources/papers-daily/hf-daily-2026-09-27.md
- sources/evals/artificial-analysis-2026-09-27.md
- sources/newsletters/import-ai-474-2026-W40.md