$ cat wiki/models/glm-5-2.md
GLM-5.2
modelupdated 2026-09-06created 2026-07-13
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Z.ai (Z.ai / Zhipu AI) |
| Released | 2026-06-13 (subscribers) / 2026-06-16 (open weights + API) |
| Announced | unknown |
| Context window | 1,000,000 tokens |
| Pricing | $1.40/M input · $4.40/M output — Z.ai's own endpoint on the OpenRouter catalogue, read 2026-08-01. Earlier revisions of this row quoted ≈$0.62–$0.77, which was not a price change: OpenRouter serves this model from 33 providers between $0.72 and $2.31 on input, and /models reports whichever one is routing at that instant. Four alerts in a day named $0.97, $1.19, $1.12 and $0.76 — Alibaba, SiliconFlow, an intermediate route, and StreamLake. Reseller quotes are real prices but they are not this model's price; quote the first-party figure, or quote a reseller with its name attached. |
| License | MIT (unrestricted, "no regional limits") |
| Availability | GLM Coding Plan (highest subscription tier), standalone API, open weights, NVIDIA NIM hosted |
| Total parameters | 744B |
| Active parameters | ~40B per token (Mixture-of-Experts) |
| Max output tokens | 131,072 |
| Reasoning modes | high, max effort levels |
Release Date
- June 13, 2026: Available to Z.ai GLM Coding Plan subscribers
- June 16, 2026: Open weights published + standalone API launched
⚡ Note: Captured 27 days late (July 13, 2026 ingest). GLM-5.2's release period overlapped with the Grok V9-Medium (June 16) launch and was not surfaced in Tier-1 source polling during that week.
Architecture: IndexShare
GLM-5.2 introduces IndexShare — an architectural optimization to sparse attention that reduces per-token compute by 2.9× at the full 1-million-token context length compared to standard sparse attention. This is the primary reason the 1M-context capability is practical (not just theoretical).
Benchmarks
| Benchmark | GLM-5.2 | Notes |
|---|---|---|
| Code Arena | #2 globally | Trails Claude Opus 4.8 by ~1pp on 3 long-horizon coding evals |
| Long-horizon coding (vs GPT-5.5) | Beats GPT-5.5 | At ~1/6th API cost |
| LMArena (2026-09-06) | 6.23% ±0.77%, rank #10 | (Max). The board's own percentage with a confidence interval, not Elo |
| 2026-09-06 — first appearance on LMArena's visible leaderboard. GLM 5.2 (Max) | ||
| enters at #10, displacing Claude Opus 4.7. **No GLM row appears in any prior | ||
| top-10 snapshot this repo holds** (2026-08-09, 08-16, 08-23, 08-30 all have zero), | ||
| so this is a first, not a rise. The ±0.77% interval is the second-narrowest in | ||
| the top ten, most rows sitting at ±1.5% to ±2.1% — a statement about vote volume, | ||
| not about quality. The capture sees **only the top 10 the page renders | ||
| server-side**, so it says nothing about where GLM-5.3 or | ||
| GLM-5.3-Flash rank, and nothing read explains why the older model is the | ||
| one that surfaced (source). |
- Described by Interconnects.ai as "the step change for open agents"
- Data risk note: TechTimes (June 17) — the Z.ai hosted API routes queries through Zhipu AI's China-based infrastructure. For data-sensitive workloads, self-hosting the open weights is recommended.
Use Cases
- Long-horizon agentic coding tasks (primary target — Code Arena #2)
- Cost-effective alternative to Opus 4.8 for organizations that can accept the licensing/data-routing tradeoffs
- Self-hosted deployment for organizations that need frontier-class open weights without data-routing risk
Compared To
| Model | Code Arena / SWE-Pro | License | Cost |
|---|---|---|---|
| GLM-5.2 | #2 (long-horizon coding) | MIT | ~1/6th GPT-5.5 API cost |
| Claude Opus 4.8 | #1 (long-horizon) | Closed | $15/$75 per Mtok |
| GPT-5.5 | Below GLM-5.2 (coding) | Closed | Standard OpenAI API |
| Mistral Large 3 | TBD (July 2026 GA) | Apache 2.0 | TBD |
| Llama 4 | Below GLM-5.2 | Meta custom | Free (self-hosted) |
Referenced by
AI-Enabled CyberattacksAnthropicClaude Mythos PreviewDeepSeek V4Eval Harness ConfigurationFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157)GLM-5.3GLM-5.3-FlashJune 2026 — Monthly DigestKimi K3Liquid AIMilitary and Intelligence Capability EvalsMiniMax M3Open-Weights Policy FightPost-Training ScalingT1: Terminal Agent Reinforcement Learning for Long-Horizon TasksWeekly Synthesis — 2026-W29 (2026-07-13 ~ 2026-07-19)Weekly Synthesis — W34 (2026-08-17 → 2026-08-23)Weekly Synthesis — W36 (2026-08-31 → 2026-09-06)Z.ai
Sources
- sources/evals/lmarena-2026-09-06.md
- sources/blogs/zai-2026-06-16-glm-5-2.md
- https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost
- https://aiweekly.co/alerts/zais-glm-52-brings-1m-token-context-to-open-weight-coding
- https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm
- https://openrouter.ai/z-ai/glm-5.2