$ cat wiki/models/glm-5-2.md
GLM-5.2
modelupdated 2026-07-28created 2026-07-13
Compared with
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Claude Opus 5
- Laguna S 2.1
- Kimi K3
- Inkling
- GPT-5.6 Sol
- LongCat-2.0
- MiniMax M3
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Grok 4.5
- Claude Sonnet 5
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
Spec
| Attribute | Value |
|---|---|
| Developer | Z.ai (Z.ai / Zhipu AI) |
| Released | 2026-06-13 (subscribers) / 2026-06-16 (open weights + API) |
| Announced | unknown |
| Context window | 1,000,000 tokens |
| Pricing | $1.40/M input · $4.40/M output — Z.ai's own endpoint on the OpenRouter catalogue, read 2026-08-01. Earlier revisions of this row quoted ≈$0.62–$0.77, which was not a price change: OpenRouter serves this model from 33 providers between $0.72 and $2.31 on input, and /models reports whichever one is routing at that instant. Four alerts in a day named $0.97, $1.19, $1.12 and $0.76 — Alibaba, SiliconFlow, an intermediate route, and StreamLake. Reseller quotes are real prices but they are not this model's price; quote the first-party figure, or quote a reseller with its name attached. |
| License | MIT (unrestricted, "no regional limits") |
| Availability | GLM Coding Plan (highest subscription tier), standalone API, open weights, NVIDIA NIM hosted |
| Total parameters | 744B |
| Active parameters | ~40B per token (Mixture-of-Experts) |
| Max output tokens | 131,072 |
| Reasoning modes | high, max effort levels |
Release Date
- June 13, 2026: Available to Z.ai GLM Coding Plan subscribers
- June 16, 2026: Open weights published + standalone API launched
⚡ Note: Captured 27 days late (July 13, 2026 ingest). GLM-5.2's release period overlapped with the Grok V9-Medium (June 16) launch and was not surfaced in Tier-1 source polling during that week.
Architecture: IndexShare
GLM-5.2 introduces IndexShare — an architectural optimization to sparse attention that reduces per-token compute by 2.9× at the full 1-million-token context length compared to standard sparse attention. This is the primary reason the 1M-context capability is practical (not just theoretical).
Benchmarks
| Benchmark | GLM-5.2 | Notes |
|---|---|---|
| Code Arena | #2 globally | Trails Claude Opus 4.8 by ~1pp on 3 long-horizon coding evals |
| Long-horizon coding (vs GPT-5.5) | Beats GPT-5.5 | At ~1/6th API cost |
- Described by Interconnects.ai as "the step change for open agents"
- Data risk note: TechTimes (June 17) — the Z.ai hosted API routes queries through Zhipu AI's China-based infrastructure. For data-sensitive workloads, self-hosting the open weights is recommended.
Use Cases
- Long-horizon agentic coding tasks (primary target — Code Arena #2)
- Cost-effective alternative to Opus 4.8 for organizations that can accept the licensing/data-routing tradeoffs
- Self-hosted deployment for organizations that need frontier-class open weights without data-routing risk
Compared To
| Model | Code Arena / SWE-Pro | License | Cost |
|---|---|---|---|
| GLM-5.2 | #2 (long-horizon coding) | MIT | ~1/6th GPT-5.5 API cost |
| Claude Opus 4.8 | #1 (long-horizon) | Closed | $15/$75 per Mtok |
| GPT-5.5 | Below GLM-5.2 (coding) | Closed | Standard OpenAI API |
| Mistral Large 3 | TBD (July 2026 GA) | Apache 2.0 | TBD |
| Llama 4 | Below GLM-5.2 | Meta custom | Free (self-hosted) |
Referenced by
Sources
- sources/blogs/zai-2026-06-16-glm-5-2.md
- https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost
- https://aiweekly.co/alerts/zais-glm-52-brings-1m-token-context-to-open-weight-coding
- https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm
- https://openrouter.ai/z-ai/glm-5.2