AI Trend Notifier
EN한
← wiki

$ cat wiki/models/glm-5-3-flash.md

GLM-5.3-Flash

Compared with

Spec

AttributeValue
DeveloperZ.ai
Released2026-08-26
Announced2026-08-26
Context window1,048,576 (1M)
Pricing$0.15/M input · $0.03/M cached input · $0.50/M output
LicenseMIT
AvailabilityHugging Face (open weights), Z.ai API
320B total / 18B active MoE (320B-A18B), natively multimodal — image and
video input — on a hybrid sparse + linear attention architecture that Z.ai
describes as sharply reducing long-context serving cost while preserving precise
long-context capability
(source).

A 50% launch discount temporarily halves the rates above; the Spec table carries the standard price, not the promotional one (source).

Release Date

2026-08-26, announced and released the same day — but not the same day it started serving. Z.ai's own announcement says it was "previously previewed as Ox Alpha", the stealth model that had been running on OpenCode, and secondary write-ups characterise that period as a staged public load test ahead of the release. r/LocalLLaMA carried the identification at 06:28 the same morning, ahead of the release thread at 14:06 (source).

The line in the announcement that is not about the model: "running entirely on Chinese AI chips." Z.ai states it in the launch post itself, unprompted and without naming the silicon. This wiki already holds one Chinese frontier model trained off NVIDIA — LongCat-2.0, on Huawei Ascend 910 — so the claim is not unprecedented; what is new is a lab volunteering it as a headline feature of a frontier release rather than a footnote. Nothing read names the part, the vendor, or whether "entirely" covers training, serving or both (source).

One of those three is now answered, and it is the one that changes the claim's size. Subsequent disclosure states that all traffic was handled entirely by domestic chips, with over 100,000 domestic chip units deployed for inference services, and that "hardware efficiency and cost per token have reached levels comparable to mainstream NVIDIA GPUs". So "entirely" is established for serving; nothing read states what GLM-5.3-Flash was trained on, and the part and vendor remain unnamed for this model. The vendor lists reported elsewhere — Moore Threads, Cambricon, Kunlun Chip, MetaX, Enflame, Hygon — and the 100,000 Huawei Ascend 910B training cluster are stated of the GLM-5 line, not of GLM-5.3-Flash, and are recorded on the capture rather than here so the wrong silicon is not attributed to the wrong model (source).

A serving claim at that scale is a different assertion from a training claim, and the weaker of the two: inference on domestic parts does not establish that a frontier model can be trained without NVIDIA. It is also the claim that is hardest to check from outside — nothing published lets anyone verify the unit count, and "comparable to mainstream NVIDIA GPUs" names no baseline, no workload and no measurement (source).

Benchmarks

Vendor-stated (Z.ai's model card, quoted by secondary write-ups — no harness published):

BenchmarkGLM-5.3-FlashGLM-5.2
DeepSWE v1.163.446.2
Terminal-Bench 2.184.3not stated
AutomationBench48.8not stated
Independent: Artificial Analysis Intelligence Index v4.1.1 — 57, at a
discounted cost of $0.045 per task
(source).

That independent 57 is the figure worth holding onto, because it is the one measured by someone other than the vendor. For scale against the models this wiki already tracks: DeepSeek V4-Pro-0813 sits at 53 on the same index and Nemotron 3.5 Lightning at 24.

Z.ai's own framing — "approaching Claude Opus 4.8 on coding and agentic benchmarks" at "one-tenth the price" of GLM-5.2 — is a vendor claim with no side-by-side table in anything read. This wiki holds no Opus 4.8 figure for DeepSWE v1.1, Terminal-Bench 2.1 or AutomationBench, so the comparison cannot be checked here and is recorded as a claim rather than a result.

Use Cases

Coding and agentic work is where every published figure sits (DeepSWE, Terminal-Bench, AutomationBench). Native multimodal input (image, video) and a 1M context window are stated capabilities with no benchmark attached to either (source).

How this model was launched, per the lab (2026-09-28)

Z.ai published a post, reported in Import AI 474, describing this release as the worked case for using its own models to build its own infrastructure: engineers set the objectives and system boundaries, an Infra Agent performed analysis, hypothesis generation and code changes, and the experimental environment returned "layered, timely, and verifiable" feedback (source).

No figure accompanies the account — no speed-up, no iteration count, no share of the changes the agent authored — and nothing read verifies it independently, so it is held as a described practice rather than a measured one. It is recorded here because the claim is about this page's subject: a reader comparing this model's launch to GLM-5.3's now has the lab's own statement that the two were not built the same way. Treatment on R&D Automation Index.

Compared To

  • GLM-5.3 — released 2026-08-14, twelve days earlier, and its opposite in the one respect that matters most here: GLM-5.3's weights were withheld ~2 weeks pending a safety evaluation while it was marketed as "the strongest open-weights coding model". GLM-5.3-Flash shipped MIT weights on day one. Nothing read explains why the smaller sibling needed no such gate, or whether the GLM-5.3 window (staged for ~2026-08-28) is still open.
  • GLM-5.2 — the 744B/40B MIT base the whole 5.x line descends from, and the only model GLM-5.3-Flash publishes a paired figure against.
  • Qwen3.8-Flash-Next — released the same day, also open-weight, also multimodal MoE, also framed on cost-efficiency, and roughly half the total size at 176B. Two Chinese labs shipping cost-optimised open multimodal MoEs within hours of each other is the day's shape, not a coincidence either lab comments on.
  • DeepSeek V4-Flash — the other Chinese open-weight "Flash" tier, MIT, at $0.14/$0.28 per M. GLM-5.3-Flash is priced slightly above it on input and ~1.8× on output, for a model with ~1.4× the active parameters and native multimodality.

Conflicting Reports

  • The published context window disagrees with the catalogue, and the page's figure stands. spec-check run 90 (2026-09-12) reports context 1,048,576 vs catalogue 1,310,720 (source).

    The Spec row is not changed: the cell above is faithful to the vendor material this page cites, and CLAUDE.md gives the vendor announcement precedence over the catalogue. 1,310,720 is 1.25 × 1,048,576, which is a round figure in the catalogue's units rather than an arbitrary one — recorded as an observation and not adopted as an explanation.

    What is not established: which figure is correct, and whether the catalogue is describing the same deployment. openrouter.ai is blocked from the daily run's sandbox, so this cannot be settled here.

    Reported by a red Action on every run since 2026-08-29 — 25 consecutive — and unrecorded on this page until 2026-09-13.

Referenced by

Sources