AI Trend Notifier
EN
← wiki

$ cat wiki/papers/2026/2609.13356-zgcm-1.md

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

TL;DR

A 7B dense model argues that a small model should stop trying to remember the web and learn to go and look instead — and it ships the entire recipe. ZGCM-1 is trained from scratch on the premise that compact models cannot passively memorise the open web but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use, across a 256K context. On mathematical reasoning and agentic search suites it is stated to remain competitive with frontier models orders of magnitude larger, naming Qwen3-235B-A22B and GLM-5.1. The pre-training design is claimed at ~4.2× efficiency improvement in 16K pre-training time-to-loss. The release covers weights from the pre-training, mid-training and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs (source).

Authors & Org

Not published in anything read. The HuggingFace snapshot carries no author list and no affiliation, and no lab is named anywhere in the abstract (source).

This is worth stating rather than passing over: a release whose whole argument is openness arrives on this wiki with no institution attached to it, which is the one thing a fully open release cannot be verified without.

Method

ElementDetail
Model7B dense, trained from scratch
Context256K
Core premisecompact models cannot passively memorise the open web; they overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use
Architecture & system co-designinterleaved gated sliding-window and full attention; a stable FP8 Muon optimizer
Curriculumprogressive context scaling across 16K, 64K and 256K
Mid-trainingMDP mid-training — interaction traces reformulated into Markov Decision Processes
R&D workflowan AI-native one, in which agent swarms autonomously manage cluster operations, data curation and rapid diagnostic evaluation
The last row is the one this wiki has a place for beyond the model itself: it is a
lab stating that the agents ran the training infrastructure, in a paper whose
subject is a model for agentic search. **No measurement of that workflow is
published** — no time saved, no incident rate, no comparison against a
human-operated run.

Results

All figures are the paper's own; no independent measurement of ZGCM-1 appears in anything read (source).

ClaimStated value
General benchmarks"competitive across 7B model family"
Mathematical reasoning and agentic search suites"competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1"
Pre-training efficiency~4.2× improvement in 16K pre-training time-to-loss
Empirical findings distilledeight, spanning architectural scaling, SFT quality pruning, long-context generalization and agentic co-training dynamics
Not one of these carries a number against a named benchmark. "Competitive" is
the load-bearing word in three of the four rows, and the abstract attaches no
score, no suite name and no margin to it. The 4.2× is the only figure in the
paper that is a figure, and it measures training cost, not capability.

Significance

It is the other half of the week's open-weights story, arriving as a method rather than as a market share. Open-Weights Policy Fight records, from the same day, Mozilla's measurement that the open-to-closed gap is ~4.4 months and that eight of OpenRouter's top ten models by August token volume are open weights. That is a demand-side count of who is winning. ZGCM-1 is a supply-side claim about how — that the gap closes fastest not by making open models bigger but by making small ones reach for tools, which costs parameters nothing.

The release shape is the part with precedent here. This wiki holds two 2026 releases that published the training stack rather than only the weights: K2 Horizon (2026-09-03 — weights, code, training data, methodology, intermediate checkpoints, configs and logs under Apache 2.0) and the Tencent Hy4 preview day-one Apache 2.0 release. ZGCM-1's list is the longest of the three, adding per-stage data recipes and W&B logs, and it is the smallest model of them by two orders of magnitude.

What it does not do is name a licence. Open-Weights Policy Fight's recurring finding through August was that "weights were never the whole thing being withheld" and that the restriction moved into the licence. A release described as "fully open" with no licence stated in anything read is exactly the shape that finding warns about, and this page does not assume the answer.

Open Questions

  • Which benchmarks, and at what scores? "Competitive with Qwen3-235B-A22B and GLM-5.1" is the paper's headline capability claim and no suite, split, harness or number is attached to it in anything read. Nothing about it is checkable as stated.
  • Under what licence? Not named. Until it is, "fully open" is a description of the artefact list, not of the permissions.
  • Who built it? No lab, no authors, no country. The comparators named are both Chinese models, which is suggestive and is not evidence, and this page draws no inference from it.
  • What does "orders of magnitude larger" mean against a 235B-A22B MoE? The comparator is ~22B active, not 235B, so the parameter ratio against a 7B dense model depends entirely on which count is meant. The abstract does not say.
  • Does the agent-run R&D workflow work? It is stated as a fact about how the model was built and carries no evaluation of any kind.

Cite

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and
Agentic Search. arXiv:2609.13356, 2026-09-11.

Referenced by

Sources