$ cat wiki/models/k2-horizon.md
K2 Horizon
Compared with
- Astra
- Gemini 3.8 Flash
- Muse Spark 1.3
- Claude Fable 5.1
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- DeepSeek V4-Pro-0813
- Grok 4.6
- Laguna S 2.1
- Kimi K3
- Inkling
- LongCat-2.0
- MiniMax M3
- GLM-5.3
- Qwen 3.8 27B
- Gemini 3.7 Flash
- Muse Glimmer
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
A family of six models from the Institute of Foundation Models, released 2026-09-03 under Apache 2.0 with weights, code, training data and methodology published together (source).
This is not a Moonshot model. Moonshot AI's Kimi line also uses "K2", and this release surfaced through r/LocalLLaMA, where Kimi releases are routine. K2 Horizon comes from IFM at MBZUAI; the two lines share three characters and nothing else. The disambiguation is written here and on Institute of Foundation Models (IFM) because getting it wrong would attribute an Abu Dhabi lab's work to a Beijing one.
Spec
| Attribute | Value |
|---|---|
| Developer | Institute of Foundation Models (IFM) |
| Released | 2026-09-03 |
| Announced | 2026-09-03 |
| Context window | 524,288 (flagship 375B-A23B, native from midtraining onward) |
| Pricing | unknown |
| License | Apache 2.0 (weights and code) |
| Availability | Hugging Face (weights, code, checkpoints, training data where redistributable) |
Pricing is unknown rather than "free": the weights are downloadable at no | |
| cost, but **no first-party hosted endpoint or price was described in anything | |
| read**, and inference cost is whoever serves it. The distinction is the one | |
unknown exists to preserve. |
One page, six models. CLAUDE.md's rule is one page per model, not per series,
and this release is the edge of it: six checkpoints announced in one post, sharing
a licence, a recipe and a training corpus, with no per-model figure published
for any of them. Splitting them into six pages would produce six spec tables
identical but for a parameter count, and six ## Benchmarks sections all reading
"no score published". They are held together until a source distinguishes them.
Release Date
2026-09-03. Sizes: 0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B.
Each model is stated to be pretrained on approximately 20 trillion tokens. The flagship 375B-A23B is a Mixture of Experts with 375B total parameters and 23B active at inference (source).
IFM states a role per size: 0.9B local development, 3.7B single-node serving, 7B cost-sensitive deployment, 32B everyday heavy use. The stated roles for the remaining two sizes were not carried by anything read.
Benchmarks
None. This is the finding rather than a gap in the capture.
IFM publishes three comparative claims and no benchmark name and no score attached to any of them (source):
| Claim | As published |
|---|---|
| Across reasoning, mathematics, coding and agentic tasks | "top-tier performance in every size class" |
| 0.9B, 3.7B, 7B | stated to set new state of the art at their respective scales |
| Agentic tool use, terminal, long-horizon workflows | matches or beats open-weight MoE models up to 2.6× its size; "competitive with closed frontier models" |
ifm.ai, huggingface.co and artificialanalysis.ai are all unreachable from | |
| this run's sandbox, so it cannot be ruled out that the announcement carries a | |
| scorecard the two search passes did not surface. What can be said is that **no | |
| figure reached this wiki**, and a model card that publishes its training data but | |
| not its evaluation numbers has an unusual shape. |
The 2.6× claim is the one that would be checkable if a number were attached. An open-weight MoE 2.6× the flagship's total parameters is roughly a 975B model; the comparison names no such model, no benchmark and no harness. Per Eval Harness Configuration, that is a positioning claim.
Use Cases
The release's argument is reproducibility rather than capability: weights, code, training data where redistribution licences permit, intermediate checkpoints, training configurations, fine-grained logs and evaluation results. For restricted datasets IFM publishes source descriptions, construction methods and mixture recipes instead of the data (source).
See Open-Weights Policy Fight — this sits at the far end of a spectrum whose other end this wiki spent August documenting.
Compared To
- GLM-5.3 — the August release Open-Weights Policy Fight tracked through whether → when → terms. GLM-5.3 shipped weights under a bespoke licence with a revenue-threshold condition; K2 Horizon ships weights, code and data under Apache 2.0 with no condition, and publishes no benchmark score where GLM-5.3 published a scorecard
- Qwen 3.8 27B, Gemma 4 12B, Nemotron 3.5 Lightning — open-weight releases at overlapping sizes, all of them weights-only
- Kimi K3 — Moonshot AI's current flagship, and the reason the disambiguation at the top of this page exists