$ cat wiki/models/kolibri-1.md
Kolibri-1
Compared with
- Gemini 4 Argon
- GPT-6.1 Sol
- Claude Sonnet 5.5
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Ling-3.0-tiny
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- Claude Opus 5.5
- GPT-6 Luna
- GPT-6 Sol
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Gemini 3.5 Flash
- Grok Build
- Claude Opus 4.7
- Muse Spark
Aleph Alpha's 2026-10-03 open-weight release, and the first European open-weight model on this wiki whose pitch is sparsity rather than scale: 78.1B total parameters with 3.46B active per token, Apache 2.0, with the weights published on Hugging Face (source).
The ratio is the claim. 3.46B active is small enough to sit in the range this wiki tracks under Open-Weights Policy Fight as locally runnable, and Aleph Alpha asserts the model
matches models with up to four times its active parameter count
across math, coding, grounding and long-context tasks. Its own comparison table puts that against two ~120B dense-class open models rather than against any frontier closed model — a choice the page records, because it bounds what the claim covers.
Read first-party. aleph-alpha.com answered normally on 2026-10-04 and the
blog post is the source for every figure below. The accompanying technical report
was retrieved (3.3 MB PDF) but did not parse to text, so nothing here comes
from the report, and the architecture figures are at the level of detail the
blog post gives.
Spec
| Attribute | Value |
|---|---|
| Developer | Aleph Alpha |
| Released | 2026-10-03 |
| Announced | 2026-10-03 |
| Context window | 16,384 tokens native; up to 1,048,576 extended |
| Pricing | unknown |
| License | Apache 2.0 |
| Availability | full weights on Hugging Face |
| Catalogue id | unknown |
Pricing is unknown rather than absent: Aleph Alpha publishes no list price and | |
names no hosted endpoint in the announcement. Catalogue id is unknown because | |
| the model is not confirmed present in the OpenRouter catalogue from anything read — | |
openrouter.ai is blocked from this pipeline, and the daily spec-check Action is | |
| what will report a match if one appears. |
Rows with no slot in this schema, all from the same source:
| Attribute | Value |
|---|---|
| Architecture | Mixture-of-Experts Transformer |
| Total parameters | 78.1 billion |
| Active parameters per token | 3.46 billion |
| Experts | 384 total, 6 active |
| Layers | 50, all MoE with 1 shared expert |
| Attention | sliding window (512 tokens), full attention every 5th layer |
| Languages | English and German |
Release Date
2026-10-03, announced and released the same day. Weights were published to Hugging Face on that date.
Benchmarks
As published by Aleph Alpha (source). The comparison set is Aleph Alpha's own:
| Benchmark | Kolibri-1 | Nemotron 3 Super 120B | Mistral Small 4 119B |
|---|---|---|---|
| AIME 2025 (EN) | 96.9 | 91.7 | 79.8 |
| GPQA Diamond (EN) | 84.3 | 78.0 | 74.7 |
| HumanEval+ | 92.7 | 94.7 | 92.8 |
| BFCL v4 | 61.4 | 61.0 | 58.0 |
| Read the shape of the table, not just the wins. Kolibri-1 leads on the two | |||
| reasoning benchmarks by wide margins (+5.2 on AIME 2025, +6.3 on GPQA | |||
| Diamond against the stronger comparator) and is behind on HumanEval+ | |||
| (92.7 against Nemotron's 94.7) while effectively tied on tool calling | |||
| (61.4 against 61.0 on BFCL v4). So the four-times-active-parameters claim is | |||
| carried by maths and science, not by code — which is consistent with the training | |||
| mix, where code is about 14% against English ~62% and German 21.3%. |
No frontier closed model appears in the table, and no figure for any Claude, GPT, Gemini or Qwen model is stated, so this page carries no comparison against the frontier.
Use Cases
The announcement positions the model for sovereign and on-premises deployment — Apache 2.0 weights, published in full, for operators who need the model to run on their own hardware. The bilingual training mix (4.3 trillion German tokens, 21.3% of a 20 trillion token three-stage pre-training run) targets German-language work specifically, which is the gap Aleph Alpha's own earlier writing argues for.
The extended 1,048,576-token context is stated as an extension of a 16,384-token native window, so long-context use rests on that extrapolation rather than on native training length.
No deployment guidance, hardware requirement or serving configuration is stated in
the first-party post. See ## Conflicting Reports.
Compared To
- Nemotron 3 Super 120B and Mistral Small 4 119B — the two models Aleph
Alpha's own table compares against; see
## Benchmarksabove. A Mistral Small 4 release is held on this wiki under Mistral AI. - Qwen 3.8 27B — the other current open-weight model this wiki tracks in the locally-runnable class. No head-to-head figure exists in anything read; the two are not compared here.
- Sparsity as a design axis is the subject of a scaling result captured the same day — see Test-Time Compute (Inference-Time Compute Scaling) on Scaling Laws for Looped Mixture of Experts, which reports ~3× active-parameter efficiency from sparsity as a general finding. That paper does not evaluate Kolibri-1, and the two are linked as the same question, not as evidence for each other.
Sources
- Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha, 2026-10-03, read first-party 2026-10-04
Conflicting Reports
Native context window: 16,384 or 262,144 tokens. The first-party announcement states 16,384 native, extended to 1,048,576. Secondary coverage published 2026-10-03/04 — MarkTechPost, testingcatalog, alphasignal and daily.dev — states 262,144 native, extended to 1,048,576 (source).
The first-party figure is the one in ## Spec, per the CLAUDE.md priority
rule (official announcement > third-party reporting). The disagreement is recorded
rather than resolved because it is not a rounding difference — a 16× gap in native
length changes what the extrapolation is doing — and because the technical report
that would settle it did not parse.
Figures only secondary coverage states, and which are therefore not adopted onto this page: an FP8 checkpoint of about 78 GB, runnable on a single B200/B300/H200 or 2×H100 SXM5 through vLLM, and a 128,000-token vocabulary. The first-party post states no file size, no hardware and no vocabulary size.