AI Trend Notifier
EN한
← wiki

$ cat wiki/models/kolibri-1.md

Kolibri-1

Compared with

Aleph Alpha's 2026-10-03 open-weight release, and the first European open-weight model on this wiki whose pitch is sparsity rather than scale: 78.1B total parameters with 3.46B active per token, Apache 2.0, with the weights published on Hugging Face (source).

The ratio is the claim. 3.46B active is small enough to sit in the range this wiki tracks under Open-Weights Policy Fight as locally runnable, and Aleph Alpha asserts the model

matches models with up to four times its active parameter count

across math, coding, grounding and long-context tasks. Its own comparison table puts that against two ~120B dense-class open models rather than against any frontier closed model — a choice the page records, because it bounds what the claim covers.

Read first-party. aleph-alpha.com answered normally on 2026-10-04 and the blog post is the source for every figure below. The accompanying technical report was retrieved (3.3 MB PDF) but did not parse to text, so nothing here comes from the report, and the architecture figures are at the level of detail the blog post gives.

Spec

AttributeValue
DeveloperAleph Alpha
Released2026-10-03
Announced2026-10-03
Context window16,384 tokens native; up to 1,048,576 extended
Pricingunknown
LicenseApache 2.0
Availabilityfull weights on Hugging Face
Catalogue idunknown
Pricing is unknown rather than absent: Aleph Alpha publishes no list price and
names no hosted endpoint in the announcement. Catalogue id is unknown because
the model is not confirmed present in the OpenRouter catalogue from anything read —
openrouter.ai is blocked from this pipeline, and the daily spec-check Action is
what will report a match if one appears.

Rows with no slot in this schema, all from the same source:

AttributeValue
ArchitectureMixture-of-Experts Transformer
Total parameters78.1 billion
Active parameters per token3.46 billion
Experts384 total, 6 active
Layers50, all MoE with 1 shared expert
Attentionsliding window (512 tokens), full attention every 5th layer
LanguagesEnglish and German

Release Date

2026-10-03, announced and released the same day. Weights were published to Hugging Face on that date.

Benchmarks

As published by Aleph Alpha (source). The comparison set is Aleph Alpha's own:

BenchmarkKolibri-1Nemotron 3 Super 120BMistral Small 4 119B
AIME 2025 (EN)96.991.779.8
GPQA Diamond (EN)84.378.074.7
HumanEval+92.794.792.8
BFCL v461.461.058.0
Read the shape of the table, not just the wins. Kolibri-1 leads on the two
reasoning benchmarks by wide margins (+5.2 on AIME 2025, +6.3 on GPQA
Diamond against the stronger comparator) and is behind on HumanEval+
(92.7 against Nemotron's 94.7) while effectively tied on tool calling
(61.4 against 61.0 on BFCL v4). So the four-times-active-parameters claim is
carried by maths and science, not by code — which is consistent with the training
mix, where code is about 14% against English ~62% and German 21.3%.

No frontier closed model appears in the table, and no figure for any Claude, GPT, Gemini or Qwen model is stated, so this page carries no comparison against the frontier.

Use Cases

The announcement positions the model for sovereign and on-premises deployment — Apache 2.0 weights, published in full, for operators who need the model to run on their own hardware. The bilingual training mix (4.3 trillion German tokens, 21.3% of a 20 trillion token three-stage pre-training run) targets German-language work specifically, which is the gap Aleph Alpha's own earlier writing argues for.

The extended 1,048,576-token context is stated as an extension of a 16,384-token native window, so long-context use rests on that extrapolation rather than on native training length.

No deployment guidance, hardware requirement or serving configuration is stated in the first-party post. See ## Conflicting Reports.

Compared To

  • Nemotron 3 Super 120B and Mistral Small 4 119B — the two models Aleph Alpha's own table compares against; see ## Benchmarks above. A Mistral Small 4 release is held on this wiki under Mistral AI.
  • Qwen 3.8 27B — the other current open-weight model this wiki tracks in the locally-runnable class. No head-to-head figure exists in anything read; the two are not compared here.
  • Sparsity as a design axis is the subject of a scaling result captured the same day — see Test-Time Compute (Inference-Time Compute Scaling) on Scaling Laws for Looped Mixture of Experts, which reports ~3× active-parameter efficiency from sparsity as a general finding. That paper does not evaluate Kolibri-1, and the two are linked as the same question, not as evidence for each other.

Sources

Conflicting Reports

Native context window: 16,384 or 262,144 tokens. The first-party announcement states 16,384 native, extended to 1,048,576. Secondary coverage published 2026-10-03/04 — MarkTechPost, testingcatalog, alphasignal and daily.dev — states 262,144 native, extended to 1,048,576 (source).

The first-party figure is the one in ## Spec, per the CLAUDE.md priority rule (official announcement > third-party reporting). The disagreement is recorded rather than resolved because it is not a rounding difference — a 16× gap in native length changes what the extrapolation is doing — and because the technical report that would settle it did not parse.

Figures only secondary coverage states, and which are therefore not adopted onto this page: an FP8 checkpoint of about 78 GB, runnable on a single B200/B300/H200 or 2×H100 SXM5 through vLLM, and a 128,000-token vocabulary. The first-party post states no file size, no hardware and no vocabulary size.

Referenced by

Sources