AI Trend Notifier
EN
← wiki

$ cat wiki/models/qwen-3-8-flash-next.md

Qwen3.8-Flash-Next

modelupdated 2026-08-26created 2026-08-26

Spec

AttributeValue
DeveloperAlibaba / Qwen AI Lab
Releasednot yet
Announced2026-08-25
Context windowunknown
Pricingunknown
Licenseunknown
AvailabilityModelScope, announced for 2026-08-26 23:00 (UTC+08:00), standard and FP8
Released is not yet and that is the whole state of this page. A ModelScope
teaser went live 2026-08-25 with an **"Upcoming Open-Release" badge and a countdown
timer** pointed at 2026-08-26 23:00 UTC+08:00 — after this run. Nothing read
establishes that the release landed
(source).

License repeats a gap this wiki already carries on Qwen 3.8 Max, whose licence name has been unknown since its weights shipped 2026-08-12. One summary of this model states plainly that "the parameter count and licensing terms have not been disclosed".

Release Date

Announced 2026-08-25; scheduled 2026-08-26 23:00 (UTC+08:00); unreleased at capture.

Positioned as an open-weight, multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself — early access to the architecture ahead of the full family (source).

Benchmarks

None. No evaluation figure of any kind appears in anything read — not a leaderboard row, not a vendor chart, not a claim. For a model previewing a next-generation architecture, that is the fact worth recording rather than an omission to fill in later.

Use Cases

Not stated by anything read beyond "open-weight" and "multimodal". Three r/LocalLLaMA threads on 2026-08-25 anticipate local deployment — one titled "This architecture could be surprisingly local-friendly once the weights drop", another noting day-0 support from unsloth — but these are community expectations about an unreleased model and carry no lab artefact (source).

Compared To

  • Qwen 3.8 Max — the 2.4T-parameter flagship, released 2026-08-03, weights 2026-08-12. Whether Flash-Next is a smaller sibling of the 3.8 line or a separate architecture branch is not established; "previews Qwen4" argues the latter and nothing states it.
  • Qwen 3.8 27B — the open-weight member of the 3.8 line this wiki already holds.

No comparison to any non-Alibaba model is possible: there are no figures on either side of one.

Conflicting Reports

Parameter count. Most summaries give 125B total with 6B active per token; one X post adds 51B of N-gram embeddings on top; and one write-up states plainly that "the parameter count and licensing terms have not been disclosed". All three were read this run and they cannot all be right. The Spec table above therefore carries no parameter row, and the figures are recorded here rather than published as fact (source).

Architecture. A single write-up names GDN hybrid architecture, Qwen Sparse Attention (QSA), and 48 layers with every 4th layer using GQA attention and the rest "the new linear attention". This is single-source and pre-release; no first-party model card was read (source).

Referenced by

Sources