$ cat wiki/models/qwen-3-8-flash-next.md
Qwen3.8-Flash-Next
Spec
| Attribute | Value |
|---|---|
| Developer | Alibaba / Qwen AI Lab |
| Released | not yet |
| Announced | 2026-08-25 |
| Context window | unknown |
| Pricing | unknown |
| License | unknown |
| Availability | ModelScope, announced for 2026-08-26 23:00 (UTC+08:00), standard and FP8 |
Released is not yet and that is the whole state of this page. A ModelScope | |
| teaser went live 2026-08-25 with an **"Upcoming Open-Release" badge and a countdown | |
| timer** pointed at 2026-08-26 23:00 UTC+08:00 — after this run. Nothing read | |
| establishes that the release landed | |
| (source). |
License repeats a gap this wiki already carries on Qwen 3.8 Max, whose
licence name has been unknown since its weights shipped 2026-08-12. One summary
of this model states plainly that "the parameter count and licensing terms have
not been disclosed".
Release Date
Announced 2026-08-25; scheduled 2026-08-26 23:00 (UTC+08:00); unreleased at capture.
Positioned as an open-weight, multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself — early access to the architecture ahead of the full family (source).
Benchmarks
None. No evaluation figure of any kind appears in anything read — not a leaderboard row, not a vendor chart, not a claim. For a model previewing a next-generation architecture, that is the fact worth recording rather than an omission to fill in later.
Use Cases
Not stated by anything read beyond "open-weight" and "multimodal". Three r/LocalLLaMA threads on 2026-08-25 anticipate local deployment — one titled "This architecture could be surprisingly local-friendly once the weights drop", another noting day-0 support from unsloth — but these are community expectations about an unreleased model and carry no lab artefact (source).
Compared To
- Qwen 3.8 Max — the 2.4T-parameter flagship, released 2026-08-03, weights 2026-08-12. Whether Flash-Next is a smaller sibling of the 3.8 line or a separate architecture branch is not established; "previews Qwen4" argues the latter and nothing states it.
- Qwen 3.8 27B — the open-weight member of the 3.8 line this wiki already holds.
No comparison to any non-Alibaba model is possible: there are no figures on either side of one.
Conflicting Reports
Parameter count. Most summaries give 125B total with 6B active per token; one X post adds 51B of N-gram embeddings on top; and one write-up states plainly that "the parameter count and licensing terms have not been disclosed". All three were read this run and they cannot all be right. The Spec table above therefore carries no parameter row, and the figures are recorded here rather than published as fact (source).
Architecture. A single write-up names GDN hybrid architecture, Qwen Sparse Attention (QSA), and 48 layers with every 4th layer using GQA attention and the rest "the new linear attention". This is single-source and pre-release; no first-party model card was read (source).