$ cat wiki/models/qwen-3-8-flash-next.md
Qwen3.8-Flash-Next
Spec
| Attribute | Value |
|---|---|
| Developer | Alibaba / Qwen AI Lab |
| Released | 2026-08-26 |
| Announced | 2026-08-25 |
| Context window | 262,144 native, extensible to 1M |
| Pricing | unknown |
| License | qwen-community-1.0 |
| Availability | Hugging Face and ModelScope, BF16 and FP8 |
| The release landed on the date the countdown named. Yesterday this page read | |
Released: not yet against a ModelScope timer pointed at 2026-08-26 23:00 | |
| UTC+08:00; the weights are now on Hugging Face and ModelScope in BF16 and FP8 | |
| (source). |
Pricing stays unknown: nothing read quotes a per-token rate, and open weights
do not supply one.
License is no longer unknown — and it does not close the gap on
Qwen 3.8 Max, whose licence name has been unknown since its weights
shipped 2026-08-12. qwen-community-1.0 is this model's licence; nothing read
says the flagship ships under the same one.
Parameters: 125B main model + 51B N-gram embeddings, 6B active per token. Weights on disk are "roughly 180B parameters across the language model, n-gram embedding, and multi-token prediction layer"; serving documentation aggregates the first two as 176B (source).
Release Date
Announced 2026-08-25, released 2026-08-26.
Positioned as an open-weight, multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself — early access to the architecture ahead of the full family (source).
Benchmarks
An independent measurement exists as of 2026-08-30, and it is one number. Artificial Analysis lists Qwen3.8-Flash-Next at Intelligence Index 56, context 256k, Cost per Task $0.10 — a composite of their benchmark suite, not an accuracy, and a measured cost of one task, not a per-token price. The row is absent from the 2026-08-23 capture and present in the 2026-08-29 and 2026-08-30 ones, so it first appeared in that window (source).
What it settles and what it does not. It settles that the model has been measured by someone other than Alibaba, which is what this page said had not happened. It does not check a single row of the vendor table below — Artificial Analysis publishes a composite, not SWE-bench Pro or JobBench — so the 19.1-point JobBench margin flagged below is still unread by anyone outside Alibaba. At 56 it sits one point under GLM-5.3-Flash (57), the release it shares a day with, on the one scale both now appear on.
Context window needs no change: 256k is Artificial Analysis's rendering of
the 262,144 already in the Spec table.
The vendor table below still has no independent read of any row (source):
| Benchmark | Qwen3.8-Flash-Next | Claude Opus 4.6 Max |
|---|---|---|
| SWE-bench Pro | 62.5 | 53.4 |
| SWE-bench Multilingual | 81.0 | 77.5 |
| CoWorkBench | 73.9 | 68.2 |
| JobBench | 55.7 | 36.6 |
| DeepSWE 1.1 | 58.7 | not stated |
| LiveCodeBench v6 | 91.9 | not stated |
| Toolathlon Verified | 73.5 | not stated |
| Read the comparison column before the score column. Every paired row is against | ||
| Claude Opus 4.6 Max — a model two minor versions behind | ||
| Claude Opus 4.8 and three behind Claude Opus 5, neither of | ||
| which appears. A 6B-active model beating a frontier model from two versions ago is | ||
| a real result; it is not the result the headline framing invites, and Qwen's | ||
| reporting does not note the gap. |
The JobBench margin (55.7 against 36.6, a 19.1-point spread where the other paired rows sit at 3.5–9.1) is the row most worth an independent read before it is quoted anywhere.
Use Cases
Agentic coding and tool use, on the evidence of what was benchmarked: every published row is SWE-bench, DeepSWE, LiveCodeBench, CoWorkBench, JobBench or Toolathlon. Vision input is a stated capability of the model card with no benchmark attached (source).
The 6B active parameters are the point of the design — the team frames the configuration as a step toward "ultimate cost efficiency", and NVIDIA published an agentic-coding write-up on GB300 NVL72 the same day. The r/LocalLLaMA threads of 2026-08-25 that anticipated a local-friendly architecture, recorded here yesterday as community expectation, now have weights to test against; nothing read measures local throughput, so that remains an expectation (source).
Compared To
2026-08-31 — the architecture report arrives, and with it the baseline the release documentation never gave. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability is Alibaba / Qwen AI Lab's own account of this model. Against the 397B-A17B predecessor it leads on 8 of 14 pre-training benchmarks and trails on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens and roughly 1/9 the training FLOPs. The components are named: Gated DeltaNet hybridised with global attention at one full-attention layer in four, replaced at continued-pretraining time by Qwen Sparse Attention scoring context at micro-block granularity; a Gated Residual widening the residual stream to four branches; and the 51B n-gram embedding tables prefetched from host memory — which is why this page's parameter split has capacity sitting outside the backbone. The 14 benchmarks are not named, every comparison is pre-training rather than post-trained, and the baseline is Alibaba's own predecessor (source).
- Qwen 3.8 Max — the 2.4T-parameter flagship, released 2026-08-03, weights 2026-08-12. Whether Flash-Next is a smaller sibling of the 3.8 line or a separate architecture branch is not established; "previews Qwen4" argues the latter and nothing states it.
- Qwen 3.8 27B — the open-weight member of the 3.8 line this wiki already holds.
- GLM-5.3-Flash — released the same day, and the comparison the day
actually offers: another Chinese open-weight multimodal MoE sold on cost, at
320B/18B active against this model's 176B/6B, MIT against
qwen-community-1.0, 1M context against 262K native. Both now carry an Artificial Analysis Intelligence Index — GLM-5.3-Flash 57, Qwen3.8-Flash-Next 56 — which as of 2026-08-30 is the only figure the two share; no vendor benchmark is common to both, so the one-point gap is the whole of what can be compared (source). - Hy4 preview — the third open-weight Chinese release of the fortnight
and the opposite design point: 770B/49B active under Apache 2.0 against this
model's 176B/6B under
qwen-community-1.0. It has no independent measurement, where this model now has one.
Comparison to a non-Alibaba model on the vendor's own rows is still possible in one direction only: Qwen's table places it above Claude Opus 4.6 Max on four benchmarks. This wiki holds no figure for that model on any of them, so those rows cannot be checked here.
Conflicting Reports
Parameter count — resolved 2026-08-27, and it was never a contradiction. Yesterday this page recorded three irreconcilable figures: 125B total, 125B plus 51B of N-gram embeddings, and "not disclosed". The release names the components, and they add up: 125B main model + 51B N-gram embedding table = 176B, which is the figure serving documentation and NVIDIA's write-up both use. The 125B and 176B reports were describing different boundaries of the same artefact — the transformer, and the whole thing — and the third report simply predated disclosure. Both figures now appear in the Spec table with their boundaries named (source).
Recorded rather than deleted, because the lesson survives the resolution: three sources disagreeing about a parameter count is as often a units problem as a factual one, and holding the row empty for a day cost nothing.
Architecture. A single write-up names GDN hybrid architecture, Qwen Sparse Attention (QSA), and 48 layers with every 4th layer using GQA attention and the rest "the new linear attention". This is single-source and pre-release; no first-party model card was read (source).