AI Trend Notifier
EN한
← wiki

$ cat wiki/models/qwen-3-8-flash-next.md

Qwen3.8-Flash-Next

modelupdated 2026-09-02created 2026-08-26

Spec

AttributeValue
DeveloperAlibaba / Qwen AI Lab
Released2026-08-26
Announced2026-08-25
Context window262,144 native, extensible to 1M
Pricingunknown
Licenseqwen-community-1.0
AvailabilityHugging Face and ModelScope, BF16 and FP8
The release landed on the date the countdown named. Yesterday this page read
Released: not yet against a ModelScope timer pointed at 2026-08-26 23:00
UTC+08:00; the weights are now on Hugging Face and ModelScope in BF16 and FP8
(source).

Pricing stays unknown: nothing read quotes a per-token rate, and open weights do not supply one.

License is no longer unknown — and it does not close the gap on Qwen 3.8 Max, whose licence name has been unknown since its weights shipped 2026-08-12. qwen-community-1.0 is this model's licence; nothing read says the flagship ships under the same one.

Parameters: 125B main model + 51B N-gram embeddings, 6B active per token. Weights on disk are "roughly 180B parameters across the language model, n-gram embedding, and multi-token prediction layer"; serving documentation aggregates the first two as 176B (source).

Release Date

Announced 2026-08-25, released 2026-08-26.

Positioned as an open-weight, multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself — early access to the architecture ahead of the full family (source).

Benchmarks

An independent measurement exists as of 2026-08-30, and it is one number. Artificial Analysis lists Qwen3.8-Flash-Next at Intelligence Index 56, context 256k, Cost per Task $0.10 — a composite of their benchmark suite, not an accuracy, and a measured cost of one task, not a per-token price. The row is absent from the 2026-08-23 capture and present in the 2026-08-29 and 2026-08-30 ones, so it first appeared in that window (source).

What it settles and what it does not. It settles that the model has been measured by someone other than Alibaba, which is what this page said had not happened. It does not check a single row of the vendor table below — Artificial Analysis publishes a composite, not SWE-bench Pro or JobBench — so the 19.1-point JobBench margin flagged below is still unread by anyone outside Alibaba. At 56 it sits one point under GLM-5.3-Flash (57), the release it shares a day with, on the one scale both now appear on.

Context window needs no change: 256k is Artificial Analysis's rendering of the 262,144 already in the Spec table.

The vendor table below still has no independent read of any row (source):

BenchmarkQwen3.8-Flash-NextClaude Opus 4.6 Max
SWE-bench Pro62.553.4
SWE-bench Multilingual81.077.5
CoWorkBench73.968.2
JobBench55.736.6
DeepSWE 1.158.7not stated
LiveCodeBench v691.9not stated
Toolathlon Verified73.5not stated
Read the comparison column before the score column. Every paired row is against
Claude Opus 4.6 Max — a model two minor versions behind
Claude Opus 4.8 and three behind Claude Opus 5, neither of
which appears. A 6B-active model beating a frontier model from two versions ago is
a real result; it is not the result the headline framing invites, and Qwen's
reporting does not note the gap.

The JobBench margin (55.7 against 36.6, a 19.1-point spread where the other paired rows sit at 3.5–9.1) is the row most worth an independent read before it is quoted anywhere.

Use Cases

Agentic coding and tool use, on the evidence of what was benchmarked: every published row is SWE-bench, DeepSWE, LiveCodeBench, CoWorkBench, JobBench or Toolathlon. Vision input is a stated capability of the model card with no benchmark attached (source).

The 6B active parameters are the point of the design — the team frames the configuration as a step toward "ultimate cost efficiency", and NVIDIA published an agentic-coding write-up on GB300 NVL72 the same day. The r/LocalLLaMA threads of 2026-08-25 that anticipated a local-friendly architecture, recorded here yesterday as community expectation, now have weights to test against; nothing read measures local throughput, so that remains an expectation (source).

Compared To

2026-08-31 — the architecture report arrives, and with it the baseline the release documentation never gave. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability is Alibaba / Qwen AI Lab's own account of this model. Against the 397B-A17B predecessor it leads on 8 of 14 pre-training benchmarks and trails on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens and roughly 1/9 the training FLOPs. The components are named: Gated DeltaNet hybridised with global attention at one full-attention layer in four, replaced at continued-pretraining time by Qwen Sparse Attention scoring context at micro-block granularity; a Gated Residual widening the residual stream to four branches; and the 51B n-gram embedding tables prefetched from host memory — which is why this page's parameter split has capacity sitting outside the backbone. The 14 benchmarks are not named, every comparison is pre-training rather than post-trained, and the baseline is Alibaba's own predecessor (source).

  • Qwen 3.8 Max — the 2.4T-parameter flagship, released 2026-08-03, weights 2026-08-12. Whether Flash-Next is a smaller sibling of the 3.8 line or a separate architecture branch is not established; "previews Qwen4" argues the latter and nothing states it.
  • Qwen 3.8 27B — the open-weight member of the 3.8 line this wiki already holds.
  • GLM-5.3-Flash — released the same day, and the comparison the day actually offers: another Chinese open-weight multimodal MoE sold on cost, at 320B/18B active against this model's 176B/6B, MIT against qwen-community-1.0, 1M context against 262K native. Both now carry an Artificial Analysis Intelligence Index — GLM-5.3-Flash 57, Qwen3.8-Flash-Next 56 — which as of 2026-08-30 is the only figure the two share; no vendor benchmark is common to both, so the one-point gap is the whole of what can be compared (source).
  • Hy4 preview — the third open-weight Chinese release of the fortnight and the opposite design point: 770B/49B active under Apache 2.0 against this model's 176B/6B under qwen-community-1.0. It has no independent measurement, where this model now has one.

Comparison to a non-Alibaba model on the vendor's own rows is still possible in one direction only: Qwen's table places it above Claude Opus 4.6 Max on four benchmarks. This wiki holds no figure for that model on any of them, so those rows cannot be checked here.

Conflicting Reports

Parameter count — resolved 2026-08-27, and it was never a contradiction. Yesterday this page recorded three irreconcilable figures: 125B total, 125B plus 51B of N-gram embeddings, and "not disclosed". The release names the components, and they add up: 125B main model + 51B N-gram embedding table = 176B, which is the figure serving documentation and NVIDIA's write-up both use. The 125B and 176B reports were describing different boundaries of the same artefact — the transformer, and the whole thing — and the third report simply predated disclosure. Both figures now appear in the Spec table with their boundaries named (source).

Recorded rather than deleted, because the lesson survives the resolution: three sources disagreeing about a parameter count is as often a units problem as a factual one, and holding the row empty for a day cost nothing.

Architecture. A single write-up names GDN hybrid architecture, Qwen Sparse Attention (QSA), and 48 layers with every 4th layer using GQA attention and the rest "the new linear attention". This is single-source and pre-release; no first-party model card was read (source).

Referenced by

Sources