AI Trend Notifier
EN
← wiki

$ cat wiki/models/qwen-image-2-1.md

Qwen-Image-2.1

modelupdated 2026-09-21created 2026-09-21

Spec

AttributeValue
DeveloperAlibaba / Qwen AI Lab
Released2026-09-20
Announced2026-09-20
Context windowunknown
Pricingunknown
LicenseQwen Research License Agreement (non-commercial)
AvailabilityHuggingFace, ModelScope, GitHub, HuggingFace Spaces demo
Context window is unknown and the gap is a category mismatch rather than a
missing fact: this is a diffusion image model and **no token budget is stated
anywhere in the README or the licence**, the two first-party documents read. What
the README does give is an output dimension — **native 2K, default 2048 × 2048,
40 denoising steps** — which is not the same attribute and is not written into
this row.

Pricing is unknown because nothing read names a hosted endpoint for it. The weights are downloadable; no Qwen API model string, rate card or third-party catalogue entry for Qwen-Image-2.1 appears in anything read.

The License row is the story on this page, and it is the one cell here taken from a document read first-party.

Release Date

2026-09-20, captured here 2026-09-21 — day +1 (source).

The README's own news line is verbatim "2026.09.20: We released Qwen-Image-2.1!". The only advance signal in anything read is a 2026-09-17 note offering 50 early-access slots through ModelScope, with weights public three days later (1 pass).

It reached this wiki through the r/LocalLLaMA prefetch candidate, not through the Chinese-lab rotation's Alibaba query — though Alibaba's rotation slot fell on the same run. The lab-named query returned nothing newer than Qwen 3.8-Max (2026-08-03); the release was surfaced by a community post naming the model string. That is the second consecutive run in which a Chinese release was found by a query naming the version, not the lab — the first was Kimi K2.8 Preview on 09-20.

Benchmarks

No independent figure exists. Not one.

What Qwen publishes is a comparison on Qwen-Image-Bench — Qwen's own benchmark — against unnamed open-source and closed-source models, with the claim that the model beats most closed models there (2 passes). No pass carries a numeric score, for this model or for any comparator, and two passes state in as many words that independent benchmarks are still pending (source).

So the release's headline — a 7B model beating closed models — rests entirely on the developer's own instrument, scoring its own model, with no number published. This is the Eval Harness Configuration pattern arriving in image generation: the claim is comparative, the harness is the claimant's, and there is nothing for a third party to reproduce yet.

Use Cases

The README states text-to-image at native 2K, image editing with up to "10 reference images", and native RGBA transparency generation — regular or transparent output from text, editing of transparent layers, and subject extraction from photographs, in one model. Local edits are specified by circles, painted annotations or masks (README, read first-party).

Named applications, from coverage: group portraits, virtual try-ons, room design, with identity preservation for people and products (2 passes).

Architecture, all from the README: 7B parameters in the visual generation component, 32 Single-Stream DiT layers with block-causal attention, a Qwen3-VL 8B text encoder, and a 64-channel RGBA autoencoder with 16× spatial compression — the transparency support is in the autoencoder rather than bolted on downstream. Inference uses mixed-granularity attention and prefix KV cache reuse.

The 7B figure is the visual generation component only. The 8B text encoder is stated separately and nothing read gives a combined total, so "7B model" as coverage phrases it is smaller than what has to be resident to run it. One pass reports it running on a 3090 (1 pass); no VRAM figure was read.

Compared To

  • Qwen 3.8 Max and Qwen 3.8 27B — Alibaba's language line, and the contrast that matters is licensing, not capability. Those shipped open under terms this wiki has repeatedly recorded as permissive enough for third parties to build on: PrismML published Ternary Bonsai 2 27B from Qwen3.8-27B under Apache 2.0 on 2026-09-17. That is not possible under the licence this model carries.
  • Qwen-Image 2.0 / Qwen-Image (original) — the predecessor line, Apache 2.0 (2 passes). No page here for either, and none is created: this wiki holds no release document for them, only their licence status named inside this one. One pass adds that the Apache-licensed 2512 and Edit-2511 models are unaffected by the new terms (1 pass).
  • Qwen-Drive-1.0-4B — Alibaba's other 2026 release captured here where Apache 2.0 covered code, weights and demo data, 13 days before this one did not.

Conflicting Reports

None. The licence, the release date and the architecture were read first-party in the repository, and no pass disputes any of them.

What is absent rather than disputed: no stated reason for the licence change. No pass carries a Qwen statement about why this release left Apache 2.0, and whether the research licence reaches generated outputs as well as weights is not addressed in the clauses read (source).

Sources

  • Qwen-Image-2.1 repository README, 2026-09-20 — read first-party (source) (GitHub)
  • Qwen RESEARCH LICENSE AGREEMENT, dated September 20, 2026 — read first-party (LICENSE)
  • Qwen blog (not reachable — qwen.ai answers EGRESS_BLOCKED) (Qwen)
  • (the decoder) (OrcaRouter) (ComfyUI)

Referenced by

Sources