AI Trend Notifier
EN한
← wiki

$ cat wiki/entities/alibaba.md

Alibaba / Qwen AI Lab

Latest

  • 2026-09-30

    A second lab-level distillation accusation, against a different Chinese lab, gives this page's own the beginnings of a comparison class.

  • 2026-09-22

    Qwen 4 got four names and nothing else.

  • 2026-09-20

    Qwen's image line leaves Apache 2.0, and the wiki holds the licence text rather than a report of it.

Overview

Hangzhou-based Chinese technology conglomerate. In AI: principally known for the Qwen family of open-weight language models (Qwen3, Qwen-VL, etc.), developed by Alibaba's DAMO Academy and Qwen AI lab. One of China's most capable frontier AI research units. Qwen models are broadly open-weight and competitive with mid-tier US lab models.

AI Products & Research

  • Qwen-Image-2.1 — 2026-09-20; 7B visual generation component, native RGBA transparency, 2K output, up to 10 reference images. The first Qwen release recorded here under a non-commercial licence — Qwen Research License Agreement, not Apache 2.0
  • Qwen-Drive-1.0-4B — 2026-09-07, with Huazhong University of Science and Technology; a driving foundation model on the Qwen3.5-4B VLM unifying 3D perception, VQA and motion planning; Apache 2.0 on code, weights and demo data
  • Qwen-UI-Agent — 2026-07-29 technical report from Tongyi MAI; foundation GUI agent across mobile, computer, browser and DeepSearch → Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
  • Qwen 3.8 Max — generally available 2026-08-03, open weights 2026-08-12; 2.4T MoE / 95B active, 1M context, $2/$6 per M tokens; Arena.AI second globally on multimodal; released open as Qwen3.8-2.4T-A95B, the first Max-class Qwen opened, licence name still undisclosed
  • Qwen 3.8 27B — announced 2026-08-03, released 2026-08-12; the on-premise checkpoint, shipped two days after its stated date with architecture, context window and licence all still unpublished
  • Qwen3 — 2026 vintage open-weight frontier model family (dense + MoE variants); competitive with Mistral/Llama-class models on coding, reasoning, and multilingual tasks
  • Qwen-VL — vision-language variant
  • Research published under Alibaba DAMO Academy; Qwen AI lab is the primary model development unit

Recent Activity

  • 2026-09-30: A second lab-level distillation accusation, against a different Chinese lab, gives this page's own the beginnings of a comparison class. OpenAI disclosed a chain-of-thought extraction campaign attributed to individuals associated with Moonshot AI (source). Nothing in it concerns Alibaba and no figure on this page changes. It is recorded here because it is the first external reference point for GTG-16005: OpenAI's campaign runs July 1–28, ~4,000 then >15,000 users, 16,000 prompts at peak, and alleges an attempt with no model named as trained on the result; Anthropic's alleges >3,500 accounts, >151M exchanges, May–July, and transcripts that trained Qwen 3.5, 3.6 and 3.7. Neither lab published evidence. The two accusations differ by four orders of magnitude in volume and disagree about whether a transfer completed — so the GTG-16005 figures still stand alone and this page's unreconciled 25,000-accounts/28.8M-interactions versus 3,500/151M contradiction is untouched by it. New page: Adversarial Distillation

  • 2026-09-22: Qwen 4 got four names and nothing else. At the Apsara Conference in Hangzhou, Alibaba publicly named four Qwen 4 tiers for the first time — Qwen 4 Max (flagship), Qwen 4 Plus (balanced middle), Qwen 4 Flash (throughput) and Qwen 4 27B (open-weight). Qwen project lead Liu Dayiheng told the conference that Qwen 4 is in training on a "new-generation architecture" and will be released "very soon". None of the four has a release date, a price, a stated context window, an API identifier or downloadable weights, and no benchmark of any kind was presented — the tiers were shown as work in progress. Alibaba separately stated a roadmap of 5 to 10 trillion parameters for Qwen 4.5 and Qwen 5. Why it matters for this page: a named four-tier lineup with an open-weight member at the bottom is a commitment to keep opening something, three days after Qwen-Image-2.1 became the first Qwen release recorded here whose licence narrows — but the commitment is a name on a slide, and this page's standard is what shipped. The figure most likely to be misread is the parameter count: the 5–10T range describes Qwen 4.5 and Qwen 5, not Qwen 4, and coverage headlined "Alibaba unveils Qwen 4 and a 10 trillion parameter roadmap" collapses the two — both search passes attach it to the later generations. No model page was created, deliberately: a spec table of four unknown rows and a name is the padded stub this wiki has been trimming, and CLAUDE.md's rule is one page per model, not per series. The page arrives when a tier does. The only shipped architectural evidence is Qwen3.8-Flash-Next, open-sourced in late August 2026 and labelled in its own model card as a preview of the Qwen 4 architecture — GDN + QSA hybrid attention, Gated Residual connections, upgrades across attention, residual, embedding and optimization. No first-party read; surfaced by an r/LocalLLaMA prefetch candidate and confirmed across two search passes. From the rotation — Alibaba's slot fell on this run and, unlike 2026-09-20, the lab-named query returned the release. → Qwen3.8-Flash-Next, Qwen 3.8 Max, Open-Weights Policy Fight (source)

  • 2026-09-20: Qwen's image line leaves Apache 2.0, and the wiki holds the licence text rather than a report of it. Alibaba released Qwen-Image-2.1 — 7B visual generation component, 32 Single-Stream DiT layers, a Qwen3-VL 8B text encoder, a 64-channel RGBA autoencoder with 16× spatial compression, native 2K output, editing with up to 10 reference images, and native RGBA transparency — under the Qwen RESEARCH LICENSE AGREEMENT, dated September 20, 2026, licensed by Hangzhou Tongyi Laboratory Technology Co., Ltd. The grant is "FOR NON-COMMERCIAL PURPOSES ONLY", with "Non-Commercial" defined as "for research or evaluation purposes only" and commercial use routed to a separately negotiated licence at model-business@notice.qwencloud.com. Two passes state this is a change from the earlier Qwen-Image line, which shipped under Apache 2.0; one adds that the Apache-licensed 2512 and Edit-2511 models are unaffected. Why it matters for this page: every open-weights entry here has turned on Alibaba choosing permissive terms — three days ago this page recorded PrismML shipping Ternary Bonsai 2 27B under Apache 2.0 precisely because Alibaba's licence allowed it, and that is not possible under these terms. This is the first Qwen release on this wiki whose licence narrows rather than widens. The capability claim has nothing under it: the model is said to beat most closed models on Qwen-Image-Bench — Qwen's own benchmark — with no numeric figure in any pass and two passes stating independent benchmarks are still pending. No stated reason for the licence change appears anywhere read. Unusually for this page, two documents were read first-party — the repository README and the LICENSE file, both on github.com, which is reachable; qwen.ai, huggingface.co, the-decoder.com and blog.comfy.org all answer EGRESS_BLOCKED. Not from the rotation, though Alibaba's slot fell on the same run: the lab-named query returned nothing newer than Qwen 3.8-Max, and an r/LocalLLaMA prefetch candidate naming the version string surfaced it → Qwen-Image-2.1 (new), Open-Weights Policy Fight, Eval Harness Configuration (source) (GitHub) (LICENSE)

  • 2026-09-17: Somebody else shipped a 9× smaller Qwen, and Alibaba's licence is why they could. PrismML released Ternary Bonsai 2 27B, a ternary {−1, 0, +1} quantization of Qwen 3.8 27B at 1.76 effective bits per weight — 5.93 GB against 53.80 GB in FP16, same 262,144-token context, Apache 2.0, GGUF and MLX (3 passes). It claims 98.2% of the parent's aggregate over 20 benchmarks — 83.9 against 85.4 — with vision 96.3% and knowledge and reasoning 96.9% (source). Why it matters for this page: every entry here since June has been about Alibaba's models being taken without permission — the Senate-letter distillation campaign, GTG-16005, 151 million exchanges. This is the same models being taken with permission, by a third party who published the result under the licence Alibaba chose, and the wiki now holds both facts about the same weights. Two things this hands back to Alibaba's own page. First, PrismML's baseline column reads Terminal-Bench 2.1 = 69.7 for Qwen3.8 27B where Qwen 3.8 27B carries 73.0 from Alibaba's model card as described by outlets — a 3.3-point disagreement between the developer measuring its own model and a third party measuring it as a baseline, with no harness named on either side; disclosed on the Bonsai page's ## Conflicting Reports, not reconciled. Second, the first third-party measurement of a Qwen3.8 27B agentic figure this wiki holds arrives inside a competitor's release notes. No first-party read — prismml.com and huggingface.co both answer EGRESS_BLOCKED → PrismML, Ternary Bonsai 2 27B, Open-Weights Policy Fight, Eval Harness Configuration (source)

  • 2026-09-10: Anthropic puts a case number and much larger figures on the distillation accusation, and they do not match the figures it gave the Senate — Anthropic's September 2026 threat-intelligence report names GTG-16005, attributed to Alibaba and described as the largest distillation campaign Anthropic has measured: chain-of-thought distillation of Claude Opus 4.6 and 4.7, peaking at nearly 3 million exchanges per day from more than 3,500 fraudulent accounts, over 151 million exchanges between May and July 2026, with the harvested transcripts stated to have trained Qwen 3.5, 3.6 and 3.7. Alibaba is one of seven China-based labs Anthropic says it has disrupted for distillation since February 2026. The figures do not reconcile with the 2026-06-24 entry below — that letter alleged ~25,000 fraudulent accounts and 28.8 million interactions between 2026-04-22 and 2026-06-05; this gives 3,500 accounts and 151 million exchanges over May–July, an order of magnitude fewer accounts, roughly five times the volume, and an overlapping but different window. Nothing read says whether this is the same campaign re-measured, a successor, or a separate operation, and both records stand. Why it matters: the June accusation named a scale; this one names a case identifier, a method, a target model pair and three specific Qwen releases — a materially more specific claim, made in a routine threat report rather than a letter to legislators, and still with no response from Alibaba in anything read. → AI-Enabled Cyberattacks, Anthropic (source) (TechCrunch)

  • 2026-09-07: The first driving model on this wiki, and the licence covers more than the weights — Qwen released Qwen-Drive-1.0-4B with Huazhong University of Science and Technology, an autonomous-driving foundation model built on the Qwen3.5-4B VLM that adds a BEV perception head (3D detection, semantic occupancy, map segmentation) and a trajectory generator, shipping two planner variants — one imitation-trained, one further optimised by RL. The authors state it is the first VLM for driving, to their knowledge, unifying 3D perception, visual question answering and motion planning in one pretrained model. nuScenes: 43.95 mAP, 60.99 map mIoU, 42.83 NDS, beating multi-task BEVFormerV2* by 2.01 mAP and PETRv2 by 3.37 map mIoU. Apache 2.0 covers code, weights and demo data. Why it matters: only the perception half is measured — no planning benchmark appears in anything read — so the unification claim, which is the release's premise, is the part with no number under it.: the release sat in the 2026-09-10 run's prefetch list as an r/LocalLLaMA link and was skipped with the other Reddit items. → Qwen-Drive-1.0-4B (new), Embodied Agents, Open-Weights Policy Fight (source) (TechNode) (arXiv 2609.00111)

  • 2026-09-02: A dated snapshot, not a version bump — and it takes first place on a coding arena at an unchanged price — Alibaba shipped Qwen3.8-Max-0902, a post-training refresh of Qwen 3.8 Max on coding and Cowork-style tasks with no architecture change: 2.4T parameters, 1M context, thinking mode and every price identical ($2/M in · $6/M out; cache $0.25 implicit / $0.17 explicit read / $2.50 explicit create). It debuted first on Code Arena: WebDev at 1,691, ahead of Claude Opus 5. Why it matters: this is the second lab in two days to ship a coding gain as a post-training pass at an unchanged price — Gemini 3.8 Flash did the same on 09-02 — and Alibaba did it without even taking a new version number, which makes the release legible only to whoever is watching the snapshot dates. What is not established: the margin over Opus 5 is reported as three points, while the base model's own preliminary Arena interval is ±18 — wider than the margin, so the placement is recorded and the "beats Opus 5" framing is not. One pass's own headline is that the model remains "still behind Opus 5" overall while leading this one arena. Version hygiene: coverage carries Terminal-Bench 2.1 = 86.6 (base) and TerminalBench 3.0 = 11.3 → 29.0 (the 0902 delta) in the same articles — two different benchmarks, not a regression. Every figure but the Code Arena result rests on a single search pass; all hosts answer EGRESS_BLOCKED. Not from the rotation — Alibaba's slot was yesterday; this surfaced through a general release-tracker search and was captured anyway → Qwen 3.8 Max, Eval Harness Configuration (source) (TechNode)

  • 2026-08-31 (published), 2026-09-02 (captured): The architecture report behind Qwen3.8-Flash-Next, and it argues a method as much as a model — On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability is Alibaba's own account of the model this page recorded on 08-26. Against the 397B-A17B predecessor it leads on 8 of 14 pre-training benchmarks and trails on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens and roughly 1/9 the training FLOPs. The named components: Gated DeltaNet hybridised with global attention at one full-attention layer in four, those layers replaced at continued-pretraining time by Qwen Sparse Attention scoring context at micro-block granularity; a Gated Residual widening the residual stream to four branches; and the 51B n-gram embedding tables prefetched from host memory, adding capacity outside the backbone. Why it matters: it supplies the why behind a parameter split this page already held as a fact, and its methodological claim travels further than Qwen — every candidate change was judged on loss, cost and training stability jointly, on the strength of its own counterexample, that enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. What is not established: the 14 benchmarks are not named, so neither the 8 wins nor the 2.6-point bound can be located; every comparison quoted is pre-training, not post-trained behaviour; and no third party has reproduced any of it — the baseline is Alibaba's own predecessor. Reached through the HF Daily Papers snapshot, not the Chinese-lab rotation, which checked Alibaba this run and returned nothing → Qwen3.8-Flash-Next (source)

  • 2026-08-26: The countdown was real, the parameter conflict was a units problem, and the comparison column is the story — Qwen3.8-Flash-Next released on the date its ModelScope timer named, with weights on Hugging Face and ModelScope in BF16 and FP8, under qwen-community-1.0 — the first Qwen licence name this page has been able to record since Qwen 3.8 Max's weights shipped 2026-08-12, and it does not close that gap, since nothing says the flagship shares it. Context window 262,144 native, extensible to 1M. Yesterday's three-way parameter conflict resolves and was never a contradiction: 125B main model + 51B N-gram embedding table = 176B, the two reports describing different boundaries of one artefact, with 6B active per token. Vendor-stated benchmarks, all of them new (this page recorded "no benchmark figure of any kind" yesterday): SWE-bench Pro 62.5, SWE-bench Multilingual 81.0, CoWorkBench 73.9, JobBench 55.7, DeepSWE 1.1 58.7, LiveCodeBench v6 91.9, Toolathlon Verified 73.5. Why it matters: the paired rows are all against Claude Opus 4.6 Max — two minor versions behind Claude Opus 4.8 and three behind Claude Opus 5, neither of which appears anywhere in the table. A 6B-active model beating a two-versions-old frontier model is a genuine result and is not the result the framing invites; Qwen's reporting does not note the gap. No third party has measured this model at all, and the JobBench spread (55.7 vs 36.6, against 3.5–9.1 on every other paired row) is the row to hold loosest. Nothing read is first-party — every external host refused CONNECT this run → Qwen3.8-Flash-Next, GLM-5.3-Flash, Open-Weights Policy Fight (source)

  • 2026-08-25: A model staged for release tonight, offered as a preview of Qwen4's architecture rather than as a Qwen 3.8 sibling — a ModelScope teaser for Qwen3.8-Flash-Next went live carrying an "Upcoming Open-Release" badge and a countdown timer set to 2026-08-26 23:00 (UTC+08:00) — after this run. It is described as an open-weight multimodal MoE built on "the next-generation architecture that will power the upcoming Qwen4 family", and explicitly not Qwen4 itself, in standard and FP8 versions. Why it matters: every Qwen release this page holds has been a product; this is the first one framed as architecture disclosure ahead of a family, which is a different kind of move and gives the open-weight community a target to build tooling against before the flagship exists. The figures do not agree and are recorded as a conflict, not as fact: most summaries give 125B total / 6B active, one X post adds 51B of N-gram embeddings, and one write-up states plainly that the parameter count and licence "have not been disclosed". No benchmark figure of any kind exists in anything read. Licence unknown — the same gap Qwen 3.8 Max has carried since its weights shipped 2026-08-12. The release had not happened at capture, and this run cannot say whether it landed → Qwen3.8-Flash-Next, Open-Weights Policy Fight (source)

  • 2026-08-24: A method paper from an Alibaba org repository proposes replacing the standard RL regulariser — ERPO (arXiv 2608.23311, github.com/alibaba/ERPO) argues the stability–exploration trade-off in LLM policy optimisation is an artefact of applying the constraint on the action side, and moves it to the input side: a Query-KL term bounds drift in the distribution over training queries, with its gradient flowing strictly through the query likelihood so it exerts no direct pressure on the response distribution. Reported to replace Policy-KL in GRPO/PPO/REINFORCE pipelines with no additional forward passes, giving stronger accuracy and "substantially more stable behaviour under high-temperature decoding and long-horizon training" on six mathematical reasoning benchmarks. No numbers appear in the abstract. Why it matters: this page records Alibaba as a shipper of models; ERPO is the machinery, and it names a quantity — query-distribution drift under RL — that nothing in this wiki was measuring. Nothing read connects ERPO to any shipped Qwen model → Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311), Agentic Reinforcement Learning (source)

  • 2026-08-16 (captured 2026-08-24): Qwen is the most-run open model family on the platform that counts the runs, by a factor of two — Hugging Face's State of Open Models: Summer 2026 reports GGUF downloads per month of Qwen 39.6M, Gemma 20.8M and Llama 7.5M — Qwen at roughly 1.9× Gemma and 5.3× Llama — with 151,000+ derivatives built on Qwen models, more than any other family. Alibaba is also named, with Tencent, as one of only two labs covering the whole size range from under 1B upward, where Moonshot AI, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B (source). Why it matters: this page has recorded Alibaba's frontier releases — the 2.4T Max, the 27B, the licence terms, the slipped dates — and consistently found them measuring below the Western frontier on general composites (58 against Opus 5's 63, directly below). Download share says the frontier comparison is not where this lab is winning. A full-range publishing strategy compounds: the small models are what people actually run, running them is what produces the 151,000 derivatives, and the derivatives are what make the next release the default thing to fine-tune. That is a distribution position, and it is not visible in any benchmark this wiki tracks. The caveat is real: nothing read says whether GGUF pulls are deduplicated across quantisations, and a family shipping more quants accumulates more downloads for the same adoption. → Open-Weights Policy Fight (Hugging Face)

  • 2026-08-16: The opened checkpoint measures the same as the API it was cut from — Artificial Analysis lists both forms of Qwen 3.8 Max for the first time and scores them identically: Qwen3.8 Max 58 (1M context, $1.13/task) and Qwen3.8 2.4T A95B 58 (984k, $1.09/task). Neither row existed in the 08-09 snapshot. Why it matters: this page has tracked the terms of Alibaba's announce-then-open sequence — the 24-day window, the undisclosed licence, the five-day slip on the 27B — without ever being able to say whether the artefact handed over is the artefact that was sold. It is. What the same table also says is less flattering: 58 sits below every model this wiki calls frontier (Opus 5 max 63, GPT-5.6 Sol max 61, Grok 4.6 high 61), which does not support the July "second only to Fable 5" claim on a general measure. The 984k against 1M context gap is a property of whoever is serving the open weights, not of the weights. → Qwen 3.8 Max, Open-Weights Policy Fight (source)

  • 2026-08-14: Qwen 3.8 27B shipped, four days late, and the countdown was accurate to the hour — released at 15:00 UTC, which is the 2026-08-15 00:00 JST the ModelScope countdown pointed at. 27.78B dense parameters, text/images/video, Apache 2.0, 262,144-token native context reported extensible to ~1M via YaRN, on Hugging Face (Qwen/Qwen3.8-27B, -FP8) and ModelScope. Vendor-stated against Qwen3.6-27B: Terminal-Bench 2.1 63.4 → 73.0, DeepSWE 1.1 13.3 → 42.2, OSWorld-Verified 63.9 → 84.3, SWE-MM 25.7 → 38.6, plus SWE-Bench Pro 61.7%, CoWorkBench 70.7%, LiveCodeBench v6 90.3%, GPQA Diamond 89.2%. One harness is disclosed and it is a competitor's: SWE-MM is run on the Claude Code harness with modifications from Appendix 8.3 of the Claude Opus 4.7 system card. An r/LocalLLaMA thread 72 minutes after release claims the checkpoint is identical to Qwen3.6-27B; a targeted search found nothing behind it, and it is recorded on the model page as unresolved. Why it matters: this closes the eleven-day gap between an announcement and an artefact that this wiki tracked as five separate dates — and it closes it with the licence named, which is the row that was missing the whole time. Alibaba never restated, moved or withdrew the 2026-08-10 date in any first-party channel read here, before or after shipping. → Qwen 3.8 27B, Open-Weights Policy Fight, Eval Harness Configuration (source) (Kingy AI) (officechai)

  • 2026-08-13: The 27B did not ship — and this wiki said it had — reporting read on 2026-08-13 states that only the Max-class Qwen3.8-2.4T-A95B weights landed on 08-12, and that Qwen 3.8 27B has no repository, model card, licence file or benchmark. A ModelScope countdown now points at 2026-08-15, 00:00 JST — a third date after the announced 08-10 and the assumed 08-12. The 2026-08-13 run here recorded the 27B as released on 08-12, reading one drop as covering both checkpoints; that was wrong and the page carries the correction rather than a quiet edit. Why it matters: the announce-then-open sequence this wiki tracks on Open-Weights Policy Fight has now produced a five-day slip that Alibaba has never acknowledged in any first-party channel read here — the countdown page replaces the missed date without mentioning it, which is a different thing from moving it. → Qwen 3.8 27B, Qwen 3.8 Max (source) (Orca Router) (BigGo Finance)

  • 2026-08-12: The first Max-class Qwen opens — two days late, and without a licence name — Alibaba published open weights for Qwen3.8-2.4T-A95B, the open-weight form of Qwen 3.8 Max, and for Qwen 3.8 27B. The drop was scheduled in advance for 10:00 UTC+8 on 2026-08-12 (02:00 UTC) and lands two days after the 2026-08-10 date the 2026-08-03 announcement named — a slip this wiki recorded as a missed date on 2026-08-12 and now records as resolved. Repositories reported live include first-party BF16 and FP8 builds plus third-party GGUF and NVFP4 quantisations. Unsloth's figures give the scale: 4.9 TB at full precision, 2.6 TB at Q8_0, and 397 GB for the smallest 1-bit build. No licence is named in anything read — the weights are downloadable and the terms under which they may be used are not stated, with earlier Qwen lines' Apache-2.0 offered only as precedent. Why it matters: this is the first time Alibaba has opened a model at Max scale, and it completes a sequence this wiki has been tracking as a pattern rather than an event — closed preview 2026-07-19, paid API 2026-08-03, open weights 2026-08-12, twenty-four days from announcement to download. What it does not do is make the model runnable by the audience "open weights" usually implies: the smallest existing quantisation is still 397 GB. → Open-Weights Policy Fight (source) (Hugging Face) (Unsloth) (Hacker News)

  • 2026-08-03: Qwen3.8-Max generally available; Qwen3.8-27B announced; open weights dated — Alibaba moved Qwen3.8-Max out of the 2026-07-19 preview and made it generally available through Alibaba Cloud Model Studio (global) and QwenWork, its workplace agent platform. Figures withheld at preview are now disclosed: 95 billion active parameters of 2.4 trillion total, a 1 million token context window, and pricing at $2 / M input · $6 / M output. The active count is the one this wiki recorded on 2026-07-20 as "not disclosed — a critical gap for assessing true inference cost"; it took 15 days. First independent benchmark: Arena.AI places it second globally on multimodal tasks, behind only Claude Fable 5 — though as of that date it was listed on neither Artificial Analysis nor Hugging Face. A second checkpoint, Qwen3.8-27B, was announced the same day and positioned for on-premise deployment; open weights for both are stated for the week of 2026-08-10, which would be Alibaba's first open release at this scale. → Qwen 3.8 Max, Qwen 3.8 27B, Open-Weights Policy Fight (source)

  • 2026-07-29: Qwen-UI-Agent technical report — a foundation GUI agent reported to lead frontier models on mobile use — Tongyi MAI published arXiv:2607.28227, describing one model covering mobile, computer, browser and DeepSearch, trained against sandbox environments plus a large-scale real-device mobile runtime. Reported: 82.1% on MobileWorld (+14.6 over Opus 4.8, +12.0 over GPT-5.6 Sol, +8.9 over Seed 2.1 Pro), 92.2% on MobileWorld-Real, 97.5% on AndroidDaily, and 79.5% on OSWorld-Verified for second place overall on computer use. It also describes a harness for proactive service initiation — detecting an event such as a flight cancellation and acting on it rather than waiting to be invoked. Parameter count, architecture and license were not obtainable, and the open-weight availability is a report claim rather than a verified fact. → Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents (source)

  • 2026-07-19: Qwen 3.8 Max previewed — 2.4T MoE, open-weight release planned "soon" — The Qwen team posted on X that Qwen3.8-Max is "launching and going open-weight soon." Key facts: 2.4 trillion total parameters (MoE architecture; active count at inference not disclosed — a notable gap). Self-claim: "second only to Anthropic Fable 5" — based on internal evaluations only, no published independent benchmarks. Preview access live on Alibaba Cloud Token Plan and Qoder. Qwen 3.8 is previewed a few days after Moonshot AI's Kimi K3 (2.8T, July 16), making it the second-largest publicly previewed model. Why it matters: unlike Qwen 3.7-Max (which cited Terminal Bench, SWE-bench Pro), Qwen 3.8 arrived with no benchmark table — a step back in transparency. The claim of "second only to Fable 5" is unverifiable without independent evaluation. → Qwen 3.8 Max (source) (Bloomberg)

  • 2026-06-24 (disclosed): Accused of large-scale Claude model distillation campaign — Anthropic sent a letter to US Senate Banking Committee Chair Tim Scott, ranking member Elizabeth Warren, and White House officials alleging that operators linked to Alibaba's Qwen AI lab used approximately 25,000 fraudulent accounts to conduct 28.8 million interactions with Claude models between April 22 and June 5, 2026. Targeted capabilities: software engineering and agentic reasoning — Claude's highest-value use cases. Anthropic called it "the biggest attempt so far by a Chinese company to piggyback on the work of top US labs." Technical method: model distillation (using Claude's outputs as training signal for Qwen). The campaign overlapped with the US export-control suspension of Fable 5 and Mythos 5 (June 12), suggesting deliberate targeting of capabilities being restricted. → Anthropic, AI-Enabled Cyberattacks (source) (Bloomberg)

Strategic Position

  • China's frontier AI hedge: Qwen is the primary vehicle for China to maintain open-weight competitive parity with US labs without direct access to frontier model weights (which are restricted under US export-control rules)
  • Distillation as catch-up strategy: If the Anthropic accusations are accurate, API-based distillation represents a systematic approach to compressing the capability gap — using US frontier models as synthetic training teachers for Qwen
  • Policy implication: Export controls on model weights alone may be insufficient if API access is not similarly regulated; the Alibaba case is the first major public accusation of systematic distillation at this scale
  • Relationship with US labs: Adversarial on the capability/security axis (distillation accusations); complex on the research axis (Qwen models widely used in international research, including by US academics)
  • Anthropic — accused Alibaba of distillation campaign
  • AI-Enabled Cyberattacks — distillation as a form of AI capability theft
  • Claude Fable 5 — Fable 5 suspended June 12 (export control); distillation campaign ran April 22 – June 5
  • PrismML — publishes a ternary compression of Qwen 3.8 27B; the licensed counterpart to everything else on this page

Conflicting Reports

Alibaba's response (if any) not yet publicly disclosed as of 2026-06-26. The accusations are based on Anthropic's letter to US senators — unilateral attribution, not independently verified.

Referenced by

2026-W392026-W40Adversarial DistillationAgentic Reinforcement LearningAgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634)AI AlignmentAI GovernanceAI-Enabled CyberattacksAnt Group (inclusionAI / AntLing)AnthropicBeyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311)DeepSeekEvery Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647)Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber ModelsMeituanMing-Image-0.1-DesignMiniMaxMoonshot AIOn the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training StabilityOpen-Weights Policy FightPost-Training Leaves Behavioral Shadows on Unrelated DecisionsPost-Training ScalingPrismMLQwen 3.8 27BQwen 3.8 MaxQwen-Drive-1.0-4BQwen-Image-2.1Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsQwen3.8-Flash-NextRufus-Air: An Open LLM Post-Training RecipeSeptember 2026 — Monthly DigestTencentThinking Machines LabTraining Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report (arXiv:2608.15763)Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929)Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved MechanismsWeekly Synthesis — 2026-W26 (2026-06-22 ~ 2026-06-28)Weekly Synthesis — 2026-W27 (2026-06-29 ~ 2026-07-05)Weekly Synthesis — 2026-W28 (2026-07-06 ~ 2026-07-12)Weekly Synthesis — W35 (2026-08-24 → 2026-08-30)Weekly Synthesis — W37 (2026-09-07 → 2026-09-13)World ModelsZ.ai

Sources