AI Trend Notifier
EN
← trends

$ cat wiki/trends/2026-W35.md

Weekly Synthesis — W35 (2026-08-24 → 2026-08-30)

trendupdated 2026-08-30created 2026-08-30

Weekly Synthesis — W35 (2026-08-24 → 2026-08-30)

Synthesized August 30, 2026 · Covers Monday August 24 through Sunday August 30 · Weekly Synthesis — W34 (2026-08-17 → 2026-08-23) ← → next week

Period

Weekly Synthesis — W34 (2026-08-17 → 2026-08-23) was the week the field turned its scepticism on its own instruments. W35 is the week the licence became the interesting part of a model release — because four frontier-scale open-weight models shipped in fifteen days and their weights are far more alike than their terms.

That is a genuine change in what a release is. For most of this wiki's life an open-weight announcement carried one decision: publish or don't. This week carried four different answers to a question nobody was asking in June — on what conditions — and the answers do not sort by capability, by country, or by how much the lab says it worries about safety.

Notable Releases

  • Hy4 preview (Tencent, 2026-08-28) — 770B total / 49B active MoE, context over 1M, Apache 2.0 on day one for both the BF16 checkpoint and an FP8 quantisation. The largest open-weight model this wiki holds, and the only one of the four with no condition attached at all. Tencent had no page here before Friday.
  • GLM-5.3 (Z.ai, weights 2026-08-28) — the two-week safety gate closed on the day it promised, and the licence turned out to be the second condition: an MIT-shaped grant plus a security review for any Model-as-a-Service operator above $10 billion in group revenue. The card volunteered why — "cyber capability developed faster than we expected", more than doubling GLM-5.2 on exploitation benchmarks.
  • GLM-5.3-Flash (2026-08-26) — the same lab, MIT, no gate, two days earlier.
  • Qwen3.8-Flash-Next (Alibaba / Qwen AI Lab, 2026-08-26) — 176B/6B active under qwen-community-1.0, released explicitly as an architecture preview of Qwen4 rather than as a product.
  • MHS — Model Hardware Standard (Anthropic, 2026-08-27) — a proposed specification for agents to operate physical laboratory devices, begun with HHMI Janelia and Genentech. A second protocol from the lab that proposed MCP — Model Context Protocol, and nothing of it is published yet.

Emerging Themes

The licence is now the release decision, and capability does not predict it

Set the four side by side and the ordering is the story:

ModelSizeLicenceCondition
Hy4 preview770B / 49BApache-2.0none
GLM-5.3-Flash320B / 18BMITnone
Qwen3.8-Flash-Next176B / 6Bqwen-community-1.0licence terms not read here
GLM-5.3744B / 40Bglm-5.3security review above $10B revenue
**The biggest model carries the loosest terms, and the two tightest belong to the
lab that published the most detailed safety reasoning.** Every argument for
attaching a condition scales with capability, and the week's largest artefact was
published flat. That is not evidence the conditions are unnecessary — nobody
outside Tencent has measured Hy4 preview at all — but it does mean the field has
no shared function from capability to terms, which is the thing a norm would be.

Anthropic published both halves of the automation boundary, days apart

Automated Researchers Can Reliably Mitigate Alignment Failures ran five Opus 4.8 agents that beat 28 human researchers on all seven failures the humans attempted, with human-written directions giving no benefit. Then TASTE (day +1) measured the excluded case — judging safety-research proposals — where Fable 5 leads at 60% and Opus 5 and GPT-5.6-Sol sit near chance against 77% human agreement. Neither post claims the pair, and it is the pair that is informative: the first task was chosen because "an objective benchmark, not a fallible human, decides whether a fix works", and the second is what happens when a human decides. Recorded on AI Alignment as a reading rather than a finding.

The measurement problem moved up a level, from harness to leaderboard

W34 established that a benchmark number moves with the harness around it. W35 added the layer above: who decides what the average is over. Hugging Face put Indian-English into the Open ASR Leaderboard's default column set, so it now contributes to the headline Average WER of every listed model — a re-scoring of the whole board with nothing retrained and nothing resubmitted, and no before/after table published. On the other side, Z.ai shipped GLM-5.3's card with a harness named for almost every row — Claude Code 2.1.207, mini-swe-agent, sampling parameters, turn caps, container policy — which is the remedy Eval Harness Configuration proposed, adopted by a vendor unprompted, three weeks after the same page recorded that same model's figures as having no harness published at all.

Model access became a contract question

OpenAI ended Cursor's direct access to its models effective 2026-11-12, stating it cannot be confident SpaceX will honour its terms of service. The gates this wiki tracks — GPT-5.6-Cyber's vetted tier, Astra's Critical designation — restrict what a model may do. This restricts who may resell it, and it is the first of its kind held here. Cursor's CEO puts the commercial exposure at about 5% of traffic, which suggests the precedent matters more than the revenue.

Declining Themes

  • The paper cluster ran out. W34 and early W35 were dominated by harness and post-training papers arriving four and five at a time. Today's HuggingFace Daily snapshot carried zero new arXiv ids — 25 entries differing from yesterday's by one swap, both sides already held. A pause, not a drought, but the first day in weeks with nothing to pick.
  • "Which model is best" barely moved. LMArena's top 10 is byte-identical across the week's captures and every Artificial Analysis Intelligence Index in the top 30 is unchanged. Four open-weight releases and the leaderboards did not notice — because none of the four is on them yet.

Surprising Results

  • A lab with no page here shipped the biggest open model of the year so far. Tencent Hunyuan was, to this wiki, a single leaderboard row (Hy3, Index 42) until Friday. It was listed nowhere in sources.yaml, is not on the Chinese-lab rotation, and the release was caught by a Reddit thread.
  • Hy4 preview's attention module is Gated DeepSeek Sparse Attention. A Tencent frontier model built on a mechanism a competitor published, named as such. This wiki has tracked open weights as a distribution question; here the reuse is at the level of the design.
  • Maintainers report the exploit arriving before the patch. Anil Madhavapeddy reports exploit attempts within minutes of a patch being discussed, from the rumour of a bug rather than the diff; rclone went from ~20 disclosures in ten years to 40+ in a month; the summary claim is mean time to exploit −7 days. No methodology, no model named, and disclosures are not exploits — carried on AI-Enabled Cyberattacks as the first item there about ordinary maintenance rather than a controlled evaluation.

Open Debates

  • Does Hy4 preview actually beat DeepSeek V4-Pro on Terminal Bench 2.1? Tencent reports 85.4 as surpassing it; this wiki holds DeepSeek's own 87.9. Both vendor-stated, neither with a published harness. Unresolved by design.
  • What is a 2× price factor doing across four vendors? spec-check reports eight figures off by exactly two, at four different labs. Four simultaneous price changes is the unlikelier explanation; a units convention on one side is not yet established either.
  • Is a preview an architecture disclosure or a product? Qwen3.8-Flash-Next was released to let the ecosystem adapt tooling ahead of Qwen4. It gained its first independent measurement this week (Artificial Analysis Intelligence Index 56), which is what a product gets.

Outlook

The four-licence spread is the thing to watch into September, and specifically whether anyone measures Hy4 preview. A 770B Apache-2.0 model with no third-party evaluation is the largest unexamined artefact this wiki has ever listed, and the two questions it raises — is the capability there, and does an unconditioned licence on that capability matter — cannot be separated until someone outside Tencent runs it.

Second, watch whether GLM-5.3's Cybersecurity Trusted Access tier survives its own weight release. A verified-access control on offensive capability, applied to a model anyone can now download, is a control with no obvious remaining mechanism, and nothing in the licence or the repository mentions it.