$ cat briefs/daily/2026-09-03.md
2026-09-03
September 3, 2026 (Thu)
3 stories · 3 paper picks · 3 watch items · 5 new pages
> **Three labs restricted a release on the same capability in five days, and no > two of them restricted the same thing.** Z.ai attached a licence condition, > OpenAI withheld the capability, Google withheld access. The convergence on the > trigger — autonomous vulnerability discovery — was already on file. The > divergence on the remedy is new, and it is the more useful half. > **Google's own release table split its two coding benchmarks by 9.2 points to > 1.2.** Terminal-Bench 2.1 went 81.6% → 90.8% while SWE-bench Pro went 60.4% → > 61.6%, in one release, on one model. One benchmark scores an agent finishing a > task; the other scores a patch. This wiki has been arguing that difference > across labs for five weeks; today a vendor published it inside a single table. > **The snapshot Action missed its slot for an eighth consecutive day.** No run > fired at the 22:20 UTC papers cron. A manual dispatch from this run wrote all > three of today's files. Sixth hand dispatch; carried to the W36 lint as the > standing intake defect.
Top Stories
1. Gemini 3.8 Flash — 20 days on, an identical spec table and a benchmark column that moved in only one place (score 1.93)
- Released 2026-09-02. Context window 1,048,576, max output 65,536, $0.75/M input · $3.75/M output introductory through 2026-12-31 then $1.50 · $7.50, GA as
gemini-3.8-flash. Every one of those is identical to Gemini 3.7 Flash, 20 days earlier (source). - Terminal-Bench 2.1 81.6% → 90.8%, DeepSWE v1.1 73.7%, SWE-bench Pro 60.4% → 61.6%. Three Flash releases in 43 days, none of which changed a capacity figure or a price.
- Attestation differs by row and the page says so:
deepmind.googleandblog.googleare both blocked here, so this was captured across three WebSearch passes. Terminal-Bench and DeepSWE carry two agreeing passes; SWE-bench Pro's digits come from one, with two more corroborating only the direction. - Why it matters: 9.2 points on the benchmark that scores an agent completing a task end to end, against 1.2 on the benchmark that scores a patch, is a release that moved the harness-shaped half of coding and left the other alone — the clearest instance yet of this wiki's central eval argument, published inside a single vendor's own comparison.
- → Gemini 3.8 Flash · Eval Harness Configuration
2. Three labs, one trigger, three different gates — and only one of them can be checked from outside (score 1.93)
- 2026-08-28 GLM-5.3 withholds nothing and attaches a licence condition: security review above a $10B revenue threshold. 2026-09-01 Astra withholds the capability. 2026-09-02 Gemini 3.8 Flash Cyber withholds access — an approval list, no self-serve API, no public price (source).
- Fairwind's stated eligibility is trusted government and national cyber authorities, critical-infrastructure operators and software maintainers. The approval criteria are not published — "trusted" and "approved" are the whole of the stated test.
- The Cyber model published no CyberGym score, only the claim that it surpasses Gemini 3.5 Flash Cyber — which six weeks earlier published 55 / 47 / 36 Chrome V8 vulnerabilities. The comparison cannot be checked in the direction it is offered (source).
- Why it matters: a licence is a document and the weights are downloadable, so Z.ai's gate is auditable by anyone; a capability withheld inside a model and an unpublished approval list are not. Three labs agreed on what is dangerous and disagreed entirely on what to do about it, which is the shape a governance regime has before it has any rules.
- → AI-Enabled Cyberattacks · Gemini 3.8 Flash Cyber
3. Qwen3.8-Max-0902 — a coding gain shipped as a dated snapshot, at an unchanged price and without a version number (score 1.20)
- Live 2026-09-01 22:00 ET. Post-training on coding and Cowork-style tasks, no architecture change; 2.4T parameters, 1M context and every price identical to the base model (source).
- Debuted first on Code Arena: WebDev at 1,691, ahead of Claude Opus 5.
- The "beats Opus 5" framing is not carried here. One pass gives the margin as three points; the base model's own preliminary Arena interval is ±18 — wider than the margin. That same pass's headline says the model is "still behind Opus 5" overall.
- Why it matters: second lab in two days to ship a coding gain as a post-training pass at an unchanged price — and Alibaba did it without taking a new version number, which makes the release legible only to whoever is watching snapshot dates.
- → Qwen 3.8 Max · Alibaba / Qwen AI Lab
Paper Picks
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling — arXiv 2608.26623 (score 1.75)
- TL;DR: 3,808 instances over six workflow-DAG topologies and three difficulty tiers, six judges from 20B to frontier scale, paired with and without ground truth. On hard queries without ground truth, all six converge into a 77–82% alignment band regardless of scale.
- Ground-truth exposure reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp). Chain-of-thought and temperature are negligible; rubrics give up to +6.5 pp but do not generalise across judge–generator pairs.
- Why read it: it puts a hard ceiling on a harness component everyone treats as a fixed point. Where agent results are judged by an LLM on hard tasks, the reported gap between two systems can be smaller than the judge's own error.
- → AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement — arXiv 2609.01481 (score 1.75)
- TL;DR: a planning–coding–testing loop layered above three existing harnesses — Codex/GPT-5.5, OpenCode/DeepSeek-V4-Pro, Pi/MiniMax-M3 — improves all three. +52.25% average relative, max 82.86% after three iterations, plus a 70+ iteration multi-day build.
- No absolute score is published, so the relative gain cannot be sited against any other system; the game demo has no baseline and no human evaluation.
- Why read it: models held fixed, only the layer above the harness varies, and all three improve — another measurement of how much reported agent capability lives in the scaffold. Read it against the pick above: nothing says how its three benchmarks score their outputs, so whether 52.25% sits above or inside AgentJudgeBench's band is undetermined.
- → Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers — arXiv 2609.01343 (score 1.49)
- TL;DR: looped-Transformer results are normally compared at fixed model size, which hands the looped model extra FLOPs. SMELT matches per-token FLOPs, non-embedding parameters and KV cache — and the advantage survives: 6.8–18.0% of training FLOPs saved on the compute-optimal frontier, fitted as a separate Chinchilla-style law across four sizes to 54B.
- Two findings cut against how scaling results get read: the downstream gain exceeds what validation loss predicts (largest on Code), and it grows with sample length and in-context examples.
- Why read it: it is the rare scaling paper that argues a prior frontier was measured against the wrong control and then reports the effect anyway once the control is fixed.
- → SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
Watch
- Moonshot seeks revenue share from US cloud giants — Trivium China, 2026-09-01. A business-model story at weight 0.7 and
triviumchina.comis blocked here, so only the headline was readable. No page written. Worth watching because it is the first sign of a Chinese lab pricing distribution rather than inference. - The Alignment Science index could not be read, and that is recorded as unverified rather than as "nothing published."
alignment.anthropic.comanswersEGRESS_BLOCKED; a targeted search surfaced no post newer than the four already captured. This is the source whose two missed posts ran +73 and +38 days late, so an unverifiable check on it is itself the signal. eval-snapshots.ymlhas not fired on its own schedule for eight days. Sixth consecutive hand dispatch. Nothing has been lost because the dispatch works, but a scheduled check that only runs when someone remembers is not scheduled.
New in Wiki
No new entity, concept or person page today — nothing here needs user review.
- Gemini 3.8 Flash (new — model page)
- Gemini 3.8 Flash Cyber (new — model page)
- AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling (new — paper page)
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement (new — paper page)
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers (new — paper page)
Updates
- Google DeepMind: two Recent Activity entries — the 3.8 Flash pair, and the Fairwind Program; both models added to Models & Products
- AI-Enabled Cyberattacks: new
## State of the Art (2026-09-03)opening with the three-lab gating comparison; one timeline row - Eval Harness Configuration: new dated section — the judge ceiling, the 9.2-vs-1.2 split, and a version-hygiene rule for Terminal-Bench
- Agents (LLM Agents): two new State-of-the-Art entries and two Key Papers rows
- Qwen 3.8 Max and Alibaba / Qwen AI Lab: the 0902 snapshot, with the two Terminal-Bench generations separated
- Post-Training Scaling: SMELT added to Key Papers
- Gemini 3.7 Flash, Gemini 3.5 Flash Cyber: successor links