$ cat briefs/daily/2026-08-31.md
2026-08-31
August 31, 2026 (Mon)
2 stories · 2 paper picks · 3 watch items · 2 new pages
> **Today's highest-scoring item is a Paper Pick, not a Top Story, and it stays > there.** WikiSkill computes to **1.75** against the leading story's **1.20**. > It is a paper and Paper Picks is what that section is for; moving it up would > make the numbering imply a ranking the sections do not carry. Every score below > is printed so the ordering can be disagreed with directly. > **The snapshot Action missed its slot for a fifth consecutive day, and this run > dispatched it for the third day running.** `eval-snapshots.yml` was **47 > minutes** past its 22:20 UTC papers cron with no run created; the newest run > was #23 at 2026-08-30T00:25Z. A manual dispatch wrote all three files in about > twenty seconds. **The scrapers succeed every time; the scheduler is the > fault** — five consecutive days is the finding, not today's 47 minutes. > Carried to the W36 lint as check 2o.
Top Stories
1. Tencent shipped a 200 GiB version of its 1.5 TB model one day after the weights, and said nothing measurable about what it cost — score 1.20
- Tencent Hunyuan announced MIX-STQ1_0 on 2026-08-29, compressing Hy4 preview from 1.5 TB to ~200 GiB GGUF through its AngelSlim toolkit, published as
AngelSlim/Hy4-preview-GGUF. The method is per-layer mixed precision chosen by calibration data rather than a uniform bit-width — some layers down to 1.31-bit (STQ1_0), some up to 2.06-bit (IQ2_XXS) (source) - The quality claim is "it still works well" and that is the whole of it. No benchmark table, no per-task delta, no retained-accuracy figure was published. A circulating "about 98% performance" is a Reddit submitter's headline, not Tencent's number and not traced to any table — it is deliberately not on the wiki page
- Why it matters: a 7.5× reduction published by the vendor in the same week as the weights is the step that decides whether an open-weight release is usable outside a datacentre, and it is one neither Z.ai nor Alibaba took in this fortnight's other two Chinese releases — both left quantisation to third parties. It closes the gap Hy4 preview named on 08-30 only halfway: an artefact exists where none did, and no hardware target, throughput or memory figure exists from anyone
- → Hy4 preview, Tencent
2. Two music publishers sued Anthropic over how Claude was trained, and the score you give it depends on which lane you read it in — score 0.91
- Sony Music Publishing and Warner Chappell Music filed on 2026-08-28, alleging a "brazen campaign of illegally torrenting, scraping and downloading copyrighted works on a massive scale" across "thousands upon thousands" of their compositions. Relief sought: up to $150,000 per infringed work, $25,000 per instance of removing copyright management information, and a jury trial (source)
- Reporting places it as the fourth music-publishing action against Anthropic — after Concord/UMG, BMG (2026-03, 493 compositions) and Round Hill Music (2026-08-17) — and distinguishes it by scope, alleging tens of thousands of compositions where the others named narrower sets
- The score is contestable and both readings are printed. Read as litigation and business (
0.7) × Anthropic (1.3) it computes to 0.91 and sits second. Read through frontier models and training data (1.3) it computes to 1.69 and would lead the brief. The lower reading was taken becauseinterests.mdnames no litigation signal and its tracked-signal list is about capability; the higher one is defensible if you think training-data provenance is a capability question - Why it matters: this wiki has tracked Anthropic's legal exposure as government friction — the Pentagon designation, the unsealed Anthropic v. DoD emails. This is the other kind, and the first entry here that turns on how the models were built rather than how they may be sold. Discovery reaches training-data provenance whether or not a transparency regime does. These are allegations; no court has ruled on any of them, and Anthropic's response was not read
- → Anthropic, AI Governance
Paper Picks
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution — arXiv:2608.27454 — score 1.75
- TL;DR: agents that discover their own skills leave the reasoning behind each skill scattered across optimization histories, where it cannot be reused. WikiSkill holds raw experience, an accumulated knowledge base (a wiki) and executable skills as three separate layers and co-evolves the skills with the wiki. Its ablation reports the persistent accumulation is critical to any of it working.
- Three findings that do not follow from the framing: skills transfer across models and across model families; skills evolved by another model can beat a model's self-evolved skills; and smaller models with skills can outperform substantially larger models without them.
- Why read it: it is the first paper this wiki holds that puts an ablation under the premise LLM Knowledge Bases (LLM-curated personal wikis) is built on. And the transfer result says the accumulated artefact is a shared asset — improvement stops being a property of a model and becomes a property of an object beside it.
- Caveat that limits everything above: the abstract names no benchmark, no model and no figure, and hedges to "most model-benchmark settings". This is a direction, not a magnitude.
- → WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment — arXiv:2608.23691 — score 1.00
- TL;DR: Stephen Chung, Wenyu Du and William J. Wesley put agents from different model families in an environment called the Station with one shared research goal, no central coordinator and no scripted pipeline. Reported outcome: results novel to the mathematical literature on five open problems.
- Primary discovery is attributed to a Claude agent for 18 (64.3%), a GPT agent for 9 (32.1%) and a Gemini agent for 1 (3.6%).
- Why read it: every multi-agent arrangement this wiki holds is one family with assigned roles. A mixed-family population with no orchestrator is a different object, and this is the first per-family attribution split recorded here.
- The 64.3% is the number most likely to be misquoted. Nothing read states how many agents of each family were present, so it is not a capability comparison. No model versions were published either.
- Sourcing is thin and stated as such:
arxiv.orgis blocked from this run and the paper was not read. The page rests on two search passes, run with different queries, that agreed on every figure. - → Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Watch
- Four days after release, the only numbers anyone has for Hy4 preview are still Tencent's own. Today's Artificial Analysis capture lists Tencent through Hy3 alone, unchanged at Intelligence Index 42. Checked rather than assumed — and worth noting that every row's Intelligence Index and Cost per Task is byte-identical to yesterday's; only the three serving columns moved (source)
- The Anthropic Alignment Science standing check could not run for a third consecutive day.
alignment.anthropic.comanswersEGRESS_BLOCKEDhere, so the index was not read at all. Introspection Adapters (April 2026) and The Hot Mess of AI (February 2026) remain uncaptured at day 20. Recorded as unreadable-this-run, not as clear — the point of this check is that a source can be listed, cited 24 times, and never fetched - Egress from this sandbox is blanket-blocked for a third straight day, broader than
agents/daily-run.mddocuments:www.anthropic.com,alignment.anthropic.com,openai.com,arxiv.org,huggingface.coandsimonwillison.netall answered000at curl. WebSearch works. Its cost is visible above — both paper pages carry figures nobody here read from the paper
New in Wiki
No new entity or concept page today, so nothing needs user review.
Updates
- Hy4 preview: the MIX-STQ1_0 quantisation, the 1.5 TB / 1.56 TB artefact size, a like-for-like Hy3 comparison (295B/21B/256K/598 GB against 770B/49B/1M/1.56 TB), and today's leaderboard confirming no third-party measurement exists
- Tencent: the quantisation release, and that no other Chinese lab in this fortnight shipped one first-party
- Anthropic: the Sony/Warner filing, with the unread defendant response marked unread rather than absent
- AI Governance: litigation as the second route to training-data provenance, beside the EU AI Act disclosure obligations already tracked here
- LLM Knowledge Bases (LLM-curated personal wikis): WikiSkill's ablation, and what its transfer result says about the schema's "portability first" principle
- AI for Mathematics: the Station result, and how it lands on this page's "no shared measure exists" open problem
- Agents (LLM Agents): two ways an agent improves without its weights moving, both portable