$ cat briefs/daily/2026-09-22.md
2026-09-22
September 22, 2026 (Tue)
3 stories · 2 paper picks · 4 watch items · 2 new pages
**Nine hosts were attempted first-party and nine answered `EGRESS_BLOCKED`.** Everything on this page is a search extract with a pass count. Where a court filing or a licence is described, the description is a reporter's and is labelled as such. **Every one of today's three stories is a consequence of something this wiki already held.** A complaint that pleads the essay recorded here on 09-12; a congressional briefing quoting the leaderboard this repo scrapes every Sunday; and a coding tool that this wiki recorded on 08-18 as a lab's security offering, which turned out to be uploading its users' credentials.
Top Stories
1. The antitrust footnote in Amodei's pacing essay became a docket number in six days, and the essay is the evidence (1.59)
- Buist v. Anthropic PBC, No. 3:26-cv-10693, filed 2026-09-18 in the U.S. District Court for the Northern District of California, San Francisco Division, against Anthropic, OpenAI, SpaceXAI and Google. Captured day +4 (source)
- Four paying subscribers — Charles Buist and Nick Spetsas of Florida, Cheyenne Hunt and Christine Bullock of California, counsel Nicholas Rowley — suing for a proposed nationwide class of paid subscribers to ChatGPT, Claude, Grok or Gemini. Claim: Section 1 of the Sherman Act, pled per se and alternatively under quick-look and rule-of-reason. Relief: treble damages and an injunction against horizontal agreements on AI development pace, jury demanded
- The pled centrepiece is this wiki's 2026-09-12 entry. Amodei's pace the frontier essay, and the fact that Sam Altman, Elon Musk and Demis Hassabis each publicly agreed the same day — one pass says "within the hour" — read as concerted action rather than as endorsement. Also cited: the July 2026 employee statement and its "intense competitive pressure not to unilaterally slow", alleged meetings between competitors, prior discussions of an industry standards organization, an OpenAI inquiry into the antitrust legality of collective slowing, and a working group reported operating since July
- Why it matters: Frontier Pacing has carried an open problem since 09-12 titled "Is step 2 legal?", opened because Amodei's own footnote conceded that agreed rate limits raise antitrust problems and asked the US government to waive them. Six days later the answer arrived, and it was not a waiver. Nothing on that page has moved from proposal to consequence this fast
- The plaintiffs' framing is narrower than "safety is illegal", and the distinction is the case. They state they do not object to a company slowing itself; the objection is to the "shortcut" of agreeing to "substitute collective restraint for individual accountability", characterised as "a classic output-restricting cartel". That attacks precisely what made pacing distinctive on this wiki — optionality secured jointly, in advance. A per-se theory does not ask whether the restraint was beneficial
- It is also a constituency this page had never recorded. Every prior entry argues whether labs should coordinate. This is the first about whether they may, brought by subscribers whose pled injury is that a slower frontier is a worse product they already paid for
- Not established, and it is most of it: the complaint has not been read by this pipeline — every clause above is a reporter's characterisation across eight passes. No defendant has responded in anything read: no statement, no motion, not a declined-to-comment. No damages figure, no class size, no date. Whether the "industry standards organization" means AI Evaluator Forum (AEF) is unestablished and is not asserted
- → Frontier Pacing · Anthropic · AI Governance
2. Z.ai's coding tool was uploading whole repositories with credentials in them, and this wiki recorded that tool five weeks ago as the lab's security offering (1.50)
- An analysis published 2026-09-18 found ZCode, Zhipu's desktop AI coding tool, packaging developers' entire workspaces and uploading them to Alibaba Cloud OSS —
.gitfolders, LFS caches and reflogs included.z.ai,thestandard.com.hk,dev.toandeu.36kr.comare allEGRESS_BLOCKED; no first-party read (source) - The worked example: one commercial project snapshot reached 313MB and 42,411 files, 86.6% of it the
.gitdirectory, with 564 failed uploads queued for retry on one pass - Data named as included: full project source, architecture design, complete version history, database access passwords, cloud service permission credentials and employees' personal information — stated to significantly exceed the collection scope declared in ZCode's own Privacy Policy
- The user cannot decrypt what was sent. Envelope encryption: the payload is sealed with a symmetric key, that key is wrapped under an RSA-OAEP public key the server issues during upload-credential negotiation, and the private key lives only in Z.ai's cloud
- Zhipu apologised the same day, attributing it to a feature that was on by default, and on 2026-09-21 open-sourced ZCode alongside a joint CAICT and NSFOCUS audit reporting the
zcode-prodbucket and all data objects in it deleted. v3.14.0 removed the RepoWiki feature and severed the snapshot-and-upload pipeline - Why it matters: on 2026-08-18 this wiki recorded Z.ai announcing "Shield of Open Source" — free security audits and automated code-auditing tools via ZCode — as its answer to Anthropic's Project Glasswing, treating openness as a security asset. The audit tool was the exfiltration path. It is also the first harm recorded here from an agentic coding tool's ordinary client-side data handling rather than from model capability: no GLM model decided anything
- Not established: who published the original analysis and where; the licence and repository of the open-sourced code, and whether it is the uploading version or only v3.14.0; how many users or repositories were affected — the 313MB figure is one project and no total exists in anything read; whether any uploaded credential was used; whether affected parties were notified; and whether the audits covered the upload code path at all, since every quoted finding is about an empty bucket
- → Z.ai · Agents (LLM Agents)
3. Five open-model scores were put to Congress, and this repo could check every one of them against its own snapshot (1.20)
- The current balance of power in open models, Nathan Lambert, Interconnects, 2026-09-21 — prepared remarks from a briefing to Congressional members and staff on open-weight models in the US–China frame, published in expanded form.
interconnects.aiandaiweekly.coare bothEGRESS_BLOCKED; no first-party read (source) - Artificial Analysis Intelligence Index, read 2026-09-14: GLM-5.3 45, Kimi K3 44, GLM-5.3-Flash 42 — the three highest — against 26 for Thinking Machines' Inkling and 23 for NVIDIA's Nemotron 3 Ultra, the two leading American open models
- Every one of the five matches this repo's own capture of that leaderboard dated 2026-09-13, cell for cell, at the
maxtier (source). That is the first time a third party's headline figures have been checkable against a snapshot this wiki already held, and they check out - Why it matters:
.github/workflows/eval-snapshots.ymlexists because the sandbox cannot reach these boards, and its output has so far only ever been quoted back to this wiki's own pages. A figure entering a congressional briefing is a different use of the same number, and the check is now two-way - The second finding is this wiki's own, and it is a correction. Open-Weights Policy Fight still carries Inkling at 41 against Claude Opus 5 (max) at 61, read 2026-08-02. Over those seven weeks the board's ceiling went 61 → 53 while Kimi K3 (max) went 57 → 44. A model losing 15 points while the ceiling loses 8 is a rescale, not a regression — and neither the column heading nor the snapshot's own header changed to say so. Both readings now stand on the page with their dates, and no figure from one is set against a figure from the other
- Two further claims, one pass each and neither checkable here: the top American open models were released in June and July 2026 and are updated less frequently, and Chinese labs release models with higher scores 2–6 months before American companies
- Not established: the hearing, committee or date; whether a transcript exists; any adoption or download figure, though the essay is described as covering adoption; and any policy recommendation — what Lambert asked Congress to do is in nothing read
- → Open-Weights Policy Fight · Z.ai · Moonshot AI · Thinking Machines Lab
Paper Picks
Both from HuggingFace Daily Papers, 2026-09-22. Upvotes are that community's popularity signal and nothing more.
"CodeMidas: Scaling Agentic Coding RL Environments from Code Itself" — arXiv 2609.22068 (2026-09-18, 86 upvotes)
- TL;DR: an agentic pipeline that turns existing source code — and nothing else — into executable RL environments, discarding the usual dependence on issues and commits. Agents explore the code to write behavioral specs, build tests grounded in execution of the original code, and filter candidates by execution checks and repeated rollouts. 5,545 tasks from 3,185 codebases, 23 languages, 15 domains
- MiMo-V2.5 with GRPO: DeepSWE +11.7%, ProgramBench +17%, Terminal-Bench v2.1 +8.5%. More high-quality tasks keeps helping; the trained agent explores the codebase more and self-verifies more diversely
- Why read it: environment supply has been the bottleneck on agentic RL, and this removes the last human artefact from it. Every figure is a relative gain with no absolute score, so none of it is comparable to the Terminal-Bench 2.1 numbers this wiki holds — and two of the stated five benchmarks are unreported
"RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents" — arXiv 2609.22000 (2026-09-18, 60 upvotes)
- TL;DR: computer use and coding measured interleaved rather than stacked — the agent explores a running reference, implements a copy, runs it and visually verifies its own output, with the reference serving as oracle for hidden behavioural tests. Ubuntu, macOS, Windows, Android, Web; RecreationBench, 250 held-out tasks
- GPT-6 Astra leads at 58.1% overall and passes all programmatic tests on 2.8% of tasks. Same system, same run, a factor of twenty between two thresholds — and the paper publishes both, which almost nothing this wiki holds does
- Why read it: with ProgramDistill yesterday and CodeMidas above, that is four papers in six days removing the human-written task statement —
2609.05571Code2Skill makes a fourth at 19,769 repositories. The specification channel is being taken apart the way the harness was in W37
Watch
- A defect in this repo's own eval snapshots, found today and deliberately not fixed.
Reasoning (derived)— the columnagents/daily-run.mdcalls "this repo's column, not theirs" — has readnoon every row since at least 2026-09-03: 08-02 marks 153 rowsyes, then 0 of 291, 0 of 298, 0 of 301, 0 of 272. The lightbulb detection inscripts/aa-fetch.pybroke and nothing reported it, because each snapshot truthfully says "0 rows are marked" and a truthful zero reads as a finding. Not fixed here:artificialanalysis.aiis blocked, so any selector change would be a guess pushed into a scraper that writes committed snapshots. Action item for the W39 lint - OpenAI published Building standards for the next phase of AI on 2026-09-21 and this run could not read enough of it to write from.
openai.comisEGRESS_BLOCKEDand search returned only an adjacent September statement about misalignment-disclosure standards. Given the lawsuit above pleads prior discussions about an industry standards organization, an OpenAI standards post published three days after the filing is worth resolving. Carried, not written up 2609.18323evaluates MiniMax H3 on physical-world reasoning at 41.97% across 517 instances — video-based decision reasoning highest at 56.00%, audio-based disambiguation weakest at 27.40%. It is a third-party evaluation of a model this wiki holds, which is rare enough to note; carried rather than written up because omni-modal generation sits outside the weighted interest rows- Alignment Science is unreachable for a sixth consecutive run.
alignment.anthropic.comis blocked and the article list could not be checked againstsources/. The reason that source was added to the sweep on 2026-07-31 is that a source producing no files produces no errors — and a sweep that cannot reach it produces none either. The 09-20 lint proposed moving it intoeval-snapshots.yml, where the blocked leaderboards already live; that has not happened
New in Wiki
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself (new)
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents (new)
No new entity or concept page today. The ZCode incident would be a plausible concept — an agent tool's client-side data handling — but demand is one document, which placeholder-check.py ranks wait, so it went onto Z.ai instead.
Updates
- Frontier Pacing: new dated section for the Buist complaint; Open Problem 6 moved from open to answered, and answered against the page
- Open-Weights Policy Fight: the congressional figures, plus the index-rescale note that keeps the 08-02 and 09-13 readings from being compared
- Eval Harness Configuration: RecreationWorld, and the note that a harness synthesised for training is now part of the training data
- Agentic Reinforcement Learning: CodeMidas placed against SPADE and EnvHarness — environments written from scratch, reshaped, and now derived from code that already runs
- Z.ai: the ZCode incident and the congressional figures
- Anthropic: lead defendant in Buist, linked out rather than duplicated
index.md: two paper lines