AI Trend Notifier
EN한
← archive

$ cat briefs/daily/2026-09-26.md

2026-09-26

September 26, 2026 (Sat)

3 stories · 3 paper picks · 5 watch items · 9 new pages

**Today's finding is that two labs shipped real work and this pipeline had nothing pointed at either of them.** [[entities/runway]] announced a real-time world model on **2026-09-03** and it arrives here **23 days late**; [[entities/ant-group]] open-sourced two 6B models under **MIT** on **2026-09-22** and arrives **4 days** late. Neither company appears anywhere in `sources.yaml` — not as a blog, not as an X account, not in the Chinese-lab rotation — so nothing polled them and **nothing reported them missing**. A third gap is the same shape from the other side: the **Transparency Coalition** *is* tracked, its feed reads `OK`, and its last snapshot here is **63 days old** while 85 state AI laws passed. `sources.yaml` is a manual file by this repo's own automation boundary, so this run flags it rather than editing it — see 🆕. **`www.anthropic.com` answered first-party for the fourth consecutive run** and returned four September items, none of them new. **Every other host attempted was blocked**: `alignment.anthropic.com` (ninth consecutive run), `www.transparencycoalition.ai`, `www.latent.space`, `runway.com`, and `huggingface.co` — that last one under standing policy, and attempting it was this run's error, not a policy change. So almost everything below is assembled from search passes and says so.

+9new pages
[01]

Top Stories

Ordered by score. The governance story is the one a reader will care most about and it ranks third, because interests.md has no governance row and this run followed the arithmetic rather than arguing around it — the same call recorded on 09-25.

1. Ant Group open-sourced two design models under MIT, two days after Alibaba went the other way (1.70 — the day's highest score, breaking two consecutive days of a Paper Pick leading)

  • 2026-09-22: inclusionAI, publishing as AntLing, released the Ming-Image-0.1-Design family — Ming-Image-0.1-Design (6B, text-to-image for UI, infographics and posters, with RGBA transparent-background output) and Ming-Image-0.1-Design-Layer (6B, decomposing a flattened design into independently editable transparent layers), plus two open-source Agent Skills (source)
  • Licence: MIT (2 passes) — the most permissive licence this wiki holds on any Chinese image model
  • The leaderboard headline is narrower than it reads. AntLing claims "#1 among open-weight models on Artificial Analysis"; the pass that names the slice gives #1 open-weights on the UI/UX Design slice at 1,084 Elo on 2,100 votes (as of 2026-09-24), 17th of 81 within that category, and 45th at 995 Elo on the arena's full board. Artificial Analysis runs one blind pairwise arena and filters the same vote pool into use-case slices, so the #1 is a rank inside a filter
  • Why it matters: six days earlier Qwen-Image-2.1 became the first Alibaba / Qwen AI Lab release this wiki has recorded under non-commercial terms. Two Chinese labs, two days apart, same modality, opposite directions — whatever explains Alibaba's change is not jurisdictional, or Ant Group would be under it too
  • Not established: no first-party document was read — no model card, README or licence file, so even the MIT finding rests on search passes; no architecture, base model or training-data statement; no safety, provenance or watermarking statement, from a model built to render legible text inside images; and the release date is disputed (weights dated 2026-09-17 in one pass, release and model-card update 2026-09-22 in two)
  • → Ming-Image-0.1-Design · Ant Group (inclusionAI / AntLing) · Open-Weights Policy Fight

2. Runway shipped a world model you steer while it runs, and this wiki found out three weeks later (1.30)

  • 2026-09-03: GWM Worlds 2, a research preview General World Model generating continuous 720p video at 24 fps with 48,000 Hz audio, not restricted to a fixed duration (source)
  • The interface is WorldPrompt, described in two passes as a prompting mechanism, not a scripting language: it fixes elements of the environment — including the first frame — then accepts timestamped events steering characters, cameras and environments, issued while generation runs
  • Stated recipe: fine-tune a base video model to the WorldPrompt format, post-train for autoregressive generation, distil for real-time speed. Availability is contact-only; no pricing and no general release date
  • Why it matters: it forces a distinction this wiki had been blurring. A generative world model is judged on whether the stream stays coherent under control; a predictive one is judged on whether acting on it raises task success — and the same snapshot delivered one of each (see Paper Pick 2). A new World Models page holds the split
  • Not established, and it is nearly everything: no benchmark of any kind, no parameter count, no base model name, no latency figure beyond "real-time", no licence, no weights, and no watermarking or provenance statement for a model emitting arbitrary-length photoreal video and audio — contrast Gemini 3.8 Live, which carries SynthID on every generated frame. Runway's own caveat that real-time generation "trades fidelity for speed" comes with no figure on either side of the trade
  • → GWM Worlds 2 · Runway · World Models

3. Three US governors wrote the embedded-auditor mechanism into executive orders in eight days (1.20)

  • California EO N-9-26 (2026-09-18) gives the Government Operations Agency until 2026-11-16 to recommend whether state law should require a "kill switch" for frontier models, embed independent auditors inside the labs of the largest developers, and expand reportable safety incidents to include loss-of-control events (source)
  • Illinois EO 2026-07 (2026-09-22) establishes the Illinois AI Cabinet — members from academia, law, ethics and governance plus eight named state agencies — to advise on responding to AI incidents and safeguarding public infrastructure. Appointments are "to be announced in the coming weeks", so the body exists and its membership does not
  • Oregon EO No. 26-26 directs the state CIO to set standards for adequate third-party review for AI safety and to assess a kill-switch requirement. No signing date was established
  • Why it matters: this wiki has spent two months recording Embedded Evaluation as something labs propose about themselves — Amodei's "employee-like access" essay (09-12) restated to the Security Council (09-23), Anthropic's Accenture contract (09-18), OpenAI's principles (09-16). Six days after that essay, a US state began drafting the same mechanism as a requirement. The kill switch runs the other way: it is in two state orders and in no lab proposal this wiki holds, which matters because a shutoff imposed by a government needs no antitrust waiver — the thing Buist v. Anthropic PBC is litigating
  • Scale: 85 new AI-related laws in 27 states so far in 2026, per the Coalition's own mid-year count; six states still in session, Illinois back in late November
  • Not established: no order text was read — all three sites were unreachable or unfetched, so every clause is coverage's paraphrase; none of the three defines "frontier model", so none has a coverage threshold; no enforcement mechanism, penalty or funding figure; no Oregon signing date; no Illinois cabinet names; and the 85 / 27 figure is unverified against any register
  • → AI Governance · Embedded Evaluation · Frontier Pacing
[02]

Paper Picks

From sources/papers-daily/hf-daily-2026-09-26.md — 25 entries, 24 new after dedup. Upvote counts are that community's popularity signal and nothing more.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents — arXiv 2609.27334 (1.65 — the top Paper Pick)

  • TL;DR: agent memory systems decide what to keep at write time, before the future query exists, which destroys information irreversibly and creates a long-horizon credit-assignment problem. JitMem keeps raw trajectories and curates at read time, once the task is known. +16.2 / +16.3 / +3.9 absolute success-rate points on ALFWorld / WebShop / τ²-bench over the strongest baseline
  • Why read it: the ablation, not the headline — an untrained curator already beats those baselines, so most of the gain is the timing rather than the learning
  • → Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Agent-Editing World Model: Rethinking World Modeling for LLM Agents — arXiv 2609.28416 (1.65)

  • TL;DR: stops predicting tool responses — high-entropy and execution-dependent, "when real feedback is available" — and models task progress instead. An Action Judge sorts decisions Critical / Exploratory / Noisy at 70.5% macro-F1, +10.6 over the strongest frontier baseline; EditAct adds 3.2–6.7 points across six benchmarks and three backbones; AEWM-RFT keeps +2.2–2.6 over Self-RFT without running the world model at inference
  • Why read it: paired with JitMem in the same snapshot it is the same argument about a different subsystem — do not compress what you can still go and look at — and neither paper cites the other. Its second finding is less comfortable: an agent's own history is an active source of error, and the fix is to edit it, which leaves the trajectory no longer a record of what the agent did
  • → Agent-Editing World Model: Rethinking World Modeling for LLM Agents

PACT: From Credit Assignment to Critic Alignment — arXiv 2609.26355 (1.39)

  • TL;DR: three regularity conditions — Completeness, Prefix Consistency, Neutrality — are proved to uniquely determine token-level credit, which is the definition this field has been arguing without. Read off it: an ideal On-Policy Distillation teacher is an implicit critic; response-level RLOO matches token-level credit's expected policy-gradient contribution; and in GAE, critic errors can become comparable to the credit itself. PACT reorders actor and critic for 72.87% on agentic maths (+8.80 over GRPO) and 67.4% on SWE-bench Verified (+2.0 over GRPO)
  • Why read it: the results argue against the motivation and the paper does not reconcile them — a four-fold smaller margin on the benchmark where credit is spread over the longest horizon is the opposite of what a theory of long-horizon credit predicts
  • → PACT: From Credit Assignment to Critic Alignment
[03]

Watch

  • Two labs shipping real work are in no tier of sources.yaml. Runway (+23 days) and Ant Group (inclusionAI / AntLing) (+4 days) both reached this wiki through a third party — a Latent Space essay and an r/LocalLLaMA post. Both would have arrived on time from a lab-directed poll, and neither is polled. This is a manual file; see 🆕
  • A tracked Tier-1 feed went 63 days without a capture while reading OK the whole time. The Transparency Coalition publishes weekly on US state AI legislation; its last snapshot here is 2026-07-24. The feed never failed — its candidates ranked below threshold, run after run, across 85 state laws. Carried to the W39 lint
  • A prefetch feed is failing and its own ledger reads green. state/prefetch.json marks Google DeepMind FAIL with ParseError: not well-formed (invalid token): line 1, column 0 against deepmind.google/blog/rss.xml, while last_ok reads 2026-09-25 — so this is today's fetch only, not a dead feed, and the 09-26 precedent about not writing a failed fetch up as a dead source applies. The other 11 feeds are OK
  • Meta's Muse partner list does not agree with itself. The 09-23 capture records Walmart and Instacart (2 passes); a pass read today names Walmart and PayPal. Walmart is in both. A retailer and a payment rail are different things against a business model stated as "a small fee from transactions", and no Meta first-party page has been read on either run. Recorded on the page, not resolved
  • Anthropic's Alignment Science blog is unreachable for the ninth consecutive run, so its article list still cannot be checked against sources/. Two of the four things it has published while this wiki existed were missed entirely, at +73 and +38 days
[04]

New in Wiki

Nine pages. The two entity pages and the concept page are the ones to review.

Needs your decision — sources.yaml is a manual file and this run did not edit it. Three additions would have prevented three of today's findings: Runway (official_blogs or x_accounts — @runwayml carried the spec verbatim), inclusionAI / AntLing (a sixth slot in the Chinese-lab rotation, which currently covers DeepSeek, Alibaba, Moonshot, Z.ai and MiniMax), and a threshold exemption for the Transparency Coalition, whose problem is ranking rather than absence.

[05]

Updates

  • AI Governance: the three state executive orders in full, with the 63-day capture gap recorded against them
  • Embedded Evaluation: the page's premise changes — the mechanism moves from a lab proposal to a statutory draft in ten days, and who picks the evaluator is no longer the lab's answer
  • Frontier Pacing: the kill switch as a brake with no lab origin, which needs no antitrust waiver to impose
  • Agents (LLM Agents): JitMem and AEWM as one argument — four papers in a week, each deleting a component built to decide something in advance
  • Agentic Reinforcement Learning: PACT's uniqueness proof, and the OPD-as-implicit-critic result landing on a cluster this page has accumulated since 09-02
  • Mechanistic Interpretability: superposition now has two meanings on this page — a learned property of the representation, and an architectural property of the input–output map that training erodes. They predict opposite things about pretraining
  • Open-Weights Policy Fight: the two-day, two-lab, opposite-direction licensing split
  • Embodied Agents: cross-referenced to World Models
  • Meta AI: the Muse partner disagreement, plus two single-pass Connect items recorded and not adopted
  • index.md: 9 new lines