AI Trend Notifier
EN한
← archive

$ cat briefs/daily/2026-09-27.md

2026-09-27

September 27, 2026 (Sun)

2 stories · 3 paper picks · 5 watch items · 8 new pages

**Yesterday this pipeline found two labs it was not polling. Today it found a model it *was* capturing — weekly, into its own repository, for seven weeks — and never read.** [[models/ling-3-0-tiny]] was released **2026-08-06** and arrives here **52 days late**. `Ling 3.0 Tiny` is in **every Artificial Analysis snapshot in `sources/evals/` from 2026-08-09 onward — 13 consecutive captures** — and `InclusionAI` is in the **earliest** such snapshot this repo holds. The snapshot intake reads those tables to check figures for models the wiki already names; **nothing reads the `Creator` column for labs it does not**. A 269-row leaderboard captured weekly is a list of who is shipping, and this pipeline has been using it as a lookup table. **Two more of today's findings are about reading rather than fetching.** [[concepts/world-models]], written yesterday, said no instrument existed to compare the two senses of the term — **one had been published five days earlier**, and it is Paper-adjacent below. And `sources/evals/lmarena-2026-09-27.md` **is missing**: today is Sunday, both leaderboards were due, the Action committed one. Carried to the W39 lint. **Egress**: `www.anthropic.com` answered first-party for the **fifth consecutive run** and returned nothing new — though its index rendered **seven** September items where yesterday's read of the same page gave **four**, which is a dedup hazard nothing here would catch. Blocked: `alignment.anthropic.com` (**tenth consecutive run**), `www.meta.com`, `ant-ling.medium.com`. **`huggingface.co` was not attempted**, correcting yesterday's error.

+8new pages
[01]

Top Stories

Ordered by score. Both rank below today's top Paper Pick (1.65), which is the third time in four days the highest score on the page is a paper — interests.md scores items, and the day's finding is about this pipeline, which has no row.

1. A 7.9B MIT model shipped seven weeks ago, and this repository has been committing the evidence every Sunday (1.40)

  • 2026-08-06: Ling-3.0-tiny from inclusionAI / AntLing — a hybrid reasoning MoE, 7.9B total / ~1.3B active, MIT, weights in BF16, FP8, INT4 and GGUF, built for local agent use and stated as validated on NVIDIA DGX Spark, Apple Silicon MacBooks and Mac mini (source)
  • Architecture, and it is the substance if it holds: a hybrid linear-attention stack alternating KDA with MLA (2 passes; one expands KDA as Kimi Delta Attention — Moonshot AI's name on a mechanism inside another lab's model), which one pass says was "previously validated only in frontier-scale models" and is carried here down to ~1.3B active parameters. A sparse 128-expert MoE is single-pass and not adopted
  • Three published figures disagree with this repo's own captured leaderboard, and the captured page wins each time. Artificial Analysis Intelligence Index 15 * against 25 in a summary — and 25 is that same snapshot's figure for the sibling Ling 3.0 Flash, which is the obvious explanation and is stated nowhere, so it is recorded as a conflict rather than inferred away. Context 262k against 260k. 31 Median Tokens/s against a claimed 160+, which are not the same measurement. A fourth figure — "16 on the Artificial Analysis Agentic Index" — cannot be checked here even in principle, that not being one of the seven columns aa-fetch.py asserts (source)
  • The lab gained a model line it was recorded as not having. Yesterday's Ant Group (inclusionAI / AntLing) said "nothing read establishes whether there is a Ling model line". There is — five more siblings sit in today's captured table, including a 1T model, so Ling and Ring are two lines and nothing read says what separates them. None is given a page: a leaderboard row gives a name, a context window and a score, and no release date, licence or announcement
  • Why it matters: MIT, downloadable, 52 days old — and nothing in this wiki has been built on it. Ternary Bonsai 2 27B exists because someone rebuilt a Qwen derivative under Apache 2.0. A permissive licence is necessary and demonstrably not sufficient; what was missing here was anyone looking
  • Not established: no first-party document read — ant-ling.medium.com and huggingface.co both answer EGRESS_BLOCKED; no training-data statement; no named benchmark score at all, only the domains evaluated. No Catalogue id row was added, deliberately — Ming-Image-0.1-Design already carries one unverifiable one from 09-26, and a second is not what that row is for
  • → Ling-3.0-tiny · Ant Group (inclusionAI / AntLing) · Open-Weights Policy Fight

2. Meta's avatar model does watermark everything — and this wiki spent yesterday recording that it does not (1.10)

  • 2026-09-23, Meta Superintelligence Labs: Muse Realtime Avatar, "state-of-the-art embodiment technology that turns Muse Realtime Voice into expressive, interactive avatars" (verbatim, 2 passes). It pairs with Muse Realtime Voice so audio and video stay in sync while Muse replies in under a second for as long as the conversation runs, rendering a reference image — a photograph, an illustration, an animal, a household object — talking, gesturing and shifting posture in continuous video (source)
  • Specification: 448×768, 25 fps, ~870 ms end-to-end (2 passes), and all output watermarked as AI, "without adding latency". The shipped default is a cartoon — "Jolly", cream-coloured, "completely customizable" — while the stated inputs include "a photograph"
  • Why it matters, and it is about the procedure rather than the model. On 09-26 the name surfaced in a single pass and was held in Meta AI's ## Conflicting Reports as not adopted, with the observation that nothing read stated any watermarking. Three passes today give the specification. The hold's factual guess was wrong; the hold was right. A page written a day earlier would have published the opposite of the truth, cited. The entry is rewritten to record the resolution, not deleted
  • Three real-time generative video models, three different provenance answers: Gemini 3.8 Live carries SynthID on every frame, a named published scheme; this one says "watermarked as AI" with no mechanism stated; GWM Worlds 2 — unbounded photoreal video and audio — has none at all, and is the model the other two are chasing on capability
  • "Without adding latency" is the claim to hold onto and it is unverifiable here. Marking 25 frames a second inside an 870 ms budget is either a real engineering result or a rounding claim, and nothing read says whether the mark is visible, imperceptible, signed, or detectable by anyone outside Meta
  • Not established: no model size, architecture or base model; no availability, date, region, price or licence. Meta's claim that it beat Runway Characters and HeyGen LiveAvatar in head-to-head tests carries no win rate, sample size, rater description or comparator version — recorded as a vendor preference claim, not a result. No Meta first-party page was read, on this run or any other
  • → Muse Realtime Avatar · Meta AI · Content Provenance (AI output marking)
[02]

Paper Picks

From sources/papers-daily/hf-daily-2026-09-27.md — 25 entries, 17 new after dedup against pages and yesterday's one-off mentions. Upvote counts are that community's popularity signal and nothing more.

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue — arXiv 2609.26780 (1.65 — the day's highest score)

  • TL;DR: general agent memory retrieves content and loses who said it about whom. Two tracks — speaker-labelled verbatim messages and derived state in person-level and group-level views — joined at query time by entity, event and time. 47.9 / 69.2 / 61.9% binary accuracy on GroupMemBench / SocialMemBench / EverMemBench, 70.85% on all 1,986 LoCoMo questions
  • Why read it: it is the counter-example to the week's run of papers deleting components. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents argued curation should wait for the query; this argues speaker attribution must be built at write time because it cannot be recovered from an interleaved transcript afterwards. The two are the same disagreement stated precisely, published a day apart, neither citing the other. Its own headline number is the ablation: RL on the writer alone moves 57.38 → 68.20, +10.82, architecture held fixed
  • → SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms — arXiv 2609.27321 (1.65)

  • TL;DR: pipelines "construct an environment before defining its outcome rule", then align the two afterwards. VHD-Play solves a mathematical model first and has a setter render its decision process as stateful tools, so executable dynamics and trajectory scoring are inherited from the same solved model. 3,300 environments at a few cents each; Qwen3.6-35B-A3B 0.204 → 0.815 mean agentic score, transfer to eight unseen mechanism families and to function calling, travel planning and e-commerce; on E-Commerce Bench it completes every run without bankruptcy and exceeds Qwen3.7-Max
  • Why read it: the ablation, not the 4× headline. Comparing written-out problems against stateful versions, most of the learnable gap turns out to be stateful interaction rather than the underlying problem solving — the most direct evidence this wiki holds for why agentic and reasoning benchmarks come apart. The headline itself is an undefined scale: nothing read says what the agentic score measures or where its ceiling is
  • → Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds — arXiv 2609.30199 (1.49)

  • TL;DR: two sandboxes — AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24, 70) — whose rules are executable, so answers check exactly, and which conflict with familiar knowledge, so recall alone cannot solve them. Each ships a flawed manual. Across 10 systems the strongest do acquire and apply unfamiliar rules, but performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains
  • Why read it: the contamination control is the contribution — most agent benchmarks' defence against pre-training leakage is hoping the task is new, and making the ground truth deliberately alien is a control. Pair it with Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms from the same snapshot: the same scarce resource — a verifiable unfamiliar mechanism — manufactured once to train on and built once to test with, a day apart, neither citing the other. No score is reported, so the third finding is the one to carry: an agent that gets worse the longer it explores fails differently from one that plateaus
  • → ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
[03]

Watch

  • The snapshot intake reads for figures, not for names. Today's +52-day capture was in this repo's own weekly leaderboard files for 13 consecutive captures, and InclusionAI has been in every Artificial Analysis snapshot here since 2026-07-30. Yesterday's finding was "the lab is in no tier of sources.yaml" — true, and not the whole failure. Nothing sweeps the Creator column for labs the wiki does not have. Carried to the W39 lint
  • A Sunday snapshot is missing and nothing failed. sources/evals/lmarena-2026-09-27.md was due at 07:00 KST; commit a46cb05 carries artificial-analysis-2026-09-27.md and hf-daily-2026-09-27.md and no LMArena file. Last capture is 2026-09-20. The scrapers were not run from here, per standing policy. Carried to the W39 lint as check 2o
  • This wiki asserted an absence that was not there. World Models, created 2026-09-26, said there was no instrument on which the generative and predictive senses of the term could be compared. HappyWorld-Bench is dated 2026-09-21 — six capabilities over video / spatial / embodied tracks, 1,138 prompts / 300 scenes / 254 cases, 31 systems, human A/B Elo alongside behavioural metrics, spatial models at at best 70.14% placement and 73.33% edit execution. Recorded on the page as a reading failure, not a field gap
  • A first-party index listed different items on consecutive reads. www.anthropic.com/news gave seven September items today where yesterday's run recorded four from the same page. Six are held; the seventh, "The Situation Report" (09-22, Features), is not — see 🔇. An index whose contents vary between reads is a dedup hazard, and nothing in this pipeline would report it
  • Gemini 4's timing, held at arm's length. Koray Kavukcuoglu is reported (The Information, via one pass, 09-23/24) to hope to launch Gemini 4 "much earlier" than the end of 2026, with no month named. Single-pass, second-hand and undated, so it is not written into Gemini 4's Release Date — the same discipline that made yesterday's avatar hold the right call
[04]

New in Wiki

Eight pages, no new entity or concept page — today's captures landed on entities and concepts that already existed, which is the first run in four days with nothing needing your review on that front.

Still needs your decision, unchanged from 09-26 — sources.yaml is manual and this run did not edit it. Yesterday's three requests stand (Runway; inclusionAI / AntLing as a sixth Chinese-rotation slot; a threshold exemption for the Transparency Coalition). Today adds a fourth that is not a source at all: a sweep of the Creator column in each Artificial Analysis capture against the wiki's entity list. It is one comparison over a file this repo already commits, and it would have surfaced inclusionAI on 2026-07-30.

[05]

Updates

  • Ant Group (inclusionAI / AntLing): the Ling and Ring lines, five sibling leaderboard rows, and ## Strategic Position off _TBD_ for the first time — reporting a complication rather than a thesis, because a lab with a 1T model is not the specialist its two captured releases suggest
  • Meta AI: the avatar adopted into Recent Activity; the ## Conflicting Reports hold rewritten to record its own resolution, including that its factual guess was wrong. The Private Processing on glasses half is still single-pass and still not adopted — what corroborates is distribution, not on-device privacy, and they are not the same claim
  • Tencent: first entry since 2026-09-01; an 80B/13B open MoE with five domains claimed and zero scores, absent from the 269-row table this repo captures
  • Z.ai: a third party's eight-stage recipe on GLM-4.5-Air-Base claiming to beat Z.ai's own post-trained release — with no benchmark figure, on a base two minor generations behind what the leaderboards list
  • World Models: the missing instrument, found, dated five days before the page that said it was missing
  • Agents (LLM Agents): SpeakerMem-R1 as the week's counter-example, plus two independent treatments of search agents in one snapshot — IterSynth's architecture and Rufus-Air's dedicated training stage
  • Agentic Reinforcement Learning: VHD-Play removing the step where a reward gets written by hand, against PACT: From Credit Assignment to Critic Alignment proving what a correct one is
  • Content Provenance (AI output marking): the three-model provenance table, and a product detail with a consequence — a cartoon default and a photograph input inside one feature, treated identically by one watermark
  • R&D Automation Index: ExplorationBench and WhatWorkedBench, two benchmarks for the step this page's 26% figure assumes. WhatWorkedBench's result is the harder one: a Gaussian process fitted to the agent's own observations recovers effects better than the agent does — 0.632 → 0.698
  • Open-Weights Policy Fight: "open weights" used this month for MIT and for a territory-excluded custom licence — a restriction type this page had not held
  • Post-Training Scaling: the first fully documented recipe on the page, publishing no numbers
  • index.md: 8 new lines, the Ant Group (inclusionAI / AntLing) line rewritten