AI Trend Notifier
EN

$ head -3 briefs/daily/2026-08-20

full brief

August 20, 2026 (Thu)

  1. OpenAI says frontier monitoring doesn't need your data. Anthropic has been
  2. The harness stopped being a deployment choice — it is now inside the weights
  3. What a skill actually does, measured — and it is not what the word suggests

$ graph wiki/

The AI field, kept as a linked map

Every lab, model, paper, and concept worth tracking gets a page — and a link to whatever it relates to. Pull one thread and the rest comes with it.

SOURCES/ polled daily at the origin — arXiv · HF Daily Papers · Anthropic · OpenAI · Google DeepMind · Meta · xAI · Mistral · 5 Chinese labs · US Federal Register · 22 X accounts — 31 tracked feeds, every claim cited to its source.

ASPIRE: Agentic Skills Discovery for Robotics — paper, 4 linksShieldstral (arXiv:2607.25857) — paper, 3 linksENPIRE: Agentic Robot Policy Self-Improvement in the Real World — paper, 5 linksMistral Medium 3.5 — model, 3 linksHarness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008) — paper, 4 linksCosmos 3 Super — model, 6 linksLong-Horizon-Terminal-Bench (LHTB) — paper, 4 linksAgent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528) — paper, 5 linksRobostral Navigate — model, 6 linksFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) — paper, 5 linksCosmos-H-Dreams — model, 4 linksJim Fan — person, 7 linksPoolside — org, 8 linksDemystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036) — paper, 7 linksClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798) — paper, 7 linksQwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents — paper, 4 linksSelf-Distilled Agentic Reinforcement Learning — paper, 5 linksLLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867) — paper, 4 linksCook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT) — paper, 4 linksLiquid AI — org, 9 linksSpatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (arXiv:2608.12743) — paper, 7 linksMistral Large 3 — model, 7 linksASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271) — paper, 4 linksMiniMax Music 3.0 — model, 3 linksDevstral 2 — model, 6 linksBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (arXiv:2608.13417) — paper, 6 linksGLM-5.2 — model, 8 linksThinking Machines Lab — org, 10 linksShieldstral 1.0 — model, 8 linksVentor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391) — paper, 5 linksAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310) — paper, 5 linksStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089) — paper, 12 linksLaguna S 2.1 — model, 15 linksEmbodied Agents — concept, 17 linksIntern-S2-Preview: Scientific Agentic Foundation Model (arXiv:2608.13505) — paper, 7 linksMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning — paper, 5 linksScaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent — paper, 5 linksDeepSeek V4-Flash — model, 10 linksGemini Robotics 2 — model, 9 linksGemini Robotics ER 1.6 — model, 7 linksSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning — paper, 7 linksSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills — paper, 8 linksLFM2.5-2.6B — model, 9 linksMiniMax — org, 8 linksInkling — model, 15 linksGemma 4 12B — model, 6 linksMistral AI — org, 17 linksDarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545) — paper, 14 linksNVIDIA — org, 22 linksHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905) — paper, 10 linksWeak-to-Strong Generalization via Direct On-Policy Distillation — paper, 5 linksDataPrep-Bench: Benchmarking LLMs as Training Data Preparators — paper, 5 linksAgentic Reinforcement Learning — concept, 38 linksAI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307) — paper, 7 linksMiniMax M3 — model, 8 linksMuse Spark 1.2 — model, 5 linksSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277) — paper, 5 linksKimi K3 — model, 16 linksNemotron 3.5 Lightning — model, 11 linksLeanstral 1.5 — model, 5 linksDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517) — paper, 4 linksMuse Glimmer — model, 14 linksIntern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290) — paper, 6 linksGemini Robotics ER 2 — model, 8 linksOpen-Weights Policy Fight — concept, 40 linksMoonshot AI — org, 9 linksGrok Build — model, 5 linksQwen 3.8 27B — model, 16 linksMiniMax H3 — model, 12 linksDeepSeek — org, 12 linksApodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (arXiv:2608.11341) — paper, 5 linksEval Harness Configuration — concept, 51 linksThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning — paper, 5 linksModel Routing — concept, 14 linksRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning — paper, 9 linksZ.ai — org, 14 linksGoogle ADK (Agent Development Kit) — concept, 6 linksAgents (LLM Agents) — concept, 82 linksDeepSeek V4 — model, 7 linksAREX: Towards a Recursively Self-Improving Agent for Deep Research — paper, 8 linksGrok 4.1 Fast (xAI) — model, 4 linksMCP — Model Context Protocol — concept, 9 linksDeepSeek V4-Pro-0813 — model, 14 linksYann LeCun — person, 3 linksR³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033) — paper, 5 linksAlibaba / Qwen AI Lab — org, 15 linksGLM-5.3 — model, 10 linksThought-Level Beam Search for Reasoning (arXiv:2608.08020) — paper, 7 linksMAI-Code-1 / MAI-Code-1-Flash — model, 7 linksMuse Video — model, 5 linksClaude Managed Agents — concept, 8 linksAgent Data Injection Attacks are Realistic Threats to AI Agents — paper, 6 linksQwen 3.8 Max — model, 12 linksTest-Time Compute (Inference-Time Compute Scaling) — concept, 27 linksGrok Imagine Video 1.5 (Preview) — model, 5 linksClaude Sonnet 5 — model, 13 linksSoftware 3.0 — concept, 11 linksDiffusionGemma — model, 2 linksKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211) — paper, 9 linksGemini 3.7 Flash — model, 12 linksGrok Imagine Image 2.0 — model, 7 linksAMD — org, 3 linksFrontier Pacing — concept, 22 linksMuse Spark (1.0 / 1.1) — model, 8 linksMuse Image — model, 4 linksMeta AI — org, 15 linksHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review (arXiv:2608.08975) — paper, 8 linksOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677) — paper, 8 linksClaude Opus 4.8 — model, 11 linksAI Governance — concept, 28 linksReasoning Models — concept, 41 linksAndrej Karpathy — person, 8 linksMAI-Thinking-1 — model, 7 linksFull-bandwidth transformer (arXiv:2608.08888) — paper, 6 linksGrok 4.6 — model, 10 linksClaude Opus 5 — model, 14 linksAI for Mathematics — concept, 16 linksGPT-5.6 Sol (and Terra, Luna) — model, 22 linksProject Polaris — model, 7 linksSolipsistic Superintelligence is Unlikely to be Cooperative — paper, 8 linksLLM Knowledge Bases (LLM-curated personal wikis) — concept, 4 linksLyria 3.5 — model, 4 linksMicrosoft — org, 11 linksAutomated Weak-to-Strong Researcher (AAR) — paper, 9 linksRound-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675) — paper, 2 linksContent Provenance (AI output marking) — concept, 7 linksxAI — org, 12 linksAI-Enabled Cyberattacks — concept, 25 linksStealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867) — paper, 9 linksGemini 3.6 Flash — model, 11 linksGemma 3n — model, 5 linksMeituan — org, 3 linksEval Environment Containment — concept, 15 linksAI Alignment — concept, 37 linksClaude Fable 5 — model, 23 linksGoogle DeepMind — org, 48 linksConceptual Reasoning Index (CRI) — concept, 7 linksOpenAI — org, 43 linksJason Wei — person, 4 linksGemini Spark — model, 3 linksAI Control Roadmap — concept, 16 linksAnthropic — org, 55 linksAstra — model, 22 linksClaude Opus 4.7 — model, 11 links2028: Two Scenarios for Global AI Leadership — Anthropic — paper, 3 linksGrok 4.5 — model, 5 linksGemini Omni — model, 5 linksGemini 3.5 Flash-Lite — model, 4 linksPreparedness Framework — concept, 13 linksAlphaEvolve — model, 11 linksBDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888) — paper, 4 linksGemini 3.5 Pro — model, 12 linksMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (arXiv:2608.12036) — paper, 8 linksCo-Scientist (Google DeepMind) — model, 8 linksGPT-Rosalind — model, 5 linksAchieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling — paper, 3 linksSam Altman — person, 6 linksLongCat-2.0 — model, 1 linksSL2T — model, 3 linksMechanistic Interpretability — concept, 11 linksGemini 3.1 Deep Think — model, 10 linksMore than two thirds of the zeros of the Riemann zeta function lie on the critical line — paper, 8 linksOpenAI Parameter Golf — What It Taught Us — paper, 4 linksChris Olah — person, 6 linksGemini 3.5 Flash — model, 6 linksGrok V9-Medium — model, 3 linksAi2 (Allen Institute for AI) — org, 1 linksPositive Alignment: Artificial Intelligence for Human Flourishing — paper, 6 linksGPT-Live-1 — model, 5 linksSafety Monitoring and Data Retention — concept, 8 linksGemini 3.5 Flash Cyber — model, 6 linksApple — org, 5 linksImproving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv:2608.16884) — paper, 4 linksClaude Mythos Preview — model, 11 linksHarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577) — paper, 4 linksDiscovery Loop — org, 7 linksGenerative design of bacteriophages with genome language models (Science, DOI 10.1126/science.aec2657) — paper, 4 linksGemini 4 — model, 4 linksWeatherNext Cyclones — model, 3 linksGPT-5.6-Cyber — model, 6 linksAn OpenAI model has disproved a central conjecture in discrete geometry — paper, 5 linksDiffuse AI Control on Fuzzy Tasks — paper, 4 linksNoam Shazeer — person, 5 linksGRAM — Gradient-Routed Auxiliary Modules — concept, 5 linksDeep Research Max — model, 1 linksGrok Voice Think Fast 2.0 — model, 4 linksGPT-5.5 Instant — model, 5 linksGPT-Realtime-2 (OpenAI) — model, 5 linksSLEIGHT-Bench: Finding Blind Spots in AI Monitors — paper, 5 linksJohn Jumper — person, 5 linksJeff Dean — person, 6 linksNano Banana 2 Lite (Gemini 3.1 Flash Lite Image) — model, 1 linksAgentic Misalignment in Summer 2026 — paper, 3 linksClaude Science — model, 2 linksClaude Mythos PreviewGemini 3.1 Deep ThinkMechanistic InterpretabilityAlphaEvolvePreparedness FrameworkAstraAnthropicAI Control RoadmapGoogle DeepMindClaude Fable 5Gemini 3.6 FlashAI-Enabled CyberattacksxAIGrok 4.6Claude Opus 4.8Meta AIKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)Test-Time Compute (Inference-Time Compute Scaling)GLM-5.3DeepSeek V4-Pro-0813Model RoutingMiniMax H3Qwen 3.8 27BMuse GlimmerKimi K3NVIDIA
194 pages950 links77 models · 61 papers · 26 concepts · 21 orgs · 9 people

$ tail -1 briefs/daily/

Added August 20, 2026 (Thu)

4 stories · 2 papers · 4 watch items · 8 new pages

[01]

Top Stories

1. OpenAI says frontier monitoring doesn't need your data. Anthropic has been requiring it since June — and voiding zero-retention contracts to get it

  • OpenAI, 2026-08-19: Offering Zero Data Retention for frontier models restates ZDR for eligible API customers — nothing retained after processing, no personnel review, no training use without opt-in, customer-held encryption keys — and previews Private Safety Processing, said to detect misuse patterns across related interactions while sending OpenAI only a narrowly defined safety signal, without exposing the underlying prompts or responses. Enterprise and API only; rollout and a technical white paper in September (source).
  • Anthropic, since 2026-06-09 ⚡ ** traffic is retained 30 days, on first- and third-party surfaces. It overrides a negotiated zero-retention agreement with no opt-out, and Fable 5 does not support ZDR at all (source).
  • The comparison is not yet like-for-like: one policy is enforced and 72 days old, the other is a preview with a promised paper.
  • Why it matters: both labs accept the same premise and only one can be right about whether it forces content retention — the answer decides whether "zero retention" survives as a procurement category for frontier models. And neither has published a detection accuracy or false-positive rate, the third safety-relevant classifier in two weeks announced here without one.
  • Safety Monitoring and Data Retention, Claude Fable 5, OpenAI, Anthropic

2. The harness stopped being a deployment choice — it is now inside the weights

3. What a skill actually does, measured — and it is not what the word suggests

  • Demystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036) open-codes 8,135 trial records: procedural anchoring accounts for 65.7% of skill cases against 4.5% for explicit knowledge injection. A skill library is a runbook, not a knowledge base. Skills beat Workflow Memory by +6.06 points.
  • The separate failure is the day's sharpest number: actual-use precision falls 29.6% → 3.3% as the pool grows from 5 to 100.
  • Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008) reports the same shape from the memory side — one harness, 26 metrics, 3 backbones, 4 suites, and no substrate dominates: broad retrieval helps long-context factual QA and harms sequential decision-making by pulling attention off action-critical context.
  • Why it matters: this is the mechanism under the reading adopted here yesterday — harness scaling buys execution reliability, not self-assessment — reached by a third method. And it says retrieval quantity is not monotone in usefulness: every skill and memory figure this wiki holds is reported at one small pool size, on one task class, and both now look load-bearing.
  • Agents (LLM Agents), Eval Harness Configuration

4. Frontier open weights became runnable on one machine, in the week the memory to run them repriced 5×

  • FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) serves 20+ MoE models from an 8GB laptop GPU upward — 35B on a laptop, 284B on a gaming desktop, and 753B GLM-5.2 on a single workstation GPU — by remapping computation onto free resources instead of fixing an offloading strategy.
  • Against it: consumer DDR5 up as much as 485% year over year; a 128GB kit at $3,399 against a lowest tracked price near $329; 2×32GB kits $222 → $1,272. No forecast read expects relief before late 2027 (source).
  • Why it matters: FreeToken's method substitutes host memory and bandwidth for GPU capacity — it spends precisely the resource that just repriced. Also worth holding: every FreeToken figure is capacity, none is throughput, and no quantisation level or accuracy check is published. Fitting a model is not serving it.
  • Open-Weights Policy Fight, NVIDIA
[02]

Paper Picks

HarmProfile: Characterizing Harmful Distributions in Frontier LLMsarXiv:2608.14577

  • TL;DR: 80,000+ validated harmful artifacts from 23 frontier LLMs across 13 families, in 15 categories and 57 subcategories, defining the output distribution as a model-level risk profile rather than a failure rate.
  • Why read it: six days ago Anthropic published that its task-based evaluations have saturated. A distribution keeps resolution where a threshold has lost it — two models failing at the same rate can fail in different shapes. Read the limits with it: the capability measure behind "harmfulness grows with capability" is unnamed, and nothing separates produces more from we found more at the effort we spent.
  • HarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577), AI Alignment

ASI-Bench: At the Dawn of Artificial SuperintelligencearXiv:2608.17271

[03]

Watch

  • The China compute constraint switched sides. ByteDance and Tencent have each taken about 10,000 H200s — the first to reach the mainland — against US clearance for 100,000 each. It is Beijing holding back the rest, via case-by-case NDRC approval, to protect domestic chipmakers; NVIDIA is reported to hold about 500,000 built largely for Chinese buyers (source). → NVIDIA
  • A Meta AI Mac app with standing access to your ad account and Workspace mail — reported 2026-08-19, connectors to Instagram, Facebook, Meta ad campaigns and Google Workspace. Recorded at low confidence: no first-party URL, no second outlet. The permissions question has no page here (source). → Meta AI
  • Three Alignment Science posts remain uncaptured for the seventh day — AuditBench (2026-03-10), Introspection Adapters (2026-04-28), The Hot Mess of AI. Carried to the W34 lint.
  • Today's second lead came from a vendor privacy centre, not a news page. Nothing in this pipeline reads privacy or help-centre articles, which is how a contract-overriding policy went 72 days unrecorded. A wider version of the [INTAKE-1] shape from the W33 lint, still unaddressed.
[04]

New in Wiki

[05]

Updates

Get it by email

The same brief, the morning it is written. No other mail, and one click to stop.