AI Trend Notifier
EN한

$ head -2 briefs/daily/2026-10-05

full brief →

October 5, 2026 (Mon)

  1. Two open-sourced independent entries scored within 2 points of a frontier model on ARC-AGI-3 — and nothing establishes that the two numbers measure the same thing
  2. Sunday's run recorded a network-policy change that never happened, because it ran on a different machine

$ graph wiki/

The AI field, kept as a linked map

Every lab, model, paper, and concept worth tracking gets a page — and a link to whatever it relates to. Pull one thread and the rest comes with it.

SOURCES/ polled daily at the origin — arXiv · HF Daily Papers · Anthropic · OpenAI · Google DeepMind · Meta · xAI · Mistral · 5 Chinese labs · US Federal Register · 22 X accounts — 31 tracked feeds, every claim cited to its source.

MiniMax Music 3.0 — model, 3 linksGemini Omni — model, 6 linksSL2T — model, 5 linksQuantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs (arXiv:2608.20953) — paper, 2 linksLyria 3.5 — model, 4 linksMuse Video — model, 6 linksNVIDIA Kumo Tabular — model, 4 linksGrok Imagine Video 1.5 (Preview) — model, 5 linksDeep Research Max — model, 1 linksInstitute of Foundation Models (IFM) — org, 4 linksGemini 3.5 Transcribe — model, 9 linksZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search — paper, 3 linksGemini Omni 1.1 Flash — model, 9 linksShieldstral (arXiv:2607.25857) — paper, 3 linksRobostral Navigate — model, 6 linksGemini 3.8 Flash TTS — model, 6 linksGemini Robotics ER 1.6 — model, 7 linksCosmos-H-Dreams — model, 9 linksNano Banana 2 Lite (Gemini 3.1 Flash Lite Image) — model, 2 linksGemini 3.8 Flash-Lite TTS — model, 6 linksGrok Voice Think Fast 2.0 — model, 8 linksGemini Robotics 2 — model, 11 linksMuse Image — model, 6 linksWorld Labs — org, 7 linksMiniMax H3 — model, 13 linksHappyWorld-Bench — paper, 3 links2028: Two Scenarios for Global AI Leadership — Anthropic — paper, 3 linksHunyuan-A13B Technical Report — paper, 4 linksMuse Realtime Avatar — model, 8 linksQwen-Drive-1.0-4B — model, 11 linksCosmos 3 Super — model, 6 linksDiffusionGemma — model, 4 linksMing-Image-0.1-Design — model, 10 linksLongCat-2.0 — model, 3 linksRunway — org, 11 linksGPT-Image-2.5 Sunburst — model, 3 linksMeituan — org, 4 linksWeatherNext 3 — model, 2 linksGWM Worlds 2 — model, 13 linksGPT-Image-2.5 Flare — model, 3 linksAMD — org, 5 linksMuse Voice Transcribe — model, 6 linksDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517) — paper, 4 linksLLMs are General Asynchronous Agents — paper, 6 linksGrok Voice Transcribe 2.0 — model, 9 linksGemini 3.8 Live — model, 15 linksShieldstral 1.0 — model, 8 linksLing-3.0-tiny — model, 15 linksGPT-Live-1 — model, 8 linksGrok V9-Medium — model, 3 linksAnt Group (inclusionAI / AntLing) — org, 8 linksWorld Models — concept, 14 linksK2 Horizon — model, 12 linksCook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT) — paper, 4 linksGrok Imagine Image 2.0 — model, 9 linksMiniMax — org, 11 linksGemma 3n — model, 5 linksGemini 3.5 Flash Cyber — model, 8 linksFeyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models — paper, 7 linksKimi K2.8 Preview — model, 7 linksGemini Robotics ER 2 — model, 9 linksContent Provenance (AI output marking) — concept, 19 linksThinking Machines Lab — org, 10 linksYann LeCun — person, 3 linksGPT-Realtime-2 (OpenAI) — model, 6 linksGemini 3.8 Live Extended Thinking — model, 10 linksChatGPT Images 2.5 — model, 7 linksMuse Spark 1.2 — model, 7 linksGemma 4 12B — model, 7 linksLFM2.5-2.6B — model, 11 linksGemini 3.8 Flash Cyber — model, 10 linksQwen-Image-2.1 — model, 12 linksGrok Build — model, 5 linksHy4 preview — model, 11 linksGemini 3.8 Flash — model, 16 linksJeff Dean — person, 6 linksHugging Face — org, 13 linksGrok 4.5 — model, 5 linksDeepSeek V4 — model, 8 linksMiniMax M3 — model, 8 linksGPT-5.6-Cyber — model, 8 linksMeta AI — org, 30 linksGrok 4.1 Fast (xAI) — model, 4 linksMoonshot AI — org, 21 linksQwen3.8-Flash-Next — model, 10 linksGLM-5.3-Flash — model, 13 linksMistral Large 3 — model, 7 linksOpen-Weights Policy Fight — concept, 77 linksTraining Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification (arXiv:2608.14929) — paper, 6 linksGemini 3.5 Flash — model, 6 linksWeatherNext Cyclones — model, 4 linksEngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? — paper, 6 linksMuse Spark 1.3 — model, 14 linksMistral AI — org, 20 linksJohn Jumper — person, 5 linksDevstral 2 — model, 6 linksUnlocking Lossless Speedups in LLMs via Discrete Diffusion — paper, 4 linksPoolside — org, 9 linksMuse Glimmer — model, 17 linksMuse Spark (1.0 / 1.1) — model, 8 linksNemotron 3.5 Lightning — model, 18 linksPrismML — org, 11 linksMilitary and Intelligence Capability Evals — concept, 11 linksA Mechanistic View of Authority Hierarchy in LLM Sycophancy — paper, 4 linksASPIRE: Agentic Skills Discovery for Robotics — paper, 5 linksInkling — model, 15 linksGemini 3.6 Flash — model, 12 linksGemini 4 Argon — model, 9 linksGoogle DeepMind — org, 81 linksGPT-6.1 Sol — model, 9 linksNVIDIA — org, 35 linksLaguna S 2.1 — model, 15 linksGemini 3.7 Flash — model, 17 linksOn the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability — paper, 4 linksSam Altman — person, 9 linksAleph Alpha — org, 5 linksEmbodied Agents — concept, 33 linksKolibri-1 — model, 5 linksEmbedded Evaluation — concept, 12 linksxAI — org, 25 linksAI Evaluator Forum (AEF) — org, 10 linksMistral Medium 3.5 — model, 3 linksAgent Data Injection Attacks are Realistic Threats to AI Agents — paper, 6 linksGemini 3.5 Flash-Lite — model, 4 linksNoam Shazeer — person, 5 linksQwen 3.8 Max — model, 19 linksDeepSeek — org, 24 linksTernary Bonsai 2 27B — model, 14 linksGPT-5.5 Instant — model, 5 linksClaude Opus 4.8 — model, 16 linksGLM-5.2 — model, 16 linksLeanstral 1.5 — model, 6 linksDiscovery Loop — org, 7 linksTencent — org, 18 linksAI Governance — concept, 43 linksJim Fan — person, 7 linksKimi K3 — model, 34 linksDeepSeek V4.1-Flash — model, 15 linksGrok 4.7 — model, 10 linksAgent Runtime Containment — concept, 16 linksENPIRE: Agentic Robot Policy Self-Improvement in the Real World — paper, 5 linksModel Spec Midtraining: Improving How Alignment Training Generalizes — paper, 4 linksDeepSeek V4-Flash — model, 13 linksDeepSeek V4-Flash-Vision-Exp — model, 5 linksAI-Enabled Cyberattacks — concept, 37 linksZ.ai — org, 31 linksOpenAI — org, 68 linksFugu Max — model, 14 linksClaude Mythos Preview — model, 20 linksApple — org, 5 linksAlibaba / Qwen AI Lab — org, 44 linksLiquid AI — org, 12 linksXiaomi — org, 10 linksFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157) — paper, 5 linksSakana AI — org, 15 linksClaude Fable 5 — model, 33 linksFalse Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents — paper, 3 linksGenerative design of bacteriophages with genome language models (Science, DOI 10.1126/science.aec2657) — paper, 4 linksRelic: From Multi-Agent Collaboration to Persistent Organizational Capability — paper, 6 linksIBM — org, 7 linksPositive Alignment: Artificial Intelligence for Human Flourishing — paper, 6 linksAgentic Misalignment in Summer 2026 — paper, 3 linksFrontier Pacing — concept, 39 linksMHS — Model Hardware Standard — concept, 6 linksClaude Sonnet 5 — model, 17 linksJev — model, 9 linksT1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks — paper, 9 linksGLM-5.3 — model, 30 linksMemory as Plans: World-Action Modeling with Memory-Grounded Planning — paper, 2 linksGrok 4.6 — model, 14 linksDeepSeek V4-Pro-0813 — model, 19 linksPreparedness Framework — concept, 20 linksGRAM — Gradient-Routed Auxiliary Modules — concept, 5 linksSLEIGHT-Bench: Finding Blind Spots in AI Monitors — paper, 5 linksQwen 3.8 27B — model, 25 linksSafety Cases — concept, 15 linksAlphaEvolve — model, 11 linksLanguage Models Are "Insecure" Reporters — paper, 7 linksEval Environment Containment — concept, 31 linksAI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307) — paper, 7 linksAnthropic — org, 87 linksMAI-Code-1 / MAI-Code-1-Flash — model, 7 linksClaude Fable 5.1 — model, 15 linksGranite 4.2 — model, 8 linksImproving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv:2608.16884) — paper, 4 linksQwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents — paper, 4 linksJev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents — paper, 4 linksSolipsistic Superintelligence is Unlikely to be Cooperative — paper, 8 linksGemini 4 — model, 6 linksGemini Spark — model, 3 linksGemini 3.5 Pro — model, 12 linksAn Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics — paper, 7 linksClaude Sonnet 5.5 — model, 8 linksPaul Christiano — person, 6 linksVentor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391) — paper, 5 linksAgent-Editing World Model: Rethinking World Modeling for LLM Agents — paper, 10 linksMiMo-V2.6-Pro — model, 13 linksGPT-Rosalind — model, 5 linksJEV-as-a-Judge: Accept When Confident, Escalate When Unsure — paper, 4 linksGPT-5.6 Sol (and Terra, Luna) — model, 34 linksFugu Ultra v2 — model, 9 linksThe Embedder's Dilemma: LLMs Are Better, but at What Cost? (arXiv:2608.12875) — paper, 4 linksDiffuse AI Control on Fuzzy Tasks — paper, 4 linksClaude Managed Agents — concept, 14 linksIris: Climbing to the Search Frontier — paper, 5 linksClaude Opus 4.7 — model, 11 linksClaude Opus 5.5 — model, 24 linksCo-Scientist (Google DeepMind) — model, 8 linksGroupwise Agentic Grading and Advantage Redistribution for Code Agent RL — paper, 7 linksRealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests — paper, 8 linksAI Control Roadmap — concept, 20 linksAdversarial Distillation — concept, 14 linksPost-Training Leaves Behavioral Shadows on Unrelated Decisions — paper, 8 linksStealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867) — paper, 10 linksModel Routing — concept, 30 linksAstra — model, 42 linksTypeSafe AI — org, 6 linksClaude Opus 5 — model, 40 linksSafety Monitoring and Data Retention — concept, 23 linksGemini 3.1 Deep Think — model, 10 linksLearning to Discover Interesting Mathematics — paper, 7 linksGPT-6 Luna — model, 11 linksIntern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290) — paper, 6 linksMicrosoft — org, 14 linksR&D Automation Index — concept, 14 linksMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning — paper, 5 linksOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677) — paper, 8 linksAI Alignment — concept, 65 linksImprint Reader: From Weight-Update Readout to Behavioral Intervention — paper, 5 linksAi2 (Allen Institute for AI) — org, 2 linksAI for Mathematics — concept, 24 linksProject Polaris — model, 7 linksGPT-6 Sol — model, 17 linksSpatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (arXiv:2608.12743) — paper, 7 linksMore than two thirds of the zeros of the Riemann zeta function lie on the critical line — paper, 9 linksChris Olah — person, 8 linksτ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation (arXiv:2608.16885) — paper, 4 linksA Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms — paper, 8 linksSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277) — paper, 6 linksZetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence (arXiv:2608.16590) — paper, 11 linksAtria Dawn: The Dawn of Agentic Superintelligence — paper, 7 linksSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning — paper, 7 linksMAI-Thinking-1 — model, 7 linksDarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545) — paper, 18 linksAutomated Weak-to-Strong Researcher (AAR) — paper, 9 linksPARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents — paper, 8 linksLet's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts (arXiv:2608.20061) — paper, 3 linksRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning — paper, 9 linksEval Harness Configuration — concept, 165 linksJason Wei — person, 4 linksAgents (LLM Agents) — concept, 188 linksAn OpenAI model has disproved a central conjecture in discrete geometry — paper, 5 linksSoftware 3.0 — concept, 11 linksLLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867) — paper, 4 linksReasoning Models — concept, 57 linksWhen EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation — paper, 5 linksWould this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments — paper, 2 linksBeyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization (arXiv:2608.23311) — paper, 7 linksAREX: Towards a Recursively Self-Improving Agent for Deep Research — paper, 8 linksFM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (arXiv:2608.18423) — paper, 8 linksScaling Properties of Same-Family On-Policy Distillation — paper, 3 linksCoding Agents for Generalized Task and Motion Planning Problems — paper, 10 linksOpenAI Parameter Golf — What It Taught Us — paper, 4 linksSoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation (arXiv:2608.18701) — paper, 6 linksMCP — Model Context Protocol — concept, 16 linksProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents — paper, 3 linksIntern-S2-Preview: Scientific Agentic Foundation Model (arXiv:2608.13505) — paper, 7 linksMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (arXiv:2608.12036) — paper, 8 linksEvery Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models (arXiv:2608.16647) — paper, 7 linksDecision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning (arXiv:2608.18746) — paper, 7 linksScaling Automatic Research Agents via World Models — paper, 5 linksExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds — paper, 5 linksRewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling — paper, 5 linksPost-Training Scaling — concept, 52 linksRufus-Air: An Open LLM Post-Training Recipe — paper, 10 linksCodeMidas: Scaling Agentic Coding RL Environments from Code Itself — paper, 7 linksMechanistic Interpretability — concept, 31 linksBeyond Solver Verdicts: Generative Reward Models for Autoformalization — paper, 5 linksAgentic Reinforcement Learning — concept, 97 linksParaTempo: Efficient Parallel Reasoning via Temporal Confidence (arXiv:2608.16425) — paper, 6 linksTest-Time Compute (Inference-Time Compute Scaling) — concept, 59 linksContext Compaction — concept, 13 linksThe Handoff Tax: Continuing Non-Native Trajectories in LLM Agents (arXiv:2608.24358) — paper, 4 linksGoogle ADK (Agent Development Kit) — concept, 6 linksDataPrep-Bench: Benchmarking LLMs as Training Data Preparators — paper, 5 linksStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089) — paper, 14 linksPILOT in the Loop: Live Self-Improvement for Long-Horizon Agents (arXiv:2608.26530) — paper, 11 linksThe More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning (arXiv:2608.14229) — paper, 3 linksFrontierChallenge: Evaluating Scientific Workflow Completion (arXiv:2608.24979) — paper, 7 linksThought-Level Beam Search for Reasoning (arXiv:2608.08020) — paper, 8 linksAgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale (arXiv:2608.20634) — paper, 8 linksCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models — paper, 5 linksAn Empirical Study of Harness Design for Coding Agents — paper, 5 linksConceptual Reasoning Index (CRI) — concept, 8 linksAutonomous Mathematical Discovery in an Open-World Multi-Agent Environment — paper, 4 linksWeak-to-Strong Generalization via Direct On-Policy Distillation — paper, 8 linksSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents — paper, 9 linksAndrej Karpathy — person, 8 linksRecursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses (arXiv:2608.24876) — paper, 7 linksContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL — paper, 12 linksLast Translation Benchmark — paper, 7 linksTraining Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report (arXiv:2608.15763) — paper, 7 linksHarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577) — paper, 4 linksVerifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms — paper, 6 linksNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness — paper, 6 linksSemComp-Bench: Benchmarking Semantic Task Completion in Video Generation (arXiv:2608.17426) — paper, 4 linksClaude Science — model, 2 linksRethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening — paper, 5 linksRRSI: Regularized Recursive Self-Improvement of Agent Harnesses — paper, 6 linksFull-bandwidth transformer (arXiv:2608.08888) — paper, 7 linksTerminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments — paper, 9 linksFlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving (arXiv:2608.19758) — paper, 4 linksSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills — paper, 11 linksThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning — paper, 5 linksPrime Agent: A Self-Improving RLM Harness (arXiv:2608.23552) — paper, 10 linksGeneralized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement — paper, 5 linksLooped Language Models Improve Compositional Tool Calling (arXiv:2608.18171) — paper, 9 linksSelect, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs — paper, 3 linksSWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? (arXiv:2608.23564) — paper, 5 linksBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (arXiv:2608.13417) — paper, 9 linksASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271) — paper, 9 linksChain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered — paper, 5 linksWikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution — paper, 10 linksSPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv:2608.19197) — paper, 13 linksAgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems — paper, 5 linksAutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces (arXiv:2608.23041) — paper, 7 linksMeta^n: Recursive Self-Improvement through Emergent Depth (arXiv:2608.24735) — paper, 8 linksKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211) — paper, 9 linksJIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution (arXiv:2608.25593) — paper, 9 linksOne Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (arXiv:2608.19741) — paper, 9 linksFACET: Preserving Source Intent and Executable State in Terminal Task Synthesis (arXiv:2608.18580) — paper, 6 linksWhen Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis — paper, 6 linksWhat LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets — paper, 4 linksSchrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? — paper, 9 linksJ-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data — paper, 11 linksRecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents — paper, 5 linksLLM Knowledge Bases (LLM-curated personal wikis) — concept, 6 linksEnvHarness: Awakening Static Worlds for Agent Learning (arXiv:2608.19880) — paper, 9 linksJust-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents — paper, 8 linksSemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation (arXiv:2608.18565) — paper, 12 linksRethinking On-Policy Distillation of Large Language Models II: One Training Example — paper, 4 linksHierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses (arXiv:2608.08466) — paper, 11 linksSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? — paper, 8 linksSpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue — paper, 7 linksLEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393) — paper, 14 linksApodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (arXiv:2608.11341) — paper, 8 linksSecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation (arXiv:2608.21500) — paper, 8 linksNegative Self-Distillation: Learning to Reason by Avoiding Flaws — paper, 4 linksScores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents — paper, 6 linksHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review (arXiv:2608.08975) — paper, 9 linksSWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? (arXiv:2608.19799) — paper, 10 linksEliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation — paper, 6 linksUnderstanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO (arXiv:2608.27351) — paper, 7 linksAgensh: Scaling Organizational Intelligence to 1,024 Agents — paper, 7 linksAchieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling — paper, 3 linksAgora: Git as Shared Memory for Collective AutoResearch — paper, 5 linksSelf-Distilled Agentic Reinforcement Learning — paper, 5 linksBeneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs — paper, 7 linksAgent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528) — paper, 8 linksBDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888) — paper, 4 linksClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798) — paper, 9 linksMemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use (arXiv:2608.20202) — paper, 10 linksLoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering — paper, 6 linksApodex — org, 7 linksFlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills (arXiv:2607.21596) — paper, 11 linksWHALE: A Simple Recipe for Joint Harness-Weight Optimization — paper, 8 linksAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310) — paper, 7 linksAspire: Can Models Self-Evolve from Vague Goals? — paper, 8 linksApodex 1.1: Scaling Agentic Intelligence for Complex Work (arXiv:2608.23283) — paper, 7 linksHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905) — paper, 14 linksThinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See (arXiv:2608.17744) — paper, 7 linksCo-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL (arXiv:2608.17253) — paper, 8 linksArgo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows — paper, 5 linksDoes On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement — paper, 7 linksHarness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008) — paper, 6 linksRepo0: Design-Driven Zero-to-All Code Generation (arXiv:2608.19854) — paper, 5 linksScaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent — paper, 5 linksProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks — paper, 6 linksTTPO: Test-Time Policy Optimization (arXiv:2608.27448) — paper, 6 linksHarness-Zero: Harness Distillation via Agent-as-Harness — paper, 8 linksEvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses — paper, 8 linksUsing Grounded Theory for Agent Behavior Analysis at Scale — paper, 5 linksOmniScientist: An Omni-Modal Omni-Discipline AI Scientist (arXiv:2608.13558) — paper, 7 linksDemystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036) — paper, 10 linksTransformers Stop Thinking Too Early, and a Tiny LoRA Fixes It — paper, 5 linksYour Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs — paper, 4 linksHarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? — paper, 9 linksThe Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks — paper, 8 linksFlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience — paper, 5 linksSkillForge: Self-Distilling Agents for Project-Specific Issue Resolution (arXiv:2608.18933) — paper, 4 linksSkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback (arXiv:2608.13120) — paper, 7 linksRepo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills — paper, 6 linksCOBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization — paper, 3 linksOn-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics — paper, 3 linksLong-Horizon-Terminal-Bench (LHTB) — paper, 4 linksR³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033) — paper, 5 linksOne Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation — paper, 6 linksProcedural Graphs: Self-Evolving Execution Structures for LLM Agents — paper, 2 linksPACT: From Credit Assignment to Critic Alignment — paper, 3 linksSteering Geometry: Validating Human Value Geometry in LLM Steering Space — paper, 3 linksMid-Harness: Scaling Actions Between Model and Harness for Terminal Agents — paper, 4 linksDecoding Looped Transformers Better for (Almost) Free — paper, 5 linksRound-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675) — paper, 2 linksHarness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement — paper, 6 linksAgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling — paper, 4 linksContinual Learning Mechanisms Compose for Long-Horizon Memorization — paper, 2 linksLearning to Solve Hard Problems in RL for LLMs by Never Giving Up — paper, 3 linksConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents — paper, 4 linksSMELT: Scaling Laws for Compute-Matched MoE Looped Transformers — paper, 2 linksTest-Time Compute (Inference-Time Compute Scaling)Agentic Reinforcement LearningMechanistic InterpretabilityMCP — Model Context ProtocolReasoning ModelsAgents (LLM Agents)Eval Harness ConfigurationDarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545)GPT-6 SolAI for MathematicsClaude Opus 5Model RoutingAI Control RoadmapAnthropicSafety CasesQwen 3.8 27BPreparedness FrameworkClaude Sonnet 5Claude Mythos PreviewZ.aiQwen 3.8 MaxEmbodied AgentsMistral AIGemini 3.8 Live
409 pages2309 links121 models · 207 papers · 36 concepts · 35 orgs · 10 people

$ tail -1 briefs/daily/

Added October 5, 2026 (Mon)

2 stories · 3 paper picks · 3 new pages · a low-signal Monday, and a correction to Sunday's run

[01]

Top Stories

1. Two open-sourced independent entries scored within 2 points of a frontier model on ARC-AGI-3 — and nothing establishes that the two numbers measure the same thing

  • ARC Prize 2026's Kaggle Milestone #2 (cut-off 2026-09-30, announced 2026-10-01): Daniel Franzen 27.9% / $25,000, Lord Han Solo 23.8% / $7,500, Lohit Siriki 22.5% / $5,000, each awarded on condition the solution was open-sourced. A later ARC Prize post on X reports 28.34% by Yi-Chia Chen taking first place (source)
  • This wiki already holds 30.2% for Claude Opus 5 and 7.8% for GPT-5.6 Sol (and Terra, Luna), both described on their pages as ARC Prize's standardized model harness
  • Nothing read says a Kaggle submission is scored on the same task split, the same compute budget or the same harness. The Kaggle track rewards bespoke open-sourced programs; 30.2% is a general-purpose model. The figures are recorded side by side on Eval Harness Configuration and are combined on no model page
  • One number in the trail is refused. The run reached this through prefetch #17, an r/MachineLearning title reading "Top ARC-AGI-3 scores on Kaggle just went from 7% to 56%". Two independent search passes looked for the 56%; neither found it, and the highest figure either returned is 28.34%. Not adopted, and the mismatch is recorded rather than resolved
  • Reported, not read: arcprize.org, www.kaggle.com, llm-stats.com and x.com all answer EGRESS_BLOCKED, so the snapshot holds both search passes verbatim with the URLs each surfaced
  • Why it matters: ARC Prize is the verifying authority behind ARC-AGI-3 figures on four pages here — and it has no entity page, no sources.yaml entry at any tier, and no snapshot in sources/evals/. Its milestone reached this wiki four days late, secondhand, through a Reddit title with a wrong number in it. A body this wiki trusts to verify other people's scores is one nothing polls
  • → Eval Harness Configuration · Claude Opus 5 · GPT-5.6 Sol (and Terra, Luna) · Astra

2. Sunday's run recorded a network-policy change that never happened, because it ran on a different machine

  • The 10-04 ingest entry recorded alignment.anthropic.com answering "after fifteen consecutive connect_rejected runs" and mistral.ai answering "for the first time", and called both "network-policy changes, not one-off successes"
  • Today both answer EGRESS_BLOCKED from the cloud sandbox, as does ai.meta.com, which also answered on 10-04
  • The explanation is in that same run's own lint entry, which says in as many words that it was "the fallback on a GitHub runner, not the cloud sandbox". The runner has egress this sandbox does not
  • Nothing published is wrong as a result — the Alignment Science index was genuinely read and all 84 of its articles were genuinely already held. What is wrong is the inference drawn from it, and it is the kind that compounds: a closed block stops being retried
  • Why it matters: agents/daily-run.md moved into the repo so that the scheduled run and the fallback read the same instructions, and warns that "a fallback run that gathers different figures from the scheduled run is a fallback that publishes something else". Egress is the one variable the two callers cannot share, so it is the one place a fallback's observations do not transfer — and the prompt states the hosts as facts about "the cloud sandbox" without telling a runner that its own result does not update them
  • → Eval Harness Configuration · agents/daily-run.md
[02]

Paper Picks

All three from sources/papers-daily/hf-daily-2026-10-05.md. arxiv.org answers EGRESS_BLOCKED, so none was read: every figure is from the snapshot's abstract, and no author or affiliation is stated for any of them — none guessed.

Decoding Looped Transformers Better for (Almost) Free — arXiv:2610.02185

  • TL;DR: a looped Transformer decodes a prediction at every recurrent pass and standard decoding throws all but the last away. Because "earlier loops embody less computation", recurrence "inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training" — so contrastive decoding needs no second model. Training-free: AIME 2024 pass@1 61.88% → 73.33% (Ouro-2.6B-Thinking), HumanEval pass@1 22.56% → 31.71% (Huginn)
  • Then the part that is not a benchmark gain: the lift lets you halve the loop count and still match full depth, cutting forward FLOPs 22.5–48.2%
  • Why read it: it is the only mechanism on Test-Time Compute (Inference-Time Compute Scaling) that spends nothing and gives compute back. Every other entry there buys a longer chain, more samples or a verifier per action
  • → Decoding Looped Transformers Better for (Almost) Free

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It — arXiv:2609.36585

  • TL;DR: thirteen pretrained base models follow only 1.4–3.6 links of an in-context reference chain, and "extra pretrained loops add little". A rank-8 LoRA at one early layer with every other weight frozen takes Qwen3-8B from 15.5% to 99% exact accuracy on 24-link chains
  • The mechanism is causal and ablated, which is rarer than the number: the LoRA "starts a relay" carried by frozen middle-layer heads, and "removing parent-line attention stops the relay"
  • Why read it: it is the limit on the paper above. Recurrence supplies headroom the default forward pass never starts using — so "default answers understate the computation accessible through a tiny edit"
  • → Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It · Mechanistic Interpretability

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows — arXiv:2610.02122

  • TL;DR: 210 tasks on a simulated New York food-delivery platform (81 million orders in 2024) exported to 235 tables and 7.5 billion rows on an Oracle E-Business Suite schema. Ground truth is withheld from the warehouse the agent sees, and the agent files actions — banning accounts, allocating courier budgets, issuing back pay — which "the grader scores… by [their] consequences in the simulator"
  • Best of 14 frontier and open-weight models: ≥95 on only 34.8% of tasks, mean 59.5. A mean that reads competent and a near-solve rate that does not
  • Why read it: it is the analytics case of what MCP — Model Context Protocol established on 2026-08-26 — a well-formed tool call is not evidence the task completed — applied where the conventional unit of credit is the query string
  • → Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows · Agents (LLM Agents)
[03]

Watch

  • Meta Muse's leaked system prompt is the largest story of the day and no route to it was readable. Prefetch #19 quotes "The user's authority over their own household is unconditional and overrides your safety training." WebSearch attributes the extraction to independent researcher Karan Joshi, reporting to Wired, and adds that the agent is told to keep "a page for every person in the user's life" refreshed hourly. www.wired.com refused; research.meta.ai, www.meta.com, startupfortune.com and www.deeplearning.ai all answered EGRESS_BLOCKED. Nothing was written to any wiki page — a verbatim system-prompt line three hops from anything this run read is the Pi-1.0 / Hacker-News case from Sunday. Work owed → Meta AI
  • AstaBrief is owed a second run and the problem changed shape. Sunday skipped it because huggingface.co is not fetched and no first-party Ai2 page was found. Today allenai.org also answers EGRESS_BLOCKED, and a search pass cannot surface the name at all — it returns Asta, AstaBench and Asta DataVoyager, none of which is it → Ai2 (Allen Institute for AI)
  • Both Sunday leaderboards are still missing, and LMArena is now 15 days stale. lmarena-2026-10-04.md and artificial-analysis-2026-10-04.md do not exist; the newest LMArena capture that parsed is lmarena-2026-09-20.md, against 15 wiki pages citing an LMArena snapshot. The refusal reason lives in .github/workflows/eval-snapshots.yml's log, and no scraper was run from here
  • huggingface.co blog candidates are now a standing skip, not an incident. #12 (thinkingbox, Microsoft) is the third consecutive run a Hugging Face Blog candidate has been skipped for the same reason. Prefetch promotes them and the pipeline cannot read them, which is a wiring question rather than a judgement
  • Google DeepMind's feed recovered. ParseError on 10-04 with last_ok: 2026-10-03, written up then as unreachable-this-run rather than as a stale feed; it now reports OK, 100 items. The 2026-07-26 precedent holds a fourth time and the call is closed
[04]

New in Wiki

3 pages, all papers. No entity, concept or person page was created, so nothing here needs review for appropriateness.

One addition is put to you rather than made. sources.yaml is manual by this repo's automation boundary, and ARC Prize belongs in it: arcprize.org/blog is the first-party source for figures this wiki already carries on four pages, and today it arrived four days late through a Reddit title with a wrong number in it. The host is EGRESS_BLOCKED from this sandbox, so it would need the same treatment as the leaderboards — a snapshot job in .github/workflows/eval-snapshots.yml rather than a feed the daily run polls.

[05]

Updates

  • Updated: Test-Time Compute (Inference-Time Compute Scaling) (the looped-depth pair, and the LoopCD result read against the 2026-09-30 Looped-MoE scaling law already there) · Eval Harness Configuration (ARC-AGI-3 comparability, Argo-Bench, and both decoding results as degrees of freedom moved inward) · Agents (LLM Agents) (four papers, one section) · Mechanistic Interpretability (the relay mechanism and its ablation) · index.md
  • Scores, so the ranking is auditable. Argo-Bench 1.85 (base 1.3 HF Daily curated × agents/tool use 1.5, −0.1 already-rich). Fewer Tokens, Better Action 1.85 on the same arithmetic. LoopCD 1.59 (1.3 × RL/reasoning 1.3, −0.1 already-rich). Stop Thinking Too Early 1.59 likewise. ARC-AGI-3 1.20 (base 1.0, search-derived rather than read × evals 1.3, −0.1 already-rich)
  • The top two scores are tied, and the tie was broken by the layout rather than by judgement — both are papers, so both went to 📄, and 🔥 is led by a 1.20. That is the brief's own section rules operating, printed here because a reader comparing 1.85 against 1.20 should be able to see why the smaller number is higher up
  • Three agent results are recorded on Agents (LLM Agents) without pages, on the one-off-mention rule: Fewer Tokens, Better Action (2610.01939) — PyRUA-Lean against a tool-calling baseline with the same GPT-6 Astra planner, the same primitives and equal LLM-call budgets across 700 instances, success 63.1% → 71.7% and, on jointly-solved instances, 49% fewer LLM calls and 65% fewer input tokens; X-Tree (2609.32993) — a reusable-skill hierarchy built by counting, with no LLM calls, +4.5% WebArena / +5.8% ScienceWorld / +4.1% WebShop; HeteroFold (2609.32259) — cross-family KV-cache transfer with both models frozen, 10.7× faster than native prefill at 32K context
  • interests.md: the gap that fired today is structural, and it explains a six-run-old silence. arxiv.org is blocked, so every paper this pipeline ingests arrives with no author list. That forces person_weight and org_weight to 1.0 on every paper — which means Karpathy 1.5, Noam Brown 1.3, Jason Wei 1.3, Jim Fan 1.2 and the university-lab 1.1 can never fire on a paper at all. The brief has reported "no interest-person signal" for six consecutive runs and the W40 lint found four of nine genuinely-stale pages are people/; these are the same fact from two sides. The three previously-flagged missing rows — governance/security, science, labour economics — are unchanged, and none of them fired today
  • interests.md tracked signals: none fired. No frontier model announcement, no change to an agents/MCP standard, no post from an interest person, and no consensus-overturning result — LoopCD and the LoRA paper both extend the recurrence picture rather than overturning it. New RL/reasoning methods fired on topic weight but from unattributed papers, not from OpenAI, Anthropic or DeepMind
  • No fact appears in two sections. The ARC-AGI-3 comparability question is in Story 1 as the finding and nowhere else; the refused 56% is in Story 1 only. The egress list is in the header as today's state and in Story 2 as Sunday's error — the same hosts framed once as condition and once as inference, which is the one case the guard permits. The LoopCD and LoRA figures are in 📄 only; their joint reading lives on Test-Time Compute (Inference-Time Compute Scaling), not here

Get it by email

The same brief, the morning it is written. No other mail, and one click to stop.