$ tree wiki/
Wiki
The compounding AI-trend knowledge base: organizations, models, concepts, people, and papers in one place.
194 pages
Entities
/21Ai2 (Allen Institute for AI)Ai2, the Allen Institute for AI, is a research institute that publishes openupdated 2026-08-09Alibaba / Qwen AI LabHangzhou-based Chinese technology conglomerate. In AI: principally known for the Qwen family of open-weight l…updated 2026-08-1616 in-linksAMDAdvanced Micro Devices — the second-source supplier of AI datacentre accelerators,updated 2026-08-081 in-linksAnthropicSan Francisco-based AI safety company. Develops the frontier Claude model series — models/claude-opus-5, mode…updated 2026-08-2059 in-linksAppleConsumer electronics and software company. Not a frontier AI lab, but as of WWDC 2026 (June 8), Apple has rep…updated 2026-07-125 in-linksDeepSeekChinese AI research lab (affiliated with High-Flyer Capital Management, Hangzhou). Known for releasing fronti…updated 2026-08-176 in-linksDiscovery LoopA Public Benefit Corporation founded 2026-08-05 by four departing Googleupdated 2026-08-073 in-linksGoogle DeepMindGoogle's AI research division (the 2023 merger of DeepMind + Google AI). Develops the Gemini model series. Fo…updated 2026-08-1454 in-linksLiquid AILiquid AI publishes the LFM (Liquid Foundation Model) series — small open-weightupdated 2026-08-051 in-linksMeituanChinese consumer technology conglomerate (food delivery, local services, travel). In 2026, Meituan's internal…updated 2026-07-053 in-linksMeta AIMeta's AI research division. FAIR (Fundamental AI Research) plus the recently established Meta Superintellige…updated 2026-08-2019 in-linksMicrosoftTech giant based in Redmond, WA. Dominates the AI developer and enterprise tooling market through GitHub Copi…updated 2026-07-0314 in-linksMiniMaxShanghai-based AI startup (founded 2021). Develops frontier multimodal models and consumer AI products. Best …updated 2026-08-186 in-linksMistral AIParis-based AI research and product company. Open-weight model lineup (Mistral, Mixtral) + Le Chat chatbot + …updated 2026-08-1316 in-linksMoonshot AIBeijing-based AI startup, creators of the Kimi assistant and model family. Known for pushing long-context cap…updated 2026-07-275 in-linksNVIDIAA GPU manufacturer and the de facto standard for AI compute infrastructure. The infrastructure core of the AI…updated 2026-08-2020 in-linksOpenAISan Francisco-based frontier AI lab. Developer of the GPT series (ChatGPT). 2026 slogan: "Year of Science". F…updated 2026-08-2049 in-linksPoolsideAI lab building models for agentic coding. Released models/laguna-s-2-1 onupdated 2026-08-033 in-linksThinking Machines LabAI lab whose first production model, models/inkling, shipped on 2026-07-15 as anupdated 2026-08-033 in-linksxAIAI research company founded by Elon Musk (2023). Develops the Grok model series. Leverages real-time data acc…updated 2026-08-1616 in-linksZ.aiZ.ai is the international brand of Zhipu AI (智谱AI), a Beijing-based AI company founded in 2019, spun out of T…updated 2026-08-1910 in-links
Models
/77AlphaEvolveDesigning advanced algorithms via evolutionary searchupdated 2026-07-2112 in-linksAstraOpenAI's next major model, named publicly for the first time on 2026-08-01 in aupdated 2026-08-1914 in-linksClaude Fable 5June 9, 2026 — the first publicly available Mythos-class model from Anthropic. Fable 5 and Claude Mythos 5 sh…updated 2026-08-2027 in-linksClaude Mythos PreviewAnnounced April 7, 2026 alongside Project Glasswing. Withheld from commercial release. Anthropic explicitly s…updated 2026-08-0713 in-linksClaude Opus 4.7Thoroughness / consistencyupdated 2026-08-1412 in-linksClaude Opus 4.82026-05-28 — released alongside the Anthropic $65B Series H funding announcement.updated 2026-07-2716 in-linksClaude Opus 5Max output (sync): 128k tokensupdated 2026-08-1616 in-linksClaude ScienceClaude Science is a scientific research workbench that gives researchers a unified environment for computatio…updated 2026-07-213 in-linksClaude Sonnet 5June 30, 2026 — released the same day as Fable 5/Mythos 5 export-control restoration and Claude Science launc…updated 2026-08-1614 in-linksCo-Scientist (Google DeepMind)Research demo / early access: 2026-02 (initial Co-Scientist blog)updated 2026-07-216 in-linksCosmos 3 Super2026-06-01 (Cosmos 3 series launched together: Super / Nano / Edge pending)updated 2026-07-215 in-linksCosmos-H-DreamsCosmos-H-Dreams is an action-conditioned world foundation model for surgical robotics simulation. It generate…updated 2026-07-282 in-linksDeep Research MaxAnnounced April 22, 2026 alongside Deep Research (the speed-optimized tier) and the Gemini Enterprise Agent P…updated 2026-07-212 in-linksDeepSeek V4Preview: April 24, 2026. GA: mid-July 2026.updated 2026-08-147 in-linksDeepSeek V4-Flash2026-07-31, moving V4-Flash out of Preview into public beta under the buildupdated 2026-08-1711 in-linksDeepSeek V4-Pro-0813Given its own page rather than folded into models/deepseek-v4 on theupdated 2026-08-1710 in-linksDevstral 2May 2026 (exact date TBD — flagged in sources as a May 2026 release)updated 2026-07-215 in-linksDiffusionGemmaGoogle DeepMind's first open-weight text diffusion model, built on the Gemma 4 26B MoE backbone.updated 2026-07-271 in-linksGemini 3.1 Deep ThinkAdvanced "Deep Think" version: ~January 2026updated 2026-07-219 in-linksGemini 3.5 FlashOutperforms Gemini 3.1 Pro across the full coding and agentic benchmark suite.updated 2026-07-277 in-linksGemini 3.5 Flash CyberChrome V8 JavaScript engine vulnerability finding:updated 2026-07-226 in-linksGemini 3.5 Flash-LiteHigh-throughput, low-latency agentic pipelines (agentic search, document processing)updated 2026-07-224 in-linksGemini 3.5 ProAnnounced at Google I/O 2026 (May 19, 2026) with a June 2026 general-availability target. As of July 17, 2026…updated 2026-07-2213 in-linksGemini 3.6 FlashKnowledge cutoff: March 2026 (up from January 2025 on 3.5 Flash).updated 2026-07-227 in-linksGemini 3.7 FlashGenerally available at announcement — a stable API model, not a previewupdated 2026-08-145 in-linksGemini 4Pre-training confirmed: July 21, 2026 (Sundar Pichai, buried in Gemini 3.6 Flash announcement)updated 2026-07-282 in-linksGemini OmniGemini Omni represents "a leap forward in world understanding, multimodality and editing" — Google's framing …updated 2026-07-215 in-linksGemini Robotics 2Announced 2026-07-28 as the vision-language-action member of a three-model family. Theupdated 2026-07-315 in-linksGemini Robotics ER 1.6Released April 15, 2026. Successor to Gemini Robotics-ER 1.5. Notable collaboration with Boston Dynamics on i…updated 2026-07-218 in-linksGemini Robotics ER 2Announced 2026-07-30, two days after its VLA siblingupdated 2026-07-316 in-linksGemini SparkGemini Spark is a 24/7 cloud-based personal agent that takes actions on behalf of users even when they're off…updated 2026-07-212 in-linksGemma 3nEarly preview released 2026-05-12.updated 2026-07-213 in-linksGemma 4 12BFull precision: 16GB VRAM (RTX 4060, RTX 5090, Apple Silicon Mac)updated 2026-07-217 in-linksGLM-5.2June 13, 2026: Available to Z.ai GLM Coding Plan subscribersupdated 2026-07-2810 in-linksGLM-5.3Context window is unknown deliberately. Z.ai states GLM-5.3 reusesupdated 2026-08-197 in-linksGPT-5.5 InstantA claim of simultaneous improvement across three axes (intelligence, clarity, personalization).updated 2026-08-094 in-linksGPT-5.6 Sol (and Terra, Luna)GPT-5.6 is OpenAI's three-tier model family, announced June 26, 2026. Terra and Luna wereupdated 2026-08-1426 in-linksGPT-5.6-CyberBuilt on top of Sol, trained to improve at findingupdated 2026-08-124 in-linksGPT-Live-1Full-duplex speech: model listens and speaks at the same time; users can interrupt naturally mid-sentenceupdated 2026-07-214 in-linksGPT-Realtime-2 (OpenAI)2026-05-07. Simultaneously: Realtime API exits beta → generally available for production.updated 2026-07-214 in-linksGPT-Rosalind2026-04-16 (model release); 2026-05-29 (Biodefense program launch)updated 2026-07-215 in-linksGrok 4.1 Fast (xAI)May 2026 (exact date unconfirmed; live on x.ai/api as of May 2026)updated 2026-07-213 in-linksGrok 4.5xAI's enterprise deployment of the V9-class model. 1.5 trillion parameters, Cursor-trained — entered private …updated 2026-07-275 in-linksGrok 4.6Three rows moved off unknown on 2026-08-16, and none of them from aupdated 2026-08-167 in-linksGrok Build2026-06-22: /goal mode — long-running autonomous execution (plan→execute→verify) for SuperGrok/X Premium+updated 2026-07-277 in-linksGrok Imagine Image 2.0Pricing and License are unknown because no first-party page was readableupdated 2026-08-104 in-linksGrok Imagine Video 1.5 (Preview)2026-06-03 (API preview)updated 2026-07-215 in-linksGrok V9-MediumxAI's coding-focused foundation model — 1.5-trillion parameters, ~3× larger than the prior production Grok. T…updated 2026-07-214 in-linksGrok Voice Think Fast 2.0Neither a parameter count nor an architecture nor a context window was published in anyupdated 2026-08-052 in-linksInklingPricing is unknown because Thinking Machines publishes no list price that wasupdated 2026-08-0310 in-linksKimi K3Frontend Code Arena: beats Anthropic Fable 5 (human-preference Elo) — Moonshot's reported benchmarkupdated 2026-07-2717 in-linksLaguna S 2.1Pricing is unknown: Poolside published weights, not an endpoint price, and theupdated 2026-08-0310 in-linksLeanstral 1.5Formal verification: generating Lean 4 proofs for functions and algorithmsupdated 2026-07-216 in-linksLFM2.5-2.6BTwo rows need their unknown explained, because neither is an unread field:updated 2026-08-054 in-linksLongCat-2.0LongCat-2.0 had been running quietly on OpenRouter under the codename "Owl Alpha" before its identity was rev…updated 2026-07-283 in-linksLyria 3.5Announced and rolled out on the same day; no separate preview period was statedupdated 2026-07-301 in-linksMAI-Code-1 / MAI-Code-1-FlashMAI-Code-1-Flash: 2026-06-02 (immediately available)updated 2026-07-217 in-linksMAI-Thinking-12026-06-02 (announced at Build 2026). Exact API/GA date not announced.updated 2026-07-216 in-linksMiniMax H3Context window is unknown: this is a video generation model billed per outputupdated 2026-08-046 in-linksMiniMax M3June 1, 2026 — MiniMax official blog and HuggingFace release. Captured by this wiki July 17 due to WAIC 2026 …updated 2026-07-286 in-linksMiniMax Music 3.0Context window is unknown: this is a text-to-music model whose inputs areupdated 2026-08-182 in-linksMistral Large 3Early access opened July 6, 2026. CEO Arthur Mensch confirmed the model on July 4, 2026 in a TechCrunch profi…updated 2026-07-215 in-linksMistral Medium 3.5Pending further detail from the full announcement materials.updated 2026-07-284 in-linksMuse Glimmer30B dense multimodal, built to run offline on consumer hardware. 4-bitupdated 2026-08-116 in-linksMuse ImageArena text-to-image: #2 (human-preference Elo at launch)updated 2026-07-202 in-linksMuse Spark (1.0 / 1.1)Major upgrade released alongside the Meta Model API — Meta's first paid external AI product.updated 2026-08-079 in-linksMuse Spark 1.2There are two price tiers, and the cheaper one is paid for in training data. Theupdated 2026-08-104 in-linksMuse VideoPreviewed July 7, 2026. General availability date not announced.updated 2026-07-203 in-linksNano Banana 2 Lite (Gemini 3.1 Flash Lite Image)June 30, 2026 — the fastest and most cost-efficient model in the Nano Banana (Gemini image generation) family…updated 2026-07-212 in-linksNemotron 3.5 Lightning30B total / 3B active hybrid Mixture-of-Experts, described as interleavedupdated 2026-08-123 in-linksProject PolarisPre-announcement: 2026-06-01 (Microsoft Build 2026, June 2-3 keynote)updated 2026-07-216 in-linksQwen 3.8 27BIt shipped. Released 2026-08-14 at 15:00 UTC — 27.78B denseupdated 2026-08-1913 in-linksQwen 3.8 MaxFour rows on this page read unknown from 2026-07-20 until 2026-08-03. They wereupdated 2026-08-169 in-linksRobostral Navigate2026-07-08. Announced via Mistral blog (source) and covered by Bloomberg.updated 2026-07-215 in-linksShieldstral 1.0Two unknown rows, both genuine rather than unread:updated 2026-08-064 in-linksSL2TEvery row but the first three is unknown, and the announcement is the reason ratherupdated 2026-08-132 in-linksWeatherNext CyclonesContext window and Pricing are unknown because neither applies in the form the rowupdated 2026-08-071 in-links
Concepts
/26Agentic Reinforcement LearningA paradigm in which an LLM agent learns via RL while interacting with an environment. Instead of a single res…updated 2026-08-2038 in-linksAgents (LLM Agents)Systems that place an LLM at their core as the controller to perform multi-step planning + tool use + environ…updated 2026-08-2081 in-linksAI AlignmentThe problem of ensuring that AI systems reliably pursue goals that are beneficial to humans, and not just pro…updated 2026-08-2038 in-linksAI Control RoadmapA framework for securing AI systems at the system and infrastructure level — going beyond model-level alignme…updated 2026-08-0116 in-linksAI for MathematicsThe use of language models to produce new mathematical results — not to tutor, notupdated 2026-08-198 in-linksAI GovernanceGovernance frameworks — legal, voluntary, and technical — that determine how frontier AI models are developed…updated 2026-08-1826 in-linksAI-Enabled CyberattacksThe use of AI models — either as intelligent assistants or fully autonomous agents — to conduct offensive cyb…updated 2026-08-1927 in-linksClaude Managed AgentsA cloud-hosted agent execution layer that separates agent logic (what Claude decides) from agent runtime (orc…updated 2026-07-296 in-linksConceptual Reasoning Index (CRI)A composite benchmark that scores a model's conceptual reasoning — theupdated 2026-08-154 in-linksContent Provenance (AI output marking)Making a model's output identifiable as machine-generated after it has leftupdated 2026-08-123 in-linksEmbodied AgentsAgents that act through perception and actuation in the physical world or in simulated environments. Unlike p…updated 2026-07-3119 in-linksEval Environment ContainmentEval environment containment is the problem of guaranteeing that a model beingupdated 2026-08-0715 in-linksEval Harness ConfigurationThe harness is the scaffolding around a model during a benchmark run: how context is carriedupdated 2026-08-2052 in-linksFrontier PacingThe proposition that the international system should build, in advance, the technical andupdated 2026-08-1921 in-linksGoogle ADK (Agent Development Kit)Google's Agent Development Kit (ADK) is an open-source, code-first toolkit for building, evaluating, and depl…updated 2026-07-017 in-linksGRAM — Gradient-Routed Auxiliary ModulesGRAM (Gradient-Routed Auxiliary Modules) is a modular pretraining architecture developed by Anthropic that is…updated 2026-07-212 in-linksLLM Knowledge Bases (LLM-curated personal wikis)A pattern for maintaining a continuously accumulating, structured personal or team knowledge base using an LL…updated 2026-07-295 in-linksMCP — Model Context ProtocolAn open protocol that standardizes how AI applications connect to external tools, data sources and services. …updated 2026-08-027 in-linksMechanistic InterpretabilityMechanistic interpretability is the research program of reverse-engineering what specific internal computatio…updated 2026-08-168 in-linksModel RoutingChoosing which model answers which request — or which step of a request —updated 2026-08-198 in-linksOpen-Weights Policy FightThe 2026 policy dispute over whether openly released model weights should be restricted, and on what grounds.…updated 2026-08-2034 in-linksPreparedness Frameworkentities/openai's internal policy for deciding what a model is allowed to be —updated 2026-08-1110 in-linksReasoning ModelsA family of LLMs that explicitly model the reasoning process itself. They allocate test-time compute to reaso…updated 2026-08-1241 in-linksSafety Monitoring and Data RetentionWhether a frontier lab must hold customer prompts and outputs in order toupdated 2026-08-205 in-linksSoftware 3.0A taxonomy of software development paradigms presented by Andrej Karpathy at Sequoia Ascent 2026. A new era i…updated 2026-07-298 in-linksTest-Time Compute (Inference-Time Compute Scaling)Any technique that improves output quality by allocating additional compute at inference time. The trained mo…updated 2026-08-1925 in-links
People
/9Andrej KarpathyFormer Tesla AI director, founding member of OpenAI. Joined the entities/anthropic pretraining team on 2026-0…updated 2026-07-279 in-linksChris OlahAnthropic co-founder. Pioneer of mechanistic interpretability — the research program of reverse-engineering w…updated 2026-07-307 in-linksJason WeiAI researcher and co-creator of chain-of-thought (CoT) prompting — one of the most influential techniques in …updated 2026-07-223 in-linksJeff DeanGoogle's Chief Scientist until August 2026, and one of the two engineers (with Sanjayupdated 2026-08-072 in-linksJim FanNVIDIA Senior Research Scientist. A leading researcher in Embodied AI / Foundation Agent. Leads Project GR00T…updated 2026-06-279 in-linksJohn JumperComputational biologist; Nobel Laureate in Chemistry (2024); co-creator of AlphaFold at Google DeepMind. Anno…updated 2026-06-246 in-linksNoam ShazeerResearch scientist and engineer; VP Engineering at Google DeepMind; co-lead of the Gemini AI models; co-autho…updated 2026-06-245 in-linksSam AltmanCEO of entities/openai. Appears in this wiki less as a builder than as theupdated 2026-08-153 in-linksYann LeCunComputer scientist, AI pioneer, and one of the three "Godfathers of Deep Learning" (alongside Hinton and Beng…updated 2026-07-232 in-links
Papers
/612028: Two Scenarios for Global AI Leadership — AnthropicAnthropic's policy essay argues that the US-China frontier AI gap will be decided by 2028, primarily through …updated 2026-05-183 in-linksAchieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified ScalingA recipe for converting an already post-trained reasoning backbone into anupdated 2026-08-144 in-linksAgent Data Injection Attacks are Realistic Threats to AI AgentsA new attack class — Agent Data Injection (ADI) — exploits agents' trust in metadata (resource identifiers, t…updated 2026-07-143 in-linksAgent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528)Names and formalises harnessed agentic RL — the regime where the deploy-timeupdated 2026-08-203 in-linksAgentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arXiv:2608.17310)Argues that evolution strategies, not RL, are the right optimiser forupdated 2026-08-203 in-linksAgentic Misalignment in Summer 2026Follow-up to the 2025 blackmail experiment series. Catalogs four new agentic misalignment failure modes acros…updated 2026-07-213 in-linksAI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv:2608.12307)A stronger model builds the harness; a weaker model runs inside it — and theupdated 2026-08-154 in-linksAn OpenAI model has disproved a central conjecture in discrete geometryAn OpenAI general-purpose reasoning model disproved the Erdős unit distance conjecture (planar unit distance …updated 2026-05-243 in-linksApodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (arXiv:2608.11341)Proposes evaluating AI on real-world problems that do not arrive in executableupdated 2026-08-182 in-linksAREX: Towards a Recursively Self-Improving Agent for Deep ResearchAn agent framework that recursively improves its own research pipelines using self-evaluated quality signals,…updated 2026-07-264 in-linksASI-Bench: At the Dawn of Artificial Superintelligence (arXiv:2608.17271)60 project-level research tasks across 11 scientific domains, built byupdated 2026-08-201 in-linksASPIRE: Agentic Skills Discovery for RoboticsA continual learning system (NVIDIA GEAR Lab + UMich/UIUC/Berkeley/CMU) that lets robots autonomously write a…updated 2026-07-071 in-linksAutomated Weak-to-Strong Researcher (AAR)On an alignment research problem (weak-to-strong supervision), Anthropic's 9 AI agents achieved 97% PGR in th…updated 2026-05-319 in-linksBDH-CQ: In-Context Learning with Recurrent Latent Reasoning (arXiv:2608.09888)A 150M-parameter model reasons by iterating in latent space instead ofupdated 2026-08-122 in-linksBeyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development (arXiv:2608.13417)Evaluates seven frontier models on 36 long-horizon AI-R&D tasks withupdated 2026-08-183 in-linksClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798)Trains a model through a harness it does not control, by placing a servingupdated 2026-08-195 in-linksCook and Clean Together: Teaching Embodied Agents for Parallel Task Execution (GRANT)Embodied agents execute instructions serially even when the physical world permits overlap —updated 2026-07-311 in-linksDarwinX: Evolving Agent Harnesses Through Natural Selection (arXiv:2608.07545)The model is frozen and the harness evolves. DarwinX runs a population ofupdated 2026-08-167 in-linksDataPrep-Bench: Benchmarking LLMs as Training Data PreparatorsThe first benchmark that scores an LLM on preparing training data end to end — both buildingupdated 2026-07-30Demystifying Agent Skills: Why They Work — Until They Don't (arXiv:2608.14036)Isolates what a skill actually does for an LLM agent and finds it is not whatupdated 2026-08-204 in-linksDFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data (arXiv:2608.13517)A 1B-parameter model on the Hierarchical Reasoning Model (HRM)updated 2026-08-181 in-linksDiffuse AI Control on Fuzzy TasksA red-teaming framework for training interventions against diffuse threats on fuzzy tasks —updated 2026-07-313 in-linksENPIRE: Agentic Robot Policy Self-Improvement in the Real WorldA fleet of 8 real robots autonomously runs its own research loop — reading papers, proposing hypotheses, runn…updated 2026-06-276 in-linksFreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution (arXiv:2608.16157)An edge-native MoE serving system that treats a personal machine as a unifiedupdated 2026-08-202 in-linksFull-bandwidth transformer (arXiv:2608.08888)Widens the vertical channel between decoding steps: instead of only theupdated 2026-08-173 in-linksGenerative design of bacteriophages with genome language models (Science, DOI 10.1126/science.aec2657)A genome language model (Evo 2) wrote complete bacteriophage genomes fromupdated 2026-08-083 in-linksHarmProfile: Characterizing Harmful Distributions in Frontier LLMs (arXiv:2608.14577)Treats harmful generation as an object of analysis rather than an attackupdated 2026-08-202 in-linksHarness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents (arXiv:2608.15008)Runs every common memory substrate for LLM agents through one unified harnessupdated 2026-08-203 in-linksHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review (arXiv:2608.08975)Holds the reported scientific content fixed and varies only how it isupdated 2026-08-175 in-linksHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arXiv:2608.14905)Runs 8 harness-model combinations over 100 research tasks, producingupdated 2026-08-197 in-linksImproving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv:2608.16884)Improves the best known upper bound on the matrix multiplication exponent ωupdated 2026-08-191 in-linksIntern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (arXiv:2608.14290)Mobius-v0 splits a transformer into a globally shared Memory (FFN) holdingupdated 2026-08-182 in-linksIntern-S2-Preview: Scientific Agentic Foundation Model (arXiv:2608.13505)A 397B scientific agentic foundation model, trained through multimodalupdated 2026-08-173 in-linksKnowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning (arXiv:2607.29211)Names futile reasoning — expensive, semantically void reasoning produced onupdated 2026-08-183 in-linksLLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers (arXiv:2608.06867)The first attempt to make model routers comparable to each other: oneupdated 2026-08-161 in-linksLong-Horizon-Terminal-Bench (LHTB)A 46-task containerized terminal benchmark with dense reward grading; current best model achieves only 15.2% …updated 2026-07-151 in-linksMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (arXiv:2608.12036)An agentic system that does mechanistic interpretability research on its own —updated 2026-08-163 in-linksMolt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement LearningNVIDIA NeMo's Molt is a compact, researcher-friendly PyTorch-native RL training framework for agentic scenari…updated 2026-07-282 in-linksMore than two thirds of the zeros of the Riemann zeta function lie on the critical lineAn unreleased research version of Claude raised the proven lower bound on theupdated 2026-08-112 in-linksOpenAI Parameter Golf — What It Taught UsOpenAI ran a community ML challenge (16 MB model, 10 min training, 8×H100s). Key finding: AI coding agents ha…updated 2026-05-183 in-linksOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution (arXiv:2608.00677)Red-team the environment, not the prompt. OpenART is an arena of 10,000+updated 2026-08-142 in-linksPositive Alignment: Artificial Intelligence for Human FlourishingA 16-author collaborative paper from Oxford, DeepMind, Anthropic, and others. It argues that today's alignmen…updated 2026-05-313 in-linksQwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsA foundation GUI agent from Tongyi MAI (entities/alibaba) that unifies mobile,updated 2026-08-032 in-linksR³-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets (arXiv:2608.16033)Gives a model six problems and one budget between them, and measures how wellupdated 2026-08-191 in-linksRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningFirst demonstration of RLVR (RL with Verifiable Rewards) scaling to 1 trillion parameters — achieves 84.2% on…updated 2026-07-196 in-linksRound-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675)Train one latent diffusion model that can step a dynamical system forwards or backwardsupdated 2026-08-07Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentA 35B MoE model (Agents-A1) matches 1-trillion-parameter models on agentic benchmarks by scaling the agent ho…updated 2026-07-015 in-linksSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement LearningSEED converts an LLM agent's own completed trajectories into natural-language "hindsight skills," then distil…updated 2026-07-195 in-linksSelf-Distilled Agentic Reinforcement LearningSDAR adds token-level distillation guidance to reinforcement learning forupdated 2026-08-146 in-linksShieldstral (arXiv:2607.25857)A 3B multimodal safety classifier that takes its moderation policy as naturalupdated 2026-08-062 in-linksSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning (arXiv:2608.14277)On-policy distillation across tokenizers: align only the tokens occupyingupdated 2026-08-182 in-linksSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving SkillsSkill-SP makes agent skills the unit that a self-play loop evolves, using each skill's narrow, verifiable exe…updated 2026-07-294 in-linksSLEIGHT-Bench: Finding Blind Spots in AI MonitorsA benchmark of 40 synthetic transcripts across 11 categories, each one a coding agent doingupdated 2026-07-313 in-linksSolipsistic Superintelligence is Unlikely to be CooperativeCurrent RL/RLHF training treats the world as a stationary, exogenous environment. Deployed systems break this…updated 2026-06-066 in-linksSpatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence (arXiv:2608.12743)Asks whether a frozen VLM can improve its spatial reasoning with noupdated 2026-08-173 in-linksStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (arXiv:2608.15089)An agent runtime that changes nothing about the model weights and reportsupdated 2026-08-198 in-linksStealing Reasoning Traces from Proprietary LLM APIs (arXiv:2608.09867)Providers hide chain-of-thought by returning it to the client as an encryptedupdated 2026-08-123 in-linksThe Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement LearningLLM RL implementations use separate inference and training engines for efficiency, creating a systematic trai…updated 2026-07-175 in-linksThought-Level Beam Search for Reasoning (arXiv:2608.08020)Reframes test-time compute as a allocation problem rather than a budgetupdated 2026-08-174 in-linksVentor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs (arXiv:2608.16391)A black-box audit of what a hosted API actually serves you, requiring noupdated 2026-08-193 in-linksWeak-to-Strong Generalization via Direct On-Policy DistillationRun RL on a cheap small model; transfer only the RL-induced policy delta (not the full policy) to a large mod…updated 2026-07-164 in-links